Information processing device, information processing method, brain model and program

By employing an information processing apparatus with a learning unit that utilizes a loss function based on the logarithm of a normalized model function, the inefficiencies in current machine learning technologies are addressed, resulting in a more efficient and general-purpose learning technique.

JP2025083976AActive Publication Date: 2025-06-02BONITRADE LLC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2023197689
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-11-21
Publication Date
2025-06-02
Estimated Expiration
2043-11-21

AI Technical Summary

Technical Problem

Current machine learning technologies rely heavily on empirical methods due to a lack of understanding of the learning process, resulting in inefficient and context-specific algorithms.

Method used

An information processing apparatus that includes an acquisition unit, a model function, and a learning unit that executes a learning process using a loss function defined by the logarithm of the model function, satisfying a predetermined normalization condition.

Benefits of technology

This approach enables the development of a more efficient and general-purpose learning technique, reducing computational costs and improving the accuracy of predictions across various data sets.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025083976000001_ABST
    Figure 2025083976000001_ABST
Patent Text Reader

Abstract

To provide an efficient and versatile learning technique.SOLUTION: An information processing device (1) comprises an acquisition unit (11) that obtains one or more input values, and a learning unit (12) that executes a learning process by referring to a loss function defined by a logarithm of a model function, which uses the one or more input values as arguments and satisfies a predetermined normalization condition.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an information processing apparatus, an information processing method, a brain model, and a program.

Background Art

[0002] Since the usefulness of deep neural networks has been demonstrated, machine learning technologies such as deep learning have made great progress and have been applied to various fields (for example, Non-Patent Document 1). However, it is difficult to say that there has been substantial progress in the understanding of the learning process including machine learning.

Prior Art Documents

Non-Patent Documents

[0003]

Non-Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] Due to the difficulty in the essential understanding of the learning process as described above, the machine learning technologies known to date have a large empirical element rather than an essential one. For this reason, algorithms that are not necessarily efficient or algorithms that operate well only under specific circumstances have been widely used.

[0005] One aspect of the present invention is based on the essential findings of the inventors regarding the learning process, and aims to provide a learning technique that is more efficient and general-purpose than those known in the past.

Means for Solving the Problems

[0006] In order to solve the above problems, an information processing apparatus according to an aspect of the present invention includes an acquisition unit that acquires one or more input values, and a model function that takes the one or more input values as arguments, and a learning unit that executes a learning process with reference to a loss function defined by the logarithm of a model function that satisfies a predetermined normalization condition.

[0007] In order to solve the above problems, an information processing apparatus according to an aspect of the present invention includes a prediction unit that executes a prediction process using a model function that has been learned with reference to a loss function defined by the logarithm of a model function that takes one or more input values as arguments and satisfies a predetermined normalization condition, and an output unit that outputs a prediction result by the prediction process.

[0008] In order to solve the above problems, an information processing apparatus according to an aspect of the present invention includes a network including a plurality of nodes directly or indirectly connected to each other, and a learning unit that learns each of the plurality of nodes using a loss function defined to be spatially and temporally localized to the node.

[0009] In order to solve the above problems, a brain model according to an aspect of the present invention includes a plurality of cells directly or indirectly connected to each other, and each of the plurality of cells is learned using a loss function defined to be spatially and temporally localized to the cell.

[0010] In order to solve the above problems, an information processing method according to an aspect of the present invention includes an acquisition step of acquiring one or more input values, and a learning step of executing a learning process with reference to a loss function defined by the logarithm of a model function that takes the one or more input values as arguments and satisfies a predetermined normalization condition.

[0011] To solve the above problems, an information processing method according to one aspect of the present invention includes a prediction step of performing prediction processing using a model function having one or more input values as arguments, the model function being learned with reference to a loss function defined by the logarithm of a model function that satisfies a predetermined normalization condition, and an output step of outputting a prediction result by the prediction processing.

[0012] The information processing apparatus according to each aspect of the present invention may be realized by a computer. In this case, a program for realizing the information processing apparatus by operating the computer as each part (software element) included in the information processing apparatus, and a computer-readable recording medium on which the program is recorded also fall within the scope of the present invention.

Advantages of the Invention

[0013] According to one aspect of the present invention, an efficient and general-purpose learning technique can be provided.

Brief Description of the Drawings

[0014]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

Figure 14

Figure 15

Figure 16

Figure 17

Figure 18

Mode for Carrying Out the Invention

[0015] 〔Embodiment 1〕 Hereinafter, an embodiment of the present invention will be described in detail. FIG. 1 is a block diagram showing a configuration example of an information processing apparatus 1 according to the present embodiment. As shown in FIG. 1, the information processing apparatus 1 includes, as an example, a control unit 10, a storage unit 20, a communication unit 30, and an input / output unit 40.

[0016] (Communication unit 30) The communication unit 30 communicates with one or more devices external to the information processing apparatus 1. The communication unit 30 transmits the data supplied from the control unit 10 to the external device, or supplies the data received from the external device to the control unit 10. As an example, the communication unit 30 supplies the input data INPUT acquired from an external device to the control unit 10, or transmits the output data OUTPUT derived by the control unit 10 to an external device.

[0017] (Input / output unit 40) The input / output unit 40 is configured to include at least one of input / output devices such as a keyboard, a mouse, a display, a printer, and a touch panel, for example. Alternatively, input / output devices such as a keyboard, a mouse, a display, a printer, and a touch panel may be connected to the input / output unit 40. In such a configuration, the input / output unit 40 receives input of various types of information to the information processing apparatus 1 from the connected input devices. Further, the input / output unit 40 outputs various types of information to the connected output devices under the control of the control unit 10A. Examples of the input / output unit 40 include interfaces such as USB (Universal Serial Bus).

[0018] (Storage unit 20) The storage unit 20 stores various types of data referred to by the control unit 10 and various types of data generated by the control unit 10. As an example, the storage unit 20 stores · The input value (input data) INPUT acquired by the acquisition unit 11 described later · The model function MF that is the object of learning by the learning unit 12 described later · The loss function LF referred to in the learning process by the learning unit 12 described later · The output value (output data) OUTPUT generated or derived by the control unit 10 and the like. In the present embodiment, the phrase "model function MF" may include the meaning of "parameters defining the model function MF". Also, the "model function" may be simply referred to as the "model".

[0019] (Control Unit 10) As shown in FIG. 1, the control unit 10 includes an acquisition unit 11, a learning unit 12, a prediction unit 13, and an output unit 14.

[0020] (Acquisition Unit 11) The acquisition unit 11 acquires one or more input values INPUT. The input value may be simply referred to as an input or input data. Also, the input value may be algebraically expressed as an input value x using x. Further, since a plurality of input values can be regarded as constituting a vector having the input values as components, the input value may be referred to as an input vector.

[0021] (Learning Unit 12) The learning unit 12 performs learning of a model function with reference to a loss function. As an example, the learning unit 12 is a model function Φ(x) having the one or more input values x as arguments, and a loss function defined by the logarithm (natural logarithm as an example) of the model function Φ(x) that satisfies a predetermined normalization condition [Equation] to execute a learning process. Here, the predetermined normalization condition is, for example, when the input value x is discrete, [Equation] given by, and when the input value is continuous, [Equation] given by.

[0022] As an example, the learning unit 12 optimizes the model function by minimizing the loss function. Note that the appropriate choice of the loss function does not limit the present embodiment. For example, a loss function obtained by inverting the sign of the above (Equation 1) may be used. In such a case, it can be expressed that the learning process by the learning unit 12 includes a process of optimizing the model function by maximizing the loss function.

[0023] Here, the above model function Φ(x) is the object of optimization in the above learning process. Although it will be described in more detail later, according to the findings obtained by the inventor, it can be expressed that the essence of the learning process is to optimize the model function Φ(x) so as to reproduce the probability P(x) of the input value x. In other words, it can be expressed that the process of constructing a model function Φ(x) that is as close as possible to the probability P(x) of the input value x is the essence of the learning process.

[0024] Although it will be described in detail later, according to the findings obtained by the inventor, · Define the loss function L using the logarithm of the model function Φ(x), and · Minimize (maximize if the sign of the loss function is inverted) the loss function L By doing so, it is theoretically guaranteed that the above model function Φ(x) approaches the probability P(x) of the input value x.

[0025] Here, it should be noted that, as an example, the expression of the above loss function L does not include a sum over any values or indices. For example, the Kullback-Leibler divergence (KL divergence), which is often used as a loss function in the past and is sometimes called relative entropy, uses, as an example, probability distributions P(i) and Q(i) with respect to the value i

Equation

[0026] On the other hand, the loss function (Equation 1) according to the present embodiment is defined, for example, without including a sum with respect to any value or index. For this reason, by performing learning with reference to the loss function (Equation 1) according to the present embodiment, the calculation cost can be significantly suppressed as compared with the prior art.

[0027] (Prediction unit 13) The prediction unit 13 is a model function that takes one or more input values as arguments, and executes prediction processing using a model function that has been learned with reference to a loss function defined by the logarithm of a model function that satisfies a predetermined normalization condition. As an example, the prediction unit 13 executes prediction processing using the model function Φ(x) that has been learned (optimized) by the learning unit 12 described above. Then, the prediction unit 13 derives an output value (output data) OUTPUT indicating the result of the prediction processing.

[0028] (Output unit 14) The output unit 14 outputs the output value (output data) OUTPUT generated or derived by the control unit 10, for example, via the input / output unit 40. For example, the output unit 14 outputs the prediction result of the prediction processing executed by the prediction unit 13. Further, the output unit 14 may be configured to output the model function Φ(x) itself that has been learned (optimized) by the learning unit 12.

[0029] According to the prediction unit 13 and the output unit 14 described above, prediction processing can be performed using a model (model function) that has been suitably learned by the learning unit 12, and the prediction result can be suitably output.

[0030] (Flow of processing by the information processing apparatus) Next, with reference to FIGS. 2 and 3, an example of the processing flow by the information processing apparatus 1 will be described. FIG. 2 is a flowchart showing an example of the learning processing flow by the information processing apparatus 1.

[0031] (Step S11) In step S11, the acquisition unit 11 acquires one or more input values INPUT. Since the processing by the acquisition unit 11 has been described above, the description is omitted here.

[0032] (Step S12) Subsequently, in step S12, the learning unit 12 executes a learning process with reference to a loss function (Equation 1) defined by the logarithm (natural logarithm as an example) of a model function Φ(x) that takes the one or more input values x as arguments and satisfies a predetermined normalization condition. Since the processing by the learning unit 12 has been described above, the description is omitted here.

[0033] As described above, the loss function (Equation 1) according to the present embodiment is defined, for example, without including a sum related to any value or index. Therefore, according to the above processing flow for performing learning with reference to the loss function (Equation 1) according to the present embodiment, the calculation cost can be significantly reduced compared to the prior art.

[0034] FIG. 3 is a flowchart showing an example of the inference processing (prediction processing) flow by the information processing apparatus 1.

[0035] (Step S13) In step S13, the prediction unit 13 executes a prediction process using a model function that has been learned with reference to a loss function defined by the logarithm of a model function that takes one or more input values as arguments and satisfies a predetermined normalization condition. As an example, the prediction unit 13 executes a prediction process using the model function Φ(x) learned (optimized) by the above-described learning unit 12.

[0036] (Step S14) Subsequently, in step S14, the output unit 14 outputs the prediction result obtained by the prediction process executed by the prediction unit 13. According to the information processing apparatus 1 that executes the above processing flow, prediction processing can be performed using the model (model function) preferably learned by the learning unit 12, and the prediction result can be preferably output.

[0037] (Processing example) Hereinafter, various processing examples by the information processing apparatus 1 will be described. Note that the processing examples listed below are examples of processing expressed in an executable manner based on the inventor's findings, but these specific processing examples do not limit the present embodiment.

[0038] (Processing example 1) In this example, the acquisition unit 11 of the information processing apparatus 1 acquires an input value INPUT including one or more data values and a teacher label associated with the data value. As an example, the input value INPUT acquired by the acquisition unit 11 may be configured to include a set of input data {x 1 , x 2 , ···} and a set of teacher labels {t 1 , t 2 , ···}. In other words, the input value INPUT may be configured to include one or more pairs (x, t) of input data x and the corresponding teacher label t. The input value INPUT configured in this way can be used as learning data in a supervised classification problem, for example. However, this statement does not limit this example.

[0039] The model function in this example can be expressed as Φ(x, t) using the above input data x and the corresponding teacher label t. In other words, the model function Φ according to this example is, for example, a function that takes the input data x included in the above input value and the corresponding teacher label t as arguments.

[0040] The learning unit 12 according to this example is a model function Φ(x, t) that takes as arguments the input data x included in the input value INPUT and the corresponding teacher label t, and is a loss function defined by the logarithm (natural logarithm as an example) of the model function Φ(x, t) that satisfies a predetermined normalization condition.

Number

Number

[0041] As an example of the model function (also simply referred to as a model) according to this example, a neural network shown in FIG. 4 can be cited. As shown in FIG. 4, in this example, the input value x is input to the input layer, and the output layer has the same number of nodes as the type of teacher label (sometimes also called a class, etc.). In the example of FIG. 4, the sum of the values of the nodes in the output layer is normalized by the softmax function, and the model function Φ(x, t) is defined as the value of the corresponding node in the output layer. In this case, the loss function (Equation 5) according to this example is the same as the cross entropy with a one-hot vector.

[0042] (Processing Example 2) In the above-described Processing Example 1, the case where the value of the teacher label is discrete was taken as an example, but this does not limit the present embodiment. As will be described in this example, the value of the teacher label may be continuous.

[0043] As an example, the input value INPUT acquired by the acquisition unit 11 is a set of input data {x 1 , x 2, ···}, and a set of teacher labels {t 1 , t 2 , ···} is configured to include, and each teacher label t i may take continuous values. In other words, the input value INPUT may be configured to include one or more pairs (x, t) of input data x and a corresponding teacher label t that takes continuous values. The input value INPUT configured in this way can be used, for example, as learning data in a regression problem. However, this statement does not limit this example.

[0044] The learning unit 12 according to this example is a model function Φ(x, t) that takes the input data x included in the input value INPUT and the corresponding teacher label t as arguments, and is a loss function defined by the logarithm (natural logarithm as an example) of the model function Φ(x, t) that satisfies a predetermined normalization condition

Equation

Equation

[0045] Note that in this example, as described above, since the normalization condition regarding the model function Φ(x, t) includes an integral, it is not suitable to directly use the model as described in Processing Example 1.

[0046] In the following description according to this example, several approximations and assumptions are used. As an example, as the model function Φ(x, t), a Gaussian distribution function

Equation

[0047] FIG. 5 is a diagram showing the model according to this example, and shows the model in the case of n = 2. More specifically, μ is

Number

Number

[0048] And in the above example, the loss function L(Φ(x, t)) is

Number

[0049] As schematically shown in FIG. 5, the learning unit 12 according to this example, as an example, inputs the input data x included in the input value INPUT into the model (model function Φ(x, t)), and each parameter output by the model (in the above example, μ 1 , μ 2 , σ 11 , σ 12 , σ 22Determine the values of the above parameters so as to minimize (or maximize when the sign of the loss function is inverted) the loss function L(Φ(x,t)) given by Equation 12.

[0050] Thus, in this example, · The model function Φ(x,t) is, as an example, a distribution function defined including one or more parameters (in the above example, μ 1 , μ 2 , σ 11 , σ 12 , σ 22 ), and · The learning process according to this example can be expressed as including a process of determining the one or more parameters by minimizing or maximizing the loss function L(Φ(x,t)).

[0051] Incidentally, as another example, when the tilde-marked Σ is approximated by the identity matrix, the loss function of Equation 12 results in

Equation

[0052] Although it will be described in more detail later, as can be understood from the above, the conventionally used mean squared error, from the perspective of the information processing apparatus 1 according to this embodiment, is a distribution (distribution function) defined by the input data x and the corresponding t, and corresponds to approximating the true distribution P(t|x) regarding the input value by a Gaussian distribution having a certain variance. There have been cases where the mean squared error has been used even for data for which such an approximation should not be used conventionally, which has been a cause of poor learning. In this example, by selecting an appropriate probability distribution function for the data used for learning, more general-purpose learning becomes possible.

[0053] (Processing Example 3) In this processing example, a process executed by the information processing apparatus 1 and related to a parameter estimation problem will be described. However, this statement does not limit this example.

[0054] First, assume that some data {x 1 , x 2 , ···, x N} is given. When these data follow a distribution defined by some parameters, there is a known problem of how to estimate the values of these parameters (also called the parameter estimation problem). In such a problem, the maximum likelihood estimation method has often been used conventionally.

[0055] As an example, when representing a set of parameters by θ and representing the probability distribution of x parameterized by θ by P(x|θ), the above parameter estimation problem can be expressed as the problem of finding the parameter θ that best explains the above data set {x 1 , x 2 , ···, x N}.

[0056] This can be expressed as the problem of maximizing the likelihood function L * (θ), or the logarithm logL * (θ) of the likelihood function L * (θ).

[0057]

Equation

Equation

Equation

Number

Number

[0058] Figure 6 shows the model according to this example. As shown in Figure 6, the model according to this example has neither layers nor links, and can be interpreted as being the parameter θ itself (in the example shown in Figure 6, the 4-dimensional parameter θ 1 , θ 2 , θ 3 , θ 4 ).

[0059] As processing using the model function Φ(x) given as described above, the learning unit 12 according to this example

Number

[0060] In this way, the model function Φ(x) according to this example · can also be expressed as a distribution function taking one or more input values x as arguments, · can also be expressed as being the one or more parameters (θ) themselves that define the distribution function.

[0061] Comparing the conventionally used likelihood function (Equation 15) with the loss function L(Φ(x)) (Equations 16 and 19) referred to by the learning unit 12 according to this example, the conventionally used likelihood function (Equation 15) is defined including the sum over the entire dataset (in the above example, the sum from index i = 1 to N), whereas the loss function L(Φ(x)) referred to by the learning unit 12 according to this example is defined without including such a sum. In other words, the loss function L(Φ(x)) referred to by the learning unit 12 according to this example is defined for each x, and the learning unit 12 calculates the value of the loss function L(Φ(x)) for each x.

[0062] As described above, since the learning unit 12 according to this example performs learning using a loss function defined without including the sum over the dataset, the computational cost can be significantly reduced as compared with the conventional case.

[0063] (Expression of the model function by differentiation) Although it will be described in more detail later, based on the findings obtained by the inventor, the above-described model function may be expressed by differentiation. In other words, the information processing apparatus 1 according to the present embodiment may perform processing with reference to the model function Φ(x) expressed as the derivative of a certain function.

[0064] (Case of one variable) First, the case of one variable will be described as follows. In the case of one variable, as an example, the model function Φ(x) according to the present embodiment may be a function expressed as

Equation

[0065] Here, in order to give a probabilistic interpretation, as a condition for the above function y,

Number

[0066] The first condition in Equation 21 is achieved, for example, in a network model as shown in FIG. 7 by setting the weight coefficients to only positive values and the activation function at each node to a monotonically increasing function. Also, the second condition in Equation 21 is achieved, for example, by using a sigmoid function at the output node as shown in FIG. 7. A model function Φ(x) that satisfies such conditions satisfies the following normalization conditions.

Number

[0067] Note that even in the above configuration, the loss function L(Φ(x)) is

Number

[0068] FIG. 8 is a diagram showing an example of the result of the learning process by the learning unit 12 using the model function Φ(x) defined as described above, and shows the result when the input data shows a multimodal distribution. More specifically, in FIG. 8, the results for the case where the sample size of the input data is 4000 and the results for the case where the sample size of the input data is 40 are shown.

[0069] In FIG. 8, the histogram shows the distribution of the input data (multimodal distribution in the example of FIG. 8), the curve represented by the light gray indicates the true probability distribution of the input data, and the curve represented by the dark gray indicates the (estimated) probability distribution derived (estimated) by the above learning process by the learning unit 12 (the distribution shown by the model function Φ(x)).

[0070] As can be seen from FIG. 8, by executing the learning process using the above model function obtained based on the inventors' findings, not only when the sample size of the input data is relatively large (for example, in the case of the sample size = 4000 shown in FIG. 8), but also when the sample size of the input data is very small (for example, in the case of the sample size = 40 shown in FIG. 8), good estimation results are shown.

[0071] Further, FIG. 9 is a diagram showing an example of the result of the learning process by the learning unit 12 using the model function Φ(x) defined as described above, and shows the result when the input data shows a flat distribution. More specifically, in FIG. 9, the results for the case where the sample size of the input data is 1000 and the results for the case where the sample size of the input data is 10 are shown. The legend is the same as in FIG. 8.

[0072] As can be seen from FIG. 9, by executing the learning process using the above model function obtained based on the inventors' findings, not only when the sample size of the input data is relatively large (for example, in the case of the sample size = 1000 shown in FIG. 9), but also when the sample size of the input data is very small (for example, in the case of the sample size = 10 shown in FIG. 9), good estimation results are shown.

[0073] FIG. 10 is a diagram showing an example of the result of the learning process by the learning unit 12 using the model function Φ(x) defined as described above, and shows the result when the input data shows an asymmetric distribution. More specifically, in FIG. 10, the results for the case where the sample size of the input data is 10,000 and the case where the sample size of the input data is 100 are shown. The legend is the same as in FIG. 8.

[0074] As can be seen from FIG. 10, by executing the learning process using the above-described model function obtained from the inventor's findings, not only when the sample size of the input data is relatively large (for example, in the case of the sample size = 10,000 shown in FIG. 10), but also when the sample size of the input data is very small (for example, in the case of the sample size = 100 shown in FIG. 10), good estimation results are shown.

[0075] Thus, according to the configuration of the present embodiment, not only can the calculation cost in the learning process be significantly suppressed, but also a model with high estimation accuracy can be constructed even when the data size of the input data is small. In addition, for any data, high-precision learning can be performed without any prior knowledge, and the learning process is extremely versatile compared to conventional models. (Case of Multivariable (Multiple Variables) (Part 1)) Subsequently, the case of multivariable (multiple variables) will be described as follows. First, a set of input data {x 1 , x 2 , ···} is represented as x = (a 1 , a 2 , ···, a n ). Here, each a takes a continuous value as an example, and n is the dimensionality of the input data. i

[0076] In the case of multiple variables, as an example, the model function Φ(x) according to the present embodiment uses a certain function y(x) to [Equation]​ It may be a function expressed as follows. In other words, the learning unit 12 may use a model function Φ(x) expressed as the derivative (partial derivative) of a certain function y(x). In other words, the model function Φ(x) may be expressed as the derivative of a certain function with the input value (input data) as an argument, and may be expressed as a derivative obtained by a plurality of differentiations (n-th order differentiations) corresponding to the dimension (n) of the input value (input data).

[0077] Here, as an example, the function y(x) may be a network model that takes x as an input value and y as an output value, as expressed by FIG. 11.

[0078] Here, in order to give a probabilistic interpretation, as a condition for the above function y,

Equation

[0079] The first condition in Equation 25 is achieved, for example, in the network model shown in FIG. 11 by setting the weight coefficients to only positive values and the activation function at each node to a monotonically increasing function. Also, the second condition in Equation 25 is achieved, for example, by using a sigmoid function at the output node as shown in FIG. 11. A model function Φ(x) that satisfies such conditions satisfies the following normalization conditions.

Equation

[0080] Note that even in the above configuration, the loss function L(Φ(x)) is

Number

[0081] Note that since the above example in the case of multiple variables includes n-th order differentiation, as the dimension n of the input data increases, there is an aspect that the calculation cost increases. Hereinafter, an example of a process that can solve this problem will be described.

[0082] (Case of multiple variables (plural variables) (Part 2)) In the following example of a process, a process using the Jacobian (Jacobian determinant) will be described instead of the n-th order differentiation (Equation 24) used in the above example.

[0083] In this example, as an example, the model function Φ(x) according to the present embodiment is expressed using a certain function y(x) as

Number

Mathematics

[0084] Here, as an example, the function y(x) may be a network model that takes x as an input value and y as an output value, as represented by FIG. 12.

[0085] Here, in order to give a probabilistic interpretation, as a condition for the above function y,

Mathematics

Mathematics

[0086] Also, in this example, further, as a condition for the above function y,

Equation

[0087] Here, as an example, the function y(x) may be a network model that takes x as an input value and y as an output value, as represented by FIG. 13. The model has multivariate x = (a 1 , a 2 , ···, a n ) input and multivariate y(x) = (b 1 , b 2 , ···, b n ) output. As shown in FIG. 13, the model has only links (edges) that satisfy the condition i ≥ j as links (edges) from the j-th node in a certain layer to the i-th node in the next layer. Expressed using the arrangement shown in FIG. 13, there are no links (edges) from lower nodes to upper nodes. This configuration in the network shown in FIG. 13 corresponds to the condition of the first equation of Equation 32 above. Also, as shown in the second equation of Equation 32, the values of the links (edges) where i = j are all positive.

[0088] In the configuration as described above, the model function Φ(x) according to the present embodiment is

Equation

[0089] Thus, the Jacobian according to this example is expressed as the product of the diagonal components of the triangular matrix, and the loss matrix is expressed as the sum of the logarithms of each of the diagonal components of the triangular matrix, so the computational cost is of the order of n. Also, according to the second equation of Equation 32 described above, the positive definiteness of the Jacobian is guaranteed.

[0090] Therefore, according to the learning unit 12 that performs the processing according to this example, the computational cost required for the learning process can be significantly suppressed.

[0091] (Application to regression problems) Hereinafter, a processing example of applying the above-described expression of the model function by differentiation to regression problems in general will be described. Note that descriptions of matters already explained regarding the notation will be omitted as appropriate.

[0092] In this example, as an example, the model function Φ(x, t) according to this embodiment uses a certain function y(x, t) and [Number] is expressed as follows. Here, as an example, the variable t is t=(a 1 ,a 2 ,···,a n ) an n-dimensional variable expressed as such, and the variable y(x,t) is, as an example, y(x,t)=(b 1 ,b 2 ,···,b n ) an n-dimensional variable (function) expressed as such.

[0093] In this way, the learning unit 12 may use a model function Φ(x) expressed as a Jacobian defined by a plurality of input values (x,t) (where t=(a 1 ,a 2 ,···,a n )) and a plurality of functions y(x,t)=(b 1 ,b 2 ,···,b n ) that take at least a part of the plurality of input values (x,t) as arguments. In other words, the model function Φ(x) may be expressed as a Jacobian defined by at least a part of the plurality of input values (x,t) and a plurality of functions y that take at least a part of the plurality of input values (x,t) as arguments.

[0094] Here, the function y(x,t) may be, as an example, a network model that takes (x,t) as an input value and y as an output value, as expressed by FIG. 14.

[0095] Also, in order to give a probabilistic interpretation, as a condition for the above function y(x,t),

Equation

[0096] Here, the function y(x,t) may be a network model that takes (x,t) as an input value and y(x,t) as an output value, as represented by FIG. 14 as described above. The model has multivariate (x,t) (where t = (a 1 ,a 2 ,···,a n )) input and multivariate y(x,t)=(b 1 ,b 2 ,···,b n ) output. As shown in FIG. 14, in the sector where t is input (the sector corresponding to the bottom two rows in FIG. 14), the model has only links (edges) that satisfy the condition i≧j from the j-th node in a certain layer to the i-th node in the next layer. Expressed using the arrangement shown in FIG. 14, in the sector where t is input, there are no links (edges) from lower nodes to upper nodes. This configuration in the network shown in FIG. 14 corresponds to the condition of the first equation of the above formula 36. Also, as shown in the second equation of the above formula 36, in the sector where t is input, the values of the links (edges) with i = j are all positive.

[0097] In the configuration as described above, the model function Φ(x,t) according to this embodiment is

Equation

Equation

Equation

[0098] Thus, the Jacobian according to this example is expressed as the product of the diagonal components of a triangular matrix even in the process related to the regression problem, and the loss matrix is expressed as the sum of the logarithms of each of the diagonal components of the triangular matrix. Therefore, the computational cost is of the order of n. Also, the positive definiteness of the Jacobian is guaranteed by the second equation of Equation 36 described above.

[0099] Therefore, according to the learning unit 12 that performs the process according to this example, even in the learning process related to the regression problem, the computational cost required for the learning process can be significantly suppressed.

[0100] (Local loss function) By considering the above processing example, it can be seen that the inventor's finding of expressing the model function by differentiation leads to an important finding of "local loss function". More specifically, the loss functions expressed by the above Equation 34 and Equation 39 are both local derivatives specified by the index i [Number] is expressed as the sum of logarithms (sum with respect to index i). In other words, the loss functions expressed by the above equations (34) and (39) are both local loss functions specified by index i.

Number

[0101] In other words, the learning unit 12 according to the present embodiment can be expressed as applying learning processing with reference to a local loss function defined by the target node belonging to a neural network including one or more nodes in each layer and the nodes belonging to the layer adjacent to the target node.

[0102] Hereinafter, a more general processing example will be described with respect to the above local loss function. First, the intermediate layer in FIG. 13 described above is (z 1 ,···,z d ) denoted as. Here, d represents the depth of the layer of the network. Each layer has n variables. As an example, the i-th layer is z i =(c 1;i ,c 2;i ,···,c n;i ) denoted as.

[0103] As described above, the model function Φ(x) according to the present embodiment expressed as a Jacobian uses the chain rule of differentiation

Number

[0104] And the loss function L(Φ(x)) according to this embodiment is

Equation

Equation

[0105] In other words, the learning unit 12 according to this embodiment can be expressed as applying learning processing with reference to a local loss function defined by the target node (i, j) belonging to a neural network including one or more nodes in each layer, and the node (i, j + 1) belonging to the layer adjacent to the target node (i, j). Furthermore, more specifically, the local loss function is expressed using the logarithm of the derivative defined by the variable c i;j associated with the target node, and the variable c i;j+1 associated with the node belonging to the layer adjacent to the target node. Since the learning unit 12 according to this embodiment performs learning processing using such a local loss function, for example, it is not necessary to use the backpropagation method as was previously done. Therefore, according to the learning unit 12 that performs the processing according to this example, even in the learning processing related to the regression problem, the computational cost required for the learning processing can be significantly suppressed.

[0106] (Normalization by time evolution, and brain model) Hereinafter, normalization by time evolution and a brain model by time evolution based on the inventor's findings will be described. Note that the "brain model" described below may be configured to be executed by the information processing apparatus 1 in the form of an algorithm as an example, or may be configured to be realized as a physical structure using biological or chemical materials in part or in whole. Therefore, the "brain model" according to this embodiment may be expressed as the "algorithm of the brain model" or as the "brain model using biological or chemical materials".

[0107] FIG. 15 is a diagram showing an example of a network representing the brain model according to this embodiment, and is a diagram showing a fully-connected model. In each of the above examples, a neural network having a plurality of layers was given as an example. However, as a more realistic brain model, a model that does not require the concept of a "spatial layer" is suitable. As such a model, a fully-connected model as shown in FIG. 15 or a model in which some connections are deleted in the fully-connected model (sometimes called a locally-connected model) is suitable. Note that, as will be described below, the brain model according to this embodiment is expressed as a time evolution model using the concept of a "temporal layer".

[0108] First, the notation will be described. The brain model according to this embodiment is configured as a network having a plurality of temporal layers defined by each time (each time step), as shown in FIG. 16 as an example.

[0109] The brain model according to this embodiment includes n nodes (n = 4 in FIG. 16), and each node (the function associated with each node) exhibits a time-dependent change. These nodes are represented as a(t)=(a 1 (t),···a n (t)) Here, t is a variable representing time, and each a i (t) is assumed to take a real value. As an example, among the above n nodes, the nodes from the 1st to the mth (m < n) are represented using the variable x=(x 1 ,···,x m ) as (a 1 (0),···a m (0))=(x 1 ,···,x m ) and the nodes from the (m + 1)th to the nth among the above n nodes may be set as (a m+1 (0),···a n (0))=(r m+1 ,···,r n ) Here, each r i takes a random value following a predetermined probability distribution. As an example, FIG. 16 illustrates the case of n = 4 and m = 2. However, these specific settings do not limit this embodiment.

[0110] And, as shown in FIG. 16, the brain model according to this embodiment has each node connected (temporally) by one or more edges (links), and each edge has a weight taking a real value. More specifically, it is assumed that the edge from the jth node to the ith node has a weight w ij .

[0111] Using the above notation, the time evolution of each node of the brain model according to this embodiment can be expressed as

Equation

Number

Number

Number

[0112] And the above model function is expressed as

Number

Number

Number

Number

[0113] Thus, the brain model according to this embodiment · includes a plurality of cells directly or indirectly connected to each other, · Each of the plurality of cells is learned using a loss function defined locally in space and time for the cell can be expressed as such.

[0114] Also, the brain model according to this embodiment · a network including a plurality of nodes directly or indirectly connected to each other, · a learning unit (learning unit 12) that learns each of the plurality of nodes using a loss function defined locally in space and time for the node can be expressed as an information processing device having the above, or can be expressed as an algorithm executed in such an information processing device.

[0115] As described above, since the loss function (local loss function) according to this embodiment is defined locally in space and time, in the learning stage, optimization may be performed for each time and each node. Therefore, the brain model configured as described above can significantly suppress the learning cost.

[0116] Note that the concept of the local loss function according to the present embodiment, exemplified using Formula 34, Formula 39, Formula 44, and Formula 53, is based on the profound insights and considerations of the inventors, is not something known in the past, and is not something obtained by combining conventional technologies. More detailed considerations and philosophies by the inventors leading to the invention described in the present embodiment will be described later.

[0117] [Embodiment 2] Other embodiments of the present invention will be described below. For convenience of explanation, members having the same functions as the members described in the above embodiment are given the same reference numerals, and the description thereof will not be repeated.

[0118] (Information processing apparatus 2) Hereinafter, the information processing apparatus 2 according to the present embodiment will be described. FIG. 18 is a block diagram showing a configuration example of the information processing apparatus 2.

[0119] As shown in FIG. 18, the control unit 10 included in the information processing apparatus 2 does not include the learning unit 12 included in the information processing apparatus 1. Further, in the storage unit 20 included in the information processing apparatus 2, model functions MF and the like are stored in the same manner as in the information processing apparatus 1. Here, as an example, the model function is a (optimized) model learned by the learning unit 12 included in the above-described information processing apparatus 1.

[0120] The prediction unit 13 included in the information processing apparatus 2 executes processing using the model function. Then, the output unit 14 outputs the prediction result of the prediction processing executed by the prediction unit 13. As described above, the information processing apparatus 2 according to the present embodiment is a model function having one or more input values as arguments, and refers to a loss function defined by the logarithm of a model function that satisfies a predetermined normalization condition. A prediction unit (13) that executes prediction processing using the learned model function, and an output unit (14) that outputs the prediction result of the prediction processing are provided.

[0121] According to the above configuration, prediction processing can be performed using a preferably learned model (model function), and the prediction result can be preferably output.

[0122] (Applicable Examples) Hereinafter, exemplary applicable examples of the information processing apparatus 1 and the information processing apparatus 2 (hereinafter also simply referred to as the information processing apparatus) according to each of the above embodiments will be described. However, the following applicable examples do not limit the invention described in this specification. (Example 1: Application to Anomaly Detection) Techniques for detecting anomalies from image data and numerical data are very important. The information processing apparatus according to each of the above embodiments can be used for such anomaly detection. The application range of techniques for detecting anomalies from image data and numerical data is wide, including defective product detection of factory products, error judgment of equipment in airplanes and cars, early detection of diseases from MRI images, and so on. When performing such anomaly detection by machine learning, it is practically impossible to assign teacher labels to all error patterns for learning, so unsupervised learning must be performed. However, conventional unsupervised learning is theoretically incomplete in the first place, and there are many parts that are done based on empirical rules. On the other hand, it is theoretically guaranteed that the unsupervised learning described in this specification regarding the information processing apparatus according to each of the above embodiments can basically handle any case. For example, by using the model described in the above (Case of Multivariables (Multiple Variables) (2)), if image data or numerical data is learned as input, a model for estimating the probability of the input itself can be created. If the probability is high to a certain extent (higher than a predetermined value), it can be determined to be normal, and if it is low, it can be determined that some anomaly has occurred. (Example 2: Application to Stock Price Prediction and Exchange Rate Prediction) Stock price prediction and exchange rate prediction have been considered difficult with conventional machine learning. The information processing apparatus according to each of the above embodiments can be used for stock price prediction and exchange rate prediction. It is considered that the reasons why stock price prediction and foreign exchange rate prediction have been difficult with conventional machine learning are the high randomness and the large number of outliers. For example, regarding stock price prediction, assume that the price movement and the company's performance over a certain period are used as inputs, and then the subsequent price movement is used as a teacher label for learning. When the mean squared error is used as the loss function in this case, it was shown in the above (Processing Example 2) that it is equivalent to approximating the probability distribution of the price movement with a normal distribution with a fixed variance. However, since the actual probability distribution of the price movement may vary greatly or slightly depending on the time period, the variance is not a constant value. Furthermore, since there are many outliers, approximating with a normal distribution itself is incorrect in the first place. For example, distributions such as the Cauchy distribution, the t-distribution, and the Johnson SU distribution have heavier tails than the normal distribution and are probability distributions that can better represent outliers. In the information processing apparatus according to each of the above embodiments, by adopting an approximation using these probability distributions (Cauchy distribution, t-distribution, Johnson SU distribution, etc.), better prediction becomes possible. Also, in the information processing apparatus according to each of the above embodiments, if the model described in the above (Application to regression problems) is used, it is also possible to obtain the probability distribution of the price movement without even making such an approximation. (Example 3: Application to marketing) Market research is very important for conducting good marketing. The information processing apparatus according to each of the above embodiments can be used for such market research. For example, consider the case of investigating data such as the average sales price per customer and the frequency of use at a certain store. Usually, in common analysis, the average sales price per customer and the average frequency of use are calculated, but this is not always a good method. If the sales prices of customers are distributed at both extremes, high and low, calculating the average value has almost no statistical meaning. Also, when the sales price and the frequency of use are correlated, analyzing the data by taking the average of these quantities individually often leads to incorrect results. Even in such a situation, in the information processing apparatus according to each of the above embodiments, analysis can be easily performed by using the model described in the above (case of multivariable (plural variables) (the second thereof)). Specifically, a plurality of parameters such as the unit selling price and usage frequency of each customer are input, and unsupervised learning is performed in the information processing apparatus according to each of the above embodiments. The resulting model represents a multidimensional distribution of the unit selling price and usage frequency of customers, etc., and by referring to it, a much more meaningful analysis can be performed than simply using the average value. Different from conventional machine learning, in the information processing apparatus according to each of the above embodiments, the great advantage is that it can handle multimodal distributions, skewed distributions, and even a small number of data. (Example 4: Realization in Distributed Computing) Also, regarding the hardware aspect in a specific implementation, it should be mentioned. The greatest feature of this model is that optimization is performed by referring to a locally defined loss function. Conventional machine learning and computers have a centralized structure, and they use a core with powerful computing capabilities such as GPUs and CPUs to perform calculations in a batch, and optimize various parameters by referring to the results. On the other hand, in the model described in each of the above embodiments, rather, a large number of cells with a little computing power may be prepared, and they may be individually calculated to optimize the parameters. Or, if this model is implemented by regarding computers connected on the Internet as cells like in distributed computing, the network itself can be regarded as a huge AI. Such a configuration is also included in the configuration described in this specification. In other words, a network including a plurality of nodes directly or indirectly connected to each other, a learning unit that learns each of the plurality of nodes by using a loss function defined to be spatially and temporally localized to the node, and a plurality of nodes included in the above network are distributedly arranged in a plurality of information processing apparatuses such a system configuration is also included in the configuration described in this specification.

[0123] <Explanation of learning principle> In each of the above-described embodiments, each configuration and each process described are based on the findings of the learning principle discovered by the inventor. Hereinafter, the learning principle will be described. Note that the following description is not merely an explanation of the principle, but also has the position of an explanation regarding the configurations realized in each information processing apparatus and brain model described in each of the above-described embodiments.

[0124] (Summary) While deep learning has achieved remarkable success, there is no clear explanation as to why it is so effective. In order to quantitatively discuss this question, a mathematical framework for explaining what learning is in the first place is necessary. As a result of repeated considerations, the inventor has succeeded in constructing a mathematical framework that can uniformly understand all types of learning, including deep learning and learning in the brain. This is called the learning principle, and it is derived that all learning is equivalent to estimating the probability of input data. Here, the objective function used when performing learning must be defined by the logarithm of the estimated probability. The inventor has not only shown this principle but also mentioned its application to specific machine learning models. For example, by showing that supervised learning is equivalent to estimating a conditional probability distribution, machine learning can be applied to a wider field. Also, a completely new method of defining using differentiation with respect to probability estimation was proposed, and it was shown that it is essentially general-purpose. Furthermore, by applying this method and considering the temporal evolution of the model, the inventor has also succeeded in mathematically describing learning in the brain. The learning principle provides mathematical solutions to many unsolved problems in deep learning and cognitive brain science.

[0125] (Chapter 1: Introduction) Since the usefulness of deep learning was demonstrated, research has been continuously conducted, and its applications have advanced in various fields. In particular, the performance of models specialized for specific data is remarkable. Models incorporating CNNs used in image recognition and attention mechanisms used in language processing have shocked the world with their high performance. Then, why has deep learning been so successful? Currently, we are just using what has been successful in an ad-hoc manner, and there is no theoretical guarantee that deep learning will always work well. For further development in this field, an essential understanding of learning is indispensable, which can provide a clear answer to the reason why deep learning is successful.

[0126] As a rough explanation, it is said that by making the layers deeper, it becomes possible to extract complex features, which is the key factor in the success of deep learning. However, there is no mathematically rigorous definition of features, and the question of "why learning works well" has simply been replaced by the question of "why features can be extracted." We should not be satisfied with such qualitative explanations. Without evaluating using quantitative methods, we cannot gain an essential understanding. Therefore, let's reconsider from the most fundamental aspect of what a state where learning works well is like, using a quantitative method. When evaluating supervised machine learning, learning is considered to be successful if it can make predictions closer to the teacher labels. Then, where is the guarantee that this teacher label is actually correct? For example, the teacher labels for tasks such as image classification problems and language translation are assigned by humans, and there is no mathematical necessity there. Ultimately, teacher data is created as a result of humans' learning, and we must discuss from the point of why it is correct. Therefore, in order to conduct a mathematically quantitative evaluation of machine learning, we must have a framework that can uniformly discuss human learning, that is, learning by the human brain. Such a framework is called the learning principle, and clarifying it is the purpose of the <Explanation of the learning principle> in this specification. The first half until Chapter 4 describes the derivation of the learning principle, and the second half from Chapter 5 describes the application of the learning principle.

[0127] (Chapter 2: Philosophy of Learning Principles) In this chapter, we will describe the fundamental concepts underlying the derivation of the learning principles. The detailed mathematical calculations will be left to Chapter 3, and the complete definition of the learning principles will be presented in Chapter 4.

[0128] (2.1: Three Essential Elements for Learning) The goal of the inventors is to derive learning principles that can uniformly describe all types of learning, including machine learning and learning in the human brain. To this end, we will consider the elements common to all types of learning.

[0129] First, in any type of learning, there is an object to be optimized. For example, in deep learning, parameters such as weights and biases are the objects of optimization, and in the case of the brain, the way neurons are connected corresponds to this. Neural networks and the brain itself can be regarded as structures that include these objects of optimization, and hereinafter, such structures will be referred to as models using the terminology of deep learning.

[0130] Second, input data is required in any type of learning. In the case of machine learning, it is the input itself to be fed into the model, and in the case of the brain, information from the five senses corresponds to this.

[0131] Third, an optimization strategy must be defined in any type of learning. In the case of machine learning, basically, some objective function is defined in advance, and the model is optimized by minimizing or maximizing its value. In the case of the brain, such an objective function is not clear, but we will proceed with the discussion assuming that there is something corresponding to it. Hereinafter, using the terminology of deep learning, the objective function will be referred to as the loss function, and optimization will be performed by minimizing the loss function.

[0132] From the above considerations, three elements are essential for learning: a model, input data, and a loss function. Conversely, if these elements are minimally defined, learning can be carried out. In what follows, without considering specific aspects such as the internal structure of the model, the types of input data, and the computational methods for minimizing the loss function, an abstract discussion of the three essential elements will be advanced.

[0133] (2.2: Thought Experiment on the Ideal Case) To understand the essence of learning, a thought experiment is conducted on learning in an ideal situation. Here, the ideal situation refers to a situation where the model has universal approximation ability and an infinite amount of input data and computational resources can be prepared. Any loss function can be used. Here, a brief explanation of universal approximation ability will be given. Any model can be regarded as a function that receives an input and returns some output, and hereinafter such a function will be called a model function. When the model function can approximate any function by changing the internal parameters, the model is said to have universal approximation ability.

[0134] Returning to the topic, let's organize what results will be obtained when actually learning in an ideal situation. Since there is an infinite amount of input data and computational resources, the model can be optimized as much as possible, and finally a solution that minimizes the loss function can be reached. Also, assuming that the model has universal approximation ability, the solution here is a model function that truly minimizes the loss function. Hereinafter, when using the word "solution", it will refer to such a model function.

[0135] The important thing is that there is only one solution that truly minimizes the loss function. If the model has universal approximation ability, the same model function will ultimately be obtained regardless of its internal structure. That is, the solution depends only on the input data set and the loss function, and does not depend on the internal structure of the model. This is a very important result, indicating that the details of the model are irrelevant when considering the learning principle.

[0136] (2.3: Solution to Minimize the Loss Function) Now that we know that the details of the model are irrelevant to the learning principle, the next thing to consider is the relationship between the input data, the loss function, and the solution that minimizes it. First, consider the example of supervised learning in deep learning. In supervised learning, an input dataset and its corresponding teacher labels are given, and the goal is to build a model that predicts the teacher labels from the input data as accurately as possible. That is, what is expected as a solution is a model function that predicts the correct teacher labels for all input data. The loss function used for learning must be minimized for such a model function. An example of a loss function that satisfies this condition is the mean squared error, and when the actual model function always predicts the correct teacher labels, the value of the mean squared error is minimized.

[0137] Next, a more general case will be described. Since we want to find a framework that can uniformly understand all learning, consider the most general case, that is, unsupervised learning without prior knowledge about the input data. What was emphasized in the example of supervised learning earlier is that first, assume the desired solution and then consider the loss function to obtain it. Then, in the case of unsupervised learning, what model function should be assumed as the solution? Also, assuming the solution, how should the loss function to obtain it be defined? These questions will be examined in the next section.

[0138] (2.4: Probability of Input) One needs to assume something as the solution, but what kind of information can be extracted in the first place in unsupervised learning? Here, we are discussing the most common case, so such information must be definable for any input data. The answer to this question is the probability itself that the input data has. To explain it in an easier-to-understand way, an example will be given. Consider the case of learning a large amount of image data represented by white and black dots, and assume that the input is given as binary data of 0 or 1. Of course, the information that this data represents image data cannot be used. If the image data is 100 pixels, there are a total of 2 100 possibilities of patterns. When an infinite amount of input data is given, even if there are such a large number of patterns, the exact same data will be repeatedly input. At this time, there should be patterns with high and low frequencies of occurrence, and this is exactly what was described as the probability that the input data has earlier. There is a probability distribution corresponding to any input data set, and each input has its own unique probability. This probability alone is information that must be included in any input data regardless of its form.

[0139] (2.5: Rough Summary of Learning Principles) From the above, it can be said that learning is to estimate the probability that the input data has, and the solution to be obtained is a model function that returns the true probability of the input data. The loss function must be a function that is minimized for such a model function. In Chapter 3, a loss function that satisfies this condition will be derived mathematically, and it will also be shown that the normalization (standardization) of the estimated probability plays an important role. The complete form of the learning principle will be given in Chapter 4.

[0140] (Chapter 3: Derivation of the Loss Function) In this chapter, the loss function in the learning principle will be derived. Here, two patterns will be considered: the case where one wants to estimate the probability of the input itself, and the case where one wants to estimate the conditional probability defined by the input.

[0141] In this specification, a scalar may be denoted by a normal character (a non-bold character), and a vector may be represented by a bold character. Also, a matrix may be represented by a character with a tilde. However, in each description without a formula number, as long as there is no confusion, vectors and matrices may also be denoted by normal characters (neither bold nor with a tilde). This notation also applies to each of the above descriptions.

[0142] (3.1: General Case) Here, without considering the internal structure of the model, we focus only on the model function and the loss function. Consider a model that has the universal approximation property and returns a value Φ(x) for an input x. Φ(x) is a model function that is optimized as learning progresses, and we want Φ(x) to represent an estimate of the input probability P(x). Let the loss function to be used be L(Φ(x)), and consider the condition for it to be minimized. Minimizing the loss function means minimizing the expected value defined as follows. [Number] Here, when x takes continuous values, it is expressed as follows. [Number]

[0143] Next, let's consider the solution of Φ(x) that minimizes this expected value. However, if we simply try to minimize this, we can see that any Φ(x) will approach the same value that is the minimum point of L. To avoid this situation, the following normalization condition is imposed. When x takes discrete values, the normalization condition is [Number] and when x takes continuous values, the normalization condition is [Number] That is.

[0144] This normalization condition is reasonable if we consider that Φ(x) is expected to approach P(x). Next, we use the variational method to find the condition for minimizing the expected value of the loss function. First, we choose any two points x and x', and take the variation at these points as follows.

Number

Number

Number

Number

Number

Number

[0145] (3.2: Conditional Probability) In this section, we derive a loss function for estimating conditional probability. Here, we assume that the input is a discrete value, but in the case of a continuous value, it is exactly the same as below except that the sum is changed to an integral. Suppose the input x is composed of two vectors x = (a, b). Then P(x) can be decomposed as follows.

Number

Number

Number

Number

Number

Mathematics

Mathematics

Mathematics

Mathematics

[0146] (Chapter 4: Definition of Learning Principles) ----Learning Principles---- · Learning requires three elements: the object of optimization, input data, and an objective function. Borrowing terms from deep learning, the structure including the object of optimization is called a model, and the objective function is called a loss function. · The model can be considered as a function that takes an input and returns some output, which is called a model function. The model function estimates the probability of the input and must satisfy the probability normalization condition that it always takes positive values and the sum of all values is 1. · The loss function is defined by taking the logarithm of the estimated probability and attaching a negative sign. It has the same form as the mathematical formula for self-information. Optimization is performed to minimize this loss function, and it is guaranteed that the model function will automatically approach the true probability. ---------------- These are the learning principles derived in Chapters 2 and 3. All learning can be understood by this principle, and as long as this condition is met, it can be called learning. Since the form of the loss function is fixed, the only thing that can be devised is how to satisfy the normalization condition of the model function. Also, as long as the model function is always positive and satisfies the normalization condition, it can be estimated and interpreted as probability regardless of how it is defined. That is, it is not necessary for the model function to be defined in the form of the output in ordinary deep learning. The details of this point will be clarified in later chapters.

[0147] Hereafter, the discussion will focus on the normalization condition of the model function. In Chapter 5, it will be discussed how ordinary machine learning can be understood based on the learning principle, and it will be shown that supervised learning in particular can be identified with the estimation of conditional probability. In Chapter 6, a method to satisfy the normalization condition by defining the model function using differentiation will be proposed. By this method, unsupervised learning without prior knowledge for any dataset becomes possible, and machine learning with generality in an essential sense is realized. In Chapter 7, a method to satisfy the normalization condition by defining the model function based on the time evolution of a fully connected model will be proposed. In this method, a completely new concept of a loss function localized in time and space is derived, from which it can be identified that this model is a mathematical description of the learning mechanism of the brain. In Chapter 8, the results and findings brought about by the learning principle will be summarized, and their importance and validity will be reconfirmed.

[0148] (Chapter 5: Learning Principle for Some Problems) In this chapter, the learning principle will be applied to some problems to see how ordinary machine learning can be understood from the perspective of the learning principle. Before looking at individual problems, let's see how the commonly used loss functions can be understood from the perspective of the learning principle. The discussion in Chapter 3 is very general, and the relationships in (Equation 63) or (Equation 72) hold. Rewriting these relationships using the loss function gives the following. [Number] [Number] Here, let L be an arbitrary loss function, for example, the mean squared error. Here, exp(-L(x)) can be interpreted as a model function and needs to satisfy a certain normalization condition. Minimizing it using a certain loss function L(x) can be said to be equivalent to approximating the probability of the input by exp(-L(x)).

[0149] (5.1: Classification Problem) In a classification problem, usually, a set of input data {x 1 , x 2 , ···} and a set of teacher labels {t 1 , t 2 , ···} are given. If we consider the corresponding (x, t) as a set and regard it as one input data, the discussion in Section 3.2 can be applied. That is, it can be regarded as unsupervised learning for the data set {(x 1 , t 1 ), (x 2 , t 2 ), ···}. Thinking in this way, it becomes clear that the teacher label is just one numerical value included in the input data and is not an absolute correct answer. That is, different teacher labels may be assigned to exactly the same input. The combination of input data and teacher labels is considered to have a unique probability for each combination, and the goal is to clarify that probability. In traditional machine learning, there was a correct answer called the teacher label, and the goal was to create a model that made predictions close to it, but that way of thinking is fundamentally wrong. In a classification problem, the probability distribution we want to know is the conditional probability of the teacher label t when the input x is given, which is written as P(t|x). In this case, we consider a model function Φ(x, t) and define the loss function under the normalization condition as follows. [Number] Here, the normalization condition can be written as follows. [Number] This normalization condition is realized by using the following model. The model is a normal neural network as shown in Figure 4, where the input layer takes x as the input, and the output layer has the same number of nodes as the number of label types. In Figure 4, an example of classification regarding three labels is shown. The sum of the values of the nodes in the output layer is normalized by the softmax function, and Φ(x,t) is defined as the value of the corresponding node in the output layer. In this case, the loss function becomes exactly the same as the cross-entropy by the one-hot vector. Although the same loss function has been used in conventional machine learning, its true meaning was not to reduce the error from the teacher label but to estimate the conditional probability P(t|x).

[0150] (5.2: Regression Problem) The regression problem is different from the classification problem in that the teacher label is continuous. Consider the model function Φ(x,t) and define the loss function under the normalization condition as follows. [Number] Here, the normalization condition can be written as follows. [Number] (Equation 78) involves integration and the model described in Section 5.1 cannot be used. In this section, we will discuss prescriptions for satisfying this normalization condition considering several approximations and assumptions. General prescriptions without approximations and assumptions will be discussed in Section 6.4.

[0151] Let's assume that when x is fixed, P(t|x) has a unimodal distribution. Approximating this distribution by a Gaussian distribution, Φ(x,t) can be defined as follows. [Number] Here, n is the dimension of t, Σ with a tilde is an n×n matrix, and μ is a vector of the same dimension as t. Here, Σ with a tilde and μ are functions of x and are calculated through the model. An example for the case of n = 2 is shown in Figure 5. Here, μ = (μ 1 , μ 2 ). And

Number

[0152] It can be easily confirmed that (Equation 79) satisfies the normalization condition of (Equation 78). And the loss function is calculated as follows.

Number

Number

[0153] (5.3: Parameter Estimation Problem) {x 1 , ···, x NSuppose a dataset denoted as {}. If we assume that these datasets follow a specific distribution parameterized by certain variables, how should we estimate these variables? To solve such a problem, maximum likelihood estimation is often used, and its outline will be briefly explained. First, assume a probability distribution with unknown parameters. For example, if we consider that the dataset follows a Gaussian distribution, the mean and variance of the distribution would be good parameters. Let the set of parameters be θ, and let the probability distribution of x with θ as the parameter be P(x|θ). Our goal is to find the optimal θ that explains the dataset. This is done by maximizing the conditional probability of obtaining the dataset {x 1 , ···, x N} under the assumption of a certain θ. This means maximizing the likelihood L * . Here, the likelihood L * is given by

Number

Number

[0154] . Next, it will be shown that this problem can be solved based on the learning principle. In this problem, the original purpose is to find the true probability distribution P(x) that the dataset follows. That is, this problem is a kind of problem of finding probabilities by unsupervised learning described in Section 3.1. Using the conclusion there, consider the model function Φ(x) and define the loss function as follows under the normalization condition.

Number

Number

Equation

Equation

[0155] Comparing Equation 84 and Equation 85, the former takes the sum of all data in the dataset, while the latter performs calculations only for each data. The difference in signs is whether to maximize or minimize the value. According to the learning principle described in this specification, the parameter θ can be optimized one by one for each data, and there is no need to calculate the sum of the entire dataset together. Although the same θ will be obtained ultimately by either method, the method based on the learning principle is faster and more flexible because it does not require taking the sum.

[0156] (Chapter 6: Normalization by Differentiation) As we have seen so far, the loss function is always written as L(Φ(x)) = -log(Φ(x)), and the essence of the problem lies in how to make Φ(x) satisfy the normalization condition. If this normalization condition is satisfied, the learning principle can be easily achieved. When the normalization condition includes an integral, usually some assumptions or approximations are used. In this chapter, it is shown how to make Φ(x) satisfy the normalization condition without assumptions or approximations. Using the method discussed in this chapter, any probability distribution can be learned, and it can be said to be a learning with general applicability in an essential sense.

[0157] (6.1: Input by One Variable) First, for simplicity, consider the case where each input has only one variable. That is, a set of numerical values {x 1 , x 2 , ···} is given, and the probability distribution P(x) that they follow is estimated. In this case, our goal is to construct a model function Φ(x) that satisfies the following normalization condition without assumptions or approximations.

Number

Number

Number

Number

[0158] (The condition in Equation 91) is introduced to make Φ(x) positive and can be realized by using only positive weights and a monotonically increasing activation function. (The condition in Equation 92) can be realized by using the sigmoid function at the final node. And it can be confirmed that Φ(x) satisfies the normalization condition as follows.

Number

Number

[0159] To show how powerful this method is, the learning results for several datasets are shown in FIGS. 8, 9, and 10. FIG. 8 shows the probability estimation of a multimodal distribution. The sample size on the left is 4000, and the sample size on the right is 40. The true probabilities are the same for both. FIG. 9 shows the probability estimation of a flat distribution. The sample size on the left is 1000, and the sample size on the right is 10. The true probabilities are the same for both. FIG. 10 shows the probability estimation of a skewed distribution. The sample size on the left is 10000, and the sample size on the right is 100. The true probabilities are the same for both.

[0160] In these figures, samples are randomly selected from the probability distributions drawn as light gray lines and represented as histograms. The probability distributions are estimated from these samples using the method shown in this section and represented as dark gray lines. As shown in these figures, the fitting is very good regardless of the shape of the distribution and the sample size.

[0161] (6.2: Input with Multiple Variables (Multivariate)) Next, consider the case where each input contains multiple variables. Here, as the input dataset, {x 1 , x 2 , ···} is taken, and x = (a 1 , a 2 , ···, a n ) is assumed, where each a i takes continuous values. To obtain P(x) without assumptions or approximations, let's consider generalizing the method derived in the previous section for multivariate cases. Consider the model defined as shown in FIG. 11. The model shown in FIG. 11 is for inputs with multiple variables. And the model function Φ(x) is defined as follows. [Number] Here, y(x) is the value of the final layer. Similar to the previous section, the following conditions are imposed on y(x).

Number

Number

Number

Number

[0162] (6.3: Input of Multiple Variables and Output of Multiple Variables) The problem in the previous section is due to the fact that the definition of Φ(x) contains multiple derivatives. To solve this, it is proposed to use the Jacobian determinant as another generalization of (Equation 90) to multiple variables. For this purpose, consider the model shown in Figure 12. The model shown in Figure 12 is for an input with multiple variables and an output with multiple variables. And the model function Φ(x) is defined as follows.

Number

Mathematics

[0163] Here, the following conditions

Mathematics

Mathematics

Mathematics

Mathematics

Mathematics

Number

Number

Number

Number

Number

Number

[0164] (6.5: Considerations on Normalization by Differentiation) In this section, we describe notable properties of the method proposed in this chapter. Since we want to discuss the most general case, we focus on the model and loss function constructed in Section 6.3. For the intermediate layer of the model in Figure 13, (z 1 , ···, zd ) and name it. Here, d represents the depth of the layer. Each layer has n variables defined by z i =(c 1;i , c 2;i , ···, c n;i ). Using the chain rule of the Jacobian determinant, Φ(x) is calculated as follows.

Math

Math

Math

[0165] (Chapter 7: Normalization by Time Development) In this chapter, we focus on constructing a model that mimics the learning mechanism of the brain. As described in the previous chapter, learning is to estimate the probability of input, and the human brain is no exception. Considering the brain as a kind of machine learning model, there exists a model function Φ(x) for a given input x, and Φ(x) satisfies the following normalization condition. [Equation] The problem is how a model function that satisfies this normalization condition is defined in the brain, which becomes the focus of the discussion.

[0166] When creating a model that mimics the human brain, there are several conditions that must be met. First, the model must be general-purpose. The human brain processes information coming in through the five senses and automatically learns the rules behind it. That is, it is performing unsupervised learning without prior knowledge. Second, the nodes and edges are uniformly isotropic, and there should be no special nodes or special edges. This is based on the observational fact that the neurons in the brain seem to form a fully connected or locally connected network. Neurons in the brain do not necessarily form layers and may have loop structures. In that sense, the model discussed in Chapter 6, although general-purpose, cannot be a model of the human brain. Based on this requirement, consider a fully connected (locally connected) model as shown in Figure 15. As an analogy to the fact that synapses have directions in the brain, assume that each edge has a direction and information is transmitted only in that direction. Of course, the edges can be connected in any way, and there may be no edge between certain nodes or they may be connected in both directions. As will be shown later, by considering the time development of such a model, Φ(x) that satisfies the normalization condition can be constructed. During the time development, each node is affected by other nodes and its internal state changes, but this flow itself is considered as a machine learning model. In the next section, Φ(x) will be specifically constructed and it will be confirmed that the normalization condition is satisfied.

[0167] (7.1: Linear Model) Before getting into the detailed calculations, first define the variables of the model in Fig. 15. Assume that the model has n nodes, which change over time. Denote them as a(t) = (a 1 (t), ···, a n (t)). Here, t is a variable representing time, and each a i (t) takes real values. Next, each edge has a weight represented by a real number, and denote the weight of the edge from the j-th node to the i-th node as w ij . Here, self-connections of nodes are not assumed, so accordingly, w ii = 0. When representing w ij collectively as a matrix, use the symbol Ŵ. Finally, each node has a constant bias value, which is defined as b = (b 1 , ···, b n ). Among these variables, only a(t) changes over time. Ŵ and b are the objects to be optimized.

[0168] Next, define the time evolution of this model. First, the input x is input into the model at t = 0, and the time evolution starts from there. Since this model does not have a special input layer, some nodes must be assigned as inputs. Assume that x is composed of m variables represented as x = (x 1 , ···, x m ), and substitute (a 1 (0), ···, a m (0)) = (x 1, ···, x m ). Assign (a m+1 (0), ···, a n (0)) = (r m+1 , ···, r n ) to the remaining nodes. Here, r i is a random value following a certain probability distribution and is randomly selected each time.

[0169] Since the initial conditions have been established, equations to describe the time evolution are needed. They are defined as follows.

Number

Number

Number

Number

Number

Number

[0170] The loss function is calculated as follows. [Mathematics] Here, since the first term is a constant independent of a(0), it can be ignored. For example, if E(a)=a 2 , then the loss function can be written as follows. [Mathematics] And a local loss function can be defined as follows. [Mathematics] This model seems good, but actually there is a problem. That is, the transformation of a(t) is merely linear and does not have the universal approximation property. In fact, if a * (t)=a(t)+(tilde - marked W) -1 b, then the following equation is obtained. [Mathematics] This shows that a * (T) is completely linear with respect to a * (0). To solve this problem, it is necessary to add non - linearity, and the prescription for that will be described in the next section.

[0171] (7.2: Non - linear Model) In normal machine learning, nonlinearity is added using an activation function, but we will consider how to generalize the activation function in this time evolution model. For example, the commonly used sigmoid function is σ(x)=(1+exp(-x))- 1 This function has the role of converting inputs from -∞ to ∞ into outputs from 0 to 1, and from this it can be said that the essence of an activation function is to restrict the range of output values. In this model, nonlinearity is added by making a modification to restrict the value of a(t) to 0 to 1. Specifically, this is done by restricting the input value a(0) to 0 to 1, and further modifying the time evolution equation as follows:

number

[0172] ​We have added non-linearity to the time evolution equation, but this differential equation has become unsolvable and the method used in Section 7.1 is not applicable. On the other hand, since the range of a(t) is 0 < a(t) < 1, the method introduced in Chapter 6 can be applied. Therefore, the model function Φ(a(0)) is defined as follows.

Number

Number

Number

Number

Number

Number

Number

[0173] (7.3: Considerations on Normalization by Time Development) The model described here was constructed with reference to the structure of the brain. In this section, we will deepen our understanding while comparing it with the actual brain.

[0174] First, let's consider the hardware aspects of the actual brain. The brain is an object made of proteins and should function only according to physical laws. In physics, it is known that physical laws are described by local interactions, and any object is affected only locally. Neurons in the brain are no exception and should be affected only by local interactions. That is, the optimization of neurons also needs to be performed locally, and the loss function also needs to be defined locally. The localization of the loss function is required by physical laws, and our model follows this.

[0175] Next, let's consider the software aspects of the actual brain. The most amazing characteristic of the brain is its versatility. It can learn all concepts, and for example, it can also learn causal relationships such as "B occurs because A occurs". Our brains are constantly receiving information from the five senses and have come to find causal relationships by sequentially processing such temporally continuous information. The fact that information is sequentially processed means that optimization is performed for each time, and this property is common to our model. For example, if data such as a video is sequentially input into this model as temporally continuous image data, the concept of causal relationship will be automatically acquired.

[0176] Based on these, it can be said that this model mathematically realizes the learning mechanism of the brain. However, two major problems remain. One is that even if this model is mathematically correct, it is not clear how it is actually realized in the brain. Neurons in the brain contain various substances and electrical signals, and it is necessary to clarify how they are related to the variables of the model. The other is how to put this model into practical use. There is still not enough knowledge for specific implementation, such as how many nodes are sufficient, what time resolution is required for time evolution, and what is an efficient method for optimizing parameters. If these problems are solved, this model could become a general-purpose artificial intelligence that can learn from all kinds of data like humans.

[0177] (Chapter 8: Discussion) In this chapter, we will consider the meaning of the learning principle and the interpretations derived from it. First, we will explain the concept of features, which is considered important in ordinary machine learning, from the perspective of the learning principle. Based on the learning principle, all information is contained in the probability distribution of the input dataset. Moreover, features correspond to parts where the probability is maximized in a certain phase space. For example, when considering the task of handwritten digit recognition, each digit is composed of features such as how the lines are drawn. If the curved part of the digit '3' is represented by a straight line or written separately from above and below, it will become difficult to recognize it as '3'. This is due to the fact that the normal form of the digit is frequently seen, while the deformed form is rarely seen. That is, what can be regarded as a feature has a higher occurrence probability than its deformed form. What is important here is that although there are countless patterns of deformation, the probability is concentrated only in a very small number of patterns that are frequently seen. This bias in probability is the essence of features.

[0178] Next, the reasons for the success of deep learning will be explained. According to the learning principle, as long as the model function satisfies the normalization condition, it will automatically approach the true probability. In particular, if the model has the universal approximation property and there are infinite datasets and computational resources, it can approach the true probability distribution infinitely closely. That is to say, the success of deep learning is due to obtaining the universal approximation property by making the layers deeper and the improvement of the processing power of current computers. Deep learning is sometimes explained as being successful because it can capture complex features, but features are just an afterthought concept that appears when the estimation of the probability distribution is successful.

[0179] Then, what does it mean to improve the model in deep learning will be explained. According to the learning principle, as long as the model has the universal approximation property, it will ultimately reach the same solution regardless of the model structure. What changes depending on the model structure is the convergence speed to the solution. For example, if the characteristics that image data or language data may have are incorporated into the model structure in advance, the convergence speed to the solution will be overwhelmingly faster for such specific data. And in reality, since there are no infinite datasets and computational resources, learning must be interrupted halfway, and a model with a fast convergence will achieve good results. This has become a major factor in the big misunderstanding that the essence of learning lies in the model structure. Improving the model means specializing it for specific inputs.

[0180] Finally, two methods we devised will be described. Both methods are based on the learning principle and satisfy the normalization condition without assumptions or approximations. This shows that it is a general-purpose model that can be used without specifying the type of input data. In fact, the results in Figures 8, 9, and 10 show the generality of the method using differentiation. It was also shown that the method using the time evolution for the fully connected model can be regarded as a mathematical realization of the learning mechanism of the brain. These results are the consequences of the learning principle and, conversely, guarantee the correctness of the learning principle.

[0181] (Chapter 9: Conclusion) In this paper, a learning principle that uniformly describes all learning, including machine learning and brain learning, was derived. Under this principle, all learning is understood to be the probability estimation of input data. The conditions for applying this principle are that the estimated probability is always positive and satisfies the normalization condition, and optimization is performed using a loss function defined by the logarithm of the estimated probability. Conversely, if these conditions are met, anything can be regarded as learning. It was confirmed that supervised learning also satisfies the learning principle and can be regarded as a type of probability estimation of conditional probability.

[0182] In the learning principle, the loss function is always defined as the negative of the logarithm of the model function, and what can be changed is the method of satisfying the normalization condition of the model function. In this specification, two methods of satisfying the normalization condition without assumptions or approximations were proposed. Both methods are considered to be very general methods that can perform learning with high accuracy for any dataset. The first method is a method of satisfying the normalization condition using differentiation, and it was shown that it actually learns on several datasets and exhibits very good behavior. The second method is a method of satisfying the normalization condition by considering the time change of a fully connected model. This model is based on the structure of neurons and synapses in the brain, and it is significant to show that a function that satisfies the normalization condition can be constructed from such a model. Furthermore, it was shown that the loss function defined there is expressed as the sum of functions localized in space and time. This means that the optimization of parameters can be performed sequentially and locally without using the backpropagation method. In the actual brain, optimization must also be performed sequentially and locally as a consequence of the laws of physics. Considering this fact, it can be said that this method mathematically realizes the learning mechanism of the brain. This will be a major step towards the realization of general-purpose artificial intelligence.

[0183] Furthermore, from the perspective of learning principles, the conventional learning was reexamined and a new understanding was provided. In learning, all information is included in the probability distribution of the input, and concepts such as features are also defined based on probability. Improving the model corresponds to accelerating the convergence to the solution by specializing it for a specific dataset. In aiming for the generalization of machine learning, it would be important to understand learning itself based on learning principles rather than ad-hoc countermeasures.

[0184] [Example of implementation by software] The functions of information processing apparatuses 1 and 2 (hereinafter referred to as "apparatuses") can be realized by a program for causing a computer to function as the apparatus, and by a program for causing a computer to function as each control block (particularly each part included in control unit 10) of the apparatus.

[0185] In this case, the above apparatus includes, as hardware for executing the above program, a computer having at least one control device (for example, a processor) and at least one storage device (for example, a memory). By executing the above program by this control device and storage device, each function described in the above embodiments is realized.

[0186] The above program may be recorded on one or more computer-readable recording media, not temporarily. This recording medium may or may not be provided in the above apparatus. In the latter case, the above program may be supplied to the above apparatus via any wired or wireless transmission medium.

[0187] Also, part or all of the functions of each of the above control blocks can also be realized by a logic circuit. For example, an integrated circuit in which a logic circuit functioning as each of the above control blocks is formed is also included in the scope of the present invention. In addition to this, for example, it is also possible to realize the functions of each of the above control blocks by a quantum computer.

[0188] (Summary) The following configurations are described in this specification at least.

[0189] (Configuration A1) An acquisition unit that acquires one or more input values, A learning unit that executes a learning process with reference to a loss function defined by the logarithm of a model function that takes the one or more input values as arguments and satisfies a predetermined normalization condition An information processing apparatus comprising the same.

[0190] (Configuration A2) The learning process includes a process of optimizing the model function by minimizing or maximizing the loss function The information processing apparatus according to Configuration A1.

[0191] (Configuration A3) The input value includes one or more data values and a teacher label associated with the data value, The normalization condition is a sum of the model functions and includes a condition regarding the sum with respect to the teacher label in the arguments of the model functions The information processing apparatus according to Configuration A1 or A2.

[0192] (Configuration A4) The input value includes one or more data values and a teacher label associated with the data value, The normalization condition is an integral of the model functions and includes a condition regarding the integral with respect to the teacher label in the arguments of the model functions The information processing apparatus according to Configuration A1 or A2.

[0193] (Configuration A5) The model function is a distribution function defined including one or more parameters, The loss function is defined including the one or more parameters, The learning process The process of determining the one or more parameters by minimizing or maximizing the loss function is included. The information processing apparatus according to Configuration A4.

[0194] (Configuration A6) The model function is a distribution function having the one or more input values as arguments, or the one or more parameters themselves that define the distribution function The information processing apparatus according to Configuration A1 or A2.

[0195] (Configuration A7) The model function is a derivative of a certain function having the one or more input values as arguments, and is a derivative obtained by one or more order differentiations according to the dimension of the input values, The predetermined normalization condition is defined by boundary values of the certain function. The information processing apparatus according to any one of Configurations A1 to A6.

[0196] (Configuration A8) The model function is a Jacobian defined by at least a part of the plurality of input values and a plurality of functions having the plurality of input values as arguments, The predetermined normalization condition is defined by boundary values of the plurality of functions. The information processing apparatus according to any one of Configurations A1 to A6.

[0197] (Configuration A9) The Jacobian is the determinant of a triangular matrix The loss function is the sum of the logarithms of each of the diagonal components of the triangular matrix. The information processing apparatus according to Configuration A8.

[0198] (Configuration A10) The learning unit is For a target node belonging to a neural network including one or more nodes in each layer, apply learning processing with reference to a local loss function defined by the target node and nodes belonging to the layer adjacent to the target node. The information processing apparatus according to any one of Configurations A1 to A9.

[0199] (Configuration A11) The local loss function is a variable associated with the target node, and a variable associated with a node belonging to the layer adjacent to the target node and is expressed using the logarithm of the derivative defined by The information processing apparatus according to Configuration A10.

[0200] (Configuration B1) A prediction unit that executes prediction processing using a model function that is a model function having one or more input values as arguments and is learned with reference to a loss function defined by the logarithm of a model function that satisfies a predetermined normalization condition, and an output unit that outputs a prediction result of the prediction processing An information processing apparatus comprising.

[0201] (Configuration C1) A network including a plurality of nodes directly or indirectly connected to each other, and a learning unit that learns each of the plurality of nodes using a loss function defined to be spatially and temporally localized with respect to the node An information processing apparatus comprising.

[0202] (Configuration D1) Including a plurality of cells directly or indirectly connected to each other, each of the plurality of cells is learned using a loss function defined to be spatially and temporally localized with respect to the cell Brain model.

[0203] (Configuration E1) An acquisition step of acquiring one or more input values, A learning step of executing a learning process with reference to a loss function defined by the logarithm of a model function that takes the one or more input values as arguments and satisfies a predetermined normalization condition, and An information processing method including the same.

[0204] (Configuration F1) A prediction step of executing a prediction process using a model function learned with reference to a loss function defined by the logarithm of a model function that takes one or more input values as arguments and satisfies a predetermined normalization condition, and An output step of outputting a prediction result by the prediction process, and An information processing method including the same.

[0205] (Configuration G1) A program for causing a computer to function as the information processing apparatus according to Configuration A1, the program for causing the computer to function as the acquisition unit and the learning unit.

[0206] (Configuration H1) A program for causing a computer to function as the information processing apparatus according to Configuration B1, the program for causing the computer to function as the prediction unit and the output unit.

[0207] The present invention is not limited to the above-described embodiments, and various modifications are possible within the scope shown in the claims. Embodiments obtained by appropriately combining the technical means disclosed in different embodiments are also included in the technical scope of the present invention.

Explanation of Reference Numerals

[0208] 1, 2 ··· Information processing apparatus 11 ··· Acquisition unit 12 ··· Learning unit 13 ··· Prediction unit 14 ··· Output unit

Claims

1. An acquisition unit that acquires one or more input values; A learning unit that executes a learning process with reference to a loss function defined by the logarithm of a model function that takes the one or more input values as arguments and satisfies a predetermined normalization condition An information processing apparatus comprising the above.

2. The learning process includes: A process of optimizing the model function by minimizing or maximizing the loss function The information processing apparatus according to Claim 1.

3. The input values include one or more data values and teacher labels associated with the data values, The normalization condition includes: An integral of the model function, which includes a condition regarding an integral with respect to the teacher label in the arguments of the model function The information processing apparatus according to Claim 1.

4. The model function is a distribution function defined including one or more parameters, The loss function is defined including the one or more parameters, The learning process includes: A process of determining the one or more parameters by minimizing or maximizing the loss function The information processing apparatus according to Claim 3.

5. The model function is: A distribution function taking the one or more input values as arguments, or The one or more parameters themselves that define the distribution function The information processing apparatus according to Claim 1 or 2.

6. The model function is a derivative of a certain function that takes the one or more input values as arguments, and is a derivative obtained by one or more order differentials according to the dimension of the input values, The predetermined normalization condition is defined by the boundary values of the certain function The information processing apparatus according to Claim 1 or 2.

7. The model function is: A Jacobian defined by at least a part of the plurality of input values and a plurality of functions that take the plurality of input values as arguments, The predetermined normalization condition is defined by the boundary values of the plurality of functions The information processing apparatus according to Claim 1 or 2.

8. The Jacobian is the determinant of a triangular matrix The loss function is the sum of the logarithms of each of the diagonal components of the triangular matrix The information processing apparatus according to Claim 7.

9. The learning unit: Applies a learning process with reference to a local loss function defined by a target node belonging to a neural network including one or more nodes in each layer and the target node and nodes belonging to a layer adjacent to the target node The information processing apparatus according to claim 1 or 2.

10. The local loss function is the variable associated with the target node and the variable associated with the nodes belonging to the layer adjacent to the target node expressed using the logarithm of the derivative defined by The information processing apparatus according to claim 9.

11. A prediction unit that executes prediction processing using a model function having one or more input values as arguments and learned with reference to a loss function defined by the logarithm of a model function that satisfies a predetermined normalization condition, an output unit that outputs a prediction result by the prediction processing An information processing apparatus comprising:

12. A network including a plurality of nodes directly or indirectly connected to each other, a learning unit that learns each of the plurality of nodes using a loss function defined to be spatially and temporally localized to the node An information processing apparatus comprising:

13. Including a plurality of cells directly or indirectly connected to each other, each of the plurality of cells is learned using a loss function defined to be spatially and temporally localized to the cell Brain model.

14. An acquisition step of acquiring one or more input values, a learning step of executing learning processing with reference to a loss function defined by the logarithm of a model function having the one or more input values as arguments and satisfying a predetermined normalization condition An information processing method comprising:

15. A prediction step of executing prediction processing using a model function having one or more input values as arguments and learned with reference to a loss function defined by the logarithm of a model function that satisfies a predetermined normalization condition, an output step of outputting a prediction result by the prediction processing An information processing method comprising:

16. A program for causing a computer to function as the information processing apparatus according to claim 1, and a program for causing a computer to function as the acquisition unit and the learning unit.

17. A program for causing a computer to function as the information processing apparatus according to claim 11, and a program for causing a computer to function as the prediction unit and the output unit.