Calculation units, integrated circuits, machine learning devices, discrimination devices, equation utilization devices, and calculation methods.

The neural network structure addresses the challenge of representing higher-order power exponents by incorporating power exponents and learning functions, enabling accurate correlation derivation in complex systems.

JP7832427B2Active Publication Date: 2026-03-18大庭富美男
View PDF 5 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2023-02-10
Publication Date
2026-03-18

AI Technical Summary

Technical Problem

Conventional neural network structures struggle to accurately represent and derive exact solutions for phenomena involving higher-order power exponents, particularly in group systems like water flow, molecules of electric current, and wave equations with electron interference, due to the fixed power exponents and inability to handle statistical distributions.

Method used

A neural network structure incorporating power exponents and learning functions that raise input data to specific powers, combined with weighted feature amounts, allows for accurate derivation of correlations between input and output in phenomena with interference, using a transformation layer and hidden layers to process these exponents and functions.

Benefits of technology

Enables accurate handling of phenomena expressed by power exponents and learning functions, allowing for precise derivation of correlation relationships between input and output, even in complex systems with interference.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007832427000001
    Figure 0007832427000001
  • Figure 0007832427000002
    Figure 0007832427000002
  • Figure 0007832427000003
    Figure 0007832427000003
Patent Text Reader

Abstract

To provide an arithmetic device (machine learning device) which can derive an equation.SOLUTION: An input layer of an arithmetic device includes a plurality of power exponents (p0, p1,..., pN) which are respectively associated with a plurality of input data, and respectively exponentiate the plurality of input data, and a coefficient of a predetermined function for learning, as a learning parameter of a neural network structure. An output layer of the arithmetic device outputs an output value (y=f (YY0*De)), on the basis of products of the products (YY0=D0p0*D1p1*...*DNpN) of a plurality of powers (D0p0,D1p1,...,DNpN) in which the plurality of input data inputted to the input layer are respectively exponentiated by the plurality of power exponents and the function for learning (De=g(D0,D1,...,DN)).SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an arithmetic device using a neural network, an integrated circuit, a machine learning device, a discrimination device, an equation utilization device, and an arithmetic method.

Background Art

[0002] In recent years, machine learning has been applied to various fields. In particular, the neural network structure has been widely applied to both regression problems and classification problems. In such a neural network structure, weight coefficients are respectively multiplied by a plurality of input data input to an input layer, and an output value based on the result of calculating their sum is output from an output layer (see, for example, Patent Document 1, Patent Document 2, etc.). However, since there is a drawback that the output result is difficult for people to understand, Patent Document 3 that adopts an addition type operation and a differential type operation method has been published. Further, Patent Document 4 that adopts a power exponential type neural network having an equation of a product of power values in an output format has been published, thereby achieving an output format that is easy for people to understand and an improvement in output accuracy.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Patent Document 2

Patent Document 3

Patent Document 4

Non-Patent Documents

[0004]

Non-Patent Document 1

[0005] The neural network structures described in Patent Documents 1, 2, and 3 above perform machine learning by adjusting weight coefficients, but the "power exponents" for multiple input data are fixed, for example, to "1," and there was a problem that they could not represent any exact solutions to output equations with higher-order power exponents. Patent Document 4 attempts to solve this problem by incorporating power exponents into the features of the neural network. According to Patent Document 4, it is possible to derive equations from data that are expressed as the product of power values ​​of individual systems where the individuals do not interfere with each other, such as geometric problems, physics problems like inverted pendulums, and equations of motion. However, there was a problem that the exact solution equation could not be derived for group systems such as water flow, molecules of electric current, and wave equations involving the interference of electrons, because the solution includes a statistical distribution function containing Napier's number (e).

[0006] In view of the above-mentioned problems, the present invention aims to provide a computing device, integrated circuit, machine learning device, discrimination device, equation utilization device, and computing method that prepare a learning function that incorporates several types of functions for handling phenomena expressed by power exponents, and that enables accurate derivation of the correlation between input and output even in phenomena involving interference between individuals. [Means for solving the problem]

[0007] To achieve the above objective, an arithmetic device according to one aspect of the present invention is an arithmetic device that uses a neural network structure including at least an input layer and an output layer to output an output value from the output layer for a plurality of input data (D0, D1, ..., DN) input to the input layer, The aforementioned input layer is Associated with each of the plurality of input data, respectively, and having, as learning parameters of the neural network structure, a plurality of power exponents (p0, p1, …, pN) for raising each of the plurality of input data to a power, decimal weighted feature amounts (w0, w1, …, wN), and coefficients of a predetermined learning function. The output layer The plurality of power values (D0 pN+wN , p0+w0 , p1+w1 , D1 p1+w1 , …, DN pN+wN ) obtained by raising each of the plurality of input data input to the input layer to a power by the plurality of power exponents and weighted feature amounts, respectively, (D0 p0+w0 *D1 p1+w1 *…*DN pN+wN ), and based on the product of the learning function (De = g(q, Dn)), outputs the output value (ZZ = f(D0 p0+w0 *D1 p1+w1 *…*DN pN+wN *De)).

[0008] An arithmetic method according to another aspect of the present invention is an arithmetic method using an arithmetic device that operates to output an output value from an output layer for a plurality of input data (D0, D1, …, DN) input to an input layer using a neural network structure including at least the input layer and the output layer. The method includes inputting the plurality of input data (D0, D1, …, DN) to the input layer having, as learning parameters of the neural network structure, a plurality of power exponents (p0, p1, …, pN) associated with each of the plurality of input data for raising each of the plurality of input data to a power, respectively, and coefficients of a predetermined learning function. Inputting, to the arithmetic device, from the output layer The product (YY0 = D0 p0 , D1 p1 , …, DN pN ) of the plurality of power values (D0 p0 *D1 p1 *…*DN pNBased on the product of ) and the learning function (De=g(D0,D1,…,DN)), the output value (y=f(YY0*De)) is output. [Effects of the Invention]

[0009] According to the present invention, it is possible to handle phenomena expressed by power exponents and learning functions, and to accurately derive the correlation relationship that exists between the input and output in such phenomena.

[0010] Other issues, configurations, and effects will be clarified in the embodiments for carrying out the invention described later. [Brief explanation of the drawing]

[0011] [Figure 1] This figure illustrates the neural network structure 100D used by the computing device according to the basic embodiment of the present invention and its basic principle. [Figure 2] This is a block diagram showing the configuration of a computing device 1 using a neural network structure according to a basic embodiment of the present invention. [Figure 3] This diagram illustrates the problem of vanishing gradients. [Figure 4] This is an Arrhenius plot diagram according to the first embodiment, in which the average lifespan L (average time to failure) is shown as the output value on the vertical axis and the temperature (1000 / T) is shown on the horizontal axis. [Figure 5] This is a graph according to the first embodiment, with the average lifespan L (average time to failure) as the output value on the vertical axis and voltage on the horizontal axis. [Figure 6] This diagram shows a table of input data values ​​based on Figures 4 and 5 according to the first embodiment. [Figure 7] This table shows specific examples of learning functions (Eureka functions) according to the basic embodiment of the present invention. [Figure 8] This is a list of coefficients of variation CV (loss) in ascending order according to the first embodiment. [Figure 9] This is a table listing the values ​​of ZZ' calculated by the output formula (relational formula) [Equation 3-9] according to the first embodiment. [Figure 10] This is a correlation graph according to the first embodiment, in which the measured value of average lifespan L is shown on the horizontal axis and the lifespan calculation (prediction) value based on equation [Equation 3-10] is shown on the vertical axis. [Figure 11] This is a graph according to the second embodiment, with wavelength λ on the horizontal axis, intensity E on the vertical axis, and temperature T as parameters. [Figure 12] This is a table showing the experimental wavelength λ, experimental temperature T, and experimental intensity E extracted from [Figure 11] according to the second embodiment. [Figure 13] This is a list of coefficients of variation CV (loss) in ascending order according to the second embodiment. [Figure 14] In the second embodiment, the horizontal axis is wavelength λ and the vertical axis is intensity E, with the measurement data from Figure 12 displayed as dots and the values ​​of equation [Equation 3-20] as solid lines. [Figure 15] Table (1) of Iris data according to the third embodiment. [Figure 16] Table (2) of the Iris data according to the third embodiment. [Figure 17] This is a correlation graph of calyx length-calyx width and calyx length-petal length for Iris data according to the third embodiment. [Figure 18] This is a table showing the top 30 items in a descending order of reward (discrimination rate) according to the third embodiment. [Figure 19] This is a graph according to the third embodiment, in which the flower number of iris is plotted on the horizontal axis and the value of the discriminant AA is plotted on the vertical axis. [Figure 20] This is a list of the results of searching for discriminant formulas AB and AC according to the third embodiment. [Figure 21] This is a graph according to the third embodiment, in which the flower number of iris is plotted on the horizontal axis and the value of the discriminant AB is plotted on the vertical axis. [Figure 22] This is a graph according to the third embodiment, in which the flower number of iris is plotted on the horizontal axis and the value of the discriminant AC is plotted on the vertical axis. [Figure 23]This is a table showing the classification of training data for distinguishing Iris flowers using the majority-vote discriminant formula Z=MODE(AA,AB,AC) according to the third embodiment. [Modes for carrying out the invention]

[0012] The present invention will be described below with reference to the drawings, divided into a "basic form" that shows the fundamental principle of the present invention and "embodiments" for implementing the present invention by applying that basic principle.

[0013] (basic form) Figure 1 illustrates the neural network structure 100D used by the computing device according to the basic embodiment of the present invention and its basic principle.

[0014] The computing unit (reference numeral 1 in Figure 2) is a device that uses a neural network structure 100D, which includes at least an input layer 110D and an output layer 120D, to output an output value ZZ from the output layer 120D for multiple input data Dn=(D0,D1,…,DN) input to the input layer 110D.

[0015] The neural network structure 100D shown in Figure 1 consists of an input layer 110D having N+1 dimensions (where N is a natural number greater than or equal to 1) of neurons (nodes) and an output layer 120D having one neuron (node). A transformation layer 140 and a hidden layer 130 are included between the input layer 110D and the output layer 120D.

[0016] Each of the N+1 neurons in the input layer 110D is associated with an N+1-dimensional input data Dn=(D0,D1,…,DN), and each of them receives the N+1-dimensional input data Dn as input and outputs it to the transformation layer 140.

[0017] The transformation layer 140 has N+1 dimensional exponents pn=(p0,p1,···,pN) which are obtained by raising each of the N+1 dimensional input data Dn from the input layer 110D to a power, as learning parameters for the neural network structure 100D. Furthermore, the transformation layer 140 is composed of N+2 dimensional nodes, each having a learning function node 141 that takes multiple learning function values ​​De(=g(q,Dn)) as input.

[0018] The value De(=g(q, Dn)) of the learning function is the output value of the learning function, which has the coefficient q of the learning function as a learning parameter and the product of this coefficient with the N+1-dimensional input data Dn of the input layer 110D as a parameter.

[0019] The outputs of the N+1 neurons in the transformation layer 140 correspond to the power values ​​(D0) of the N+1-dimensional input data Dn=(D0,D1,…,DN), which are each raised to powers of the N+1-dimensional exponents pn=(p0,p1,…,pN). p0 ,D1 p1 ,…,DN pN It is converted to ( ). Furthermore, the output of N+1 neurons is concatenated with the value De of the learning function and converted to N+2 dimensional input data (D0 p0 ,D1 p1 ,…,DN pN It is output as ,De) to the hidden layer input node 133 of the hidden layer 130.

[0020] The hidden layer 130 has a first hidden node 131 and a second hidden node 132. The first hidden node 131 receives the input N+2 dimensional input data (D0) of the hidden layer input node 133 via N+1 dimensional weighting parameters wn=(w0,w1,…,wN) as learning parameters. p0 ,D1 p1 ,…,DN pN The first hidden node 132 receives the bias parameter b as a learning parameter and outputs the target value YY defined by the following equation [Equation 3-1] to the output layer 120D. The second hidden node 132 receives the bias parameter b as a learning parameter and outputs the additive operation output BYA defined by the following equation [Equation 3-2] to the output layer 120D.

[0021] The output layer 120D outputs an output value ZZ (=f(YY,BYA)) based on the target value YY and the additive calculation output BYA.

[0022] [Math 3-1] YY=D0 p0 *D0 p1 *…*DN pN *W0*W1*…*WN*De [Math 3-2] BYA = B * (base) (Σ[n=0→N](wn*pn*dn)) However, the parameters in the above formula are as follows: base is a positive number other than 1. Dn=base dn (n=0,1,…,N) : Input data pn(p0,p1,…,pN): Power exponent Dn pn : Power value De=g(q,Dn): Training function value wn=log base Wn(n=0,1,…,N): Weighting parameters (Wn=base wn ) W = W0 * W1 * ... * WN: Product of weight parameters b = log base B: Bias parameter (B=base b ) YY: Target value BYA: Additive arithmetic output

[0023] The N+1-dimensional exponent pn, the N+1-dimensional weighting parameter wn, the bias parameter b, and the coefficient q of the learning function are parameters that are learned by using multiple input data Dn as training data.

[0024] The N+1-dimensional exponent pn, the N+1-dimensional weighting parameter wn, the bias parameter b, and the coefficient q of the learning function are adjusted so that a predetermined loss function (coefficient of variation) is small when the N+1-dimensional input data Dn as learning data is input to the input layer 110D, and the summation operation output BYA output from the second hidden node 132 is small.

[0025] The computing unit 1 determines that a predetermined learning termination condition has been met when it has repeated the series of steps of adjusting (searching for) the learning parameters using the learning data a predetermined number of times, or when the value of the loss function (loss amount) becomes smaller than a predetermined tolerance value, and terminates learning for the learning parameters. This realizes a trained neural network structure 100D having an N+1-dimensional exponent pn, an N+1-dimensional weighting parameter wn, a bias parameter b, and a coefficient q of the learning function as learning parameters.

[0026] According to the neural network structure 100D used by the computing device in this basic form, the hidden layer 130 has a first hidden node 131 that outputs a target value defined by the above formula [Equation 3-1] to the output layer, and a second hidden node 132 that outputs an additive calculation output defined by the above formula [Equation 3-2] to the output layer, and the output layer 120D outputs an output value based on the target value and the additive calculation output. Therefore, the computing device 1 can handle phenomena expressed by exponents and learning functions, and can accurately derive the correlation relationship that holds between input and output in said phenomena.

[0027] When the learning function, which is a characteristic of this basic form, is not used, the value of the learning function De is set to 1. In this case, consistency is maintained with Patent Document 4, which does not use the learning function.

[0028] (Basic device configuration) Figure 2 is a block diagram showing the configuration of a computing device 1 using a neural network structure according to the basic embodiment of the present invention.

[0029] The computing unit 1 functions as a machine learning device 1A that generates a learning model having a basic neural network structure 100D, and a discrimination device 1B that outputs a discrimination result AA for the discrimination data BB to be discriminated using the learning model generated by the machine learning device 1A. The machine learning device 1A is used in the learning phase, and the discrimination device 1B is used in the discrimination phase (inference phase).

[0030] The arithmetic unit 1 is configured to include, as its components, a discriminator learning unit 2, a learning parameter storage unit 3, a learning data storage unit 4, a learning data processing unit 5, a discriminant result processing unit 6, a discriminant data acquisition unit 7, and a learning function storage unit 8.

[0031] The discriminator learning unit 2 comprises a learning unit 20 that learns learning parameters using a learning model having a neural network structure 100D, and a discriminant processing unit 21 that outputs a discriminant result for discriminant data using a learning model that reflects the learning parameters being learned or have been learned. The learning parameters for the basic form are an N+1-dimensional power exponent pn, an N+1-dimensional weighting parameter wn, a coefficient q of the learning function, and a bias parameter b.

[0032] The learning parameter storage unit 3 stores the learning parameters and the organization number of the learning function as learning results performed by the learning unit 20 during the learning phase. The initial values ​​of the learning parameters are stored in the learning parameter storage unit 3 through the learning parameter initialization process, and the learning parameters are sequentially updated as learning is repeatedly performed by the learning unit 20. The learning parameter storage unit 3 then stores the learning parameters and the organization number of the learning function when learning by the learning unit 20 is completed, and these are read out by the discrimination processing unit 21 during the discrimination phase (inference phase).

[0033] The learning data storage unit 4 stores multiple sets of learning data, each containing at least multiple input data. This learning data may include input data and training data associated with that input data. In the case of unsupervised learning, the learning data may consist only of input data. The training data is, for example, data corresponding to the classification result. If the classification result is represented by "0" for normal and "1" for abnormal, then either "0" or "1" is set.

[0034] The learning function memory unit 8 stores multiple pre-set learning functions and multiple sets of associated learning function reference numbers.

[0035] The learning unit 20 inputs the learning data stored in the learning data storage unit 4 and the multiple learning functions stored in the learning function storage unit 8 to the learning model via the learning data processing unit 5, and performs learning of the learning parameters, for example, so that the loss function is minimized. That is, the learning unit 20 receives the discrimination result output from the discrimination processing unit 21 and the learning data read from the learning data processing unit 5 as input, performs learning using this data, and stores the learning parameters and the organization numbers of the learning functions in the learning parameter storage unit 3. Furthermore, when performing supervised learning, the learning unit 20 performs learning to adjust the output value output from the output layer when the multiple input data included in the learning data is input to the input layer, so that the difference between the output value and the training data included in the learning data becomes small.

[0036] In the learning phase, the discrimination processing unit 21 inputs the learning data acquired by the learning data processing unit 5 into a learning model that reflects the initial values ​​or learning parameters during learning, and outputs the discrimination result to the learning unit 20 and the discrimination result processing unit 6 based on the output values ​​from the learning model.

[0037] Furthermore, in the discrimination phase (inference phase), the discrimination processing unit 21 inputs the discrimination data acquired by the discrimination data acquisition unit 7 into a learning model that reflects the learned learning parameters, and outputs the output values ​​from the learning model (e.g., features, etc.) to the discrimination result processing unit 6.

[0038] During the learning phase, the learning data processing unit 5 reads the learning data from the learning data storage unit 4 and the learning function storage unit 8, performs predetermined preprocessing, and then sends the learning data to the learning unit 20 and the discrimination processing unit 21. At that time, the learning data processing unit 5 sends the learning data to the learning unit 20 and the discrimination processing unit 21 in response to a request from the discrimination result processing unit 6.

[0039] The discrimination result processing unit 6 receives the output value from the discrimination processing unit 21 and outputs it as discrimination result AA to a predetermined output device, such as a display. In the learning phase, the discrimination result processing unit 6 calculates the coefficient of variation, discrimination rate, etc., based on the discrimination result and requests the learning data processing unit 5 to send further learning data to the learning unit 20 and the discrimination processing unit 21 according to the calculation result.

[0040] In the discrimination phase (inference phase), the discrimination data acquisition unit 7 receives discrimination data BB from a predetermined input device, performs predetermined preprocessing, and then sends the discrimination data BB to the discrimination processing unit 21.

[0041] The arithmetic unit 1 having the above configuration is made up of a general-purpose or dedicated computer. The machine learning device 1A and the discrimination device 1B may be made up of separate computers. In that case, the machine learning device 1A only needs to include at least a learning data storage unit 4, a learning function storage unit 8, a learning unit 20, and a learning parameter storage unit 3. The discrimination device 1B only needs to include at least a discrimination data acquisition unit 7 and a discrimination processing unit 21.

[0042] Among the components of the arithmetic unit 1, the learning parameter storage unit 3, the learning data storage unit 4, and the learning function storage unit 8 may be composed of storage devices such as hard disk drives (HDDs) and solid-state drives (SSDs) (internal, external, network-connected, etc.), or they may be composed of USB memory, storage media playable by a storage media playback device (CDs, DVDs, BDs), etc. Furthermore, among the components of the arithmetic unit 1, the discriminator learning unit 2, the learning data processing unit 5, the discriminant result processing unit 6, and the discriminant data acquisition unit 7 may be composed of an arithmetic unit having, for example, one or more processors (CPUs, MPUs, GPUs, etc.).

[0043] (program) The arithmetic unit 1 may function as a discriminator learning unit 2, a learning data processing unit 5, a discriminant result processing unit 6, and a discriminant data acquisition unit 7 by executing programs stored in various memory devices and storage media, or programs obtained by downloading from an external source via a network.

[0044] (integrated circuit) The neural network structure 100D may be constructed using an integrated circuit. In this case, the integrated circuit comprises an input / output unit that constitutes an input layer and an output layer, a storage unit that stores learning parameters, and an arithmetic unit that performs calculations to output the output value from the output layer based on a plurality of input data input to the input layer and the learning parameters stored in the storage unit. The integrated circuit may be constructed using, for example, an FPGA (Field Programmable Gate Array), an ASIC (Application Specific Integrated Circuit), or other hardware.

[0045] Next, we will explain the loss function, which is adjusted based on the relationship between the target value YY defined by equation [Equation 3-1] and the additive operation output BYA defined by equation [Equation 3-2].

[0046] There are at least three methods for finding the minimum value of the loss function (loss amount).

[0047] The first method uses the difference (|YY-BYA|) between the target value YY defined by equation [Equation 3-1] and the additive operation output BYA defined by equation [Equation 3-2] as the loss function, and directly calculates the loss amount by varying the exponent pn (a feature), the weighting parameter wn, the coefficients q of multiple learning functions, and the bias parameter b, and searches for the minimum value.

[0048] The second method uses the coefficient of variation CV (standard deviation / mean), defined by the ratio of the mean to the standard deviation of the target value YY as defined by equation [Equation 3-1], as the loss function. By varying the exponent pn, which is a feature, and the coefficients q of multiple learning functions, the value of the coefficient of variation CV that minimizes the loss (loss amount) is searched for. This determines the exponent pn, which is a feature, one learning function De, and the coefficients q of the learning functions.

[0049] The third method involves determining the exponent pn, a learning function De, and the coefficient q of the learning function De, using the second method to determine the feature, the coefficient of variation CV (second coefficient of variation CV') of the output formula [Equation 3-3]ZZ as the loss function, and then determining the weighting parameter wn by moving the feature using gradient descent to find the value of the coefficient of variation CV (second coefficient of variation CV') that minimizes it (loss amount).

[0050] The second method does not require the adjustment of the multiple weighting parameters wn and the bias parameter b, which are features. Furthermore, the bias parameter b in the second and third methods is uniquely determined, which is advantageous in terms of program structure. In addition, since the coefficient of variation CV, which is the loss amount, is calculated and stored with each parameter update, it is possible to create, for example, a list of loss amounts for each exponent pn (which is a feature) or a graph related to the loss amount without recalculating. In the embodiment, the second and third methods are appropriately adopted as the loss function.

[0051] The output equation ZZ (=f(YY,BYA)) output from the output layer 120D of the arithmetic unit 1 is designed as shown in the following equation [Equation 3-3]. Furthermore, according to the second and third methods of the loss function described above, the feature quantity bias parameter B (B=base b The product W of the ), and the weighting parameter is uniquely determined and can be simplified to the following equation [Equation 3-4] by substituting ZZ'=ZZ / (B*W). In other words, in the second and third methods, first, the weighting parameter wn is determined such that the coefficient of variation of YY is minimized, and this is used to calculate the bias parameter B such that the difference between YY and BYA is minimized (i.e., YY=BYA).

[0052] [Math 3-3] ZZ=f(YY,BYA) ZZ=D0 (p0+w0) *D1 (p1+w1) *…*DN (pN+wN) *De*B*W [Math 3-4] ZZ'=D0 (p0+w0) *D1 (p1+w1) *…*DN (pN+wN) *De

[0053] The output formula [Equation 3-4] ZZ' searches for an integer value for the exponent pn, which is a feature, and a decimal value that falls within ±1 for the weighting parameter wn, thereby finding the exponent of the input data Dn within the real value range of pn + wn.

[0054] The problems inherent in conventional neural network loss function minimum search methods exemplified in Patent Documents 1 and 2 will be explained using Figure 3. Figure 3 is an illustrative diagram showing weighted features w on the horizontal axis and loss L on the vertical axis. As shown in Figure 3, the neural network loss function minimum search methods exemplified in Patent Documents 1 and 2 use gradient descent to move the weighting parameter w by a small amount (Δw) to minimize the loss (error) L. In this method, even if there is a point smaller than the predicted minimum loss (error) L (local minimum), the search ends there, resulting in the problem of vanishing gradients, where the global minimum cannot be found. There are two ways to address this problem: firstly, increase the step size Δw when moving the parameter w. Secondly, instead of using the entire number of samples in the input data Dn, use sampled data (mini-batches) to reduce the amount of data and reduce the potential for the search to end at the local minimum. As a side effect, the former method carries the risk of accuracy deteriorating and missing the minimum point if the step size Δw is made too large. The latter approach does not necessarily prevent local minimums, and repeated trials are necessary when creating mini-batch data. Neither solution is perfect, and a problem exists in that users need enough knowledge and experience not only of neural networks but also of the problem they are trying to solve to be able to combine and evaluate these two methods in order to avoid falling into local minimums.

[0055] The method for finding the minimum value of the loss function of the arithmetic unit 1 of the present invention overcomes the aforementioned problems by performing a two-stage process. In the first stage, the second method described above is used to search for integer values ​​of the exponent pn and narrow down the exponent that minimizes the loss (error) L. In the second stage, the third method described above is used to search for the minimum value of the loss (error) L by using the narrowed-down integer exponent pn and slightly adjusting the weighting parameter wn using the gradient descent method. By processing in this two-stage method, it is possible to prevent the process from ending at the local minimum.

[0056] (First embodiment) Next, we will describe the first embodiment as a specific example of a method for searching for and discovering relational expressions. In the following, we will describe a specific example focusing on the feature parts that correspond to the present invention, following the procedure of the basic principle of the neural network structure 100D related to the basic form described above.

[0057] In recent years, with the miniaturization of various electronic components, including highly integrated LSIs, printed circuit boards have also seen rapid advancements in technology, such as finer patterning, narrower component pitches, and reduced conductor gaps. As a result, issues of wiring breakage and insulation reliability have come to the forefront. There is no absolute method for evaluating reliability and estimating lifespan from test results, and engineers are forced to make judgments based on repeated testing and empirical observations. To address this challenge, a more advanced method of lifespan prediction is desired, and it is crucial to derive a highly reliable lifespan prediction formula for sets of electronic components mounted on a printed circuit board.

[0058] In the first embodiment, we will describe an example in which a lifetime prediction formula was derived from data obtained from a reliability evaluation experiment of a circuit board on which electronic components were mounted on printed wiring material.

[0059] Figures 4 and 5 show graphs where the average lifespan L (average time to failure) of samples judged to have failed was plotted on the vertical axis. These graphs were conducted using a standard FR-4 substrate as a printed circuit board and mounted electronic components, with multiple samples performing current shunt tests for three parameters: temperature, humidity, and voltage. Figure 4 shows the Arrhenius plot when the voltage is fixed at 75V and the humidity is 60% and 95%. Figure 5 shows the voltage dependence graph when the number of test points at voltages of 50V, 65V, and 75V is increased.

[0060] The input data based on Figures 4 and 5 above is divided into three categories: experimental temperatures of 50°C, 75°C, and 85°C (item D0), experimental humidity of 95% and 60% (item D1), and experimental voltage of 75V, 65V, and 50V (item D2). The lifespan L, which is judged to be a failure, is then given as the value of item D3 and shown in the table in Figure 6.

[0061] The computing device 1 extracts the four-dimensional exponent pn, the four-dimensional weighting parameter wn, and the learning function De and its coefficient q, which are the features of the output formula (relation) ZZ' based on [Equation 3-4]. More specifically, the computing device 1 searches and extracts the features pn, wn, and De and their coefficient q of the output formula (relation) ZZ' based on [Equation 3-4] that hold between the input item data in the phenomenon, from the 11 four-dimensional input data Dn=(D0,D1,D2,D3) in the table of Figure 6, according to the flow of the basic form described above.

[0062] In the four-dimensional input data Dn=(D0,D1,D2,D3) of the reliability test results mentioned above, the item names D0, D1, and D2 are parameter values ​​for temperature, humidity, and voltage, respectively, and D3 is the output value of the average lifetime L (average time to failure) judged as a one-dimensional failure. Since the output formula for item name D3 is the lifetime prediction formula for the average lifetime L that we want to derive, the output formula (relationship) ZZ' based on [Equation 3-4] can be expressed as the following formula [Equation 3-5] by fixing the feature p3 to, for example, -1 and w3 to 0, and the formula with the average lifetime L D3 as the solution can be expressed as [Equation 3-6].

[0063] [Math 3-5] ZZ'=D0 (p0+w0) *D1 (p1+w1) *D2 (p2+w2) *D3 (-1) *De [Math 3-6] D3=D0 (p0+w0) *D1 (p1+w1) *D2 (p2+w2) *De / ZZ'

[0064] As in the example above, when using the learning function De in the output expression ZZ', the arithmetic unit 1 includes a function that allows specifying one of the input data Dn to contain the desired output item, and fixing the feature quantities pn to -1 and 1, and wn to 0. This function contributes to reducing the computer's feature search time.

[0065] Here, we will explain examples of functions that can be used as the learning function De=g(q,Dn). The learning function is the function used in the learning function node 141, and a typical example is the Maxwell-Bolzman distribution, a statistical distribution function used in classical mechanics. This type can be expressed as De=exp(q / (D0)). The feature quantity q, which is an element in exp(), is obtained by scanning the width of values ​​within a set predetermined range with a predetermined step size. D0 is composed of, for example, D0 from the input data Dn=(D0,D1,…,DN). Typical examples in quantum physics include the Fermi-Dirac distribution function De=1 / (1+exp(q / (D0))), the distribution function of Bose particles De=1 / (exp(q / (D0))-1), and the standard normal distribution function De=exp(-q*D0 2 There is a type of exp() function. In many cases, the elements within exp() of such a typical type use one item from the input data Dn, but this is not always the case. Therefore, multiple combinations of products and ratios of multiple items of the input data Dn are registered in advance as types of learning functions, and the most appropriate learning function De is searched for and extracted. An example of a learning function using two-dimensional input items D0 and D1 of the input data Dn is shown in Figure 7. The number of dimensions can be the combination of the maximum number of input items, or the number of items can be narrowed down based on the characteristics of the input data items. In addition to statistical distribution functions, dipolar trigonometric functions that can express wave oscillations, or functions that represent the convergence process to Napier's number (e) can also be used as learning functions. The latter type of function that represents the convergence process to Napier's number (e) is the same type as the compound interest calculation formula De=(1+q / D0)^D1 when q is the interest rate (e.g., annual interest rate), D0 is the number of divisions of the interest rate, and D1 is the number of investments, and it is possible to grasp the characteristics of the formula inherent in economic activity from the input data. These are just a few examples of functions that can be used as training functions De; many other types of functions can also be used as training functions. Hereafter, these functions will be collectively referred to as Eureka functions.

[0066] The first stage involves inputting 4-dimensional input data Dn=(D0,D1,D2,D3) into the input layer 110D of the neural network structure 100D, and searching for the most appropriate learning function De with integer exponents.

[0067] The 4-dimensional input data Dn=(D0,D1,D2,D3) is passed to the transformation layer 140, and the power values ​​Dn pn It is converted and passed to the hidden layer input node 133 of the hidden layer 130. The initial value of the 4-dimensional integer exponent pn, which is the feature, is set to have a range of -9 to 9, and it is scanned one step at a time from the starting value of (-9,-9,-9,-9) to (0,0,0,0). The initial value of the integer exponent pn can be set arbitrarily by the user, and may be shortened to -6 to 6, etc. As mentioned above, in this example, by specifying the item name D3 in the output formula of the mean lifetime L, the feature p3 is fixed to -1 and 1. Also, (0,0,0,0) is a singularity and is excluded.

[0068] The De node of the learning function node 141 in the transformation layer 140 has multiple registered learning functions set in parallel and is passed to the hidden layer input node 133 of the hidden layer 130.

[0069] The respective power values ​​Dn set at the hidden layer input node 133 of the hidden layer 130. pn And the multiple learning functions De are processed by the calculation of equation [3-1] YY at the first hidden layer node 131. As described above, in the first stage, all the weighting parameters wn, which are the features of equation [3-1] YY, are fixed to 0, and the product W of the weighting parameters is fixed to 1, and the product of the power values ​​of the multiple learning functions De is expressed as YY = D0 p0 *D1 p1 *D2 p2 *D3 p3 *De is calculated.

[0070] The arithmetic unit 1 scans the q value within a predetermined range and increment to determine a single learning function De and its coefficient q that minimize the coefficient of variation CV (loss) from the calculated value of equation [Equation 3-1] YY. Once a single learning function De and its q value that minimize the coefficient of variation CV (loss) are determined, the learning parameter storage unit 3 of the machine learning device 1A stores the integer exponent pn, the sorting number of the learning function De, the feature quantities of the coefficient q of the learning function, and the coefficient of variation CV (loss).

[0071] As described above, the integer exponent pn is repeatedly set from the initial value (-9,-9,-9,-9) down to (0,0,0,0), and the learning parameter storage unit 3 of the machine learning device 1A stores the integer exponent pn, the organization number of the learning function De, the feature quantities of the coefficient q of the learning function, and the coefficient of variation CV (loss amount). The first stage is completed by searching for the most appropriate learning function De for the integer exponent. Next, the discriminant processing unit 21 receives the values ​​of the integer exponent pn, the organization number of the learning function De, the feature quantities of the coefficient q of the learning function, and the coefficient of variation CV (loss amount) from the learning parameter storage unit 3 of the machine learning device 1A, creates a list sorted by the smallest coefficient of variation CV (loss amount), and outputs it to a display or the like.

[0072] Figure 8 shows a list of the coefficients of variation (loss) output in ascending order. However, for the sake of explanation, the list shown is for when p3 = -1. As can be seen from Figure 8, the list in Figure 8 is output in order of the coefficient of variation value, with the power exponents being ranked in descending order of importance.

[0073] From the list of coefficients of variation CV (loss) in ascending order in Figure 8, the smallest feature combination is pn=(0,-3,-1,-1), the learning function is exp(q / D0), and the coefficient q=10116. From this, the output equation (relationship) [Equation 3-5] ZZ' can be expressed in the following equation [Equation 3-7], and the equation with D3 of the average lifetime L as the solution can be expressed in [Equation 3-8].

[0074] [Math 3-7] ZZ'=D0 (0+w0) *D1 (-3+w1) *D2 (-1+w2) *D3 (-1) *exp(q / D0) [Math 3-8] D3=D0 (0+w0) *D1 (-3+w1) *D2 (-1+w2) *exp(q / D0) / ZZ'

[0075] In the second stage, the four-dimensional input data consisting of the integer exponent pn with the smallest coefficient of variation CV (loss), the sorting number of the learning function De, and the features of the coefficient q of the learning function is output from the discrimination processing unit 21 to the learning data processing unit 5 via the discrimination result processing unit 6, and the features of the weighting parameter wn are searched.

[0076] The process of receiving the features obtained in the first stage and proceeding to the second stage can be a continuous process of passing the feature with the smallest coefficient of variation CV (loss amount) to the second stage, or it can be a process of stopping, specifying a feature from the list of coefficients of variation CV (loss amount) data, and manually inputting it into the second stage of processing.

[0077] The learning data processing unit 5 receives the learning function De corresponding to the sorting number of the learning function De from the learning function storage unit 8, and outputs the integer exponent pn, the learning function corresponding to the sorting number of the learning function De, and the coefficient q of the learning function to the learning unit 20.

[0078] The data transfer flow from the learning data processing unit 5 to the learning unit 20 is the flow from the input layer 110D to the conversion layer 140 of the neural network structure 100D, where the input data is converted and set in the hidden layer input node 133.

[0079] The data from the hidden layer input node 133 is passed to the first hidden layer node 131 and the second hidden layer node 132, respectively. The coefficient of variation CV (loss amount) is calculated using the third method, which is the loss function calculation method described above, and the feature quantities of the weighting parameter wn are determined.

[0080] The weighting parameter wn, which is a characteristic when the coefficient of variation CV (loss) is at its minimum value, is extracted as (0,0,-0.4). ZZ' in the output equation (relationship) [Equation 3-7] is expressed in [Equation 3-9] and approximates the average value of ZZ'. If we define this average value of ZZ' as ZZ'(average), the equation (second equation) with D3 of the average lifespan L as the solution is expressed in [Equation 3-10].

[0081] [Math 3-9] ZZ'=D1 (-3) *D2 (-1。4) *D3 (-1) *exp(q / D0) [Math 3-10] D3=D1 (-3) *D2 (-1.4) *exp(q / D0) / ZZ'(average)

[0082] Figure 9 shows the values ​​of ZZ' calculated using the output formula (relationship) [Equation 3-9] corresponding to the 11 samples in Figure 6. From this table, it can be seen that the mean of the output formula (relationship) ZZ' is 1.76E+09 and the standard deviation is 7.83+E07, indicating that a constant approximation with little variability has been achieved. If we replace the mean of the aforementioned ZZ' (mean) with A = 1 / ZZ' (mean), equation [Equation 3-10] can be expressed as equation [Equation 3-11]. Note that the values ​​in the ZZ' column in Figure 9 are denoted in E. [Math 3-11] D3 = A * D1 (-3) *D2 (-1。4) *exp(q / D0) q = 10116 constant A = 5.68E-10 constant

[0083] The equation for which D3 of the average lifespan L is the solution is automatically calculated by a computer algorithm following the procedure described above. Furthermore, a correlation function graph of D3 of the average lifespan L is displayed, and in Figure 10, the measured value of D3 of the average lifespan L is shown on the horizontal axis, and the calculated (predicted) lifespan value using equation [Equation 3-10] is shown on the vertical axis. The correlation coefficient is 0.999, indicating a very strong correlation.

[0084] Thus, the calculation device 1 of the present invention can derive a highly accurate lifetime prediction formula from a small amount of input data. Furthermore, the calculation device 1 of the present invention is also effective in analyzing the obtained lifetime prediction formula. Equation 3-11 can be expressed as equation [Equation 3-12] (the second equation) by replacing the item name Dn with the measurement name.

[0085] [Math 3-12] L=A*H (-3) *V (-1、4) *exp(Eb / kT) A = 5.68E-10 constant n=1.4 (Inverse) n-th power voltage acceleration coefficient k = 8.6157E-5eV / K Boltzmann constant Eb = 0.872 eV Activation energy

[0086] Equation [Equation 3-12] identifies the Arrhenius model equation, an empirical rule known in chemical reaction theory that describes the relationship between electronic component failures and the inverse n-power law of type exp(Eb / kT) with respect to voltage V, as the most suitable equation for measurement data. It is easy for people to understand and is useful for advancing the analysis and interpretation of the model.

[0087] In proceeding with the analysis of the model, the list of coefficients of variation CV (loss) in ascending order shown in Figure 8 is useful. Referring to Figure 8, it can be seen that the losses for pn = (0 to -9, -3, -1, -1) occupy the top 10 positions and are concentrated in the range of -1.111 to -1.097. In other words, the loss amount does not differ significantly regardless of the value of the power exponent p0 of temperature T from 0 to -9. To an engineer, when no significant difference is found among multiple equations that express the measurement results, the Eyring model equation, a known theoretical formula, is the second highest in the list in Figure 8, and this can also be considered as a criterion for deciding which equation to adopt. From the list in Figure 8, the pn in the Eyring model equation is represented as (-1,-3,-1,-1), and using the feature with a learning function coefficient q of 9772, starting from the second stage described above, searching for the feature of the weighting parameter wn yields a value equivalent to (0,0,-0.4), and the lifetime prediction equation for the Eyring model can be expressed as the following equation [Equation 3-13]. The correlation coefficient is calculated to be 0.999, which is comparable to the correlation coefficient in [Equation 3-12], indicating that both equations yield highly accurate model equations.

[0088] [Math 3-13] L=A*T (-1) *H (-3) *V (-1、4) *exp(Eb / kT) A=5.32E-7 constant n=1.4 (Inverse) n-th power voltage acceleration coefficient k = 8.6157E-5eV / K Boltzmann constant Eb = 0.842 eV Activation energy of the Eyring model

[0089] As described above, by utilizing the present invention, the problem that engineers are forced to make judgments based on repeated testing and empirical observations when deriving a lifetime prediction formula is solved, and a highly accurate lifetime prediction formula (second equation) that can be logically explained can be automatically fitted and derived with high accuracy through consistent computer processing. Furthermore, by equipping the computer with a memory unit (equation memory unit) that stores the lifetime prediction formula (second equation) derived in this way, and a calculation unit that calculates the output for input values ​​using this second equation, an equation utilization device for predicting the lifetime of electronic components can be configured.

[0090] (Second embodiment) Next, as a second specific example of a method for searching for and discovering relational equations, a second embodiment will be described. Dr. Planck, a German physicist, is known for finding a very good experimental equation for a spectral distribution by interpolating (fitting) the entire wavelength λ range based on temperature, wavelength, and intensity data of the blackbody radiation spectral distribution.

[0091] In the second embodiment, following the first embodiment, a method for easily and automatically deriving an empirical formula for the spectral distribution from a small amount of data by extracting data from a graph of the blackbody radiation spectral distribution found in the literature and using the present invention will be explained.

[0092] Figure 11 is the third figure in "The Origins of Thermal Radiation Theory and Quantum Theory" by Kiyoshi Amano, and is a graph with wavelength λ on the horizontal axis, intensity E on the vertical axis, and temperature T as parameters, with observed values ​​displayed as dots. Figure 12 is a table extracted from the data in Figure 11, with the experimental wavelength λ divided into item name D0, the experimental temperature T into item name D1, and the experimental intensity E into item name D2.

[0093] From the 21-sample 3D input data Dn=(D0,D1,D2) in the table of Figure 12, the output equation (relationship) that holds between the input item data in the phenomenon is extracted by extracting the 3D power exponent pn, the 3D weighting parameter wn, and the learning function De and its coefficient q, which are the features of the aforementioned equation [Equation 3-4]ZZ', and finally derive an equation that expresses the experimental intensity E (equation for item name D2). The procedure of this method is the same as in the first embodiment, and only the feature part will be explained.

[0094] In the aforementioned 3D input data Dn=(D0,D1,D2), the item names D0 and D1 are the parameter values ​​for wavelength and temperature, respectively, and D2 is the output value of the one-dimensional intensity E. Since the output formula for item name D2 is the empirical formula for the spectral distribution we want to derive, the output formula (relation) ZZ' based on [Equation 3-4] can be expressed as the following formula [Equation 3-15] when the feature quantities p2 are fixed to -1 and 1, and w2 to 0, and the formula with intensity E as the solution can be expressed as [Equation 3-16]. However, for the sake of explanation, the case where p2=-1 is shown.

[0095] [Math 3-15] ZZ'=D0 (p0+w0) *D1 (p1+w1) *D2 (-1) *De [Math 3-16] D2=D0 (p0+w0) *D1 (p1+w1) *De / ZZ'

[0096] The first step involves inputting the 3D input data Dn=(D0,D1,D2) into the input layer 110D of the neural network structure 100D, and searching for the most appropriate learning function De with integer exponents. The initial value of the 3D integer exponents pn, which are the features, is set to have a range of -9 to 9, starting from (-9,-9,-9) and scanning one step at a time down to (0,0,0).

[0097] Upon completion of the first stage, the minimum value of the coefficient of variation (loss), which is the integer exponent pn, and the most suitable learning function De and its coefficient q are determined. A list of the coefficients of variation in ascending order is then created and output to a display or similar device.

[0098] Figure 13 shows a list of the coefficients of variation (CV) for this case, ordered from smallest to largest. As can be seen from Figure 13, the list in Figure 13 is: statistical distribution The list displays the order of importance for each function type, sorted by the value of the coefficient of variation.

[0099] From the list of coefficients of variation CV in ascending order in Figure 13, the smallest feature combination is pn=(-5,0,-1), the learning function is 1 / (exp(q / D0 / D1)-1), and the coefficient q=13964 is extracted. From this, the output equation (relationship) [Equation 3-15]ZZ' can be expressed as the following equation [Equation 3-17], and the equation with intensity E and D2 as the solution can be expressed as [Equation 3-18].

[0100] [Math 3-17] ZZ'=D0 (-5) *D2 (-1) / (exp(q / D0 / D1)-1) [Math 3-18] D2=D0 (-5) / (exp(q / D0 / D1)-1) / ZZ'

[0101] In the second stage, the procedure involves searching for the features of the weighting parameter wn from the features of the integer exponent pn=(-5,0,-1), the learning function, 1 / (exp(q / D0 / D1)-1), and the coefficient q=14125 of the learning function.

[0102] As a result of the above procedure, (0,0) is extracted as the weighting parameter wn, which is the characteristic when the coefficient of variation CV (loss) is at its minimum value. In other words, the optimal value for both the decimal weighting features w0 and w1 is 0, and it is found that ZZ' in equation [Equation 3-17] is composed of a simple integer power exponent, which is shown in [Equation 3-19], and the average value of ZZ', ZZ' (mean), is output as 3.4E-6. For convenience, if we replace c1=1 / ZZ' and q=c2, and replace D0 with the item name wavelength λ, D1 with temperature T, and D2 with intensity E, the equation with intensity E as the solution is shown in [Equation 3-20].

[0103] [Math 3-19] ZZ'=D0 (-5) *D2 (-1) / (exp(q / D0 / D1)-1) q=1.4 E4 ZZ' (mean value) = 3.4E-6 constant [Math 3-20] E=c1*λ (-5) / (exp(c² / (λ*T))-1) c1 = 2.8E5 constant (1 / ZZ') c² = 1.4 E4 Constant (q)

[0104] The equation [Equation 3-20] above is the same as the experimental formula for the spectral distribution of blackbody radiation discovered by Dr. Planck.

[0105] For confirmation, following Figure 11, the horizontal axis is plotted on wavelength λ and the vertical axis on intensity E. Figure 14 shows a graph with the measurement data from Figure 12 displayed as dots and the values ​​from equation [Equation 3-20] as solid lines. The correlation coefficient is 0.997, indicating good agreement between the experimental values ​​and the empirical formula.

[0106] As shown in the first and second embodiments described above, the present invention performs a two-stage exponential search, in which the integer exponent pn is searched for by a one-shift blind scan in the first step from the measurement data, and the decimal exponent wn is searched for using gradient descent in the second step, thereby preventing the amount of loss from remaining at the local minimum. Moreover, in nature, exponents are often simple integer values, making this an effective method for equation discovery. Furthermore, searching only for integer exponents for a number of dimensions equal to the input item has little impact on the search time, but a gradient descent search based on integer values ​​may also be used from the first step.

[0107] Furthermore, this invention has an advantage not found in other neural networks: the ability to use the coefficient of variation CV as the loss function. As shown in the example above, a ranked list of coefficients of variation CV (loss amount) is created, making it possible to visualize and analyze the relationship between the exponent and the coefficient of variation CV (loss amount). This advantage also makes it worthwhile to select a blind scan search with a one-shift integer exponent pn in the first step.

[0108] The name of the exponential neural network incorporating the Eureka function (learning function) described above will be standardized and referred to as the exponential neural network below.

[0109] (Third embodiment)

[0110] Next, we will explain a method for deriving a function that obtains a high discrimination rate by using the exponential neural network, which is the basis of this invention, for object discrimination.

[0111] Below, we consider a method for classifying objects, for example, three different types A, B, and C, by selecting four data items (4-dimensional data) D0, D1, D2, and D3 to distinguish between them. In this case, two methods can be considered for distinguishing and classifying the different A, B, and C.

[0112] First, let's explain the first classification method. Initially, using training data from groups A, B, and C of objects to be classified, we search for the function with the smallest error in the relationships between data within each group. Next, we set thresholds for each of the three functions discovered. For example, applying the formula for classification B to 4D data belonging to classification A yields outliers based on the threshold. Using these functions, it is possible to accurately classify an unknown object from 4D data obtained from that object.

[0113] The method described above utilizes the search procedure for relational equations of an exponential neural network to determine highly accurate functions representing group A, group B, and group C, set thresholds for each equation, and use them as discriminant equations.

[0114] The second classification method involves searching for an equation that maximizes the discriminant rate for groups A, B, and C of the objects to be classified, and then using that equation as the discriminant formula.

[0115] The second classification method will be explained using an example of classifying the flowers of the iris (Iris) as the object to be classified. Below, we will explain a method in which multiple discriminant formulas are selected from among discriminant formulas with different characteristics, based on the discriminative performance of the said discriminant formula and the value of the coefficient of variation, and these discriminant formulas are used to classify the flowers of the iris (Iris) as the object to be classified.

[0116] In object discrimination, the discrimination rate can be treated in the same way as the reward in reinforcement learning. Similar to how the ascending rank list with the coefficient of variation CV as the loss amount was used in the function search described above, the reward function (discriminant formula) that yields the maximum reward (discrimination rate) is determined by using a descending rank list of rewards (discrimination rates) obtained by searching for integer exponents pn.

[0117] The third embodiment will be described in detail below using an example of distinguishing between three types of irises (setosa, versicolor, and virginica). Figures 15 and 16 are tables showing examples of data for the three types of irises (hereinafter referred to as "Iris data"). In the examples in Figures 15 and 16, 50 four-dimensional data sets each were prepared as Iris data, with the calyx side length and width and the petal length and width associated. These were then assigned serial numbers to the following groups: Group A No. 1-50 (setosa), Group B No. 51-100 (versicolor), and Group C No. 101-150 (virginica).

[0118] The basic set of Iris data mentioned above is frequently cited as a foundation for machine learning and is described in Non-Patent Document 1.

[0119] Figure 17 shows the correlation graphs between calyx length-calyx width and calyx length-petal length for the Iris data in Figures 15 and 16. From the data comparison between species, setosa can be easily distinguished from the other two species because the data distances are far apart, but versicolor and virginica have close data distances, making differentiation difficult. In this case, we will explain a specific example that maximizes the discrimination rate between versicolor, virginica, and setosa.

[0120] Using the training data of versicolor, virginica, and setosa input to the input layer 110D of the neural network structure 100D in Figure 1, the values ​​of the equation YY, which is the product of the power values ​​of the weighted feature wn=0 and the training function De=1, are calculated at the first hidden layer node 131 using the aforementioned equation [Equation 3-1], which consists of the exponent pn.

[0121] The exponent pn is scanned from (-9,-9,-9,-9) to (0,0,0,0), and the value of the product of the exponents of each exponent pn, YY, is compared with the values ​​of the versicolor group, the virginica group, and the setosa group. The upper and lower thresholds for the value of the product of the exponents YY that yields the maximum reward (discrimination rate) are calculated, and the reward (discrimination rate) corresponding to each exponent pn is stored and listed.

[0122] In this case, it is preferable to exclude from the calculation exponents pn if none of the absolute values ​​of the elements of the exponent pn are 1 or prime absolute values, and exponents pn if all of the absolute values ​​of the prime absolute values ​​are the same. For example, when the exponent pn = (-2, -1, 3, 1), values ​​that are 2 times, 3 times, etc., such as (-4, -2, 6, 2) and (-6, -3, 9, 3), and values ​​corresponding to the exponent pn = (-1, -1, -1, -1), such as (-2, -2, -2, -2) and (-3, -3, -3, -3), are excluded because their mathematical rewards (discrimination rates) are equivalent. Doing so has the advantage of reducing computation time and simplifying the list.

[0123] Furthermore, in this case, it is clear that we are searching for a discriminant that uses all four elements of the training data: calyx length, calyx width, calyx length, and petal length. Therefore, we are excluding exponents pn that contain at least one element of 0.

[0124] Figure 18 shows the top 30 in a descending list of the rewards (discrimination rates) calculated above. From Figure 18, the highest reward (discrimination rate) is 99.3%, and there are four such values. If the rewards (discrimination rates) are tied, the groups are listed in descending order of their coefficient of variation and ranked accordingly. This allows us to rank the groups based on the strength of their exponents, with the exponents pn representing the least variation in individual values ​​within the formula YY, i.e., the most robust exponents.

[0125] In Figure 18, the highest accuracy rate (99.3%, 149 out of 150) was achieved when pn = (-2, -1, 4, 1), with only one misclassification (No. 84). The formula (discriminant formula AA), threshold, and judgment value are shown below in [Equation 3-21].

[0126] [Math 3-21] YY=D2 4 *D3 / (D0 2 *D1) Discriminant AA Upper threshold 8.8 Lower threshold 0.2 Judgment values: A < 0.2, 0.2 ≤ B ≤ 8.8, C > 8.8

[0127] Figure 19 shows a graph plotting the flower number of iris on the horizontal axis and the discriminant AA value on the vertical axis. It can be seen that misclassification occurs when versicolor No. 84 is misclassified as virginica.

[0128] Using the discriminant formula AA obtained in the above procedure, unknown input data for classifying flowers can be classified.

[0129] Next, we will explain how to further improve the classification rate. In this example, a classification rate of 99.3% (149 out of 150) has already been achieved, and the method aims to proceed with the search with the goal of achieving 100% classification.

[0130] First, I will explain the classification concept that underlies this invention. Classification involves setting some criteria, creating multiple categories based on those criteria, and placing individual items into one of these categories, and therefore involves arbitrariness. The classification of the Iris flowers in this example is not absolute. Since the classification criteria are set by humans, the classification formula is not represented by a single equation based on natural laws as shown in the first and second embodiments described above. Therefore, we will incorporate a simulation of majority-vote decision-making by multiple people, which is one of the behavioral patterns that humans should adopt, into the classification formula. Below, I will explain how to incorporate majority-vote decision-making using multiple discriminant formulas into the classification formula using a power-law neural network.

[0131] For convenience, the aforementioned discriminant [Equation 3-21] will be denoted as discriminant AA. Furthermore, if we define unknown discriminants AB and AC, the majority vote discrimination can be expressed using the MODE function as follows:

[0132] [Math 3-22] Z=MODE(AA,AB,AC)

[0133] The MODE function of the aforementioned multiple discriminant [Equation 3-22] outputs versicolor (symbol B) when, for example, the discriminant AA is virginica (symbol C), the discriminant AB is versicolor (symbol B), and the discriminant AC is versicolor (symbol B). Alternatively, if the discriminant results of AA, AB, and AC are all different, or if there are multiple maximum numbers, the symbol of the first maximum number is output. Note that the number of discriminant expressions can be further increased, such as Z = MODE(AA, AB, AC, AD, AE).

[0134] Next, we will explain with specific examples how to search for the aforementioned discriminant formulas AB and AC.

[0135] One of the misclassifications of discriminant AA is that it mistakes flower No. 84 from versicolor (symbol B) to virginica (symbol C). Therefore, the condition for improving the discrimination rate is that discriminant AB correctly identifies flower No. 84, although it may make mistakes with other flower Nos, but it should select those with a high discrimination rate. Next, the condition for discriminant AC is that it correctly identifies the flower Nos that were misclassified by discriminant AA and discriminant AB, although it may make mistakes with other flower Nos, but it should select those with a high discrimination rate.

[0136] Based on the descending list of discrimination rates shown in Figure 18, the results of searching for discriminant formulas AB and AC that meet the above conditions are shown in the list in Figure 20.

[0137] The discrimination rate of the first item in the list of Figure 20 with pn = (-1, 3, -2, 5) is 98.7% (148 / 150), and the misclassifications are 2, No. 134 and 135, where C is misclassified as B. Figure 21 shows a graph with the iris flower No. on the horizontal axis and the value of discriminant AB on the vertical axis. The discriminant AB is shown in [Equation 3-23] below.

[0138] [Equation 3-23] YY / W = D1 3 / (D0 * D2 2 * D3 5 ) Discriminant AB Upper threshold value 21.0 Lower threshold value 1.11E-2 Judgment value A > 21.0, 1.11E-2 ≤ B ≤ 21.0, C < 1.11E-2

[0139] The discrimination rate of the second item in the list of Figure 20 with pn = (-1, 2, 7, -1) is 90.7% (136 / 150), and the misclassifications are 14, No. 102, 107, 114, 115, 120, 122, 124, 127, 128, 139, 142, 143, 146, 147, where C is misclassified as B. Figure 22 shows a graph with the iris flower No. on the horizontal axis and the value of discriminant AC on the vertical axis. The discriminant AC is shown in [Equation 3-24] below. Although the discrimination performance of discriminant AC is not good, it is a formula with a sharp view that can give the correct answer where discriminant AA and discriminant AB make mistakes.

[0140] [Equation 3-24] YY / W = D1 2 * D2 7 / (D0 * D3) Discriminant AC Upper threshold value 69.3E3 Lower threshold value 1.8E3 Judgment value A < 1.8E3, 1.8E3 ≤ B ≤ 69.3E3, C > 69.3E3

[0141] As mentioned above, Figure 23 shows the results of classifying the training data for identifying Iris flowers using the majority-vote discriminant Z=MODE(AA,AB,AC), which utilizes three discriminant formulas: discriminant formula AA[Equation 3-21], discriminant formula AB[Equation 3-23], and discriminant formula AC[Equation 3-24]. The classification rate achieved is 100%.

[0142] The combination of the three formulas has the following characteristics: Each formula is a combination of products of power values ​​that utilize different quadrants separated by an exponent of 0, and the order of the elements of the exponent pn in each formula includes elements with different signs (±).

[0143] Furthermore, as a means of discrimination that does not use majority voting discrimination, a method can also be adopted that includes means for converting the output value of the formula represented in the list in Figure 18 according to Patent Document 3 into points (scores) corresponding to the variance or standard deviation of the output value, and means for summing the points, and then weighting the points (scores) of each discriminator output so that the judgment result of the object to be discriminated using the summed points (scores) obtains the highest discrimination rate. In comparison with these methods, the majority voting discrimination method described in this example has a simple classification formula, and the method according to Patent Document 3 has the advantage of not using a threshold.

[0144] In other words, the present invention naturally explores, discovers, and combines multiple optimal formulas that offer different perspectives on quadrants of the power exponent, thereby presenting a highly accurate classification formula that is easy for humans to understand.

[0145] As demonstrated in the procedure described above, the machine learning process, data, and model creation methods are transparent. Conventional AI (Artificial Intelligence) and machine learning solutions are often complex and opaque, raising concerns that those with unclear processes and creation methods cannot be reliably used in practical applications. This invention addresses such concerns by constructing an exponential neural network capable of deriving equations, and by preparing and combining equations of product of power values ​​and learning functions as components of the model equation, thereby solving problems that humans cannot interpret and realizing transparent AI. [Explanation of symbols]

[0146] 1...Calculation unit, 1A...Machine learning unit, 1B...Discrimination unit, 2...Discriminator learning unit, 3...Learning parameter storage unit, 4...Learning data storage unit, 5...Learning data processing unit, 6...Discrimination result processing unit, 7...Discrimination data acquisition unit, 8...Learning function memory unit, 20...Learning unit, 21...Discrimination processing unit, 100D... Neural network structure, 110D…Input layer, 120D…Output layer 130...Hidden layer, 131...First hidden node, 132...Second hidden node 133...Hidden layer input node, 140...Transformation layer, 141...Tracking function node

Claims

1. A computing device that uses a neural network structure including at least an input layer and an output layer to output an output value from the output layer for a plurality of input data (D0, D1, ..., DN) input to the input layer, The aforementioned input layer is The neural network structure has, as learning parameters, a plurality of exponents (p0, p1, ..., pN) that are associated with each of the plurality of input data and each of the plurality of input data raised to a power, and the coefficients of a predetermined learning function. The output layer is, Multiple power values ​​(D0) are obtained by raising each of the multiple input data inputs to the input layer to the power of a plurality of exponents. p0 D1 p1 ,…,DN pN ) product (YY0 = D0 p0 *D1 p1 *...* DN pN Based on the product of ) and the learning function (De = g(D0, D1, ..., DN)), the output value (y = f(YY0 * De)) is output. Computing device.

2. The multiple exponents used as learning parameters and the coefficient parameters of the learning function are, A parameter that is learned by using multiple sets of the aforementioned input data as training data, The coefficient of variation of the output value (y) output when the multiple input data as training data are input to the input layer is adjusted to be small. The computing device according to claim 1.

3. Based on the value of the coefficient of variation of the output value (y) described in claim 2, the power exponents are displayed in a list in order of their coefficient of variation, in descending order of priority. The arithmetic device according to claim 2.

4. Based on the coefficient of variation of the output value (y), the superiority / inferiority order of each type of learning function is displayed as a list in order of the coefficient of variation. The arithmetic device according to claim 2.

5. The aforementioned neural network structure is, A hidden layer is further included between the input layer and the output layer. The aforementioned hidden layer is Multiple input data are input via multiple weighting parameters (w0, w1, ..., wN) as learning parameters, and a first hidden node outputs a target value (YY) defined by the following equation [Equation 1] to the output layer, The system includes a second hidden node to which multiple input data are each input via the multiple weighting parameters, and a bias parameter (b) as a learning parameter is input, and which outputs an additive operation output (BYA) defined by the following equation [Equation 2] to the output layer, The output layer is, Based on the target value (YY) and the additive calculation output (BYA), the output value (y = f(YY,BYA)) is output. The computing device according to claim 1. [Mathematics 1] YY=D0 p0 *D1 p1 *…*DN pN *W0*W1*…*WN*De (=YY0*W0*W1*...*WN*De) [Math 2] BY=B*(base) (Σ[n=0→N](wn*pn*dn)) however, base is a positive number other than 1. Dn=base dn (n=0,1,…,N) Wn=base wn (n=0,1,…,N) B=base b That is the case.

6. The coefficient of variation is a value defined by the ratio of the mean value to the standard deviation of the target value YY, as defined in [Equation 3] below. After determining the weighting parameter wn such that the coefficient of variation is minimized, the bias parameter B (B = base) is determined using the weighting parameter wn such that the difference between YY defined in [Equation 3] below and BYA defined in [Equation 4] below is minimized. b ) calculate The arithmetic device according to claim 2. [Mathematics 3] YY=D0 p0 *D0 p1 *…*DN pN *W0*W1*…*WN*De [Math 4] BY=B*(base) (Σ[n=0→N](wn*pn*dn)) however, base is a positive number other than 1. Dn=base dn (n=0,1,…,N) Wn=base wn (n=0,1,…,N) B=base b

7. After determining the learning parameters, the second coefficient of variation CV' of ZZ, defined by the output formula [Equation 5] below, is used as the loss function. The weighting parameter wn, which is a feature, is moved using gradient descent to search for the value of the second coefficient of variation CV' that minimizes it (loss amount), and the weighting parameter wn is determined. The computing device according to claim 6. [Math 5] ZZ=f(YY,BYA) ZZ=D0 (p0+w0) *D1 (p1+w1) *…*DN (pN+wN) *De*B*W

8. The multiple power exponents, multiple weighting parameters, and bias parameters used as learning parameters are, A parameter that is learned by using multiple sets of the aforementioned input data as training data, When the multiple input data as training data are input to the input layer, the difference (|YY - BYA|) between the target value (YY) output from the first hidden node and the additive operation output (BYA) output from the second hidden node is adjusted to be small. The arithmetic device according to claim 5.

9. The multiple exponents used as learning parameters and the coefficients of the learning function are, A parameter learned by using multiple sets of training data, each set including multiple input data and training data associated with the multiple input data, The output value output from the output layer when multiple input data included in the training data are input to the input layer is adjusted so that the difference between the output value and the training data included in the training data becomes small. The computing device according to claim 1.

10. The multiple exponents used as learning parameters and the coefficients of the learning function are, A parameter learned by using multiple sets of training data, each set including multiple input data and training data associated with the multiple input data, The output value output from the output layer when multiple input data included in the training data are input to the input layer is adjusted so that the difference between the output value and the training data included in the training data becomes small. The arithmetic device according to claim 5.

11. The aforementioned multiple power values ​​(D0 p0 D1 p1 ,…,DN pN ) product (YY0 = D0 p0 *D1 p1 *...* DN pN From the first equation using the learning function (De = g(D0, D1, ..., DN)), A second equation is created, representing the portion corresponding to the aforementioned input data using variables and a constant term, and presented to the user. The arithmetic device according to any one of claims 1 to 10.

12. An integrated circuit comprising the neural network structure used by the computing device according to any one of claims 1 to 10, The input / output section comprising the input layer and the output layer, A storage unit for storing the learning parameters, The system includes a calculation unit that performs calculations to output an output value from the output layer based on a plurality of input data input to the input layer and the learning parameters stored in the storage unit, Integrated circuit.

13. A machine learning device for generating a learning model having the neural network structure used by the computing device according to any one of claims 1 to 10, A learning data storage unit that stores learning data including at least multiple input data, A learning unit that learns the learning parameters by inputting the learning data stored in the learning data storage unit into the learning model, The system includes a learning parameter storage unit that stores the learning parameters as a result of the learning by the learning unit. Machine learning device.

14. A discrimination device that outputs a discrimination result for discrimination data using the learning model generated by the computing device described in claim 13, A discrimination data acquisition unit that acquires the aforementioned discrimination data, The system includes a discrimination processing unit that inputs the discrimination data acquired by the discrimination data acquisition unit into the learning model and outputs the discrimination result based on the output value from the learning model. Discrimination device.

15. Using the computing device described in any one of claims 1 to 10, a function is searched for that holds true for the relationships between data within the group of objects to be classified. A discrimination device that performs classification of an unknown object by inputting data of the unknown object into the aforementioned function.

16. From among the discriminant expressions with different characteristics derived by the computing device described in claim 2, a plurality of discriminant expressions are selected based on the discrimination performance of the discriminant expression and the value of the coefficient of variation. A discrimination device that classifies unknown objects to be classified using the aforementioned multiple discrimination formulas.

17. An equation storage unit that stores a second equation derived by the arithmetic device according to claim 11, A calculation unit that calculates the output for an input value using the second equation, An equation utilization device equipped with the following features.

18. A calculation method using a computing device that operates by using a neural network structure including at least an input layer and an output layer, and outputting an output value from the output layer for a plurality of input data (D0, D1, ..., DN) input to the input layer, The input layer of the neural network structure has, as learning parameters for the neural network structure, a plurality of exponents (p0, p1, ..., pN) which are associated with each of the plurality of input data and which are powers of each of the plurality of input data, and coefficients of a predetermined learning function. The plurality of input data (D0, D1, ..., DN) is input to the input layer. The aforementioned computing device receives from the output layer, Multiple power values ​​(D0) are obtained by raising each of the multiple input data inputs to the input layer to the power of a plurality of exponents. p0 D1 p1 ,…,DN pN ) product (YY0 = D0 p0 *D1 p1 *...* DN pN Based on the product of ) and the learning function (De = g(D0, D1, ..., DN)), the output value (y = f(YY0 * De)) is output. Calculation method.

19. The multiple exponents used as learning parameters and the coefficient parameters of the learning function are, A parameter that is learned by using multiple sets of the aforementioned input data as training data, The calculation unit adjusts the coefficient of variation of the output value (y) output when the multiple input data as training data are input to the input layer so that the coefficient of variation becomes small. Calculation method according to claim 18

20. The calculation device outputs and displays a list of the exponents in order of their coefficient of variation, based on the value of the coefficient of variation of the output value (y). The calculation method according to claim 19.

21. The calculation device outputs and displays a list of statistical distribution function types in order of their relative importance, based on the value of the coefficient of variation of the output value (y), in the order of the coefficient of variation. The calculation method according to claim 19.

22. The aforementioned neural network structure is, A hidden layer is further included between the input layer and the output layer. The aforementioned hidden layer is Multiple input data are input via multiple weighting parameters (w0, w1, ..., wN) as learning parameters, and a first hidden node outputs a target value (YY1) defined by the following equation [Equation 1] to the output layer, The system includes a second hidden node to which multiple input data are each input via the multiple weighting parameters, and a bias parameter (b) as a learning parameter is input, and which outputs an additive operation output (BYA) defined by the following equation [Equation 2] to the output layer, The output layer is, Based on the target value (YY1) and the additive calculation output (BYA), the output value (y = f(YY1,BYA)) is output. The calculation method according to claim 18. [Mathematics 1] YY1=D0 p0 *D1 p1 *…*DN pN *W0*W1*…*WN*De (=YY0*W0*W1*...*WN*De) [Math 2] BYA=B*(base) (Σ[n=0→N](wn*pn*dn)) however, base is a positive number other than 1. Dn=base dn (n=0,1,…,N) Wn=base wn (n=0,1,…,N) B = base b That is the case.

23. The coefficient of variation is a value defined by the ratio of the mean value to the standard deviation of the target value YY, as defined in [Equation 3] below. The calculation device determines the weighting parameter wn such that the coefficient of variation is minimized, and then uses the weighting parameter wn to determine B (B = base) such that the difference between YY defined in [Equation 3] below and BYA defined in [Equation 4] below is minimized. b ) Calculate backwards, The calculation method according to claim 19. [Mathematics 3] YY=D0 p0 *D0 p1 *…*DN pN *W0*W1*…*WN*De [Math 4] BY=B*(base) (Σ[n=0→N](wn*pn*dn)) however, base is a positive number other than 1. Dn=base dn (n=0,1,…,N) Wn=base wn (n=0,1,…,N) B=base b

24. After determining the learning parameters, the computing device further uses the second coefficient of variation CV' of ZZ, defined by the following output formula [Equation 5], as the loss function, and searches for the value of the second coefficient of variation CV' (loss amount) that minimizes the feature quantity, wn, by gradient descent, thereby determining the weighting parameter wn. The calculation method according to claim 23. [Math 5] ZZ=f(YY,BYA) ZZ=D0 (p0+w0) *D1 (p1+w1) *…*DN (pN+wN) *De*B*W

25. The multiple power exponents, multiple weighting parameters, and bias parameters used as learning parameters are, A parameter that is learned by using multiple sets of the aforementioned input data as training data, The arithmetic unit adjusts the input layer so that the difference (|YY1 - BYA|) between the target value (YY1) output from the first hidden node and the additive arithmetic output (BYA) output from the second hidden node becomes small when the multiple input data as training data are input to the input layer. The calculation method according to claim 22.

26. The multiple exponents used as learning parameters and the coefficients of the learning function are, A parameter learned by using multiple sets of training data, each set including multiple input data and training data associated with the multiple input data, The computing device adjusts the output value output from the output layer when multiple input data included in the training data are input to the input layer, so that the difference between the output value and the training data included in the training data becomes small. The calculation method according to claim 18.

27. The multiple exponents used as learning parameters and the coefficients of the learning function are, A parameter learned by using multiple sets of training data, each set including multiple input data and training data associated with the multiple input data, The computing device adjusts the output value output from the output layer when multiple input data included in the training data are input to the input layer, so that the difference between the output value and the training data included in the training data becomes small. The calculation method according to claim 22.

28. The aforementioned multiple power values ​​(D0 p0 D1 p1 ,…,DN pN ) product (YY0 = D0 p0 *D1 p1 *...* DN pN From the first equation using the learning function (De = g(D0, D1, ..., DN)), A second equation is created, representing the portion corresponding to the aforementioned input data using variables and a constant term, and presented to the user. The calculation method according to any one of claims 18 to 27.

Citation Information

Patent Citations

  • Learning method and device for neural network

    JP1995141315A

  • Method for operating neural network and data classification system

    JP2021124979A

  • Arithmetic apparatus, integrated circuit, machine learning apparatus, and discrimination apparatus

    JP2022161099A

  • Battery model construction method and battery degradation prediction device

    JP2023018289A

  • Hierarchical neural network device, learning method for determination device, and determination method

    WO2015118686A1