Classification method, classification program, and classification apparatus, as well as model generation method, model generation program, and model generation apparatus

The classification method addresses the increasing computational load in neural networks by using weights from a probability distribution for classification without adjustment, maintaining accuracy with a reduced number of hidden layers and employing the extended Fisher discrimination criterion for output value selection.

JP2025096036APending Publication Date: 2025-06-26OKINAWA INST OF SCI & TECH SCHOOL
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2023212494
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-12-15
Publication Date
2025-06-26

AI Technical Summary

Technical Problem

Neural network models face an increase in computational load when adjusting weights to improve classification accuracy, especially with multiple layers or classes, which needs to be suppressed.

Method used

A classification method that uses weighted data inputted into activation functions, where the weights are generated from a predetermined probability distribution, allowing for classification without adjusting the weights, thereby reducing the computational load.

Benefits of technology

The method effectively suppresses the increase in computational load while maintaining classification accuracy, even with a reduced number of hidden layers, by using a model with one hidden layer and selecting a subset of output values based on the extended Fisher discrimination criterion.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025096036000001_ABST
    Figure 2025096036000001_ABST
Patent Text Reader

Abstract

To provide a classification method, a classification program, and a classification apparatus, as well as a model generation method, a model generation program, and a model generation apparatus, configured to suppress the increase in a computational load.SOLUTION: A classification method is executed for classifying target data into at least two classes, the method including: inputting weighted data to a plurality of activating functions generated by applying a plurality of sets of weight obtained from a predetermined probability distribution to training data; calculating the sum of a subset of output values from the multiple activating functions, for each of the at least two classes; calculating, as class discrimination factors of the at least two classes, the product of the sum of the subset of the output values and the prior probability that the target data belongs to one of the at least two classes; and classifying the target data into a class where the class discrimination factor is the maximum value.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to a classification method, a classification program, and a classification device, as well as a model generation method, a model generation program, and a model generation device.

Background Art

[0002] As a model for classifying data such as images, voices, or texts into a plurality of classes, for example, neural network models such as RBFN (Radial Basis Function Network) described in Non-Patent Document 1 or ELM (Extreme Learning Machine) described in Non-Patent Document 2 are known.

Prior Art Documents

Non-Patent Documents

[0003]

Non-Patent Document 1

Non-Patent Document 2

Summary of the Invention

Problems to be Solved by the Invention

[0004] The neural network model has multiple layers. The connection between each layer of the neural network model is represented by weights. In order to improve the accuracy of classifying data into classes by the neural network model, the weights are adjusted by learning. When adjusting the weights, the computational load increases according to an increase in the number of layers included in the model or the number of classes classified by the model. It is required to suppress the increase in the computational load.

[0005] An object of the present disclosure is to provide a classification method, a classification program, and a classification device capable of suppressing an increase in computational load, as well as a model generation method, a model generation program, and a model generation device.

Means for Solving the Problem

[0006] (1) The classification method according to an embodiment of the present disclosure is executed by one or more processors to classify target data into at least two classes. The classification method includes inputting weighted data obtained by weighting the target data with a plurality of sets of weights to a plurality of activation functions corresponding to each of the at least two classes, the plurality of sets of weights being generated by applying to training data belonging to each of the at least two classes a plurality of sets of weights obtained from a predetermined probability distribution; calculating, for each of the at least two classes, a sum of a subset of output values of the plurality of activation functions; calculating, as a class discrimination element for each of the at least two classes, a product of a prior probability that the target data belongs to one of the at least two classes and the sum of the subset of the output values; and classifying the target data into a class in which the class discrimination element of the class is the maximum value.

[0007] (2) The classification method described in (1) above may further include, when classifying the target data into two classes, calculating the distribution of the weighting data corresponding to each of the plurality of activation functions for each of the two classes, calculating the extended Fisher discrimination criterion corresponding to each of the plurality of activation functions based on the distribution of the weighting data corresponding to each of the plurality of activation functions, and selecting, from the output values of the plurality of activation functions, the output values of a selected number from the larger ones of the values of the corresponding extended Fisher discrimination criteria as a subset of the output values.

[0008] (3) In the classification method described in (1) or (2) above, the prior probability may be estimated as the ratio of each class in the training data.

[0009] (4) In the classification method described in any one of (1) to (3) above, the activation function may be generated as a sum of a plurality of kernels. The kernel may be generated by applying the set of weights to each of the plurality of samples included in the training data.

[0010] (5) In the classification method described in any one of (1) to (4) above, the training data and the target data may be represented in vector form. The set of weights may be represented in matrix form.

[0011] (6) A classification program according to an embodiment of the present disclosure causes one or more processors to execute the classification method described in any one of (1) to (5) above.

[0012] (7) A classification device according to an embodiment of the present disclosure includes one or more processors that execute the classification method described in any one of (1) to (5) above.

[0013] (8) The model generation method according to an embodiment of the present disclosure is executed by one or more processors to generate a model used to classify target data into at least two classes. The model generation method includes obtaining a plurality of sets of weights from a predetermined probability distribution, and generating, as the model, a plurality of activation functions corresponding to each class by applying the plurality of sets of weights to training data belonging to each of the at least two classes.

[0014] (9) The model generation program according to an embodiment of the present disclosure causes one or more processors to execute the model generation method described in (8) above.

[0015] (10) The model generation device according to an embodiment of the present disclosure includes one or more processors that execute the model generation method described in (8) above. [[Effect of the Invention]]

[0016] According to the classification method, classification program, and classification device, as well as the model generation method, model generation program, and model generation device according to an embodiment of the present disclosure, an increase in the computational load is suppressed. [[Brief Description of the Drawings]]

[0017]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11A

Figure 11B

Mode for Carrying Out the Invention

[0018] (Configuration Example of Classification System 1) As shown in FIG. 1, a classification system 1 according to an embodiment of the present disclosure includes a model generation device 10 and a classification device 20. The classification system 1 is configured to classify target data into a plurality of classes. The target data may include various data such as image data, audio data, or text data. For example, the classes for classifying image data may be set to distinguish the objects shown in the image data. The classes for classifying audio data may be set to distinguish the content being spoken, or may be set to distinguish whether the audio is a conversation or music. The classes for classifying text data may be set to distinguish the content of the text.

[0019] The model generation device 10 generates an activation function corresponding to each class based on training data and weights obtained from a predetermined probability distribution, and generates a neural network model to which the generated activation function is applied. The neural network model is hereinafter also simply referred to as a model. The training data includes data classified into each class. The model is configured to output, when target data is input, the result of classifying the target data into a plurality of classes. The model may be configured to output the probability that the target data is classified into each class. The model may also be configured to output the class into which the target data should be classified.

[0020] The classification device 20 inputs target data to the model and classifies the target data into a plurality of classes based on the output of the model.

[0021] In the present disclosure, by generating an activation function based on weights obtained from a predetermined probability distribution, a model is generated without adjusting the weights. Further, in the model to which the activation function generated in this way is applied, even if the number of hidden layers is reduced to one layer, the accuracy of classifying target data into a plurality of classes is less likely to decrease. As a result, an increase in the computational load is suppressed.

[0022] In the classification system 1, the model generation device 10 and the classification device 20 may be integrally configured. That is, the classification device 20 may include the model generation device 10. When the classification device 20 includes the model generation device 10, the classification device 20 may acquire training data and generate a model based on the training data. The classification device 20 may input the target data to be classified into classes to the model generated by the classification device 20 itself and output the output of the model as a classification result.

[0023] The classification system 1 may not include the model generation device 10. When the classification system 1 does not include the model generation device 10, the classification device 20 may acquire a model from an external device, input the target data to be classified into classes to the model, and output the output of the model as a classification result.

[0024] The model generation device 10 or the classification device 20 may be implemented as a cloud service or in an on-premises environment.

[0025] Hereinafter, a configuration example of the classification system 1 will be described.

[0026] <Model generation device 10> The model generation device 10 includes a generation unit 12, a storage unit 14, and an interface 16.

[0027] The generation unit 12 may be configured to include at least one processor. The processor may execute a program that realizes the functions of the generation unit 12. The processor may be realized as a single integrated circuit (IC). The processor may be realized as a plurality of communicably connected integrated circuits and discrete circuits. The processor may be configured to include a CPU (Central Processing Unit). The processor may be configured to include a DSP (Digital Signal Processor) or a GPU (Graphics Processing Unit). The processor may be realized based on various other known technologies.

[0028] The storage unit 14 may be configured to include an electromagnetic storage medium such as a magnetic disk, or may be configured to include a memory such as a semiconductor memory or a magnetic memory. The storage unit 14 may be configured as an HDD (Hard Disk Drive) or an SSD (Solid State Drive). The storage unit 14 stores various information and programs executed by the processor. The storage unit 14 may function as a work memory of the generation unit 12. At least a part of the storage unit 14 may be included in the generation unit 12, or may be configured as a storage device separate from the model generation device 10.

[0029] The interface 16 may be configured to include a communication device capable of communicating wired or wirelessly. The communication device may be configured to communicate based on various communication standards such as a LAN (Local Area Network) or RS-232C or RS-485. The model generation device 10 may output the generated model to the classification device 20 through the interface 16.

[0030] The interface 16 may have an input device or an output device.

[0031] The input device may be configured to include, for example, a touch panel or a touch sensor, or a pointing device such as a mouse. The input device may be configured to include physical keys. The input device may be configured to include a voice input device such as a microphone. The input device may be configured to include a camera. The input device is not limited to these examples and may be configured to include various other devices.

[0032] The output device may be configured to include a display device. The display device may be configured to include, for example, a liquid crystal display (LCD), an organic EL (Electro-Luminescence) display or an inorganic EL display, or a plasma display panel (PDP), etc. The display device is not limited to these displays and may be configured to include various other types of displays. The display device may be configured to include a light emitting device such as an LED (Light Emitting Diode). The display device may be configured to include various other devices. The output device may be configured to include an audio output device such as a speaker that outputs auditory information such as sound. The output device is not limited to these examples and may be configured to include various other devices.

[0033] <Classification device 20> The classification device 20 includes a control unit 22, a storage unit 24, and an interface 26.

[0034] The control unit 22 may be configured to include at least one processor. The control unit 22 may be configured to be the same as or similar to the generation unit 12 of the model generation device 10.

[0035] The storage unit 24 may be configured to be the same as or similar to the storage unit 14 of the model generation device 10. The storage unit 24 may function as a working memory of the control unit 22. At least a part of the storage unit 24 may be included in the control unit 22, or may be configured as a storage device separate from the classification device 20. When the storage unit 14 and the storage unit 24 are configured as external storage devices of the model generation device 10 and the classification device 20, the storage unit 14 and the storage unit 24 may be integrally configured.

[0036] The interface 26 may be configured to be the same as or similar to the interface 16 of the model generation device 10. The classification device 20 may acquire the model generated by the model generation device 10 through the interface 26.

[0037] (Operation example of the classification system 1) Hereinafter, on the premise that the classification system 1 includes the model generation device 10 and the classification device 20, the procedure for the model generation device 10 to generate a model based on training data and the procedure for the classification device 20 to classify target data into classes using the model will be described respectively.

[0038] <Model generation method> The generation unit 12 of the model generation device 10 may execute a model generation method including the procedure illustrated as a flowchart in FIG. 2. The model generation method is executed to generate a model used to classify target data into at least two classes. The model generation method may include a procedure for generating an activation function corresponding to each of at least two classes for classifying target data. The model generation method may be realized as a model generation program to be executed by the generation unit 12 of the model generation device 10. The model generation program may be stored in a non-transitory computer-readable medium.

[0039] In this embodiment, a model is generated based on the following assumptions. Assume that the number of classes for classifying target data is Nc. Each class is distinguished by a number. Each class is represented as class #k. k is a set of integers greater than or equal to 0 and less than or equal to Nc-1. Each class is distinguished into Nc classes from class #0 to class #Nc-1. When the number of classes for classifying target data is 2, k = {0, 1}. Assume that the target data and the training data are vectors having n elements. That is, the target data and the training data may be represented in vector form. Assume that the number of activation functions corresponding to each class is m. Assume that the number of samples included in the training data used to generate the activation function corresponding to each class is v. The number of samples included in the training data may be the same for each class or different for each class.

[0040] Hereinafter, an example of the procedure of the model generation method will be described.

[0041] The generation unit 12 generates a set of weights used to generate m activation functions (step S1). The set of weights may be represented in matrix form. When the set of weights is represented in matrix form, it is represented by an m×n matrix. In other words, the matrix representing the set of weights includes, as row elements, m row vectors having n column elements. Each of the m row vectors corresponds to each of the m activation functions. The row vector of the set of weights used to generate the i-th activation function among the m activation functions is represented by w i as. The elements of the row vector w i are represented by w ij . The set of weights is commonly used for all classes to generate activation functions.

[0042] Specifically, the generation unit 12 obtains a set of weights from a predetermined probability distribution. The generation unit 12 obtains m sets so as to correspond to m activation functions as row vectors of the set of weights. That is, the generation unit 12 obtains a plurality of sets of weights from a predetermined probability distribution. The predetermined probability distribution may be, for example, a standard normal distribution. When a plurality of sets of weights are obtained from the standard normal distribution, each element w ij is represented as w ij ~N(0,1).

[0043] The generation unit 12 sets the value of k to 0 (step S2). When the value of k is set to 0, the activation function corresponding to class #0 is generated by the procedures of steps S3 and S4 described later.

[0044] The generation unit 12 obtains the training data of class #k (step S3). For example, when k = 0, the generation unit 12 obtains the training data of class #0. For example, when k = 1, the generation unit 12 obtains the training data of class #1. As described above, the training data is a column vector having n elements. By multiplying the column vector of the training data on the right by the row vector representing the set of weights, the inner product between the vector of the set of weights and the vector of the training data is calculated. In other words, the set of weights is applied to the training data. The number of samples included in the training data is v. The q-th sample among the v samples of the training data is represented by x q with q as the index. q is a natural number less than or equal to v.

[0045] The generation unit 12 generates an activation function corresponding to class #k (step S4). Specifically, the generation unit 12 calculates the inner product of a set of weights and training data by multiplying an m×n matrix representing a plurality of sets of weights by v samples of training data that are column vectors having n elements. The inner product of the set of weights and the training data is used as a parameter of the activation function. The number of inner products of each set of weights corresponding to each of the m activation functions and each of the v training data is v×m. In other words, v inner products are used as parameters to generate one activation function. The v inner products are represented by z1~z v The z1~z v that are parameters for generating the i-th activation function are represented by Equation (1) as the z q corresponding to the training data x q . [Equation]

[0046] As shown in Equation (2), the generation unit 12 generates the sum of v Gaussian functions to which z1~z v are respectively applied as parameters as the i-th activation function f i (z). [Equation]

[0047] Here, β is the bandwidth and is represented by the following Equation (3). The σ on the right side is the spread of z1~z v . The spread may be the standard deviation of z1~z v , but in this embodiment, it is the median absolute deviation of z1~z v . [Equation]

[0048] The Gaussian function may be used as the kernel of the kernel density estimator. That is, the activation function may be generated as the sum of a plurality of kernels. The kernel K(u) is represented by the following formula (4).

Equation

[0049] The kernel density estimator f(z) is represented by the following formula (5). The argument u of the kernel K(u) includes z q z q is calculated by multiplying each of a plurality of samples included in the training data by a set of weights and adding the results. That is, the kernel may be generated by applying a set of weights to each of a plurality of samples included in the training data. The application of the set of weights to the training data is not limited to multiplication and may be performed by various operations.

Equation

[0050] An example of an activation function generated as the sum of v Gaussian functions is shown in FIG. 3. The horizontal axis represents z. The vertical axis represents the output value of the activation function. z1 to z v correspond to the peaks of the output values of the respective Gaussian functions.

[0051] In the procedure of step S4 described above, the generation unit 12 sequentially sets i from 1 to m and generates m activation functions. In other words, the generation unit 12 generates m activation functions corresponding to class #k as the model of class #k.

[0052] The generation unit 12 adds 1 to k (step S5). The generation unit 12 determines whether k is less than or equal to Nc - 1 (step S6). That is, the generation unit 12 determines whether there remains a class for which an activation function is to be generated. When k is less than or equal to Nc - 1 (step S6: YES), the generation unit 12 returns to the procedure of step S3 in order to generate an activation function corresponding to the next class. The generation unit 12 generates, as a model, an activation function corresponding to class #k according to the value of k by changing the value of k and executing the procedures of steps S3 and S4. That is, the generation unit 12 generates, as models, activation functions corresponding to at least two classes to be classified respectively.

[0053] When k is not less than or equal to Nc - 1 (step S6: NO), the generation unit 12 determines that the generation of activation functions corresponding to all classes is completed, and outputs, as a model, a set of m activation functions corresponding to each class (step S7). After executing the procedure of step S7, the generation unit 12 ends the execution of the flowchart in FIG. 2.

[0054] In the model generation method described above, the activation function is generated as a sum of Gaussian functions. The activation function may be generated, for example, by applying the maximum entropy method based on the prior probability of each class found from training data. The activation function is not limited to these algorithms and may be calculated by various other algorithms.

[0055] <Classification method> The control unit 22 of the classification device 20 may execute a classification method including the procedure illustrated as a flowchart in FIG. 4. The classification method is executed to classify target data into at least two classes. The classification method may be realized as a classification program for causing the control unit 22 of the classification device 20 to execute. The classification program may be stored in a non-transitory computer-readable medium.

[0056] The control unit 22 weights the target data (step S11). The target data is a column vector having n elements. The control unit 22 multiplies the column vector of the target data from the right by an m×n matrix representing a plurality of sets of weights, thereby calculating a column vector having, as elements, m inner products between each vector of the plurality of sets of weights and the vector of the target data. The plurality of sets of weights used in step S11 are the same as the plurality of sets of weights generated in the procedure of step S1 of the model generation method in FIG. 2. The inner product between each vector of the plurality of sets of weights and the vector of the target data is also referred to as weighted data. Each of the m inner products of the weighted data corresponds to one of the m activation functions.

[0057] The control unit 22 sets the value of k to 0 (step S12). When the value of k is set to 0, a class discrimination element y0 of class #0 is calculated in the procedures of steps S13 to S16 described later. As will be described later, the class discrimination element is a value referred to for classifying the target data into classes.

[0058] The control unit 22 inputs the weighted data to a plurality of activation functions corresponding to class #k (step S13). For example, when k = 0, the control unit 22 inputs the corresponding weighted data to each of the m activation functions corresponding to class #0. For example, when k = 1, the control unit 22 inputs the corresponding weighted data to each of the m activation functions corresponding to class #1.

[0059] The control unit 22 selects a subset of the output values of the plurality of activation functions corresponding to class #k (step S14). Specifically, the control unit 22 selects, as a subset, a number of output values less than m from the output values of the m activation functions. The number of output values selected as the subset is also referred to as the selection number and may be set as appropriate. For example, when m = 10000, 10 output values may be selected as the subset.

[0060] The control unit 22 can select a subset using various algorithms. For example, the control unit 22 may select a subset using a norm. When the control unit 22 classifies the target data into two classes, it may select a subset using the Fisher discrimination criterion or may select a subset using the extended Fisher discrimination criterion described later.

[0061] The control unit 22 calculates the sum of the subset of the output values of the plurality of activation functions corresponding to class #k (step S15).

[0062] The control unit 22 k calculates the class discrimination element y of class #k (step S16). Specifically, the control unit 22 multiplies the sum of the subset calculated in step S15 by the prior probability of class #k to calculate the class discrimination element y of class #k. k

[0063] The control unit 22 adds 1 to k (step S17). The control unit 22 determines whether k is less than or equal to Nc - 1 (step S18). That is, the control unit 22 determines whether there is a class for which the class discrimination element is to be calculated. When k is less than or equal to Nc - 1 (step S18: YES), the control unit 22 returns to the procedure of step S13 to calculate the class discrimination element of the next class. By changing the value of k and executing the procedures from step S13 to S16, the control unit 22 calculates the class discrimination element y of class #k according to the value of k. That is, the control unit 22 calculates the class discrimination element of each of at least two classes to be classified. k

[0064] When k is not less than or equal to Nc - 1 (step S18: NO), the control unit 22 determines that the calculation of the class discrimination element for all classes is completed, and classifies the target data into the class for which the class discrimination element is the maximum value (step S19). The function for specifying the number of the class for which the class discrimination element y is the maximum value is represented by argmax(y). k k

[0065] The control unit 22 outputs the classification result of the class from the interface 26 (step S20). The control unit 22 may display the classification result of the class or output it as data to an external device. The control unit 22 may output the classification result of the class in various modes. After executing the procedure of step S20, the control unit 22 ends the execution of the flowchart in FIG. 4.

[0066] <parentheses> In the classification system 1 according to the present disclosure, the model generation device 10 can generate an activation function using weights obtained from a predetermined probability distribution. Using this approach, classification is performed without tuning the weights. This reduces the computational load.

[0067] The classification device 20 obtains an output value by weighting the target data and inputting it into the activation function, selects a subset from the output value, and calculates a class discrimination element by multiplying the sum of the subset of the output value by the prior probability of the class. Then, the classification device 20 classifies the target data into the class for which the class discrimination element is the maximum value. The function of weighting the target data corresponds to the input layer of the neural network model. The function of obtaining the output value of the activation function corresponds to the hidden layer of the neural network model. The function of calculating the class discrimination element and classifying the target data into a class corresponds to the output layer of the neural network model. In the model according to the present disclosure, the output value of the activation function is obtained in one layer. That is, the number of hidden layers in the model according to the present disclosure is one layer. By using a model with one hidden layer, the computational load of the model is reduced.

[0068] The generation unit 12 of the model generation device 10 may generate m activation functions based on the training data, appropriately set the number of selections, and select the activation functions of the number of selections from the m activation functions to generate a model. In this case, the control unit 22 of the classification device 20 may calculate only the activation functions of the number of selections included in the model to obtain the output values. Since the model includes only the activation functions selected from the m activation functions, the data size of the model is reduced. In addition, the computational load for inputting the target data to the model and obtaining the output values of the activation functions is reduced.

[0069] The number of selections may be appropriately adjusted within a range less than m. The maximum value of the range in which the number of selections is adjusted is also referred to as a predetermined number. The generation unit 12 of the model generation device 10 may generate m activation functions and select a predetermined number of activation functions from the m activation functions to generate a model. In this case, the control unit 22 of the classification device 20 sets the number of selections, calculates only the activation functions of the predetermined number included in the model to obtain the output values, and selects the output values of the number of selections from among the output values of the predetermined number. Since the number of activation functions included in the model is reduced to a number less than m, the data size of the model is reduced. In addition, the computational load for inputting the target data to the model and obtaining the output values of the activation functions is reduced.

[0070] (Extended Fisher Discriminant Criterion) As described above, in the classification method according to the present disclosure, when classifying target data into two classes (class #0 and #1), the extended Fisher discriminant criterion may be used to select a subset of the output values of the activation functions. The extended Fisher discriminant criterion is a criterion obtained by extending the Fisher discriminant criterion. Hereinafter, the extended Fisher discriminant criterion will be described.

[0071] First, it is assumed that the probability distribution P(x|C k ) of class #k is represented by the following equation (6). N represents a Gaussian distribution. μ k represents the mean value of the Gaussian distribution. σ k represents the standard deviation of the Gaussian distribution.

Equation

[0072] Also, it is assumed that the relationship between the prior probability that the target data is classified into class #k and the prior probability that the target data is classified into class #1-k is represented by the following inequality (7).

Number

Number

[0073] The reason why the above inequality (7) holds is as follows. As shown in FIG. 5, it is assumed that the probability distributions of class #0 and class #1 have the same shape. f0(x) is the probability distribution of class #0, that is, P(x|C0). f1(x) is the probability distribution of class #1, that is, P(x|C1).

[0074] When x follows P(x|C0), the following inequality (9) holds.

Number

[0075] The product of two Gaussian functions becomes a single Gaussian function. The mean value μ~, standard deviation σ~, and scaling factor of the single Gaussian function calculated as the product of two Gaussian functions are as shown in the following formulas (10A), (10B), and (10C).

Mathematics

[0076] By applying Equation (10A), Equation (10B), and Equation (10C) to Equation (8), the following Equation (11) is derived.

Mathematics

[0077] Similarly, the following Equation (12) is derived.

Mathematics

[0078] By applying Equation (11) and Equation (12) to Inequality (7), the following Inequality (13) is derived.

Mathematics

[0079] Here, F is the Fisher discrimination criterion and is defined by the following Equation (14).

Mathematics

[0080] And the extended Fisher discrimination criterion D k proposed in the present disclosure k is defined by the following Equation (15). T

Mathematics

[0081] The larger the value of the extended Fisher discrimination criterion D k , the higher the discriminability of the characteristics of the target data. D k is valuable in capturing the difference between the probability distribution of class #0 and the probability distribution of class #1. D kThe difference in the probability distributions captured depends not only on the mean value but also on the relative variance. Therefore, even when the mean value μ0 of the probability distribution of class #0 and the mean value μ1 of the probability distribution of class #1 are approximately equal (when μ0 ~ μ1 or F ~ 0), D k can measure the dissimilarity of the probability distributions.

[0082] When the control unit 22 of the classification device 20 classifies the target data into two classes, in the selection procedure of the subset of step S14 of the classification method illustrated in the flowchart of FIG. 4, the extended Fisher discrimination criterion defined by Equation (15) may be used.

[0083] The inequality (7) assumed to derive the extended Fisher discrimination criterion may be realized as a neural network model as illustrated in FIG. 6. For simplicity, the model in FIG. 6 has a single node in the hidden layer for each class. Also, the prior probability of the class is realized by the activation function.

[0084] The upper block in FIG. 6 represents the network of class #0. The lower block in FIG. 6 represents the network of class #1. f i k (·) represents the activation function. The weights (w 11 ~w n1 ) used in the network of class #0 and the weights (w 11 ~w n1 ) used in the network of class #1 are the same.

[0085] In the model of FIG. 6, an input sample belonging to class #0 is input to both the network of class #0 and the network of class #1. The network of class #0 outputs the value of the class discrimination element y0. The network of class #1 outputs the value of the class discrimination element y1. The class of the network that generates the maximum output value as the class discrimination element represents the class of the input sample.

[0086] As described above, the model illustrated in FIG. 6 realizes the inequality (7). The class discrimination element y0 corresponds to the left side of the inequality (7). The class discrimination element y1 corresponds to the right side of the inequality (7). Therefore, when an input sample belonging to class #0 is input to the model, generally y0 > y1 holds.

[0087] When the class of the input sample to the model is unknown, the model converts the input sample with the activation function corresponding to each class, and obtains the class discrimination element y k for each class as the output. The input sample is classified into the class with the maximum class discrimination element.

[0088] As shown in FIG. 7, the model corresponding to class #k has m nodes as hidden layers. Each of the m nodes corresponds to m activation functions generated by executing the model generation method of FIG. 2. The n elements (x1, ···, x n ) of the vector x representing the input data are weighted and input to the hidden layer. The weights applied to the input data are represented as elements of an m×n matrix and are the same as the weights applied to the training data for generating the activation function.

[0089] The output values a1 k ~a m k of each of the m activation functions are input to the output layer. In the output layer, N output values are selected as a subset Ω k ~a m k from the output values a1 N k . That is, the number of elements of the subset Ω N k is N. The number of output values selected as the subset Ω N k , that is, the number of elements of the subset Ω N k is also referred to as the selection number. φ(·) selects a subset Ω k ~a m k from the output values a1 N kSelect and calculate the sum of the subset Ω N k represents a function. In φ(·), the subset Ω N k is selected using the extended Fisher discrimination criterion as the selection criterion.

[0090] As shown in FIG. 8, the model includes a network corresponding to class #0 and a network corresponding to class #1 to classify the input data into class #0 and class #1.

[0091] The upper block in FIG. 8 represents the network of class #0. The lower block in FIG. 8 represents the network of class #1. f i k (·) represents an activation function. The value of i is a natural number less than or equal to m. The value of k is 0 or 1. The weights (w 11 ~w n1 ) used in the network of class #0 and the weights (w 11 ~w n1 ) used in the network of class #1 are the same. In the model of FIG. 8, an input sample of an unknown class is input to both the network of class #0 and the network of class #1. The network of class #0 outputs the value of the class discrimination element y0. The network of class #1 outputs the value of the class discrimination element y1. The class of the network that generates the maximum output value as the class discrimination element represents the class of the input sample.

[0092] In FIG. 8, the subset Ω N k is selected in descending order of the values of the extended Fisher discrimination criteria D1 k ~a m k calculated for each activation function from the output values a1 k ~D m k . The values of the extended Fisher discrimination criteria D1 k ~D m k are calculated for each activation function using Equation (15). The extended Fisher discrimination criterion D1k ~D m k The set of ~D k is represented as D. That is, D k ={D1 k , ···, D m k}.

[0093] Referring to FIG. 9, a method for calculating the set D k of the extended Fisher discrimination criterion will be described. For the training data of each of class #0 and class #1, weighted data obtained by weighting the input data is generated. The weighted data is generated corresponding to each of the m activation functions of each class. That is, the number of elements of the weighted data corresponding to each of the m activation functions of each class matches the number of samples of the training data. The average value of the weighted data corresponding to each activation function of class #0 is represented by μ0, and the variance is represented by σ0 2 . The average value of the weighted data corresponding to each activation function of class #1 is represented by μ1, and the variance is represented by σ1 2 . In the present disclosure, the variance is the mean absolute deviation, but it may be the standard deviation instead.

[0094] For class #0, by applying the average value μ0 and variance σ0 2 of the weighted data corresponding to each activation function to the definition formula (15) of the extended Fisher discrimination criterion, the value of the extended Fisher discrimination criterion corresponding to each activation function of class #0 is calculated. Also, for class #1, by applying the average value μ1 and variance σ1 2 of the weighted data corresponding to each activation function to the definition formula (15) of the extended Fisher discrimination criterion, the value of the extended Fisher discrimination criterion corresponding to each activation function of class #1 is calculated.

[0095] The m activation functions are the extended Fisher discrimination criteria D1 k ~D m kThey may be sorted in descending order of the value of m . The index of the activation function with the largest value of the extended Fisher discriminant criterion is represented by ω1. The index of the activation function with the second largest value of the extended Fisher discriminant criterion is represented by ω2. The index of the activation function with the m-th largest value, that is, the activation function with the smallest value of the extended Fisher discriminant criterion, is ω m . The set of ω1 to ω k representing the order of the m activation functions is denoted as Ω k . That is, Ω m = {ω1, ω2, ···, ω

[0096] For example, for three activation functions f1 k , f2 k , f3 k , if the values of the extended Fisher discriminant criterion calculated for them are D2 k > D3 k > D1 k , then the three activation functions are sorted in the order of f2 k , f3 k , f1 k . In this case, the order of the activation functions is represented by Ω k = {2, 3, 1}.

[0097] Returning to FIG. 8, the sum of the subsets of the output values of the activation functions selected in descending order of the value of the extended Fisher discriminant criterion is calculated by the function φ(·) for each of class #0 and class #1. Then, the class discriminant element y0 of class #0 is calculated by multiplying the sum of the subset of the output values of the activation functions for class #0 by the prior probability P(C0) of class #0. Also, the class discriminant element y1 of class #1 is calculated by multiplying the sum of the subset of the output values of the activation functions for class #1 by the prior probability P(C1) of class #1. The input data is classified into the class with the larger class discriminant element. Specifically, the input data is classified into class #0 when y0 > y1, and classified into class #1 when y0 < y1.

[0098] <Example of Classification by a Model Applying the Extended Fisher Discriminant Criterion> In order to evaluate the accuracy of classifying target data into classes by the model according to the present disclosure described above, a model having 10,000 activation functions in the hidden layer is generated. The model according to this embodiment is configured to select a subset having 10 output values from the output values of the 10,000 activation functions using the extended Fisher discriminant criterion in the output layer and calculate the respective class discrimination elements for class #0 and class #1. That is, the number of selections is set to 10.

[0099] Weights are generated in 10 patterns to report the classification results in a 95% confidence interval. Target data is classified into classes by a model to which the weights of each of the 10 patterns are applied. The accuracy rate of the result of classifying the target data into classes represents the accuracy of the classification result.

[0100] The target data to be classified into classes by the model according to this embodiment is adjusted so that the average value μ0 of the output values of the 10,000 activation functions of class #0 and the average value μ1 of the output values of the 10,000 activation functions of class #1 are substantially the same (become μ0~μ1). In other words, the target data is adjusted so that Δμ (=μ1 - μ0) becomes a value obtained from the standard normal distribution. When Δμ~0, as shown in FIG. 10, the peak of the probability distribution of class #0 and the peak of the probability distribution of class #1 overlap. The model according to this embodiment can classify the target data into classes by the difference in the dispersion of the probability distribution being reflected in the value of the extended Fisher discriminant criterion even when the peaks of the probability distributions substantially coincide.

[0101] In the model according to this embodiment, the interquartile mean is used as the average value of the output values in order to process the outliers of the output values when calculating the dispersion of the output values of the 10,000 activation functions. Also, the median absolute deviation η is used as the standard deviation σ of the output values, and is represented as σ~η / 0.6745. The prior probability of each class is estimated as the ratio of each class in the training data.

[0102] In the model according to this embodiment, the accuracy when classifying the data included in the MNIST dataset and the Fashion-MNIST dataset is evaluated. Further, as a comparative example, in a model that classifies data into classes using the Fisher discrimination criterion, the accuracy when classifying the data included in the MNIST dataset and the Fashion-MNIST dataset is evaluated.

[0103] The MNIST dataset is a dataset of images of 10 digits from 0 to 9. Each of the 10 digits corresponds to 10 classes from class #0 to #9. In the model according to this embodiment, the accuracy rate of the result of classifying 2 out of the 10 digits in the MNIST dataset from each other was calculated.

[0104] The Fashion-MNIST dataset is a dataset of images of 10 types of clothing. Each of the 10 types of clothing corresponds to 10 classes from class #0 to #9. In the model according to this embodiment, the accuracy rate of the result of classifying 2 out of the 10 types of clothing in the Fashion-MNIST dataset from each other was calculated.

[0105] As shown in the table in FIG. 11A, for the MNIST dataset, the result of classifying data into two classes using the model according to this embodiment is compared with the result of classifying data into two classes using the model of the comparative example. Also, as shown in the table in FIG. 11B, for the Fashion-MNIST dataset, the result of classifying data into two classes using the model according to this embodiment is compared with the result of classifying data into two classes using the model of the comparative example. In the tables of FIGS. 11A and 11B, for example, the upper numerical value described in the cell where the row is class #0 and the column is class #1 is the accuracy rate when classifying data into two types of classes, class #0 and class #1, using the model according to this embodiment. The lower numerical value is the accuracy rate when classifying data into two types of classes, class #0 and class #1, using the model according to the comparative example. The accuracy rate is represented in a form where an error in the 95% confidence interval is added to the average value obtained by attempting classification with each of the 10 models generated with weights of 10 patterns.

[0106] According to the table in FIG. 11A, the accuracy rate when classifying the MNIST dataset using the model according to this embodiment is 25.50% higher than the accuracy rate when classifying the MNIST dataset using the model according to the comparative example. According to the table in FIG. 11B, the accuracy rate when classifying the Fashion-MNIST dataset using the model according to this embodiment is 27.97% higher than the accuracy rate when classifying the Fashion-MNIST dataset using the model according to the comparative example. From these results, it can be seen that the classification accuracy is improved by using the extended Fisher discrimination criterion.

[0107] The reason for the low classification accuracy rate of the model according to the comparative example is that, by adjusting the target data so that Δμ~0 in this embodiment, the value of the Fisher discrimination criterion defined by the above formula (14) approaches 0. In contrast to the comparative example, the model to which the extended Fisher discrimination criterion according to this embodiment is applied can improve the classification accuracy rate of the target data where Δμ~0, that is, the target data that is difficult to classify by the Fisher discrimination criterion.

[0108] In the above results, the extended Fisher discriminant criterion has significantly improved the classification accuracy of the Fashion-MNIST dataset compared to that of the MNIST dataset. This is presumably because the output values of the activation function in the Fashion-MNIST dataset are more widely spread than those in the MNIST dataset.

[0109] <Parentheses> As described above, by using the model to which the extended Fisher discriminant criterion according to this embodiment is applied, even when the difference between the peaks of the probability distributions of two classes is small, the accuracy rate when classifying target data into the two classes can be increased.

[0110] Also, just by changing the criterion for extracting a subset of the output values of the activation function to the extended Fisher discriminant criterion while keeping the model with one hidden layer, the classification accuracy rate can be increased. That is, the classification accuracy rate can be increased without complicating the model. As a result, an increase in the processing load of classification is suppressed.

[0111] When using the model to which the extended Fisher discriminant criterion is applied, when the generation unit 12 of the model generation device 10 generates m activation functions based on the training data, it calculates the value of the extended Fisher discriminant criterion corresponding to each activation function. The generation unit 12 may generate a model that combines the m activation functions and the values of the extended Fisher discriminant criterion corresponding to each of the m activation functions. In this case, the control unit 22 of the classification device 20 may calculate the class discrimination element by selecting the output values of the selection number from the output values of each of the m activation functions in descending order of the value of the extended Fisher discriminant criterion. Since the model includes m activation functions, the control unit 22 can adjust the classification accuracy by changing the selection number without changing the model.

[0112] The generation unit 12 of the model generation device 10 generates m activation functions, calculates the values of the extended Fisher discrimination criteria for each of the m activation functions, and then selects the activation functions of the selection number in descending order of the values of the extended Fisher discrimination criteria, and may generate a model including only the selected activation functions. In this case, the control unit 22 of the classification device 20 may calculate only the activation functions of the selection number included in the model to obtain the output values. Since the model includes only the activation functions selected from the m activation functions, the data size of the model becomes smaller. In addition, the computational load for inputting the target data to the model and obtaining the output values of the activation functions is reduced.

[0113] In the model using the extended Fisher discrimination criterion, the selection number may be appropriately adjusted within a range less than m. The generation unit 12 of the model generation device 10 generates m activation functions, calculates the values of the extended Fisher discrimination criteria for each of the m activation functions, and then selects a predetermined number of activation functions in descending order of the values of the extended Fisher discrimination criteria, and may generate a model including only the selected activation functions. The generation unit 12 may set the predetermined number to a number less than m and greater than or equal to the selection number that the control unit 22 of the classification device 20 may set. In this case, the control unit 22 of the classification device 20 sets the selection number, calculates only the predetermined number of activation functions included in the model to obtain the output values, and selects the output values of the selection number from among the predetermined number of output values. Since the number of activation functions included in the model is reduced to a number less than m, the data size of the model becomes smaller. In addition, the computational load for inputting the target data to the model and obtaining the output values of the activation functions is reduced.

[0114] (Another example of feature selection) It is useful to select features to enhance the model performance for classification. The features may be selected from the original dataset. Also, new features may be generated from the original dataset via a mapping function. And a feature selection scheme may be employed to select a subset of features from the newly generated features. The process of generating new features from the features of the original dataset is referred to as "feature construction". The feature selection following feature construction is referred to as "feature extraction".

[0115] Based on the methods adopted for feature selection, the literature on feature selection is classified into three categories: filter, wrapper, and embedded methods. Here, in the filter method, features are ranked according to a given criterion. Since this approach excludes features with low predictive power, it is called the filter method. Some filter methods are described below.

[0116] <Pearson correlation coefficient> The Pearson correlation coefficient is a linear correlation measurement where the feature and the class label are treated as random variables. The strength of the correlation between these two random variables is used to rank the features.

[0117] <Information-theoretic distance> Similar to the Pearson correlation coefficient, the feature and the class label are treated as random variables in information-theoretic measurements. However, a major difference from the Pearson correlation coefficient is that the information-theoretic distance can capture the non-linear dependencies between random variables.

[0118] <Kolmogorov-Smirnov test> The Kolmogorov-Smirnov test is a hypothesis test for determining whether two-class samples are generated from the same distribution. The lower the probability of the null hypothesis, the higher the likelihood that the feature is beneficial for classification.

[0119] <ReliefF method> In this approach, sample x is randomly selected from the training set without replacement. Then, two distances of the given feature (1) d(x, x s ) and (2) d(x, x d ) are measured. (1) d(x, x s ): The distance to the closest x s in the same class as x (2) d(x, x d ): The distance to the closest x d in a different class from x

[0120] To obtain the relevance index J i given by the following equation (16), the process of randomly selecting x and measuring the two distances is repeated n times.

Equation

[0121] J i A large value of J indicates that the feature has a high relevance to classification.

[0122] <Overlap amount> The overlap amount is measured as the amount by which the tail parts of the conditional probability distributions of the two classes overlap for a given feature. The smaller the overlap amount, the higher the ranking of the feature.

[0123] <Fisher discriminant criterion> This criterion is the difference (μ1 - μ0) between the two classes with respect to the sum of the within-class variances σ1 2 + σ0 2 as defined by the above equation (14). 2is the ratio. This ratio is used in Fisher discriminant analysis that projects high-dimensional data onto a straight line. The goal of this projection is to find a straight line such that the Fisher ratio is maximized. After finding the straight line that maximizes the Fisher ratio, classes are discriminated using a linear classifier. The Fisher discriminant criterion evaluates the separation of the probability distributions of classes for which features are ranked. The larger the value of the Fisher discriminant criterion, the more suitable the features are for classification.

[0124] The difference between the present disclosure and the conventional examples is that for a setting where the mean values of the conditional probability distributions of classes are very close to each other, that is, there is a possibility that μ1 - μ0 ~ 0, a filtering method using the extended Fisher discriminant criterion defined by the above formula (10) has been developed.

[0125] Although the embodiments according to the present disclosure have been described based on the drawings and examples, it should be noted that those skilled in the art can make various modifications or alterations based on the present disclosure. Therefore, it should be noted that these modifications or alterations are included in the scope of the present disclosure. For example, the functions etc. included in each component etc. can be rearranged so as not to be logically contradictory, and it is possible to combine a plurality of components etc. into one or divide them.

[0126] All of the constituent elements described in the present disclosure, and / or all of the disclosed methods, or all of the steps of the processes, can be combined in any combination except combinations where these features are mutually exclusive. Also, each of the features described in the present disclosure can be replaced with an alternative feature that serves the same purpose, an equivalent purpose, or a similar purpose, unless explicitly negated. Therefore, unless explicitly negated, each of the disclosed features is merely an example of a comprehensive series of identical or equivalent features.

[0127] Furthermore, embodiments according to the present disclosure are not limited to any specific configurations of the above-described embodiments. Embodiments according to the present disclosure can be extended to all novel features described in the present disclosure, or combinations thereof, or all novel methods described, or processing steps, or combinations thereof.

Description of Reference Numerals

[0128] 1 Classification system 10 Model generation device (12: Generation unit, 14: Storage unit, 16: Interface) 20 Classification device (22: Control unit, 24: Storage unit, 26: Interface)

Claims

1. A classification method executed by one or more processors to classify target data into at least two classes, Inputting weighted data obtained by weighting the target data with a plurality of sets of weights into a plurality of activation functions corresponding to each class, the plurality of activation functions being generated by applying a plurality of sets of weights obtained from a predetermined probability distribution to training data belonging to each of the at least two classes; Calculating the sum of a subset of the output values of the plurality of activation functions for each of the at least two classes; Calculating, as a class discrimination element for one of the at least two classes, the product of the prior probability that the target data belongs to one of the at least two classes and the sum of the subset of the output values; Classifying the target data into the class for which the class discrimination element of the class is the maximum value A classification method comprising:

2. When classifying the target data into two classes, calculating, for each of the two classes, the distribution of the weighted data corresponding to each of the plurality of activation functions; Calculating an extended Fisher discrimination criterion corresponding to each of the plurality of activation functions based on the distribution of the weighted data corresponding to each of the plurality of activation functions; Selecting, as the subset of the output values, a selected number of output values from the output values of the plurality of activation functions, the output values being the larger ones of the values of the corresponding extended Fisher discrimination criteria The classification method according to claim 1, further comprising:

3. The classification method according to claim 1 or 2, wherein the prior probability is estimated as the ratio of each class in the training data.

4. The activation function is generated as a sum of a plurality of kernels, The kernel is generated by applying the set of weights to each of a plurality of samples included in the training data. The classification method according to claim 1 or 2.

5. The training data and the target data are represented in vector form, The set of weights is represented in matrix form. The classification method according to claim 1 or 2.

6. A classification program for causing one or more processors to execute the classification method according to claim 1 or 2.

7. A classification device comprising one or more processors that execute the classification method according to claim 1 or 2.

8. A model generation method executed by one or more processors to generate a model for classifying target data into at least two classes, comprising: obtaining a plurality of sets of weights from a predetermined probability distribution; and generating, as the model, a plurality of activation functions corresponding to each class by applying the plurality of sets of weights to training data belonging to each of the at least two classes. **Claim 9** A model generation program for causing a processor to execute the model generation method according to claim 8. **Claim 10** A model generation apparatus comprising a processor that executes the model generation method according to claim 8.