Learning device, learning method, and learning program, as well as classification device, classification method, and classification program
By minimizing the area above the ROC curve (pAOC) using a differentiable function and hill-descent method, the method addresses the instability of pAUC-based parameter settings, enhancing classification accuracy.
Patent Information
- Application Number
- JP2023556011
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2021-10-29
- Publication Date
- 2025-05-21
- Estimated Expiration
- 2041-10-29
AI Technical Summary
Existing learning methods using partial AUC (pAUC) for two-class classification are prone to falling into local solutions, leading to unstable high accuracy in parameter settings.
The method minimizes the area above the ROC curve (pAOC) to set parameters, using a differentiable function and hill-descent method to avoid local solutions, focusing on the entire ROC curve rather than a partial region.
This approach stabilizes parameter settings, reducing the likelihood of local solutions and achieving higher classification accuracy.
Smart Images

Figure 0007680659000007 
Figure 0007680659000008 
Figure 0007680659000009
Abstract
Description
[Technical field]
[0001] The present invention relates to a learning device, a learning method, and a learning program for setting parameters included in a score function used for two-class classification, and also to a classification device, a classification method, and a classification program for performing two-class classification using a score function including the parameters. [Background technology]
[0002] In two-class classification, a score function is used that takes data as input and outputs a score. Data with a score above a threshold is classified into a positive class, and data with a score below the threshold is classified into a negative class. In order to perform two-class classification with high accuracy, parameters included in the score function are set by machine learning. Two-class classification is used, for example, in video surveillance, fault diagnosis, inspection, medical image diagnosis, and the like that utilize image data.
[0003] As a machine learning method for two-class classification, a learning method in which the parameters of a score function are set to maximize the area under the curve (AUC) (hereinafter, also referred to as a "learning method using AUC"). Here, AUC refers to the area under the receiver operating characteristic (ROC) curve in a graph with the false positive rate on the horizontal axis and the true positive rate on the vertical axis. In two-class classification, a situation may occur in which there is extremely little training data for the positive class compared to the training data for the negative class. Even in such a case, it is known that a learning method using AUC can obtain a score function with high classification accuracy.
[0004] Furthermore, when it is necessary to improve the true positive rate while limiting the false positive rate to a predetermined threshold value α or less, a learning method has been proposed in which parameters of a score function are set so as to maximize the area of pAUC (partial AUC) (hereinafter, also referred to as a "learning method using pAUC"). Here, pAUC refers to the region of AUC where the false positive rate is equal to or less than the threshold value α (the region to the left of the line representing the false positive rate = α). For example, Patent Document 1 is an example of a prior art document that discloses a learning method using pAUC. [Prior art documents] [Patent documents]
[0005] [Patent Document 1] JP 2017-102540 A Summary of the Invention [Problem to be solved by the invention]
[0006] However, the training method using pAUC has a problem in that the parameters tend to fall into a local solution when repeatedly updating the parameters by the hill-climbing method. Therefore, there is a problem that a stable high accuracy cannot be obtained in two-class classification using a score function whose parameters are set by the training method using pAUC.
[0007] One aspect of the present invention has been made in consideration of the above problems, and aims to realize a learning technique in which parameters are less likely to fall into a local solution, and a classification technique that can stably obtain high accuracy. [Means for solving the problem]
[0008] A learning device according to one aspect of the present invention includes a learning means for setting parameters included in a score function for performing two-class classification of data, and the learning means sets the parameters so as to minimize the area of a region above a Receiver Operating Characteristic (ROC) curve obtained from a training data group, in which the false positive rate is equal to or less than a given threshold, in a square with the false positive rate on the horizontal axis and the true positive rate on the vertical axis.
[0009] Furthermore, a learning method according to one aspect of the present invention is a learning method in which a learning device sets parameters included in a score function for performing two-class classification of data, and the parameters included in the score function are set so as to minimize the area of the region above a Receiver Operating Characteristic (ROC) curve obtained from a training data group, in which the false positive rate is equal to or less than a given threshold, in a square with the false positive rate on the horizontal axis and the true positive rate on the vertical axis.
[0010] A learning program according to one aspect of the present invention is a learning program that causes a computer to operate as the above-mentioned learning device, and causes the computer to function as each of the means provided in the learning device.
[0011] A classification device according to one aspect of the present invention includes a classification means for performing two-class classification of data using a score function, and parameters included in the score function are set so as to minimize the area of a region above a Receiver Operating Characteristic (ROC) curve obtained from a training data group, in which the false positive rate is equal to or less than a given threshold, in a square with the false positive rate on the horizontal axis and the true positive rate on the vertical axis.
[0012] A classification method according to one aspect of the present invention is a classification method in which a classification device classifies data into two classes using a score function, and parameters included in the score function are set so as to minimize the area of a region above a Receiver Operating Characteristic (ROC) curve obtained from a training data group, in which the false positive rate is equal to or lower than a given threshold, in a square with the false positive rate on the horizontal axis and the true positive rate on the vertical axis.
[0013] A program according to one aspect of the present invention is a classification program that causes a computer to operate as the classification device described above, and causes the computer to function as each of the means included in the classification device. Effect of the Invention
[0014] According to one aspect of the present invention, it is possible to realize a learning technique in which parameters are unlikely to fall into a local solution. Also, according to one aspect of the present invention, it is possible to realize a classification technique that can stably obtain high accuracy. [Brief description of the drawings]
[0015] [Figure 1] 1 is a block diagram showing a configuration of a learning device according to an embodiment of the present invention; [Diagram 2] 2 is a flowchart showing the flow of a learning method performed by the learning device shown in FIG. 1. [Diagram 3] FIG. 2 is a diagram illustrating an effect of the learning device shown in FIG. [Figure 4] FIG. 2 is a diagram illustrating an effect of the learning device shown in FIG. [Diagram 5] 1 is a block diagram showing a configuration of a classification device according to an embodiment of the present invention. [Figure 6] 6 is a flowchart showing the flow of a classification method performed by the classification device shown in FIG. 5. [Figure 7] FIG. 1 is a block diagram showing a configuration of a learning and classification device according to an embodiment of the present invention. [Figure 8]8 is a block diagram showing a hardware configuration of a computer that functions as the learning device shown in FIG. 1, the classification device shown in FIG. 5, or the learning and classification device shown in FIG. 7. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0016] [Definition of terms] The score function f refers to a function whose domain is the data set X and whose range is the real number R. The score function f includes a parameter θ. The parameter θ may be a scalar or a vector. The value of the score function f for data x belonging to the data set X is denoted as f(x;θ). The score function f is used for two-class classification of the data x. In two-class classification, for example, data x whose score f(x;θ) exceeds a threshold η is determined to belong to a positive class, and data x whose score f(x;θ) is below the threshold η is determined to belong to a negative class.
[0017] In machine learning of a score function f, we use N positively labeled + x data i ∈X(i is 1 to N + (a natural number below) and negatively labeled N - x data j ∈X(j is 1 to N - The following natural numbers are used as training data. i Let x be the positive example. + i and the negatively labeled training data x j Let x be the negative example. - j Also, the positive example x + i The score of f(x + i ;θ) to s + i and the negative example x - j The score of f(x - j ;θ) to s - j Let us denote it as follows. The set of training data {x+ i |1≦i≦N +}∪{x - j |1≦j≦N -} is called the training data set D1.
[0018] False positive rate P η By P η = (score s - j Negative examples x that exceed the threshold η - j ) / (number of negative examples x - j The total number of N - ) is a real number between 0 and 1. η Q η = (score s + i Positive example x exceeds the threshold η + i (Number of positive examples x + i The total number of N + ) is a real number between 0 and 1. In the square [0,1]×[0,1], we change η and find the point (P η ,Q η ) is plotted, a rising curve {(P η ,Q η )|-∞<η<+∞} is obtained. This curve is called the ROC (Receiver Operating Characteristic) curve. The ROC curve divides the square [0,1]×[0,1] into two regions.
[0019] The area under the ROC curve in the square [0,1]×[0,1] is called the AUC (Area Under the Curve). In addition, the false positive rate P η The area where P η The area to the left of the line representing =α is called pAUC (partial AUC). AUC is N + x positive examples + i and N- negative examples x - j With reference to the above, it can be calculated using the following formula (1): Here, I(·) is a function that takes the value 1 when · is true and takes the value 0 when · is false.
[0020]
number
[0021] In addition, the area of pAUC S pAUC is N + x positive examples + i and top score αN - negative examples x - j Here, the negative example x - j is assumed to be sorted in descending score order. That is, s - 1 >s - 2 >…>s - N - In order to normalize the area of the region to the left of the threshold α to 1, N + N - , N + αN - may be replaced with.
[0022]
number
[0023] The area above the ROC curve in the square [0,1]×[0,1] is called the AOC (Area Over the Curve). In addition, the false positive rate P η is less than a given threshold α (P η The area to the left of the line representing =α is called pAOC (partial AOC). AOCis the bottom p positive examples x + i and the top n negative examples x - j Here, the negative example x - j are sorted in descending order of score, and the positive example x + i The maximum negative example score s - 1 Positive example x with a lower score than + i The number of,p,is,p,and the minimum positive score,s, + 1 A negative example with a higher score than x - j The number of s is n. That is, - 1 >…>s - n >s + 1 >s - n+1 >…>s - N - , and s + 1 <… + p - 1 + p+1 <… + N + The above conditions shall be satisfied.
[0024]
number
[0025] In addition, the area of pAOC S pAOC is the bottom p positive examples x + i and top score αN - negative examples x - j In order to normalize the area of the region to the left of the threshold value α to 1, N + N - , N + αN - may be replaced with.
[0026]
number
[0027] Example 1 A first exemplary embodiment of the present invention will be described with reference to the drawings. This exemplary embodiment is a base exemplary embodiment for each of the exemplary embodiments described below.
[0028] (Learning device configuration) The configuration of a learning device 1 according to this exemplary embodiment will be described with reference to Fig. 1. Fig. 1 is a block diagram showing the configuration of the learning device 1.
[0029] The learning device 1 is a device for optimizing a score function f for performing two-class classification of data by machine learning. As shown in Fig. 1, the learning device 1 includes a learning unit 11. The learning unit 11 is one form of the learning means described in the claims.
[0030] The learning unit 11 sets parameters included in a score function for performing two-class classification of data so as to minimize the area of a region above a receiver operating characteristic (ROC) curve obtained from a training data group in a square with the false positive rate on the horizontal axis and the true positive rate on the vertical axis. The functions of the learning unit 11 will be specifically described below.
[0031] The learning unit 11 calculates the area S of pAOC determined from the training data set D1 and the threshold α. pAOC This is a means for setting the parameter θ included in the score function f so as to minimize
[0032] The learning unit 11 may, for example, determine the area S pAOC A differentiable function S that approximates pAOC (θ) to set the parameter θ by the hill-descent method. That is, the learning unit 11 sets the parameter θ by the function S pAOC The gradient of (θ) ∂S pAOC The parameter θ is set by repeating the process of changing the parameter θ in the direction in which (θ) / ∂θ is minimized. For example, if the SpAOC is given by (4) above, the function S pAOC (θ) is given by the following equation (5): where g(·) is a differentiable monotonically increasing function, for example a sigmoid function or a hinge function.
[0033]
number
[0034] Here, the area of pAOC S pAOC A differentiable function S that approximates pAOC The gradient of (θ) ∂S pAOC The direction in which (θ) / ∂θ is minimized and the area S of pAUC pAUC A differentiable function S that approximates pAUC The gradient of (θ) ∂S pAUC The directions in which (θ) / ∂θ is maximized do not coincide with each other. Therefore, the function S pAOC (θ) to set the parameter θ by the hill-descent method, and the function S pAUC This is a process that is substantially different from the process of setting the parameter θ by the hill-climbing method using (θ).
[0035] The learning device 1 may further include a training data set storage unit that stores the training data set D1. The learning device 1 may further include a threshold storage unit that stores the threshold value α. The learning device 1 may further include a threshold setting unit that sets the threshold value α in response to a user operation.
[0036] (Learning device operation) A specific example of the learning method S1 performed by the learning device 1 will be described with reference to Fig. 2. Fig. 2 is a flow diagram showing the flow of the learning method S1.
[0037] Learning method S1 is a learning method in which a learning device sets parameters included in a score function for performing two-class classification of data, and the parameters included in the score function are set so as to minimize the area of a region above a ROC (Receiver Operating Characteristic) curve obtained from a training data group, in which the false positive rate is equal to or less than a given threshold, in a square with the false positive rate on the horizontal axis and the true positive rate on the vertical axis. Learning method S1 will be specifically described below.
[0038] 2, the learning method S1 includes a score calculation process S11, a pair creation process S12, a function creation process S13, a parameter update process S14, and a termination determination process S15. The score calculation process S11, the pair creation process S12, the function creation process S13, the parameter update process S14, and the termination determination process S15 are executed by, for example, the learning unit 11 described above.
[0039] In the score calculation process S11, a score function f is used to calculate N + x positive examples + i Score + i =f(x + i ;θ) and N - negative examples x - j Score - j =f(x - j ;θ) is calculated.
[0040] The pair creation process S12 is the area S of pAOC. pAOC In the pair creation process S12, the learning unit 11 executes, for example, the following steps. First, the learning unit 11 creates a pair of scores to be referenced in order to calculate +i Sort the scores in ascending order and sort the negative examples x - j Secondly, the learning unit 11 sorts the positive examples x + i and top score αN - negative examples x - j By combining p×αN - Pairs of (x + i ,x - j ), where p is the number of + p - 1 + p+1 is a natural number that satisfies.
[0041] The function creation process S13 is the p×αN - Pairs of (x + i ,x - j ) to find the area S pAOC A differentiable function S that approximates pAOC (θ) is created. Here, the function S pAOC (θ) is given, for example, by the above equation (5).
[0042] The parameter update process S14 is a process for updating the function S pAOC The gradient of (θ) ∂S pAOC In the parameter update process S14, the learning unit 11 executes, for example, the following steps. First, the parameter update process S14 updates the function S pAOC The gradient of (θ) ∂S pAOC (θ) / ∂θ is derived. Secondly, in a parameter updating process S14, the parameter θ is updated according to the following equation (6) using a predetermined small positive real number ε.
[0043]
number
[0044] The termination determination process S15 is a process for determining whether or not the parameter θ updated in the parameter update process S14 satisfies a predetermined termination condition. The learning unit 11 repeats the score calculation process S11, pair creation process S12, function creation process S13, and parameter update process S14 described above until a parameter θ that satisfies the termination condition is obtained. Then, when a parameter θ that satisfies the termination condition is obtained, the learning unit 11 ends the learning method S1.
[0045] (Effects of learning devices) The learning method S1 using pAOC has the advantage that the parameter θ is less likely to fall into a local solution than the learning method using pAUC. This advantage will be described with reference to Figs. 3 and 4.
[0046] FIG. 3 is a diagram for explaining a learning method using pAUC.
[0047] The upper part of Fig. 3 shows AUC and pAUC in a square [0,1] x [0,1]. In the upper part of Fig. 3, the area with stripe hatching is pAUC. The area with stripe hatching and the area with dot hatching are combined to form AUC. The lower part of Fig. 3 shows a list of positive example / negative example pairs. In the lower part of Fig. 3, the pairs with stripe hatching are pairs that are referenced in parameter settings that maximize pAUC. In addition, both the pairs with stripe hatching and the pairs with dot hatching are pairs that are referenced in parameter settings that maximize AUC.
[0048] Referring to the upper part of FIG. 3, the parameter θ should be set by taking into account the entire AUC, but in learning using pAUC, the area considered is limited to pAUC. Here, the area not considered, i.e., the area obtained by subtracting pAUC from AUC, is larger than the area considered pAUC. For this reason, in learning using pAUC, the parameter θ is likely to fall into a local solution. Note that this undesirable tendency becomes more pronounced as the threshold α becomes smaller.
[0049] FIG. 4 is a diagram for explaining a learning method using pAOC.
[0050] The upper part of Fig. 4 shows the AOC and pAOC in a square [0,1] × [0,1]. In the upper part of Fig. 4, the area with stripe hatching is the pAOC. The area with stripe hatching and the area with dot hatching are combined to form the AOC. The lower part of Fig. 4 shows a list of positive example / negative example pairs. In the lower part of Fig. 4, the pairs with stripe hatching are pairs that are referenced in parameter setting to minimize the pAOC. In addition, both the pairs with stripe hatching and the pairs with dot hatching are pairs that are referenced in parameter setting to minimize the AOC.
[0051] Referring to the upper part of Figure 4, the parameter θ should be set by taking into account the entire AOC, but in learning using pAOC, the area considered is limited to pAOC. However, the area not considered, that is, the area obtained by excluding pAOC from AOC, is smaller than the area considered pAOC. Therefore, in learning using pAOC, the parameter θ is less likely to fall into a local solution.
[0052] Exemplary embodiment 2 A second exemplary embodiment of the present invention will be described with reference to the drawings. Note that components having the same functions as those described in the first exemplary embodiment are denoted by the same reference numerals, and the description thereof will be omitted.
[0053] (Classification device configuration) The configuration of the classification device 2 according to this exemplary embodiment will be described with reference to Fig. 5. Fig. 5 is a block diagram showing the configuration of the classification device 2.
[0054] The classification device 2 is a device for performing two-class classification of data using a score function f. As shown in Fig. 5, the classification device 2 includes a classification unit 21. The classification unit 21 is one form of the classification means described in the claims.
[0055] The classification unit 21 is a means for performing two-class classification of data x belonging to the test data set D2 using a score function f. Here, the score function f is the area S of pAOC where the parameter θ is determined from the training data set D1. pAOC is the score function set to minimize
[0056] The classification device 2 acquires the parameter θ or a score function f including the parameter θ from, for example, the above-mentioned learning device 1. As a result, the area S of pAOC where the parameter θ is determined from the training data set D1 is pAOC It is possible to perform two-class classification using a score function f set to minimize
[0057] The classification device 2 may further include a test data group storage unit that stores the test data group D2.
[0058] (Classification device operation) A specific example of the classification method S2 performed by the classification device 2 will be described with reference to Fig. 6. Fig. 6 is a flow diagram showing the flow of the classification method S2.
[0059] Classification method S2 is a classification method in which a classification device classifies data into two classes using a score function, and the parameters included in the score function are set so as to minimize the area of a region above a ROC (Receiver Operating Characteristic) curve obtained from a training data group, in which the false positive rate is equal to or less than a given threshold, in a square with the false positive rate on the horizontal axis and the true positive rate on the vertical axis. Classification method S2 will be specifically described below.
[0060] The classification method S2 includes a score calculation process S21, a comparison process S22, a positive determination process S23, and a negative determination process S24, as shown in Fig. 6. The score calculation process S21, the comparison process S22, the positive determination process S23, and the negative determination process S24 are executed by, for example, the classification unit 21 described above.
[0061] The score calculation process S21 is a process for calculating a score s=f(x;θ) by inputting data x into a score function f.
[0062] The comparison process S22 is a process for comparing the score s calculated in the score calculation process S21 with a threshold value η.
[0063] If the score s exceeds the threshold η (for example, s≧η), a positive determination process S23 is executed. The positive determination process S23 is a process for determining that the data x belongs to the positive class.
[0064] If the score s is below the threshold η (for example, s<η), a negative determination process S24 is executed. The negative determination process S24 is a process for determining that the data x belongs to the negative class.
[0065] (Effect of classification device) The parameters included in the score function f referred to by the classifier 2 are set so as to minimize the area of pAUC. Therefore, the classifier 2 can perform two-class classification stably and with high accuracy.
[0066] Exemplary embodiment 3 A third exemplary embodiment of the present invention will be described with reference to the drawings. Note that components having the same functions as those described in the first and second exemplary embodiments are denoted by the same reference numerals, and the description thereof will be omitted.
[0067] The configuration of the learning and classification device 3 according to this exemplary embodiment will be described with reference to Fig. 7. Fig. 7 is a block diagram showing the configuration of the learning and classification device 3.
[0068] The learning and classification device 3 is a device that combines the functions of the above-mentioned learning device 1 and classification device 2. As shown in FIG.
[0069] As described above, the learning unit 11 determines the area S of pAOC determined from the training data set D1. pAOCThe classification unit 21 is a means for setting a parameter θ included in the score function f so as to minimize the above. As described above, the classification unit 21 is a means for performing two-class classification of the data x belonging to the test data group D2 using the score function f.
[0070] In the learning / classification device 3, (1) the learning unit 11 executes the above-mentioned learning method S1 to set the parameter θ included in the score function f, and (2) the classification unit 21 executes the above-mentioned classification method S2 to perform two-class classification of the data x. Therefore, both learning and classification can be achieved by a single device.
[0071] [Software implementation example] Some or all of the functions of the learning device 1, the classification device 2, and the learning / classification device 3 (hereinafter referred to as "learning devices, etc.") may be realized by hardware such as an integrated circuit (IC chip), or by software.
[0072] In the latter case, the learning device etc. is realized, for example, by a computer that executes instructions of a program, which is software that realizes each function. An example of such a computer (hereinafter, referred to as computer C) is shown in FIG. 8. The computer C has at least one processor C1 and at least one memory C2. The memory C2 stores a program P for operating the computer C as a learning device etc. In the computer C, the processor C1 reads and executes the program P from the memory C2, thereby realizing each function of the learning device etc.
[0073] The processor C1 may be, for example, a central processing unit (CPU), a graphic processing unit (GPU), a digital signal processor (DSP), a micro processing unit (MPU), a floating point number processing unit (FPU), a physics processing unit (PPU), a microcontroller, or a combination of these. The memory C2 may be, for example, a flash memory, a hard disk drive (HDD), a solid state drive (SSD), or a combination of these.
[0074] The computer C may further include a RAM (Random Access Memory) for expanding the program P during execution and for temporarily storing various data. The computer C may further include a communication interface for transmitting and receiving data to and from other devices. The computer C may further include an input / output interface for connecting input / output devices such as a keyboard, a mouse, a display, and a printer.
[0075] Furthermore, the program P can be recorded on a non-transitory tangible recording medium M that can be read by the computer C. Such a recording medium M can be, for example, a tape, a disk, a card, a semiconductor memory, or a programmable logic circuit. The computer C can obtain the program P via such a recording medium M. Furthermore, the program P can be transmitted via a transmission medium. Such a transmission medium can be, for example, a communication network or broadcast waves. The computer C can also obtain the program P via such a transmission medium.
[0076] [Additional Note 1] The present invention is not limited to the above-described embodiment, and various modifications are possible within the scope of the claims. For example, embodiments obtained by appropriately combining the technical means disclosed in the above-described embodiment are also included in the technical scope of the present invention.
[0077] [Additional Note 2] Some or all of the above-described embodiments can be described as follows. However, the present invention is not limited to the following described aspects.
[0078] (Appendix 1) A learning device comprising: a learning means for setting parameters included in a score function for performing two-class classification of data, the learning means setting the parameters so as to minimize the area of a region above a Receiver Operating Characteristic (ROC) curve obtained from a training data group, in which the false positive rate is equal to or less than a given threshold, in a square with the false positive rate on the horizontal axis and the true positive rate on the vertical axis.
[0079] According to the above configuration, it is possible to realize a learning technique in which parameters are unlikely to fall into a local solution.
[0080] (Appendix 2) 2. The learning device according to claim 1, wherein the learning means sets the parameters by a hill-descent method using a differentiable function that approximates the area.
[0081] According to the above configuration, it is possible to realize a learning technique in which parameters are unlikely to fall into a local solution.
[0082] (Appendix 3) The learning means (1) uses the score function to + x positive examples + i Score + i and N - negative examples x -j Score - j (2) score calculation process to calculate the positive example x + i Sort the scores in ascending order and sort the negative examples x - j Sort by score in descending order, and then s + p - 1 + p+1 Let p be a natural number that satisfies the above, the threshold value be α, and let x + i and top score αN - negative examples x - j By combining p×αN - Pairs of (x + i ,x - j ), (3) p × αN - Pairs of (x + i ,x - j 3. The learning device according to claim 2, further comprising: (a) a function creation process for creating a differentiable function that approximates the area using a gradient of the function; and (b) a parameter update process for updating the parameters using a gradient of the function, the parameter being set by repeating the steps of: (1) creating a function that approximates the area using a gradient of the function; and (2) updating the parameters using a gradient of the function, until a predetermined termination condition is satisfied.
[0083] According to the above configuration, appropriate parameters can be set with high accuracy.
[0084] (Appendix 4) A learning method in which a learning device sets parameters included in a score function for performing two-class classification of data, the learning method being characterized in that the parameters included in the score function are set so as to minimize the area of a region above a Receiver Operating Characteristic (ROC) curve obtained from a training data group, in which the false positive rate is equal to or less than a given threshold, in a square with the false positive rate on the horizontal axis and the true positive rate on the vertical axis.
[0085] According to the above method, it is possible to implement a learning technique in which parameters are less likely to fall into a local optimum.
[0086] (Appendix 5) A learning program that causes a computer to operate as the learning device according to any one of claims 1 to 3, and causes the computer to function as each of the means provided in the learning device.
[0087] (Appendix 6) A classification device comprising: a classification means for classifying data into two classes using a score function; and parameters included in the score function are set so as to minimize the area of a region above a Receiver Operating Characteristic (ROC) curve obtained from a training data group, in which the false positive rate is equal to or lower than a given threshold, in a square with the false positive rate on the horizontal axis and the true positive rate on the vertical axis.
[0088] According to the above configuration, a classification technique that can stably obtain high accuracy can be realized.
[0089] (Appendix 7) A classification method in which a classification device classifies data into two classes using a score function, the parameters of the score function being set so as to minimize the area of a region above a Receiver Operating Characteristic (ROC) curve obtained from a training data group, in a square with the false positive rate on the horizontal axis and the true positive rate on the vertical axis, where the false positive rate is equal to or less than a given threshold.
[0090] According to the above method, a classification technique that can stably obtain high accuracy can be realized.
[0091] (Appendix 8) A classification program that causes a computer to operate as the classification device described in appendix 6, and causes the computer to function as each of the means provided in the classification device.
[0092] (Appendix 9) A method for generating a score function for performing two-class classification of data, comprising setting parameters included in the score function so as to minimize the area of a region above a Receiver Operating Characteristic (ROC) curve obtained from a training data group, in a square with the false positive rate on the horizontal axis and the true positive rate on the vertical axis, where the false positive rate is equal to or less than a given threshold.
[0093] According to the above method, a classification technique that can stably obtain high accuracy can be realized.
[0094] (Appendix 10) A score function for causing a computer to perform two-class classification of data, the parameters included in the score function being set so as to minimize the area of a region above a Receiver Operating Characteristic (ROC) curve obtained from a training data group, in a square with the false positive rate on the horizontal axis and the true positive rate on the vertical axis, where the false positive rate is equal to or less than a given threshold.
[0095] According to the above configuration, a classification technique that can stably obtain high accuracy can be realized.
[0096] (Appendix 11) A learning device comprising: a learning means for setting parameters included in a score function for performing two-class classification of data, the learning means setting the parameters so as to minimize the area of the region above a Receiver Operating Characteristic (ROC) curve obtained from a training data group in a square with the false positive rate on the horizontal axis and the true positive rate on the vertical axis.
[0097] According to the above configuration, it is possible to realize a learning technique in which parameters are unlikely to fall into a local solution.
[0098] (Appendix 12) A learning method in which a learning device sets parameters included in a score function for performing two-class classification of data, the learning method being characterized in that the parameters included in the score function are set so as to minimize the area of the region above a Receiver Operating Characteristic (ROC) curve obtained from a training data group in a square with the false positive rate on the horizontal axis and the true positive rate on the vertical axis.
[0099] According to the above method, it is possible to realize a learning technique in which parameters are unlikely to fall into a local optimum.
[0100] (Appendix 13) A classification device comprising: a classification means for classifying data into two classes using a score function; and parameters included in the score function are set so as to minimize the area of the region above a Receiver Operating Characteristic (ROC) curve obtained from a training data group in a square with the false positive rate on the horizontal axis and the true positive rate on the vertical axis.
[0101] According to the above configuration, a classification technique that can stably obtain high accuracy can be realized.
[0102] (Appendix 14) A classification method in which a classification device classifies data into two classes using a score function, the parameters of the score function being set so as to minimize the area of the region above a Receiver Operating Characteristic (ROC) curve obtained from a training data group in a square with the false positive rate on the horizontal axis and the true positive rate on the vertical axis.
[0103] According to the above configuration, a classification technique that can stably obtain high accuracy can be realized.
[0104] [Additional Note 3] A part or all of the above-described embodiments can be further expressed as follows.
[0105] A learning device comprising at least one processor, the processor executing a learning process for setting parameters included in a score function for performing two-class classification of data, the learning process setting the parameters so as to minimize the area of a region above a Receiver Operating Characteristic (ROC) curve obtained from a training data group, in a square with the false positive rate on the horizontal axis and the true positive rate on the vertical axis, where the false positive rate is equal to or less than a given threshold.
[0106] The learning device may further include a memory, and the memory may store a program for causing the processor to execute the learning process. The program may be recorded in a computer-readable, non-transitory, tangible recording medium. [Explanation of symbols]
[0107] 1. Learning device 11. Learning Department 2...Classification device 21...Classification section 3. Learning and classification device
Claims
1. A learning means for setting parameters included in a score function for performing two-class classification of data, The learning means sets the parameters so as to minimize the area of a region above an ROC (Receiver Operating Characteristic) curve obtained from a training data group, in which the false positive rate is equal to or less than a given threshold, in a square with the false positive rate on the horizontal axis and the true positive rate on the vertical axis. A learning device characterized by:
2. the learning means sets the parameters by a hill-descent method using a differentiable function that approximates the area.
2. The learning device according to claim 1 .
3. The learning means (1) uses the score function to + Positive examples x + i Score of + i and N included in the training data - negative examples x - j Score of - j (2) a score calculation process for calculating a positive example x + i Sort the scores in ascending order and sort the negative examples x - j Sorting in descending order of score, s + p <s - 1 <s + p+1 Let p be a natural number that satisfies the above, the threshold value be α, and let x + i and top score αN - negative examples x - j By combining p×αN - Pairs (x + i , x - j (3) pair creation process to create p × αN - Pairs (x + i , x - j (3) a function creation process for creating a differentiable function that approximates the area using a gradient of the function; and (4) a parameter update process for updating the parameters using a gradient of the function, thereby setting the parameters.
3. The learning device according to claim 2.
4. A learning method in which a learning device sets parameters included in a score function for performing two-class classification of data, the method comprising the steps of: In a square with the false positive rate on the horizontal axis and the true positive rate on the vertical axis, the parameters included in the score function are set so as to minimize the area of the region above an ROC (Receiver Operating Characteristic) curve obtained from the training data group, where the false positive rate is equal to or less than a given threshold. A learning method comprising:
5. A learning program for causing a computer to operate as the learning device according to any one of claims 1 to 3, the learning program causing the computer to function as each of the means provided in the learning device.
6. A classification means for classifying data into two classes using a score function, The parameters included in the score function are set so as to minimize the area of a region above an ROC (Receiver Operating Characteristic) curve obtained from a training data group, in a square with the false positive rate on the horizontal axis and the true positive rate on the vertical axis, where the false positive rate is equal to or less than a given threshold value. A classification device comprising:
7. A classification method in which a classification device classifies data into two classes using a score function, The parameters included in the score function are set so as to minimize the area of a region above an ROC (Receiver Operating Characteristic) curve obtained from a training data group, in a square with the false positive rate on the horizontal axis and the true positive rate on the vertical axis, where the false positive rate is equal to or less than a given threshold value. A classification method characterized by:
8. 7. A classification program for causing a computer to operate as the classification device according to claim 6, the classification program causing the computer to function as each of the means included in the classification device.
Citation Information
Patent Citations
Classification device, method, and program
JP2017102540A
Classification device, classification method and classification program
JP2020071708A
Time series data analysis method, time-series data analyzer and computer program
JP2020170214A
Media Content Selection
US20180173400A1