Data Classification Method, Device, Medium and Electronic Device Based on Base Classifiers
By calculating the accuracy and diversity indicators of the base classifier combination, selecting the base classifier combination with the largest evaluation indicators, solving the problem of unbalanced data sets and data deviations, improving the accuracy and robustness of the integrated classifiers, and improving classification efficiency.
Patent Information
- Application Number
- CN201910314159.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2019-04-18
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2040-01-15
AI Technical Summary
When selecting a base classifier, the prior art faces the problem of poor classification results due to evaluation indicators that are not suitable for unbalanced data sets, and poor classification prediction results due to inconsistent data deviations and data dimensions.
By calculating the accuracy and diversity indicators of multiple base classifier combinations, the area (AUC) and diversity indicators under the subject's working characteristic curve are used, combined with the weighted harmonization average, the base classifier combination with the largest evaluation indicators is determined for data classification.
Improves the accuracy and robustness of integrated classifiers, is suitable for unbalanced data sets, quickly builds integrated classifiers, and improves classification efficiency and performance.
Smart Images

Figure CN110097098B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of machine learning technology, and in particular, to a data classification method based on a base classifier, a data classification device based on a base classifier, a computer-readable medium, and an electronic device. Background Art
[0002] In the scenario of data classification, especially in some data mining competitions, it is necessary to fuse multiple classification models to improve the classification effect. In this case, Ensemble Learning came into being. Stacking (also known as Stacked Generalization) is an important implementation framework in Ensemble Learning and has broad application prospects.
[0003] The base classifiers of Stacking often need to satisfy both accuracy and diversity simultaneously. The selection of base classifiers greatly affects the performance of the Stacking ensemble classifier. When selecting base classifiers, it is usually necessary to evaluate the base classifiers using evaluation metrics.
[0004] Currently, on the one hand, the metrics used to evaluate base classifiers are not suitable for imbalanced data sets with obvious data differences, resulting in poor classification effects of the selected base classifiers. On the other hand, due to the problems of data deviation and inconsistent data dimensions of the data to be classified, it is difficult to uniformly process the data, and the classification prediction results are not ideal.
[0005] It should be noted that the information disclosed in the above background art section is only used to enhance the understanding of the background of the present disclosure, and thus may include information that does not constitute the prior art known to those of ordinary skill in the art. Summary of the Invention
[0006] The purpose of the present disclosure is to provide a data classification method based on a base classifier, a data classification device based on a base classifier, a computer-readable medium, and an electronic device, so as to at least to some extent overcome the problem of poor classification effects of the selected base classifiers caused by the limitations and defects of the related art.
[0007] According to one aspect of the present disclosure, there is provided a data classification method based on base classifiers, which is applied to an ensemble classifier including a plurality of base classifiers. The data classification method based on base classifiers includes: determining a plurality of combinations of base classifiers, and calculating an accuracy index and a diversity index for each combination of base classifiers; wherein each combination of base classifiers is composed of a predetermined number of base classifiers; sorting each combination of base classifiers according to the accuracy index of each combination of base classifiers, and determining an accuracy ranking value for each combination of base classifiers; sorting each combination of base classifiers according to the diversity index of each combination of base classifiers, and determining a diversity ranking value for each combination of base classifiers; calculating an evaluation index for each combination of base classifiers according to the accuracy ranking value and the diversity ranking value of each combination of base classifiers, determining the base classifiers in the combination of base classifiers with the largest evaluation index as the selected base classifiers, and using the selected base classifiers to classify the data to be classified.
[0008] According to one aspect of the present disclosure, there is provided a data classification device based on base classifiers, which is applied to an ensemble classifier including a plurality of base classifiers. The data classification device based on base classifiers includes an index calculation module, a first sorting module, a second sorting module, and a data classification module.
[0009] Specifically, the index calculation module is configured to determine a plurality of combinations of base classifiers, and calculate an accuracy index and a diversity index for each combination of base classifiers; wherein each combination of base classifiers is composed of a predetermined number of base classifiers; the first sorting module is configured to sort each combination of base classifiers according to the accuracy index of each combination of base classifiers, and determine an accuracy ranking value for each combination of base classifiers; the second sorting module is configured to sort each combination of base classifiers according to the diversity index of each combination of base classifiers, and determine a diversity ranking value for each combination of base classifiers; the data classification module is configured to calculate an evaluation index for each combination of base classifiers according to the accuracy ranking value and the diversity ranking value of each combination of base classifiers, determine the base classifiers in the combination of base classifiers with the largest evaluation index as the selected base classifiers, and use the selected base classifiers to classify the data to be classified.
[0010] Optionally, the index calculation module includes an accuracy index determination unit.
[0011] Specifically, the accuracy index determination unit is configured to calculate the accuracy index of the base classifiers included in the combination of base classifiers; calculate the average value of the accuracy indexes of the base classifiers included in the combination of base classifiers as the accuracy index of the combination of base classifiers.
[0012] Optionally, the accuracy index of the base classifier is the area under the receiver operating characteristic curve.
[0013] Optionally, the accuracy metric determination unit is configured to perform: determining classification results of each test sample in the test set using a base classifier, determining a receiver operating characteristic curve based on the classification results of each test sample and a preset threshold corresponding to each test sample, and calculating an area under the receiver operating characteristic curve as an accuracy metric of the base classifier.
[0014] Optionally, the metric calculation module includes a diversity metric determination unit.
[0015] Specifically, the diversity metric determination unit is used to calculate diversity metrics between pairs of base classifiers in a base classifier ensemble, and calculate an average of the diversity metrics between pairs of base classifiers as a diversity metric of the base classifier ensemble.
[0016] Optionally, the base classifier ensemble includes a first base classifier and a second base classifier; wherein, the diversity metric determination unit is configured to perform: inputting a test sample into the first base classifier and the second base classifier respectively, determining classification results of the first base classifier and the second base classifier respectively, and calculating a diversity metric between the first base classifier and the second base classifier based on the classification results of the first base classifier and the second base classifier.
[0017] Optionally, the data classification module includes an evaluation metric calculation unit.
[0018] Specifically, the evaluation metric calculation unit is configured to perform: determining the number of base classifier ensembles; determining a first value using the number of base classifier ensembles and an accuracy ranking value of the base classifier ensemble, and determining a second value using the number of base classifier ensembles and a diversity ranking value of the base classifier ensemble; calculating a weighted harmonic mean of the first value and the second value as an evaluation metric of the base classifier ensemble.
[0019] Optionally, the evaluation metric calculation unit is configured to perform: determining the number of base classifiers in an ensemble classifier; determining the number of base classifier ensembles based on the number of base classifiers in the ensemble classifier and a predetermined number.
[0020] Optionally, the evaluation metric calculation unit is configured to perform: configuring a weight of the first value as a first weight and configuring a weight of the second value as a second weight; wherein, both the first weight and the second weight are greater than or equal to 0, and the sum of the first weight and the second weight is 1; calculating a weighted harmonic mean of the first value and the second value using the first weight and the second weight such that the weighted harmonic mean is maximum, and determining the maximum weighted harmonic mean as an evaluation metric of the base classifier ensemble.
[0021] Optionally, the data classification module further includes a classification result determination module.
[0022] Specifically, the classification result determination module is configured to perform: inputting the data to be classified into the selected base classifier to determine an intermediate classification result; and inputting the intermediate classification result into the secondary classifier layer of the integrated classifier to determine the classification result of the data to be classified.
[0023] According to one aspect of the present disclosure, there is provided a computer-readable medium having a computer program stored thereon, and when the computer program is executed by a processor, the above-described data classification method based on a base classifier is implemented.
[0024] According to one aspect of the present disclosure, there is provided an electronic device, including: one or more processors; and a storage device for storing one or more programs, and when the one or more programs are executed by the one or more processors, the one or more processors are caused to implement the data classification method based on a base classifier as described above.
[0025] In the technical solutions provided in some embodiments of the present disclosure, the accuracy index and the diversity index of each combination of base classifiers under the integrated classifier are calculated, the accuracy ranking and the diversity ranking are respectively performed on each combination of base classifiers according to the accuracy index and the diversity index of each combination of base classifiers, the accuracy ranking value and the diversity ranking value of each combination of base classifiers are respectively determined, the evaluation index of each combination of base classifiers is calculated according to the accuracy ranking value and the diversity ranking value of the combination of base classifiers, the base classifier is selected based on the evaluation index, and the selected base classifier is used to classify the data to be classified. On the one hand, the present disclosure calculates the evaluation index by using the accuracy ranking value and the diversity ranking value, avoiding the problem of incorrect calculation of the evaluation index caused by data deviation and / or inconsistent data dimensions, and helping to improve the classification accuracy of the integrated classifier; on the other hand, the present disclosure comprehensively considers the accuracy and diversity of the base classifier, making the integrated classifier more robust; on the still other hand, the solution described in the present disclosure has universality, can quickly assist in building an integrated classifier, and improves the efficiency and performance of the integrated classifier applied to classification problems.
[0026] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] The accompanying drawings herein are incorporated into the specification and form a part of the specification, showing embodiments consistent with the present disclosure, and are used together with the specification to explain the principles of the present disclosure. Obviously, the drawings in the following description are only some embodiments of the present disclosure, and those of ordinary skill in the art can obtain other drawings without creative efforts based on these drawings. In the drawings:
[0028] Figure 1A schematic diagram showing an exemplary system architecture to which the data classification method or data classification apparatus based on a base classifier according to an embodiment of the present disclosure can be applied;
[0029] Figure 2 A schematic diagram showing the structure of a computer system of an electronic device suitable for implementing an embodiment of the present disclosure;
[0030] Figure 3 A schematic diagram showing the Stacking ensemble classification architecture;
[0031] Figure 4 A flowchart schematically showing a data classification method based on a base classifier according to an exemplary embodiment of the present disclosure;
[0032] Figure 5 A schematic diagram showing a receiver operating characteristic curve (ROC curve) according to an exemplary embodiment of the present disclosure;
[0033] Figure 6 A schematic diagram showing that a base classifier combination includes multiple base classifiers;
[0034] Figure 7 A block diagram schematically showing a data classification apparatus based on a base classifier according to an exemplary embodiment of the present disclosure;
[0035] Figure 8 A block diagram schematically showing an index calculation module according to an exemplary embodiment of the present disclosure;
[0036] Figure 9 A block diagram schematically showing an index calculation module according to another exemplary embodiment of the present disclosure;
[0037] Figure 10 A block diagram schematically showing a data classification module according to an exemplary embodiment of the present disclosure;
[0038] Figure 11 A block diagram schematically showing a data classification module according to another exemplary embodiment of the present disclosure. Detailed implementation manners
[0039] Example embodiments will now be described more fully with reference to the accompanying drawings. However, the example embodiments can be implemented in various forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided so that this disclosure will be thorough and complete, and will fully convey the concept of the example embodiments to those skilled in the art. The features, structures, or characteristics described may be combined in any suitable manner in one or more embodiments. In the following description, numerous specific details are provided to give a thorough understanding of the embodiments of the present disclosure. However, those skilled in the art will recognize that the technical solutions of the present disclosure may be practiced without one or more of the specific details, or may be implemented using other methods, components, devices, steps, etc. In other cases, well-known technical solutions are not shown or described in detail to avoid obscuring aspects of the present disclosure.
[0040] In addition, the accompanying drawings are only schematic illustrations of the present disclosure and are not necessarily drawn to scale. The same reference numerals in the drawings denote the same or similar parts, and thus repeated descriptions thereof will be omitted. Some of the block diagrams shown in the drawings are functional entities and do not necessarily correspond to physically or logically independent entities. These functional entities may be implemented in software, or in one or more hardware modules or integrated circuits, or in different networks and / or processor devices and / or microcontroller devices.
[0041] The flowcharts shown in the accompanying drawings are merely illustrative and do not necessarily include all steps. For example, some steps may be decomposed, while some steps may be combined or partially combined, so the actual execution order may change according to the actual situation. Additionally, the following terms "first" and "second" are for the purpose of distinction only and should not be construed as a limitation of the content of the present disclosure.
[0042] Figure 1 A schematic diagram showing an exemplary system architecture to which the data classification method or data classification apparatus based on a base classifier according to embodiments of the present disclosure can be applied;
[0043] As Figure 1 shown, the system architecture 1000 may include one or more of the terminal devices 1001, 1002, 1003, a network 1004, and a server 1005. The network 1004 serves as a medium for providing a communication link between the terminal devices 1001, 1002, 1003 and the server 1005. The network 1004 may include various connection types, such as wired, wireless communication links, or fiber optic cables, etc.
[0044] It should be understood that Figure 1The numbers of the terminal devices, network, and server in it are merely illustrative. According to the implementation requirements, there can be any number of terminal devices, network, and server. For example, the server 1005 can be a server cluster composed of multiple servers, etc.
[0045] Users can use the terminal devices 1001, 1002, 1003 to interact with the server 1005 through the network 1004 to receive or send messages, etc. The terminal devices 1001, 1002, 1003 can be various electronic devices with a display screen, including but not limited to smartphones, tablets, portable computers, desktop computers, and so on.
[0046] The server 1005 can be a server that provides various services. For example, the server 1005 can calculate the accuracy index and diversity index of the combination of each base classifier in the integrated classifier, sort the combinations of each base classifier according to the accuracy index of each base classifier combination to determine the accuracy sorting value of each base classifier combination, sort the combinations of each base classifier according to the diversity index of each base classifier combination to determine the diversity sorting value of each base classifier combination, calculate the evaluation index of each base classifier combination according to the accuracy sorting value and diversity sorting value of each base classifier combination, determine the base classifier in the base classifier combination with the largest evaluation index as the selected base classifier, and use the selected base classifier to classify the data to be classified.
[0047] Specifically, the area under the ROC (Receiver Operating Characteristic) curve can be used as the accuracy index of the base classifier, and this index can be expressed as AUC (Area Under Curve).
[0048] After constructing an integrated classifier based on the selected base classifier, the server 1005 can obtain the data to be classified from the terminal devices 1001, 1002, 1003 through the network 1004. Subsequently, the server 1005 uses the integrated classifier to perform classification processing on the data to be classified to obtain a classification result. In some scenarios that require feedback, the server 1005 can feedback the classification result to the terminal devices 1001, 1002, 1003.
[0049] It should be noted that the data classification method based on the base classifier provided by the exemplary embodiment of the present disclosure is generally executed by the server 1005. Correspondingly, the data classification device based on the base classifier described below is generally configured in the server 1005.
[0050] However, the following data classification method based on the base classifier can also be implemented by the terminal devices 1001, 1002, 1003, and the present disclosure does not make special restrictions on this.
[0051] Figure 2 The structural diagram of a computer system of an electronic device suitable for implementing the exemplary embodiments of the present disclosure is shown.
[0052] It should be noted that Figure 2 The computer system 200 of the electronic device shown is only an example and should not impose any limitations on the functions and usage scope of the embodiments of the present disclosure.
[0053] As Figure 2 shown, the computer system 200 includes a central processing unit (CPU) 201, which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 202 or the program loaded from the storage section 208 into the random access memory (RAM) 203. In the RAM 203, various programs and data required for system operations are also stored. The CPU 201, ROM 202, and RAM 203 are connected to each other via a bus 204. The input / output (I / O) interface 205 is also connected to the bus 204.
[0054] The following components are connected to the I / O interface 205: an input section 206 including a keyboard, a mouse, etc.; an output section 207 including a cathode ray tube (CRT), a liquid crystal display (LCD), etc. and a speaker, etc.; a storage section 208 including a hard disk, etc.; and a communication section 209 including a network interface card such as a LAN card, a modem, etc. The communication section 209 performs communication processing via a network such as the Internet. A drive 210 is also connected to the I / O interface 205 as needed. A removable medium 211, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 210 as needed so that a computer program read from it can be installed into the storage section 208 as needed.
[0055] Specifically, according to the embodiments of the present disclosure, the processes described below with reference to the flowcharts can be implemented as computer software programs. For example, the embodiments of the present disclosure include a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes program codes for performing the methods shown in the flowcharts. In such an embodiment, the computer program can be downloaded and installed from the network through the communication section 209, and / or installed from the removable medium 211. When the computer program is executed by the central processing unit (CPU) 201, various functions defined in the system of the present application are executed.
[0056] It should be noted that the computer-readable medium shown in this disclosure can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. A computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of a computer-readable storage medium can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In this disclosure, a computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. And in this disclosure, a computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium can also be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on a computer-readable medium can be transmitted using any appropriate medium, including but not limited to: wireless, wire, optical fiber, RF, etc., or any suitable combination of the above.
[0057] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram can represent a module, a program segment, or a part of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the block can occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks shown can actually be executed substantially in parallel, and they can sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram or flowchart, as well as the combination of blocks in the block diagram or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.
[0058] The units involved in the embodiments of the present disclosure can be implemented in software or in hardware, and the described units can also be provided in a processor. Among them, the names of these units do not constitute a limitation to the unit itself in some cases.
[0059] As another aspect, the present application also provides a computer-readable medium, which can be included in the electronic device described in the above embodiments; or can exist alone without being assembled into the electronic device. The above computer-readable medium carries one or more programs, and when the one or more programs are executed by an electronic device, the electronic device is caused to implement the methods described in the following embodiments.
[0060] In the scenario of classifying data using machine learning, in order to improve the classification effect, an ensemble learning method can be adopted to process the data to be classified. As an important implementation framework in ensemble learning, Stacking has broad application prospects.
[0061] Figure 3 FIG. shows a schematic diagram of a Stacking ensemble classification architecture. Stacking is a hierarchical model ensemble framework. Although Figure 3 only shows a two-layer structure including a base classifier layer (Base-level) and a secondary classifier layer (Meta-level), it should be understood that other classifier layers can also be included in the Stacking framework, and no special limitation is made in this exemplary embodiment.
[0062] Taking the two-layer structure shown in the figure as an example, the base classifier layer can be composed of multiple base classifiers. The data set passes through some base classifiers, and the results are output to the secondary classifier layer. The secondary classifier layer can further classify the data output by the base classifier layer to determine the classification result corresponding to the input data set.
[0063] Those skilled in the art can easily understand that, on the one hand, the flexible framework structure of Stacking enables it to not only integrate classifiers of the same type, but also integrate classifiers of different types, making it more suitable for actual industrial applications; on the other hand, Stacking can re-learn the results output by the classifier, which makes Stacking have more robust classification performance.
[0064] Although the data classification method based on the base classifier of the present disclosure can be applied to the Stacking architecture, it should be noted that the concept of determining the classifier evaluation index and selecting the classifier described below can be applied to other scenarios of classification prediction using multiple classifiers, and no special limitation is made in this exemplary embodiment.
[0065] It is easy for those skilled in the art to understand that Figure 3 The base classifier layer includes multiple base classifiers. However, in practical applications, only several of these base classifiers are often used. Thus, the selection of base classifiers greatly affects the performance of the ensemble classifier.
[0066] In order to determine the required base classifiers, evaluation metrics can be used to evaluate the base classifiers to be selected. The present disclosure can use a confusion matrix to assist in evaluating the performance of the classifier. As shown in Table 1, the confusion matrix gives the prediction results obtained by the classifier and the class information in the actual situation.
[0067] Table 1
[0068]
[0069] The solution described in the embodiments of the present disclosure can use the WAD (Weighted Accuracy and Diversity) metric to select base classifiers. The WAD metric is defined as the weighted harmonic mean between the accuracy metric (Acc) and the diversity metric (Div), and can be specifically expressed as Equation 1:
[0070]
[0071] Where α and β respectively represent the weights controlling the importance of diversity and accuracy.
[0072] The accuracy metric Acc can be determined using the confusion matrix. Specifically, the accuracy metric Acc can be determined using Equation 2:
[0073]
[0074] However, in practical applications, imbalanced datasets are often encountered. The imbalanced datasets mentioned here specifically refer to the situation where the data in the dataset varies greatly. For example, in the modeling samples of a credit scoring model, the ratio of users with good credit to users with bad credit is 1:99. For such imbalanced datasets, using the accuracy Acc to measure the classifier will have problems.
[0075] For example, referring to the confusion matrix in Table 1 above, the actual situation is as follows: among 10,000 data, 9,900 data are negative examples and only 100 are positive examples. Suppose the classifier model predicts all 10,000 data as negative examples. That is to say, the classifier model predicts all negative examples accurately and all positive examples wrongly. In this case, TP = FP = 0, TN = 9,900 / 10,000, FN = 100 / 10,000, and Acc = (TP + TN) / (TP + TN + FP + FN) = TN / (TN + FN) = TN = 9,900 / 10,000. That is, since the proportion of negative example samples is very high and the prediction of negative examples is accurate, the value of the Acc index is very large. However, the classifier model does not correctly predict any positive example. In practical applications, the prediction ability of positive examples is often more concerned. Therefore, the method of selecting a base classifier using the Acc index is not suitable for unbalanced classification datasets.
[0076] In addition, due to data deviation and / or inconsistent data dimensions, the problem of incomparable data is caused, which in turn leads to poor results in dataset classification. Among them, the data deviation mentioned here can refer to the appearance of abnormal data in the dataset, and these abnormal data may deviate from the range of normal cognition. For example, for a dataset of monthly consumption times, the monthly consumption times of most users range from 0 to 1,000 times, but a user's monthly consumption times is 100,000 times, then it can be considered that the data of this user has deviated. In some existing technologies, there is no good method to solve such problems.
[0077] In view of this, a new data classification method based on a base classifier is needed to solve the above problems.
[0078] Figure 4 The flowchart of the data classification method based on a base classifier according to an exemplary embodiment of the present disclosure is schematically shown. This data classification method can be applied to an ensemble classifier including multiple base classifiers. Refer to Figure 4 , the data classification method based on a base classifier may include the following steps:
[0079] S42. Determine multiple combinations of base classifiers, and calculate the accuracy index and diversity index of each combination of base classifiers; wherein, each combination of base classifiers is composed of a predetermined number of base classifiers.
[0080] In an exemplary embodiment of the present disclosure, the base classifier is a machine learning model for classification processing, and its types may include, but are not limited to, LR (Logistic Regression) model, RF (Random Forest) model, SVM (Support Vector Machine) model, LDA (Linear Discriminant Analysis) model, QDA (Quadratic Discriminant Analysis) model, DT (Decision Tree) model, GNB (Gaussian Naive Bayes) model, KNN (K-Nearest Neighbor classification) model, etc.
[0081] The base classifiers included in the ensemble classifier can be base classifiers of different types. For example, the base classifier 1 corresponding to classification algorithm 1 can be an LR model, and the base classifier 2 corresponding to classification algorithm 2 can be an SVM model.
[0082] In addition, it should be noted that even if the same type of base classifier is used, the classification results may not be the same. For example, for the same LR model, if the model parameters are different, the classification results may also vary greatly. Therefore, there can also be diversity among base classifiers of the same type.
[0083] The base classifier combination described in the present disclosure can be composed of base classifiers to be selected, and the number of base classifiers included in each base classifier combination is the same. Specifically, each base classifier combination can be composed of a predetermined number of base classifiers. If there are a total of M base classifiers in the ensemble classifier, then each base classifier combination can be composed of a predetermined number m of base classifiers, where M is usually a positive integer greater than 2, and m is a positive integer and 1 < m ≤ M. Usually, the predetermined number m can be set manually. However, a mapping table related to the classification task can also be constructed in advance. Using this mapping table, the predetermined number can be directly determined through the classification task. No special limitation is made in this exemplary embodiment.
[0084] In some other embodiments, the predetermined number may not be limited to a single value. That is to say, the number of base classifiers in each base classifier combination can be different, and these different numbers can all be set manually.
[0085] In addition, it is easy to understand that when the predetermined number m and the total number M of base classifiers are determined, the number of base classifier combinations is also determined. Specifically, the number of base classifier combinations
[0086] The server can determine multiple combinations of base classifiers. According to some embodiments of the present disclosure, the multiple combinations of base classifiers can be N combinations of base classifiers. That is to say, all possible combinations of base classifiers under the ensemble classifier are the objects processed in the following steps.
[0087] According to some other embodiments of the present disclosure, considering the reasonable allocation of system resources and improving the classification speed, the server can determine several combinations of base classifiers from the N combinations of base classifiers as the multiple combinations of base classifiers determined in step S42. Specifically, a random algorithm can be used to determine a certain number (for example, 5) of combinations of base classifiers from the N combinations of base classifiers.
[0088] After determining the multiple combinations of base classifiers, the server can calculate the accuracy index and diversity index of each combination of base classifiers.
[0089] Procedure for calculating the accuracy index of each combination of base classifiers:
[0090] In an exemplary embodiment of the present disclosure, taking the calculation of the accuracy index of a combination of base classifiers as an example, first, the server can calculate the accuracy index of the base classifiers included in the combination of base classifiers; next, the server can calculate the average value of the accuracy indexes of these base classifiers and use this average value as the accuracy index of the combination of base classifiers.
[0091] According to some embodiments of the present disclosure, the accuracy index of the base classifier can be the area under the receiver operating characteristic curve, denoted as AUC.
[0092] AUC is an index used to measure the performance of a classifier, and its main analysis tool is an ROC curve plotted on a two-dimensional plane. Refer to Figure 5 , the abscissa of this two-dimensional plane is the false positive rate (FPr), and the ordinate is the true positive rate (TPr). The calculation methods of the false positive rate and the true positive rate are shown in Formulas 3 and 4 respectively:
[0093]
[0094]
[0095] For a base classifier, a point pair including the false positive rate and the true positive rate can be determined by using the classification prediction result of the test sample in the test set and the preset threshold corresponding to the test sample, and this point pair is mapped into Figure 5A point on the plane shown. Among them, the preset threshold can be set artificially, and for different classification tasks, the preset threshold is also different. By adjusting the preset threshold, a curve passing through (0, 0) and (1, 1) can be determined, and this curve is the ROC curve. Generally, the ROC curve is above the line connecting (0, 0) and (1, 1). In this case, the area AUC under the ROC curve can be calculated, and the AUC is used as the accuracy index of the base classifier. In addition, according to the above description, the value of AUC usually ranges between 0.5 and 1.0. The larger the value of AUC, the better the accuracy of the base classifier.
[0096] After determining the accuracy indexes of all the base classifiers included in the base classifier combination, their average value can be calculated to determine the accuracy index of this base classifier combination.
[0097] Similarly, the server can calculate the accuracy indexes of all the base classifier combinations, and the specific process will not be elaborated here.
[0098] It should be noted that in the classification scenario, especially for the binary classification problem, in the case of sample one (the gap between positive and negative samples is small) and sample two (too many negative samples), the shape of the ROC curve changes little. That is to say, using AUC as the accuracy index can be well applied to both balanced datasets and unbalanced datasets.
[0099] In addition, the exemplary unbalanced dataset in the present disclosure can be understood as: a dataset in which the gap between positive and negative samples is greater than a preset balance threshold. For example, if the positive samples only account for 5% of the total samples, then the gap between positive and negative samples is 90%. If the preset balance threshold is configured to 20%, it can be clearly determined that this dataset is an unbalanced dataset.
[0100] In addition, Acc as shown in Formula 2 above can also be used to represent the accuracy index, and the present disclosure does not make special restrictions on this.
[0101] Procedure for calculating the diversity index of each combination of base classifiers:
[0102] In the exemplary embodiment of the present disclosure, still taking the calculation of the accuracy index of a base classifier combination as an example, first, the server can calculate the diversity indexes between each pair of base classifiers in the base classifier combination; next, the server can calculate the average value of the diversity indexes between each pair of base classifiers as the diversity index of this base classifier combination.
[0103] Regarding the process of calculating the diversity indexes between each pair of base classifiers, specifically, taking the first base classifier and the second base classifier as an example to illustrate the calculation of the diversity index. Refer to Figure 6, for the base classifier combination A, which may include base classifier 1, base classifier 2, …, base classifier m, where the first base classifier herein may be any base classifier in the base classifier combination A, and the first base classifier is denoted as m i , and the second base classifier may be any base classifier in the base classifier combination A that is different from the first base classifier, and the second base classifier is denoted as m j , where both i and j are positive integers less than or equal to m, and i is not equal to j. For example, take Figure 6 the base classifier 1 in as the first base classifier, and the base classifier 2 as the second base classifier.
[0104] First, the server may input the test samples into the first base classifier and the second base classifier respectively, and determine the classification results of the first base classifier and the classification results of the second base classifier respectively; next, the server may calculate the diversity index between the first base classifier and the second base classifier according to the classification results of the first base classifier and the classification results of the second base classifier.
[0105] For example, Table 2 exemplarily shows the classification results of the first base classifier m i and the second base classifier m j .
[0106] Table 2
[0107]
[0108] Among them, the sum of N 11 , N 10 , N 01 , N 00 is the number of samples.
[0109] The diversity index between any two base classifiers described in the exemplary embodiments of the present disclosure may include, but is not limited to, Q statistic, correlation coefficient, disagreement measure, double error measure, Kappa statistic, balanced disagreement measure, entropy, KW variable, difficulty measure, generalized diversity, consistent failure diversity, etc. Below, taking the Q statistic and the correlation coefficient as examples, the calculation of the diversity index between any two base classifiers will be described. However, those skilled in the art should understand that any one of the above other metrics or a combination of multiple metrics may also be used to represent the diversity index described in the present disclosure, and these should all fall within the concept of the present invention.
[0110] In an example where the diversity index is characterized by the Q statistic, the Q statistic Q i between the first base classifier m j and the second base classifier m i,j can be calculated using Equation 5:
[0111]
[0112] where, -1 ≤ Q i,j ≤ 1. It is easy to understand that if the classification results of the first base classifier m i and the second base classifier m j are the same, the Q statistic is positive; if the classification results of the first base classifier m i and the second base classifier m j are different, the Q statistic is negative. In addition, when the Q statistic is 0, the diversity between the first base classifier m i and the second base classifier m j is the strongest, indicating that the first base classifier m i and the second base classifier m j are not relevant.
[0113] In an example where the diversity index is characterized by the correlation coefficient, the correlation coefficient ρ i between the first base classifier m j and the second base classifier m i,j can be calculated using Equation 6:
[0114]
[0115] where, -1 ≤ ρ i,j ≤ 1, and the smaller the absolute value of ρ i,j , the stronger the diversity between the first base classifier m i and the second base classifier m j . When ρ i,j = 0, the diversity is the strongest, indicating that the first base classifier m i and the second base classifier m j are not relevant.
[0116] The above only describes the determination process of the diversity index between the first base classifier m i and the second base classifier m j . For any two base classifiers in the base classifier combination, the above process can be used to determine their diversity indices respectively. For example, if the base classifier combination includes a different base classifiers, it is easy to understand that a(a - 1) / 2 diversity indices can be determined. Subsequently, the average value of these a(a - 1) / 2 diversity indices can be calculated to obtain the diversity index of the base classifier combination.
[0117] Similarly, the server can calculate the diversity indices of all base classifier combinations, and the specific process will not be elaborated here.
[0118] S44. Sort the base classifier combinations according to the accuracy indices of the base classifier combinations, and determine the accuracy ranking values of the base classifier combinations.
[0119] In an exemplary embodiment of the present disclosure, as described above, the number of base classifier combinations is denoted as N. The accuracy metrics of the N base classifier combinations can be sorted from largest to smallest using the reverse sorting method, and the obtained sorting values are sequentially labeled as 1, 2, …, N. The AUC Rank can be used to represent the determined accuracy sorting value.
[0120] It is easy to understand that although the exemplary embodiment of the present disclosure is described by taking reverse sorting as an example, however, the present disclosure can also use other sorting algorithms to achieve the accuracy sorting. For example, bubble sort, selection sort, insertion sort, etc. No special limitation is made in this exemplary embodiment.
[0121] S46. Sort the base classifier combinations according to the diversity metrics of each base classifier combination, and determine the diversity sorting values of each base classifier combination.
[0122] In an exemplary embodiment of the present disclosure, as described above, the number of base classifier combinations is denoted as N. The diversity metrics of the N base classifier combinations can be sorted from largest to smallest using the reverse sorting method, and the obtained sorting values are sequentially labeled as 1, 2, …, N. The Div Rank can be used to represent the determined diversity sorting value.
[0123] It is easy to understand that although the exemplary embodiment of the present disclosure is described by taking reverse sorting as an example, however, the present disclosure can also use other sorting algorithms to achieve the diversity sorting. For example, bubble sort, selection sort, insertion sort, etc. No special limitation is made in this exemplary embodiment.
[0124] S48. Calculate the evaluation metrics of each base classifier combination according to the accuracy sorting value and the diversity sorting value of each base classifier combination, determine the base classifiers in the base classifier combination with the largest evaluation metric as the selected base classifiers, and use the selected base classifiers to classify the data to be classified.
[0125] In an exemplary embodiment of the present disclosure, still taking the calculation of the evaluation metric of a base classifier combination as an example, first, the first value can be determined using the number of base classifier combinations and the accuracy metric sorting value of the base classifier combination, and the second value can be determined using the number of base classifier combinations and the diversity sorting value of the base classifier combination; next, the weighted harmonic mean of the first value and the second value can be calculated as the evaluation metric of the base classifier combination.
[0126] For the process of determining the weights of the first value and the second value, according to some embodiments of the present disclosure, first, the weight of the first value can be configured as the first weight, and the weight of the second value can be configured as the second weight; wherein, both the first weight and the second weight are greater than or equal to 0, and the sum of the first weight and the second weight is 1; next, the weighted harmonic mean of the first value and the second value is calculated using the first weight and the second weight, such that the weighted harmonic mean is maximized, and the first weight and the second weight at this time are the weights after parameter tuning. Additionally, the maximum weighted harmonic mean can be used as the evaluation index for the combination of base classifiers.
[0127] Specifically, formula 7 can be referred to for calculating the evaluation index of the combination of base classifiers.
[0128]
[0129] Wherein, β is the above-mentioned first weight, and α is the above-mentioned second weight.
[0130] After determining the evaluation index of each combination of base classifiers through the exemplary formula 7 constructed, the base classifier in the combination of base classifiers with the maximum evaluation index can be determined as the base classifier selected by the method of the present disclosure.
[0131] According to some other embodiments of the present disclosure, after determining the evaluation index of each combination of base classifiers, each combination of base classifiers and the corresponding evaluation index can be fed back to the terminal device for the user to refer to, and finally the user selects the base classifier to be used.
[0132] In addition, after selecting the base classifier, the server can receive the data to be classified sent by the terminal device and input the data to be classified into the selected base classifier to determine the intermediate classification result. Refer to Figure 3 , and then the server can input the intermediate classification result into the secondary classifier layer of the ensemble classifier to determine the final classification result of the data to be classified.
[0133] Next, the data classification method based on base classifiers of the present disclosure will be described with reference to an example.
[0134] There are a total of 5 alternative base classifiers in the Stacking ensemble classifier, namely LDA, QDA, GNB, DT, and KNN. Three of these 5 base classifiers are to be selected as the base classifiers for actual application. Then, it is easy to understand that there are a total of 10 combinations of base classifiers Referring to Table 3, the specific situation of this example is listed.
[0135] Table 3
[0136]
[0137] For these 10 combinations of base classifiers in Table 3, first, the accuracy index and diversity index of each combination of base classifiers can be calculated using the method described in step S42, where the AUC can be used to characterize the accuracy index of the base classifier. Next, the accuracy ranking value and diversity ranking value of each combination of base classifiers can be determined using the methods described in steps S44 and S46. Subsequently, the evaluation index of each combination of base classifiers can be calculated using Formula 7, and the base classifiers in the combination of base classifiers with the largest evaluation index are selected as the base classifiers selected by the method of the present disclosure.
[0138] For example, when it is determined that the evaluation index of QDA+DT+KNN is the largest, QDA, DT, and KNN can be determined as the selected base classifiers for actual classification applications.
[0139] In addition, for scenarios that require high recall performance, the recall rate can be used to replace the above-mentioned AUC to achieve the effect of evaluating the combination of base classifiers using the recall rate.
[0140] Specifically, the recall rate can be calculated using Formula 8:
[0141]
[0142] Next, similarly, the recall rates of each combination of base classifiers can be ranked. Combining the content of Formula 7 and using the recall rate Recall to replace the AUC, the evaluation index based on the recall rate and diversity can be calculated, and the specific process will not be elaborated.
[0143] For example, in the scenario of identifying abnormal users (such as users with frequent logins, users with abnormal transactions, etc.), the base classifiers selected using the evaluation index based on the recall rate can better screen out abnormal users from the user group.
[0144] In addition, F1-measure can also be used to replace the AUC. Among them, F1-measure is also known as F-score, which is the weighted harmonic mean of precision and recall.
[0145] Precision is also known as the precision rate. The higher the precision, the better the prediction performance of the classifier for positive examples. Precision can be the ratio of true positive examples to all samples predicted as positive examples, and can be calculated using Formula 9:
[0146]
[0147] Precision and recall often restrict each other. That is to say, if the precision is high, the recall is often low; while when the recall is high, the precision is often low. The following Formula 10 can be used to determine the F1-measure:
[0148]
[0149] Next, similarly, the F1-measures of the combinations of the base classifiers can be sorted to obtain sorting values, and in combination with the content of Formula 7 and replacing AUC with F1-measure, an evaluation index based on precision, recall, and diversity can be calculated. The specific process will not be elaborated here.
[0150] According to the data classification method based on base classifiers of the exemplary embodiment of the present disclosure, on the one hand, by using the accuracy sorting value and the diversity sorting value to calculate the evaluation index, it avoids the problem of incorrect calculation of the evaluation index caused by data deviation and / or inconsistent data dimensions, and helps to improve the classification accuracy of the ensemble classifier; on the other hand, using the AUC index to determine the evaluation index can be applied to the classification scenario for imbalanced data sets; moreover, the present disclosure comprehensively considers the accuracy and diversity of the base classifiers, making the ensemble classifier more robust; on the further hand, the solution described in the present disclosure has universality and can quickly assist in building an ensemble classifier, improving the efficiency and performance of the ensemble classifier in classification problems.
[0151] It should be noted that although the steps of the method in the present disclosure are described in a specific order in the drawings, this does not require or imply that these steps must be executed in that specific order, or that all the steps shown must be executed to achieve the desired result. Additionally or alternatively, some steps may be omitted, multiple steps may be combined into one step for execution, and / or one step may be decomposed into multiple steps for execution, etc.
[0152] Furthermore, in the present exemplary embodiment, a data classification device based on base classifiers is also provided.
[0153] Figure 7 A block diagram of the data classification device based on base classifiers of the exemplary embodiment of the present disclosure is schematically shown. This data classification device is applied to an ensemble classifier including multiple base classifiers. Refer to Figure 7 , according to the data classification device 7 based on base classifiers of the exemplary embodiment of the present disclosure, it may include an index calculation module 71, a first sorting module 73, a second sorting module 75, and a data classification module 77.
[0154] Specifically, the metric calculation module 71 can be used to determine multiple combinations of base classifiers, and calculate the accuracy metric and diversity metric for each combination of base classifiers; wherein, each combination of base classifiers is composed of a predetermined number of base classifiers; the first sorting module 73 can be used to sort each combination of base classifiers according to the accuracy metric of each combination of base classifiers, and determine the accuracy sorting value of each combination of base classifiers; the second sorting module 75 can be used to sort each combination of base classifiers according to the diversity metric of each combination of base classifiers, and determine the diversity sorting value of each combination of base classifiers; the data classification module 77 can be used to calculate the evaluation metric for each combination of base classifiers according to the accuracy sorting value and diversity sorting value of each combination of base classifiers, determine the base classifiers in the combination of base classifiers with the largest evaluation metric as the selected base classifiers, and use the selected base classifiers to classify the data to be classified.
[0155] In the data classification device based on base classifiers provided by the present disclosure, on the one hand, the present disclosure calculates the evaluation metric using the accuracy sorting value and diversity sorting value, avoiding the problem of incorrect calculation of the evaluation metric due to data deviation and / or inconsistent data dimensions, which helps to improve the accuracy of the integrated classifier classification; on the other hand, the present disclosure comprehensively considers the accuracy and diversity of the base classifiers, making the integrated classifier more robust; on the other hand, the solution described in the present disclosure is universal, and can quickly assist in building an integrated classifier, improving the efficiency and performance of the integrated classifier in classification problems.
[0156] According to an exemplary embodiment of the present disclosure, with reference to Figure 8 , the metric calculation module 71 may include an accuracy metric determination unit 801.
[0157] Specifically, the accuracy metric determination unit 801 can be used to calculate the accuracy metric of the base classifiers included in the combination of base classifiers; calculate the average value of the accuracy metrics of the base classifiers included in the combination of base classifiers as the accuracy metric of the combination of base classifiers.
[0158] According to an exemplary embodiment of the present disclosure, the accuracy metric of the base classifier is the area under the receiver operating characteristic curve.
[0159] In this embodiment, by using the AUC metric to determine the evaluation metric, it can be applied to the classification scenario for imbalanced data sets
[0160] According to an exemplary embodiment of the present disclosure, the accuracy metric determination unit 801 can be configured to perform: using the base classifier to determine the classification results of each test sample in the test set, determining the receiver operating characteristic curve according to the classification results of each test sample and the preset threshold corresponding to each test sample, and calculating the area under the receiver operating characteristic curve as the accuracy metric of the base classifier.
[0161] According to an exemplary embodiment of the present disclosure, with reference to Figure 9 , the metric calculation module 71 further includes a diversity metric determination unit 901.
[0162] Specifically, the diversity metric determination unit 901 can be used to calculate the diversity metrics between each pair of base classifiers in the base classifier combination, and calculate the average value of the diversity metrics between each pair of base classifiers as the diversity metric of the base classifier combination.
[0163] According to an exemplary embodiment of the present disclosure, the base classifier combination includes a first base classifier and a second base classifier; wherein, the diversity metric determination unit 901 can be configured to perform: inputting a test sample into the first base classifier and the second base classifier respectively, determining the classification results of the first base classifier and the second base classifier respectively, and calculating the diversity metric between the first base classifier and the second base classifier according to the classification results of the first base classifier and the second base classifier.
[0164] According to an exemplary embodiment of the present disclosure, with reference to Figure 10 , the data classification module 77 may include an evaluation metric calculation unit 101.
[0165] Specifically, the evaluation metric calculation unit 101 can be configured to perform: determining the number of base classifier combinations; determining a first value by using the number of base classifier combinations and the accuracy ranking value of the base classifier combination, and determining a second value by using the number of base classifier combinations and the diversity ranking value of the base classifier combination; calculating the weighted harmonic mean of the first value and the second value as the evaluation metric of the base classifier combination.
[0166] According to an exemplary embodiment of the present disclosure, the evaluation metric calculation unit 101 can be configured to perform: determining the number of base classifiers in the ensemble classifier; determining the number of base classifier combinations based on the number of base classifiers in the ensemble classifier and a predetermined number.
[0167] According to an exemplary embodiment of the present disclosure, the evaluation metric calculation unit 101 can be configured to perform: configuring the weight of the first value as a first weight and configuring the weight of the second value as a second weight; wherein, both the first weight and the second weight are greater than or equal to 0, and the sum of the first weight and the second weight is 1; calculating the weighted harmonic mean of the first value and the second value by using the first weight and the second weight, such that the weighted harmonic mean is maximized, and determining the maximum weighted harmonic mean as the evaluation metric of the base classifier combination.
[0168] According to an exemplary embodiment of the present disclosure, with reference to Figure 11 , the data classification module 77 may include a classification result determination module 111.
[0169] Specifically, the classification result determination module 111 can be configured to perform: input the data to be classified into the selected base classifier to determine an intermediate classification result; input the intermediate classification result into the secondary classifier layer of the integrated classifier to determine the classification result of the data to be classified.
[0170] Since each functional module of the program running performance analysis device according to the embodiments of the present invention is the same as that in the method invention embodiments described above, it will not be described in detail herein.
[0171] Through the description of the above embodiments, those skilled in the art can easily understand that the exemplary embodiments described herein can be implemented by software, or by a combination of software and necessary hardware. Therefore, the technical solutions according to the embodiments of the present disclosure can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, including several instructions to enable a computing device (such as a personal computer, a server, a terminal device, or a network device, etc.) to execute the method according to the embodiments of the present disclosure.
[0172] In addition, the above drawings are only schematic illustrations of the processes included in the method according to the exemplary embodiments of the present invention, rather than for limiting purposes. It is easy to understand that the processes shown in the above drawings do not indicate or limit the time sequence of these processes. Additionally, it is also easy to understand that these processes can be executed synchronously or asynchronously in, for example, multiple modules.
[0173] It should be noted that although several modules or units of the devices for action execution are mentioned in the above detailed description, such a division is not mandatory. In fact, according to the embodiments of the present disclosure, the features and functions of the two or more modules or units described above can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided and embodied by multiple modules or units.
[0174] After considering the specification and practicing the invention disclosed herein, those skilled in the art will readily think of other embodiments of the present disclosure. This application is intended to cover any variations, uses, or adaptations of the present disclosure, which follow the general principles of the present disclosure and include the common general knowledge or conventional technical means in the technical field not disclosed in the present disclosure. The specification and embodiments are only regarded as exemplary, and the true scope and spirit of the present disclosure are pointed out by the claims.
[0175] It should be understood that the present disclosure is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present disclosure is only limited by the appended claims.
Claims
1. A data classification method based on base classifiers, which is applied to an ensemble classifier including multiple base classifiers, and is characterized in that, Including: Obtain the total number of base classifiers to be selected M , M is a positive integer greater than 2; Using A pre-constructed mapping table related to the classification task, determining a predetermined quantity through the classification task m , m is a positive integer and 1 < m ≤ M ; According to the total number M of the base classifiers and the predetermined quantity m determine the number of combinations of base classifiers, and the combination of base classifiers is composed of m base classifiers; Classify the modeling samples of the credit scoring model respectively according to N combinations of the base classifiers to obtain the prediction results of the modeling samples of the credit scoring model; The modeling samples of the credit scoring model include user samples with good credit and user samples with bad credit, and the quantity gap between the user samples with good credit and the user samples with bad credit is greater than a preset balance threshold; Calculating an accuracy index and a diversity index for each combination of the base classifiers according to a confusion matrix, where the confusion matrix includes the prediction results and the class information in the real situation; Sorting each combination of the base classifiers according to the accuracy index of each combination of the base classifiers, and determining the accuracy ranking value of each combination of the base classifiers; Sorting each combination of the base classifiers according to the diversity index of each combination of the base classifiers, and determining the diversity ranking value of each combination of the base classifiers; Determining the number of combinations of the base classifiers; Determining a first value using the number of combinations of the base classifiers and the accuracy ranking value of the combination of the base classifiers, and determining a second value using the number of combinations of the base classifiers and the diversity ranking value of the combination of the base classifiers; Configuring the weight of the first value as a first weight, and configuring the weight of the second value as a second weight; where both the first weight and the second weight are greater than or equal to 0, and the sum of the first weight and the second weight is 1; Calculating the weighted harmonic mean of the first value and the second value using the first weight and the second weight, so that the weighted harmonic mean is maximized, and determining the maximum weighted harmonic mean as the evaluation index of the combination of the base classifiers; the calculation formula of the weighted harmonic mean is as follows: Among them, WAD_Rank α,β ( AUC Rank , Div Rank ) is the weighted harmonic mean, N is the number of the base classifier combinations, AUC Rank is the accuracy ranking value, Div Rank is the diversity ranking value, ( N + 1 - AUC Rank ) is the first value, ( N + 1 - Div Rank ) is the second value, β is the first weight, α is the second weight; Determining the base classifiers in the combination of the base classifiers with the maximum evaluation index as the selected base classifiers, and classifying the data to be classified using the selected base classifiers to obtain the credit scoring result of the data to be classified.
2. The data classification method based on a base classifier according to claim 1, wherein Calculating the accuracy index of the combination of the base classifiers includes: Calculating the accuracy index of the base classifiers included in the combination of the base classifiers; Calculating the average value of the accuracy indexes of the base classifiers included in the combination of the base classifiers as the accuracy index of the combination of the base classifiers.
3. The data classification method based on a base classifier according to claim 2, characterized in that The accuracy index of the base classifier is the area under the receiver operating characteristic curve.
4. The data classification method based on a base classifier according to claim 3, wherein Calculating the accuracy index of the base classifier includes: Using the base classifier to determine the classification results of the modeling samples of each credit scoring model in the test set; Determining the receiver operating characteristic curve according to the classification results of the modeling samples of each credit scoring model and the preset threshold corresponding to each credit scoring model modeling sample; Calculating the area under the receiver operating characteristic curve as the accuracy index of the base classifier.
5. The data classification method based on a base classifier according to claim 1, characterized in that Calculating the diversity index of the combination of the base classifiers includes: Calculating the diversity indexes between each pair of the base classifiers in the combination of the base classifiers; Calculating the average value of the diversity indexes between each pair of the base classifiers as the diversity index of the combination of the base classifiers.
6. The data classification method based on a base classifier according to claim 5, wherein The combination of the base classifiers includes a first base classifier and a second base classifier; where calculating the diversity index between each pair of the base classifiers includes: Inputting the modeling samples of the credit scoring model into the first base classifier and the second base classifier respectively, and respectively determining the classification result of the first base classifier and the classification result of the second base classifier; Calculate the diversity index of the first base classifier and the second base classifier according to the classification results of the first base classifier and the classification results of the second base classifier.
7. The data classification method based on a base classifier according to claim 1, characterized in that, Using the selected base classifiers to classify the data to be classified includes: Input the data to be classified into the selected base classifier to determine an intermediate classification result; Input the intermediate classification result into the secondary classifier layer of the ensemble classifier to determine the classification result of the data to be classified.
8. A data classification device based on a base classifier, applied to an ensemble classifier including a plurality of base classifiers, characterized in that Includes: An index calculation module, configured to obtain the total number of base classifiers to be selected M , M is a positive integer greater than 2; using a pre-constructed mapping table related to the classification task, determining a predetermined number through the classification task m , m is a positive integer and 1 < m ≤ M ; according to the total number of the base classifiers M and the predetermined number m determining the number of combinations of base classifiers , the combination of base classifiers is composed of m base classifiers; classifying the modeling samples of the credit scoring model respectively according to N combinations of the base classifiers to obtain the prediction results of the modeling samples of the credit scoring model; the modeling samples of the credit scoring model include user samples with good credit and user samples with bad credit, and the quantity gap between the user samples with good credit and the user samples with bad credit is greater than a preset balance threshold; Calculate the accuracy index and diversity index of each combination of base classifiers according to the confusion matrix, where the confusion matrix includes the prediction results and the category information in the actual situation; A first sorting module, configured to sort each combination of base classifiers according to the accuracy index of each combination of base classifiers, and determine the accuracy sorting value of each combination of base classifiers; A second sorting module, configured to sort each combination of base classifiers according to the diversity index of each combination of base classifiers, and determine the diversity sorting value of each combination of base classifiers; A data classification module, configured to determine the number of combinations of base classifiers; determine a first value using the number of combinations of base classifiers and the accuracy sorting value of the combination of base classifiers, and determine a second value using the number of combinations of base classifiers and the diversity sorting value of the combination of base classifiers; configure the weight of the first value as the first weight, and configure the weight of the second value as the second weight; where the first weight and the second weight are both greater than or equal to 0, and the sum of the first weight and the second weight is 1; calculate the weighted harmonic mean of the first value and the second value using the first weight and the second weight, so that the weighted harmonic mean is the largest, and determine the largest weighted harmonic mean as the evaluation index of the combination of base classifiers; the calculation formula of the weighted harmonic mean is as follows: Among them, WAD_Rank α,β ( AUC Rank , Div Rank ) is the weighted harmonic mean, N is the number of the base classifier combinations, AUC Rank is the accuracy ranking value, Div Rank is the diversity ranking value,( N +1 - AUC Rank ) is the first value,( N +1 - Div Rank ) is the second value, β is the first weight, α is the second weight; Determine the base classifiers in the combination of base classifiers with the largest evaluation index as the selected base classifiers, and use the selected base classifiers to classify the data to be classified to obtain the credit score result of the data to be classified.
9. The data classification device based on a base classifier according to claim 8, characterized in that The index calculation module includes: A diversity index determination unit, configured to calculate the diversity index between each pair of base classifiers in the combination of base classifiers; calculate the average value of the diversity index between each pair of base classifiers as the diversity index of the combination of base classifiers.
10. The data classification device based on the base classifier according to claim 9, wherein The combination of base classifiers includes a first base classifier and a second base classifier; Wherein, the diversity index determination unit is configured to perform: input the credit scoring model modeling samples into the first base classifier and the second base classifier respectively, and determine the classification results of the first base classifier and the classification results of the second base classifier respectively; calculate the diversity index of the first base classifier and the second base classifier according to the classification results of the first base classifier and the classification results of the second base classifier.
11. A computer-readable medium having a computer program stored thereon, characterized in that, When the program is executed by a processor, it implements the data classification method based on base classifiers according to any one of claims 1 to 7.
12. An electronic device, characterized in that, Comprising: One or more processors; A storage device for storing one or more programs, which when executed by the one or more processors, cause the one or more processors to implement the data classification method based on a base classifier according to any one of claims 1 to 7.
13. A computer program product, characterized in that, A computer program carried on a computer-readable medium, which when executed by a processor, implements the data classification method based on a base classifier according to any one of claims 1 to 7.
Citation Information
Patent Citations
A classification and prediction method based on multi-stage hybrid model
CN109242021A