Model evaluation methods, apparatus, computer-readable storage media and electronic devices
By standardizing the model output values and using a unified evaluation method, the problem of low evaluation efficiency for a single model in existing technologies is solved, and efficient evaluation of multiple models is achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-11-25
- Publication Date
- 2026-03-13
AI Technical Summary
Existing technologies only support evaluation of individual models, resulting in low evaluation efficiency and an inability to perform unified model evaluation.
By standardizing the output values of each model and evaluating each model based on a unified evaluation method, the output values and true results of each model on the input data are obtained. Data adjustment rules are determined based on the model type, and model indicator values are calculated using preset indicators to evaluate the abnormality level of the model.
It enables unified evaluation of different models, improves model evaluation efficiency, avoids the need to develop different evaluation standards for different models, and improves evaluation efficiency.
Smart Images

Figure CN115936493B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of artificial intelligence, and more specifically, to a model evaluation method, apparatus, computer-readable storage medium, and electronic device. Background Technology
[0002] With the development of big data computing and storage architectures, the artificial intelligence (AI) technologies built upon them are becoming increasingly mature. Financial institutions are increasingly demanding models for related business operations, such as customer risk assessment and fraud detection. Model lifecycle management includes seven stages: requirement initiation, model development, model validation, model review, model release, model evaluation, and model optimization / exit. Model performance evaluation is one of the most critical stages. Traditional model evaluation systems are developed and processed for individual models based on big data architectures, which makes unified model evaluation impossible and results in low evaluation efficiency.
[0003] There is currently no effective solution to the above problems. Summary of the Invention
[0004] The present invention provides a model evaluation method, apparatus, computer-readable storage medium, and electronic device to at least solve the technical problem of low evaluation efficiency caused by the prior art only supporting the evaluation of a single model.
[0005] According to one aspect of the present invention, a model evaluation method is provided, comprising: acquiring the output value obtained by each model in at least one model after predicting input data, and the true result corresponding to the input data; determining a data adjustment rule corresponding to each model type based on the model type of each model; standardizing the output value corresponding to each model according to the data adjustment rule corresponding to each model type to obtain a standardized output value corresponding to each model, wherein the standardization process is used to map the output value corresponding to at least one model to a preset numerical range; processing the standardized output value and the true result corresponding to each model using a preset index determination rule to determine the index values of multiple model indices corresponding to each model; and evaluating the anomaly level of each model based on the index values of the multiple model indices corresponding to each model, wherein the anomaly level is used to determine whether the model is abnormal.
[0006] Furthermore, the model evaluation method also includes: combining different standardized output values of the target model with different true results of the target model to obtain multiple target combination terms, wherein the target model is any one of at least one model; determining the frequency of occurrence of the combined content corresponding to each target combination term of the target model based on the correspondence between the standardized output values and the true results of the target model; and using preset index determination rules to process the standardized output values, true results, and frequency of occurrence of the combined content corresponding to each target combination term of the target model to determine the index values of multiple model indices corresponding to the target model.
[0007] Furthermore, the model evaluation method also includes: determining a first evaluation result for the target model based on multiple indicator values corresponding to the target model and at least one target threshold, wherein the target threshold corresponds one-to-one with the model indicators; determining a second evaluation result for the target model based on multiple indicator values corresponding to the target model and the target prediction model; determining a third evaluation result for the target model based on multiple indicator values corresponding to the target model and the target time series prediction algorithm; and evaluating the anomaly level of the target model based on the first evaluation result, the second evaluation result, and the third evaluation result.
[0008] Furthermore, the model evaluation method also includes: determining the target threshold corresponding to each indicator value of the target model based on the correspondence between model indicators and target thresholds; comparing each indicator value corresponding to the target model with the target threshold corresponding to that indicator value to obtain the comparison result corresponding to each indicator value of the target model; and determining the first evaluation result corresponding to each model based on the comparison result corresponding to the target model.
[0009] Furthermore, the model evaluation method also includes: predicting the indicator change threshold of each model indicator corresponding to the target model based on the target indicator value of each model indicator corresponding to the target model and the target time series prediction algorithm, wherein the indicator change threshold represents the range of change allowed for the indicator value of the model indicator, each model indicator corresponding to the target model corresponds to multiple indicator values, and the multiple indicator values corresponding to the same model indicator are determined at different times, and the target indicator value is the indicator value determined before a preset time; and determining the third evaluation result corresponding to the target model based on the indicator change threshold of each model indicator corresponding to the target model and the multiple indicator values corresponding to the target model.
[0010] Furthermore, the model evaluation method also includes: when a target evaluation result exists among the first evaluation result, second evaluation result, and third evaluation result corresponding to the target model, the anomaly level of the target model is determined as the first anomaly level, wherein the target evaluation result is an evaluation result that represents the anomaly probability of the target model being greater than a preset probability, and the first anomaly level represents the model as an anomaly model; when a target evaluation result does not exist among the first evaluation result, second evaluation result, and third evaluation result corresponding to the target model, the anomaly level of the target model is determined as the second anomaly level, wherein the second anomaly level represents the model as a non-anomaly model.
[0011] Furthermore, the model evaluation method also includes: after evaluating the anomaly level of each model based on the index values of multiple model indicators corresponding to each model, determining whether there is an anomalous model in at least one model based on the anomaly level; if there is an anomalous model in at least one model, then performing stability tests on multiple input variables of the anomalous model to obtain test results; and determining the anomalous input variables based on the test results.
[0012] According to another aspect of the present invention, a model evaluation apparatus is also provided, comprising: an acquisition module, configured to acquire the output value obtained by each model in at least one model after predicting input data, and the actual result corresponding to the input data, and determine a data adjustment rule corresponding to each model type based on the model type of each model; a first processing module, configured to perform standardization processing on the output value corresponding to each model according to the data adjustment rule corresponding to each model type to obtain a standardized output value corresponding to each model, wherein the standardization processing is used to map the output value corresponding to at least one model to a preset numerical range; a second processing module, configured to process the standardized output value and the actual result corresponding to each model using a preset index determination rule to determine the index values of multiple model indices corresponding to each model; and an evaluation module, configured to evaluate the anomaly level of each model based on the index values of the multiple model indices corresponding to each model, wherein the anomaly level is used to determine whether the model is abnormal.
[0013] According to another aspect of the present invention, a computer-readable storage medium is also provided, wherein a computer program is stored in the computer-readable storage medium, and the computer program is configured to execute the above-described model evaluation method at runtime.
[0014] According to another aspect of the present invention, an electronic device is also provided, the electronic device including one or more processors; a memory for storing one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors are configured to run the programs, wherein the programs are configured to execute the model evaluation method described above during runtime.
[0015] In this embodiment of the invention, the output values of each model are standardized, and then each model is evaluated based on a unified evaluation method. This involves obtaining the output value of each model after predicting the input data, along with the corresponding true results. Based on the model type, a data adjustment rule is determined for each model type. Then, according to the data adjustment rule, the output value of each model is standardized to obtain a standardized output value. Next, a preset index determination rule is used to process the standardized output value and the true results for each model, determining the index values of multiple model indices for each model. Based on these index values, the anomaly level of each model is evaluated. The standardization process maps the output value of at least one model to a preset numerical range, and the anomaly level is used to determine whether the model is abnormal.
[0016] In the above process, by standardizing the output values corresponding to each model, it is possible to achieve output values with the same numerical value representing the same actual meaning for different models, thereby facilitating the input requirements of the same indicator determination rule. Furthermore, by calculating the indicator for each model based on a unified indicator determination rule, and evaluating whether the model is abnormal based on the indicator calculation results, the efficiency of model evaluation is improved. This avoids the problem of low evaluation efficiency in related technologies that require different evaluation standards for different models and need to process data from different models based on different evaluation standards.
[0017] Therefore, the solution provided in this application achieves the goal of standardizing the output values of each model and then evaluating each model based on a unified evaluation method, thereby improving the technical effect of model evaluation efficiency and solving the technical problem of low evaluation efficiency caused by only supporting the evaluation of a single model in the prior art. Attached Figure Description
[0018] The accompanying drawings, which are included to provide a further understanding of the invention and form part of this application, illustrate exemplary embodiments of the invention and, together with their description, serve to explain the invention and do not constitute an undue limitation thereof. In the drawings:
[0019] Figure 1 This is a schematic diagram of an optional model evaluation method according to an embodiment of the present invention;
[0020] Figure 2 This is a schematic diagram of the relationship of at least one optional model according to an embodiment of the present invention;
[0021] Figure 3 This is a flowchart of an optional model evaluation method according to an embodiment of the present invention;
[0022] Figure 4 This is a schematic diagram of an optional model evaluation device according to an embodiment of the present invention;
[0023] Figure 5 This is a schematic diagram of an optional electronic device according to an embodiment of the present invention. Detailed Implementation
[0024] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.
[0025] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0026] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for display, data used for analysis, etc.) involved in this disclosure are all information and data authorized by the user or fully authorized by all parties.
[0027] Example 1
[0028] According to an embodiment of the present invention, an embodiment of a model evaluation method is provided. It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Furthermore, although a logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than that shown here.
[0029] Figure 1 This is a schematic diagram of an optional model evaluation method according to an embodiment of the present invention, such as... Figure 1 As shown, the method includes the following steps:
[0030] Step S101: Obtain the output value of each model in at least one model after predicting the input data, as well as the true result corresponding to the input data, and determine the data adjustment rules corresponding to each model type based on the model type of each model.
[0031] In step S101, the aforementioned output values and actual results can be obtained through electronic devices, application systems, servers, etc. In this application, the aforementioned output values and actual results are obtained through a model evaluation system. Specifically, at least one of the aforementioned models is a machine learning model, which can be a credit card risk assessment model, a user pre-loan risk assessment model, etc. Figure 2 As shown, each model in at least one model can come from a different model development platform, for example, Figure 2 Models m1 and m2 come from model development platform S1, models m3 and m4 come from model development platform S2, and models m5 and m6 come from model development platform S3. Furthermore, different models can be used to process different data, and the input data for different models may be the same or completely different.
[0032] The aforementioned output values represent the direct result of the model output and are predictions, typically expressed as probabilities or scores. They indirectly represent the distinction or degree of the input data type. For example, when the model assesses user risk, a value of 1 indicates the user is at risk, and 0 indicates the user is not at risk. The aforementioned true results directly represent the distinction of the input data type and are the actual results corresponding to the input data. For instance, if a user applies for a loan at a target financial institution with a one-month repayment period, a model can determine whether the user has a risk of loan default when applying for the loan, obtaining an output value of 0.8. Based on this output value, it can be determined that the user is likely not at risk. Further, one month later, the target financial institution can determine whether the user has actually defaulted, thus determining whether the user has a loan default risk based on the default result. This determined result is the aforementioned true result, which can be either "the user is at risk" or "the user is not at risk."
[0033] Optionally, before the model evaluation system obtains the aforementioned output values and true results, such as Figure 2As shown, relevant data for each model can be uniformly written to a data lake or data warehouse. Real-time data related to the models can be written to the data warehouse in real-time or near real-time using Flink and Spark Streaming, while offline data can be written to the unified data warehouse in batches at scheduled times using Hive or Spark. This ensures that the computational data for the model evaluation system originates from a unified data address, reducing the complexity and difficulty of model evaluation access. The aforementioned relevant data includes, but is not limited to, the model's input data, output values, and actual results. Flink is a framework and distributed processing engine used for stateful computation on unbounded and bounded data streams; Spark Streaming is a real-time stream computing framework built on Spark for real-time stream data processing; Spark is a general-purpose big data analytics engine; and Hive is a data warehouse tool based on Hadoop.
[0034] Furthermore, such as Figure 3 As shown, the model evaluation system can obtain relevant data for each model from a data lake or data warehouse, and determine the corresponding data adjustment rules for each model based on its model type. The aforementioned data adjustment rules correspond one-to-one with the model type. Optionally, the model evaluation system can determine the model type based on the model's characteristics, operating mechanism, usage method, and operating log files, and determine the corresponding data adjustment rules for each model based on a preset correspondence between model types and data adjustment rules. Alternatively, the model evaluation system can also determine the model type for each model based on a predefined correspondence between models and model types defined by relevant personnel, and thus determine the corresponding data adjustment rules for each model based on the preset correspondence between model types and data adjustment rules.
[0035] Step S102: According to the data adjustment rules corresponding to each model type, the output value corresponding to each model is standardized to obtain the standardized output value corresponding to each model. The standardization process is used to map the output value corresponding to at least one model to a preset numerical range.
[0036] Optionally, different models may output different values. For example, model A outputs data of probability type with values between 0 and 1, model B outputs data of rating type with values between 0 and 100, and model C outputs data of rating type with values between 0 and 10.
[0037] Therefore, as Figure 3As shown, in step S102, the model evaluation system can standardize the output value of a model according to the data adjustment rules corresponding to the model type, thereby obtaining the standardized output value corresponding to the model. The goal of each data adjustment rule is to map the output value of the model to a preset numerical range. For example, the output values of models A, B, and C are mapped to a score between 0 and 20. If the output value of model B is 5, even though it is already within the 0-20 range, because there is a relative concept between each output value and its range, the output value of model B still needs to be processed according to the data adjustment rules corresponding to model B. It should be noted that although the aforementioned actual results may not be the same, they all represent results of the "yes / no" category, which only have two result types. Therefore, standardization of the actual results is not necessary.
[0038] Furthermore, such as Figure 3 As shown, after determining the standardized output value corresponding to each model, the standardized output value, actual results and other data corresponding to each model can be compressed and summarized into the same target intermediate table T, so that different monitoring and evaluation components in the model evaluation system can perform real-time online calculations and flexible displays based on this intermediate table T.
[0039] It should be noted that by standardizing the output values corresponding to each model, the same numerical value is used to represent the same meaning for different models, so as to meet the input requirements of the indicator determination rules and facilitate the calculation of indicators for each model based on unified indicator determination rules.
[0040] Step S103: Using preset index determination rules, process the standardized output value and real results corresponding to each model to determine the index values of multiple model indicators corresponding to each model.
[0041] In step S103, the model evaluation system can use the same index determination rules to process the standardized output values and true values corresponding to each model, thereby determining the index values of multiple model indicators for each model. The index determination rules can include sub-rules, and different sub-rules can be used to determine the index values of different model indicators. Indicator determination rules can be used to calculate model performance before determining the true results of the model, such as model call volume, concentration, maximum, minimum, average, and percentile of model results. They can also be used to calculate model stability, for example, by comparing current model data with model development data, initial model data, and model data to calculate population stability index (PSI) and feature stability index (CSI). Optionally, indicator determination rules can also be used after determining the true results of the model to evaluate whether the model's effectiveness and performance have deteriorated. For example, they can assess model improvement, good-to-bad ratio, bad sample rate, score correction, Kolmogorov-Smirnov (KS) value, AUC (area under the curve), and Gini coefficient. Furthermore, they can calculate MSE (mean squared error) and R² using regression models. 2 Indicators such as MAE (Mean Absolute Error).
[0042] Optionally, the model evaluation system can calculate the values of multiple model indicators for each model only once, or it can periodically calculate the values of multiple model indicators for each model. Since the input data processed by each model will be updated over time, and the performance of each model may also change, the value of the same model indicator for the same model may differ at different points in time when the model indicator values are calculated periodically.
[0043] It should be noted that by using unified indicator determination rules to calculate the indicators for each model, the problem of low evaluation efficiency in related technologies, which require different evaluation standards for different models and to process data from different models based on different evaluation standards, is avoided.
[0044] Step S104: Based on the index values of multiple model indicators corresponding to each model, evaluate the anomaly level of each model, where the anomaly level is used to determine whether the model is abnormal.
[0045] In step S104, the model evaluation system can evaluate the index values of multiple model indicators corresponding to each model according to preset evaluation rules, so as to evaluate the anomaly level of each model. The preset evaluation rules can be one or more.
[0046] It should be noted that by evaluating the values of multiple model metrics corresponding to each model, an accurate assessment of the degree of model anomaly is achieved.
[0047] Based on the scheme defined in steps S101 to S104 above, it can be understood that in this embodiment of the invention, the output values of each model are standardized, and then each model is evaluated based on a unified evaluation method. This involves obtaining the output value of each model after predicting the input data, as well as the corresponding true result of the input data, and determining the data adjustment rules for each model type. Then, according to the data adjustment rules for each model type, the output values of each model are standardized to obtain the standardized output value for each model. Next, a preset index determination rule is used to process the standardized output value and the true result for each model to determine the index values of multiple model indices for each model. Based on the index values of these multiple model indices, the anomaly level of each model is evaluated. Specifically, the standardization process maps the output values of at least one model to a preset numerical range, and the anomaly level is used to determine whether the model is abnormal.
[0048] It is noteworthy that, in the above process, by standardizing the output values corresponding to each model, the same numerical value is used to represent the same practical meaning for different models, thus facilitating the input requirements of the same indicator determination rule. Furthermore, by calculating the indicator for each model based on a unified indicator determination rule, and evaluating whether the model is abnormal based on the indicator calculation results, the efficiency of model evaluation is improved. This avoids the low evaluation efficiency inherent in related technologies that require different evaluation standards for different models and the processing of data from different models based on these standards.
[0049] Therefore, the solution provided in this application achieves the goal of standardizing the output values of each model and then evaluating each model based on a unified evaluation method, thereby improving the technical effect of model evaluation efficiency and solving the technical problem of low evaluation efficiency caused by only supporting the evaluation of a single model in the prior art.
[0050] In an optional embodiment, since the model's output value for the input data can be obtained in real time, but the actual result corresponding to the input data is delayed, in the process of the model evaluation system acquiring the output value obtained by each model in at least one model after predicting the input data, and the actual result corresponding to the input data, the model evaluation system can first acquire the output value obtained by each model in at least one model after predicting the input data. Simultaneously with acquiring the output value, it can acquire the input data processed by each model and the prediction object corresponding to the input data. Here, input data refers to the data used to input into the model, and the prediction object refers to the object to which the data input into the model belongs. For example, when the model is used to assess whether a user has risk, the input data can be user A's asset information, consumption information, etc., and the prediction object is user A. Similarly, when the model is used to assess the creditworthiness of a credit card, the input data can be the usage information of credit card B, and the prediction object is credit card B. Furthermore, the model evaluation system can fill the target data table with the detailed information such as the acquired input data, output value, and prediction object corresponding to each model.
[0051] Furthermore, when the model evaluation system populates the target data table with the aforementioned information but has not yet obtained the actual results, it can first standardize the output values of each model according to the data adjustment rules corresponding to each model type, thus obtaining standardized output values for each model. During this process, the model evaluation system can first template the identifiers of the prediction objects corresponding to different models; in simpler terms, it unifies the identifiers of the prediction objects corresponding to different models into the same column in the target data table, for example, into a column named Target_Key. Then, the model evaluation system can standardize the output values corresponding to each model and unify the standardized output values of each model into the same column in the target data table, for example, into a column named Score.
[0052] Furthermore, after obtaining the standardized output values for each model, the model evaluation system can acquire the actual results corresponding to the input data of each model in real time, and populate the aforementioned target data table with these actual results. In addition, the model evaluation system can also populate the target data table with the total business volume and business call volume for each model. The total business volume includes the data that needs to be analyzed from the business corresponding to the current model; the business corresponding to the model is the business to which the data to be processed by the current model belongs; and the business call volume includes the data processed by the current model from the aforementioned total business volume that needs to be analyzed.
[0053] It should be noted that the model evaluation system can also, after obtaining the actual results corresponding to the input data of each model, fill the target data table with detailed information such as the input data, output value, prediction object, and actual result of each model, and then calculate the aforementioned standardized output value.
[0054] In one optional embodiment, after determining the target data table, the model evaluation system processes the standardized output values and actual results corresponding to each model using preset indicator determination rules to determine the indicator values of multiple model indicators for each model. In this process, the system can combine different standardized output values and actual results corresponding to the target model to obtain multiple target combination terms. Then, based on the correspondence between the standardized output values and actual results, the system determines the frequency of occurrence of the combined content corresponding to each target combination term for the target model. Thus, using preset indicator determination rules, the system processes the standardized output values, actual results, and frequency of occurrence of the combined content corresponding to each target combination term for the target model to determine the indicator values of multiple model indicators for the target model. The target model can be any one of at least one model.
[0055] Optionally, if the target standardized output value corresponding to the target model includes 10 points, 20 points, 30 points, 40 points, and 50 points, and the corresponding real results are "the user has risk" and "the user does not have risk", then the model evaluation system can combine them to obtain target combination entries such as "10 points - the user has risk", "10 points - the user does not have risk", "20 points - the user has risk", and "20 points - the user does not have risk".
[0056] Furthermore, the model evaluation results can be based on the target standardized output value and the target true result corresponding to each input data in the input data processed by the target model, to determine the frequency of occurrence of the combined content corresponding to each target combination term of the target model. For example, if a target model A outputs 200 scores of 10 after processing all the input data, and among these 200 scores, 170 scores correspond to the target true result of "the user is not at risk," and 30 scores correspond to the target true result of "the user is at risk," then the frequency of occurrence of the combined content corresponding to the target combination term "10 points - the user is at risk" is 30, and the frequency of occurrence of the combined content corresponding to the target combination term "10 points - the user is not at risk" is 170.
[0057] Furthermore, the model evaluation results can be used as the primary key to generate a target intermediate table T, with the standardized output value corresponding to each model, the occurrence frequency of each target combination term, and the actual result as the main elements. Optionally, the design structure of the target intermediate table T can be {Model_id, Score, Weight, Target}, where Model_id represents a unique model, Score represents the standardized output value, Target represents the actual result, and Weight represents the occurrence frequency of the target combination term generated by the corresponding Score and Target combination. For example, for the aforementioned target model A, it can have data such as {Target Model A, 10, 30, User has risk} and {Target Model A, 10, 170, User does not have risk} stored in the target intermediate table.
[0058] Optionally, before generating the target intermediate table, the model evaluation system can also filter the data to be recorded in the target intermediate table based on the total business volume and business call volume defined in the target data table corresponding to the model. For example, if user A is not recorded in the total business volume and business call volume, but the prediction target corresponding to the input data actually processed by the model includes user A, then when counting the occurrence times of each target combination term in the target intermediate table, the data corresponding to user A will not be counted.
[0059] Optionally, after determining the target intermediate table, the model evaluation system can store the target intermediate table in the Hadoop Distributed File System (HDFS). Preferably, such as... Figure 3 As shown, the model evaluation system can synchronize the target intermediate table to a regular database to overcome the resource and time consumption problems of traditional calculation logic, which uses mainstream offline computing engines such as Hadoop or Spark to batch calculate model index values. By synchronizing the aforementioned target intermediate table to MySQL or other memory-based high-speed cache databases, the latency of program data retrieval and calculation can be reduced.
[0060] Furthermore, such as Figure 3 As shown, the model evaluation system can determine the values of multiple model indicators corresponding to the target model based on the data corresponding to the target model obtained in the aforementioned intermediate target table, using preset indicator determination rules. These model indicators can include model call volume, concentration, maximum, minimum, average, and percentile values of model results, as well as population stability indicators such as PSI and feature stability CSI. Optionally, model indicators can also include model improvement, good-to-bad ratio, bad sample rate, score correction, KS value, AUC, Gini coefficient, MSE, and R². 2Indicators such as MAE are used. The methods for calculating these indicators can be based on methods already disclosed in related technologies. It is important to emphasize that although related technologies utilize disclosed methods to calculate indicators based on detailed data corresponding to the target model, the model evaluation system does not directly use the detailed data for indicator calculation. Instead, it first statistically analyzes the detailed data to obtain information such as the number of different output values for each model and the number of true results corresponding to each different output value. Then, it calculates the indicators based on the statistical results (equivalent to the frequency of occurrence of each target combination term in this application). Therefore, although detailed data is no longer used for calculation in this application, the indicators can still be calculated based on disclosed methods. Furthermore, when calculating evaluation indicators, a high-performance key-value database, Redis, can be used between the database where the target intermediate table is stored and the model evaluation system to cache frequently queried data, improve data collision efficiency, and accelerate calculation speed.
[0061] For example, one possible method for calculating the KS index is as follows:
[0062] 1. Input the following into the KS index calculation model: number of segments (Bins), calculation method (equal frequency, equal interval).
[0063] If it is an equal-frequency calculation:
[0064] (1) It is necessary to first calculate the total number of data sets S based on the occurrence frequency of each target combination term;
[0065] (2) Based on the number of segments and the total number S, the total number of data in each segment Seg_Cnt = total number S / number of segments Bins;
[0066] (3) If it is not divisible, record the remainder as M, and set the number of segments at the end as Seg_Cnt+M, or add 1 to the data of the first M segments respectively, as the final total number of segments.
[0067] If it is an equal interval calculation:
[0068] (1) First, find the maximum and minimum values of the model score (result);
[0069] (2) Divide the spacing of each segment according to the number of segments (Bins).
[0070] 2. Sort each data entry.
[0071] For example, first sort the corresponding standardized output values from smallest to largest. Then, to ensure consistency and stability of the results in each calculation, sort them again from smallest to largest (e.g., from 0 to 1) according to the actual results, ensuring the accuracy of subsequent loop structure calculations.
[0072] 3. Use loop statements to calculate indicators such as the cumulative number of good and bad accounts and the percentage of good and bad accounts within each segment.
[0073] 4. Use formulas to calculate the values of monitoring indicators, such as KS and PSI.
[0074] It should be noted that by determining the frequency of occurrence of the combined content corresponding to each target term in the target model, the data required for model indicator calculation is extracted. Then, using pre-defined indicator determination rules, the standardized output values, actual results, and frequency of occurrence of the combined content corresponding to each target term in the target model are processed to determine the indicator values of multiple model indicators for the target model. This avoids the need to store various detailed data (such as input data) for the model, as required in related technologies. Combining this detailed data with the model indicator calculations achieves effective compression of the data used for calculating indicator values. For example, taking the credit card risk model as an example, nearly 100 million detailed data entries can be compressed into less than 2,000 data entries, achieving a compression efficiency of nearly 99.9%, thus greatly saving data storage resources. Furthermore, by compressing the target model data, more calculation methods can be applied during the calculation of indicator values for each model. This avoids the problems of traditional model evaluation requiring offline batch calculations using Hadoop / Spark big data engines, which are computationally expensive and time-consuming. This further reduces system resource consumption and improves the efficiency of model evaluation.
[0075] In one optional embodiment, during the process of evaluating the anomaly level of each model based on the index values of multiple model indicators corresponding to each model, the model evaluation system can process the index values of multiple model indicators corresponding to each model based on at least three evaluation components to evaluate the anomaly level of each model. Different evaluation components are used to implement different evaluation methods, and different evaluation results characterize the anomaly probability corresponding to the target model under different evaluation methods.
[0076] Optionally, in this application, the model evaluation system can determine a first evaluation result for the target model based on multiple indicator values corresponding to the target model and at least one target threshold. Specifically, in determining the first evaluation result, the model evaluation system can determine the target threshold corresponding to each indicator value of the target model based on the correspondence between model indicators and target thresholds, and then compare each indicator value of the target model with the target threshold corresponding to that indicator value to obtain the comparison result corresponding to each indicator value of the target model. Based on the comparison result corresponding to the target model, the first evaluation result for each model is determined.
[0077] In this system, there is a one-to-one correspondence between model indicators and target thresholds. These indicators can be preset by relevant personnel or dynamically determined by the model evaluation indicators. For example, the target threshold is the sum of the indicator value of the model indicator in the previous period and the preset value. Specifically, after determining the correspondence between model indicators and target thresholds, the model evaluation system can determine the target threshold corresponding to each indicator value of the target model based on the correspondence between the model indicators corresponding to each target model and the indicator values corresponding to the target model. Then, it compares each indicator value of the target model with the target threshold corresponding to that indicator value to obtain the comparison result corresponding to each indicator value of the target model. The comparison result can be that indicator value 1 corresponding to model indicator A is greater than the target threshold, indicator value 2 corresponding to model indicator B is less than the target threshold, and so on.
[0078] Furthermore, the model evaluation system can determine the first evaluation result of a model corresponding to a target model if, among multiple comparison results, there is a comparison result with a characterization index value higher than the target threshold, the probability of anomaly corresponding to the target model is greater than a preset probability. Conversely, if none of the comparison results correspond to the target model have a characterization index value higher than the target threshold, the probability of anomaly corresponding to the first evaluation result of the model corresponding to the first evaluation result of ... second evaluation result of the first evaluation result of the first evaluation result of the second evaluation result of the first evaluation result of the first evaluation result of the first evaluation result of the second evaluation result of the first evaluation result of the first evaluation result of the first evaluation result of the second evaluation result of the first evaluation result of the first evaluation result of the first evaluation result of the first evaluation result of the second evaluation result of the first evaluation result of the first evaluation result of the first evaluation result of the first evaluation result of the first evaluation result of the first evaluation result of the second evaluation result of the first evaluation result of the first evaluation result of the first evaluation result of the first evaluation result of the first evaluation result of the first evaluation result of the first evaluation result of the first evaluation result of the second evaluation result of the first evaluation result of the first evaluation result of the first evaluation result of the first evaluation result of the first evaluation result of the first evaluation result of the first evaluation result of the first evaluation result of the first evaluation result of the first evaluation result of the first evaluation result of the first evaluation result of the first evaluation result of the first evaluation result of the first evaluation result of
[0079] It should be noted that by using the threshold determination method, an effective assessment of the model's anomaly probability was achieved.
[0080] Optionally, in this application, the model evaluation system can also determine a second evaluation result corresponding to the target model based on multiple indicator values corresponding to the target model and the target prediction model.
[0081] Optionally, the aforementioned target prediction model can be an isolated forest model using an anomaly detection algorithm. The target prediction model can predict and score multiple index values corresponding to each model based on the isolated forest anomaly detection algorithm to obtain an anomaly score corresponding to each model. Then, when the anomaly score is greater than a preset anomaly score, the second evaluation result corresponding to the model is determined to be that the anomaly probability corresponding to the target model is greater than the preset probability. When the anomaly score is less than or equal to the preset anomaly score, the second evaluation result corresponding to the model is determined to be that the anomaly probability corresponding to the target model is less than or equal to the preset probability.
[0082] Optionally, in this application, the model evaluation system can also determine a third evaluation result for the target model based on multiple indicator values corresponding to the target model and the target time series prediction algorithm. Specifically, in determining the third evaluation result, the model evaluation system can predict the indicator change threshold for each model indicator corresponding to the target model based on the target indicator value of each model indicator corresponding to the target model and the target time series prediction algorithm. Then, based on the indicator change threshold for each model indicator corresponding to the target model and the multiple indicator values corresponding to the target model, the third evaluation result for the target model is determined. Here, the indicator change threshold represents the range of allowable changes in the indicator value of the model indicator. Each model indicator corresponding to the target model corresponds to multiple indicator values, and the multiple indicator values corresponding to the same model indicator are determined at different times. The target indicator value is the indicator value determined before a preset time.
[0083] Specifically, the performance of a model declines over time. Therefore, the model evaluation system can periodically calculate the values of multiple model indicators for each model, obtaining the indicator values of the same model indicator at different time points. The combination of the indicator values of a particular model indicator of the target model at each time point constitutes a time series. Thus, the model evaluation system can filter the indicator values of a specific model indicator corresponding to the target model from the target intermediate table before a preset time, and then use the FBPROPHET time series prediction algorithm to process the data before the preset time, thereby predicting the upper and lower bounds of the indicator value. The range between the upper and lower bounds can be directly used as the aforementioned indicator variation threshold, or the calculation result can be used as the aforementioned indicator variation threshold after certain mathematical calculations. For example, a range exceeding the upper and lower bounds by less than 15% can be used as the aforementioned indicator variation threshold.
[0084] Furthermore, once the threshold for the change of the model indicator is determined, the model evaluation system can compare the indicator value of a certain model indicator corresponding to the target model after a preset time with the threshold for the change of the indicator. Thus, if there is an indicator value that exceeds the threshold for the change of the indicator, the third evaluation result corresponding to the model is determined to be that the probability of anomaly corresponding to the target model is greater than the preset probability. If there is no indicator value that exceeds the threshold for the change of the indicator, the third evaluation result corresponding to the model is determined to be that the probability of anomaly corresponding to the target model is less than or equal to the preset probability.
[0085] It should be noted that, as the input data processed by each model will be updated over time, and the performance of each model may also change, using time series forecasting to evaluate the model can achieve better evaluation results.
[0086] Optional, such as Figure 3 As shown, after determining the first, second, and third evaluation results, the model evaluation system can assess the anomaly level of the target model based on these results. The target threshold corresponds one-to-one with the model indicators.
[0087] Specifically, the model evaluation system can determine the anomaly level of the target model as the first anomaly level when the target evaluation result exists among the first, second, and third evaluation results corresponding to the target model; and determine the anomaly level of the target model as the second anomaly level when the target evaluation result does not exist among the first, second, and third evaluation results corresponding to the target model. Here, the target evaluation result is the evaluation result indicating that the anomaly probability of the target model is greater than a preset probability; the first anomaly level indicates that the model is an anomalous model; and the second anomaly level indicates that the model is a non-anomaly model.
[0088] In layman's terms, if any one of the first, second, or third evaluation results indicates that the target model is highly likely to be abnormal, the target model is determined to be an abnormal model that needs to be investigated. If all three evaluation results indicate that the target model is likely to be a normal model, the target model is determined to be a non-abnormal model, thereby achieving an accurate judgment on whether the target model is an abnormal model.
[0089] It should be noted that in this application, by simultaneously evaluating each model using the aforementioned three evaluation methods, the model indicators are analyzed from multiple dimensions, thereby effectively identifying abnormal models and improving the recognition accuracy and efficiency of this application.
[0090] In one optional embodiment, after evaluating the anomaly level of each model based on the index values of multiple model indicators corresponding to each model, the model evaluation system can determine whether there is an anomalous model in at least one model based on the anomaly level. If there is an anomalous model in at least one model, a stability test is performed on multiple input variables of the anomalous model to obtain the test results. Then, based on the test results, the anomalous input variables are determined.
[0091] Optional, such as Figure 3 As shown, after the model evaluation system determines that an anomaly has occurred, it can trace and attribute the anomaly to its source, thereby quickly locating the problem and facilitating business personnel and model developers to resolve the anomaly rapidly, reducing losses caused by model anomalies.
[0092] Specifically, when a model anomaly occurs, i.e., when an anomalous model exists, if the stability of the output value of the anomalous model is abnormal, the model evaluation system can quickly calculate whether the stability of each input variable has changed, obtain the model stability index (PSI) value for each input variable, and then sort the PSI values for each input variable. The input variable with the largest fluctuation relative to the value determined in the previous period is then identified as the anomalous input variable. It is important to emphasize that if the anomalous conditions of the anomalous models are different, the aforementioned method can be used to first detect the input variables of the model, and then combined with other methods for more targeted anomaly localization.
[0093] It should be noted that by automatically locating the cause of anomalies, the window period for model optimization is increased, which facilitates model developers in quickly locating problems and iterating the model, and reduces the losses caused by troubleshooting and optimizing model problems.
[0094] Therefore, the solution provided in this application achieves the goal of standardizing the output values of each model and then evaluating each model based on a unified evaluation method, thereby improving the technical effect of model evaluation efficiency and solving the technical problem of low evaluation efficiency caused by only supporting the evaluation of a single model in the prior art.
[0095] Example 2
[0096] According to an embodiment of the present invention, an embodiment of a model evaluation apparatus is provided, wherein, Figure 4 This is a schematic diagram of an optional model evaluation apparatus according to an embodiment of the present invention, such as... Figure 4 As shown, the device includes:
[0097] The acquisition module 401 is used to acquire the output value obtained by each model in at least one model after predicting the input data, as well as the true result corresponding to the input data, and to determine the data adjustment rules corresponding to each model type based on the model type of each model.
[0098] The first processing module 402 is used to standardize the output value of each model according to the data adjustment rules corresponding to each model type, so as to obtain the standardized output value of each model. The standardization process is used to map the output value of at least one model to a preset numerical range.
[0099] The second processing module 403 is used to process the standardized output value and the real result corresponding to each model using preset index determination rules, and determine the index values of multiple model indicators corresponding to each model.
[0100] Evaluation module 404 is used to evaluate the anomaly level of each model based on the index values of multiple model indicators corresponding to each model, wherein the anomaly level is used to determine whether the model is abnormal.
[0101] It should be noted that the above-mentioned acquisition module 401, first processing module 402, second processing module 403 and evaluation module 404 correspond to steps S101 to S104 in the above embodiments. The four modules and the corresponding steps implement the same examples and application scenarios, but are not limited to the content disclosed in the above embodiment 1.
[0102] Optionally, the second processing module includes: a combination submodule, used to combine different standardized output values of the target model with different true results of the target model to obtain multiple target combination terms, wherein the target model is any one of at least one model; a first determination submodule, used to determine the frequency of occurrence of the combined content corresponding to each target combination term of the target model based on the correspondence between the standardized output values and the true results of the target model; and a second determination submodule, used to process the standardized output values, true results, and frequency of occurrence of the combined content corresponding to each target combination term of the target model using preset indicator determination rules to determine the indicator values of multiple model indicators corresponding to the target model.
[0103] Optionally, the evaluation module includes: a third determination submodule, used to determine a first evaluation result corresponding to the target model based on multiple indicator values corresponding to the target model and at least one target threshold, wherein the target threshold corresponds one-to-one with the model indicators; a fourth determination submodule, used to determine a second evaluation result corresponding to the target model based on multiple indicator values corresponding to the target model and the target prediction model; a fifth determination submodule, used to determine a third evaluation result corresponding to the target model based on multiple indicator values corresponding to the target model and the target time series prediction algorithm; and an evaluation submodule, used to evaluate the anomaly level of the target model based on the first evaluation result, the second evaluation result, and the third evaluation result corresponding to the target model.
[0104] Optionally, the third determining submodule further includes: a first determining unit, used to determine the target threshold corresponding to each indicator value of the target model based on the correspondence between the model indicators and the target threshold; a comparison unit, used to compare each indicator value corresponding to the target model with the target threshold corresponding to that indicator value to obtain the comparison result corresponding to each indicator value of the target model; and a second determining unit, used to determine the first evaluation result corresponding to each model based on the comparison result corresponding to the target model.
[0105] Optionally, the fifth determining submodule further includes: a prediction unit, used to predict the indicator change threshold of each model indicator corresponding to the target model based on the target indicator value of each model indicator corresponding to the target model and the target time series prediction algorithm, wherein the indicator change threshold represents the range of change allowed for the indicator value of the model indicator, each model indicator corresponding to the target model corresponds to multiple indicator values, and the multiple indicator values corresponding to the same model indicator are determined at different times, and the target indicator value is the indicator value determined before a preset time; and a third determining unit, used to determine the third evaluation result corresponding to the target model based on the indicator change threshold of each model indicator corresponding to the target model and the multiple indicator values corresponding to the target model.
[0106] Optionally, the evaluation submodule further includes: a fourth determining unit, used to determine the anomaly level of the target model as the first anomaly level when a target evaluation result exists among the first evaluation result, second evaluation result, and third evaluation result corresponding to the target model, wherein the target evaluation result is an evaluation result that represents the anomaly probability of the target model being greater than a preset probability, and the first anomaly level represents the model as an anomaly model; and a fifth determining unit, used to determine the anomaly level of the target model as the second anomaly level when a target evaluation result does not exist among the first evaluation result, second evaluation result, and third evaluation result corresponding to the target model, wherein the second anomaly level represents the model as a non-anomaly model.
[0107] Optionally, the model evaluation device further includes: a first determining module, used to determine whether an anomalous model exists in at least one model based on the anomaly level; a testing module, used to perform stability tests on multiple input variables of the anomalous model if an anomalous model exists in at least one model, and obtain test results; and a second determining module, used to determine the anomalous input variables based on the test results.
[0108] Example 3
[0109] According to another aspect of the present invention, a computer-readable storage medium is also provided, wherein a computer program is stored in the computer-readable storage medium, wherein the computer program is configured to execute the above-described model evaluation method at runtime.
[0110] Example 4
[0111] According to another aspect of the present invention, an electronic device is also provided, wherein, Figure 5 This is a schematic diagram of an optional electronic device according to an embodiment of the present invention, such as... Figure 5 As shown, the electronic device includes one or more processors; and a memory for storing one or more programs, which, when executed by one or more processors, enable the one or more processors to run the programs, wherein the programs are configured to execute the model evaluation method described above during runtime.
[0112] The sequence numbers of the above embodiments of the present invention are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.
[0113] In the above embodiments of the present invention, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments.
[0114] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. The device embodiments described above are merely illustrative; for example, the division of units can be a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the displayed or discussed mutual couplings, direct couplings, or communication connections may be through some interfaces; indirect couplings or communication connections between units or modules may be electrical or other forms.
[0115] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0116] Furthermore, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0117] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, read-only memory (ROM), random access memory (RAM), portable hard drives, magnetic disks, or optical disks.
[0118] The above are merely preferred embodiments of the present invention. It should be noted that those skilled in the art can make various improvements and modifications without departing from the principle of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.
Claims
1. A model evaluation method, characterized by, The method comprises the following steps: obtaining output values obtained after each model of at least one model predicts input data, and real results corresponding to the input data, and determining data adjustment rules corresponding to each model type based on the model type of each model, wherein the model is a credit card risk assessment model, and the input data of the model comprises asset information and consumption information of a user; performing standardization processing on the output values corresponding to each model according to the data adjustment rules corresponding to each model type, to obtain standardized output values corresponding to each model, wherein the standardization processing is used to map the output values corresponding to the at least one model to a preset numerical range, and the output values of the model are used to represent default risk; determining index values of a plurality of model indexes corresponding to each model by processing the standardized output values corresponding to each model and the real results according to a preset index determination rule; evaluating an abnormality level of each model based on the index values of the plurality of model indexes corresponding to each model, wherein the abnormality level is used to determine whether the model is abnormal; determining index values of a plurality of model indexes corresponding to each model by processing the standardized output values corresponding to each model and the real results according to a preset index determination rule, comprising: combining different target standardized output values corresponding to a target model with different target real results corresponding to the target model to obtain a plurality of target combination entries, wherein the target model is any one of the at least one model; determining the number of occurrences of combination contents corresponding to each target combination entry corresponding to the target model according to the corresponding relationship between the target standardized output values and the target real results; determining index values of a plurality of model indexes corresponding to the target model by processing the standardized output values corresponding to the target model, the real results, and the number of occurrences of the combination contents corresponding to each target combination entry according to a preset index determination rule.
2. The method of claim 1, wherein, evaluating an abnormality level of each model based on the index values of the plurality of model indexes corresponding to each model, comprising: determining a first evaluation result corresponding to the target model based on the plurality of index values corresponding to the target model and at least one target threshold value, wherein the target threshold value corresponds to the model index one by one; determining a second evaluation result corresponding to the target model based on the plurality of index values corresponding to the target model and a target prediction model; determining a third evaluation result corresponding to the target model based on the plurality of index values corresponding to the target model and a target time series prediction algorithm; evaluating an abnormality level of the target model based on the first evaluation result, the second evaluation result, and the third evaluation result corresponding to the target model.
3. The method of claim 2, wherein, determining a first evaluation result corresponding to the target model based on the plurality of index values corresponding to the target model and at least one target threshold value, comprising: determining a target threshold value corresponding to each index value of the target model based on the corresponding relationship between the model index and the target threshold value; The target model corresponding to each index value is compared with a target threshold value corresponding to the index value, to obtain a comparison result corresponding to each index value of the target model; Based on the comparison result corresponding to the target model, a first evaluation result corresponding to each model is determined.
4. The method of claim 2, wherein, Based on the target model corresponding to multiple index values and a target time series prediction algorithm, a third evaluation result corresponding to the target model is determined, including: Based on the target index value of each model index corresponding to the target model and the target time series prediction algorithm, an index variation threshold value of each model index corresponding to the target model is predicted, wherein the index variation threshold value represents a variation range allowed for the index value of the model index, each model index corresponding to the target model corresponds to multiple index values, and the multiple index values corresponding to the same model index are determined at different times, and the target index value is an index value determined before a preset time; Based on the index variation threshold value of each model index corresponding to the target model and the multiple index values corresponding to the target model, the third evaluation result corresponding to the target model is determined.
5. The method of claim 2, wherein, Based on the first evaluation result, the second evaluation result and the third evaluation result corresponding to the target model, the abnormal level of the target model is evaluated, including: When there is a target evaluation result in the first evaluation result, the second evaluation result and the third evaluation result corresponding to the target model, it is determined that the abnormal level of the target model is a first abnormal level, wherein the target evaluation result is an evaluation result representing that the abnormal probability of the target model is greater than a preset probability, and the first abnormal level represents that the model is an abnormal model; When there is no target evaluation result in the first evaluation result, the second evaluation result and the third evaluation result corresponding to the target model, it is determined that the abnormal level of the target model is a second abnormal level, wherein the second abnormal level represents that the model is a non-abnormal model.
6. The method of claim 5, wherein, After evaluating the abnormal level of each model based on the index values of multiple model indexes corresponding to each model, the method further includes: Based on the abnormal level, it is determined whether there is an abnormal model in the at least one model; If there is an abnormal model in the at least one model, a stability test is performed on multiple input model variables of the abnormal model, to obtain a test result; Based on the test result, an abnormal input model variable is determined.
7. A model evaluation apparatus characterized by comprising: The method includes: An acquisition module is configured to acquire an output value obtained after each model in at least one model predicts input data, and a true result corresponding to the input data, and determine a data adjustment rule corresponding to each model type based on a model type of each model, wherein the model is a credit card risk assessment model, and the input data of the model includes asset information and consumption information of a user; The first processing module is configured to perform standardization processing on the output value corresponding to each model according to a data adjustment rule corresponding to each model type, to obtain a standardized output value corresponding to each model, wherein the standardization processing is configured to map the output value corresponding to the at least one model to a preset numerical range, and the output value of the model is used to represent a default risk; The second processing module is configured to perform processing on the standardized output value corresponding to each model and a real result according to a preset index determination rule, to determine an index value of a plurality of model indexes corresponding to each model; The evaluation module is configured to evaluate an abnormal level of each model based on the index value of the plurality of model indexes corresponding to each model, wherein the abnormal level is used to determine whether the model is abnormal. The second processing module includes: a combination submodule configured to combine different target standardized output values corresponding to a target model and different target real results corresponding to the target model to obtain a plurality of target combination entries, wherein the target model is any one of the at least one model; a first determination submodule configured to determine a number of occurrences of combination content corresponding to each target combination entry corresponding to the target model according to a corresponding relationship between the target standardized output value and the target real result; and a second determination submodule configured to perform processing on the standardized output value corresponding to the target model, the real result, and the number of occurrences of the combination content corresponding to each target combination entry according to a preset index determination rule, to determine the index value of the plurality of model indexes corresponding to the target model.
8. A computer-readable storage medium, characterized in that, The computer readable storage medium stores a computer program, wherein the computer program is set to execute the model evaluation method in any one of claims 1 to 6 when running.
9. An electronic device, comprising: The electronic device includes one or more processors; The memory is configured to store one or more programs, and when the one or more programs are executed by the one or more processors, the one or more processors are caused to implement a program for running, wherein the program is set to execute the model evaluation method in any one of claims 1 to 6 when running.
Citation Information
Patent Citations
Consumption credit scene risk assessment method based on random forest algorithm
CN112037009A
Model evaluation method, model evaluation device, electronic device and storage medium
CN113052509A