Risk control model stability improving method

Through the identification, analysis and adjustment strategies of the risk control model, combined with clustering algorithms and hyperparameter optimization, the target risk control model is generated, and the stability problem of the risk control model when changes in the economic environment is solved, achieving higher stability and adaptability.

CN120235697AActive Publication Date: 2025-07-01HANGYIN CONSUMER FINANCE CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510727522.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-03
Publication Date
2025-07-01
Estimated Expiration
2045-06-03

AI Technical Summary

Technical Problem

When the economic environment changes, the existing risk control model has low stability in analyzing customer group data, and there are insufficient existing evaluation indicators and hyperparameter optimization methods, resulting in the model being insufficiently comprehensive and stable in actual applications.

Method used

By identifying and analyzing the preset risk control model of the enterprise, determining the credit evaluation map, obtaining regulatory demand information and customer group information, using the preset stability map generation and adjustment strategy, adjusting the data processing algorithm and process of the risk control model, combining clustering algorithms and hyperparameter optimization methods, a target risk control model is generated.

Benefits of technology

The stability and generalization capabilities of the risk control model are improved, ensuring that the model can better analyze customer group data when the economic environment changes and meet the specific needs of the enterprise.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120235697A_ABST
    Figure CN120235697A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of credit risk assessment, in particular to a risk control model stability improvement method, which comprises the following steps: identifying and analyzing a preset risk control model of an enterprise to determine a credit assessment map of the preset risk control model; performing comparative analysis on the credit assessment atlas of the existing preset risk control model of the enterprise based on a preset stability atlas, and finding a specific reason for the lack of stability of the preset risk control model currently used by the enterprise; and then an available algorithm and a data processing flow are determined in a targeted manner according to specific reasons, so that the stability of the preset risk control model is improved, a reasonable evaluation limit is given to the customer, and the problem that the stability of analyzing the customer group data by the existing risk control model is relatively low when the economic environment changes is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of credit risk assessment, and particularly to a method for improving the stability of a risk control model. Background Art

[0002] In the field of credit risk assessment, the gradient boosting algorithm is widely used due to its excellent performance. At the algorithm principle level, such serial algorithms learn the complex non-linear relationship between features and labels by continuously correcting the prediction deviation on the training samples. Compared with traditional logistic regression and random forest algorithms, they inherently have a higher risk of overfitting. At the practical application level, when the economic environment changes, the data distribution of the customer group will change greatly, thus impacting the stability of such risk control models. To ensure the smooth development of business, improving the stability of risk control models has always been a key concern for financial institutions.

[0003] Before training a machine learning model, the user needs to independently set the evaluation function and hyperparameters. Ideal evaluation functions and hyperparameters can significantly improve the stability of the risk control model. The current mainstream risk control model evaluation indicators in the industry are usually KS, AUC, or the top capture rate. KS reflects the discrimination ability of the risk control model for good and bad customers by calculating the maximum absolute difference between the cumulative distributions of good and bad customers predicted by the model. However, this indicator only focuses on the maximum value point of the discrimination ability of the risk control model and ignores the overall distribution, so the overall performance evaluation of the risk control model is not comprehensive enough. AUC comprehensively reflects the performance of the risk control model at all possible thresholds by calculating the area under the ROC curve. However, this indicator is sensitive to extreme values, and since it does not consider cost-benefit factors, it cannot directly reflect the performance of the risk control model in practical applications. The top capture rate reflects the identification ability of the risk control model among high-risk customers by calculating the proportion of actual bad customers among the top 10% or 20% of customers with the highest predicted scores by the model. However, this indicator overly relies on the prediction of the risk control model for high-risk customers and may ignore the prediction accuracy for other customer groups.

[0004] The current mainstream hyperparameter optimization methods in the industry include grid search, random search, etc. Grid search is an exhaustive search strategy that can ensure finding the best combination of hyperparameters within the specified range. However, when the hyperparameter value range is large, its computational cost is too high to be applied in practical scenarios. Different from grid search, random search randomly selects hyperparameter values from a pre-defined hyperparameter distribution for evaluation. This strategy can more effectively utilize computational resources, but its results have a certain degree of randomness and do not consider historical evaluation information.

[0005] Therefore, the present invention provides a method for improving the stability of a risk control model to solve the above problems. Summary of the Invention

[0006] In view of the above situation, to overcome the defects of the existing technology, the present invention provides a method for improving the stability of a risk control model to solve the problem that when the economic environment changes, the stability of the existing risk control model in analyzing customer group data is relatively low.

[0007] To achieve the above object, the technical solution adopted by the present invention is as follows: A method for improving the stability of a risk control model, comprising: identifying and analyzing a preset risk control model of an enterprise to determine a credit evaluation map of the preset risk control model; obtaining the regulatory requirement information of the enterprise, and based on the customer group information and the regulatory requirement information, analyzing the credit evaluation map by using a preset stability map to generate an adjustment strategy; Adjusting the preset risk control model based on the adjustment strategy to obtain an initial risk control model; Obtaining the dynamic credit evaluation data of the customer group and supplementary data related to the adjustment strategy; Based on preset dimension indicators, using a clustering algorithm to process the supplementary data and the dynamic credit evaluation data to establish a multi-dimensional indicator data set; and dividing the multi-dimensional indicator data set into a training set, a test set and multiple validation sets; using the training set, the test set and the multiple validation sets to process the initial risk control model to obtain a target risk control model.

[0008] Preferably, the identifying and analyzing the preset risk control model of the enterprise to determine the credit evaluation map of the preset risk control model includes: obtaining the construction record information and the current program information of the preset risk control model; determining the construction framework of the preset risk control model according to the construction record information; determining the processing program corresponding to the framework based on the construction framework and the note information in the program information; analyzing the processing program to determine the data processing algorithm and the data processing flow; and determining the credit evaluation map according to the construction framework, the data processing algorithm and the data processing flow.

[0009] Preferably, the analyzing the credit evaluation map by using a preset stability map based on the customer group information and the regulatory requirement information to generate an adjustment strategy includes: extracting features from the customer group information and the regulatory requirement information to obtain key information; comparing the preset stability map and the credit evaluation map in the manner from the framework to the data processing flow to obtain the evaluation missing features lacking in the credit evaluation map; adjusting the evaluation missing features according to the key information to obtain initial adjustment features; and obtaining the corresponding construction framework and / or data processing algorithm and / or data processing flow from the preset stability map according to the initial adjustment features to generate an adjustment strategy.

[0010] Preferably, adjusting the preset risk control model according to the adjustment strategy to obtain an initial risk control model includes: generating a data processing program according to the data processing algorithm and / or data processing flow in the adjustment strategy; generating an associated structure document of the data processing program according to the relationship between the program documents in the preset risk control model; adding the associated structure program with the data processing program to the corresponding position of the preset risk control model to obtain an initial risk control model.

[0011] Preferably, obtaining supplementary data related to the adjustment strategy includes: determining the supplementary data type and supplementary data name according to the program corresponding to the adjustment strategy; analyzing the supplementary data type and supplementary data name to determine the source location of the supplementary data; collecting the information corresponding to the supplementary data type and supplementary data name according to the preset data collection method corresponding to the source location of the supplementary data to obtain supplementary data.

[0012] Preferably, processing the supplementary data and dynamic credit assessment data using a clustering algorithm based on preset dimension indicators to establish a multi-dimensional indicator dataset includes: performing clustering analysis on the supplementary data and dynamic credit assessment data using a hierarchical clustering algorithm and a K-Means clustering algorithm to obtain a first clustering type, and obtaining a dimension dataset of the same type according to the first clustering type and each preset clustering type of the preset dimension indicators; using the STING algorithm to perform deletion or addition processing on the data different from the preset clustering type in the first clustering type to obtain a multi-dimensional indicator dataset.

[0013] Preferably, dividing the multi-dimensional indicator dataset into a training set, a test set, and multiple validation sets according to a preset classification rule includes: selecting a preset number of datasets with the same time span and similar data distribution from the multi-dimensional indicator dataset and dividing them into a training set and a test set according to a preset ratio; selecting data with different performance periods and different overdue degrees from the multi-dimensional indicator dataset as validation sets.

[0014] Preferably, processing the initial risk control model using multiple validation sets to obtain a target risk control model includes: optimizing the hyperparameters of the initial risk control model using a Bayesian optimization algorithm and a grid search method based on a greedy strategy.

[0015] The beneficial effects of the present invention are:

[0016] 1. The present invention compares and analyzes the credit assessment maps of the existing preset risk control models of the enterprise based on the preset stability map, and finds out the specific reasons for the lack of stability of the preset risk control models currently used by the enterprise; then, based on the specific reasons, the available algorithms and data processing processes are targeted to adjust the preset risk control models currently used by the enterprise, and a target risk control model with enhanced stability is obtained, thereby solving the problem of low stability of the existing risk control models of the enterprise, and helping to solve the problem of low stability of the existing risk control models in analyzing customer group data when the economic environment changes.

[0017] 2. The present invention can analyze the credit assessment map based on customer group information and regulation demand information using a preset stability map to generate an adjustment strategy; by introducing the enterprise's regulation demand information, the preset risk control model can be adjusted in a targeted manner so that the final target risk control model fits the needs of the enterprise.

[0018] 3. When using the validation set to optimize the hyperparameters of the initial risk control model after the test, the Bayesian optimization algorithm and grid search method are used based on the greedy strategy to optimize the hyperparameters of the initial risk control model. The combination of the greedy strategy, the Bayesian optimization algorithm and the grid search method can effectively optimize the hyperparameters of the initial risk control model, which helps to improve the stability and generalization ability of the model, and thus helps to solve the problem of low stability of the existing risk control model in analyzing customer group data when the economic environment changes. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] Figure 1 The present invention is a schematic flow chart of a method for improving the stability of a risk control model. DETAILED DESCRIPTION

[0020] The following will refer to the attached Figure 1 The embodiments of the present invention are described in detail. It should be understood by those skilled in the art that these embodiments are only used to explain the technical principles of the present invention and are not intended to limit the protection scope of the present invention.

[0021] A method for improving the stability of risk control models, as shown in the attached Figure 1 As shown, the following steps are included:

[0022] Step S11: Identify and analyze the enterprise's preset risk control model to determine the credit assessment map of the preset risk control model.

[0023] Step S12: Obtain the enterprise's regulation demand information, and based on the customer group information and the regulation demand information, use a preset stability map to analyze the credit assessment map to generate an adjustment strategy.

[0024] Among them, the content of the preset stability map includes the data types for evaluating user credit, the data acquisition methods, the analysis process for evaluating user credit, the algorithms and processing procedures that can be used in each process step, etc. In the preset stability map, it is displayed and analyzed in a layered and folded manner, from general to specific, so as to quickly compare and analyze the credit evaluation map obtained based on the preset risk control model and the preset stability map, and quickly obtain the difference results between the two.

[0025] Similarly, the architecture and content of the credit evaluation map are similar to those of the preset stability map.

[0026] Step S13: Based on the adjustment strategy, adjust the preset risk control model to obtain an initial risk control model.

[0027] Step S14: Obtain the dynamic credit evaluation data of the customer group and the supplementary data related to the adjustment strategy.

[0028] Among them, the supplementary data are the data not analyzed in the enterprise's preset risk control model, including but not limited to industry dynamic reports, government support policy information, etc.

[0029] Step S15: Based on the preset dimension indicators, use the clustering algorithm to process the supplementary data and the dynamic credit evaluation data to establish a multi-dimensional indicator data set; and divide the multi-dimensional indicator data set into a training set, a test set, and multiple validation sets according to the preset classification rules; use the training set, the test set, and multiple validation sets to process the initial risk control model to obtain the target risk control model.

[0030] Among them, the preset dimension indicators include but not limited to user information dimension indicators, financial dimension indicators, credit history dimension indicators, user behavior preference dimension indicators, social responsibility dimension indicators, risk feature dimension indicators, and development prospect dimension indicators.

[0031] Through the setting method of this embodiment, the present invention can compare and analyze the credit evaluation map of the enterprise's existing preset risk control model based on the preset stability map, and find out the specific reasons for the lack of stability of the preset risk control model currently used by the enterprise; and then specifically determine the available algorithms and data processing processes according to the specific reasons to improve the stability of the preset risk control model and give customers a reasonable evaluation amount.

[0032] Furthermore, based on the preset stability map, a comparative analysis is conducted on the credit assessment maps of the enterprise's existing preset risk control models to find out the specific reasons for the lack of stability of the preset risk control models currently used by the enterprise; then, based on the specific reasons, the available algorithms and data processing processes are targeted to adjust the preset risk control models currently used by the enterprise, and a target risk control model with enhanced stability is obtained, thereby solving the problem of low stability of the enterprise's existing risk control models, and helping to solve the problem of low stability of customer group data analyzed by existing risk control models when the economic environment changes.

[0033] In one embodiment of the present invention, an enterprise's preset risk control model is identified and analyzed to determine a credit assessment map of the preset risk control model, including: obtaining construction record information and current program information of the preset risk control model; determining a construction framework of the preset risk control model based on the construction record information; determining a processing program corresponding to the framework based on the construction framework and the remark information in the program information; analyzing the processing program to determine a data processing algorithm and a data processing flow; and determining a credit assessment map based on the construction framework, the data processing algorithm, and the data processing flow.

[0034] The construction record information of the preset risk control model includes but is not limited to the update information of each version of the preset risk control model, the initial construction plan, the construction algorithm, and the debugging information; the current program information of the preset risk control model is the front-end program information, back-end program, database program information, and test program information involved in the entire risk control model. The data processing flow includes the selection logic, judgment logic, and loop logic of the data processing in the program.

[0035] When determining the construction framework of the preset risk control model based on the construction record information, a text processing model is used to perform comprehensive recognition and analysis on the construction record information to obtain the overall construction framework of the preset risk control model. The text processing model can be a bag-of-words model, a text feature extraction model based on a neural network, etc., which can perform text feature extraction to obtain a model of the construction framework.

[0036] Furthermore, when part of the program information does not contain remark information, the package import information and the specific program in the part of the program will be analyzed to determine which part of the information in the construction framework the program belongs to; and the content of the specific program will be analyzed to determine the data processing algorithm and data processing flow of the part of the program.

[0037] Analyzing the processing program to determine the data processing algorithm and data processing flow can be: using one or more of the abstract syntax tree (AST) analysis method, the control flow graph (CFG) analysis method, the data flow analysis method, etc. to analyze the processing program to determine the specific data processing algorithm and data processing flow of the processing program, thereby being able to comprehensively and accurately construct a credit assessment map of the preset risk control model.

[0038] Through the setting method of this embodiment, the present invention can identify and analyze the preset risk control model of the enterprise to determine the credit assessment map of the preset risk control model, so that the determined credit assessment map is accurate and comprehensive, and thus the assessment missing features obtained by subsequent comparative analysis are more accurate and reliable.

[0039] In one embodiment of the present invention, based on customer group information and regulation demand information, a preset stability map is used to analyze the credit assessment map to generate an adjustment strategy, including: extracting features from the customer group information and regulation demand information to obtain key information; comparing the preset stability map and the credit assessment map in a framework-to-data processing flow manner to obtain assessment missing features that are lacking in the credit assessment map; adjusting the assessment features according to the key information to obtain initial adjustment features; and obtaining the corresponding construction framework and / or data processing algorithm and / or data processing flow from the preset stability map according to the initial adjustment features to generate an adjustment strategy.

[0040] Among them, the acquisition methods of control demand information include text acquisition, voice acquisition or video acquisition; the acquisition method of customer group information is obtained by analyzing the customer information stored in the database. When extracting features from customer group information and control demand information to obtain key information, a text feature extraction model is used to extract features from these customer group information and control demand information, and determine the target analysis features of these information, so as to determine the focus of the adjustment of the preset risk control model required by the enterprise; for example, the data processing of the adjusted preset risk control model obtains more reasonable results, the data analysis capability of the preset risk control model is more stable, the data analysis speed of the preset risk control model is faster, and the program protection of the preset risk control model is safer.

[0041] When comparing the preset stability map and the credit assessment map in the manner of framework to data processing flow, a comparative analysis will be conducted in the order from the type in the framework to the specific implementation method, so as to comprehensively determine the framework types and specific implementation methods that are missing in the credit assessment map.

[0042] Through the above implementation, the present invention can analyze the credit assessment map based on customer group information and regulation demand information using a preset stability map, generate an adjustment strategy that is comprehensive and in line with the needs of the enterprise corresponding to the credit assessment map, and thus make the preset risk control model adjusted by the adjustment strategy more in line with the actual needs of the enterprise, the credit assessment result is more reasonable, and the assessment stability of the preset risk control model is stronger. And by introducing the regulation demand information of the enterprise, the preset risk control model can be adjusted in a targeted manner, so that the final target risk control model fits the needs of the enterprise.

[0043] In an embodiment of the present invention, based on an adjustment strategy, a preset risk control model is adjusted to obtain an initial risk control model, including: generating a data processing program according to the data processing algorithm and / or data processing flow in the adjustment strategy; generating an associated structure document of the data processing program according to the relationship between each program document in the preset risk control model; adding the associated structure program with the data processing program to the corresponding position of the preset risk control model to obtain an initial risk control model.

[0044] In this embodiment, according to the corresponding adjustment content in the adjustment strategy, the corresponding data processing program is found from the preset program corresponding library, and then, based on the generated associated structure document, the data processing program is added to the associated structure document to form an adjustment document corresponding to this part of the adjustment strategy; then the adjustment document is added to the corresponding program item of the corresponding module in the preset risk control model.

[0045] After each adjustment document is added to the preset risk control model, debugging is performed according to the program document associated with the adjustment document, and the call between program documents is tested to ensure that there is no conflict between the debugging document and the existing program documents.

[0046] Through the above implementation manner, the present invention can perform targeted processing on the preset risk control model based on the adjustment strategy, generate an adjustment document that fits the preset risk control model, and obtain an initial risk control model. Compared with the traditional manual inspection and adjustment method, through the setting method of this embodiment, the present invention can automatically generate an adjustment document corresponding to the preset risk control model according to the adjustment strategy, saving the time of manual processing, improving the data processing speed, and avoiding the interference of human factors.

[0047] In an embodiment of the present invention, obtaining supplementary data related to the adjustment strategy includes: determining the supplementary data type and supplementary data name according to the program corresponding to the adjustment strategy; analyzing the supplementary data type and supplementary data name to determine the source location of the supplementary data; collecting the information corresponding to the supplementary data type and supplementary data name according to the preset data collection method corresponding to the source location of the supplementary data to obtain the supplementary data.

[0048] Among them, after obtaining the initial risk control model, in order to train and adjust the parameters of the initial risk control model, first, the newly added supplementary data type and supplementary data name in the adjustment strategy are obtained; then the source location of these data is determined so as to be classified into the corresponding data collection method; these data are obtained through the corresponding data collection method.

[0049] For example, when additional support policy information needs to be obtained, a web crawler method is used to obtain the corresponding support policy information at the corresponding time from official websites with high authenticity; when additional information on the current industry status needs to be obtained, a web crawler method is used to obtain articles, videos, etc. in the industry within a preset time period with high evaluation authenticity, and then a natural language processing model is used to analyze these articles and videos to obtain the current status information of the industry. When additional personal information is needed, the customer is asked to supplement the relevant information, which is stored in the corresponding personal information database after being obtained.

[0050] Through the setting method of this embodiment, the present invention can accurately obtain the adjustment data corresponding to the adjustment strategy. After obtaining these supplementary data, the data can be processed and divided into a training set, a test set, and a validation set, and then the initial risk control model is trained in sequence, the trained initial risk control model is tested, and the validation set is used for verification. After passing the verification, the final target risk control model that meets the enterprise requirements is obtained.

[0051] In an embodiment of the present invention, based on preset dimension indicators, a clustering algorithm is used to process the supplementary data and dynamic credit assessment data to establish a multi-dimensional indicator data set, including: using a hierarchical clustering algorithm and a K-Means clustering algorithm to perform clustering analysis on the supplementary data and dynamic credit assessment data to obtain a first clustering type, and obtaining a dimension data set of the same type according to the first clustering type and each preset clustering type of the preset dimension indicators; using the STING algorithm to perform deletion or addition processing on the data in the first clustering type that is different from the preset clustering type to obtain a multi-dimensional indicator data set.

[0052] In this embodiment, the processing process of the clustering method that combines the hierarchical clustering algorithm and the K-Means clustering algorithm is as follows: The hierarchical clustering algorithm is used to perform preliminary clustering on the data. Since the hierarchical clustering does not require specifying the number of clusters in advance, it can flexibly process data with different distributions. During the hierarchical clustering process, an appropriate agglomerative or divisive method is selected according to the data characteristics and clustering requirements to improve the clustering accuracy. The clustering tree obtained through hierarchical clustering can visually observe the clustering structure of the data, and based on this, determine the appropriate number of clusters K. This helps to avoid the local optimal solution problem caused by randomly selecting the initial clustering centers in the K-Means algorithm. According to the determined number of clusters K, K representative clustering centers are selected from the results of hierarchical clustering as the initial clustering centers of the K-Means algorithm. These centers can more accurately reflect the actual distribution of the data. Using the selected initial clustering centers, the K-Means algorithm is executed. Since the selection of the initial clustering centers is more reasonable, the K-Means algorithm can converge to the global optimal solution faster, thus obtaining the final first clustering type. Through the combination of the hierarchical clustering algorithm and the K-Means clustering algorithm, the present invention can give full play to the advantages of both, improving the accuracy and stability of clustering.

[0053] When using the STING algorithm to perform deletion or addition processing on the data in the first clustering type that is different from the preset clustering type to obtain a multi-dimensional index data set, for the data points in the first clustering type that are different from the preset clustering type, STING will, according to the statistical information of the grid cell where the data point is located, assign it to the most similar cluster, or regard it as noise and perform deletion processing.

[0054] Through the above method, the present invention can accurately and comprehensively classify the supplementary data and dynamic credit assessment data. Through the combination of the hierarchical clustering algorithm and the K-Means clustering algorithm, the present invention can give full play to the advantages of both, improving the accuracy and stability of clustering. And the STING algorithm is used to process the data that does not conform to the preset clustering type, reducing the deletion of data, which helps to quickly accumulate the training volume of the initial risk control model data.

[0055] In an embodiment of the present invention, dividing the multi-dimensional index data set into a training set, a test set, and multiple validation sets according to a preset classification rule includes: selecting a preset number of data sets with the same time span and similar data distribution from the multi-dimensional index data set, and dividing them into a training set and a test set according to a preset ratio; selecting data with different performance periods and different overdue degrees from the multi-dimensional index data set as the validation sets.

[0056] Among them, after each type of data reaches the preset data volume, the training set and the test set are divided according to the preset ratio of 7:3. During the process of using the training set, the test set, and multiple validation sets to process the initial risk control model to obtain the target risk control model, the test set is used to evaluate the performance of the trained initial risk control model, that is, the training situation of the initial risk control model is evaluated by calculating indicators such as accuracy, recall rate, and F1 value. After passing the test of the test set, the validation set is used for cross-validation: the optimal hyperparameter combination is found through multiple trainings and validations to obtain the final target risk control model, and the target risk control model is deployed to the actual environment.

[0057] During the process of using the validation set for hyperparameter tuning, based on the greedy strategy, the Bayesian optimization algorithm and the grid search method are used to optimize the hyperparameters of the initial risk control model; optimizing the existing evaluation functions and hyperparameters in the initial risk control model after testing can improve the stability of the risk control model. When optimizing the evaluation function: by adding different evaluation data sets at different time periods and selecting appropriate evaluation indicators, and then weighted adjustment of the evaluation function for hyperparameter optimization, so as to improve the cross-time stability of the model while meeting the business application requirements. When optimizing the hyperparameters, through the combined design method of Bayesian optimization and grid search, the time for tuning the parameters of the risk control model is reduced and the performance of the risk control model is improved. Through the above methods, the present invention can greatly shorten the time for model training and parameter tuning and effectively improve the generalization and stability of the model on cross-cycle samples.

[0058] Among them, when optimizing the evaluation function, the underlying indicators in the evaluation function are selected according to the business application requirements : when it is necessary to improve the full-sample risk identification ability of the risk control model, AUC is selected; when it is necessary to improve the maximum discrimination ability of the risk control model for good and bad samples, KS is selected; when it is necessary to improve the identification ability of the risk control model for high-risk customers, the head bad capture rate is selected.

[0059] Among them, the evaluation function for hyperparameter optimization is: , Among them, is an adjustable penalty coefficient, is a test index, are all underlying indicators of the evaluation function, and the total number of underlying indicators is n, is a training index, is the test index evaluation function, and the test index evaluation function is used to evaluate the degree of parameter optimization according to the data corresponding to each test index.

[0060] When optimizing the hyperparameters, in a relatively large hyperparameter space Maximize the objective function using the Bayesian optimization method within, and construct a smaller hyperparameter space based on the hyperparameter results of each round of optimization. . Finally, gradually optimize each hyperparameter based on the greedy strategy using grid search.

[0061] Through the setting method of this embodiment, the present invention can optimize the existing evaluation function and hyperparameters in the initial risk control model after testing, and improve the stability of the target risk control model.

[0062] Furthermore, based on the greedy strategy, the hyperparameters of the initial stability model are optimized by combining the Bayesian optimization algorithm and the grid search method. One implementation method is as follows: Define the search space of hyperparameters (including the name, type, and value range of each hyperparameter). Initialize the surrogate model of the Bayesian optimization algorithm (such as the Gaussian process); use the greedy strategy to select the optimal hyperparameter combination for evaluation according to the hyperparameter performance predicted by the surrogate model; update the surrogate model to reflect the new observed data; repeat the above steps until the preset number of Bayesian optimization iterations is reached or a satisfactory hyperparameter combination is found. Define a finer grid near the optimal hyperparameter combination found by Bayesian optimization. Use the grid search method to try all possible hyperparameter combinations one by one within the defined grid. Select the hyperparameter combination with the best performance as the final result; evaluate the performance of the finally selected hyperparameter combination. According to the evaluation result, select whether to accept this hyperparameter combination as the final configuration of the model.

[0063] Specifically: Define the hyperparameter search space, and determine the hyperparameters to be optimized and their possible value ranges. For example, for a support vector machine model, the hyperparameters may include the regularization parameter C, the kernel function type, etc. Initialize the Bayesian optimization algorithm, select a suitable acquisition function (such as Expected Improvement), and set the initial state of the surrogate model. In each step, select the current optimal hyperparameter combination for actual evaluation according to the hyperparameter performance predicted by the surrogate model. Evaluate the selected hyperparameter combination using another validation set to obtain the actual performance metrics. Update the surrogate model, and use the new observed data (hyperparameter combination and its corresponding performance metrics) to update the surrogate model to improve its prediction accuracy. Repeat the above steps until the preset number of iterations is reached or a satisfactory hyperparameter combination is found. Define a finer grid near the optimal hyperparameter combination found by Bayesian optimization. The granularity of the grid can be adjusted according to the actual situation of the problem and the computing resources. Try all possible hyperparameter combinations one by one within the defined grid, and record the performance metrics of each combination. Compare the performance metrics of all hyperparameter combinations within the grid, and select the combination with the best performance as the final result.

[0064] The finally selected hyperparameter combination is evaluated using an independent validation set to verify its generalization ability. If the evaluation results meet the requirements, this hyperparameter combination is used for the final deployment of the model; otherwise, further adjustment of the hyperparameter search space or optimization algorithm can be considered.

[0065] Through the above steps, combining the greedy strategy, Bayesian optimization algorithm, and grid search method, the hyperparameters of the initial risk control model can be effectively optimized, improving the performance and generalization ability of the model.

[0066] The various embodiments of the systems and techniques described above in this document can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on a chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a dedicated or general-purpose programmable processor, receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting the data and instructions to the storage system, the at least one input device, and the at least one output device.

[0067] The program code for implementing the methods of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a dedicated computer, or other programmable data processing devices, such that when the program codes are executed by the processor or controller, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The program codes can be executed entirely on the machine, partially on the machine, executed partially on the machine as an independent software package and partially on a remote machine, or executed entirely on a remote machine or server.

[0068] In the context of this disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0069] In order to provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, speech input, or tactile input).

[0070] The systems and techniques described herein can be implemented in a computing system including backend components (e.g., as a data server), or a computing system including middleware components (e.g., an application server), or a computing system including frontend components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system including any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), and the Internet.

[0071] A computer system can include a client and a server. The client and the server are generally remote from each other and typically interact through a communication network. The client-server relationship is generated by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, a server of a distributed system, or a server incorporating a blockchain.

[0072] So far, the technical solution of the present invention has been described in conjunction with the preferred embodiments shown in the accompanying drawings. However, it is easily understood by those skilled in the art that the protection scope of the present invention is obviously not limited to these specific embodiments. Without departing from the principle of the present invention, those skilled in the art can make equivalent changes or substitutions to the relevant technical features, and the technical solutions after these changes or substitutions will all fall within the protection scope of the present invention.

Claims

1. A method for improving the stability of a risk control model, characterized in that, Including: Identifying and analyzing a preset risk control model of an enterprise to determine a credit assessment graph of the preset risk control model; Obtaining the regulatory requirement information of the enterprise, and analyzing the credit assessment graph by using a preset stability graph based on the customer group information and the regulatory requirement information to generate an adjustment strategy; Adjusting the preset risk control model based on the adjustment strategy to obtain an initial risk control model; Obtaining the dynamic credit assessment data of the customer group and supplementary data related to the adjustment strategy; Based on preset dimension indicators, using a clustering algorithm to process the supplementary data and the dynamic credit assessment data to establish a multi-dimensional indicator data set; and dividing the multi-dimensional indicator data set into a training set, a test set and multiple validation sets; using the training set, the test set and multiple validation sets to process the initial risk control model to obtain a target risk control model.

2. The method for improving the stability of the risk control model according to claim 1, wherein The identifying and analyzing the preset risk control model of the enterprise to determine the credit assessment graph of the preset risk control model includes: obtaining the construction record information and the current program information of the preset risk control model; determining the construction framework of the preset risk control model according to the construction record information; determining the corresponding processing program of the framework based on the construction framework and the note information in the program information; analyzing the processing program to determine the data processing algorithm and the data processing flow; determining the credit assessment graph according to the construction framework, the data processing algorithm and the data processing flow.

3. The method for improving the stability of the risk control model according to claim 1, wherein The analyzing the credit assessment graph by using a preset stability graph based on the customer group information and the regulatory requirement information to generate an adjustment strategy includes: extracting features from the customer group information and the regulatory requirement information to obtain key information; comparing the preset stability graph and the credit assessment graph in the manner from the framework to the data processing flow to obtain the evaluation missing features lacking in the credit assessment graph; adjusting the evaluation missing features according to the key information to obtain initial adjustment features; and obtaining the corresponding construction framework and / or data processing algorithm and / or data processing flow from the preset stability graph according to the initial adjustment features to generate an adjustment strategy.

4. The method for improving the stability of the risk control model according to claim 1, wherein The adjusting the preset risk control model based on the adjustment strategy to obtain an initial risk control model includes: Generating a data processing program according to the data processing algorithm and / or data processing flow in the adjustment strategy; generating an associated structure document of the data processing program according to the relationship between the program documents in the preset risk control model; and adding the associated structure program with the data processing program to the corresponding position of the preset risk control model to obtain an initial risk control model.

5. The method for improving the stability of the risk control model according to claim 1, wherein The obtaining the supplementary data related to the adjustment strategy includes: determining the supplementary data type and the supplementary data name according to the program corresponding to the adjustment strategy; analyzing the supplementary data type and the supplementary data name to determine the source location of the supplementary data; and collecting the information corresponding to the supplementary data type and the supplementary data name according to the preset data collection method corresponding to the source location of the supplementary data to obtain the supplementary data.

6. The method for improving the stability of the risk control model according to claim 1, wherein Based on the preset dimension indicators, a clustering algorithm is used to process the supplementary data and dynamic credit assessment data to establish a multi-dimensional indicator dataset, including: using the hierarchical clustering algorithm and the K-Means clustering algorithm to perform clustering analysis on the supplementary data and dynamic credit assessment data to obtain the first clustering type, and obtaining the dimension datasets of the same type according to the first clustering type and each preset clustering type of the preset dimension indicators; using the STING algorithm to perform deletion or addition processing on the data in the first clustering type that is different from the preset clustering type to obtain a multi-dimensional indicator dataset.

7. The method for improving the stability of the risk control model according to claim 1, wherein Dividing the multi-dimensional indicator dataset into a training set, a test set, and multiple validation sets according to the preset classification rules, including: selecting datasets with a preset quantity, the same time span, and similar data distributions from the multi-dimensional indicator dataset, and dividing them into a training set and a test set according to a preset ratio; selecting data with different performance periods and different overdue degrees from the multi-dimensional indicator dataset as the validation sets.

8. The method for improving the stability of the risk control model according to claim 1, wherein Using multiple validation sets to process the initial risk control model to obtain the target risk control model, including: based on the greedy strategy, using the Bayesian optimization algorithm and the grid search method to optimize the hyperparameters of the initial risk control model.

Citation Information

Patent Citations

  • Risk cause identification method and device and storage medium

    CN113393155A

  • Enterprise credit risk scoring method and device, equipment and storage medium

    CN114202223A

  • Risk control method and device, medium and equipment

    CN116595183A

  • Multi-target risk control strategy optimization method and system

    CN117132001A

  • Optimized risk control strategy generation processing method and system

    CN118410319A