Overdue risk identification method and device and computer storage medium

Through hyperparameter optimization of the overdue risk model, the optimized model is generated using the simulated annealing algorithm, and the problems of low accuracy and low efficiency in the existing technology are solved, and efficient and automated overdue risk identification are achieved.

CN120146984APending Publication Date: 2025-06-13BEIJING WUBA MANXIN INFORMATION TECHNOLOGY CO LTD
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202510018555.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-06
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

The prior art has low accuracy and low efficiency when identifying the risk of overdue repayment by users, and relies on manual observation and fixed rules to determine statistical models, resulting in low recognition efficiency and low degree of automation.

Method used

By obtaining the overdue risk model and sample data to be optimized, determining the hyperparameters and target indicators to be optimized, using simulated annealing algorithm for hyperparameter optimization, generating the optimized overdue risk model, and then realizing the identification and automated return of overdue risk.

Benefits of technology

It improves the accuracy and efficiency of overdue risk identification, realizes automatic identification, reduces the impact of human operations, and improves the quality and efficiency of identification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120146984A_ABST
    Figure CN120146984A_ABST
Patent Text Reader

Abstract

The invention provides an overdue risk identification method and device and a computer storage medium, and the method comprises the steps that an identification terminal obtains a to-be-optimized overdue risk model and sample data, and the sample data comprises overdue sample data and non-overdue sample data; and the identification terminal determines at least one to-be-optimized hyper-parameter of the overdue risk model and at least one target index corresponding to the sample data, wherein the hyper-parameter is located in a preset iterative search range. And based on the sample data and the at least one target index, the identification terminal performs iterative optimization calculation on the hyper-parameter by using a simulated annealing algorithm to obtain at least one optimized hyper-parameter. And the identification terminal generates an optimized overdue risk model based on the optimized hyper-parameter, and carries out overdue risk identification operation to obtain an identification result. And finally, the identification result is returned to the request terminal. According to the embodiment, the recognition accuracy and the automation degree of the overdue risk model are improved through continuous optimization of the hyper-parameters.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data processing, and in particular, to a method, device, and computer storage medium for identifying overdue risks. Background Art

[0002] A user's overdue repayment refers to the act that a borrower fails to repay the loan principal and / or interest in full and on time after the agreed repayment date. This situation not only affects the capital liquidity and income of financial institutions, but may also have a negative impact on the borrower's own credit record.

[0003] Currently, the risk level of a user's overdue repayment is often judged by manually observing the user's relevant information, which not only has a low accuracy rate, but also reduces the efficiency of overdue risk identification. Summary of the Invention

[0004] Embodiments of the present invention provide a method, device, and computer storage medium for identifying overdue risks, so as to improve the accuracy rate and efficiency of overdue risk identification to a certain extent.

[0005] In a first aspect, an embodiment of the present invention provides a method for identifying overdue risks, including:

[0006] In response to an identification request for overdue risks sent by a requesting terminal, the identification terminal obtains an overdue risk model to be optimized and sample data for performing model optimization operations, where the sample data includes overdue sample data and non-overdue sample data;

[0007] The identification terminal determines at least one hyperparameter to be optimized of the overdue risk model and at least one target index corresponding to the sample data, where the hyperparameter to be optimized is within a preset iterative search range;

[0008] The identification terminal performs iterative calculation on the at least one hyperparameter to be optimized within the iterative parameter range based on the sample data, the at least one target index, and according to the simulated annealing algorithm, to obtain at least one optimized hyperparameter;

[0009] The identification terminal determines an optimized overdue risk model based on the at least one optimized hyperparameter, performs an identification operation on overdue risks using the optimized overdue risk model, obtains an identification result of the overdue risk, and returns the identification result to the requesting terminal.

[0010] In a second aspect, an embodiment of the present invention provides an apparatus for identifying overdue risks, including:

[0011] A first acquisition module, configured to acquire an overdue risk model to be optimized and sample data for performing model optimization operations;

[0012] A first determination module, configured to determine at least one hyperparameter to be optimized of the overdue risk model and at least one target metric corresponding to the sample data;

[0013] A first calculation module, configured to perform iterative calculation on the at least one hyperparameter to be optimized within the range of the iterative parameters based on the sample data, the at least one target metric, and according to the simulated annealing algorithm, to obtain at least one optimized hyperparameter;

[0014] A first processing module, configured to determine an optimized overdue risk model based on the at least one optimized hyperparameter, perform an identification operation of overdue risk by using the optimized overdue risk model, obtain an identification result of the overdue risk, and return the identification result.

[0015] In a third aspect, an embodiment of the present invention provides an electronic device, including: a memory, a processor, and a communication interface; wherein, the memory is configured to store one or more computer instructions, and when the one or more computer instructions are executed by the processor, the method for identifying overdue risk as described in the first aspect is implemented.

[0016] In a fourth aspect, an embodiment of the present invention provides a computer storage medium, the computer storage medium stores a computer program, wherein when the computer program runs, it controls the device where the computer-readable storage medium is located to execute the method for identifying overdue risk as described in the first aspect.

[0017] In a fifth aspect, an embodiment of the present invention provides a computer program product, including a computer program, and when the computer program is executed by a processor, the method for identifying overdue risk as described in the first aspect is implemented.

[0018] The method, device, and computer storage medium for identifying overdue risks provided by the embodiments of the present invention, after obtaining an identification request for overdue risks sent by a requesting terminal, obtain an overdue risk model to be optimized and sample data through an identification terminal, determine at least one hyperparameter to be optimized of the overdue risk model and a target indicator corresponding to the sample data, and perform iterative optimization calculations on the hyperparameters based on the sample data, the target indicator, and using the simulated annealing algorithm to obtain optimized hyperparameters; subsequently, generate an optimized overdue risk model based on the optimized hyperparameters, thereby effectively realizing the optimization operation of the overdue risk model using the model annealing algorithm, improving the optimization quality and effect of the overdue risk model to a certain extent; then perform an overdue risk identification operation based on the optimized overdue risk model to obtain an identification result of the overdue risk, and the identification result can be returned to the requesting terminal; in this way, it effectively realizes the overdue risk identification operation through the optimized overdue risk model in an application scenario where overdue risk identification is required, not only effectively improving the quality and efficiency of the overdue risk; moreover, the above overdue risk identification operations can all be automated without manual review or intervention, thereby reducing the influence degree of human operations on overdue risk identification and further improving the practicability of this method. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] To more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0020] Figure 1 It is a schematic diagram of the scenario of a method for identifying overdue risks provided by the embodiments of the present invention;

[0021] Figure 2 It is a schematic flowchart of a method for identifying overdue risks provided by the embodiments of the present invention;

[0022] Figure 3 It is a schematic flowchart of obtaining sample data for model optimization operations provided by the embodiments of the present invention;

[0023] Figure 4 It is a schematic flowchart of determining variable indicators in the original sample data provided by the embodiments of the present invention;

[0024] Figure 5 It is a schematic flowchart of determining failed variable data based on the screened index values provided by the embodiments of the present invention;

[0025] Figure 6Schematic flowchart of determining optimized hyperparameters based on the simulated annealing algorithm provided by the application embodiment of the present invention;

[0026] Figure 7 Schematic flowchart of associating display index values and reasons for model differences provided by the embodiment of the present invention;

[0027] Figure 8 Schematic diagram of the full - process technical solution for optimizing the overdue risk model provided by the embodiment of the present invention;

[0028] Figure 9 Schematic structural diagram of an overdue risk identification device provided by the embodiment of the present invention;

[0029] Figure 10 Schematic structural diagram of the electronic device corresponding to the method for identifying overdue risk provided by the embodiment of the present invention. Detailed implementation manners

[0030] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Apparently, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention. Additionally, the sequence of steps in the following method embodiments is only an example and is not strictly limited.

[0031] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in the embodiments of the present invention are all information and data authorized by the user or fully authorized by all parties. And the collection, use, and processing of relevant data need to comply with the relevant laws, regulations, and standards of relevant countries and regions, and corresponding operation entrances are provided for users to select authorization or rejection.

[0032] Term definition:

[0033] Simulated annealing algorithm: A random search algorithm that combines the advantages of local search and random search and can quickly find the global optimal solution in a wide solution space. Its basic idea is to simulate the annealing process in physics and perform random walks in the search space to find the global optimal solution. In machine learning and optimization problems, the simulated annealing algorithm is widely used to optimize model parameters to improve the performance and accuracy of the model.

[0034] Hyperparameters: In machine learning, the selection of model hyperparameters is crucial for the performance and accuracy of the model. Many models have a set of hyperparameters that need to be preset before model training. However, the selection of these hyperparameters often has no clear rules to follow and requires experimentation and adjustment to find the optimal settings.

[0035] To facilitate the understanding of the technical solutions provided by the embodiments of the present invention by those skilled in the art, the related technologies will be briefly described below:

[0036] With the rapid development of science and technology, the overdue risk identification technology has gradually matured, and thus it has achieved success in practical applications. However, the existing overdue risk identification operations still have the following defects:

[0037] (1) Limitations;

[0038] For network models in the field of overdue risk identification, traditional methods use simple parameter tuning methods (such as grid search or random search) to optimize and adjust the network models.

[0039] However, the above methods have certain limitations: Grid search needs to try all possible parameter combinations one by one, with high computational costs and easy to fall into local optima; Although random search reduces the number of searches, the optimization effect on the network model depends on randomness and it is difficult to find the global optimal solution.

[0040] (2) Low recognition ability and efficiency;

[0041] In the scenario of overdue risk identification, it mostly relies on manual experience or statistical models based on fixed rules to implement the overdue risk identification operation.

[0042] However, for high-dimensional or dynamically changing data, when using the statistical models in the above methods for overdue risk identification operations, problems such as overfitting or underfitting are likely to occur, thus reducing the recognition accuracy. In addition, since the above methods require a large amount of time for data screening, feature extraction, and parameter adjustment, the overall efficiency is low and cannot meet the requirements of real-time and high efficiency.

[0043] (3) Low degree of automation;

[0044] Traditional overdue risk identification methods usually rely on manual operations, including data preparation, parameter adjustment, and model verification, etc.

[0045] However, the above method is not only easily restricted by human experience and capabilities, but also has relatively high operation costs and time consumption. In addition, manual operation is difficult to achieve rapid processing of large-scale data and cannot dynamically adapt to changes in risk characteristics, resulting in slow model updates and difficulty in promptly responding to complex and rapidly changing market environments. Such a low level of automation not only reduces the identification efficiency but also may miss important risk signals.

[0046] To solve the above technical problems, the following will, in conjunction with the accompanying drawings, elaborate on some embodiments of the present invention. In the case of no conflict between the embodiments, the following embodiments and the features in the embodiments can be combined with each other. Additionally, the step timings in the following method embodiments are only examples and are not strictly limited.

[0047] Figure 1 A schematic diagram of the scenario of a method for identifying overdue risks provided by an embodiment of the present invention. Among them, the execution entity of the method for identifying overdue risks can be an identification terminal, and the identification terminal can be communicatively connected to a request terminal. Specifically, as Figure 1 shown:

[0048] The request terminal is used to send an identification request for overdue risks to the identification terminal. Among them, the request terminal can be any computing device with file transfer capabilities. The basic structure of the request terminal can include: at least one processor. The number of processors depends on the configuration and type of the client. The request terminal can also include a memory, which can be volatile, such as RAM, or non-volatile, such as read-only memory (ROM), flash memory, etc., or can also include both types at the same time. Usually, an operating system (OS), one or more application programs, and program data, etc. are stored in the memory. In addition to the processing unit and the memory, the request terminal also includes some basic configurations, such as a network card chip, an IO bus, a display component, and some peripheral devices, etc. Optionally, some peripheral devices can include, for example, a keyboard, a mouse, a stylus, a printer, etc. Other peripheral devices are well known in the art and will not be elaborated here. Optionally, the client can be a PC (personal computer) terminal, an application terminal with the ability to send overdue risks, etc.

[0049] An identification terminal is used to respond to an identification request for overdue risk sent by a requesting terminal, obtain an overdue risk model to be optimized and sample data for model optimization operations, then determine at least one hyperparameter to be optimized for the overdue risk model and at least one target metric corresponding to the sample data, and based on the sample data, at least one target metric, and in accordance with the simulated annealing algorithm, perform iterative calculations on at least one hyperparameter to be optimized within the iterative parameter range to obtain at least one optimized hyperparameter. Finally, based on at least one optimized hyperparameter, determine the optimized overdue risk model, use the optimized overdue risk model to perform an identification operation for overdue risk, obtain the identification result of the overdue risk, and return the identification result to the requesting terminal. Herein, the identification terminal refers to a device that can have the function of identifying overdue risk. In specific applications, it can be implemented as a server, usually referring to a server that uses a network for information planning. Physically, the identification terminal can be any device that can provide computing services, respond to service requests, and perform processing, such as a conventional server, cloud server, cloud host, virtual center, etc. The composition of the identification terminal mainly includes a processor, hard disk, memory, system bus, etc., which is similar to a general computer architecture.

[0050] Based on the above-mentioned identification process of overdue risk, it effectively realizes the optimization operation of the overdue risk model using the model annealing algorithm, and to a certain extent improves the optimization quality and effect of the overdue risk model; then based on the optimized overdue risk model to perform the identification operation of overdue risk. In this way, it can not only effectively improve the automation degree of overdue risk identification, without manual review or intervention, thereby reducing the influence degree of human operation on overdue risk identification, but also can improve the quality and efficiency of overdue risk identification to a certain extent, and can effectively ensure the user experience, further improving the practicality of this method.

[0051] Figure 2 It is a schematic flowchart of a method for identifying overdue risk provided by an embodiment of the present invention; specifically, refer to the appendix Figure 2 As shown, this embodiment provides a method for identifying overdue risk. The execution subject of this method is an identification terminal for overdue risk (hereinafter referred to as the identification terminal). Herein, the identification terminal for overdue risk can be implemented as software, or a combination of software and hardware. When the identification terminal for overdue risk is implemented as hardware, it can specifically be various electronic devices capable of performing the identification operation of overdue risk, including but not limited to personal computers, servers, etc. When the identification terminal for overdue risk is implemented as software, it can be installed in the above-listed electronic devices. Based on the above-mentioned identification terminal for overdue risk, the identification operation of overdue risk can be realized. Specifically, this method for identifying overdue risk may include:

[0052] Step 101: In response to an identification request for overdue risk sent by a requesting terminal, the identification terminal obtains an overdue risk model to be optimized and sample data for performing model optimization operations, where the sample data includes overdue sample data and non-overdue sample data.

[0053] When a user has a need to identify overdue risk, an identification request for overdue risk (hereinafter referred to as "identification request") can be generated or obtained through the requesting terminal and sent to the identification terminal. After the identification terminal receives the identification request sent by the requesting terminal, the identification terminal can obtain the overdue risk model to be optimized and sample data based on this identification request.

[0054] Among them, the overdue risk model to be optimized and the sample data can be stored in a preset area within the identification terminal or in a preset device communicatively connected to the identification terminal. The overdue risk model to be optimized and the sample data can be obtained by accessing the preset area or the preset device.

[0055] For the overdue risk model to be optimized, it can be implemented as an overdue risk model that has not undergone any training operations, or it can also be implemented as a historical version of the overdue risk model. Users can make flexible selections or configurations according to specific application scenarios.

[0056] For the sample data used for model optimization operations, it can specifically include a basic data set for optimizing model performance. In some instances, the sample data can include at least one of the following: user data (collectible identity information authorized by the user), income data (salary income, investment income, or other income, etc.), consumption data (purchase transaction records, purchased commodity information, consumption platform records, etc.), and credit data. In order to improve the accuracy of optimizing the overdue risk model to a certain extent, the sample data can include two types of samples: overdue sample data and non-overdue sample data. Among them, the overdue sample data can refer to data related to users who have had overdue behaviors at historical moments, that is, negative sample data, and the non-overdue sample data refers to data related to users who have not had overdue behaviors at historical moments, that is, positive sample data. The sample data including positive sample data and negative sample data together constitutes a complete sample data set. The above sample data can provide sufficient sample difference for the optimization operation of the overdue risk model to be optimized, which can ensure the quality and efficiency of optimizing the overdue risk model to a certain extent and is conducive to improving the accuracy of the optimized model in performing overdue risk identification operations.

[0057] Step 102: The identification terminal determines at least one hyperparameter to be optimized for the overdue risk model and at least one target indicator corresponding to the sample data, where the hyperparameter to be optimized is within a preset iterative search range.

[0058] After the identification terminal obtains the overdue risk model to be optimized and the sample data, at least one hyperparameter to be optimized for the overdue risk model and at least one target metric corresponding to the sample data can be determined. Among them, the hyperparameter to be optimized refers to the key parameter used to control the model performance during model construction and optimization. For example, when the overdue risk model to be optimized is an XGBoost model, at least one hyperparameter to be optimized can include at least one of the following: learning rate, maximum depth, regularization coefficient, decision tree depth, etc. For the above hyperparameters to be optimized, to ensure the efficiency of hyperparameter optimization and the reliability of the results, an iterative search range corresponding to the hyperparameters can be preset according to business experience, and an iterative search range is provided for the hyperparameters within the preset hyperparameter range. The above iterative search range can be a reasonable interval set by the user according to actual application requirements and historical experience. When performing the optimization operation of the overdue risk model based on the hyperparameters to be optimized, iterative calculations can be performed within the iterative search range corresponding to each hyperparameter to be optimized, which can avoid problems such as excessive computational overhead, non-convergence of optimization results, or degradation of model performance caused by the hyperparameters being in too large or too small a range.

[0059] To perform the optimization operation on the overdue risk model, not only at least one hyperparameter to be optimized for the overdue risk model needs to be determined, but also at least one target metric corresponding to the sample data can be determined. The target metric can refer to the key metric used to evaluate the model performance. In different scenarios, the at least one target metric corresponding to the sample data can be the same or different, that is, the at least one target metric can be determined based on the scenario type of the application scenario. For example, in the scenario of overdue risk prediction, the target metric can include at least one of the following: Area Under the Curve (AUC value), Kolmogorov-Smirnov statistic (KS statistic), LIFT, recall rate, accuracy, etc. More commonly, the target metric can include: AUC value, KS statistic, LIFT value. The above target metrics provide a clear direction for hyperparameter optimization, ensuring that the optimization process aims to improve the model performance.

[0060] Step 103: The identification terminal performs iterative calculations on at least one hyperparameter to be optimized based on the sample data, at least one target metric, and within the iterative parameter range according to the simulated annealing algorithm, to obtain at least one optimized hyperparameter.

[0061] After the recognition terminal determines at least one hyperparameter to be optimized and at least one target metric, it can perform iterative calculation operations on at least one hyperparameter to be optimized within the iterative parameter range based on the acquired sample data, at least one target metric, and by using the simulated annealing algorithm, so as to obtain at least one optimized hyperparameter.

[0062] In some instances, the hyperparameters to be optimized for iterative calculation can be all of the at least one hyperparameter to be optimized. At this time, iterative calculation can be performed on all the hyperparameters to be optimized based on the sample data, at least one target metric, and in accordance with the simulated annealing algorithm within the iterative parameter range, so as to obtain the optimized hyperparameters corresponding to each hyperparameter to be optimized; or, the hyperparameters to be optimized for iterative calculation can be a part of the at least one hyperparameter to be optimized. At this time, iterative calculation can be performed on some of the hyperparameters to be optimized among the at least one hyperparameter to be optimized, so as to obtain the optimized hyperparameters corresponding to some of the hyperparameters to be optimized.

[0063] Specifically, when performing iterative calculation on at least one hyperparameter to be optimized within the iterative parameter range, the recognition terminal performs iterative calculation on the hyperparameter to be optimized within the iterative parameter range according to the rules set by the simulated annealing algorithm. Among them, the simulated annealing algorithm is an optimization algorithm with global search ability. Its basic principle simulates the process of a substance annealing at high temperature. Combining probability randomness and the evaluation value of the objective function, it can effectively jump out of the local optimum in the search space and find the global optimum solution; therefore, after performing multiple rounds of iterative calculation on the hyperparameter to be optimized within the iterative parameter range according to the simulated annealing algorithm, the optimized hyperparameter can be obtained within the iterative parameter range. The obtained optimized hyperparameter can refer to the hyperparameter that meets the preset iterative calculation conditions. Compared with the hyperparameter to be optimized, the optimized hyperparameter can significantly improve the performance of the overdue risk model.

[0064] Step 104: The recognition terminal determines the optimized overdue risk model based on at least one optimized hyperparameter, uses the optimized overdue risk model to perform the recognition operation of the overdue risk, obtains the recognition result of the overdue risk, and returns the recognition result to the requesting terminal.

[0065] After completing the optimization operation on the hyperparameter to be optimized, the recognition terminal generates the optimized overdue risk model based on the obtained at least one optimized hyperparameter. Specifically, it can determine the optimized overdue risk model corresponding to the optimized hyperparameter based on the mapping relationship between the hyperparameter to be optimized and the expected risk model. Specifically, the optimized hyperparameter can be the optimal hyperparameter set obtained through iterative search by the simulated annealing algorithm. For example: the optimized learning rate, regularization coefficient, or model depth, etc. These optimized hyperparameters are used to determine the optimized overdue risk model.

[0066] The optimized overdue risk model is used to implement the identification operation of overdue risk. After obtaining the optimized overdue risk model, the optimized overdue risk model can be used to perform the identification operation of overdue risk. Specifically, the optimized overdue risk model can analyze and identify the data to be analyzed with overdue risk, so as to obtain the identification result of the overdue risk corresponding to the data to be analyzed. The identification result of the overdue risk is used to identify the overdue risk probability of the data to be analyzed.

[0067] For example, the data to be analyzed may include features such as the basic data, income data, consumption data, and credit data of User A. When there is a need to identify overdue risk, the optimized overdue risk model can be used to perform the identification operation of overdue risk on the above data to be analyzed, and then the risk identification result output by the optimized overdue risk model can be obtained. In some instances, the risk identification result can be implemented as the overdue risk level of User A, such as "high risk", "medium risk", "low risk", or the risk identification result can be implemented as the overdue risk score of User A, such as "80 points", "90 points", "10 points", etc. It can be understood that the higher the overdue risk score, the greater the overdue risk; the lower the overdue risk score, the lower the overdue risk.

[0068] After the identification terminal obtains the risk identification result, in order to enable the user to obtain the identification result of the expected risk in a timely manner, the risk identification result can be returned to the request terminal, so as to output or display the identification result of the overdue risk for the user through the request terminal.

[0069] In practical applications, the loan review system of a certain platform can be used as the request terminal. Then, the request terminal can send an overdue risk identification request to the identification terminal. The expected risk identification request may include the basic information, credit information, income information, loan application information, etc. of a certain user. The identification terminal uses the optimized overdue risk model to analyze and identify the repayment overdue risk of the user and outputs an identification result of "high risk". Then, the identification result can be returned to the request terminal, and the loan review system can make a decision to reject the loan or require additional collateral based on the identification result, thereby improving the quality and efficiency of the loan review operation to a certain extent.

[0070] The overdue risk identification method provided by the embodiments of the present invention, after receiving an identification request for overdue risk sent by a requesting terminal, obtains an overdue risk model to be optimized and sample data through an identification terminal, determines at least one hyperparameter to be optimized for the overdue risk model and a target index corresponding to the sample data, and iteratively optimizes and calculates the hyperparameters based on the sample data, the target index, and by using a simulated annealing algorithm to obtain optimized hyperparameters; subsequently, an optimized overdue risk model is generated based on the optimized hyperparameters, thus effectively realizing the optimization operation of the overdue risk model by using the model annealing algorithm, improving the optimization quality and effect of the overdue risk model to a certain extent; then, an overdue risk identification operation is performed based on the optimized overdue risk model to obtain an identification result of the overdue risk, and the identification result can be returned to the requesting terminal; in this way, it effectively realizes the overdue risk identification operation through the optimized overdue risk model in an application scenario where overdue risk identification is required, not only effectively improving the quality and efficiency of the overdue risk; moreover, the above-mentioned overdue risk identification operations can all be automated without manual review or intervention, thereby reducing the influence degree of human operation on overdue risk identification and further improving the practicability of this method.

[0071] Figure 3 It is a schematic flowchart of obtaining sample data for model optimization operation provided by the embodiments of the present invention; on the basis of the above embodiments, continue to refer to the attached Figure 3 As shown, for the sample data used to implement the model optimization operation, it can not only obtain the sample data by accessing a preset area or a preset device, but also obtain the sample data for implementing the model optimization operation by performing data statistics on the original sample data. At this time, obtaining the sample data for model optimization operation in this embodiment may include:

[0072] Step 201: Obtain the original sample data for model optimization operation.

[0073] Among them, in order to obtain the sample data for implementing the model optimization operation to a certain extent, the original sample data for model optimization operation can be obtained first. It can be the original record data without any processing or screening, and can include at least one of the following: user data, income data, consumption data, credit data, or historical overdue information, etc., without any processing or screening.

[0074] In some instances, the original sample data can be obtained through a preset data platform. At this time, obtaining the original sample data for model optimization operation may include: determining a preset data platform communicatively connected to the identification terminal, where the original sample data is stored in the preset data platform; performing data communication operations actively or passively through the preset data platform, and the original sample data stored in the preset data platform can be obtained.

[0075] In some other instances, the original sample data may come from different data sources and may correspond to a valid time period. At this time, obtaining the original sample data for model optimization operations may include: obtaining at least one data source, where the data source includes at least part of the original sample data; determining the valid time period for collecting the original sample data; and determining the original sample data for model optimization operations based on the at least one data source and the valid time period.

[0076] Among them, the data source can be a platform, device, database, etc. structure that can collect or store at least part of the sample data required for model optimization operations, and it can include at least one of the following: database, terminal device, data collection device, data interaction platform, server, or file system, etc.

[0077] In order to obtain as completely as possible the original sample data for model optimization operations, the above-mentioned at least one data source can be communicatively connected to the identification terminal, so that the identification terminal can obtain at least part of the original sample data included in each data source.

[0078] In addition, since the data volume of the original sample data can directly affect the screening efficiency and quality of the final sample data for model optimization operations, therefore, in order to ensure the quality and efficiency of sample data acquisition to a certain extent, the original sample data can be obtained in combination with the valid time period. At this time, the valid time period for collecting the original sample data can be determined. In some instances, the valid time period can be determined through human-computer interaction operations; or, the valid time period usually depends on the optimization cycle of the model and the dynamic change speed of the target problem. For example, for an overdue risk assessment model, data from the most recent 6 months to 3 years may be the most effective; while for a model with high real-time requirements, only data within the past 1 month may be needed. Or, the valid time period can be determined through a pre-configured sample collection cycle, and the valid time period can be the time period corresponding to one sample collection cycle. Among them, the time period corresponding to the sample collection cycle can be based on business requirements and the update frequency of the data source. For example, for a model that is iteratively optimized on a weekly basis, the data collection frequency can be set to once a week, and the new data of the previous week is collected each time. If the model optimization cycle is long (such as one month), the time period needs to cover a longer range. By setting a reasonable valid time period, while reducing the interference of invalid data, it can ensure that the data can comprehensively reflect the characteristics of the target problem.

[0079] Since the original sample data included in each data source may include: original sample data corresponding to the valid time period and original sample data corresponding to the non-valid time period, in order to make the sample data required for model optimization timely, the original sample data corresponding to the non-valid time period may no longer reflect the characteristics or trends of the current business scenario. After determining the valid time period for collecting the original sample data, the original sample data for performing the model optimization operation can be determined based on at least one data source and the valid time period, that is, the original sample data corresponding to the valid time period in each data source is determined as the final original sample data for implementing the model optimization operation. In this way, the diversity and effectiveness of the original sample data are guaranteed to a certain extent.

[0080] In some other instances, the original sample data can be data from different time periods of the same data source. At this time, obtaining the original sample data for performing the model optimization operation may include: determining the valid time period for collecting the original sample data; based on the valid time period, determining the original sample data for performing the model optimization operation. Or, the original sample data can also refer to data from all time periods in different data sources. At this time, obtaining the original sample data for performing the model optimization operation may include: obtaining at least one data source, where the data source includes at least part of the original sample data; based on at least one data source, determining the original sample data for performing the model optimization operation. Those skilled in the art can flexibly adjust and configure the acquisition method of the original sample data according to the specific application scenario or application requirements, as long as the accuracy and reliability of the acquisition of the original sample data can be guaranteed.

[0081] Further, after determining the original sample data for performing the model optimization operation through the data source and / or the valid time period, that is, the preliminary acquisition operation of the original sample data is realized. At this time, the original sample data often includes invalid data such as duplicate data, blank data, and interference data. In order to ensure the quality and effect of model optimization to a certain extent, the collected original sample data can be preliminarily cleaned and sorted. For example, data that does not meet the regulations, duplicate data, blank data, or interference data, etc. can be deleted, data that has become invalid due to too long a time span can be ignored, and sensitive data or privacy data can be excluded. In this way, the effectiveness of the original sample data can be improved to a certain extent.

[0082] Step 202: Perform data statistics on the variable indicators of the original sample data to obtain the index statistical information of the variable indicators.

[0083] After obtaining the original sample data, statistical analysis can be performed on the original sample data, so as to obtain the variable indicators and quantitative indicators included in the original sample data. Among them, the variable indicator refers to the index parameter whose numerical value will change, and the quantitative indicator refers to the index parameter whose numerical value will not change. Since the variable indicator can reflect the parameter change trend and adjustment of the original sample data and has a direct impact on the optimization operation of the overdue risk model, therefore, statistical analysis can be performed on the numerical values of the variable indicators in the original sample data, so as to obtain the index statistical information of the variable indicators.

[0084] In some examples, the variable indicator may refer to the univariate indicator in the original sample data. At this time, the index statistical information of the variable indicator may include at least one of the following: the variable type of the univariate, the number of univariates, the missing rate used to identify the missing proportion of the univariate in the original sample data, the unique value used to identify the number of unique values in the original sample data, the mode in the original sample data, the mean in the original sample data, the median in the original sample data, the minimum value in the original sample data, the maximum value in the original sample data, the range in the original sample data, the quartiles in the original sample data, the percentiles in the original sample data, the frequency used to identify the number of occurrences of each value in the original sample data in the dataset, the frequency ratio used to identify the degree of asymmetry of the original sample data distribution, the variance in the original sample data, the standard deviation in the original sample data, the skewness used to identify the number of occurrences of each value in the original sample data in the dataset, the kurtosis used to identify the steepness of the original sample data, the percentage coefficient of variation used to measure the relative dispersion degree of the original sample data, the weight of evidence (WOE) used to measure the prediction ability of the current binning for default risk, the information value (IV) used as a reference for feature screening before model training, the KS statistic used to distinguish good and bad samples, and the population stability index (PSI) used to measure the deviation between the predicted value and the actual value of the model.

[0085] Step 203: Determine the invalid variable data in the original sample data based on the index statistical information.

[0086] After obtaining the index statistical information, invalid variable data such as outliers, invalid variables, and potentially important variables in the original sample data can be identified through the index statistical information; in some examples, the invalid variable data can be determined by analyzing and processing the index statistical information through a pre-trained neural network model. At this time, determining the invalid variable data in the original sample data based on the index statistical information includes: obtaining a pre-trained neural network model; inputting the index statistical information and the original sample data into the neural network model for analysis and processing, and obtaining the invalid variable data included in the original sample data output by the neural network model.

[0087] In other instances, the failure variable data can also be determined by analyzing and comparing a preset threshold and index statistical information. At this time, based on the index statistical information, determining the failure variable data in the original sample data can include: obtaining the index threshold corresponding to the index statistical information; determining the failure variable data in the original sample data based on the index threshold and the index statistical information. For example, if the missing rate of a certain variable index in the original sample data is too high (such as exceeding 50%), this variable index can be marked as unavailable; if the IV value of a certain variable index is low, indicating that its predictive ability for the target variable is insufficient, it can also be excluded, thus ensuring the accuracy and reliability of obtaining the failure variable data to a certain extent.

[0088] In addition, in addition to determining the failure variable data based on the above index statistical information, variables that have no predictive ability, are unstable, or are highly correlated with other features can also be excluded based on a preset algorithm or preset rules. These variables may introduce noise or redundant information, affecting the accuracy and interpretability of the label. For example, the performance of a certain variable can be as follows: too high missing rate: when the null value rate of a certain variable exceeds a certain threshold (such as 50%), this variable may have an adverse impact on the model because a high null value rate is usually difficult to effectively supplement through methods such as interpolation. Low information value: The IV value of a variable is used to measure the discrimination ability of the variable for the target variable. If the IV value is less than 0.02, it is generally considered that its contribution to the model's predictive ability is insufficient. Poor stability: PS I (Population Stability Index) is used to measure the distribution stability of a variable in different datasets. A too high PS I value (such as greater than 0.25) may indicate that the variable is unstable in terms of time or sample changes, having a negative impact on the model's generalization ability.

[0089] Generally speaking, through the analysis operations of any one or several of the above implementation methods, it can be determined which feature variables have strong predictive ability for the target variable, and the most informative features can be selected for developing the label.

[0090] Step 204: Delete the failure variable data in the original sample data to obtain the sample data.

[0091] In the overdue risk identification scenario, determining the ineffective variable data based on the index statistical information is an important part of variable screening. Among them, the determined ineffective variable data refers to the variable information with low contribution to the model prediction ability or no significant correlation with the target variable. The existence of this ineffective variable data will not only increase the model complexity and introduce noise, but even affect the optimization performance of the model. Therefore, after determining the ineffective variable data in the original sample data, it can be removed, so that the sample data without ineffective variable data can be obtained. In this way, when optimizing the overdue risk model based on the sample data, the optimization quality of the overdue risk model can be improved to a certain extent, and the accuracy and stability of the model can be enhanced.

[0092] In this embodiment, by obtaining the original sample data for model optimization operations, performing data statistics on the variable indicators of the original sample data, obtaining the index statistical information of the variable indicators, and determining and deleting the ineffective variable data in the original sample data based on the index statistical information, the sample data for optimizing the model can be obtained. In this way, the interference of the ineffective variable data in the sample data to the model optimization operation is avoided, and the calculation redundancy can be reduced. Moreover, the deletion process must strictly follow the data cleaning rules to ensure that the remaining data meets the integrity and consistency requirements.

[0093] Figure 4 It is a flow chart showing the determination of the variable indicators in the original sample data provided by the embodiment of the present invention; on the basis of the above embodiment, continue to refer to the appendix Figure 4 As shown, for the variable indicators included in the original sample data, they can be determined not only by the numerical change characteristics in the original sample data, but also by deriving the variable indicators in the original sample data. At this time, the determination of the variable indicators in the original sample data in this embodiment may include:

[0094] Step 301: Obtain the existing variable indicators in the original sample data.

[0095] Step 302: Perform a derivation operation on the existing variable indicators to obtain derived variable indicators.

[0096] In the overdue risk identification scenario, the quality and quantity of variable indicators directly determine the prediction ability of the model. When the variable indicators in the original sample data are not sufficient to comprehensively describe the user's behavior characteristics and risk characteristics, by performing a derivation operation on the existing variable indicators in the original sample data, where the existing variable indicators can be determined by the change characteristics of the numerical values, more efficient characteristic variables can be mined based on the existing variable indicators. In this way, when optimizing the model by combining the derived characteristic variables and the existing variable indicators, it is beneficial to improve the discrimination ability and prediction effect of the model.

[0097] In some instances, the derivative operation on existing variable metrics can be achieved through preset rules or pre-trained deep learning models. At this time, performing a derivative operation on existing variable metrics to obtain derivative variable metrics may include: obtaining the preset rules or pre-trained deep learning models for implementing the metric derivative operation; using the preset rules or deep learning models to perform derivative processing on the existing variable metrics, so as to obtain derivative variable metrics, and the obtained derivative variable metrics can be one or more.

[0098] In other instances, the derivative operation on existing variable metrics can be achieved through the variable type and the corresponding variable derivative method. At this time, performing a derivative operation on existing variable metrics to obtain derivative variable metrics may include: obtaining the variable type of the existing variable metrics; determining the variable derivative method for performing the derivative operation on the existing variable metrics according to the variable type; using the variable derivative method to perform the derivative operation on the existing variable metrics to obtain derivative variable metrics.

[0099] The existing variable metrics in the original sample data are the basis for model construction. These existing variables usually come from users' basic information, transaction records, historical behavior data, etc., such as users' age, income level, repayment times, overdue days, etc. The extraction and analysis of existing variable metrics require a full understanding of the variable types, such as numerical type, categorical type, time type, etc., and determining their potential value in the model in combination with the business objectives. Since different types of variable metrics can have different statistical characteristics and processing methods, for example, numerical variables are suitable for mathematical operations and polynomial expansion, while categorical variables are more suitable for encoding or grouped statistical analysis, etc. Therefore, in order to accurately perform derivative operations on existing variable metrics, the variable type of the existing variable metrics can be obtained first, and this variable type can be determined by the data format or data type of the existing variable metrics.

[0100] After obtaining the variable type, the variable derivative method for performing the derivative operation on the existing variable metrics can be determined through the preset mapping relationship between the variable type and the variable derivative method. Among them, the derivative variable is a new feature obtained by performing mathematical transformation, interaction operation, feature extraction, etc. on the existing variable metrics. The purpose of the derivative operation is to expand the feature space of the data and enhance the expression ability of the model. In some instances, the variable derivative method may include at least one of the following: mathematical transformation, aggregation statistic, time feature, text feature, interaction feature, binning processing, missing value processing, polynomial feature, encoding processing, embedding variable, tree model feature.

[0101] Among them, mathematical transformation can perform addition, subtraction, multiplication, division, logarithmic transformation, square root transformation, etc. on numerical variables to narrow the value range of variables or enhance their discrimination ability. Aggregate statistics can calculate the mean, standard deviation, maximum value, minimum value, etc. for time series or grouped data to extract the overall trend of user behavior characteristics. Time feature extraction can extract specific information from time variables, such as extracting year, month, day, hour, or working day / holiday, etc. Text features can extract keywords, word frequencies, sentiment scores, etc. from text data. Interaction features can generate cross-products or ratios between variables to capture the non-linear relationships between variables. Binning can discretize continuous variables into intervals. For example, age can be divided into intervals such as "18 - 25" or "36 - 45". Missing value processing can convert the missing information of variables into new features. For example, for the variable "education level" with missing data, a new feature "education level missing flag" can be added to indicate that the user may have a higher risk due to missing education information. Polynomial features can generate polynomial features such as quadratic and cubic terms of variables. Encoding processing can perform one-hot encoding, target encoding, etc. on categorical variables. Encoded variable generation can extract features from high-dimensional data through neural networks; tree model features can use decision tree models to generate features. For example, leaf node encodings of each sample can be generated through xgboost training and converted into one-hot encoded features.

[0102] By performing derivative operations on the existing variable indicators in one or any combination of the above ways, new efficient features, namely derivative variable indicators, can be systematically mined, which increases the number of variable indicators to a certain extent. When combining derivative variable indicators and existing variable indicators for model optimization operations, it is beneficial to improve the prediction ability and stability of the model in the overdue risk identification scenario.

[0103] Step 303: Determine the existing variable indicators and derivative variable indicators as the variable indicators in the original sample data.

[0104] Among them, existing variable indicators refer to the variables directly extracted from the original sample data, and these variables usually have a high business relevance. For example, in overdue risk identification, common existing variables include the user's age, gender, income level, loan amount, number of repayment times, and historical overdue records, etc. Derivative variable indicators refer to those generated by performing mathematical transformation, interaction operations, binning, etc. on existing variable indicators, aiming to expand the data feature space. Derivative variables often reveal the complex relationships or non-linear characteristics between existing variables and can significantly improve the prediction ability of the model. For example, the user's "debt ratio" and "month at application" are commonly used derivative variables, which respectively reflect the user's repayment ability and behavior habits.

[0105] In the overdue risk identification scenario, the variable metrics include two categories: the existing variable metrics directly obtained from the original data and the derived variable metrics generated through derivation operations. To ensure that these features can be comprehensively utilized for model optimization operations, the existing variable metrics and the derived variable metrics can be integrated into the variable metrics in the final original sample data for subsequent model training and optimization.

[0106] In this embodiment, by obtaining the existing variable metrics in the original sample data, performing derivation operations on the existing variable metrics to obtain the derived variable metrics, and then determining the existing variable metrics and the derived variable metrics as the variable metrics in the original sample data, to a certain extent, the diversity and sufficient quantity of the variable metrics are ensured, which is beneficial to improving the quality and efficiency of subsequent model training and optimization using the variable metrics.

[0107] Figure 5 It is a schematic flow diagram for determining invalid variable data based on screening index values provided by an embodiment of the present invention; on the basis of the above embodiment, continue to refer to the appendix Figure 5 As shown, this embodiment provides an implementation method for determining invalid variable data in combination with the screening parameter type. At this time, determining invalid variable data based on the screening index values in this embodiment may include:

[0108] Step 401: Obtain the screening parameter type for performing screening operations on the original sample data.

[0109] The original sample data usually contains a large number of variables or features, but not all variables contribute significantly to the prediction ability of the target variable. Some variables may be redundant, unstable, or introduce noise. Therefore, screening out valuable variables for the model and eliminating invalid variables are key steps in model optimization and feature selection.

[0110] In order to accurately determine the invalid variable data in the original sample data, the screening parameter type for performing screening operations on the original sample data can be obtained first. In some instances, the screening parameter type can be obtained through human-computer interaction operations, or the screening parameter type can be the default parameters pre-configured by the system, etc. In some instances, the screening parameter type may include at least one of the following:

[0111] The null value rate used to measure the proportion of missing values in the variable and judge the integrity of the variable;

[0112] The maximum frequency ratio used to measure the proportion of a certain value in the variable and judge whether the variable has significant skewness;

[0113] The weight of evidence (WOE) used to evaluate the discrimination ability and interpretability of the variable for the target variable;

[0114] Population Stability Index (PSI) for evaluating the stability of variables across different datasets or time periods;

[0115] Information Value (IV) for measuring the predictive ability of a variable with respect to a target variable;

[0116] Correlation parameters for measuring the linear relationship between a variable and a target variable, which may specifically include: Pearson correlation coefficient, Spearman rank correlation coefficient;

[0117] Multicollinearity parameters for detecting high correlations between variables, which may specifically include: variance inflation factor, etc. The selection of these parameter types can be combined with business requirements, model types, and data characteristics to ensure that the selected variables play a positive role in model optimization. Moreover, for any selected parameter type, there is a series of indicators and methods for evaluating the effectiveness of feature variables.

[0118] Step 402: Determine the screening parameter values of variable indicators based on the index statistical information and screening parameter types.

[0119] Step 403: Determine the failed variable data in the original sample data based on the screening index values.

[0120] After clarifying the screening parameter types, the corresponding screening parameter values can be automatically calculated based on the index statistical information of the original sample data, and then the screening index values of variable indicators can be analyzed to determine the failed variable data in the original sample data.

[0121] Example 1, when the screening parameter type is the null value rate, the null value rate refers to the proportion of the number of samples with missing values in a variable to the total sample size. Variables with a high null value rate may be incomplete in the data collection process or have a low correlation with the business objective. Usually, a threshold is set (generally set to 60%), and when the null value rate of a variable indicator exceeds this threshold, it is considered that the variable indicator does not have sufficient information and may affect the training and prediction of the model, so it needs to be excluded. Specifically, when dealing with variable indicators with a high null value rate, one can choose to directly remove them or fill in the missing values. If the missing values have specific business meanings, the missing information can be converted into new features. For example, if the "educational information" is missing, a new feature or flag such as "educational missing flag" can be added.

[0122] Example 2. When the screening parameter type is the maximum frequency ratio, the maximum frequency ratio refers to the sample proportion of the value with the highest occurrence frequency in the variable index. This index reflects the distribution concentration of the variable index. If the maximum frequency ratio of a certain variable index exceeds the set threshold (usually set at 60%), it indicates that the values of this variable index are too concentrated and it is difficult to provide effective discrimination ability, which may lead to a decline in the discrimination effect of the model on the samples. For example, if "Beijing" accounts for 60% in "city where the user is located" and the sample proportion of other cities is relatively small, this may cause the model to give the same result for most samples during prediction. For such variable indexes that exceed the preset maximum frequency ratio, they can be directly determined as invalid variable data, or they can be binned, or potential information can be mined through interaction features.

[0123] Example 3. When the screening parameter type is the Weight of Evidence (WOE), since WOE is an index used to measure the discrimination ability of each bin (or group) in a variable for the target variable (such as overdue or not), its core idea is to evaluate the prediction ability of this bin for the target variable by performing logarithmic transformation on the proportion of good samples and the proportion of bad samples in each bin. Among them, the proportion of good samples can represent the proportion of good samples in the current bin among all good samples, and the proportion of bad samples represents the proportion of bad samples in the current bin among all bad samples. The WOE value of each bin reflects the discrimination ability of this bin for the target variable. If the WOE value is positive, it means that the proportion of bad samples in the current bin is relatively high, indicating that the characteristics of this bin tend to be overdue defaults; if the WOE value is negative, it means that the proportion of good samples in the current bin is relatively high, indicating that the characteristics of this bin tend to be normal repayments; if the WOE value shows a distribution with positive at both ends and negative in the middle, it indicates that this feature may have strong discrimination ability for extreme samples, and it is necessary to further judge whether to include it in the model in combination with the business background. In addition, the WOE value can not only measure the discrimination ability of a single bin, but also evaluate the prediction ability of the entire variable index in combination with other types of screening parameters (such as: IV value) to accurately identify invalid variable data.

[0124] Example 4. When the screening parameter type is the Information Value (IV), since IV is an important index to measure the degree of association between a characteristic variable and the target variable and is used to evaluate the prediction ability of the variable index, the IV value can be calculated by combining the value distribution of the variable index and the WOE value. By calculating the IV value, variable indexes with strong prediction ability for the target variable can be effectively screened. Therefore, invalid variable data can be screened based on the IV value and the preset threshold, and removing variable indexes with lower IV values can reduce the redundant features of the model.

[0125] Example 5. When the screening parameter type is the Population Stability Index (PSI), PSI is an index that measures the distribution difference of model input variables or prediction results between different datasets or different time periods. Specifically, it can measure the stability of the model on different datasets or time periods by comparing the distribution differences of variable indicators on the training set and the test set (or different time periods), and can determine whether there are data drift or stability problems in the model. Among them, data drift may lead to a decline in the model's prediction performance. Therefore, the PSI value is an important reference index in model performance monitoring. In the embodiments of the present invention, the consistency of certain variable indicators on the training and test datasets can be judged through the PSI value, and the variable indicators with too large distribution changes can be eliminated.

[0126] Example 6. When the screening parameter type is a correlation parameter, it can be specifically implemented as the Pearson correlation coefficient and / or the Spearman rank correlation coefficient. Among them, the Pearson correlation coefficient is used to measure the linear relationship between two variable indicators and is applicable to continuous variable indicators; the Spearman rank correlation coefficient is used to measure the monotonic relationship between two variable indicators and is applicable to categorical variable indicators or data with non-linear relationships. If the feature variable and the target variable show a strong correlation, it indicates that the feature variable has a high prediction ability for the target variable. Then, the above-mentioned feature variable is determined as valid variable data, otherwise it is invalid variable data.

[0127] In the embodiments of the present invention, feature variables that may be related to the target variable can be selected first. In the overdue risk identification scenario, common feature variables may include at least one of the following: loan amount, number of repayments, income level. Then, for each feature variable, the Pearson correlation coefficient and the Spearman rank correlation coefficient with the target variable are calculated respectively. If the results of the Pearson correlation coefficient and the Spearman rank correlation coefficient show that the feature quantity and the target variable show a strong correlation, the performance indicators of the model are observed to judge the actual effect of the variable indicator on the model performance. If the variable indicator contributes limitedly to the improvement of the model performance in multiple iterations, the variable indicator can be considered for elimination.

[0128] Example 7, when the screening parameter type is a multicollinearity parameter, it can be specifically implemented as a correlation coefficient matrix or a variance inflation factor (VIF). The correlation coefficient matrix or VIF can identify the high correlation existing between feature variables. This relationship will make it difficult for the model to accurately estimate the independent effects of each variable index during the training process, and will make it difficult to distinguish which variable index has a greater impact on the target result and lead to unstable model parameter estimation, resulting in large fluctuations due to small input changes. For example, when it is necessary to analyze the correlation between "the number of times of applying for a bank loan in the past 1 month", "the number of times of applying for a bank loan in only 3 months", and "the number of times of applying for a bank loan in only 6 months". Then, it is possible to simply judge whether there is a problem of multicollinearity between feature variables through a correlation coefficient matrix or a variance inflation factor. If there are multiple collinear variable indicators, variable indicators with higher IV values can be selected, or further constraints can be imposed on the collinear variables through feature selection or regularization methods to automatically compress or eliminate redundant invalid variable indicators.

[0129] Specifically, the existing invalid variable data can be processed by the regularization method of the L1 penalty term. Among them, the L1 penalty term can be used to sparsify feature weights. It introduces the L1 norm (i.e., the sum of the absolute values of feature weights) as a penalty term in the loss function, making some feature weights in the model become zero. In this way, those feature weights that do not make significant contributions to the model prediction can be set to zero, thus achieving the effect of feature selection.

[0130] Alternatively, the invalid variable data can also be processed by forward and backward screening. This forward and backward screening is a method of gradually constructing and pruning a feature subset. It gradually adds or deletes features until an optimal feature subset is found. Forward and backward screening is a method of gradually optimizing a feature subset, used to screen out the optimal feature subset from the original feature set. Its purpose is to find a feature combination that reduces redundant features while maintaining or improving the model prediction performance by dynamically adjusting the feature set. Among them, forward screening means starting from an empty feature set and adding one feature that is most relevant to the target variable each time, gradually expanding the feature set; backward screening means starting from the complete feature set containing all features and deleting one feature that contributes the least to the target variable each time, gradually shrinking the feature set. Both methods improve the model performance through feature selection, while reducing the number of features to reduce complexity and enhance the generalization ability of the model.

[0131] In this embodiment, by eliminating invalid, unstable, or highly correlated failure variable data with other features, the stability and prediction ability of the model can be significantly improved, while reducing the model complexity. In the overdue risk identification scenario, this process can ensure that the model is more accurate and reliable, providing a solid foundation for subsequent model optimization and application.

[0132] Figure 6 This is a schematic flowchart of determining optimized hyperparameters based on the simulated annealing algorithm provided by the embodiments of the present invention; on the basis of the above embodiments, continue to refer to the attached Figure 6 As shown, the determination of optimized hyperparameters based on the simulated annealing algorithm in this embodiment may include:

[0133] Step 501: Based on sample data, at least one target metric, and perform iterative calculations on at least one hyperparameter to be optimized within the iterative parameter range according to the simulated annealing algorithm to obtain at least one intermediate hyperparameter.

[0134] In the overdue risk identification scenario, in order to build an efficient and accurate overdue risk model, it is necessary to perform iterative calculations on the hyperparameters of the overdue risk model to continuously optimize its performance on the target metrics. Among them, the target metrics may include at least one of the following: accuracy rate, recall rate, AUC value, KS value, LIFT value, etc. Hyperparameter optimization based on the simulated annealing algorithm is to simulate the energy change during the cooling process of a physical system, gradually search for the global optimal parameter combination, and thus obtain the model with the best performance.

[0135] First, before hyperparameter optimization, it is necessary to construct a basic model using the existing sample data and randomly initialize the hyperparameters of the model. An automated tuning process will be performed each time a basic model is selected, and the optimal model can be selected for application. At this time, the initial values of the hyperparameters may not be in the optimal state and only serve as the starting point of the simulated annealing algorithm. In the business scenario of overdue risk identification, the sample data may include at least one of the following: user data, income data, consumption data, credit data, etc. These data are used to train the model, and the discrimination ability of the model is measured by the target metrics. To ensure the rationality of the search range, an empirical search range is set for each hyperparameter during the optimization process. For example, taking the XGBoost model as an example, the learning rate can be set within the search range of 0.01 to 0.3, the maximum depth can be set within the search range of 3 to 10, and the ranges of 0 to 1 are set for the regularization parameters λ and α respectively.

[0136] In the initial stage of optimization, the simulated annealing algorithm builds a model based on the current hyperparameter settings and calculates the target metric value on the test set. Then, it randomly generates a new combination of hyperparameters according to the simulated annealing algorithm and uses these new parameters to train the model, recalculating the target metric value. If the target metric value of the new parameters is better than that of the current parameters, the new parameters are directly accepted as the new intermediate hyperparameters; if the target metric value of the new parameters is worse, it is decided whether to accept the new parameters according to a certain probability. This probability is determined by a control factor and gradually decreases as the number of iterations increases. This design allows the algorithm to explore more possibilities in the initial stage and avoid falling into local optima; while in the later stage, it gradually converges to a better solution and focuses on the accuracy of the optimization process.

[0137] For example, in the optimization process of the overdue risk identification model, the initially set learning rate is 0.1, the maximum depth is 6, and the subsampling rate is 0.8. The overdue risk identification model is trained based on these parameters and obtains an AUC value of 0.78 on the test set. The simulated annealing algorithm then randomly generates a new combination of parameters. For example, the learning rate is adjusted to 0.15, the maximum depth is adjusted to 7, and the subsampling rate is adjusted to 0.85. The AUC value of the trained model on the test set is increased to 0.81, so these new parameters are accepted as the intermediate hyperparameters and enter the next iteration. In another iteration, the new combination of hyperparameters may cause the AUC value to drop to 0.76, but due to being in the initial stage of optimization, the algorithm may still accept these parameters with a certain probability to explore a wider solution space. As the number of iterations increases, the algorithm gradually reduces the acceptance probability of parameters with performance degradation, making the optimization process gradually converge to the global optimal parameters.

[0138] In the overdue risk identification scenario, the model optimized by the simulated annealing algorithm can more accurately distinguish between users with normal repayments and users who may be overdue, improving metrics such as the AUC value, and ultimately providing a more reliable risk assessment tool for users.

[0139] Step 502: Determine the model performance difference caused by the intermediate hyperparameters and the corresponding hyperparameters to be optimized for the overdue risk model.

[0140] Among them, after obtaining the intermediate hyperparameters, the hyperparameters to be optimized and the intermediate hyperparameters can be analyzed and processed to determine the model performance differences constituted by the intermediate hyperparameters and the hyperparameters to be optimized for the overdue risk model. In some examples, determining the model performance differences constituted by the intermediate hyperparameters and the corresponding hyperparameters to be optimized for the overdue risk model may include: determining the intermediate overdue risk model corresponding to the intermediate hyperparameters; respectively using the overdue risk model and the intermediate overdue risk model to analyze and process the same set of data, obtaining the first processing result output by the overdue risk model and the second processing result output by the intermediate overdue risk model; by analyzing the first processing result and the second processing result, the model performance differences can be determined.

[0141] In other examples, determining the model performance differences constituted by the intermediate hyperparameters and determining the hyperparameters to be optimized for the overdue risk model may include: determining the first objective function value based on the intermediate hyperparameters and the target metrics; determining the second objective function value based on the target metrics and the hyperparameters to be optimized corresponding to the intermediate hyperparameters; determining the difference between the first objective function value and the second objective function value as the model performance difference.

[0142] First, during the optimization process of the simulated annealing algorithm, the initial model parameters are set to the "initial temperature" T 0 , that is, the hyperparameters to be optimized. These hyperparameters may include at least one of the following: learning rate, maximum depth, sample sampling rate, etc. For example, in the model training for overdue risk identification, assume that the initial intermediate hyperparameters are set to a learning rate of 0.1, a maximum depth of 6, and a sample sampling rate of 0.8. Train the model with these parameters and calculate its target metrics on the test set, and use this as the "ending temperature" T k , that is, the first objective function value.

[0143] Next, let T = T 0 = kT, where k is a constant between 0 and 1 used to control the annealing rate. The smaller k is, the slower the annealing speed, and the easier it is to find the optimal solution, but at the same time, the calculation time will also be longer. At the same time, the simulated annealing algorithm will apply a random perturbation to the current solution x t , that is, the intermediate hyperparameters. Randomly select a new combination of hyperparameters to be optimized within the parameter space. For example, the learning rate changes from 0.1 to 0.15, and the maximum depth changes from 6 to 7, and retrain the model based on the new parameters and calculate its target metric value. This value is defined as the second objective function value.

[0144] Then, the algorithm calculates the model performance difference, that is, the difference ΔE between the first objective function value and the second objective function value can be expressed by the following formula:

[0145] ΔE = E(x t+1 ) - E(x t )

[0146] Among them, E(x t ) is the first objective function value corresponding to the hyperparameter to be optimized, and E(x t+1 ) is the second objective function value corresponding to the intermediate hyperparameter.

[0147] For example, in the overdue risk identification scenario, assume that the intermediate hyperparameters of the current model are set to a learning rate of 0.1, a maximum depth of 6, and a sample sampling rate of 0.8, and the corresponding AUC value is 0.78, which is the first objective function value. The simulated annealing algorithm randomly perturbs the hyperparameters, adjusts the learning rate to 0.15, and the maximum depth to 7, and the corresponding AUC value becomes 0.81, which is the second objective function value. The performance difference is calculated as:

[0148] ΔE = 0.81 - 0.78 = 0.03

[0149] By calculating the difference between the first objective function value and the second objective function value, the simulated annealing algorithm can quantify the model performance difference between the current solution and the new solution, and decide whether to accept the new solution based on the difference.

[0150] Step 503: Determine the optimized hyperparameters based on the model performance difference and the intermediate hyperparameters.

[0151] Provide another implementation method for determining the optimized hyperparameters. The optimized hyperparameters can be directly determined through the performance difference result and the preset mapping rule.

[0152] In some other instances, the optimized hyperparameters can also be determined through the iteration termination condition. At this time, determining the optimized hyperparameters based on the model performance difference and the intermediate hyperparameters can include: when the model performance difference is greater than zero, the intermediate hyperparameters are determined as the iteration hyperparameters for the next iteration calculation; iteratively calculate the iteration hyperparameters for the sample data, at least one target metric, and within the iteration parameter range according to the simulated annealing algorithm to obtain the optimized hyperparameters, and the optimized hyperparameters satisfy the preset iteration termination condition.

[0153] The simulated annealing algorithm can continuously iterate and adjust the hyperparameters of the model to seek the hyperparameter combination that maximizes the target metric. When the calculated model performance difference is greater than zero, it means that the new solution is better than the current solution, and the new solution is accepted as the iteration hyperparameter for the next iteration; moreover, the optimization process also sets an iteration termination condition, and the iteration termination condition can be reaching the maximum number of iterations, the objective function value no longer changing, reaching the preset objective function value, etc. When the termination condition is met, the algorithm determines the current iteration hyperparameters as the optimized hyperparameters.

[0154] In addition, at this time, based on the model performance difference and intermediate hyperparameters, determining the optimized hyperparameters may include: when the model performance difference is less than zero, determining the probability information of taking the intermediate hyperparameters as the hyperparameters for the next iterative calculation; based on the probability information, intermediate hyperparameters, and hyperparameters to be optimized, determining the iterative hyperparameters for the next iterative calculation; performing iterative calculation on the sample data, at least one target metric, and the iterative hyperparameters within the iterative parameter range according to the simulated annealing algorithm to obtain the optimized hyperparameters, and the optimized hyperparameters satisfy the preset iterative termination condition.

[0155] When the model performance difference is less than zero, that is, the new hyperparameters do not significantly improve the model performance, the simulated annealing algorithm does not immediately reject the new parameters. Instead, it calculates the probability information to decide whether to accept the current solution as the hyperparameters for the next iteration. The probability information is calculated based on the acceptance rule of the simulated annealing algorithm. Specifically, the probability information is determined based on the acceptance probability formula of the simulated annealing algorithm. The probability information gradually decreases as the number of iterations increases. This design ensures that the algorithm explores more solution spaces in the early stage and converges to the global optimum in the later stage. Among them, the formula for the probability information P is determined by the following formula:

[0156]

[0157] Once the acceptance probability information is determined, the simulated annealing algorithm combines this probability with the intermediate hyperparameters and the hyperparameters to be optimized to generate the iterative hyperparameters for the next iterative calculation. The new iterative hyperparameters can be randomly selected between the current solution and the new solution, gradually converging to the global optimum while retaining the exploration space. In the next step, the algorithm recalculates the objective function value for the sample data and target metrics using the new iterative hyperparameters, updates the model performance, and determines whether the preset termination condition is satisfied. When the termination condition is met, the simulated annealing algorithm outputs the final optimized hyperparameters.

[0158] For example, the intermediate hyperparameters of the current model are a learning rate of 0.1 and a maximum depth of 6, and the corresponding target metric (such as the AUC value) is 0.82. The newly generated hyperparameters to be optimized are: a learning rate of 0.12 and a maximum depth of 7, and the corresponding AUC value is 0.8. The performance difference ΔE = 0.80 - 0.82 = -0.02. At this time, the algorithm calculates the probability of accepting the new solution as If kT = 0.1 at this time, the probability P is relatively high, and the algorithm is more inclined to accept the new solution; if kT = 0.01 at this time, the probability P is relatively low, and the new solution is more difficult to be accepted.

[0159] In this embodiment, through the above optimization process, the simulated annealing algorithm can balance exploration and convergence in a large search space, ensuring that the globally optimal hyperparameter combination is finally found, that is, to a certain extent, ensuring the accuracy and reliability of determining the optimized hyperparameters. In this way, in the overdue risk identification scenario, the prediction performance and stability of the model can be significantly improved.

[0160] Figure 7 It is a schematic flowchart of the process for associating and displaying the index values and the reasons for model differences provided by the embodiments of the present invention; on the basis of the above embodiments, continue to refer to the attached Figure 7 As shown, after determining the optimized overdue risk model, the index values and the reasons for model differences can be associated and displayed. Specifically, the method in this embodiment may include:

[0161] Step 601: Obtain model evaluation metrics for evaluating the model performance.

[0162] Step 602: Determine the first index value of the overdue risk model for the model evaluation metric and the second index value of the optimized overdue risk model for the model evaluation metric.

[0163] Step 603: Based on the first index value and the second index value, determine the reasons for the model differences between the optimized overdue risk model and the overdue risk model.

[0164] Step 604: Associatively display the model evaluation metric, the first index value, the second index value, and the reasons for the model differences.

[0165] To ensure the continuous optimization and improvement of the model performance, it is necessary to conduct a detailed performance comparison and difference analysis of different versions of the model. This process is achieved by obtaining key model evaluation metrics, comparing the performance changes of the model before and after optimization, and analyzing the reasons for model differences.

[0166] In overdue risk identification, the commonly used model evaluation metrics may include the KS value, accuracy, recall rate, and F1 score. Among them, the KS value can reflect the ability of the model to distinguish good and bad samples; the accuracy can represent the overall prediction accuracy of the model; the recall rate can be used to measure the ability of the model to identify overdue users; the F1 score can be used to balance the recall rate and the precision rate. Users can flexibly configure the model evaluation metrics that need to be obtained according to the application scenario or design requirements. These metrics can comprehensively measure the performance of the model, ensuring that the optimized overdue risk model can meet the business requirements in different dimensions.

[0167] For the overdue risk model before optimization, its evaluation metric values (i.e., the first metric values) on the test set can be recorded. For example, the KS value is 0.78, the accuracy rate is 85%, the recall rate is 70%, and the F1 score is 0.75. For the optimized model, its evaluation metric values (i.e., the second metric values) are also calculated. For example, the KS value is improved to 0.82, the accuracy rate is 87%, the recall rate is 72%, and the F1 score is 0.78. By comparing these metrics, the improvement extent of the optimized model in prediction performance can be intuitively reflected.

[0168] By analyzing and comparing the obtained first metric values and second metric values, the reasons for the model differences between the models can be obtained. Specifically, the reasons for the model differences can be determined by analyzing and processing the first metric values and second metric values through a pre-trained neural network model. The reasons for the model differences can be reflected by at least one of the following parameters: hyperparameter configuration, model weights, training data, etc. For example, it may be found that the optimized model uses a deeper network structure (the maximum depth increases from 6 to 8), adjusts the learning rate (changes from 0.1 to 0.05), or increases the regularization strength (the L1 and L2 penalty coefficients change). These changes in structure and parameters directly affect the performance of the model.

[0169] To enhance the intuitiveness and interpretability of the comparison, the optimization results can be displayed in an associated manner. By constructing a data visualization interface, the model evaluation metrics, the first metric values, the second metric values, and the difference analysis results are integrated and displayed. For example, a line chart can be used to show the performance change trends of multiple model versions to help users understand whether the optimized model continues to improve; a table can be used to list the differences between the models before and after optimization in terms of variable selection, weight adjustment, hyperparameter configuration, etc. to clarify the improvement points. At the same time, provide an explanation of the reasons for the difference analysis, such as "the optimized model uses a lower learning rate, resulting in more stable training convergence, thus improving the KS value"; or "removing the low-correlation variable user device model reduces data noise and improves the F1 score".

[0170] In this embodiment, through the above means, automatic comparison of model performance, comparison of historical versions can be achieved, and the comparison results can be displayed or output. This not only facilitates users to timely understand the differences between models, but also helps to track the improvement and optimization of models, and improve the reliability, accuracy, and efficiency of models.

[0171] In specific applications, such as Figure 8 shown, this application embodiment provides a method for implementing an optimized overdue risk model, and the method may include:

[0172] Step 1: Data import.

[0173] Step 2: Conduct variable statistical analysis on the imported original data to obtain statistical data of existing variables.

[0174] Step 3: Perform variable derivation operations on the existing variables to obtain derived variables.

[0175] Step 4: Conduct variable validity analysis on the existing variables and the derived variables to obtain sample data for optimizing the overdue risk model.

[0176] Step 5: Select a model algorithm and define an objective function.

[0177] Step 6: Generate a parameter search range for the hyperparameters.

[0178] Step 7: Perform iterative calculation operations on the hyperparameters of the overdue risk model based on the model algorithm, the objective function, the generated parameter search range, and the simulated annealing algorithm to achieve parameter optimization operations.

[0179] Among them, the specific implementation methods and implementation effects of the above steps in this application embodiment are similar to the similar steps or the same specific implementation methods and implementation effects in the above embodiments. For details, reference can be made to the above statements, which will not be elaborated here.

[0180] Step 8: When the hyperparameters after iterative calculation reach the iterative termination condition and the algorithm termination condition, the optimized hyperparameters are obtained.

[0181] Specifically, the hyperparameters can be optimized through the simulated annealing algorithm. This algorithm sets the initial temperature and the objective function value, randomly perturbs the hyperparameters, and gradually searches the parameter space to find the hyperparameter combination that maximizes the objective function value. In each iteration, the model is trained according to the new hyperparameters and the target metric value is calculated. If the new metric value is better than the current value, the hyperparameters are accepted as the new intermediate hyperparameters; otherwise, it is decided whether to retain the current solution through the acceptance probability formula. As the number of iterations increases, the algorithm gradually converges, and finally the globally optimal hyperparameter combination is found, that is, the optimized hyperparameters can be accurately determined within the parameter search range. After obtaining the optimized hyperparameters, an optimized overdue risk model can be determined based on the optimized hyperparameters. In this way, when performing overdue risk identification operations based on the overdue risk model, the accuracy of overdue risk identification can be improved to a certain extent.

[0182] Step 9: Automatically compare the differences between the old and new versions of the overdue risk model and determine whether it meets the preset requirements for model optimization operations.

[0183] Specifically, after obtaining the optimized overdue risk model, the optimized model will be comprehensively evaluated on the test set, the performance metrics after optimization will be recorded, and compared with the initial model. If the optimization effect is significant, it will enter the model difference analysis stage; if the performance does not meet the expectations, the process will return to the parameter optimization stage to re-adjust the parameter search range or the model structure. In the difference analysis stage, by comparing the parameter configurations, feature importances, and model structures of the new and old models, the reasons for the performance improvement are clarified, and targeted optimization suggestions are generated based on this. Finally, the optimization results are presented through a model report, clearly presenting the performance comparison, contribution analysis of key features, and optimization suggestions to the user, and integrating the optimization results into the continuous integration / continuous deployment process to achieve automated model evaluation and deployment.

[0184] In the embodiment of this application, the overdue risk model can continuously perform automatic optimization, continuously improve the ability to identify users' overdue risks, reduce the impact of manual operations on overdue risk identification, and improve the practicability of the overdue risk identification model; in addition, by obtaining the overdue optimization effect, analyzing the reasons for model differences, and generating optimization suggestions, it is possible to efficiently adjust the sample data, target metrics, and iterative parameter range. This process not only improves the performance and stability of the model but also provides an actionable direction for model optimization, further strengthening the overdue risk identification ability.

[0185] Figure 9 It is a schematic structural diagram of an overdue risk identification device provided by an embodiment of the present invention; refer to the appendix Figure 9 As shown: This embodiment provides an overdue risk identification device, and this overdue risk identification device can execute the above-mentioned Figure 2 corresponding overdue risk identification method. Specifically, this overdue risk identification device may include:

[0186] The first acquisition module 11 is used to acquire the overdue risk model to be optimized and the sample data for model optimization operations;

[0187] The first determination module 12 is used to determine at least one hyperparameter to be optimized for the overdue risk model and at least one target metric corresponding to the sample data;

[0188] The first calculation module 13 is used to perform iterative calculations on at least one hyperparameter to be optimized within the iterative parameter range based on the sample data, at least one target metric, and according to the simulated annealing algorithm to obtain at least one optimized hyperparameter;

[0189] The first processing module 14 is used to determine the optimized overdue risk model based on at least one optimized hyperparameter, perform overdue risk identification operations using the optimized overdue risk model, obtain the identification result of the overdue risk, and return the identification result.

[0190] Figure 9 The device shown can execute Figures 2 - 8 the method of the embodiment shown. For parts not described in detail in this embodiment, reference can be made to the relevant description of the Figures 2 - 8 embodiment shown. For the execution process and technical effects of this technical solution, refer to the description in the Figures 2 - 8 embodiment shown, which will not be elaborated here.

[0191] In a possible design, Figure 2 the method for identifying the overdue risk shown can be implemented as an electronic device, which can be various devices such as a controller, a personal computer, a server, etc. As Figure 10 shown, the electronic device may include: a first processor 21 and a first memory 22. Among them, the first memory 22 is used to store a program for the corresponding electronic device to execute the method for identifying the overdue risk provided in the Figures 2 - 8 embodiment shown above, and the first processor 21 is configured to execute the program stored in the first memory 22.

[0192] The program includes one or more computer instructions. When one or more computer instructions are executed by the first processor 21, the following steps can be implemented: in response to an identification request for overdue risk sent by a requesting terminal, obtain an overdue risk model to be optimized and sample data for performing model optimization operations, where the sample data includes overdue sample data and non-overdue sample data; determine at least one hyperparameter to be optimized of the overdue risk model and at least one target index corresponding to the sample data, where the hyperparameter to be optimized is within a preset iterative search range; based on the sample data, at least one target index, and perform iterative calculation on at least one hyperparameter to be optimized within the iterative parameter range according to the simulated annealing algorithm to obtain at least one optimized hyperparameter; based on at least one optimized hyperparameter, determine an optimized overdue risk model, use the optimized overdue risk model to perform an identification operation of overdue risk, obtain an identification result of overdue risk, and return the identification result to the requesting terminal.

[0193] Furthermore, the first processor 21 is also used to execute all or part of the steps in the Figures 2 - 8 embodiment shown above.

[0194] Among them, the structure of the electronic device may further include a first communication interface 23 for the electronic device to communicate with other devices or communication networks.

[0195] In addition, an embodiment of the present invention provides a computer storage medium for storing computer software instructions used by an electronic device, which includes a program related to the method for identifying the overdue risk in the method embodiment shown above. Figures 2 - 8

[0196] In addition, an embodiment of the present invention provides a computer program product, including: a computer program which, when executed by a processor of an electronic device, causes the processor to execute the above Figures 2 - 8 steps in the method for identifying overdue risks in the method embodiment shown above.

[0197] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of adding a necessary general hardware platform, and of course, it can also be implemented by a combination of hardware and software. Based on such an understanding, the above technical solutions, in essence, or the part that contributes to the prior art can be embodied in the form of a computer product. The present invention can adopt the form of a computer storage medium implemented on one or more computer-usable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) containing computer-usable program codes.

[0198] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for identifying overdue risk, characterized in that: Applied to an identification terminal, the identification terminal is communicatively connected to a requesting terminal; the method comprises: In response to the overdue risk identification request sent by the request terminal, the identification terminal obtains the overdue risk model to be optimized and sample data for performing model optimization operations, wherein the sample data includes overdue sample data and non-overdue sample data; The identification terminal determines at least one hyperparameter to be optimized of the overdue risk model and at least one target indicator corresponding to the sample data, wherein the hyperparameter to be optimized is within a preset iterative search range; The identification terminal performs iterative calculation on the at least one hyperparameter to be optimized within the iterative parameter range based on the sample data and the at least one target indicator according to a simulated annealing algorithm to obtain at least one optimized hyperparameter; The identification terminal determines an optimized overdue risk model based on the at least one optimized hyperparameter, performs an overdue risk identification operation using the optimized overdue risk model, obtains an overdue risk identification result, and returns the identification result to the requesting terminal.

2. The method according to claim 1, characterized in that Get sample data for model optimization operations, including: Obtaining original sample data for model optimization operations; Performing data statistics on the variable indicators of the original sample data to obtain indicator statistical information of the variable indicators; Based on the indicator statistical information, determining failure variable data in the original sample data; The invalid variable data in the original sample data is deleted to obtain the sample data.

3. The method according to claim 2, characterized in that Get the original sample data for model optimization, including: Acquire at least one data source, wherein the data source includes at least part of the original sample data; Determining a valid time period for collecting the original sample data; Based on the at least one data source and the valid time period, original sample data for performing a model optimization operation is determined.

4. The method according to claim 2, characterized in that: After obtaining the original sample data for performing the model optimization operation, the method further includes: Obtaining existing variable indicators in the original sample data; Performing a derivative operation on the existing variable index to obtain a derived variable index; The existing variable indicators and the derived variable indicators are determined as variable indicators in the original sample data.

5. The method according to claim 4, characterized in that Performing a derivative operation on the existing variable index to obtain a derived variable index includes: Obtaining the variable type of the existing variable indicator; Determine, according to the variable type, a variable derivation method for performing a derivation operation on the existing variable indicator; The variable derivation method is used to perform a derivation operation on the existing variable indicator to obtain the derived variable indicator.

6. The method according to claim 2, characterized in that Determining the failure variable data in the original sample data based on the indicator statistical information includes: Acquire a screening parameter type for performing a screening operation on the original sample data; Determining a screening parameter value of the variable indicator based on the indicator statistical information and the screening parameter type; Based on the screening index value, failure variable data in the original sample data is determined.

7. The method according to claim 6, characterized in that The screening parameter type includes at least one of the following: The null value rate is the ratio of the number of samples with null indicators of the identification variable to the total samples; The maximum frequency ratio of the number of samples with the highest value in the indicator variable to the total number of samples; The weight of evidence and IV value used for the contribution of the identification variable indicator to the prediction of the target indicator; A stability index PSI for measuring the distribution difference between different sample groups in the original sample data; A correlation parameter used to identify the degree of association between the variable indicator and the target indicator; The multicollinearity parameter is used to identify the correlation between any two variable indicators.

8. The method according to claim 1, characterized in that Based on the sample data and the at least one target indicator, the at least one hyperparameter to be optimized is iteratively calculated within the iterative parameter range according to a simulated annealing algorithm to obtain at least one optimized hyperparameter, including: Based on the sample data and the at least one target indicator, the at least one hyperparameter to be optimized is iteratively calculated within the iterative parameter range according to a simulated annealing algorithm to obtain at least one intermediate hyperparameter; Determine the model performance difference between the intermediate hyperparameter and the corresponding hyperparameter to be optimized on the overdue risk model; The optimized hyperparameters are determined based on the model performance difference and the intermediate hyperparameters.

9. The method according to claim 8, characterized in that Determining the model performance difference between the intermediate hyperparameter and the corresponding hyperparameter to be optimized on the overdue risk model includes: Determining a first objective function value based on the intermediate hyperparameter and the target indicator; Determining a second objective function value based on the target indicator and the hyperparameter to be optimized corresponding to the intermediate hyperparameter; A difference between the first objective function value and the second objective function value is determined as the model performance difference.

10. The method according to claim 8, characterized in that Determining the optimized hyperparameter based on the model performance difference and the intermediate hyperparameter includes: When the model performance difference is greater than zero, the intermediate hyperparameter is determined as the iterative hyperparameter for the next iterative calculation; The sample data, the at least one target indicator, and the iterative hyperparameter are iteratively calculated within the iterative parameter range according to a simulated annealing algorithm to obtain the optimized hyperparameter, and the optimized hyperparameter satisfies a preset iteration termination condition.

11. The method according to claim 8, characterized in that Determining the optimized hyperparameter based on the model performance difference and the intermediate hyperparameter includes: When the model performance difference is less than zero, determining the probability information of making the intermediate hyperparameter become the hyperparameter of the next iterative calculation; Determining an iterative hyperparameter for a next iterative calculation based on the probability information, the intermediate hyperparameter, and the hyperparameter to be optimized; The sample data, the at least one target indicator, and the iterative hyperparameter are iteratively calculated within the iterative parameter range according to a simulated annealing algorithm to obtain the optimized hyperparameter, and the optimized hyperparameter satisfies a preset iteration termination condition.

12. The method according to any one of claims 1 to 11, characterized in that: After determining the optimized overdue risk model, the method further includes: Get model evaluation metrics for evaluating model performance; Determining a first indicator value of the overdue risk model for the model evaluation indicator, and a second indicator value of the optimized overdue risk model for the model evaluation indicator; Determining the reason for the model difference between the optimized overdue risk model and the overdue risk model based on the first indicator value and the second indicator value; The model evaluation index, the first index value, the second index value and the reason for the model difference are displayed in association.

13. The method according to claim 12, characterized in that After determining the cause of the model difference between the optimized overdue risk model and the overdue risk model, the method further includes: Obtaining an overdue optimization effect for optimizing the overdue risk model; Based on the overdue optimization effect and the reasons for the model differences, recommendation information for model optimization operations is generated, and the recommendation information is used to assist in determining whether to adjust at least part of the sample data, target indicators, and iterative parameter ranges during the overdue risk model optimization operation.

14. An electronic device, characterized in that: include: A memory and a processor; wherein the memory is used to store one or more computer instructions, wherein the one or more computer instructions, when executed by the processor, implement the method described in any one of claims 1 to 13.

15. A computer storage medium, characterized in that Used to store a computer program, which enables a computer to implement the method of any one of claims 1 to 13 when executed.

Citation Information

Patent Citations

  • A credit risk monitoring method integrating a deep belief network and an isolated forest algorithm

    CN109685653A

  • Small and medium-sized enterprise risk early warning model optimization method and device, equipment and storage medium

    CN113506175A

  • Hyper-parameter determination method and device

    CN113870005A

  • Method, system and device for predicting overdue user of power conversion package and medium

    CN116862078A

  • Credit risk assessment method, device and equipment

    CN117057899A