Method and device for constructing risk prediction model, computer device and storage medium

By binning and spatially dividing the risk prediction variables of financial institutions, selecting target variables, and constructing a risk prediction model, the problem that traditional risk prediction rules cannot be combined with business operations is solved, thus achieving accurate risk prediction and reducing false alarms.

CN117077028BActive Publication Date: 2026-03-20CHINA CONSTRUCTION BANK +1
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-08-24
Publication Date
2026-03-20

AI Technical Summary

Technical Problem

Traditional risk prediction rules cannot be closely integrated with the business operations of financial institutions, resulting in insufficient accuracy of risk prediction, easy false alarms, and impact on business approval efficiency.

Method used

By acquiring sample data of multiple variables related to risk prediction, binning is performed to select target variables that meet the risk prediction capability standards, target binning is determined, sample space is divided, and a risk prediction model is constructed. The risk prediction function of the subspace is then used to predict the risk status.

Benefits of technology

It has improved the accuracy of risk prediction, reduced false alarms, ensured that risk prediction rules are closely aligned with the business needs of financial institutions, and enhanced the scientific nature and efficiency of business management.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117077028B_ABST
    Figure CN117077028B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of big data analysis, in particular to a risk prediction model construction method and device, computer equipment, a storage medium and a computer program product. The method comprises the following steps: acquiring sample data corresponding to a plurality of variables related to risk prediction, performing binning processing on each sample data respectively to obtain a plurality of bins corresponding to each variable; selecting target variables with up-to-standard risk prediction capability from the plurality of variables; taking a bin with the highest abnormal proportion of sample objects in the plurality of bins corresponding to the target variables as a target bin of the target variable; determining a sample space containing each target bin, dividing the sample space into a plurality of subspaces, and determining a risk prediction function of a sample object in each subspace; and constructing a risk prediction model based on the risk prediction functions corresponding to each subspace. The method can construct a risk prediction model that can closely combine the business of a financial institution and accurately predict risks.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of big data analysis, and in particular to a risk prediction model construction method and device, computer equipment, a storage medium and a computer program product. BACKGROUND

[0002] For a financial institution, it is of great significance to accurately predict the risk state (such as the probability of abnormal behavior) of a user, to scientifically and reasonably manage and support the business. In the traditional technology, the financial institution generally directly mines rules related to risk prediction based on historical data to predict the risk state of the user.

[0003] However, only a small part of users usually have abnormal behavior, resulting in less related data of the user having abnormal behavior in the historical data of the financial institution, but more related data of the user not having abnormal behavior. If the rules related to risk prediction are directly mined based on the historical data, it is easy to produce prediction rules biased towards non-risk events, resulting in risk prediction rules that cannot effectively predict risks. In addition, the risk prediction rules of the financial institution are different from general rules, and need to be closely combined with the business of the financial institution to avoid risk false positives and affect the business approval efficiency of the financial institution.

[0004] Therefore, the way of mining risk prediction rules in the traditional technology cannot ensure that the risk prediction rules can be closely combined with the business of the financial institution and accurately predict risks. SUMMARY

[0005] Therefore, it is necessary to provide a risk prediction model construction method, device, computer equipment, computer readable storage medium and computer program product capable of constructing a risk prediction model closely combined with the business of the financial institution and accurately predicting risks, in view of the above technical problems.

[0006] In a first aspect, the present application provides a risk prediction model construction method. The method comprises:

[0007] Obtaining sample data corresponding to each of a plurality of variables related to risk prediction, respectively performing binning processing on each sample data to obtain a plurality of bins corresponding to each variable; and the sample data belongs to different sample objects;

[0008] Selecting target variables with up-to-standard risk prediction ability from the plurality of variables;

[0009] For each target variable, the bin with the highest abnormality proportion of sample objects in the plurality of bins corresponding to the target variable is taken as a target bin of the target variable;

[0010] determine a sample space containing each target bin, divide the sample space into a plurality of subspaces, and determine a risk prediction function of a sample object in each subspace respectively;

[0011] construct a risk prediction model based on the risk prediction function corresponding to each subspace respectively; the risk prediction model is used to determine a target subspace to which a to-be-predicted object belongs, and predict a risk state in which the to-be-predicted object is located based on the risk prediction function of the target subspace.

[0012] In a second aspect, the present application provides a device for constructing a risk prediction model, the device comprising:

[0013] a bin processing module configured to obtain sample data corresponding to each of a plurality of variables related to risk prediction, and perform bin processing on each sample data to obtain a plurality of bins corresponding to each variable respectively; the sample data belongs to different sample objects;

[0014] a variable screening module configured to screen target variables with up-to-standard risk prediction ability from the plurality of variables;

[0015] a bin determination module configured to, for each target variable, determine a target bin of the target variable as a bin in which a sample object abnormality proportion of the target variable is the highest among the plurality of bins corresponding to the target variable;

[0016] a space processing module configured to determine a sample space containing each target bin, divide the sample space into a plurality of subspaces, and determine a risk prediction function of a sample object in each subspace respectively;

[0017] a model construction module configured to construct a risk prediction model based on the risk prediction function corresponding to each subspace respectively; the risk prediction model is used to determine a target subspace to which a to-be-predicted object belongs, and predict a risk state in which the to-be-predicted object is located based on the risk prediction function of the target subspace.

[0018] In a third aspect, the present application further provides a computer device. The computer device comprises a memory and a processor, the memory stores a computer program, and the processor implements the steps of the method for constructing a risk prediction model when executing the computer program.

[0019] In a fourth aspect, the present application further provides a computer readable storage medium. The computer readable storage medium stores a computer program, and the computer program implements the steps of the method for constructing a risk prediction model when executed by a processor.

[0020] In a fifth aspect, the present application further provides a computer program product. The computer program product comprises a computer program, and the computer program implements the steps of the method for constructing a risk prediction model when executed by a processor.

[0021] The method, device, computer device, storage medium and computer program product for constructing a risk prediction model, first acquire sample data corresponding to each of a plurality of variables related to risk prediction, i.e., first closely combine the business of a financial institution to acquire sample data related to risk prediction, then perform binning processing on each sample data to obtain a plurality of bins corresponding to each variable, to improve the business interpretability of the sample data, then select target variables with up-to-standard risk prediction capability from the plurality of variables, to avoid variables with poor risk prediction capability affecting the risk prediction effect of the risk prediction model, to reduce the risk of false positives, for each target variable, the bin with the highest proportion of abnormal sample objects in the plurality of bins corresponding to the target variable is taken as the target bin of the target variable, and a sample space containing each target bin is determined, i.e., a sample space is constructed based on high-risk sample data with up-to-standard risk prediction capability, to avoid the generation of prediction rules for non-risk items when the risk prediction model is constructed based on the sample space subsequently, which is conducive to effective risk prediction, further, the sample space is divided into a plurality of subspaces, and a risk prediction function of sample objects in each subspace is determined, so that a risk prediction model that can closely combine the business of a financial institution and accurately predict risks is constructed based on the risk prediction function corresponding to each subspace, specifically, the target subspace to which a to-be-predicted object belongs can be determined based on the risk prediction model, and the risk prediction function of the target subspace is used to closely combine the business of a financial institution and accurately predict the risk state of the to-be-predicted object. BRIEF DESCRIPTION OF DRAWINGS

[0022] Figure 1 An application environment diagram of the method for constructing a risk prediction model in an embodiment;

[0023] Figure 2 A flowchart of the method for constructing a risk prediction model in an embodiment;

[0024] Figure 3 A flowchart of the binning processing method in an embodiment;

[0025] Figure 4 A flowchart of the target variable screening method in an embodiment;

[0026] Figure 5 A flowchart of the space division method in an embodiment;

[0027] Figure 6 A diagram of a plurality of subspaces in an embodiment;

[0028] Figure 7 A flowchart of the space merging method in an embodiment;

[0029] Figure 8A flowchart of a method for constructing a survival analysis function in an embodiment;

[0030] Figure 9 A flowchart of a method for constructing a risk prediction function in an embodiment;

[0031] Figure 10 A flowchart of a method for constructing a risk prediction model in another embodiment;

[0032] Figure 11 A block diagram of a construction device for a risk prediction model in an embodiment;

[0033] Figure 12 An internal structure diagram of a computer device in an embodiment. DETAILED DESCRIPTION

[0034] In order to make the purposes, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and not to limit the present application.

[0035] It should be noted that the user (sample object) information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in the present application are all information and data authorized by each end user or fully authorized by each party, and the collection, use and processing of related data need to comply with relevant laws, regulations and standards of relevant countries and regions.

[0036] The method for constructing a risk prediction model provided in the embodiments of the present application can be applied to, for example, Figure 1The application environment shown. Among them, the terminals 102 of the plurality of branch offices of the financial institutions can communicate with the server 104 of the financial institutions through the network. The data storage system can store the data required by the server 104 to process. The data storage system can be integrated on the server 104, or placed on the cloud or other network servers. The server 104 can obtain sample data corresponding to each of the plurality of variables related to risk prediction by communicating with the terminals 102 of the plurality of branch offices of the financial institutions, so that the amount of sample data obtained is large enough, wherein the sample data belongs to different sample objects, and further, the server 104 can perform binning processing on each sample data to obtain a plurality of bins corresponding to each variable, and filter target variables with risk prediction capability from the plurality of variables, so that for each target variable, the server 104 can take the bin with the highest abnormal proportion of sample objects in the plurality of bins corresponding to the target variable as the target bin of the target variable, and further determine the sample space containing each target bin. Further, the server 104 can divide the sample space into a plurality of subspaces, determine the risk prediction function of the sample object in each subspace, and then construct a risk prediction model based on the risk prediction function corresponding to each subspace, so that the server 104 can subsequently determine the target subspace to which the to-be-predicted object belongs based on the risk prediction model, and predict the risk state of the to-be-predicted object based on the risk prediction function of the target subspace.

[0037] Among them, the financial institution refers to a financial intermediary engaged in financial services. The terminal 102 can be, but is not limited to, various personal computers, notebook computers, smart phones, tablet computers, Internet of Things devices, etc., and the terminal 102 of each branch office of the financial institution stores sample data of a plurality of sample objects obtained by the branch office when processing historical businesses. The server 104 of the financial institution can be implemented by an independent server or a server cluster composed of a plurality of servers.

[0038] In one embodiment, as Figure 2 shown, a risk prediction model construction method is provided. Taking the server in Figure 1 as an example, the method includes the following steps:

[0039] Step 202, obtaining sample data corresponding to each of the plurality of variables related to risk prediction, and performing binning processing on each sample data to obtain a plurality of bins corresponding to each variable; the sample data belongs to different sample objects.

[0040] The sample object specifically refers to a user who has completed registration in a financial institution and has handled business. The binning processing specifically can be feature binning, the core idea of which is to divide data into several bins according to certain rules, so that in the process of data mining and analysis, the distribution of data can be better understood, and the business interpretability of data can be improved.

[0041] For example, the financial institution has resource storage qualification and resource borrowing qualification. The variables related to risk prediction include but are not limited to continuous variables and discrete variables related to risk prediction. The continuous variable specifically can be the amount of resources held by the sample object, the amount of resources borrowed by the sample object from the financial institution, the amount of resources returned by the sample object, and the like. The discrete variable specifically can be the number of resource storage / borrowing accounts handled by the sample object in the financial institution, the number of times of resource borrowing handled by the sample object in the financial institution, the resource return condition corresponding to each resource borrowing of the sample object (whether the resource is returned on time), the registration information of the sample object (such as gender), and the like.

[0042] Optionally, the server can first determine a plurality of variables related to risk prediction, and then collect and obtain sample data corresponding to each of the plurality of variables related to risk prediction through communication with the terminals of a plurality of branch institutions of the financial institution. Then, the sample data is subjected to binning processing according to a set binning rule, and a plurality of bins corresponding to each of the variables are obtained. For each variable, the sample data belonging to different sample objects and related to the variable is stored in each bin.

[0043] For example, the server can determine a plurality of variables with high correlation with abnormal behavior according to historical data stored in the financial institution, such as sample data corresponding to sample objects with abnormal behavior, based on association rule mining algorithm, and determine the variables as variables related to risk prediction.

[0044] For example, the server can also obtain and analyze various documents and materials for predicting the probability of occurrence of abnormal behavior of the user through the network, determine a plurality of variables related to risk prediction, or obtain a plurality of variables set by the staff based on experience method in response to operation instructions of the staff, so as to determine the variables set by the staff as variables related to risk prediction.

[0045] Step 204: screening a target variable with risk prediction capability up to standard from the plurality of variables.

[0046] The target variable with risk prediction capability up to standard represents that the value and change trend of the sample data corresponding to the target variable have a greater impact on the probability of abnormal behavior of the sample object, have risk prediction value, and have practicality for risk prediction business of the financial institution.

[0047] Optionally, the server can closely integrate the business of the financial institution, and evaluate the risk prediction ability of each variable according to the set evaluation rule, and screen out target variables with risk prediction ability meeting the standard from the plurality of variables.

[0048] Step 206, for each target variable, the target bin with the highest abnormality proportion of sample objects in the plurality of bins corresponding to the target variable is taken as the target bin of the target variable.

[0049] Among them, for each variable, each bin has a respective abnormality proportion of sample objects.

[0050] Optionally, for each target variable, the server can determine the abnormality proportion of sample objects in each bin corresponding to the target variable according to the set abnormality proportion calculation method, and take the bin with the highest abnormality proportion of sample objects in the plurality of bins corresponding to the target variable as the target bin of the target variable, so as to determine the target bin corresponding to each of the plurality of target variables.

[0051] Step 208, determine the sample space containing each target bin, divide the sample space into a plurality of subspaces, and determine the risk prediction function of sample objects in each subspace.

[0052] Among them, the sample objects located in the same subspace have relatively close probability of abnormal behavior. The risk prediction function can be used to predict the probability of abnormal behavior of sample objects, and predict the trend of change of the risk state of sample objects.

[0053] Optionally, the server can only retain the target bin corresponding to each of the plurality of target objects, and construct the sample space based on the target bin corresponding to each of the plurality of target objects, so as to realize dimension reduction of data and improve the efficiency of subsequent iterative division and merging of the sample space.

[0054] Optionally, the server can perform the steps of subspace division and subspace merging on the sample space in a cycle according to the set subspace division rule and subspace merging rule, until the cycle stopping condition is reached, to obtain the plurality of subspaces after division, and further, the server can determine the risk prediction function of sample objects in each subspace based on the set risk prediction function fitting rule. Among them, only the target bin with the highest bin value is retained, and the sample space is constructed based on the target bin, which can avoid limiting the distribution of sample data in the sample space.

[0055] Step 210, construct a risk prediction model based on the risk prediction function corresponding to each subspace; the risk prediction model is used to determine the target subspace to which the to-be-predicted object belongs, and predict the risk state of the to-be-predicted object based on the risk prediction function of the target subspace.

[0056] Optionally, the server can construct a risk prediction model based on each subspace and the risk prediction function corresponding to each subspace.

[0057] Illustratively, if it is needed to predict the risk state in which the to-be-predicted object is located, the server can determine the target subspace to which the to-be-predicted object belongs based on the data related to the to-be-predicted object, and predict the probability of the to-be-predicted object currently performing abnormal behavior based on the risk prediction function of the target subspace.

[0058] In the method for constructing the risk prediction model, first, sample data corresponding to each variable related to risk prediction is obtained, that is, the sample data related to risk prediction is obtained in close combination with the business of the financial institution, then each sample data is subjected to binning processing to obtain a plurality of bins corresponding to each variable, so as to improve the business interpretability of the sample data, then target variables with up-to-standard risk prediction capability are selected from the plurality of variables, so as to avoid the influence of variables with poor risk prediction capability on the risk prediction effect of the risk prediction model, thereby reducing the risk of false positives, for each target variable, the bin with the highest abnormality ratio of sample objects in the plurality of bins corresponding to the target variable is taken as the target bin of the target variable, and a sample space containing each target bin is determined, that is, the sample space is constructed based on sample data with high risk and up-to-standard risk prediction capability, so as to avoid the generation of prediction rules of non-risk items when the risk prediction model is constructed based on the sample space in the subsequent step, which is conducive to effective risk prediction. Further, the sample space is divided into a plurality of subspaces, and a risk prediction function of sample objects in each subspace is determined, so as to construct a risk prediction model that can closely combine the business of the financial institution and accurately predict risks based on the risk prediction function corresponding to each subspace. Specifically, the target subspace to which the to-be-predicted object belongs can be determined based on the risk prediction model, and the risk state in which the to-be-predicted object is located can be accurately predicted based on the risk prediction function of the target subspace in close combination with the business of the financial institution.

[0059] In one embodiment, as shown in FIG. 3, Figure 3 the binning processing of each sample data to obtain a plurality of bins corresponding to each variable includes:

[0060] In step 302, for each variable related to risk prediction, the sample data corresponding to the variable is subjected to initial binning processing to obtain a plurality of initial bins corresponding to the variable.

[0061] Optionally, for each variable related to risk prediction, the server can randomly perform initial binning processing on the sample data corresponding to the variable to obtain a plurality of initial bins corresponding to the variable. Each initial bin stores sample data related to the variable and from different sample objects.

[0062] Exemplarily, for continuous variables, the server can implement the binning of the continuous variables by discretizing the variables. For discrete variables, the server can convert them into numerical features to implement the binning of the discrete variables.

[0063] At step 304, for each initial bin, the server determines a bin value of the initial bin based on an abnormality proportion and a non-abnormality proportion of the sample objects in the initial bin to which the sample data belong.

[0064] The bin value is specifically a WOE (Weight of Evidence) value of the bin. Each bin has a respective abnormality proportion and a respective non-abnormality proportion. For each bin, the abnormality proportion can be understood as a proportion of the sample objects in the bin that have abnormal behaviors among all the sample objects that have abnormal behaviors; and the non-abnormality proportion can be understood as a proportion of the sample objects in the bin that do not have abnormal behaviors among all the sample objects that do not have abnormal behaviors. That is, for each bin, the abnormality proportion can be specifically a number of the sample objects in the bin that have abnormal behaviors divided by a number of all the sample objects that have abnormal behaviors; and the non-abnormality proportion can be specifically a number of the sample objects in the bin that do not have abnormal behaviors divided by a number of all the sample objects that do not have abnormal behaviors.

[0065] Optionally, for each initial bin, the server can determine an abnormality proportion corresponding to the initial bin based on a number of the sample objects in the initial bin that have abnormal behaviors and a number of all the sample objects that have abnormal behaviors. The server can also determine a non-abnormality proportion corresponding to the initial bin based on a number of the sample objects in the initial bin that do not have abnormal behaviors and a number of all the sample objects that do not have abnormal behaviors. Further, the server can determine a bin value of the initial bin based on the abnormality proportion and the non-abnormality proportion of the initial bin.

[0066] Exemplarily, assuming that a number of the sample objects in an i-th bin of a variable that have abnormal behaviors is Bad i , a number of the sample objects in the i-th bin that do not have abnormal behaviors is Good i , a number of all the sample objects that have abnormal behaviors is Bad T , and a number of all the sample objects that do not have abnormal behaviors is Good T , an abnormality proportion corresponding to the i-th bin is and a non-abnormality proportion corresponding to the i-th bin is The server can specifically determine a bin value of the i-th bin by formula (1):

[0067]

[0068] wherein WOE i is the bin value of the i-th bin of a variable.

[0069] Step 306, adjusting the plurality of initial bins so as to make the bin value of the bin positively correlated with the abnormality proportion of the sample object to which the sample data in the bin belongs, to obtain the plurality of bins corresponding to the variable.

[0070] Optionally, for the variable without non-monotonic business explanation, the server can adjust the plurality of initial bins so as to make the bin value of the bin positively correlated with the abnormality proportion of the sample object to which the sample data in the bin belongs, until the change amplitude of the plurality of bins corresponding to the variable is small enough in the adjustment process, to obtain the plurality of adjusted bins corresponding to the variable.

[0071] Illustratively, for the variable without non-monotonic business explanation, the server can also verify the monotonicity of the WOE value of the plurality of bins corresponding to each variable based on the pre-constructed data set, to ensure that the monotonicity of the bins does not reverse.

[0072] Illustratively, for a small part of relatively special variables with explicit non-monotonic business explanation, the server does not need to adjust the bins of such variables so as to make the bin value of the bin positively correlated with the abnormality proportion of the sample object to which the sample data in the bin belongs. The variable with explicit non-monotonic business explanation can be specifically the number of resource storage / borrowing accounts handled by the sample object in the financial institution, which has explicit non-monotonic business explanation and can be understood as the probability of abnormal behavior of the sample object, which does not increase with the increase of the number of resource storage / borrowing accounts, but shows a trend of first increasing and then decreasing.

[0073] In this embodiment, the binning operation is performed on the sample data corresponding to the variable so as to make the bin value of the bin positively correlated with the abnormality proportion of the sample object to which the sample data in the bin belongs, which can ensure that the WOE value of the plurality of bins of each variable has monotonicity, and can ensure that the larger the WOE value is, the larger the proportion of the sample object with abnormal behavior in the bin is, so that the bin with the highest WOE value can be quickly determined as the bin with the highest abnormality proportion of the sample object in the plurality of bins according to the monotonicity of the WOE value in the subsequent process.

[0074] In one embodiment, as shown in Figure 4 , the target variable with risk prediction capability up to the standard is selected from the plurality of variables, including:

[0075] Step 402, for each variable, evaluating the risk prediction capability of the variable based on the bin value of each bin corresponding to the variable, to obtain the prediction capability evaluation value of the variable.

[0076] The prediction capability evaluation value of the variable is specifically an information value IV of the variable, which is a weighted sum of the bin values WOE of the plurality of bins corresponding to the variable. The WOE is closely related to an index, and is mainly used to evaluate the prediction capability of the variable to quickly screen the variable.

[0077] Optionally, for each variable, the server can evaluate the risk prediction capability of the variable based on the weighted sum of the bin values of the plurality of bins corresponding to the variable.

[0078] For example, the server can specifically determine the prediction capability evaluation value IV of a variable by formula (2):

[0079]

[0080] wherein n is the total number of bins corresponding to a variable, WOE i is the bin value of the i th bin corresponding to the variable, Bad i is the number of sample objects with abnormal behavior in the i th bin of the variable, Good i is the number of sample objects without abnormal behavior in the i th bin of the variable, Bad T is the number of sample objects with abnormal behavior in all sample objects, Good T is the number of sample objects without abnormal behavior in all sample objects.

[0081] Step 404, determining the variable with the prediction capability evaluation value greater than the evaluation threshold as the target variable with the risk prediction capability meeting the requirement.

[0082] The evaluation threshold can be flexibly configured according to actual needs.

[0083] Optionally, the server can determine the variable with the prediction capability evaluation value greater than the evaluation threshold as the target variable with the risk prediction capability meeting the requirement, to screen the variable with practicality for the risk prediction business of the financial institution.

[0084] In this embodiment, by screening the variable, on the one hand, it is ensured that the sample space is constructed based on the high-risk sample data corresponding to the variable with good prediction capability, ensuring that the risk prediction rules mined subsequently are all rules with practical value in the risk prediction scene, and it can also avoid missing the low-frequency high-risk rules representing major risks, and it can also avoid the influence of the variable with poor risk prediction capability on the risk prediction effect of the risk prediction model, so as to reduce the risk of false positives. On the other hand, only the target bin corresponding to the target variable is retained, the number of sample data is reduced, the dimensionality reduction of data is realized, and it is beneficial to improve the efficiency and accuracy of subsequent iterative division of the sample space.

[0085] In one embodiment, as shown in Figure 5 the sample space is divided into a plurality of subspaces, including:

[0086] Step 502, constructing a survival analysis function for determining the survival rate of the sample objects in the space.

[0087] The survival rate of the space can represent the risk level of the sample objects in the space. The higher the survival rate of the subspace, the higher the probability of abnormal behavior of the sample objects in the subspace, and the higher the risk level.

[0088] Optionally, the server can introduce an evaluation index, i.e., the survival rate, which is more suitable for the risk prediction business of the financial institution, based on the survival analysis algorithm, and construct a survival analysis function for determining the survival rate of the sample objects in the space. The survival analysis algorithm can be used to analyze the survival time of the sample objects. In the field of financial technology, the survival time of the sample objects can be the duration of the sample objects continuously not occurring abnormal behavior. Therefore, in this embodiment, the survival analysis function for determining the survival rate of the sample objects in the space can be constructed based on the survival analysis algorithm, so as to closely analyze the probability of abnormal behavior of the sample objects in the space.

[0089] Step 504, taking the sample space as a parent space, maximizing the increase of the survival rate after division as an objective, performing space division on the parent space to obtain two subspaces.

[0090] The PRIM (Patient Rule Induction Method) model can be combined with the survival analysis function of the space to realize the space division. The PRIM model can divide the sample space into a plurality of rectangular regions by setting the threshold values of continuous variables and classification variables. In this embodiment, the survival rate increase threshold value and the sample object increase threshold value in the parent space can be set to iteratively perform the steps of "peeling" and "merging" to realize the division of the sample space. After dividing a space, the increase of the survival rate = the sum of the survival rates of the two subspaces obtained by division - the survival rate of the divided space.

[0091] Optionally, the server can first take the sample space as a parent space, and randomly generate a plurality of schemes for dividing the parent space, and then maximize the increase of the survival rate after division as an objective, and screen the scheme with the largest increase of the survival rate from the plurality of randomly generated schemes, so as to divide the parent space according to the screened scheme to obtain two subspaces.

[0092] Step 506, the subspace with smaller survival rate is stripped from the parent space, the subspace with larger survival rate is reserved, and the subspace with larger survival rate is taken as a new parent space.

[0093] Optionally, after obtaining two subspaces, the server can compare the survival rates of the two subspaces respectively, strip the subspace with smaller survival rate from the parent space, reserve the subspace with larger survival rate, and take the subspace with larger survival rate as a new parent space.

[0094] Step 508, the process of spatial division of the parent space is cycled to maximize the increase of the survival rate after division; during the cycling process, if there is a stripped subspace adjacent to the parent space and the difference between the survival rates of the parent space and the stripped subspace is less than the difference threshold value, the stripped subspace and the parent space are merged.

[0095] The difference threshold value can be flexibly configured according to actual application scenarios. The difference between the survival rates = the survival rate of the merged space - the sum of the survival rates of the two spaces before merging.

[0096] Optionally, the server can cycle the process of spatial division of the parent space to maximize the increase of the survival rate after division, and during the cycling process, after performing the spatial division step once, it performs a judgment to determine whether there is a stripped subspace adjacent to the parent space and the difference between the survival rates of the parent space and the stripped subspace is less than the difference threshold value. If there is a stripped subspace adjacent to the parent space and the difference between the survival rates of the parent space and the stripped subspace is less than the difference threshold value, step 510 is performed to merge the stripped subspace and the parent space, so that the number of sample objects in the parent space increases, otherwise, step 508 is continued.

[0097] Step 512, when the loop stop condition is reached, the divided sample space is obtained.

[0098] The loop stop condition can be specifically: setting a maximum number of spatial divisions, and if the number of spatial divisions reaches the set maximum number, it is determined that the loop stop condition is reached. Or, when the subspace is divided or merged, the change amplitude of the subspace is small enough, and it is also determined that the loop stop condition is reached.

[0099] Optionally, when the loop stop condition is reached, the server can output the divided sample space, obtain a plurality of stripped subspaces, and each subspace has a corresponding survival rate. Since the spatial division process is to maximize the increase of the survival rate after division, the finally reserved subspace is the subspace with the largest survival rate and the highest risk rate.

[0100] Exemplarily, taking "the change range of the subspace is small enough after the subspace is divided or combined" as the loop stopping condition, the server can preset a survival rate increase threshold and a sample object increase threshold. If in a certain loop process, the survival rate increase after the space is divided is less than the survival rate increase threshold, and the sample object increase in the parent space after the space is combined is less than the sample object increase threshold, the server can determine that the loop stopping condition is reached, and output the divided sample space.

[0101] Exemplarily, taking Figure 6 as an example, a schematic diagram of a plurality of subspaces obtained after the sample space is divided is provided. First, the sample space can be regarded as a large rectangular area. For the original sample space, the server can at least randomly generate Figure 6 B1 and b 11- , B1 and b 11+ , B1 and b 12- , B1 and b 12+ These four space division modes, B1 and b 11+ This division mode makes the survival rate increase the most. For example, the server can divide the sample space into B2 and b1 * , and further, because the survival rate corresponding to B2 is greater than the survival rate corresponding to b1 * , the server will retain B2 as a new parent space and strip b1 * . Furthermore, the server can also divide B2 into B3 and b2 * according to the same space division rule, and because the survival rate corresponding to B3 is greater than the survival rate corresponding to b2 * , the server will retain B3 as a new parent space and strip b2 * . Finally, through multiple rounds of space division and space combination, the divided sample space is obtained. Among them, the sample space is divided into a plurality of subspaces b1 * , b2 * , b3 * , b4 * , b5 * , b6 * , b7 * , b8 * , B9, b1 * ~ b8 * are stripped subspaces, B9 is the finally retained subspace, and all the finally obtained subspaces do not overlap and can be spliced into the original sample space.

[0102] In this embodiment, through the steps of space division and space merging of the sample space, the survival rate of each subspace and the number of sample objects covered in each subspace can be continuously adjusted, the fitting quality of the PRIM model is improved, the survival rate of each subspace and the number of sample objects covered tend to be balanced, the prediction accuracy and coverage of each subspace tend to be balanced, and finally multiple subspaces corresponding to different survival rates are obtained, the sample object groups with different survival rates are accurately divided, and the risk prediction of the to-be-predicted object can be accurately and quickly performed based on the subspace to which the to-be-predicted object belongs.

[0103] In one embodiment, as shown in FIG. 7, Figure 7 the method further includes:

[0104] In step 702, if there are multiple peeled subspaces adjacent to the parent space and the difference between the survival rate of each peeled subspace and the survival rate of the parent space is less than the difference threshold, the number of sample objects in each peeled subspace is determined.

[0105] Optionally, after completing the space division in step 508, the server will perform a judgment to determine whether there are peeled subspaces adjacent to the parent space and the difference between the survival rate of each peeled subspace and the survival rate of the parent space is less than the difference threshold. If there are multiple peeled subspaces adjacent to the parent space and the difference between the survival rate of each peeled subspace and the survival rate of the parent space is less than the difference threshold, step 702 is performed to determine the number of sample objects in each peeled subspace.

[0106] For example, if only one peeled subspace is adjacent to the parent space and the difference between the survival rate of the peeled subspace and the survival rate of the parent space is less than the difference threshold after the judgment step is performed, the step of merging the peeled subspace and the parent space is directly performed. If there is no peeled subspace adjacent to the parent space and the difference between the survival rate of each peeled subspace and the survival rate of the parent space is less than the difference threshold after the judgment step is performed, step 508 is directly returned.

[0107] In step 704, the peeled subspace with the largest number of sample objects is merged with the parent space.

[0108] Optionally, the server can compare the number of sample objects in the multiple peeled subspaces determined in step 702, and take the peeled subspace with the largest number of sample objects as the to-be-merged subspace, so that in step 704, the peeled subspace with the largest number of sample objects is merged with the parent space, so that the increase (change) of the number of sample objects in the parent space after the space merging is maximized.

[0109] In this embodiment, the space merging is performed with the objective of maximizing the increase of the sample objects in the parent space after the space merging, which is beneficial to jump out of the local to find the global optimal solution, so that the prediction accuracy and coverage of each subspace in the final output are balanced, and the sample object groups with different survival rates are accurately divided, so that the risk prediction of the to-be-predicted object can be accurately and quickly performed based on the subspace to which the to-be-predicted object belongs.

[0110] In one embodiment, as shown in FIG. 8, Figure 8 the method further includes:

[0111] At step 802, for each sample object, the number of days during which the sample object continuously does not have abnormal behavior in a target time period is taken as the survival days of the sample object, and the target time period is a time period from the starting time point at which the sample object generates the resource transfer record to the ending time point at which the sample data of the sample object is acquired.

[0112] Optionally, for each sample object, the server can take the number of days during which the sample object continuously does not have abnormal behavior in a target time period as the survival days of the sample object, that is, the survival length of the sample object, that is, in combination with the actual business scenario of the financial institution, the data related to the abnormal behavior of the sample object is combined with the survival analysis algorithm to mine high-quality and practical rules that can predict the probability of the abnormal behavior of the sample object based on the business characteristics of the financial institution.

[0113] At step 804, the value of the survival label of the sample object is determined by judging whether the sample object has abnormal behavior until the ending time point.

[0114] Optionally, for each sample object, the server can judge whether the sample object has abnormal behavior from the starting time point to the ending time point, and configure different values for the survival label of the sample object according to the judgment result.

[0115] For example, if the sample object has abnormal behavior from the starting time point to the ending time point, the server can assign a value of 0 to the survival label of the sample object; if the sample object does not have abnormal behavior from the starting time point to the ending time point, the server can assign a value of 1 to the survival label of the sample object.

[0116] At step 806, a survival analysis function for determining the survival rate of each sample object in the space is constructed based on the sum of the survival labels of each sample object in the space and the sum of the survival days of each sample object.

[0117] The space specifically can include: an original sample space, and a parent space and a subspace generated in the space division process.

[0118] Optionally, the server can construct a survival analysis function for determining the survival rate of each sample object in the space, by taking the sum of survival labels of each sample object in the space as the numerator, and taking the sum of survival days of each sample object in the space as the denominator.

[0119] For example, for a certain space, the survival rate of each sample object in the space can be calculated by formula (3):

[0120]

[0121] where f(t, δ) is the survival rate of each sample object in the space, m is the total number of sample objects in the space, δ i is the survival label of the i-th sample object in the space, t i is the survival day of the i-th sample object in the space.

[0122] In this embodiment, the survival analysis function is constructed by closely combining the risk prediction business of the financial institution and by fusing the survival day and the survival label of the sample object, so as to determine the survival rate of each sample object in different spaces. Thus, the sample space can be divided based on the survival rate of each space, the sample object groups with different survival rates can be divided, and the survival rate corresponding to each sample object group is determined. Therefore, the risk prediction of the to-be-predicted object can be accurately and quickly performed based on the sub-space to which the to-be-predicted object belongs.

[0123] In one embodiment, as shown in Figure 9 , the risk prediction function of each sample object in each sub-space is determined, including:

[0124] In step 902, for each sub-space, the abnormal time points of the sample objects with abnormal behaviors in the sub-space are obtained, and the number of sample objects with abnormal behaviors at each abnormal time point is counted.

[0125] Optionally, for each sub-space, the server can obtain the abnormal time points of the sample objects with abnormal behaviors in the sub-space based on the sample data of each sample object in the sub-space, and count the number of sample objects with abnormal behaviors at each abnormal time point.

[0126] In step 904, the trend of the number of sample objects with abnormal behaviors changing with time is fitted to obtain the risk prediction function of the sample objects in the sub-space.

[0127] Optionally, the server can determine the number of sample objects in each abnormal time point that do not exhibit abnormal behavior based on the number of sample objects in each abnormal time point that exhibit abnormal behavior, and fit the trend of the number of sample objects that exhibit abnormal behavior over time based on the number of sample objects in each abnormal time point that exhibit abnormal behavior and the number of sample objects that do not exhibit abnormal behavior, to obtain a risk prediction function of the sample objects in the subspace.

[0128] For example, for each subspace, the server can input the number of sample objects that exhibit abnormal behavior and the number of sample objects that do not exhibit abnormal behavior corresponding to each abnormal time point in the subspace into formula (4) to fit the risk prediction function corresponding to the subspace, and formula (4) is as follows:

[0129]

[0130] Where h(t) is the probability (risk rate) of sample objects in the space exhibiting abnormal behavior in the time period [t, t+Δt]. f(t) is the number of sample objects in the space that have exhibited abnormal behavior in the time period [t, t+Δt]. S(t) is the number of sample objects in the space that have not exhibited abnormal behavior in the time period [t, t+Δt]. t is any time point, and Δt is a very short time interval.

[0131] For example, after fitting the risk prediction function corresponding to a space, the server can determine the probability of sample objects in the space exhibiting abnormal behavior and the trend of the probability of sample objects exhibiting abnormal behavior over time based on the risk prediction function. If the risk state of a to-be-predicted object needs to be predicted at a certain time point, the server can predict the probability of the to-be-predicted object exhibiting abnormal behavior based on the time point and the risk prediction function of the target subspace to which the to-be-predicted object belongs.

[0132] For example, at the time point of predicting the risk of a to-be-predicted object, if the risk prediction function of the target subspace to which the to-be-predicted object belongs is a non-decreasing function, that is, the output value of the risk prediction function increases as time increases, which indicates that the overall credit quality of the sample object group to which the to-be-predicted object belongs has a downward trend, that is, the probability of the to-be-predicted object exhibiting abnormal behavior is relatively large, that is, the to-be-predicted object is in a high-risk state.

[0133] For example, if the risk prediction function of the target subspace to which the to-be-predicted object belongs is a non-increasing function, that is, the output value of the risk prediction function decreases as time increases, which indicates that the overall credit quality of the sample object group to which the to-be-predicted object belongs has an upward trend, that is, the probability of the to-be-predicted object exhibiting abnormal behavior is relatively small, that is, the to-be-predicted object is in a low-risk state.

[0134] Exemplarily, if the risk prediction function of the target subspace to which the to-be-predicted object belongs is a constant, that is, the value output by the risk prediction function is basically unchanged as time increases, it indicates that the overall credit quality of the sample object group to which the to-be-predicted object belongs is relatively stable.

[0135] In this embodiment, the risk prediction function is introduced, so that the risk trend of any sample object group at any time point can be predicted. Specifically, based on the risk prediction function configured for the subspace, the risk prediction function of the to-be-predicted object can be accurately and quickly predicted at any time point, in combination with the risk prediction business of the financial institution.

[0136] In one embodiment, as shown in Figure 10 Another method for constructing a risk prediction model is provided, which is applied to the application environment as shown in Figure 1 The method mainly includes the following steps:

[0137] After the server obtains the sample data corresponding to each of the variables related to risk prediction, the server can perform step 1002 to perform initial binning processing on the sample data corresponding to each of the variables related to risk prediction, respectively, to obtain a plurality of initial bins corresponding to each of the variables. For each initial bin, the server can perform step 1004 to determine the bin value of the initial bin based on the abnormal proportion and the non-abnormal proportion of the sample objects in the initial bin, so that the bin value of the bin is positively correlated with the abnormal proportion of the sample objects in the bin. The server can perform step 1006 to adjust the plurality of initial bins to obtain a plurality of bins corresponding to each variable, so that for each variable, the server can perform step 1008 to evaluate the risk prediction ability of the variable based on the bin values of the plurality of bins corresponding to the variable, to obtain a prediction ability evaluation value of the variable, and then perform step 1010 to determine the target variable whose risk prediction ability meets the standard as the variable whose prediction ability evaluation value is greater than the evaluation threshold.

[0138] Further, for each target variable, the server can perform step 1012 to determine the bin with the highest abnormal proportion of sample objects in the plurality of bins corresponding to the target variable as the target bin of the target variable, and then perform step 1014 to determine the sample space containing each target bin.

[0139] After the sample space is constructed, the server can perform step 1016, taking the sample space as a parent space, and maximizing the increase of survival rate after division as an objective, step 1018, performing space division on the parent space to obtain two subspaces, and then performing step 1020, peeling off the subspace with smaller survival rate from the parent space, and taking the subspace with larger survival rate as a new parent space, and then, in a loop, performing space division on the parent space with the objective of maximizing the increase of survival rate after division. In the loop, if there are multiple peeled subspaces adjacent to the parent space and the difference between the survival rates of the parent space and the peeled subspaces is less than a difference threshold, step 1022 is performed to combine the peeled subspace with the largest number of sample objects with the parent space, until a loop stop condition is reached, and a completed sample space is obtained.

[0140] After the division of the sample space is completed, for each subspace, the server can perform step 1024, obtaining the abnormal time points of the sample objects with abnormal behaviors in the subspace, and counting the number of sample objects with abnormal behaviors at each abnormal time point, and then performing step 1026, fitting the trend of the number of sample objects with abnormal behaviors over time to obtain a risk prediction function of the sample objects in the subspace, and finally performing step 1028, constructing a risk prediction model based on the respective risk prediction functions of the subspaces.

[0141] It should be understood that, although each step in the flowchart involved in each embodiment described above is displayed in sequence according to the arrow, these steps are not necessarily executed in sequence according to the arrow. Unless otherwise specified herein, the execution of these steps is not strictly limited in sequence, and these steps can be executed in other sequences. Moreover, at least part of the steps in the flowchart involved in each embodiment described above can include multiple steps or stages, which are not necessarily executed at the same time, but can be executed at different times, and the execution sequence of these steps or stages is not necessarily sequential, but can be executed in rotation or alternation with at least part of other steps or stages or steps or stages in other steps.

[0142] Based on the same inventive concept, the embodiments of the present application also provide a risk prediction model construction device for implementing the risk prediction model construction method described above. The implementation scheme for solving the problem provided by the device is similar to the implementation scheme described in the above method, so the specific limitations in one or more risk prediction model construction device embodiments provided below can refer to the limitations of the risk prediction model construction method described above, and will not be repeated here.

[0143] In one embodiment, as Figure 11As shown, a risk prediction model construction device is provided, comprising: a binning processing module 1102, a variable screening module 1104, a binning determination module 1106, a space processing module 1108, and a model construction module 1110, wherein:

[0144] The binning processing module is configured to obtain sample data corresponding to each of a plurality of variables related to risk prediction, and perform binning processing on each sample data to obtain a plurality of bins corresponding to each of the variables; the sample data belongs to different sample objects;

[0145] The variable screening module is configured to screen target variables with up-to-standard risk prediction capability from the plurality of variables;

[0146] The binning determination module is configured to, for each target variable, determine a bin with the highest abnormality proportion of sample objects among the plurality of bins corresponding to the target variable as a target bin of the target variable;

[0147] The space processing module is configured to determine a sample space containing the target bins, divide the sample space into a plurality of subspaces, and determine a risk prediction function of sample objects in each of the subspaces;

[0148] The model construction module is configured to construct a risk prediction model based on the risk prediction function corresponding to each of the subspaces; the risk prediction model is configured to determine a target subspace to which a to-be-predicted object belongs, and predict a risk state of the to-be-predicted object based on the risk prediction function of the target subspace.

[0149] In the device for constructing the risk prediction model, sample data corresponding to each of a plurality of variables related to risk prediction is first obtained, i.e., the sample data related to risk prediction is first obtained in close combination with the business of the financial institution, each sample data is then subjected to binning processing to obtain a plurality of bins corresponding to each variable, so as to improve the business interpretability of the sample data, target variables with up-to-standard risk prediction capability are then selected from the plurality of variables, so as to avoid the influence of variables with poor risk prediction capability on the risk prediction effect of the risk prediction model, reduce the risk of false positives, for each target variable, the bin with the highest abnormality proportion of sample objects in the plurality of bins corresponding to the target variable is taken as the target bin of the target variable, and a sample space containing each target bin is determined, i.e., a sample space is constructed based on sample data with high risk and up-to-standard risk prediction capability, so as to avoid the generation of prediction rules of non-risk items when the risk prediction model is constructed based on the sample space, which is conducive to effective risk prediction, and further, the sample space is divided into a plurality of subspaces, and a risk prediction function of sample objects in each subspace is determined, so that a risk prediction model capable of accurately predicting risks in close combination with the business of the financial institution is constructed based on the risk prediction function corresponding to each subspace, specifically, the target subspace to which a to-be-predicted object belongs can be determined based on the risk prediction model, and the risk state of the to-be-predicted object can be accurately predicted in close combination with the business of the financial institution based on the risk prediction function of the target subspace.

[0150] In one of the embodiments, the binning processing module is further configured to: for each variable related to risk prediction, perform initial binning processing on sample data corresponding to the variable to obtain a plurality of initial bins corresponding to the variable; for each initial bin, determine a bin value of the initial bin based on the abnormality proportion and the non-abnormality proportion of sample objects to which sample data in the initial bin belong; and adjust the plurality of initial bins so that the bin value of the bin is positively correlated with the abnormality proportion of sample objects to which sample data in the bin belong, to obtain a plurality of bins corresponding to the variable.

[0151] In one of the embodiments, the variable screening module is further configured to: for each variable, evaluate the risk prediction capability of the variable based on the bin values of the plurality of bins corresponding to the variable, to obtain a prediction capability evaluation value of the variable; and determine a variable with a prediction capability evaluation value greater than an evaluation threshold as a target variable with up-to-standard risk prediction capability.

[0152] In one of the embodiments, the space processing module is further configured to: construct a survival analysis function for determining the survival rate of the sample objects in the space; divide the sample space into two sub-spaces by performing space division on the sample space, with the aim of maximizing the increase in the post-division survival rate; separate the sub-space with the lower survival rate from the sample space, retain the sub-space with the higher survival rate, and take the sub-space with the higher survival rate as a new sample space; perform the space division on the sample space with the aim of maximizing the increase in the post-division survival rate in a loop; in the loop, if there is a separated sub-space adjacent to the sample space and the difference between the survival rates of the separated sub-space and the sample space is less than a difference threshold, merge the separated sub-space and the sample space; and obtain the divided sample space when a loop stop condition is reached.

[0153] In one of the embodiments, the risk prediction model construction apparatus further comprises a space merging module configured to: if there are multiple separated sub-spaces adjacent to the sample space and the difference between the survival rates of the separated sub-spaces and the sample space is less than a difference threshold, determine the number of sample objects in each separated sub-space; and merge the separated sub-space with the largest number of sample objects and the sample space.

[0154] In one of the embodiments, the risk prediction model construction apparatus further comprises a survival analysis function construction module configured to: for each sample object, take the number of days during which the sample object continuously does not perform abnormal behavior within a target time period as the survival days of the sample object; the target time period is a time period from the start time of the sample object generating the resource transfer record to the end time of obtaining the sample data of the sample object; determine the value of the survival label of the sample object by judging whether the sample object performs abnormal behavior by the end time; and construct a survival analysis function for determining the survival rate of each sample object in the space based on the sum of the survival labels of the sample objects in the space and the sum of the survival days of the sample objects.

[0155] In one of the embodiments, the model construction module is further configured to: for each sub-space, obtain the abnormal time points of the sample objects performing abnormal behavior in the sub-space, and count the number of sample objects performing abnormal behavior at each abnormal time point; and fit the trend of the number of sample objects performing abnormal behavior over time to obtain a risk prediction function of the sample objects in the sub-space.

[0156] The modules in the risk prediction model construction apparatus described above can be all or partially implemented by software, hardware, or a combination thereof. The modules described above can be embedded in or independent of the processor in the computer device in hardware form, or stored in the memory in the computer device in software form, so as to be called and executed by the processor to perform the operations corresponding to the modules.

[0157] In one embodiment, a computer device is provided, which can be a server, and an internal structure diagram of the computer device can be as shown in FIG. 1. Figure 12 The computer device includes a processor, a memory, an input / output interface (I / O), and a communication interface. The processor, the memory, and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. The processor of the computer device is configured to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for running the operating system and the computer program in the non-volatile storage medium. The database of the computer device is configured to store construction data of a risk prediction model. The input / output interface of the computer device is configured to exchange information between the processor and external devices. The communication interface of the computer device is configured to communicate with terminals outside through a network connection. The computer program is executed by the processor to implement a method for constructing a risk prediction model.

[0158] Those skilled in the art can understand that Figure 12 The structure shown in FIG. 1 is only a block diagram of part of the structure related to the scheme of the present application, and does not constitute a limitation on the computer device to which the scheme of the present application is applied. The specific computer device can include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.

[0159] In one embodiment, a computer device is provided, which includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the steps in the above method embodiments.

[0160] In one embodiment, a computer readable storage medium is provided, which stores a computer program. The computer program is executed by a processor to implement the steps in the above method embodiments.

[0161] In one embodiment, a computer program product is provided, which includes a computer program. The computer program is executed by a processor to implement the steps in the above method embodiments.

[0162] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer readable storage medium, and when the computer program is executed, the processes of the above-mentioned embodiments of the methods can be included. Any reference to memory, database or other medium used in the embodiments provided in the present application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (Read-Only Memory, ROM), magnetic tape, floppy disk, flash memory, optical storage, high-density embedded non-volatile memory, resistive memory (ReRAM), magnetoresistive random access memory (Magnetoresistive Random Access Memory, MRAM), ferroelectric memory (Ferroelectric Random Access Memory, FRAM), phase change memory (Phase Change Memory, PCM), graphene memory, etc. Volatile memory can include random access memory (Random Access Memory, RAM) or external cache memory, etc. As an illustration but not limitation, RAM can be in various forms, such as static random access memory (Static Random Access Memory, SRAM) or dynamic random access memory (Dynamic Random Access Memory, DRAM), etc. The database involved in the embodiments provided in the present application can include at least one of a relational database and a non-relational database. The non-relational database can include a distributed database based on a block chain, etc., without being limited thereto. The processor involved in the embodiments provided in the present application can be a general-purpose processor, a central processing unit, a graphics processing unit, a digital signal processor, a programmable logic device, a data processing logic device based on quantum computing, etc., without being limited thereto.

[0163] Any combination of the technical features of the above embodiments can be made. In order to make the description simple, all possible combinations of the technical features in the above embodiments are not described, however, as long as the combination of the technical features does not exist contradictory, it should be considered as the scope of the present application.

[0164] The above embodiments only express several implementation manners of the present application, and the description is more specific and detailed, but it should not be understood as a limitation on the scope of the patent of the present application. It should be pointed out that for ordinary skilled in the art, without departing from the concept of the present application, a number of modifications and improvements can be made, which are within the scope of protection of the present application. Therefore, the protection scope of the present application should be subject to the appended claims.

Claims

1. A method for constructing a risk prediction model, characterized in that, The method includes: Obtain sample data corresponding to multiple variables related to risk prediction, and perform binning on each sample data to obtain multiple bins corresponding to each variable; the sample data belong to different sample objects. Select target variables that meet the risk prediction capability criteria from among the multiple variables mentioned; For each target variable, the bin with the highest anomaly rate among the multiple bins corresponding to the target variable is taken as the target bin for the target variable; Determine the sample space containing each of the target bins, and construct a survival analysis function to determine the survival rate of sample objects in the space; Using the sample space as the parent space, and with the goal of maximizing the increase in survival rate after partitioning, the parent space is divided into two subspaces. The subspace with a lower survival rate is separated from the parent space, while the subspace with a higher survival rate is retained, and the subspace with a higher survival rate is used as the new parent space. The process of partitioning the parent space is aimed at maximizing the increase in survival rate after partitioning. During the process, if there is a stripped subspace that is adjacent to the parent space and the difference in survival rate between the subspace and the parent space is less than the difference threshold, the stripped subspace is merged with the parent space. When the loop stopping condition is met, the partitioned sample space is obtained, and the risk prediction function for the sample objects in each subspace is determined respectively. A risk prediction model is constructed based on the risk prediction function corresponding to each of the subspaces. The risk prediction model is used to determine the target subspace to which the object to be predicted belongs, and to predict the risk state of the object to be predicted based on the risk prediction function of the target subspace.

2. The method according to claim 1, characterized in that, The step of binning each of the sample data to obtain multiple bins corresponding to each variable includes: For each variable related to risk prediction, the sample data corresponding to the variable is subjected to initial binning to obtain multiple initial bins corresponding to the variable. For each initial bin, the binning value of the initial bin is determined based on the proportion of abnormal and non-abnormal samples belonging to the sample objects in the initial bin. With the goal of making the binning value positively correlated with the anomaly ratio of the sample object to which the sample data in the bin belongs, the multiple initial bins are adjusted to obtain multiple bins corresponding to the variable.

3. The method according to claim 2, characterized in that, From the multiple variables mentioned above, target variables that meet the risk prediction capability criteria are selected, including: For each variable, the risk prediction capability of the variable is evaluated based on the binning values ​​of the multiple bins corresponding to the variable, and the prediction capability evaluation value of the variable is obtained. Variables whose predictive ability assessment values ​​are greater than the assessment threshold are identified as target variables for achieving the risk prediction ability standard.

4. The method according to claim 1, characterized in that, The method further includes: If there are multiple stripped subspaces adjacent to the parent space, and the difference in survival rate between them and the parent space is less than the difference threshold, determine the number of sample objects in each stripped subspace. Merge the largest number of stripped subspaces with the parent space.

5. The method according to claim 1, characterized in that, The method further includes: For each of the aforementioned sample objects, the number of days during which the sample object does not exhibit any abnormal behavior within the target time period is taken as the survival days of the sample object; the target time period is the time period consisting of the start time when the sample object generates resource transfer records and the end time when the sample data of the sample object is acquired; The value of the survival tag of the sample object is determined by judging whether the sample object has exhibited abnormal behavior up to the end time. The construction of the survival analysis function for determining the survival rate of sample objects in space includes: Based on the sum of survival labels and the sum of survival days of each sample object in the space, a survival analysis function is constructed to determine the survival rate of each sample object in the space.

6. The method according to claim 1, characterized in that, The step of determining the risk prediction function for each sample object in each subspace includes: For each subspace, obtain the abnormal time point of the sample object that exhibited abnormal behavior in the subspace, and count the number of sample objects that exhibited abnormal behavior at each abnormal time point; By fitting the trend of the number of sample objects exhibiting abnormal behavior over time, a risk prediction function for the sample objects in the subspace is obtained.

7. A device for constructing a risk prediction model, characterized in that, The device includes: The binning module is used to acquire sample data corresponding to multiple variables related to risk prediction, and to perform binning processing on each sample data to obtain multiple bins corresponding to each variable; the sample data belong to different sample objects. The variable filtering module is used to filter out target variables that meet the risk prediction capability standards from multiple variables; The binning determination module is used to select the bin with the highest anomaly rate among the multiple bins corresponding to the target variable as the target bin for each target variable. A spatial processing module is used to determine the sample space containing each target bin, construct a survival analysis function for determining the survival rate of sample objects in the space, use the sample space as the parent space, and divide the parent space into two subspaces with the goal of maximizing the increase in survival rate after partitioning; the subspace with the lower survival rate is separated from the parent space, the subspace with the higher survival rate is retained, and the subspace with the higher survival rate is used as the new parent space; the process of dividing the parent space with the goal of maximizing the increase in survival rate after partitioning is repeated; during the loop, if there is a separated subspace adjacent to the parent space and the difference in survival rate between the two is less than a difference threshold, the separated subspace is merged with the parent space; when the loop stops, the partitioned sample space is obtained, and the risk prediction function for the sample objects in each subspace is determined respectively. The model building module is used to build a risk prediction model based on the risk prediction function corresponding to each of the subspaces; the risk prediction model is used to determine the target subspace to which the object to be predicted belongs, and to predict the risk state of the object to be predicted based on the risk prediction function of the target subspace.

8. The apparatus according to claim 7, characterized in that, The binning module is also used to: perform initial binning on the sample data corresponding to each variable related to risk prediction, so as to obtain multiple initial bins corresponding to the variable. For each initial bin, the binning value of the initial bin is determined based on the proportion of abnormal and non-abnormal samples belonging to the sample objects in the initial bin. With the goal of making the binning value of the bin positively correlated with the proportion of abnormal samples belonging to the sample objects in the bin, the multiple initial bins are adjusted to obtain multiple bins corresponding to the variable.

9. The apparatus according to claim 7, characterized in that, The variable screening module is further configured to: for each variable, evaluate the risk prediction capability of the variable based on the binning values ​​of the multiple bins corresponding to the variable, and obtain the prediction capability evaluation value of the variable; and identify variables whose prediction capability evaluation values ​​are greater than the evaluation threshold as target variables for achieving the risk prediction capability standard.

10. The apparatus according to claim 7, characterized in that, The device further includes: The space merging module is used to determine the number of sample objects in each stripped subspace if there are multiple stripped subspaces adjacent to the parent space and the difference in survival rate between them and the parent space is less than a difference threshold; and merge the stripped subspace with the largest number of objects with the parent space.

11. The apparatus according to claim 7, characterized in that, The device further includes: The survival analysis function construction module is used to, for each sample object, take the number of days during which the sample object does not exhibit any abnormal behavior within a target time period as the survival days of the sample object; the target time period is the time period consisting of the start time when the sample object generates resource transfer records and the end time when the sample data of the sample object is acquired; by determining whether the sample object has exhibited abnormal behavior up to the end time, the value of the survival label of the sample object is determined; based on the sum of the survival labels of all sample objects in the space and the sum of the survival days of all sample objects, a survival analysis function is constructed to determine the survival rate of each sample object in the space.

12. The apparatus according to claim 7, characterized in that, The model building module is further configured to: for each subspace, obtain the abnormal time points of sample objects that have abnormal behavior in the subspace, count the number of sample objects that have abnormal behavior at each abnormal time point; fit the trend of the number of sample objects that have abnormal behavior over time to obtain the risk prediction function of the sample objects in the subspace.

13. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 6.

14. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 6.

15. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Feature binning method, electronic equipment and storage medium

    CN113052222A

  • Credit risk assessment method and device, storage medium and equipment

    CN113177839A