Method and apparatus for training a risk control model

By obtaining samples from the target organization and building a scene simulation sample set, and training the risk control model, the problem of poor performance of the risk control model in predicting the risk level of the institutional sample is solved, and better prediction effect and model adaptability are achieved.

CN114819695BActive Publication Date: 2025-06-24ANT BLOCKCHAIN TECHNOLOGY (SHANGHAI) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210509617.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-05-11
Publication Date
2025-06-24
Estimated Expiration
2042-05-11

AI Technical Summary

Technical Problem

When predicting the risk level of institutional samples, the risk control model is affected by the difference in the data distribution between the own sample and the institutional sample, resulting in poor prediction results, especially in the case of insufficient sample size and overfitting of the model.

Method used

By obtaining the target sample set from the target organization, traversing the risk control strategy library to determine the strategy set, filtering the training party's own risk sample library, building a scene simulation sample set, and using this sample set for two-stage training to improve the prediction effect of the risk control model.

Benefits of technology

It significantly improves the prediction effect of the risk control model on institutional samples, enhances the model's scenario adaptability, alleviates the problem of sample imbalance, and improves the generalization ability of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114819695B_ABST
    Figure CN114819695B_ABST
Patent Text Reader

Abstract

An embodiment of this specification provides a method and device for training a risk control model. The risk control model is to be provided for a target institution to use. The method is executed by a training party and includes: obtaining a target sample set from the target institution, including a plurality of first samples, where the first samples correspond to business objects of the target institution; traversing each risk control strategy in the target strategy library to determine a strategy set composed of at least one risk control strategy that all the first samples in the target sample set satisfy; the risk control strategy constrains at least one attribute feature of the business object; screening the samples included in the risk sample library owned by the training party, and selecting a number of second samples that satisfy the risk control strategies in the strategy set to form a scenario simulation sample set; the second samples are from the business object information collected by the training party and have risk labels; using the scenario simulation sample set to train the risk control model. It can improve the prediction effect of the risk control model obtained by the training party for institutional samples.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] One or more embodiments of the present specification relate to the field of computers, and in particular, to methods and devices for training risk control models. Background Art

[0002] The risk control model can be generated by the training party using its own samples for training and provided to other institutions for use in predicting the risk level of the institutional samples of the institution. Due to the differences in data distribution between the own samples and the institutional samples, the prediction effect of the risk control model for the institutional samples may not be good. In order to solve the contradiction between general products and specific customer needs to a certain extent, generally speaking, customized modeling can be used to solve it. However, in actual work, customized modeling work has problems such as insufficient sample size and overfitting of the model, and the prediction effect of the model is difficult to guarantee. For example, the small and micro business loan business of many institutions is in the cold start stage, the accumulated sample size is too small, and the risk performance is not sufficient. If the institutional samples are directly used for modeling, it is often difficult to guarantee the generalization ability of the model. Summary of the invention

[0003] One or more embodiments of this specification describe a method and device for training a risk control model, which can improve the prediction effect of the risk control model obtained by the trainer on institutional samples.

[0004] In a first aspect, a method for training a risk control model is provided, wherein the risk control model is to be provided to a target institution for use, and the method is performed by a training party, comprising:

[0005] Acquire a target sample set from the target organization, the target sample set comprising a plurality of first samples, the first samples corresponding to the business objects of the target organization;

[0006] Traversing each risk control strategy in the target strategy library, determining a strategy set consisting of at least one risk control strategy satisfied by each first sample in the target sample set; the risk control strategy constrains at least one attribute feature of the business object;

[0007] Screening the samples included in the risk sample library owned by the training party, and selecting a number of second samples that meet the risk control strategies in the strategy set to form a scenario simulation sample set; the second samples come from the business object information collected by the training party and have risk labels;

[0008] The risk control model is trained using the scenario simulation sample set.

[0009] In a possible implementation, the risk label is used to indicate the risk level of the business object, and different risk levels represent differences in the level of risk of the business object.

[0010] In a possible implementation, the business object includes an enterprise entity, and the at least one attribute feature includes several of the following: enterprise registration duration, enterprise owner information, enterprise location, and enterprise industry.

[0011] In a possible implementation, training the risk control model using the scenario simulation sample set includes:

[0012] Using the scenario simulation sample set to perform the first-stage training on the risk control model;

[0013] Determine the sample gradients of each second sample in the scenario simulation sample set for the risk control model after the first-stage training, bin the sample gradients to obtain several gradient bins, and calculate the gradient density of each gradient bin;

[0014] In the second-stage training, for any second sample, determine the sample loss of the second sample according to the prediction loss of the second sample and the gradient density corresponding to its sample gradient; update the risk control model according to the sample losses of each second sample.

[0015] Further, using the scenario simulation sample set to perform the first-stage training on the risk control model includes:

[0016] For any second sample, input the feature data of the second sample corresponding to each attribute feature into the risk control model, and output the risk control score of the second sample through the risk control model;

[0017] According to the risk control score and the risk label of the second sample, use the first loss function to calculate the prediction loss of the second sample;

[0018] According to the prediction losses of each second sample, use the gradient descent method to adjust the model parameters of the risk control model with the goal of minimizing the total prediction loss.

[0019] Further, updating the risk control model according to the sample losses of each second sample includes:

[0020] According to the sample losses of each second sample, use the gradient descent method to adjust the model parameters of the risk control model with the goal of minimizing the total sample loss.

[0021] Further, the sample loss of the second sample is positively correlated with the prediction loss of the second sample and negatively correlated with the gradient density corresponding to its sample gradient.

[0022] Further, the first-stage training includes repeating the training N times using the scenario simulation sample set, where the value of N is a preset value; the second-stage training includes repeating the training M times using the scenario simulation sample set, where the value of M is determined by the convergence of model parameters.

[0023] In a possible implementation manner, the method further includes:

[0024] Sending the trained risk control model to the target institution, so that the target institution uses the feature data corresponding to each attribute feature of the third sample it has and the risk control model to determine the risk level of the third sample.

[0025] In a second aspect, there is provided an apparatus for training a risk control model, where the risk control model is to be provided for use by a target institution, and the apparatus is disposed in a training party and includes:

[0026] An acquisition unit, configured to acquire a target sample set from the target institution, where the target sample set includes a plurality of first samples, and the first samples correspond to the business objects of the target institution;

[0027] A determination unit, configured to traverse each risk control policy in a target policy library, and determine a policy set composed of at least one risk control policy that all the first samples in the target sample set acquired by the acquisition unit satisfy; the risk control policy restricts at least one attribute feature of a business object;

[0028] A selection unit, configured to screen the samples included in the risk sample library owned by the training party, and select a plurality of second samples that satisfy the risk control policies in the policy set determined by the determination unit to form a scenario simulation sample set; the second samples are from the business object information collected by the training party and have risk labels;

[0029] A training unit, configured to train the risk control model using the scenario simulation sample set selected by the selection unit.

[0030] In a third aspect, there is provided a computer-readable storage medium, on which a computer program is stored. When the computer program is executed on a computer, the computer is made to execute the method of the first aspect.

[0031] In a fourth aspect, there is provided a computing device, including a memory and a processor, where an executable code is stored in the memory, and when the processor executes the executable code, the method of the first aspect is implemented.

[0032] Through the methods and devices provided in the embodiments of this specification, first, the training party obtains a target sample set from the target institution. The target sample set includes a plurality of first samples, and the first samples correspond to the business objects of the target institution. Then, each risk control policy in the target policy library is traversed to determine a policy set composed of at least one risk control policy that is satisfied by each of the first samples in the target sample set. The risk control policy restricts at least one attribute feature of the business object. Next, the samples included in the risk sample library owned by the training party are screened, and several second samples that satisfy the risk control policies in the policy set are selected to form a scenario simulation sample set. The second samples are from the business object information collected by the training party and have risk labels. Finally, the risk control model is trained using the scenario simulation sample set. As can be seen from the above, compared with the training party directly using its own samples for modeling, the embodiments of this specification add a screening process for its own samples. Through scenario simulation, based on its own sample data, samples approximately conforming to the institutional scenario in its own risk sample library are specifically mined, thereby explicitly increasing the scale of training samples in the customized modeling process, enhancing the scenario adaptation ability of the model, and being able to improve the prediction effect of the risk control model obtained by the training party for institutional samples. BRIEF DESCRIPTION OF THE DRAWINGS

[0033] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for the description of the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present invention, and those of ordinary skill in the art can obtain other drawings without creative efforts based on these drawings.

[0034] Figure 1 Schematic diagram of an implementation scenario of an embodiment disclosed in this specification;

[0035] Figure 2 Flowchart showing a method for training a risk control model according to an embodiment;

[0036] Figure 3 Schematic diagram showing a way to determine a scenario simulation sample set according to an embodiment;

[0037] Figure 4 Schematic block diagram showing an apparatus for training a risk control model according to an embodiment. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0038] The solutions provided in this specification will be described below with reference to the drawings.

[0039] Figure 1 Schematic diagram of an implementation scenario of an embodiment disclosed in this specification. This implementation scenario relates to training a risk control model. Refer to Figure 1, the entity executing model training is the training party, and the trained risk control model is to be provided for use by the target institution. The training party has a risk sample library, and each of the proprietary samples in the risk sample library comes from the business object information collected by the training party and has risk labels; the target institution has institutional samples, and the institutional samples correspond to the business objects of the target institution, usually only having the feature data of the attribute characteristics of the business objects and not having risk labels, and the risk level of the institutional samples can be predicted through the trained risk control model.

[0040] Among them, the risk control model can be understood as a mathematical model, and the probability output by the data-driven mathematical model is used as a score to judge the risk degree of the evaluation subject.

[0041] In the embodiments of this specification, the business field of the training party may be different from that of the target institution, or rather, the business field of the training party may be wider than that of the target institution, so that there are more or less differences in the data distribution of their samples.

[0042] In the embodiments of this specification, the samples included in the training party's own risk sample library are screened, and the samples approximately meeting the institutional scenario are selected, and the selected samples are used to train the risk control model, so as to improve the prediction effect of the risk control model for institutional samples.

[0043] Among them, the aforementioned training party or target institution can be implemented as any device, platform, server or device cluster with computing and processing capabilities.

[0044] Figure 2 The flowchart of the method for training a risk control model according to an embodiment is shown. The risk control model is to be provided for use by the target institution, and the method is executed by the training party. This method can be based on Figure 1 the implementation scenario shown. As Figure 2 shown, the method for training the risk control model in this embodiment includes the following steps: Step 21, obtain a target sample set from the target institution, where the target sample set includes a plurality of first samples, and the first samples correspond to the business objects of the target institution; Step 22, traverse each risk control strategy in the target strategy library, and determine a strategy set composed of at least one risk control strategy that all the first samples in the target sample set satisfy; the risk control strategy restricts at least one attribute characteristic of the business object; Step 23, screen the samples included in the training party's own risk sample library, and select several second samples that satisfy the risk control strategies in the strategy set to form a scenario simulation sample set; the second samples come from the business object information collected by the training party and have risk labels; Step 24, use the scenario simulation sample set to train the risk control model. The specific execution methods of the above steps are described below.

[0045] First, in step 21, a target sample set is obtained from the target institution. The target sample set includes a plurality of first samples, and each first sample corresponds to a business object of the target institution. It can be understood that the first sample can be equivalent to the aforementioned institution sample, usually without a risk label, and is adapted to the business type of the target institution. The feature values of the various attribute features of the first sample have corresponding data distributions.

[0046] In the embodiments of this specification, as the evaluation subject of risk control, the corresponding business object of the sample can be an individual, an enterprise, a group, etc.

[0047] Generally, the target institution has at least one pre-set risk control strategy, and the risk control strategy restricts at least one attribute feature of the business object. The first sample corresponds to a business object that satisfies the aforementioned at least one risk control strategy.

[0048] In one example, the risk label is used to indicate the risk level of the business object, and different risk levels represent differences in the risk levels of the business object.

[0049] In the embodiments of this specification, the setting of the risk level is relatively flexible. For example, it can be set to two risk levels, high and low, or three risk levels, high, medium, and low, or more risk levels.

[0050] In one example, the business object includes an enterprise entity, and the at least one attribute feature includes the following several: enterprise registration duration, enterprise owner information, enterprise location, and enterprise industry.

[0051] In the embodiments of this specification, the setting of the risk control strategy is also relatively flexible. For example, the limitation of enterprise owner information, whether the enterprise has been pre-determined as a preset category, etc. For example, in a specific risk control strategy, it is limited that the enterprise should not be listed in the list of dishonest executors.

[0052] Then, in step 22, each risk control strategy in the target strategy library is traversed, and a strategy set composed of at least one risk control strategy that each first sample in the target sample set satisfies is determined; the risk control strategy restricts at least one attribute feature of the business object. It can be understood that the target strategy library can be sorted out by the training party and includes risk control strategies that each institution may use. For example, it includes risk control strategies that institutions A, B, and C may use respectively.

[0053] In the embodiments of this specification, each institution may use one risk control strategy or multiple risk control strategies, and different institutions may use the same risk control strategy. The strategy set composed of at least one risk control strategy that each first sample in the target sample set satisfies is the strategy set used by the target institution.

[0054] It can be understood that for samples that do not meet the risk control strategy, it means that these samples are determined to have risks through the risk control strategy, or belong to a certain higher risk level; the first samples have been screened using the risk control strategies in the aforementioned strategy set, and these risk control strategies cannot further determine the risk level of the first samples. That is to say, these risk control strategies have no discrimination ability for the first samples.

[0055] Then in step 23, the samples included in the risk sample library owned by the training party are screened, and several second samples that meet the risk control strategies in the strategy set are selected to form a scenario simulation sample set; the second samples are from the business object information collected by the training party and have risk labels. It can be understood that the business object information collected by the training party does not necessarily belong to the business objects of the target institution. Therefore, the data distribution of the second samples is usually different from that of the first samples. Through the above screening process, the risk control process of the target institution can be simulated, and a scenario simulation sample set can be produced, which is more in line with the business logic in the sample screening.

[0056] In the embodiments of this specification, the training party sorts out the real risk samples generated in its own business and attaches corresponding risk labels to obtain the aforementioned risk sample library. For example: if an enterprise has an overdue period of more than 30 consecutive days within 12 months from the date of disbursement, it is determined as a black sample; if the overdue period is 0 days, it is determined as a white sample; if the overdue period is 1 to 29 days, it is determined as a gray sample, etc. It can be understood that black samples, gray samples, and white samples can be used as risk labels, and the risks are in descending order as black samples, gray samples, and white samples.

[0057] Figure 3 Shows a schematic diagram of the determination method of a scenario simulation sample set according to an embodiment. Refer to Figure 3 , first match the target sample set with the risk control strategies in the target strategy library, traverse each risk control strategy in the target strategy library, and determine a strategy set composed of at least one risk control strategy that each first sample in the target sample set meets. Then use this strategy set to screen the samples included in the risk sample library owned by the training party, and select several second samples that meet the risk control strategies in the strategy set to form a scenario simulation sample set; among them, the process of sample screening can also be called strategy cleaning, which is a process of filtering out some samples from the risk sample library.

[0058] Among them, the number of samples in the scenario simulation sample set is less than that in the risk sample library. Samples that do not meet the risk control strategies in the strategy set will be filtered out from the risk sample library to obtain the scenario simulation sample set.

[0059] Finally, in step 24, the risk control model is trained using the scenario simulation sample set. It can be understood that the risk control model can be trained using the usual training method, or some improvements can be made on the basis of the usual training method to train the risk control model, so as to further improve the training effect of the model.

[0060] In one example, the training of the risk control model using the scenario simulation sample set includes:

[0061] The risk control model is trained in the first stage using the scenario simulation sample set;

[0062] Determine the sample gradients of each second sample in the scenario simulation sample set for the risk control model after the first stage of training, bin the sample gradients to obtain several gradient bins, and calculate the gradient density of each gradient bin;

[0063] In the second stage of training, for any second sample, determine the sample loss of the second sample according to the prediction loss of the second sample and the gradient density corresponding to its sample gradient; update the risk control model according to the sample losses of each second sample.

[0064] In this example, two-stage training is adopted. In the first stage of training, the conventional training method is used. After the first stage of training, the magnitude of the sample gradient can reflect the difficulty of classifying the corresponding sample. The smaller the sample gradient, the easier the corresponding sample is to classify, and the larger the sample gradient, the more difficult the corresponding sample is to classify. By binning the sample gradients, each sample gradient with a relatively close magnitude can be grouped into a gradient bin. Generally, the number of samples that are easy to classify is large, and the number of samples that are difficult to classify is small. For samples that are easy to classify, the gradient is already very low, the number of sample gradients in the gradient bin is large, and the gradient density is high. For samples that are difficult to classify, the gradient is large, the number of sample gradients in the gradient bin is small, and the gradient density is low. In the second stage of training, for any second sample, determine the sample loss of the second sample according to the prediction loss of the second sample and the gradient density corresponding to its sample gradient, which can make the model training pay more attention to the difficult-to-classify samples, so as to achieve the effect of difficult sample mining.

[0065] In the embodiments of this specification, since the risk sample library owned by the training party has gone through the scenario simulation link, that is, it has been screened by the pre-risk control strategy, the data distribution of black and white samples will gradually change, generally resulting in a decrease in the number of samples, and at the same time, a lower black-and-white sample ratio, that is, a lower proportion of black samples. The risk control problem is inherently a data imbalance problem, and it will be more serious after using scenario simulation. Therefore, by using the training method of difficult sample mining, the impact brought by the sample imbalance problem can be alleviated, thereby improving the training effect of the model.

[0066] Further, the first-stage training of the risk control model using the scenario simulation sample set includes:

[0067] For any second sample, input the feature data corresponding to each attribute feature of the second sample into the risk control model, and output the risk score of the second sample through the risk control model;

[0068] According to the risk score and the risk label of the second sample, use the first loss function to calculate the prediction loss of the second sample;

[0069] According to the prediction losses of each second sample, use the gradient descent method to adjust the model parameters of the risk control model with the goal of minimizing the total prediction loss.

[0070] It can be understood that after the first-stage training, for samples that are easily classified, the gradient is already very low, while for samples that are difficult to classify, the gradient is relatively large.

[0071] Further, the updating of the risk control model according to the sample losses of each second sample includes:

[0072] According to the sample losses of each second sample, use the gradient descent method to adjust the model parameters of the risk control model with the goal of minimizing the total sample loss.

[0073] It can be understood that the total sample loss is the sum of the sample losses of each second sample. To achieve the goal of minimizing the total sample loss, more attention needs to be paid to the second samples with larger sample losses, that is, more attention needs to be paid to the samples that are difficult to classify.

[0074] Further, the sample loss of the second sample is positively correlated with the prediction loss of the second sample and negatively correlated with the gradient density corresponding to its sample gradient.

[0075] It can be understood that the gradient density corresponding to the samples that are difficult to classify is relatively low, and their sample losses will be amplified to a greater extent based on their prediction losses, making the sample losses of the samples that are difficult to classify larger than those of the samples that are easily classified.

[0076] Further, the first-stage training includes repeating the training N times using the scenario simulation sample set, where N is a pre-set value; the second-stage training includes repeating the training M times using the scenario simulation sample set, where M is determined by the convergence of the model parameters.

[0077] It can be understood that one training corresponds to one epoch, which refers to the process of sending all the data in the scenario simulation sample set into the risk control model to complete one forward calculation and backpropagation process. If the amount of data is too large to send all the data into the model at one time, the data can be sent into the model in batches. During training, it is not enough to iterate all the data once, and it needs to be repeated many times to fit and converge.

[0078] In the embodiment of this specification, the training process based on hard sample mining in step 24 may include the following processing procedures:

[0079] Step 241, in the first N epochs, perform normal training and normally calculate the prediction loss loss = L(i) corresponding to each sample. Here, L represents the loss function and i represents the sample identifier.

[0080] Step 242, before the start of the next epoch, calculate the gradients of all samples, bin the gradients, and calculate the density of each gradient bin, denoted as the gradient density gr(i). It can be understood that the gradient density of any sample is the gradient density of the gradient bin corresponding to this sample.

[0081] Step 243, perform the next epoch and calculate the sample loss loss = L(i) / gr(i) corresponding to each sample.

[0082] Among them, the processing procedures of step 242 and step 243 are looped until the model converges.

[0083] In the embodiment of this specification, in the training process based on hard sample mining, the gradient density of each sample will be calculated after N epochs. At this time, for the samples that are easy to classify, the gradients are already very low and the gradient densities are very high, and they are aggregated and distributed in the lower partition of the gradient. For the samples that are not easy to classify, the gradients are larger and the gradient densities are lower, and they are scattered and distributed in the higher partition of the gradient. Therefore, the sample losses of the hard-to-classify samples will be magnified to a greater extent, enabling the model to pay more attention to the hard-to-classify samples, thus achieving the effect of hard sample mining.

[0084] In one example, the method further includes:

[0085] Send the trained risk control model to the target institution, so that the target institution can use the feature data corresponding to each attribute feature of the third sample it has and the risk control model to determine the risk level of the third sample.

[0086] It can be understood that the third sample corresponds to the business object of the target institution.

[0087] In the embodiments of this specification, the training party can produce a customized model trained based on its own data to adapt to the data distribution of the target institution, and its effect is better than directly using the samples of the target institution for customization, or using the samples in its own samples that have the same distribution as the target samples in certain basic features for customization and other methods.

[0088] Through the method provided by the embodiments of this specification, first, the training party obtains a target sample set from the target institution, and the target sample set includes a plurality of first samples, and the first samples correspond to the business objects of the target institution; then, traverse each risk control strategy in the target strategy library, and determine a strategy set composed of at least one risk control strategy that all the first samples in the target sample set satisfy; the risk control strategy restricts at least one attribute feature of the business object; then, screen the samples included in the training party's own risk sample library, and select several second samples that satisfy the risk control strategies in the strategy set to form a scenario simulation sample set; the second samples are from the business object information collected by the training party and have risk labels; finally, use the scenario simulation sample set to train the risk control model. As can be seen from the above, in the embodiments of this specification, compared with the training party directly using its own samples for modeling, the screening process for its own samples is increased. Through scenario simulation, based on its own sample data, samples in its own risk sample library that approximately conform to the institutional scenario are specifically mined, thereby explicitly increasing the scale of training samples in the customized modeling process, enhancing the scenario adaptation ability of the model, and being able to improve the prediction effect of the risk control model obtained by the training party for institutional samples.

[0089] According to an embodiment of another aspect, there is also provided a device for training a risk control model, and the risk control model is to be provided for use by a target institution. The device is set in the training party, and this device is used to execute the method provided by the embodiments of this specification. Figure 4 The schematic block diagram of a device for training a risk control model according to an embodiment is shown. As Figure 4 shown, the device 400 includes:

[0090] An obtaining unit 41, configured to obtain a target sample set from the target institution, where the target sample set includes a plurality of first samples, and the first samples correspond to the business objects of the target institution;

[0091] A determining unit 42, configured to traverse each risk control strategy in the target strategy library, and determine a strategy set composed of at least one risk control strategy that all the first samples in the target sample set obtained by the obtaining unit 41 satisfy; the risk control strategy restricts at least one attribute feature of the business object;

[0092] A selection unit 43 is configured to screen the samples included in the risk sample library owned by the training party, and select a number of second samples that meet the risk control policies in the policy set determined by the determination unit 42 to form a scenario simulation sample set; the second samples are from the business object information collected by the training party and have risk labels.

[0093] A training unit 44 is configured to train the risk control model by using the scenario simulation sample set selected by the selection unit 43.

[0094] Optionally, as an embodiment, the risk label is used to indicate the risk level of the business object, and different risk levels represent differences in the risk levels of the business object.

[0095] Optionally, as an embodiment, the business object includes an enterprise entity, and the at least one attribute feature includes several of the following: enterprise registration duration, enterprise owner information, enterprise location, and enterprise industry.

[0096] Optionally, as an embodiment, the training unit 44 includes:

[0097] A first training subunit is configured to perform a first-stage training on the risk control model by using the scenario simulation sample set.

[0098] A density calculation subunit is configured to determine the sample gradients of each second sample in the scenario simulation sample set with respect to the risk control model after the first-stage training obtained by the first training subunit, perform binning on each sample gradient to obtain a number of gradient bins, and calculate the gradient density of each gradient bin.

[0099] A second training subunit is configured to, in the second-stage training, for any second sample, determine the sample loss of the second sample according to the prediction loss of the second sample and the gradient density corresponding to the sample gradient obtained by the density calculation subunit; update the risk control model according to the sample losses of each second sample.

[0100] Furthermore, the first training subunit includes:

[0101] A prediction module is configured to, for any second sample, input the feature data of the second sample corresponding to each attribute feature into the risk control model, and output the risk control score of the second sample through the risk control model.

[0102] A calculation module is configured to calculate the prediction loss of the second sample by using a first loss function according to the risk control score obtained by the prediction module and the risk label of the second sample.

[0103] A parameter adjustment module, configured to adjust model parameters of the risk control model by using the gradient descent method with the goal of minimizing the total prediction loss according to the prediction losses of each second sample obtained by the calculation module.

[0104] Further, the second training subunit is specifically configured to adjust model parameters of the risk control model by using the gradient descent method with the goal of minimizing the total sample loss according to the sample losses of each second sample.

[0105] Further, the sample loss of a second sample is positively correlated with the prediction loss of the second sample and negatively correlated with the gradient density corresponding to the sample gradient of the second sample.

[0106] Further, the first-stage training includes repeating training N times by using the scenario simulation sample set, where the value of N is a preset value; the second-stage training includes repeating training M times by using the scenario simulation sample set, where the value of M is determined by the convergence of model parameters.

[0107] Optionally, as an embodiment, the device further includes:

[0108] A sending unit, configured to send the trained risk control model obtained by the training unit 44 to the target institution, so that the target institution determines the risk level of a third sample by using the feature data corresponding to each attribute feature of the third sample owned by the target institution and the risk control model.

[0109] Through the device provided in the embodiments of this specification, first, an acquisition unit 41 of the training party acquires a target sample set from the target institution. The target sample set includes multiple first samples, and the first samples correspond to business objects of the target institution. Then, a determination unit 42 traverses each risk control strategy in the target policy library and determines a policy set composed of at least one risk control strategy that each first sample in the target sample set satisfies. The risk control strategy restricts at least one attribute feature of the business object. Next, a selection unit 43 screens the samples included in the risk sample library owned by the training party and selects several second samples that satisfy the risk control strategies in the policy set to form a scenario simulation sample set. The second samples are from the business object information collected by the training party and have risk labels. Finally, a training unit 44 trains the risk control model by using the scenario simulation sample set. As can be seen from the above, compared with the training party directly using its own samples for modeling, the embodiments of this specification add a screening process for its own samples. Through scenario simulation, based on its own sample data, samples approximately conforming to the institution scenario in its own risk sample library are specifically mined, thereby explicitly increasing the scale of training samples in the customized modeling process, enhancing the scenario adaptation ability of the model, and being able to improve the prediction effect of the risk control model obtained by the training party for institution samples.

[0110] According to an embodiment of another aspect, there is also provided a computer-readable storage medium having stored thereon a computer program which, when executed on a computer, causes the computer to execute the method in combination with Figure 2 the method described.

[0111] According to an embodiment of still another aspect, there is also provided a computing device including a memory and a processor, wherein the memory stores executable code, and when the processor executes the executable code, the method described in combination with Figure 2 is implemented.

[0112] Those skilled in the art should be able to realize that in one or more of the above examples, the functions described in the present invention can be implemented by hardware, software, firmware, or any combination thereof. When implemented using software, these functions can be stored in a computer-readable medium or transmitted as one or more instructions or codes on a computer-readable medium.

[0113] The specific embodiments described above further elaborate on the objectives, technical solutions, and beneficial effects of the present invention. It should be understood that the above description is only for the specific embodiments of the present invention and is not used to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made on the basis of the technical solutions of the present invention shall be included within the protection scope of the present invention.

Claims

1. A method for training a risk control model, the risk control model to be provided for use by a target institution, the method being executed by a training party, comprising: Obtaining a target sample set from the target institution, the target sample set including a plurality of first samples, the first samples corresponding to business objects of the target institution; Traversing each risk control strategy in the target strategy library, and determining a strategy set composed of at least one risk control strategy that all the first samples in the target sample set satisfy; the risk control strategy constrains at least one attribute feature of the business object; Screening the samples included in the risk sample library owned by the training party, and selecting a number of second samples that satisfy the risk control strategies in the strategy set to form a scenario simulation sample set; the second samples are from the business object information collected by the training party and have risk labels; Training the risk control model using the scenario simulation sample set.

2. The method according to claim 1, wherein The risk label is used to indicate the risk level of the business object, and different risk levels represent differences in the risk levels of the business objects.

3. The method according to claim 1, wherein The business object includes an enterprise entity, and the at least one attribute feature includes the following several: enterprise registration duration, enterprise owner information, enterprise location, enterprise industry.

4. The method according to claim 1, wherein, The training the risk control model using the scenario simulation sample set includes: Performing a first-stage training on the risk control model using the scenario simulation sample set; Determining the sample gradient of each second sample in the scenario simulation sample set for the risk control model after the first-stage training, binning each sample gradient to obtain several gradient bins, and calculating the gradient density of each gradient bin; In the second-stage training, for any second sample, determining the sample loss of the second sample according to the prediction loss of the second sample and the gradient density corresponding to its sample gradient; updating the risk control model according to the sample losses of each second sample.

5. The method according to claim 4, wherein, The performing a first-stage training on the risk control model using the scenario simulation sample set includes: For any second sample, inputting the feature data corresponding to each attribute feature of the second sample into the risk control model, and outputting the risk control score of the second sample through the risk control model; Calculating the prediction loss of the second sample using a first loss function according to the risk control score and the risk label of the second sample; According to the prediction losses of each second sample, using the gradient descent method, with the goal of minimizing the total prediction loss, adjusting the model parameters of the risk control model.

6. The method according to claim 4, wherein The updating the risk control model according to the sample losses of each second sample includes: According to the sample losses of each second sample, using the gradient descent method, with the goal of minimizing the total sample loss, adjusting the model parameters of the risk control model.

7. The method according to claim 4, wherein The sample loss of the second sample is positively correlated with the prediction loss of the second sample and negatively correlated with the gradient density corresponding to its sample gradient.

8. The method according to claim 4, wherein, The first-stage training includes repeating the training N times using the scenario simulation sample set, where the value of N is a preset value; the second-stage training includes repeating the training M times using the scenario simulation sample set, where the value of M is determined by the convergence of the model parameters.

9. The method according to claim 1, wherein The method further includes: Send the trained risk control model to the target institution so that the target institution can use the feature data corresponding to each attribute feature of the third sample it has and the risk control model to determine the risk level of the third sample.

10. An apparatus for training a risk control model, the risk control model to be provided for use by a target institution, the apparatus being provided at a training party and comprising: An acquisition unit configured to acquire a target sample set from the target institution, the target sample set including a plurality of first samples, the first samples corresponding to business objects of the target institution; A determination unit configured to traverse each risk control strategy in a target policy library and determine a set of strategies composed of at least one risk control strategy that is satisfied by each of the first samples in the target sample set acquired by the acquisition unit; the risk control strategy constrains at least one attribute feature of a business object; A selection unit configured to screen the samples included in the risk sample library owned by the training party and select a plurality of second samples that satisfy the risk control strategies in the set of strategies determined by the determination unit to form a scenario simulation sample set; the second samples are from business object information collected by the training party and have risk labels; A training unit configured to train the risk control model by using the scenario simulation sample set selected by the selection unit.

11. The apparatus according to claim 10, wherein, The risk label is used to indicate the risk level of a business object, and different risk levels represent differences in the risk levels of business objects.

12. The device according to claim 10, wherein, The business object includes an enterprise entity, and the at least one attribute feature includes several of the following: enterprise registration duration, enterprise owner information, enterprise location, and enterprise industry.

13. The apparatus according to claim 10, wherein, The training unit includes: A first training subunit configured to perform a first-stage training on the risk control model by using the scenario simulation sample set; A density calculation subunit configured to determine the sample gradients of each of the second samples in the scenario simulation sample set with respect to the risk control model after the first-stage training obtained by the first training subunit, perform binning on each of the sample gradients to obtain a plurality of gradient bins, and calculate the gradient density of each gradient bin; A second training subunit configured to, in a second-stage training, for any second sample, determine the sample loss of the second sample according to the prediction loss of the second sample and the gradient density corresponding to the sample gradient of the second sample obtained by the density calculation subunit; update the risk control model according to the sample losses of each of the second samples.

14. The apparatus according to claim 13, wherein, The first training subunit includes: A prediction module configured to, for any second sample, input the feature data corresponding to each attribute feature of the second sample into the risk control model and output the risk control score of the second sample through the risk control model; A calculation module configured to calculate the prediction loss of the second sample by using a first loss function according to the risk control score obtained by the prediction module and the risk label of the second sample; A parameter adjustment module configured to adjust the model parameters of the risk control model by using the gradient descent method with the goal of minimizing the total prediction loss according to the prediction losses of each of the second samples obtained by the calculation module.

15. The device according to claim 13, wherein, The second training subunit is specifically configured to adjust the model parameters of the risk control model by using the gradient descent method with the goal of minimizing the total sample loss according to the sample losses of each second sample.

16. The device according to claim 13, wherein, The sample loss of the second sample is positively correlated with the predicted loss of the second sample and negatively correlated with the gradient density corresponding to its sample gradient.

17. The device according to claim 13, wherein, The first-stage training includes repeating N times of training by using the scenario simulation sample set, where the value of N is a preset value; the second-stage training includes repeating M times of training by using the scenario simulation sample set, and the value of M is determined by the convergence of the model parameters.

18. The device according to claim 10, wherein, The device further includes: A sending unit, configured to send the trained risk control model obtained by the training unit to the target institution, so that the target institution uses the feature data corresponding to each attribute feature of the third sample and the risk control model it has to determine the risk level of the third sample.

19. A computer-readable storage medium, on which a computer program is stored. When the computer program is executed on a computer, the computer is made to execute the method according to any one of claims 1-9.

20. A computing device, including a memory and a processor. An executable code is stored in the memory. When the processor executes the executable code, the method according to any one of claims 1-9 is implemented.

Citation Information

Patent Citations

  • Risk control model training method and device and risk category prediction method and device

    CN114997317A

  • Method and device for training risk control model

    CN115907965A