Financial risk control model incremental learning method and device based on XGBoost, equipment and medium

By merging and replacing the splitting points of the XGBoost financial risk control model and using new samples to train the model, the problem that the old tree feature splitting points are not applicable to new samples is solved, and the model's prediction accuracy and generalization ability on new data are improved.

CN120782010APending Publication Date: 2025-10-14CHONGQING YUYIN FINANCIAL TECH CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510958949.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-11
Publication Date
2025-10-14

AI Technical Summary

Technical Problem

During the XGBoost incremental learning process, the splitting points of the old tree features are not applicable to new samples, which greatly reduces the learning effect.

Method used

The split points of all trees in the original financial risk control model are collected, arranged and merged to determine the candidate split points. Based on the group stability index and gain, the reference split points are selected to replace the original split points, and the model is trained using the new samples.

Benefits of technology

The model's prediction accuracy and generalization ability on new data are improved, and the performance and practicality of the model are enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120782010A_ABST
    Figure CN120782010A_ABST
Patent Text Reader

Abstract

The invention discloses a financial risk control model incremental learning method and device based on XGBoost, equipment and a medium, and relates to the technical field of computers, and the method comprises the steps: collecting each split point of all trees in an original financial risk control model for a target financial risk feature, carrying out the arrangement processing of each split point, and obtaining a target financial risk feature; merging the processed split points meeting a preset merging condition to obtain merged split points; binning the old sample and the new sample based on each group of candidate splitting points determined by quantile points of the target financial risk feature in a preset range and the combined splitting points, and determining a reference splitting point based on group stability indexes and gains in each binning, replacing original splitting points for splitting the target financial risk features in the original financial risk control model; and training the obtained replaced financial risk control model based on a training set determined by the new sample to obtain a target financial risk control model. And inapplicability of old tree feature splitting points to new samples in an incremental learning process is avoided.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer technology, and in particular to an XGBoost-based incremental learning method, device, equipment, and medium for a financial risk control model. Background Art

[0002] The XGBoost algorithm, a commonly used machine learning algorithm, is used in many scenarios in the financial sector, including credit scoring modeling, fraud detection, and precision marketing. With the increasing number of online users, model updates are becoming increasingly frequent, even shifting from full training to incremental learning. One approach to incremental learning with XGBoost is to maintain the structure of the trained tree—that is, the features and split points—and only update the values ​​of the leaf nodes. However, the XGBoost algorithm is a greedy algorithm that considers all possible features and split points in the training sample, calculates the gain of each split point, and then selects the split point with the largest gain as the optimal split point. This results in the split point being an overfit of the old training sample. When this approach is used for incremental learning on new samples, the feature split points of the old tree are no longer optimal for the new sample, significantly reducing the effectiveness of incremental learning.

[0003] As can be seen from the above, how to avoid the situation where the split points of old tree features are not applicable to new samples during incremental learning is an urgent problem that needs to be solved. Summary of the Invention

[0004] In view of this, the purpose of the present invention is to provide a method, device, equipment, and medium for incremental learning of financial risk control models based on XGBoost, which can avoid the situation where the split points of old tree features are not applicable to new samples during the incremental learning process. The specific scheme is as follows:

[0005] In the first aspect, this application provides an incremental learning method for financial risk control models based on XGBoost, including:

[0006] Collecting all splitting points of the target financial risk characteristics from all trees in the original financial risk control model, arranging the splitting points to obtain processed splitting points, and merging the processed splitting points that meet preset merging conditions to obtain merged splitting points; the financial risk control model is a financial risk control model determined based on XGBoost;

[0007] Determine each group of candidate splitting points based on the quantile points of the target financial risk feature within a preset range and the merged splitting points, use each group of candidate splitting points to bin the old samples and the new samples, and then determine the reference splitting point of the target financial risk feature based on the population stability index and gain within each bin; the old samples are the historical financial sample data used to train the original financial risk control model; the new samples are the new financial sample data collected after the original financial risk control model is trained;

[0008] Determining an original splitting point in the original financial risk control model for splitting the target financial risk feature, and replacing the original splitting point based on the reference splitting point to obtain a replaced financial risk control model;

[0009] A training set is determined based on the new sample, and the replaced financial risk control model is trained using the training set. If the trained financial risk control model meets a preset termination condition, a target financial risk control model is obtained.

[0010] Optionally, collecting the splitting points of all trees in the original financial risk control model for the target financial risk characteristics and arranging the splitting points to obtain processed splitting points includes:

[0011] Traversing all trees in the original financial risk control model to collect splitting points of all trees for the target financial risk characteristics, and deduplicating the splitting points to obtain deduplicated splitting points;

[0012] The deduplicated splitting points are sorted in ascending order to obtain processed splitting points.

[0013] Optionally, merging the processed split points that meet a preset merging condition to obtain a merged split point includes:

[0014] Determining an interval between two adjacent processed split points, and judging whether the interval meets a preset merging condition;

[0015] If the interval satisfies the preset merging condition, merging two adjacent processed split points that satisfy the preset merging condition to obtain a merged split point;

[0016] Among them, the preset merging condition is that the interval is smaller than the target interval threshold; the target interval threshold is an interval threshold determined based on the difference between the first quantile and the second quantile of the target financial risk feature in all samples; the first quantile is smaller than the second quantile.

[0017] Optionally, determining each group of candidate splitting points based on the quantile points of the target financial risk feature within a preset range and the merged splitting points includes:

[0018] Determining a quantile value based on quantile points within a preset range in the new sample and the old sample of the target financial risk characteristic;

[0019] The number of corresponding splitting points is determined using the merged splitting points, and then the quantile values ​​corresponding to the number of splitting points are selected as the values ​​to be combined, and the values ​​to be combined are combined to obtain each group of candidate splitting points.

[0020] Optionally, binning the old samples and the new samples using each group of candidate splitting points, and then determining the reference splitting point of the target financial risk feature based on the group stability index and gain in each bin, includes:

[0021] Binning the old samples and the new samples using the candidate split points in each group, and determining a population stability index based on the proportion of the old samples and the new samples in each bin;

[0022] The gains of the old samples and the new samples at each group of the candidate splitting points are determined, and the candidate splitting point with the smallest group stability index and the largest gain is determined as the reference splitting point of the target financial risk feature.

[0023] Optionally, determining an original splitting point in the original financial risk control model for splitting the target financial risk feature, and replacing the original splitting point based on the reference splitting point to obtain a replaced financial risk control model includes:

[0024] Determining an original splitting point in the original financial risk control model for splitting the target financial risk characteristics, and comparing the reference splitting point with the original splitting point to obtain a comparison result;

[0025] Based on the comparison result and the proximity principle, the value of the reference splitting point closest to the original splitting point is determined as a replacement value, and the value corresponding to the original splitting point is replaced based on the replacement value to obtain a replaced financial risk control model.

[0026] Optionally, determining a training set based on the new sample, using the training set to train the replaced financial risk control model, and obtaining a target financial risk control model if the trained financial risk control model satisfies a preset termination condition, includes:

[0027] Dividing the new sample based on a preset division ratio to obtain a training set and a validation set;

[0028] The replaced financial risk control model is trained using the training set, and during the training process, the performance of the trained financial risk control model is evaluated at a preset frequency to obtain an evaluation result;

[0029] If the evaluation result shows that the performance of the trained financial risk control model has not improved within the preset number of training times, the training of the trained financial risk control model is stopped, and the target financial risk control model is obtained.

[0030] In a second aspect, the present application provides an incremental learning device for a financial risk control model based on XGBoost, comprising:

[0031] A split point processing module is used to collect the split points of all trees in the original financial risk control model for the target financial risk characteristics, arrange the split points to obtain processed split points, and merge the processed split points that meet the preset merging conditions to obtain merged split points; the financial risk control model is a financial risk control model determined based on XGBoost;

[0032] A sample binning module is configured to determine groups of candidate splitting points based on the quantile points of the target financial risk characteristic within a preset range and the merged splitting points, bin the old and new samples using the candidate splitting points in each group, and then determine a reference splitting point for the target financial risk characteristic based on the population stability index and gain within each bin; the old samples are the historical financial sample data used to train the original financial risk control model; and the new samples are the new financial sample data collected after the original financial risk control model is trained.

[0033] a splitting point replacement module, configured to determine an original splitting point in the original financial risk control model for splitting according to the target financial risk feature, and replace the original splitting point based on the reference splitting point to obtain a replaced financial risk control model;

[0034] The model training module is used to determine a training set based on the new sample, and use the training set to train the replaced financial risk control model. If the trained financial risk control model meets the preset termination condition, the target financial risk control model is obtained.

[0035] In a third aspect, the present application provides an electronic device, comprising:

[0036] Memory, used to store computer programs;

[0037] A processor is used to execute the computer program to implement the aforementioned XGBoost-based incremental learning method for the financial risk control model.

[0038] In a fourth aspect, the present application provides a computer-readable storage medium for storing a computer program, wherein, when the computer program is executed by a processor, it implements the aforementioned XGBoost-based financial risk control model incremental learning method.

[0039] This application collects all splitting points of the target financial risk characteristics from all trees in the original financial risk control model, and arranges each splitting point to obtain a processed splitting point, and merges the processed splitting points that meet the preset merging conditions to obtain a merged splitting point; determines each group of candidate splitting points based on the quantile points of the target financial risk characteristics within a preset range and the merged splitting points, uses each group of candidate splitting points to bin the old samples and the new samples, and then determines the reference splitting point of the target financial risk characteristics based on the group stability index and gain in each bin; the old samples are the sample data used to train the original financial risk control model; the new samples are the new sample data collected after the training of the original financial risk control model is completed; determines the original splitting point for the target financial risk characteristics in the original financial risk control model, and replaces the original splitting point based on the reference splitting point to obtain a replaced financial risk control model; determines a training set based on the new samples, and uses the training set to train the replaced financial risk control model. If the trained financial risk control model meets the preset termination conditions, the target financial risk control model is obtained.

[0040] As can be seen from the above, the present application collects the splitting points of all trees in the original financial risk control model for the target financial risk characteristics, arranges these splitting points, and merges the splitting points that meet the preset merging conditions, which can reduce the redundancy of the splitting points and make the splitting points more concentrated and representative; then, based on the quantile points of the target financial risk characteristics within the preset range and the merged splitting points, each group of candidate splitting points is determined, and the old samples and new samples are binned using each group of candidate splitting points. Then, based on the group stability index (PSI) and gain in each bin, the reference splitting point of the target financial risk characteristics is determined. By comparing the original splitting points for the target financial risk characteristics in the original financial risk control model with the reference splitting points, the optimized splitting points are introduced into the original financial risk control model to obtain the replaced financial risk control model. In this way, the replaced financial risk control model is trained using the training set determined by the new samples, so that the obtained target financial risk control model can better adapt to the feature distribution of the new samples, thereby improving the prediction accuracy and generalization ability of the model on new data, and enhancing the performance and practicality of the model. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are merely embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying any creative work.

[0042] Figure 1This is a flow chart of an incremental learning method for a financial risk control model based on XGBoost disclosed in this application;

[0043] Figure 2 This is a schematic diagram of the structure of an incremental learning device for a financial risk control model based on XGBoost disclosed in this application;

[0044] Figure 3 This is a structural diagram of an electronic device disclosed in this application. DETAILED DESCRIPTION

[0045] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0046] At present, one way of XGBoost incremental learning is to keep the structure of the trained tree unchanged, that is, the features and splitting points remain unchanged, and only update the values ​​of the leaf nodes. However, the XGBoost algorithm is a greedy algorithm that considers all possible features and splitting points on the training sample, calculates the gain of each splitting point, and then selects the splitting point with the largest gain as the optimal splitting point, resulting in the splitting point found being essentially an overfitting of the old training sample. When this method is used to perform incremental learning on new samples, the feature splitting points of the old tree are actually not the optimal splitting points for the new samples, resulting in a significant reduction in the incremental learning effect. To this end, the present application provides an incremental learning method for a financial risk control model based on XGBoost, which uses a training set determined by the new sample to train the replaced financial risk control model, so that the obtained target financial risk control model can better adapt to the feature distribution of the new sample, thereby improving the prediction accuracy and generalization ability of the model on new data, and enhancing the performance and practicality of the model.

[0047] See also Figure 1 As shown, the embodiment of the present invention discloses an incremental learning method for a financial risk control model based on XGBoost, comprising:

[0048] Step S11: Collect the splitting points of all trees in the original financial risk control model for the target financial risk characteristics, and arrange the splitting points to obtain processed splitting points, and merge the processed splitting points that meet the preset merging conditions to obtain merged splitting points; the financial risk control model is a financial risk control model determined based on XGBoost.

[0049] In this embodiment, a financial risk control model for the financial field is determined based on XGBoost, and all trees in the original financial risk control model are traversed to collect the splitting points of all trees for the target financial risk characteristics, and each of the splitting points is deduplicated and arranged in ascending order to obtain the processed splitting points. Specifically, the collection of the splitting points of all trees in the original financial risk control model for the target financial risk characteristics and the arrangement of each of the splitting points to obtain the processed splitting points include: traversing all trees in the original financial risk control model to collect the splitting points of all trees for the target financial risk characteristics, and deduplicating each of the splitting points to obtain the deduplicated splitting points; and arranging the deduplicated splitting points in ascending order to obtain the processed splitting points.

[0050] It can be understood that after obtaining the processed splitting point, the interval between two adjacent processed splitting points is determined, and the interval threshold is determined based on 1 / 20 of the difference between the 10% quantile and the 90% quantile of the target financial risk feature in all samples. If the interval is less than the interval threshold, it indicates that the interval meets the preset merging condition, and the average value of the two adjacent processed splitting points that meet the preset merging condition is determined as the merged splitting point. Specifically, merging the processed splitting points that meet the preset merging condition to obtain the merged splitting point includes: determining the interval between two adjacent processed splitting points, and judging whether the interval meets the preset merging condition; if the interval meets the preset merging condition, merging the two adjacent processed splitting points that meet the preset merging condition to obtain the merged splitting point; wherein the preset merging condition is that the interval is less than the target interval threshold; the target interval threshold is the interval threshold determined based on the difference between the first quantile and the second quantile of the target financial risk feature in all samples; the first quantile is less than the second quantile.

[0051] Step S12: determine each group of candidate splitting points based on the quantile points of the target financial risk characteristics within a preset range and the merged splitting points, use each group of the candidate splitting points to bin the old samples and the new samples, and then determine the reference splitting points of the target financial risk characteristics based on the group stability index and gain in each bin; the old samples are the historical financial sample data used to train the original financial risk control model; the new samples are the new financial sample data collected after the training of the original financial risk control model is completed.

[0052] In this embodiment, after merging the processed splitting points, the number of splitting points corresponding to the merger is determined, and then the historical financial sample data used to train the original financial risk control model is determined as the old sample, and the new financial sample data collected after the training of the original financial risk control model is determined as the new sample. In a specific implementation, the values ​​of the target financial risk feature at the quantile points of 0.05 to 0.95 (interval of 0.05) in all samples are determined, that is, the values ​​of the target financial risk feature at 19 quantile points of 0.05, 0.10, 0.15...0.95 in all samples are determined to obtain 19 quantile values, and then the quantile values ​​corresponding to the number of splitting points are selected from the 19 quantile values ​​to determine the values ​​to be combined, and based on the values ​​to be combined and using different combination methods, each group of candidate splitting points is obtained. Specifically, the determining of each group of candidate splitting points based on the quantile points of the target financial risk characteristics within a preset range and the merged splitting points includes: determining the quantile values ​​based on the quantile points of the target financial risk characteristics within a preset range in the new sample and the old sample; using the merged splitting points to determine the corresponding number of splitting points, and then selecting the quantile values ​​corresponding to the number of splitting points as the values ​​to be combined, and combining the values ​​to be combined to obtain each group of candidate splitting points.

[0053] It can be understood that the old samples and the new samples are binned using each group of candidate splitting points, and a population stability index (PSI) is determined based on the proportion of the old samples and the new samples in each bin, the gains of the old samples and the new samples in each group of candidate splitting points are determined, and the reference splitting point of the target financial risk feature is determined based on the sum of the population stability index and the gains. In a specific embodiment, if there are two groups of candidate splitting points, the first group is [10, 20] and the second group is [15, 25], wherein the population stability index of the first group is 0.15, the gain of the old samples is 0.36, the gain of the new samples is 0.3, and the overall gain is 0.33; the population stability index of the second group is 0.08, the gain of the old samples is 0.4, the gain of the new samples is 0.38, and the overall gain is 0.39; the candidate splitting point with the smallest population stability index and the largest gain is determined as the reference splitting point of the target financial risk feature, that is, the second group of candidate splitting points is determined as the reference splitting point of the target financial risk feature.

[0054] Specifically, the old samples and the new samples are binned using the candidate splitting points in each group, and then the reference splitting point of the target financial risk feature is determined based on the group stability index and gain in each bin, including: using the candidate splitting points in each group to bin the old samples and the new samples, and determining the group stability index based on the proportion of the old samples and the new samples in each bin; determining the gain of the old samples and the new samples at the candidate splitting points in each group, and determining the candidate splitting point with the smallest group stability index and the largest gain as the reference splitting point of the target financial risk feature.

[0055] Step S13: determining an original splitting point in the original financial risk control model for splitting the target financial risk feature, and replacing the original splitting point based on the reference splitting point to obtain a replaced financial risk control model.

[0056] In this embodiment, the original splitting point for the target financial risk characteristic in the original financial risk control model is determined, and the reference splitting point is compared with the original splitting point. Based on the comparison result and the principle of proximity, the reference splitting point is replaced with the original splitting point to obtain a replaced financial risk control model. In one specific embodiment, if the original splitting point is 30 and the reference splitting point is [15, 25], the 25 reference splitting points closest to the original splitting point are selected to replace the original splitting point. Specifically, the determining of the original splitting point in the original financial risk control model for the target financial risk characteristics, and replacing the original splitting point based on the reference splitting point to obtain the replaced financial risk control model includes: determining the original splitting point in the original financial risk control model for the target financial risk characteristics, and comparing the reference splitting point with the original splitting point to obtain a comparison result; based on the comparison result and the proximity principle, determining the value in the reference splitting point closest to the original splitting point as the replacement value, and replacing the value corresponding to the original splitting point based on the replacement value to obtain the replaced financial risk control model.

[0057] Step S14: determining a training set based on the new sample, and using the training set to train the replaced financial risk control model. If the trained financial risk control model meets a preset termination condition, a target financial risk control model is obtained.

[0058] In this embodiment, after obtaining the replaced financial risk control model, the new samples are divided based on a preset division ratio to obtain a training set and a validation set; the replaced financial risk control model is trained using the training set, and during the training process, the performance of the trained financial risk control model is evaluated at a preset frequency to obtain an evaluation result. In a specific embodiment, the performance of the trained financial risk control model can be evaluated by monitoring the AUC (Area Under the Curve, a performance metric) of the training set and the validation set. If the evaluation result shows that the performance of the trained financial risk control model has not improved within the preset number of training times, the training of the trained financial risk control model is stopped to prevent overfitting, thereby obtaining the target financial risk control model.

[0059] Specifically, the training set is determined based on the new sample, and the replacement financial risk control model is trained using the training set. If the trained financial risk control model meets the preset termination condition, the target financial risk control model is obtained, including: dividing the new sample based on a preset division ratio to obtain a training set and a validation set; using the training set to train the replacement financial risk control model, and during the training process, evaluating the performance of the trained financial risk control model at a preset frequency to obtain an evaluation result; if the evaluation result is that the performance of the trained financial risk control model has not improved in the preset number of training times, then stopping the training of the trained financial risk control model and obtaining the target financial risk control model. It is worth mentioning that the preset frequency and the preset number of training times can be adjusted according to actual conditions and are not specifically limited here.

[0060] As can be seen from the above, the present application collects the splitting points of all trees in the original financial risk control model for the target financial risk characteristics, arranges these splitting points, and merges the splitting points that meet the preset merging conditions, which can reduce the redundancy of the splitting points and make the splitting points more concentrated and representative; then, based on the quantile points of the target financial risk characteristics within the preset range and the merged splitting points, each group of candidate splitting points is determined, and the old samples and new samples are binned using each group of candidate splitting points. Then, based on the group stability index (PSI) and gain in each bin, the reference splitting point of the target financial risk characteristics is determined. By comparing the original splitting points for the target financial risk characteristics in the original financial risk control model with the reference splitting points, the optimized splitting points are introduced into the original financial risk control model to obtain the replaced financial risk control model. In this way, the replaced financial risk control model is trained using the training set determined by the new samples, so that the obtained target financial risk control model can better adapt to the feature distribution of the new samples, thereby improving the prediction accuracy and generalization ability of the model on new data, and enhancing the performance and practicality of the model.

[0061] Accordingly, see Figure 2 As shown, the present application also provides an incremental learning device for a financial risk control model based on XGBoost, comprising:

[0062] A split point processing module 11 is used to collect the split points of all trees in the original financial risk control model for the target financial risk characteristics, arrange the split points to obtain processed split points, and merge the processed split points that meet the preset merging conditions to obtain merged split points; the financial risk control model is a financial risk control model determined based on XGBoost;

[0063] The sample binning module 12 is configured to determine groups of candidate splitting points based on the quantile points of the target financial risk characteristic within a preset range and the merged splitting points, bin the old samples and the new samples using the candidate splitting points in each group, and then determine the reference splitting point of the target financial risk characteristic based on the population stability index and gain within each bin; the old samples are the historical financial sample data used to train the original financial risk control model; and the new samples are the new financial sample data collected after the original financial risk control model is trained.

[0064] A splitting point replacement module 13 is configured to determine an original splitting point in the original financial risk control model for splitting the target financial risk feature, and replace the original splitting point based on the reference splitting point to obtain a replaced financial risk control model;

[0065] The model training module 14 is used to determine a training set based on the new sample, and use the training set to train the replaced financial risk control model. If the trained financial risk control model meets the preset termination condition, the target financial risk control model is obtained.

[0066] As can be seen from the above, the present application collects the splitting points of all trees in the original financial risk control model for the target financial risk characteristics, arranges these splitting points, and merges the splitting points that meet the preset merging conditions, which can reduce the redundancy of the splitting points and make the splitting points more concentrated and representative; then, based on the quantile points of the target financial risk characteristics within the preset range and the merged splitting points, each group of candidate splitting points is determined, and the old samples and new samples are binned using each group of candidate splitting points. Then, based on the group stability index (PSI) and gain in each bin, the reference splitting point of the target financial risk characteristics is determined. By comparing the original splitting points for the target financial risk characteristics in the original financial risk control model with the reference splitting points, the optimized splitting points are introduced into the original financial risk control model to obtain the replaced financial risk control model. In this way, the replaced financial risk control model is trained using the training set determined by the new samples, so that the obtained target financial risk control model can better adapt to the feature distribution of the new samples, thereby improving the prediction accuracy and generalization ability of the model on new data, and enhancing the performance and practicality of the model.

[0067] In some specific implementations, the splitting point processing module 11 may specifically include:

[0068] A splitting point deduplication unit is used to traverse all trees in the original financial risk control model to collect splitting points of all trees for the target financial risk characteristics, and to deduplicate the splitting points to obtain deduplicated splitting points;

[0069] The splitting point arrangement unit is configured to arrange the deduplicated splitting points in ascending order to obtain processed splitting points.

[0070] In some specific implementations, the splitting point processing module 11 may specifically include:

[0071] a compartment determination unit, configured to determine an interval between two adjacent processed split points and determine whether the interval satisfies a preset merging condition;

[0072] A splitting point merging unit, configured to merge two adjacent processed splitting points that meet the preset merging condition if the interval meets the preset merging condition, to obtain a merged splitting point.

[0073] In some specific embodiments, the sample binning module 12 may specifically include:

[0074] A quantile value determining unit, configured to determine a quantile value based on quantile points of the target financial risk feature within a preset range in the new sample and the old sample;

[0075] A value combining unit is configured to determine the corresponding number of splitting points using the merged splitting points, then select the quantile values ​​corresponding to the number of splitting points as the values ​​to be combined, and combine the values ​​to be combined to obtain each group of candidate splitting points.

[0076] In some specific embodiments, the sample binning module 12 may specifically include:

[0077] a stability index determination unit, configured to bin the old samples and the new samples using each group of candidate split points, and determine a population stability index based on the proportion of the old samples and the new samples in each bin;

[0078] The reference splitting point determination unit is used to determine the gains of the old samples and the new samples in each group of the candidate splitting points, and determine the candidate splitting point with the smallest group stability index and the largest gain as the reference splitting point of the target financial risk feature.

[0079] In some specific implementations, the splitting point replacement module 13 may specifically include:

[0080] a splitting point comparison unit, configured to determine an original splitting point for splitting the target financial risk feature in the original financial risk control model, and compare the reference splitting point with the original splitting point to obtain a comparison result;

[0081] A replacement value determining unit is configured to determine, based on the comparison result and a proximity principle, a value in the reference splitting point that is closest to the original splitting point as a replacement value, so as to replace the value corresponding to the original splitting point based on the replacement value to obtain a replaced financial risk control model.

[0082] In some specific implementations, the model training module 14 may specifically include:

[0083] A new sample division unit, configured to divide the new sample based on a preset division ratio to obtain a training set and a validation set;

[0084] a performance evaluation unit, configured to train the replaced financial risk control model using the training set, and during the training process, evaluate the performance of the trained financial risk control model at a preset frequency to obtain an evaluation result;

[0085] The target model determination unit is used to stop training the trained financial risk control model and obtain a target financial risk control model if the evaluation result shows that the performance of the trained financial risk control model has not improved within a preset number of training times.

[0086] Furthermore, the embodiment of the present application also discloses an electronic device, Figure 3This is a structural diagram of an electronic device 20 according to an exemplary embodiment. The content in the diagram cannot be considered as any limitation on the scope of use of this application. The electronic device 20 may specifically include: at least one processor 21, at least one memory 22, a power supply 23, a communication interface 24, an input / output interface 25, and a communication bus 26. The memory 22 is used to store a computer program, which is loaded and executed by the processor 21 to implement the relevant steps in the incremental learning method of the financial risk control model based on XGBoost disclosed in any of the aforementioned embodiments. In addition, the electronic device 20 in this embodiment may specifically be an electronic computer.

[0087] In this embodiment, the power supply 23 is used to provide operating voltage for each hardware device on the electronic device 20; the communication interface 24 can create a data transmission channel between the electronic device 20 and the external device. The communication protocol it follows is any communication protocol that can be applied to the technical solution of this application and is not specifically limited here; the input and output interface 25 is used to obtain external input data or output data to the outside world. Its specific interface type can be selected according to specific application needs and is not specifically limited here.

[0088] In addition, the memory 22, as a carrier for resource storage, can be a read-only memory, random access memory, disk or CD, etc. The resources stored thereon can include an operating system 221, a computer program 222, etc., and the storage method can be temporary storage or permanent storage.

[0089] The operating system 221 is used to manage and control the hardware devices and computer program 222 on the electronic device 20, and can be Windows Server, Netware, Unix, Linux, etc. In addition to including a computer program capable of implementing the XGBoost-based incremental learning method for financial risk control models executed by the electronic device 20 as disclosed in any of the aforementioned embodiments, the computer program 222 may further include computer programs capable of completing other specific tasks.

[0090] Furthermore, this application also discloses a computer-readable storage medium for storing a computer program; wherein, when executed by a processor, the computer program implements the aforementioned XGBoost-based incremental learning method for financial risk control models. The specific steps of this method can be found in the corresponding content disclosed in the aforementioned embodiments and will not be repeated here.

[0091] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on its differences from the other embodiments. Reference can be made to the descriptions of the identical or similar parts between the various embodiments. For the devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the descriptions are relatively simple, and the relevant parts can be referred to the descriptions of the methods.

[0092] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the above description has generally described the components and steps of each example according to their functions. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0093] The steps of the methods or algorithms described in conjunction with the embodiments disclosed herein may be implemented directly using hardware, a software module executed by a processor, or a combination of the two. The software module may be placed in random access memory (RAM), internal memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, a hard disk, a removable disk, a CD-ROM, or any other form of storage medium known in the art.

[0094] Finally, it should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of additional identical elements in the process, method, article, or device comprising the element.

[0095] The above is a detailed introduction to the technical solution provided by the present application. Specific examples are used herein to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only used to help understand the method of the present application and its core idea. At the same time, for those skilled in the art, according to the ideas of the present application, there may be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as a limitation on the present application.

Claims

1. An incremental learning method for financial risk control models based on XGBoost, characterized in that: include: Collecting all splitting points of the target financial risk characteristics from all trees in the original financial risk control model, arranging the splitting points to obtain processed splitting points, and merging the processed splitting points that meet preset merging conditions to obtain merged splitting points; the financial risk control model is a financial risk control model determined based on XGBoost; Determine each group of candidate splitting points based on the quantile points of the target financial risk feature within a preset range and the merged splitting points, use each group of candidate splitting points to bin the old samples and the new samples, and then determine the reference splitting point of the target financial risk feature based on the population stability index and gain within each bin; the old samples are the historical financial sample data used to train the original financial risk control model; the new samples are the new financial sample data collected after the original financial risk control model is trained; Determining an original splitting point in the original financial risk control model for splitting the target financial risk feature, and replacing the original splitting point based on the reference splitting point to obtain a replaced financial risk control model; A training set is determined based on the new sample, and the replaced financial risk control model is trained using the training set. If the trained financial risk control model meets a preset termination condition, a target financial risk control model is obtained.

2. The XGBoost-based incremental learning method for financial risk control models according to claim 1 is characterized in that: The collecting of the splitting points of all trees in the original financial risk control model for the target financial risk characteristics and the arranging of the splitting points to obtain the processed splitting points includes: Traversing all trees in the original financial risk control model to collect splitting points of all trees for the target financial risk characteristics, and removing duplicates from the splitting points to obtain deduplicated splitting points; The deduplicated splitting points are sorted in ascending order to obtain processed splitting points.

3. The XGBoost-based incremental learning method for financial risk control models according to claim 1 is characterized in that: Merging the processed split points that meet the preset merging condition to obtain merged split points includes: Determining an interval between two adjacent processed split points, and judging whether the interval meets a preset merging condition; If the interval satisfies the preset merging condition, merging two adjacent processed split points that satisfy the preset merging condition to obtain a merged split point; Among them, the preset merging condition is that the interval is smaller than the target interval threshold; the target interval threshold is an interval threshold determined based on the difference between the first quantile and the second quantile of the target financial risk feature in all samples; the first quantile is smaller than the second quantile.

4. The XGBoost-based incremental learning method for financial risk control models according to claim 1, characterized in that: The determining of each group of candidate splitting points based on the quantile points of the target financial risk characteristics within a preset range and the merged splitting points includes: Determining a quantile value based on quantile points within a preset range in the new sample and the old sample of the target financial risk characteristic; The number of corresponding splitting points is determined using the merged splitting points, and then the quantile values ​​corresponding to the number of splitting points are selected as the values ​​to be combined, and the values ​​to be combined are combined to obtain each group of candidate splitting points.

5. The XGBoost-based incremental learning method for financial risk control models according to claim 1, characterized in that: The method of binning the old samples and the new samples using each group of candidate splitting points, and then determining the reference splitting point of the target financial risk feature based on the group stability index and gain in each bin, includes: Binning the old samples and the new samples using the candidate split points in each group, and determining a population stability index based on the proportion of the old samples and the new samples in each bin; The gains of the old samples and the new samples at each group of the candidate splitting points are determined, and the candidate splitting point with the smallest group stability index and the largest gain is determined as the reference splitting point of the target financial risk feature.

6. The XGBoost-based incremental learning method for financial risk control models according to claim 1, characterized in that: The determining of an original splitting point for splitting the target financial risk feature in the original financial risk control model, and replacing the original splitting point based on the reference splitting point to obtain a replaced financial risk control model, includes: Determining an original splitting point in the original financial risk control model for splitting the target financial risk characteristics, and comparing the reference splitting point with the original splitting point to obtain a comparison result; Based on the comparison result and the proximity principle, the value of the reference splitting point closest to the original splitting point is determined as a replacement value, and the value corresponding to the original splitting point is replaced based on the replacement value to obtain a replaced financial risk control model.

7. The XGBoost-based incremental learning method for financial risk control models according to any one of claims 1 to 6, characterized in that: The step of determining a training set based on the new sample, training the replaced financial risk control model using the training set, and obtaining a target financial risk control model if the trained financial risk control model satisfies a preset termination condition, includes: Dividing the new sample based on a preset division ratio to obtain a training set and a validation set; The replaced financial risk control model is trained using the training set, and during the training process, the performance of the trained financial risk control model is evaluated at a preset frequency to obtain an evaluation result; If the evaluation result shows that the performance of the trained financial risk control model has not improved within the preset number of training times, the training of the trained financial risk control model is stopped, and the target financial risk control model is obtained.

8. An incremental learning device for financial risk control models based on XGBoost, characterized in that: include: A split point processing module is used to collect the split points of all trees in the original financial risk control model for the target financial risk characteristics, arrange the split points to obtain processed split points, and merge the processed split points that meet the preset merging conditions to obtain merged split points; the financial risk control model is a financial risk control model determined based on XGBoost; A sample binning module is configured to determine groups of candidate splitting points based on the quantile points of the target financial risk characteristic within a preset range and the merged splitting points, bin the old and new samples using the candidate splitting points in each group, and then determine a reference splitting point for the target financial risk characteristic based on the population stability index and gain within each bin; the old samples are the historical financial sample data used to train the original financial risk control model; and the new samples are the new financial sample data collected after the original financial risk control model is trained. a splitting point replacement module, configured to determine an original splitting point in the original financial risk control model for splitting according to the target financial risk feature, and replace the original splitting point based on the reference splitting point to obtain a replaced financial risk control model; The model training module is used to determine a training set based on the new sample, and use the training set to train the replaced financial risk control model. If the trained financial risk control model meets the preset termination condition, the target financial risk control model is obtained.

9. An electronic device, characterized in that: include: Memory, used to store computer programs; A processor, configured to execute the computer program to implement the XGBoost-based incremental learning method for a financial risk control model as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that Used to store a computer program, wherein when the computer program is executed by a processor, it implements the incremental learning method of the financial risk control model based on XGBoost as described in any one of claims 1 to 7.