Financial risk control model training optimization system based on machine learning
Through sample clustering and online learning modules, the problem of high computational complexity and poor adaptability in the financial risk control model is solved, and efficient and dynamic training optimization of financial risk control model is achieved.
Patent Information
- Application Number
- CN202510465505.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-15
- Publication Date
- 2025-07-08
AI Technical Summary
Traditional principal component analysis algorithms (PCA) have high computational complexity in financial risk control models and are sensitive to outliers, which are difficult to adapt to dynamic data changes, and cannot meet the needs of high-frequency trading and real-time risk control.
The sample clustering module is used for clustering, high-density and low-density sample sets are obtained, samples are sorted using local estimated density and Mahayana distance, representative samples are selected, and real-time updates and training are performed through the online learning module.
The online learning efficiency of financial risk control models has been improved, ensuring that the model maintains the optimal feature expression ability in a dynamic environment, and achieving efficient and dynamic financial risk control model training optimization.
Smart Images

Figure CN120277415A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of financial risk control monitoring, and particularly relates to a financial risk control model training and optimization system based on machine learning. Background Art
[0002] With the rapid development of fintech, financial institutions increasingly rely on data-driven risk control models in aspects such as credit approval, fraud detection, and market risk management. The data in the financial market has characteristics such as high dimensionality, time series, non-linearity, and dynamic changes. How to efficiently process and optimize massive financial data has become a key challenge in the construction of financial risk control models. In order to improve prediction accuracy and reduce redundant information, feature dimensionality reduction techniques are widely applied to the online learning of financial risk control models.
[0003] Traditional dimensionality reduction methods such as the principal component analysis algorithm (PCA) can reduce the computational complexity while maintaining the main information of the data. Its applications in financial risk control mainly include dimensionality reduction and compression, removing feature redundancy, and improving computational efficiency. PCA calculates the covariance matrix of the data, performs eigenvalue decomposition, selects the principal components to construct a new low-dimensional feature space, thereby reducing the data dimension while retaining as much information as possible. This method has good performance on static data sets. However, PCA has problems such as high computational complexity, sensitivity to outliers, and difficulty in adapting to dynamic changes in data. When the data scale increases, PCA needs to perform eigenvalue decomposition on the entire data set, with a large amount of calculation, which cannot meet the requirements of high-frequency trading and real-time risk control, and is difficult to meet the requirements of the online learning of financial risk control models. Summary of the Invention
[0004] In order to solve the above technical problems, the purpose of the present invention is to provide a financial risk control model training and optimization system based on machine learning, and the specific technical solutions adopted are as follows: An embodiment of the present invention provides a financial risk control model training and optimization system based on machine learning, and the system includes: A sample clustering module, configured to obtain the financial feature data of each person as the samples corresponding to each person to form the current sample set; and cluster the samples to obtain different clustering clusters; A clustering cluster sample division module, configured to obtain the local estimated density of each sample according to the Euclidean distance between the neighbor points and each sample in the clustering cluster and the standard deviation of the distance from each sample to the clustering center; and divide the samples in the clustering cluster according to the local estimated density to obtain a high-density sample set and a low-density sample set; A high-density sample set processing module, configured to sort the samples in the high-density sample set by using the Mahalanobis distance between the samples in the high-density sample set and the clustering center and the Mahalanobis distance between each sample; and obtain representative samples according to the sorted high-density sample set; The low-density sample set processing module is used to calculate the probability density of each sample in the low-density sample set according to the Euclidean distance from each sample to the cluster center; sort the samples in the low-density sample set according to the probability density, and obtain representative samples; The online learning module is used to reduce the dimension of the representative samples in the current sample set, train the financial risk control model using the dimension-reduced representative samples, and update the current sample set in real time, and perform online learning using the updated sample set.
[0005] Preferably, obtaining the financial feature data of each person as the samples corresponding to each person to form the current sample set, including: The financial feature data of each person includes the transaction amount, transaction frequency, credit score and income within one month of each person within a preset period; the preprocessed financial feature data of each person is used as the sample corresponding to each person to form the current sample set.
[0006] Preferably, obtaining the local estimated density of each sample, including: Take a preset number of samples closest to a sample within the cluster as the neighbor points of the sample; obtain the square of the Euclidean distance between a sample and its neighbor points within the cluster, and compare it with the square of the preset multiple of the standard deviation of the distances from each sample to the cluster center to obtain a first ratio; use the exponential function with the natural constant as the base to perform a negative correlation mapping on the first ratios corresponding to each neighbor point of the sample and sum them to obtain the local estimated density of the sample.
[0007] Preferably, dividing the samples in the cluster according to the local estimated density to obtain a high-density sample set and a low-density sample set, including: Obtain the local estimated density of each sample in the cluster, denoted as the average local estimated density, obtain the samples whose Euclidean distance from the cluster center is greater than the average local estimated density, and screen out the sample with the smallest Euclidean distance from the cluster center among them, denoted as the reference sample, and the Euclidean distance between the reference sample and the cluster center is the reference distance; the samples in the cluster whose Euclidean distance from the cluster center is less than or equal to the reference distance form the high-density sample set, and the samples greater than the reference distance form the low-density sample set.
[0008] Preferably, it is used to sort the samples in the high-density sample set by using the Mahalanobis distance between the samples in the high-density sample set and the cluster center and the Mahalanobis distance between each sample, including: Taking the cluster center as the base point, calculate the Mahalanobis distance between each sample in the cluster and the base point, and select the point with the largest Mahalanobis distance and the base point to form the base set; calculate the sum of the Mahalanobis distances between a sample in the high-density sample set except the base set and two points in the base set, which is denoted as the Mahalanobis distance sum corresponding to this sample, and add the sample with the largest Mahalanobis distance sum to the base set to form a new base set, and so on, until all samples in the high-density sample set are added to the base set; sort the samples in the high-density sample set according to the order in which the samples in the high-density sample set are added to the base set to obtain the sorted high-density sample set.
[0009] Preferably, obtaining representative samples according to the sorted high-density sample set includes: Dividing the sorted high-density sample set proportionally into a preset number of parts, and taking a set number of samples in each part in the order from front to back as representative samples.
[0010] Preferably, the calculation formula for the probability density of each sample is: , where, represents the probability density of the i-th sample in the low-density sample set, e represents the natural constant; represents the Euclidean distance from the i-th sample in the low-density sample set to the cluster center; α represents the hyperparameter, and β represents the adjustment parameter.
[0011] Preferably, sorting the samples in the low-density sample set according to the probability density and obtaining representative samples includes: Sorting the samples in the low-density sample set in descending order based on the probability density of each sample in the low-density sample set to obtain the sorted low-density sample set, dividing the sorted low-density sample set proportionally into a preset number of parts, and taking a set number of samples in each part in the order from front to back as representative samples.
[0012] Preferably, the current sample set is updated in real time, and online learning is performed using the updated sample set, including: Setting a sliding time window with a preset time length, sliding with a preset sliding step size, sliding once, updating the current sample set once to obtain the updated sample set, screening the updated sample set to obtain representative samples, and training the financial risk control model using the representative samples corresponding to the updated sample set.
[0013] The embodiments of the present invention have at least the following beneficial effects: By clustering the financial feature data (samples) of each person, all samples are divided into different types of financial transaction patterns. Then, the local estimated density of each sample is obtained according to the standard deviation of the Euclidean distance between the nearest neighbor points and each sample in the clustering cluster and the distance from each sample to the clustering center. The samples in each clustering cluster are divided into a high-density sample set and a low-density sample set. Furthermore, the samples in the high-density sample set and the low-density sample set are sorted respectively to obtain representative samples, and the representative samples corresponding to each clustering cluster are obtained. Representative samples are selected from the center and the edge part of the clustering cluster, taking into account each type of financial transaction pattern, so that the structural integrity and completeness of the selected representative samples are higher and the information contained is more complete. Further, the representative samples in the current sample set are dimensionally reduced, and the financial risk control model is trained using the dimensionally reduced representative samples, and the current sample set is updated in real time. The financial risk control model is online learned using the representative samples in the updated current sample set, ensuring that the financial risk control model always maintains the optimal feature expression ability while improving the online learning efficiency of the financial risk control model, and realizing an efficient and dynamic financial risk control model training optimization scheme. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] In order to more clearly illustrate the technical solutions and advantages in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0015] Figure 1 FIG. is a system block diagram of a financial risk control model training optimization system based on machine learning provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0016] In order to further elaborate on the technical means and effects adopted by the present invention to achieve the predetermined invention purpose, the following will, in conjunction with the drawings and preferred embodiments, detail the specific implementation manners, structures, features and effects of a financial risk control model training optimization system based on machine learning proposed according to the present invention. In the following description, different "one embodiment" or "another embodiment" do not necessarily refer to the same embodiment. In addition, the specific features, structures, or characteristics in one or more embodiments can be combined in any suitable form.
[0017] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present invention belongs.
[0018] The following specifically describes the specific solution of a financial risk control model training and optimization system based on machine learning provided by the present invention in conjunction with the accompanying drawings.
[0019] Embodiment: The main application scenario of the present invention is as follows: This system uses PCA for dimensionality reduction under representative data, combines the online learning method, realizes an efficient and dynamic financial risk control model training and optimization solution, and selects representative data to participate in the online learning and training of the financial risk control model.
[0020] Please refer to Figure 1 , which shows a system block diagram of a financial risk control model training and optimization system based on machine learning provided by an embodiment of the present invention. The system includes the following modules: A sample clustering module, configured to obtain the financial feature data of each person as the samples corresponding to each person to form the current sample set; and cluster the samples to obtain different clusters.
[0021] In order to eliminate possible problems in the original data of each person, the present invention first needs to perform preprocessing operations on the collected user data. The financial risk control model in this application can be a loan risk monitoring model, a repayment rate stratification prediction model, an anti-fraud model, etc. When different models perform online learning, the types of data required may be different, and it can be determined according to the actual situation.
[0022] Real-time obtain the transaction flow data of each person through the financial trading system and the third-party payment platform, including basic information such as transaction amount and transaction frequency. Cooperate with the credit scoring platform (such as the bank credit rating system or the third-party credit data interface) to collect data such as the user's credit score, repayment record, and loan history of the user. At the same time, use publicly available Internet data sources (such as social media platforms, industry reports, financial data platforms) to collect auxiliary features related to the user's financial status, such as the user's occupation, income level, and social activity frequency.
[0023] In an embodiment of the present invention, the transaction amount, transaction frequency of each person within a preset period, and the credit score and income within one month are obtained as the financial feature data of each person, that is, the financial feature data of each person includes the transaction amount, transaction frequency of each person within a preset period, and the credit score and income within one month. The preset period is one week. It should be noted that the implementer can select different types of data according to the situation. Some data need to be quantified. For example, for the occupation, it can be quantified using coding, which is a well-known technology and will not be elaborated here. And the setting of the preset period can also be adjusted by the implementer according to the actual situation.
[0024] Furthermore, preprocess the financial feature data of each person, including using outlier detection (such as Z-score) to identify and remove extreme trading outliers to prevent extreme values from having too much impact on the model, performing missing value filling (mean filling) to ensure data integrity and reduce information loss caused by missing data, and solving the problem of inconsistent feature dimensions through Min-Max normalization to ensure the balanced influence of different features and improve the model convergence speed. Use the preprocessed financial feature data of each person as the corresponding sample for each person, and all the samples form the current sample set.
[0025] During the online learning process of the financial risk control model, in order to reduce the computational complexity while ensuring the representativeness of the sample data, K-Means clustering and PCA can be combined for dynamic feature dimensionality reduction and optimization. First, perform K-Means clustering on the large-scale original samples, divide the data into several categories, and each category represents different types of user behavior patterns or credit risk levels.
[0026] Let the current sample set be , where each sample contains multiple features, including the transaction amount, transaction frequency of each person within a preset period, and the credit score and income within one month. Use K-means for clustering, and find the optimal cluster partition by minimizing the within-cluster sum of squares to obtain the clustering clusters.
[0027] The cluster sample partitioning module is used to obtain the local estimated density of each sample according to the Euclidean distance between the nearest neighbor points and each sample in the cluster and the standard deviation of the distance from each sample to the cluster center; divide the samples in the cluster according to the local estimated density to obtain a high-density sample set and a low-density sample set.
[0028] After clustering each sample in the current sample set, it is necessary to screen out the most representative sample points in each clustering cluster (category) to ensure that the dimension-reduced samples can still cover the main distribution characteristics of the overall data. In the feature dimensionality reduction stage, perform dimensionality reduction on the selected samples so that the data reduces the dimension while maintaining the main feature information, improving the computational efficiency. This method can avoid the memory overhead of standard PCA when processing large-scale data and support online learning, enabling the model to dynamically adjust the clustering and dimensionality reduction results when the samples are continuously updated. Use the dimension-reduced features to screen out representative samples and then perform real-time model training optimization, which can improve the efficiency of model training and save resource consumption.
[0029] After performing K-means clustering on a large-scale current dataset, the samples are divided into different categories, and each category represents a transaction behavior pattern or credit risk level. After clustering, the samples of each category will be distributed around the center point of the cluster. Therefore, the representativeness of the samples can be measured based on the distance from the samples to the cluster center (clustering center).
[0030] After clustering, the samples within each clustering cluster are distributed around the clustering center So, it is necessary to calculate the Euclidean distance from each sample in the clustering cluster to the clustering center. The calculation formula of the Euclidean distance is as follows: , where represents the Euclidean distance from the j-th sample within the clustering cluster to the clustering center of this clustering cluster, and respectively represent the k-th eigenvalue of the j-th sample and the clustering center of the i-th clustering cluster, that is, the value of the k-th dimension of the sample.
[0031] Furthermore, in order to ensure that the selected representative samples can not only reflect the clustering characteristics but also maintain data diversity, the standard deviation of the distances from each sample to the clustering center is obtained and denoted as .
[0032] Furthermore, calculate the local estimated density of each sample within the clustering cluster, and obtain the local estimated density of each sample based on the Euclidean distance between the nearest neighbor points of each sample in the clustering cluster and each sample and the standard deviation of the distances from each sample to the clustering center.
[0033] Specifically, within the clustering cluster, a preset number of samples closest to a sample are taken as the nearest neighbor points of this sample. The distance used when taking the nearest neighbor points is the Euclidean distance. The reference value of the preset number can be 10, and the implementer can adjust it according to the actual situation.
[0034] Furthermore, obtain the square of the Euclidean distance between a sample within the clustering cluster and its nearest neighbor points, and compare it with the square of the preset multiple of the standard deviation of the distances from each sample to the clustering center to obtain the first ratio; use the exponential function with the natural constant as the base to perform a negative correlation mapping on the first ratios corresponding to each nearest neighbor point of this sample and sum them to obtain the local estimated density of this sample; the specific calculation formula is as follows: , where represents the local estimated density of the j-th sample within the clustering cluster, represents the set of the nearest neighbor points of the j-th sample. Within the clustering cluster, a preset number of samples closest to a sample are taken as the nearest neighbor points of this sample. The reference value of the preset number is 10, and the implementer can adjust it according to the actual situation, denotes the p-th nearest neighbor of the j-th sample, and e denotes the natural constant. denotes the first ratio. is a preset multiple of the square of the standard deviation of the distances from each sample to the cluster center. The first ratio represents the contribution of each nearest neighbor among the nearest neighbors, with the contribution being greater for closer distances and smaller for farther distances.
[0035] When dealing with different transaction behavior patterns (such as small-amount high-frequency transactions and large-amount low-frequency transactions), the point distributions of each cluster are quite different. The standard deviation replacement can be adaptively adjusted according to the characteristics of each cluster, avoiding the errors caused by a unified bandwidth. In an online learning scenario with real-time updates, as the data stream changes, the standard deviation can promptly respond to the characteristics of the latest transaction behavior, making the density calculation continuously effective in a dynamic environment. Thus, the distribution of each sample within the cluster can be obtained.
[0036] Furthermore, obtain the local estimated density of each sample within the cluster, denoted as the average local estimated density. Obtain the samples whose Euclidean distance from the cluster center is greater than the average local estimated density, and screen out the sample with the smallest Euclidean distance from the cluster center among them, denoted as the reference sample. The Euclidean distance between the reference sample and the cluster center is the reference distance; the samples within the cluster whose Euclidean distance from the cluster center is less than or equal to the reference distance form a high-density sample set, and the samples greater than the reference distance form a low-density sample set.
[0037] During the K-Means clustering process, data points are assigned to the nearest cluster center, but K-Means itself does not explicitly distinguish between high-density and low-density regions, but divides the data according to the principle of minimum squared error. Therefore, directly selecting representative data based on K-Means may result in too much data near some cluster centers and insufficient data far from the centers. Thus, the samples within each cluster can be divided into a high-density sample set and a low-density sample set.
[0038] A high-density sample set processing module is used to sort the samples in the high-density sample set by using the Mahalanobis distance between the samples in the high-density sample set and the cluster center and the Mahalanobis distance between the samples; obtain representative samples according to the sorted high-density sample set.
[0039] After obtaining the high-density sample set and the low-density sample set of the cluster, screen out representative samples from the high-density sample set and then use PCA for dimensionality reduction. First, in the high-density sample set, due to the aggregation of samples, many sample information is similar. Select the sample with the largest information difference (the most representative) in each region, rather than simply selecting the sample closest to the cluster center. This can ensure that the data retains the core features and does not over-concentrate on a certain point.
[0040] It is used to sort the samples in the high-density sample set by using the Mahalanobis distance between the samples in the high-density sample set and the cluster center, as well as the Mahalanobis distance between each sample. Specifically, the cluster center is used as the base point, the Mahalanobis distance between each sample in the cluster and the base point is calculated, and the point with the largest Mahalanobis distance and the base point are used to form the base set; calculate the sum of the Mahalanobis distances between a sample in the high-density sample set except the base set and two points in the base set, which is denoted as the Mahalanobis distance sum corresponding to this sample, and add the sample with the largest Mahalanobis distance sum to the base set to form a new base set, and so on, until all the samples in the high-density sample set are added to the base set; sort the samples in the high-density sample set according to the order in which the samples in the high-density sample set are added to the base set to obtain the sorted high-density sample set.
[0041] The Mahalanobis distance can quantify the similarity or difference between data points in a multi-dimensional space. By calculating the Mahalanobis distance between samples, those points that are quite different from the existing samples can be selected. These points may be located at the boundary of the data distribution or in less-covered areas, which helps to ensure that the selected data can cover the entire data distribution, rather than just focusing on the central point or high-density area. The covariance matrix used to calculate the Mahalanobis distance is the inverse matrix of the covariance matrix of the high-density sample set.
[0042] Furthermore, the sorted high-density sample set is divided proportionally into a preset number of parts. In each part of the samples, a set number of samples are taken in the order from front to back as representative samples. Among them, the proportional division in this application is based on the proportion of the samples in the sorted high-density sample set, and is divided into within 20%, 20% to 40%, 40% to 60%, 60% to 80%, 80% to 100%. The preset number of parts is 5, and the samples are sorted when dividing. The set number of samples is 2% of the number of samples in each part, that is, the first 2% of the sample points in each part, and a total of 10% of the samples in the high-density sample set are taken.
[0043] This is done to find a small part of samples that can represent the overall data in a large number of samples, reduce the complexity of subsequent PCA dimensionality reduction processing. The high-density sample set is uniform, and taking samples proportionally is simple and efficient, and it has strong representativeness in a large number of samples.
[0044] The low-density sample set processing module is used to calculate the probability density of each sample in the low-density sample set according to the Euclidean distance from each sample in the low-density sample set to the cluster center; sort the samples in the low-density sample set according to the probability density and obtain representative samples.
[0045] If only relying on the data in the high-density sample set, the model may overfit the characteristics of the aggregation area, resulting in poor prediction performance in the sparse area. Equal-proportion screening can balance the characteristics of different density areas and enable the model to have better generalization ability on diverse data sets.
[0046] In the low-density sample set, due to the dispersion of samples and being greatly affected by noise points, a dynamic screening ratio can be set according to the distance from the cluster center, so that samples far from the center still have a certain probability of being selected, while ensuring their uniform distribution and not being affected by extreme samples in terms of overall representativeness.
[0047] When sampling the low-density area, since the samples in the low-density area are sparse and there are more noise points, in order to ensure that subsequent dimensionality reduction is not interfered by excessive noise points, this solution uses the method of density attenuation to calculate the selection probability (probability density) of each low-density data point, and selects the corresponding samples according to the probability density.
[0048] Calculate the probability density of each sample according to the Euclidean distance from each sample in the low-density sample set to the clustering center. The specific calculation formula for the probability density of each sample is: , Among them, represents the probability density of the i-th sample in the low-density sample set, e represents the natural constant; represents the Euclidean distance from the i-th sample in the low-density sample set to the clustering center; α represents the hyperparameter, and β represents the adjustment parameter.
[0049] α represents the hyperparameter that controls the sampling tendency. To amplify the sampling probability of farther samples, this value can be taken as 0.5. β represents the adjustment parameter used to adjust the probability of points, and the reference value can be taken as 2. represents allowing points at a long distance to retain a certain probability to ensure that boundary points are not completely lost in extreme cases. represents the non-linear mapping function, which can enhance the boundary screening ability of the low-density area. represents the selection probability of the i-th sample.
[0050] Furthermore, based on the probability density of each sample in the low-density sample set, sort the samples in the low-density sample set in descending order to obtain the sorted low-density sample set. Divide the sorted low-density sample set into equal proportions and divide it into a preset number of parts. Take a set number of samples in each part in the order from front to back as representative samples. The method of screening representative samples in the sorted low-density sample set is the same as that in the sorted high-density sample set.
[0051] Since the samples in the low-density sample set are often accompanied by more noise and outliers, directly sampling a large number of samples from the low-density region may cause the model to be affected by noise. Through proportional screening, the interference of noise on model training can be effectively reduced.
[0052] Thus, for each clustering cluster of the current sample set, it can be divided into a high-density sample set and a low-density sample set, and then representative samples can be selected from them, that is, representative samples, to train the financial risk control model.
[0053] The online learning module is used to reduce the dimension of the representative samples in the current sample set, train the financial risk control model using the representative samples after dimension reduction, and update the current sample set in real time, and perform online learning using the updated sample set.
[0054] After obtaining the representative samples in the current sample set, the representative samples need to be dimensionally reduced. The PCA method is used for dimensional reduction to obtain the covariance matrix of the representative samples, perform eigenvalue decomposition on it, select the top m most important principal components, and construct a dimensionality reduction transformation matrix. The value of m is determined by the cumulative variance contribution rate. In general scenarios, the threshold of the cumulative variance contribution rate can be set to 0.95 for high-precision requirements and 0.85 for limited computing resource requirements to ensure that the main information of the data is retained. Then, the financial risk control model is trained and learned using the representative samples after dimension reduction.
[0055] The representative samples under the new features after dimension reduction are used for the training of the financial risk control model. The data under the reduced features is representative, which can save a large amount of computing resource consumption and greatly improve the efficiency of model training.
[0056] Furthermore, after the new financial feature data flows into the system, the current sample set needs to be updated. The present invention updates the current sample set in the form of a sliding window. A sliding time window with a preset time length is set, and it slides with a preset sliding step. Each time it slides, the current sample set is updated once to obtain an updated sample set. The updated sample set is screened to obtain representative samples, and then the financial risk control model is trained using the representative samples corresponding to the updated sample set to achieve the purpose of online learning of the financial risk control model. Each time the window slides, the sample set is updated once.
[0057] The preset time length of the specific sliding time window is 1 day, and the sliding step is 12 hours. After each slide, 12 hours of historical data are retained, and 12 hours of new samples are added. The new samples may be the financial feature data of newly added users, or the new financial feature data generated by the change of the data of previous users within 12 hours. The implementer of the preset time length and the sliding step can adjust according to the actual situation. Taking one day as a sliding window is because financial risk control data generally has strong periodicity and timeliness within one day, and can comprehensively cover the fluctuation characteristics of transaction behaviors. Taking 12 hours as the sliding step can balance real-time performance and stability. It can not only quickly respond to the changes in the latest trading patterns, but also avoid the noise accumulation and overfitting problems caused by frequent updates, thus ensuring that the model has high robustness and accuracy during real-time updates. It ensures that new data gradually affects the training and historical data does not become invalid, guaranteeing the real-time update of the financial risk control model and its accuracy on this basis.
[0058] In summary, in the process of real-time training and update of this application, screening representative samples needs to dynamically adapt to the changes in transaction behaviors. First, the sliding window mechanism is adopted to maintain the timeliness of financial feature data, and through online K-Means clustering, the cluster centers are adjusted with new data to ensure that the clustering results always fit the latest patterns. During screening, the local density is updated in real time, and the division of high- and low-density sample sets is dynamically adjusted. Representative samples of each clustering cluster are screened, and then PCA is used for dimensionality reduction. The financial risk control model is trained using the features after dimensionality reduction to achieve real-time update. Finally, combined with the screening strategy, it is ensured that the selected data always reflects the latest trading patterns, making the financial risk control model efficient and accurate during real-time training.
[0059] It should be noted that the above sequence of the embodiments of the present invention is only for description and does not represent the advantages or disadvantages of the embodiments. And the above specific embodiments of this specification have been described. In addition, the processes depicted in the drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multi-task processing and parallel processing are also possible or may be advantageous.
[0060] Each embodiment in this specification is described in a progressive manner. The same or similar parts among the embodiments can be referred to each other, and the key points of each embodiment are the differences from other embodiments.
[0061] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A financial risk control model training and optimization system based on machine learning, characterized in that, The system includes: A sample clustering module, configured to obtain the financial feature data of each person as samples corresponding to each person to form the current sample set; cluster the samples to obtain different clustering clusters; A clustering cluster sample division module, configured to obtain the local estimated density of each sample according to the Euclidean distance between the nearest neighbor points and each sample in the clustering cluster and the standard deviation of the distances from each sample to the cluster center; divide the samples in the clustering cluster according to the local estimated density to obtain a high-density sample set and a low-density sample set; A high-density sample set processing module, configured to sort the samples in the high-density sample set by using the Mahalanobis distance between the samples in the high-density sample set and the cluster center and the Mahalanobis distance between the samples; obtain representative samples according to the sorted high-density sample set; A low-density sample set processing module, configured to calculate the probability density of each sample according to the Euclidean distance from each sample in the low-density sample set to the cluster center; sort the samples in the low-density sample set according to the probability density, and obtain representative samples; An online learning module, configured to reduce the dimension of the representative samples in the current sample set, train a financial risk control model by using the reduced-dimensional representative samples, and update the current sample set in real time, and perform online learning by using the updated sample set.
2. The financial risk control model training and optimization system based on machine learning according to claim 1, characterized in that, The obtaining the financial feature data of each person as samples corresponding to each person to form the current sample set includes: The financial feature data of each person includes the transaction amount, transaction frequency of each person within a preset period, and the credit score and income within one month; the preprocessed financial feature data of each person is used as the sample corresponding to each person to form the current sample set.
3. A financial risk control model training and optimization system based on machine learning according to claim 1, characterized in that, The obtaining the local estimated density of each sample includes: In the clustering cluster, take a preset number of samples closest to a sample as the nearest neighbor points of the sample; obtain the square of the Euclidean distance between a sample in the clustering cluster and its nearest neighbor points, and compare it with the square of the preset multiple of the standard deviation of the distances from each sample to the cluster center to obtain a first ratio; use the exponential function with the natural constant as the base to perform a negative correlation mapping on the first ratios corresponding to the nearest neighbor points of the sample and sum them to obtain the local estimated density of the sample.
4. A machine learning-based financial risk control model training and optimization system according to claim 1, characterized in that The dividing the samples in the clustering cluster according to the local estimated density to obtain a high-density sample set and a low-density sample set includes: Obtain the local estimated density of each sample in the clustering cluster, denoted as the average local estimated density, obtain the samples whose Euclidean distance from the cluster center is greater than the average local estimated density, and screen out the sample with the smallest Euclidean distance from the cluster center among them, denoted as the reference sample, and the Euclidean distance between the reference sample and the cluster center is the reference distance; the samples in the clustering cluster whose Euclidean distance from the cluster center is less than or equal to the reference distance form a high-density sample set, and the samples greater than the reference distance form a low-density sample set.
5. A training and optimization system for a financial risk control model based on machine learning according to claim 1, characterized in that, The using the Mahalanobis distance between the samples in the high-density sample set and the cluster center and the Mahalanobis distance between the samples to sort the samples in the high-density sample set includes: Taking the cluster center as the base point, calculate the Mahalanobis distance between each sample in the cluster and the base point, and select the point with the largest Mahalanobis distance and the base point to form the base set; calculate the sum of the Mahalanobis distances between a sample in the high-density sample set except the base set and two points in the base set, denoted as the Mahalanobis distance sum corresponding to the sample, and add the sample with the largest Mahalanobis distance sum to the base set to form a new base set, and so on, until all samples in the high-density sample set are added to the base set; sort the samples in the high-density sample set according to the order in which the samples in the high-density sample set are added to the base set to obtain the sorted high-density sample set.
6. A training and optimization system for a financial risk control model based on machine learning according to claim 1, wherein The obtaining of representative samples according to the sorted high-density sample set includes: Dividing the sorted high-density sample set proportionally into a preset number of parts, and taking a set number of samples in each part in the order from front to back as representative samples.
7. A machine learning-based financial risk control model training and optimization system according to claim 1, characterized in that, The calculation formula for the probability density of each sample is: , Among them, represents the probability density of the i-th sample in the low-density sample set, and e represents the natural constant; represents the Euclidean distance from the i-th sample in the low-density sample set to the cluster center; α represents a hyperparameter, and β represents an adjustment parameter.
8. An optimization system for training a financial risk control model based on machine learning according to claim 1, characterized in that, The sorting of the samples in the low-density sample set according to the probability density and the obtaining of representative samples include: Sorting the samples in the low-density sample set in descending order based on the probability density of each sample in the low-density sample set to obtain the sorted low-density sample set, dividing the sorted low-density sample set proportionally into a preset number of parts, and taking a set number of samples in each part in the order from front to back as representative samples.
9. A training and optimization system for a financial risk control model based on machine learning according to claim 1, characterized in that The real-time update of the current sample set and the use of the updated sample set for online learning include: Setting a sliding time window with a preset time length, sliding with a preset sliding step size, sliding once to update the current sample set once to obtain the updated sample set, screening the updated sample set to obtain representative samples, and training the financial risk control model with the representative samples corresponding to the updated sample set.
Citation Information
Patent Citations
A parameter adaptive clustering method based on density
CN109271424A
Intelligent grouping method and device for similar patients, equipment and storage medium
CN111739634A
Training sample selection method based on clustering and active learning
CN116662832A
Small sample event element intelligent extraction method based on clustering algorithm
CN118193738A
Training device, detection system, training method, and training program
US20220269779A1