Method, device, equipment and storage medium for processing bank risk control data
By preprocessing and multi-dimensionally analyzing risk control data, combined with improved detection methods, the problem of traditional methods being unable to identify abnormal behavior has been solved, more efficient abnormal behavior detection and risk warning have been achieved, and the bank's risk control data processing capabilities have been improved.
Patent Information
- Application Number
- CN202510082327.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-20
- Publication Date
- 2025-09-23
- Estimated Expiration
- 2045-01-20
AI Technical Summary
Traditional risk control data processing methods have difficulty identifying abnormal behaviors from risk control data of different types and characteristics, resulting in some abnormal behaviors being unable to be identified and unable to meet the banks' high requirements for risk control data processing.
After obtaining risk control data, we perform preprocessing, including data cleaning, conversion, integration and verification, and then conduct multi-dimensional analysis and abnormal behavior detection. We use improved mean, standard deviation, clustering and machine learning methods to detect abnormal behavior and trigger risk warnings when anomalies are detected.
It improves the accuracy of detecting abnormal behaviors, can identify abnormal behaviors more accurately, and ensure the security and compliance of bank risk control data.
Smart Images

Figure CN119941372B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of bank risk control data processing, and in particular relates to a method and apparatus, equipment and storage medium for processing bank risk control data. Background Art
[0002] With the increasing complexity of banking operations and the surge in data volumes, data access logging and monitoring have become crucial for ensuring data security and compliance. Traditional risk control data processing methods struggle to identify abnormal behavior within risk control data of varying types and characteristics, making them unable to meet banks' high risk control data processing requirements. This results in only identifying some abnormal behaviors while leaving others unidentified, making it difficult to meet banks' stringent risk control data processing requirements. Summary of the Invention
[0003] The present invention provides a method and apparatus for processing bank risk control data, equipment, and storage medium, which can analyze risk control data from multiple perspectives and more accurately determine abnormal behavior.
[0004] On the one hand, a method for processing bank risk control data is provided, comprising:
[0005] Obtain risk control data;
[0006] Pre-process risk control data to form standard risk control data;
[0007] Conduct multi-dimensional analysis and abnormal behavior detection on standard risk control data to identify abnormal behavior;
[0008] When abnormal behavior is detected, a risk warning is triggered.
[0009] Optionally, preprocess the risk control data, including:
[0010] Perform data cleaning, including removing duplicate data from risk control data, filling missing values in risk control data, and removing outliers in risk control data;
[0011] Perform data conversion, including data type conversion, data standardization, and data encoding of risk control data;
[0012] Conduct data integration, including data merging and data correlation of risk control data;
[0013] Perform data verification, including data integrity verification, data consistency verification, and data rationality verification of risk control data.
[0014] Optionally, perform multi-dimensional analysis on standard risk control data, including:
[0015] Standard risk control data is analyzed in the time, user, business, and data content dimensions. In the time dimension, the number of user visits in the next time period is analyzed based on the distribution of the number of user visits in different time periods. In the user dimension, a user behavior model is constructed and user behavior is analyzed based on the user behavior model. In the business dimension, risk points in the bank's key businesses are identified and data monitoring is performed on these risk points. In the data content dimension, the number of user visits to sensitive content is analyzed.
[0016] When the analysis results of the time dimension, user dimension, business dimension and data content dimension do not meet the risk indicator threshold, it is determined to be abnormal behavior.
[0017] Optionally, perform abnormal behavior detection on standard risk control data, including:
[0018] Standard risk control data is tested in turn using statistical-based anomaly detection methods, rule-based anomaly detection methods, and machine learning-based anomaly detection methods to identify abnormal behavior.
[0019] Optionally, the standard risk control data is detected according to the statistical-based anomaly detection method, which is to detect the standard risk control data according to the improved mean anomaly detection algorithm, the improved standard difference anomaly detection algorithm, and the improved cluster anomaly detection algorithm in sequence;
[0020] The improved mean anomaly detection algorithm is used to analyze standard risk control data, including:
[0021] Use the sliding window method to obtain standard risk control data and calculate the weighted average of the standard risk control data within the sliding window;
[0022] When the absolute value of the difference between the risk control data and the weighted average value does not meet the first range, the risk control data is determined to be an outlier; wherein, each time the sliding window method is used to determine the weighted average value within the sliding window, the calculation weight of the weighted average value and the size of the first range are dynamically adjusted according to the changing trend of the standard risk control data through an adaptive algorithm;
[0023] The improved standard deviation detection algorithm is used to analyze standard risk control data, including:
[0024] Analyze data using kernel density estimation or mixture distribution models to identify abnormal behavior;
[0025] The improved clustering anomaly detection algorithm is used to analyze standard risk control data, including:
[0026] Conduct multiple cluster analyses on standard risk control data, using different initial cluster centers each time, to determine the optimal initial cluster center;
[0027] In the process of analyzing standard risk control data using the optimal initial cluster center, the K value is determined using the elbow method or silhouette coefficient method;
[0028] Abnormal clusters formed by clustering are determined through local density estimation and relative distance metrics to identify abnormal behaviors.
[0029] Optionally, abnormal behavior detection is performed on standard risk control data using anomaly detection methods based on machine learning, including:
[0030] Increase the normal samples of standard risk control data through data augmentation technology, and increase the abnormal samples of standard risk control data by simulating abnormal generation;
[0031] Building a machine learning model, wherein the machine learning model includes one of a support vector machine, a random forest, and a neural network;
[0032] Normal samples and abnormal samples are used to train the machine learning model; the learning method includes active learning or semi-supervised learning; regularization methods, ensemble learning methods or optimization of the hyperparameters of the machine learning model are used during the training process of the machine learning model to prevent overfitting of the machine learning model, and the hyperparameter optimization method includes any one of grid search, random shrinkage and Bayesian optimization.
[0033] Optionally, the bank risk control data processing method further includes:
[0034] Periodically generate analysis reports based on multi-dimensional analysis results and abnormal behavior detection results in historical time periods.
[0035] In another aspect, a device for processing bank risk control data is provided, comprising:
[0036] Acquisition module, used to obtain risk control data;
[0037] The preprocessing module is used to preprocess the risk control data to form standard risk control data;
[0038] Anomaly detection module, used to perform multi-dimensional analysis and abnormal behavior detection on standard risk control data to identify abnormal behavior;
[0039] The early warning module is used to trigger risk warnings when abnormal behavior is detected.
[0040] On the other hand, an electronic device is provided, which is the device for processing bank risk control data as described above.
[0041] On the other hand, a computer-readable storage medium is provided, in which at least one program code is stored. The program code is executed by a processor to implement the bank risk control data processing method as described in any one of the above items.
[0042] The beneficial effects brought about by the technical solution provided by the present invention are:
[0043] The present invention provides a method for processing risk control data. This method obtains and preprocesses risk control data to generate standardized risk control data. Multi-dimensional analysis and abnormal behavior detection are then performed on the standardized risk control data. By employing different methods and analyzing the standardized risk control data from different perspectives, the problem of some abnormal behaviors not being detected due to the complexity and varying characteristics of the risk control data can be resolved, thereby improving the accuracy of abnormal behavior detection. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0045] Figure 1 A flow chart of a method for processing bank risk control data provided by the present invention;
[0046] Figure 2 This is a structural block diagram of a bank risk control data processing device provided by the present invention;
[0047] Figure 3 This is a structural block diagram of an electronic device provided by the present invention.
[0048] 11: Acquisition module; 12: Preprocessing module; 13: Anomaly detection module; 14: Early warning module;
[0049] 21: Processor; 22: Memory. DETAILED DESCRIPTION
[0050] To make the objectives, technical solutions, and advantages of the present invention more clear, the technical solutions of the present invention will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the embodiments described are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts shall fall within the scope of protection of the present invention.
[0051] Figure 1This is a flow chart of a method for processing bank risk control data provided by the present invention. Figure 1 ,include:
[0052] S101. Obtain risk control data.
[0053] In step S101, risk control data in the business system includes data from business modules such as deposits, loans, and credit cards, as well as transaction information and account status. Risk control data in the channel system includes data from channel systems such as online banking, mobile banking, and self-service terminals. Channel systems record customer operations performed through various channels, such as logins, inquiries, and transfers. Risk control data from external data sources includes data provided by credit reporting agencies and anti-fraud service providers to obtain more comprehensive risk information.
[0054] Risk control data collection methods include database log collection. For relational databases, data collection can be performed using the database's logging function. Database logs record all database operations, including inserts, updates, and deletes. By parsing database logs, detailed information about changes in business data can be obtained.
[0055] Risk control data collection also includes system log collection. Banks' various business and channel systems typically generate system logs, recording system status and operational information. Log collection tools such as Flume and Logstash are used to collect and process system logs.
[0056] Risk control data collection also includes network traffic collection. For data that cannot be obtained through system logs or database logs, network traffic collection technology can be considered. By deploying traffic collection equipment on the network, network packets can be captured and the required data extracted from them.
[0057] S102: Preprocess the risk control data to form standard risk control data.
[0058] In one example, step S102 includes:
[0059] Step 1: Perform data cleaning, including removing duplicate data in risk control data, filling missing values in risk control data, and removing outliers in risk control data.
[0060] The collected data may contain duplicate records, which need to be deduplicated to ensure data accuracy and uniqueness. This can be done using hash algorithms, database deduplication functions, and other methods.
[0061] Identify missing values in your dataset and address them as appropriate. For missing values in important fields, you can fill them by querying other data sources or using default values. For missing values that cannot be filled, consider marking the record as an outlier or deleting it, but exercise caution to avoid compromising the accuracy of your data analysis.
[0062] Remove outliers. Detect outliers in a data set, such as unusually large or small transaction amounts, unusual transaction times, etc. Use statistical methods (such as mean, standard deviation, and boxplots) to identify outliers and address them based on business rules, such as deleting them, performing data corrections, or marking them as abnormal.
[0063] Step 2: Perform data conversion, including data type conversion, data standardization, and data encoding of risk control data.
[0064] Among them, the data formats of different data sources may be different, and format conversion is required to enable unified storage and analysis.
[0065] For example, standardizing date formats to a specific standard format and converting string data to numeric types can be done. Data can be standardized to make data from different data sources comparable. For example, transaction amounts can be normalized and customer ratings can be standardized. Statistical methods such as Z-score normalization and Min-Max normalization can be used for data standardization.
[0066] Step 3: Perform data integration, including data merging and data association of risk control data.
[0067] This involves merging data from different data sources to obtain more comprehensive information. For example, data such as basic customer information, transaction records, and risk assessment results can be merged. This data merging can be done using database join operations (such as inner joins and outer joins) or data integration tools.
[0068] Data association establishes relationships between data to facilitate multi-dimensional analysis and monitoring. For example, it can link customer transaction records with their risk ratings, or link the time and location of transactions with risk events. Data association can be established using database indexes, views, or data mining algorithms.
[0069] Step 4: Perform data verification, including data integrity verification, data consistency verification, and data rationality verification of risk control data.
[0070] Data integrity verification involves checking the integrity of key fields in a dataset, such as customer number, transaction time, and transaction amount. This can be done using database constraints (such as non-null constraints and unique constraints) or data validation tools.
[0071] Data consistency verification verifies the consistency of data within a dataset. For example, whether basic customer information is consistent across different data sources, or whether the sum of transaction records matches the account balance, can be performed using data comparison tools or data auditing tools.
[0072] Data rationality verification checks whether the data complies with business rules and logic. For example, whether the transaction amount is within a reasonable range, whether the transaction time complies with business processes, etc. Data rationality verification can be performed using a business rule engine or a data rationality checking tool.
[0073] By implementing the above data preprocessing, we can effectively improve data quality and availability, providing a reliable data foundation for the multi-dimensional data access recording and monitoring system of the bank's risk control data mart. At the same time, data preprocessing implementation also needs to be adjusted and optimized according to specific business needs and data characteristics to ensure data processing efficiency and accuracy.
[0074] S103. Perform multi-dimensional analysis and abnormal behavior detection on standard risk control data to determine abnormal behavior.
[0075] In step S103, a data warehouse is first constructed to store data access records from different data sources. The data warehouse should have efficient data storage and query capabilities and be able to support large-scale data processing and analysis.
[0076] Design a reasonable data warehouse architecture, including fact tables and dimension tables. Fact tables store specific data access records, while dimension tables provide different analysis perspectives, such as time, user, and business dimensions. Data is extracted from various data sources through data ETL (Extract Transform Load), cleansed, transformed, and loaded into the data warehouse.
[0077] In step S103, a multi-dimensional analysis is performed on the standard risk control data, including:
[0078] The first step is to analyze standard risk control data in the time, user, business, and data content dimensions. Within the time dimension, analyze the number of user visits in the next time period based on the distribution of user visits in different time periods. Within the user dimension, build a user behavior model and analyze user behavior based on the user behavior model. Within the business dimension, identify risk points in the bank's key businesses and conduct data monitoring on these risk points. Within the data content dimension, analyze the number of times users visit sensitive content.
[0079] In the second step, when the analysis results of the time dimension, user dimension, business dimension and data content dimension do not meet the risk indicator threshold, it is determined to be abnormal behavior.
[0080] Exemplarily, performing multi-dimensional analysis includes:
[0081] Time Dimension Analysis: Analyze the distribution of data access records at different time points, such as daily, weekly, and monthly visit trends. Use time series analysis to predict future visit trends and proactively implement risk prevention measures. Detect unusual temporal patterns, such as high volumes of data access during non-business hours or holidays, which may indicate potential risk events.
[0082] User-centric analysis: Analyze the data access behavior of different user groups. Users can be categorized based on factors such as role, permissions, and geographic location to study their access habits and risk profiles. Identify abnormal user behavior, such as frequent access to sensitive data and unusual access time distribution. Build user behavior models using user profiling technology to promptly identify abnormal users and implement appropriate risk control measures.
[0083] Business-Dimensional Analysis: Analyze data access records across different business areas, based on the bank's business processes and risk factors. For example, analyze access to loan, credit card, and fund transactions to identify potential risk areas. Monitor changes in business indicators, such as loan approval rates and credit card delinquency rates, and correlate these with data access records to identify potential risk factors.
[0084] Data content analysis: Analyzes the specific data content in data access records. For example, this includes analyzing the types of sensitive data accessed and the distribution of data values. This also detects abnormal data content, such as large amounts of access to specific sensitive data or data values outside normal ranges. Using data mining techniques, this helps identify potential data leaks and unusual transaction behaviors.
[0085] Risk indicator setting: Based on the results of multi-dimensional analysis, a series of risk indicators are established to monitor the risk level of data access behavior in real time. Risk indicators can include abnormal growth in access volume, the proportion of abnormal user behavior, and the frequency of sensitive data access. Thresholds for risk indicators are determined, and when an indicator exceeds the threshold, a risk warning is triggered.
[0086] Data Analysis Technologies: A variety of data analysis techniques, such as data mining, machine learning, and statistical analysis, can be used to conduct in-depth analysis of data access records. These techniques can help identify potential risk patterns and abnormal behavior, improving the accuracy and timeliness of risk warnings. Select data analysis tools and algorithms that are appropriate for the bank's data characteristics and analysis needs, such as Python, R, SAS, and SPSS.
[0087] Data visualization tools: Use data visualization tools to present the results of multi-dimensional analysis in intuitive charts and graphs. Data visualization can help stakeholders better understand data and risk profiles, improving decision-making efficiency and accuracy. Data visualization tools include Tableau, PowerBI, and Echarts.
[0088] Real-time monitoring technology: To achieve real-time risk warnings, real-time monitoring technology is required to monitor data access records. Stream processing technologies such as Apache Kafka and Apache Flink can be used to process and analyze real-time data. Establishing a real-time monitoring system ensures timely detection and response to risk events. In summary, the multi-dimensional data access records and monitoring system designed for bank risk control data marts require a comprehensive approach to data storage and management, multi-dimensional analysis methods, and risk warning and reporting technologies to achieve comprehensive monitoring and risk analysis of data access behavior. The implementation of this analysis layer provides strong support for banks' risk management and ensures the security and stability of their operations.
[0089] In step S103, abnormal behavior detection is performed on the standard risk control data, including:
[0090] Standard risk control data is tested in turn using statistical-based anomaly detection methods, rule-based anomaly detection methods, and machine learning-based anomaly detection methods to identify abnormal behavior.
[0091] In one example, the standard risk control data is detected according to the statistical-based anomaly detection method, which is to detect the standard risk control data in sequence according to the improved mean anomaly detection algorithm, the improved standard difference anomaly detection algorithm, and the improved clustering anomaly detection algorithm.
[0092] Among them, the improved mean anomaly detection algorithm is used to analyze standard risk control data, including:
[0093] Use the sliding window method to obtain standard risk control data and calculate the weighted average of the standard risk control data within the sliding window;
[0094] When the absolute value of the difference between the risk control data and the weighted average value does not meet the first range, the risk control data is determined to be an outlier; wherein, each time the sliding window method is used to determine the weighted average value within the sliding window, the calculation weight of the weighted average value and the size of the first range are dynamically adjusted according to the changing trend of the standard risk control data through an adaptive algorithm.
[0095] In one example, the improved mean anomaly detection algorithm is used to analyze standard risk control data, including:
[0096] First, use robust statistics instead of or in combination with them.
[0097] Median: The median is the value in the middle of the sorted data, unaffected by extreme values. It can be used instead of the mean as an estimate of the central location. For example, when calculating residential housing prices, using the median can prevent the central price estimate from being affected by a small number of luxury home prices. When detecting outliers, data points that are too far from the median can be considered outliers. For example, data points with a value greater than the median plus three times the median absolute deviation (MAD, which is the median of the distances from the data point to the median) can be considered outliers.
[0098] Weighted mean: Each data point is assigned a weight based on its reliability or importance, and then a weighted mean is calculated. For example, in sensor network data, more accurate sensor data is given a higher weight, while less accurate sensor data is given a lower weight. This reduces the impact of inaccurate data on the mean, making the mean more representative of the true data center trend, thereby improving the accuracy of outlier detection based on the mean.
[0099] Second, make adjustments based on data distribution characteristics.
[0100] Data preprocessing and distribution fitting: If the data exhibit a clearly non-normal distribution, preprocessing can be performed, such as using a Box-Cox transformation to convert the data to an approximate normal distribution. The mean can then be calculated to detect outliers. Furthermore, distribution fitting can be performed on the data, such as using a normal distribution, lognormal distribution, or other appropriate distribution model. A more reasonable outlier range can be determined based on the fitted distribution parameters and the actual data. For example, for some economic data with a positively skewed distribution, a logarithmic transformation can be used to bring it closer to a normal distribution. The outlier range can then be defined based on the mean and standard deviation of the transformed data.
[0101] Consider quantile information: In addition to the mean, quantiles can be used to gain a more comprehensive understanding of the data distribution. For example, the first quartile (Q1) and third quartile (Q3) of the data can be calculated and the interquartile range (IQR = Q3 - Q1) determined. Data points less than Q1 - 1.5 * IQR or greater than Q3 + 1.5 * IQR can be identified as outliers. This quantile-based method (such as the boxplot method) is also effective in detecting anomalies in skewed data and can be combined with the mean. When the anomaly detection results based on the mean disagree with those based on quantiles, further data analysis can be performed to determine whether anomalies are truly present.
[0102] Third, dynamically update the mean and outlier judgment criteria.
[0103] Sliding Window Method: This method is used when processing time series or dynamically changing data. For example, for stock price data, a fixed-length sliding window is set and only the mean of the data within the window is calculated. As the window slides, the mean is updated to reflect the latest data changes. Furthermore, the outlier determination criteria can be dynamically adjusted based on the fluctuations of the data within the window, such as determining a dynamic outlier range based on the standard deviation of the data within the window.
[0104] Adaptive algorithms: Use adaptive algorithms to automatically adjust mean and outlier criteria based on data trends. For example, the exponentially weighted moving average (EWMA) algorithm gives more weight to recent data and can quickly adapt to data changes. When detecting outliers, the dynamic mean calculated by the EWMA and the standard deviation, which is adaptively adjusted based on data fluctuations, can be used to determine whether the data is abnormal.
[0105] The improved standard deviation detection algorithm is used to analyze standard risk control data, including:
[0106] Analyze data using kernel density estimation or mixture distribution models to identify anomalous behavior.
[0107] In this embodiment, an improved standard deviation abnormality detection algorithm is used to analyze standard risk control data, including:
[0108] First, use non-parametric methods or distribution adaptive technology.
[0109] Kernel Density Estimation (KDE): A nonparametric method for estimating the probability density function of data without relying on specific distributional assumptions. KDE allows you to identify data points in low-density areas as outliers based on the actual distribution of the data, rather than the standard deviation of a normal distribution.
[0110] Mixture Distribution Models: When data exhibits a complex distribution (such as a multimodal distribution), a mixture distribution model, such as the Gaussian Mixture Model (GMM), can be used to fit the data. The GMM assumes that the data is composed of a mixture of multiple Gaussian distributions. By estimating the parameters of each Gaussian distribution (including the mean and standard deviation) and the proportion of their mixture, it can more accurately describe the data distribution. When detecting outliers, the probability of each data point belonging to each Gaussian distribution can be used to determine whether it is an anomaly, rather than relying solely on a single standard deviation.
[0111] Second, the local standard deviation and local anomaly factor are used in combination.
[0112] Calculate local standard deviation: In order to reduce the impact of local changes on the overall standard deviation, you can calculate the local standard deviation.
[0113] Incorporating the Local Outlier Factor (LOF): LOF is a density-based method for detecting local anomalies. The local standard deviation can be combined with LOF to detect anomalies by comparing the dispersion and density differences between a data point and its local neighborhood.
[0114] Third, combine semantic information and feature engineering.
[0115] Introducing contextual features: Adding data-related contextual features to the dataset.
[0116] Multivariate Analysis and Principal Component Analysis (PCA): If your data contains multiple variables, you can perform multivariate analysis. PCA is a commonly used method that transforms multiple correlated variables into a small number of uncorrelated principal components. In this new principal component space, you can recalculate the standard deviation or other statistics to identify anomalies.
[0117] The improved clustering anomaly detection algorithm is used to analyze standard risk control data, including:
[0118] Conduct multiple cluster analyses on standard risk control data, using different initial cluster centers each time, to determine the optimal initial cluster center;
[0119] In the process of analyzing standard risk control data using the optimal initial cluster center, the K value is determined using the elbow method or silhouette coefficient method;
[0120] Abnormal clusters formed by clustering are determined through local density estimation and relative distance metrics to identify abnormal behaviors.
[0121] In this embodiment, an improved clustering anomaly detection algorithm is used to analyze standard risk control data, including:
[0122] First, optimize clustering algorithm parameters and initialization process
[0123] Multiple runs and evaluations: For algorithms like K-Means that are sensitive to initial values, you can run the algorithm multiple times, using different initial cluster centers each time. For example, using the K-Means++ initialization method can select more appropriate initial centers, or by randomly selecting initial centers multiple times and comparing the clustering results, you can select the optimal clustering solution to improve the accuracy of outlier detection.
[0124] Second, automatic parameter selection: Use methods to automatically determine the parameters of the clustering algorithm. For example, when determining the K value, methods such as the Elbow Method and the Silhouette Coefficient can be used. The Elbow Method observes the curve of the clustering error as it changes with the K value and finds the "elbow" point of the curve. The K value corresponding to this point is usually a relatively appropriate number of clusters. The Silhouette Coefficient measures the closeness of each data point to its own cluster and adjacent clusters. The appropriate K value is selected by maximizing the Silhouette Coefficient.
[0125] Third, ensemble clustering methods combine the results of multiple clustering algorithms with different parameters or assumptions. For example, using K-Means clustering results with different K values, or combining K-Means with density-based clustering (such as DBSCAN), can identify outliers through comprehensive judgment. This ensemble approach can reduce errors caused by parameter or assumption problems with a single clustering algorithm.
[0126] Fourth, combine local density and distance information to distinguish abnormal clusters and abnormal points.
[0127] Local density estimation: Calculates the local density around each data point. Data points that are in low-density areas and far from high-density areas are more likely to be outliers. For example, the Local Outlier Factor (LOF) algorithm is combined with clustering to measure the local density difference between a data point and its neighbors. After clustering, the LOF value is calculated for the data points within each cluster. Data points with high LOF values may be outliers even within the cluster. For small, low-density clusters, further analysis can be performed on their relationship with other clusters and the distribution of data points within them to determine whether they are normal small clusters or abnormal clusters.
[0128] Relative distance metrics: In addition to considering the absolute distance of a data point from the cluster center, we can also consider its relative distance to other data points within the cluster. For example, if the distance from a data point to all other data points in a cluster is significantly greater than the average distance within the cluster, then this data point may be an outlier. We can calculate a relative distance matrix between data points and combine it with the clustering results to detect outliers.
[0129] Fifth, adopt efficient clustering algorithms and approximation techniques.
[0130] Sampling-based clustering: For large datasets, sampling can be performed first, followed by clustering of the sampled small datasets. Outliers in the population can then be inferred based on the clustering results and the relationship between the sample and the population. For example, simple random sampling or stratified sampling can be used to reduce computational costs while ensuring representative samples. Incremental clustering algorithms can also be used, which can quickly update clustering results as new data is added, without re-clustering the entire dataset, improving the efficiency of outlier detection.
[0131] Approximate clustering algorithms: These use approximate computational methods to reduce computing resource consumption. For example, locality-sensitive hashing (LSH) can be used to approximate the distance between data points, accelerating the clustering process. By sacrificing some accuracy, clustering and outlier detection can be completed on large datasets within an acceptable timeframe.
[0132] In step S103, abnormal behavior detection is performed on the standard risk control data according to an abnormality detection method based on machine learning, including:
[0133] The first step is to increase the normal samples of standard risk control data through data augmentation technology, and increase the abnormal samples of standard risk control data by simulating abnormal generation;
[0134] The second step is to build a machine learning model, which includes one of a support vector machine, a random forest, and a neural network;
[0135] The third step is to train the machine learning model using normal samples and abnormal samples; the learning method includes active learning or semi-supervised learning; during the training process of the machine learning model, regularization methods, ensemble learning methods or optimization of the hyperparameters of the machine learning model are used to prevent the machine learning model from overfitting, and the hyperparameter optimization method includes any one of grid search, random shrinkage and Bayesian optimization.
[0136] In this embodiment, abnormal behavior detection is performed on standard risk control data according to an abnormality detection method based on machine learning, including:
[0137] First, improve the quality and diversity of training data.
[0138] Data augmentation techniques: Data augmentation can be used to address insufficient data. For example, in image anomaly detection, operations such as rotation, flipping, scaling, and adding noise can be used to increase the number of normal samples. Anomalous samples can be augmented by simulating anomalies. For example, in financial fraud detection, fraudulent transactions can be simulated by modifying certain features of normal transaction data.
[0139] Active learning and semi-supervised learning: Active learning allows the model to proactively select the most valuable samples for labeling, reducing its reliance on large amounts of labeled data. Semi-supervised learning utilizes a large amount of unlabeled data alongside a small amount of labeled data to train the model. Data resampling: Resampling techniques are used to address data distribution bias. For minority classes (anomalies), oversampling can be performed using the SMOTE (Synthetic Minority Over-sampling Technique) algorithm to generate synthetic samples to increase the number of anomaly samples. For majority classes (normal classes), undersampling can be performed to balance the distribution of training data, enabling the model to better learn anomaly characteristics.
[0140] Second, optimize the model structure and training process to avoid overfitting and underfitting.
[0141] Regularization methods: Regularization techniques such as L1 and L2 regularization, Dropout, etc. are used during model training. In neural networks, Dropout randomly drops some neurons during training to prevent the model from over-relying on certain features, thereby reducing overfitting.
[0142] Model selection and ensemble learning: Select the appropriate model complexity through methods such as cross-validation. For example, when selecting the depth of a decision tree, use a validation set to evaluate the performance of the model at different depths and select the depth with the best performance. Ensemble learning can combine multiple different models to improve the model's generalization and robustness. For example, combining multiple neural networks with different structures or combining decision trees with support vector machines (SVMs) and using voting or weighted averaging to identify anomalies can reduce the overfitting or underfitting problems of a single model.
[0143] Adjust hyperparameters: Properly adjust model hyperparameters, such as the learning rate, number of layers, and number of neurons per layer in the neural network. You can use grid search, random search, or more advanced hyperparameter optimization algorithms (such as Bayesian optimization) to find the optimal hyperparameter combination to balance the model's fit and generalization capabilities.
[0144] Third, improve model interpretability.
[0145] Explainable AI (XAI) methods employ feature importance analysis. For example, in decision tree models, you can directly view the importance ranking of features. In neural networks, methods such as LIME (Local Interpretable Model-agnostic Explanations) or SHAP (SHapley Additive Explanations) can be used. These methods can help explain the factors that determine data anomalies. For example, in credit risk assessment anomaly detection, SHAP can explain how the model determines whether credit risk is abnormal based on factors such as a customer's age, income, and credit history.
[0146] Build simple and interpretable models to assist in explaining complex models: In addition to using XAI methods, you can also build simple and highly interpretable models (such as linear models and decision trees) to assist in explaining complex machine learning models.
[0147] In another example, abnormal behavior detection is performed on standard risk control data, including:
[0148] The improved method for detecting outliers based on deep learning is used to detect abnormal behavior in standard risk control data, including the following steps:
[0149] First, data enhancement and optimization
[0150] Data augmentation techniques increase the amount of data by transforming and synthesizing existing data. For example, in image anomaly detection, normal images can be rotated, flipped, scaled, and noisy to generate new samples. For time series data, data augmentation can be achieved using sliding windows and resampling. Furthermore, methods such as generative adversarial networks (GANs) can be used to generate virtual data with a distribution similar to real data to increase the number of anomaly samples.
[0151] Active learning and semi-supervised learning: Active learning effectively reduces reliance on large amounts of labeled data by enabling the model to proactively select the most valuable unlabeled samples for labeling. Semi-supervised learning, on the other hand, utilizes both large amounts of unlabeled data and small amounts of labeled data to train the model, enabling it to better learn the distribution characteristics of the data and improve its ability to detect anomalies. Data resampling: To address data imbalance, oversampling and undersampling techniques are employed. For anomalous samples in the minority class, oversampling algorithms such as SMOTE can be used to generate synthetic samples. For normal samples in the majority class, undersampling can be performed to balance the ratio of normal and anomalous samples in the training data, allowing the model to more evenly learn the characteristics of both classes.
[0152] Second, model improvement and regularization.
[0153] Regularization methods: Regularization methods, such as L1 and L2 regularization and dropout, are added during training to limit model complexity and prevent overfitting. For example, dropout is used in a neural network to randomly drop some neurons, making the model less dependent on certain features and thus improving generalization.
[0154] Model selection and ensemble learning: Use methods such as cross-validation to select models of appropriate complexity, and use ensemble learning to combine multiple different models. For example, multiple neural networks with different structures or different types of anomaly detection models (such as autoencoders and generative adversarial networks) can be integrated to comprehensively judge anomalies through strategies such as voting and weighted averaging, thereby improving the model's robustness and generalization capabilities.
[0155] Adjust hyperparameters: Optimize model hyperparameters, such as the learning rate, number of layers, number of neurons per layer, and convolution kernel size. Use hyperparameter optimization algorithms such as grid search, random search, and Bayesian optimization to find the optimal hyperparameter combination to balance the model's fit and generalization capabilities.
[0156] Third, computing efficiency is improved.
[0157] Model compression and quantization: Model compression techniques, such as pruning to remove unimportant connections or parameters and quantizing parameters to lower-precision data types, are used to reduce model storage space and computational complexity while minimizing model performance. Hardware acceleration: Utilizing more powerful computing hardware, such as GPUs, FPGAs, and ASICs, accelerates model training and inference. Furthermore, distributed training and parallel computing techniques can be employed to improve computational efficiency and shorten training time.
[0158] Improved optimization algorithms: Select more efficient optimization algorithms, such as Adagrad, Adadelta, RMSProp, and Adam, which use adaptive learning rate optimization algorithms. These algorithms can automatically adjust the learning rate based on the characteristics of the problem, speeding up model convergence and reducing training time.
[0159] Fourth, the model’s interpretability is enhanced.
[0160] Explainable AI (XAI) methods use feature importance analysis methods. For example, in decision tree models, feature importance ranking can be directly viewed. In neural networks, methods such as LIME and SHAP can be used to explain the factors that determine data as abnormal, helping people understand the model's decision-making process.
[0161] Build simple, interpretable models to help explain complex models: In addition to using XAI methods, you can also build simple, highly interpretable models (such as linear models and decision trees) to help explain complex deep learning models. By comparing the decision results and explanations of two models, you can improve your understanding of anomaly detection results.
[0162] S104. When abnormal behavior is detected, a risk warning is triggered.
[0163] Risk Warning Mechanism: Establish a real-time risk warning mechanism to promptly send warning information to relevant personnel when risk indicators reach warning thresholds. Warning information can be sent via email, SMS, system notifications, etc. Warning information is categorized and prioritized to enable relevant personnel to quickly respond to high-risk events.
[0164] S105. Periodically generate an analysis report based on the multi-dimensional analysis results and abnormal behavior detection results in the historical time period.
[0165] Risk Report Generation: Risk reports are regularly generated to summarize the analysis results and risk status of data access records. These reports should include charts and data from multi-dimensional analyses, trends in risk indicators, and detailed descriptions of risk events. These reports can provide decision support for bank management, helping them develop effective risk control strategies.
[0166] The multi-dimensional data access record and monitoring system report generation layer of the bank risk control data mart is implemented as follows:
[0167] The report includes:
[0168] First, overview.
[0169] System Operation Overview: This section briefly describes the overall operation of the data access logging and monitoring system during a specific time period, including key indicators such as system availability and data collection integrity. For example, you might mention that the data collection success rate reached [X]% and that the system maintained stable operation for [specific duration].
[0170] Risk Assessment Summary: This section provides a summary of the bank's risk profile, including the current risk level (e.g., low, medium, high) and key risk trends (whether it's increasing, decreasing, or stable). This section can be based on a preliminary analysis of multi-dimensional data, providing readers with a quick overview of the bank's risk profile.
[0171] Second, multi-dimensional data analysis report section.
[0172] Time Dimension Analysis Report: This report presents data access patterns at different time scales. For example, it displays daily, weekly, and monthly access trend charts, analyzing the timing of peaks and troughs and their relationship to banking operations. It also reports data access anomalies at specific times (such as holidays and peak business periods). For example, if data access suddenly increases by [X]% during a holiday, analyze the potential risks associated with this change.
[0173] User-Dimensional Analysis Report: This report categorizes and analyzes data access behavior based on factors such as user type (e.g., individual customer, corporate customer, bank employee), user role (e.g., teller, credit manager), and user location. It details the access frequency of different user groups, the primary data types accessed, and any abnormal access behaviors. For example, the report indicates that corporate customers in a certain region accessed loan approval data an unusually high [X] times during a specific time period.
[0174] Business Dimension Analysis Report: This report analyzes the relationship between data access and business operations for each of the bank's major business lines (such as savings, loans, credit cards, and wealth management). It displays the volume of data accessed for each business area, key data points accessed, and potential risk factors. For example, within the loan business, a recent increase in access to high-risk loan customer data may indicate potential credit risk.
[0175] Data Content Analysis Report: This report provides an in-depth analysis of the specific content of accessed data, including access to sensitive data. This report lists frequently accessed sensitive data types (e.g., customer ID numbers, account passwords, etc.), along with information such as the source and time of access. It also reports data content anomalies, such as large-scale, concentrated access to a specific type of highly confidential customer information.
[0176] Third, the anomaly detection report section.
[0177] Abnormal Event Statistics: This summarizes the number of abnormal events detected during the monitoring period and categorizes them by abnormality type (such as abnormal access frequency, abnormal access path, abnormal data download, etc.). For example, a report might show that in the past month, [X] abnormal access frequency events and [Y] abnormal access path events were detected.
[0178] Detailed Analysis of Major Abnormal Events: We conduct in-depth analysis of major abnormal events with high risk levels, including the time of occurrence, users involved, data accessed, potential loss assessment, and preliminary cause analysis. For example, we provide a detailed description of an incident involving the abnormal access of a large amount of customer financial information, analyzing whether it was caused by an external attack or internal misconduct.
[0179] Fourth, risk warning and response report section.
[0180] Risk Warning Status Report: This report lists the indicators and thresholds that trigger risk warnings, and reports the number of warning triggers and their accuracy during the monitoring period. For example, if an alert is triggered when data access exceeds [X] times within an hour, and the alert was triggered [Y] times in the past week, the accuracy rate of the alert is [Z]%.
[0181] Response Implementation Report: This report describes the response measures taken in response to different risk warnings (such as restricting user access, strengthening security verification, initiating an investigation process, etc.) and an evaluation of the effectiveness of these measures. For example, the report may indicate that restricting the access rights of a suspicious user effectively curbed the related abnormal access behavior.
[0182] The report generation process is as follows:
[0183] First, data integration and processing.
[0184] Data extraction: Extract various data related to the report content from the data warehouse, including pre-processed and analyzed multi-dimensional data access records, anomaly detection results, risk warning information, etc. Use data query languages (such as SQL) or specialized data extraction tools to obtain the required data from the storage system according to predetermined rules and conditions.
[0185] Data association and integration: Data extracted from different data sources is associated and integrated to ensure report consistency and completeness. For example, user-level data can be associated with business-level data using user IDs to comprehensively display user behavior across different businesses in reports. Effective data integration can be achieved using data integration tools or custom association algorithms.
[0186] Second, the report generation techniques and tools are as follows:
[0187] Template Engine: Use a template engine (such as Jinja2 or Velocity) to define the report template structure. The template defines the report's sections, formatting (such as titles, paragraphs, charts, and tables), and where to populate the data. By populating the template with extracted and processed data, you can quickly generate a uniformly formatted, aesthetically pleasing report.
[0188] Data analysis and visualization tools: Combine data analysis tools (such as Python's Pandas and NumPy) with visualization tools (such as Matplotlib, Seaborn, and Tableau) to generate charts and data visualizations for reports. These tools can present complex data in intuitive graphics (such as bar charts, line charts, pie charts, and heat maps) and tables, making them easier for readers to understand and analyze. For example, a bar chart can be used to compare data access volume across different business lines, while a line chart can be used to display access trends over time.
[0189] Text generation technology: Leveraging natural language generation technology, we automatically generate report text descriptions based on data and analysis results. We can use rule-based text generation methods or machine learning-based text generation models (such as those used in localization applications like GPT) to generate fluent and accurate text content. Furthermore, we manually review and adjust the generated text to ensure professionalism and logicality.
[0190] Third, report distribution and storage.
[0191] Distribution channels: Determine the report's distribution channels, including those for the bank's risk management department, senior management, business departments, and other relevant personnel. This can be done via email or internal office systems to ensure timely access to the report information. For risk situations requiring immediate attention, you can also configure real-time push notifications within the system.
[0192] Storage Management: Store generated reports in a dedicated document management system or data repository for subsequent query, audit, and historical comparative analysis. Manage report storage by category, establishing a reasonable storage directory structure based on factors such as report type (e.g., daily, weekly, special reports), and timeframe. Ensure the security of report storage and set appropriate access permissions to prevent information leakage.
[0193] In this embodiment, a user interface may be further configured. The user interface layer is implemented as follows:
[0194] First, interface design goals.
[0195] Intuitive: The user interface should be intuitive and easy to understand, enabling users from all backgrounds (including risk managers, technical staff, and senior management) to quickly understand and operate it. This should be done through a clear layout, explicit icons, and concise labels to reduce the learning curve for users.
[0196] Real-time: Provide real-time data display and updates to ensure users are aware of data access and risk status. The interface should be able to quickly respond to data changes so that users can make decisions quickly.
[0197] Interactivity: Supports user interaction with the system, such as querying specific data, setting parameters, generating reports, etc. The interaction method should be simple and easy to use to improve user work efficiency.
[0198] Visualization: Using visualization technology, complex data can be displayed in the form of intuitive charts, graphs, and maps to help users better understand data and risk trends.
[0199] Security: Ensure the security of the user interface to prevent unauthorized access and operation. Adopt strict user authentication and authorization mechanisms to protect sensitive data and system functions.
[0200] Second, interface layout and functional modules
[0201] The data overview module includes a real-time data dashboard and a time trend chart.
[0202] Real-time data dashboard: Displays real-time data of key indicators, such as data access volume, number of abnormal events, risk level, etc. Using intuitive charts and numbers, users can understand the overall status of the system at a glance.
[0203] Time Trend Chart: Displays the changing trends of data access volume and risk indicators over time. Users can select different time frames to analyze long-term and short-term risk trends.
[0204] The multi-dimensional analysis module includes: time dimension analysis, user dimension analysis, business dimension analysis, and data content dimension analysis.
[0205] Time Dimension Analysis: Provides data access analysis by time dimension. Users can select a specific time period to view data access patterns, peak and trough times, and the distribution of abnormal events within that time period.
[0206] User dimension analysis: Allows users to analyze data access based on factors such as user type, role, and geographic location. Displays access behavior characteristics and anomalies of different user groups.
[0207] Business Dimension Analysis: This tool displays the relationship between data access and business operations for each of the bank's business areas. This allows users to gain in-depth insights into the risk profile and data access patterns of specific businesses.
[0208] Data content analysis: Analyzes the specific content of accessed data, including access to sensitive data. Users can view frequently accessed sensitive data types and access sources.
[0209] The anomaly detection module includes abnormal event list, abnormal event analysis,
[0210] Abnormal Event List: This lists all abnormal events detected by the system, including detailed information such as the time, type, users involved, and data content. Users can categorize, filter, and sort abnormal events to quickly locate and address high-risk events.
[0211] Abnormal Event Analysis: Provides in-depth analysis of specific abnormal events. Users can view event details, relevant data access records, and possible cause analysis. The system should also provide recommended response measures to help users promptly address abnormal events.
[0212] The risk warning module includes warning indicator settings and warning notifications.
[0213] Warning indicator settings: Allow users to set risk warning indicators and thresholds. Users can customize warning rules based on business needs and risk preferences to ensure that the system can detect potential risks in a timely manner.
[0214] Warning notifications: When the system detects an event that meets the warning criteria, it promptly sends notifications to users. Notifications can be sent via email, SMS, or system pop-up windows, ensuring users can quickly respond to risk events.
[0215] The report generation module includes custom reports and report export.
[0216] Customized Reports: Users can customize the report content and format to meet their specific needs, generating personalized risk reports. Reports can include data overviews, multi-dimensional analysis, anomaly detection, and risk warnings, providing comprehensive risk assessment and decision-making support.
[0217] Report export: Supports exporting generated reports to PDF, Excel and other formats, making it convenient for users to share and archive with other personnel.
[0218] The user management module includes user authentication and authorization and user rights management.
[0219] User authentication and authorization: Ensure that only authorized users can access the system. Use strict user authentication mechanisms, such as username and password, digital certificates, etc. to protect system security.
[0220] User rights management: Different rights are assigned to users based on their roles and responsibilities. For example, risk managers can view and process all risk data, while ordinary users can only view certain data and reports.
[0221] Third, application of visualization technology.
[0222] Charts and Graphs: Use common chart types such as bar charts, line charts, and pie charts to display the distribution and changing trends of indicators such as data access volume, number of abnormal events, and risk level. Utilize visualization techniques such as heat maps and scatter plots to analyze data access hotspots and the distribution of abnormal events. Interactive charts allow users to access more data information by hovering or clicking on them.
[0223] Map visualization: For data analysis involving geographic location, map visualization technology can be used to display the distribution of users' visit locations and the geographic distribution of risk events. Maps support zooming, panning, and layer switching, making it easier for users to query and analyze geographic information.
[0224] Data Dashboard: Design a dashboard that displays key metrics and charts on a single page, allowing users to quickly understand the overall status of the system. The dashboard should support custom layouts and metric selection to meet the needs of different users.
[0225] Fourth, interaction design.
[0226] Menus and Navigation: Design concise and clear menus and navigation structures to help users quickly find the functional modules they need. Use drop-down menus, sidebars, and other methods to organize menus and improve interface usability.
[0227] Search and Filter: Provides a search function that allows users to quickly find specific data records and abnormal events. Supports filtering functions, allowing users to filter and sort data according to different conditions to better analyze and handle risk events.
[0228] Operation prompts and feedback: Provide clear operation prompts and feedback information when users are performing operations to ensure that users know whether their operations are successful. For incorrect operations, provide clear error prompts and provide solution suggestions.
[0229] Fifth, safety considerations
[0230] User authentication and authorization: Implement strict user authentication and authorization mechanisms to ensure that only authorized users can access the system. Regularly update user passwords and strengthen password strength requirements to prevent password leaks.
[0231] Data encryption: Encrypt sensitive data for storage and transmission to protect user privacy and data security. Use encryption protocols such as SSL / TLS to ensure data security during network transmission.
[0232] Access control: Strictly control the access rights of different users to prevent unauthorized users from accessing sensitive data and system functions. Record user operation logs for audit and traceability.
[0233] Through the implementation of the above user interface layer, an intuitive, efficient and secure user interaction environment can be provided for the multi-dimensional data access recording and monitoring system of the bank risk control data mart, helping users to better manage and control risks.
[0234] Finally, the beneficial effects of this application include: First, improving data security: Through multi-dimensional recording and anomaly detection, potential security threats can be promptly discovered and prevented. Second, enhancing compliance: Regular compliance assessments are conducted to ensure that data access behavior complies with laws, regulations, and industry standards. Third, improving recording efficiency: Automating recording and report generation reduces the workload of administrators and improves recording efficiency. Fourth, enhancing traceability: Detailed records of user access behavior provide strong evidence for security incident investigations.
[0235] Figure 2 This is a structural block diagram of a bank risk control data processing device provided by the present invention. Figure 2 ,include:
[0236] Acquisition module 11, used to obtain risk control data;
[0237] The preprocessing module 12 is used to preprocess the risk control data to form standard risk control data;
[0238] Anomaly detection module 13, used to perform multi-dimensional analysis and abnormal behavior detection on standard risk control data to determine abnormal behavior;
[0239] The early warning module 14 is used to trigger a risk early warning when abnormal behavior is detected.
[0240] Figure 3 This is a structural block diagram of an electronic device provided by the present invention. Figure 3 , electronic devices may include Figure 2 The bank risk control data processing device generally includes a processor 21 and a memory 22 .
[0241] Processor 21 may include one or more processing cores, such as a quad-core processor or an octa-core processor. Processor 21 may be implemented in hardware using at least one of the following: a DSP (Digital Signal Processing), an FPGA (Field-Programmable Gate Array), or a PLA (Programmable Logic Array). Processor 21 may also include a main processor and a coprocessor. The main processor is used to process data in the awake state, also known as a CPU (Central Processing Unit); the coprocessor is a low-power processor used to process data in the standby state. Memory 22 may include one or more computer-readable storage media, which may be non-transitory. Memory 22 may also include high-speed random access memory and non-volatile memory, such as one or more disk storage devices or flash memory storage devices. In some embodiments, the non-transitory computer-readable storage medium in memory 22 is used to store at least one instruction, which is executed by processor 21 to implement the bank risk control data processing method provided by an electronic device in the method embodiments of this application.
[0242] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention.
Claims
1. A method for processing bank risk control data, characterized in that: include: Obtain risk control data; Pre-process risk control data to form standard risk control data; Conduct multi-dimensional analysis and abnormal behavior detection on standard risk control data to identify abnormal behavior; Detect abnormal behavior on standard risk control data, including: Standard risk control data is tested using statistical, rule-based, and machine learning-based anomaly detection methods to identify abnormal behavior. The standard risk control data is tested according to the statistical anomaly detection method, which is to test the standard risk control data according to the improved mean anomaly detection algorithm, the improved standard difference anomaly detection algorithm, and the improved cluster anomaly detection algorithm in sequence; The improved mean anomaly detection algorithm is used to analyze standard risk control data, including: Use the sliding window method to obtain standard risk control data and calculate the weighted average of the standard risk control data within the sliding window; When the absolute value of the difference between the risk control data and the weighted average value does not meet the first range, the risk control data is determined to be an outlier; wherein, each time the sliding window method is used to determine the weighted average value within the sliding window, the calculation weight of the weighted average value and the size of the first range are dynamically adjusted according to the changing trend of the standard risk control data through an adaptive algorithm; The improved standard deviation detection algorithm is used to analyze standard risk control data, including: Analyze data using kernel density estimation or mixture distribution models to identify abnormal behavior; The improved clustering anomaly detection algorithm is used to analyze standard risk control data, including: Conduct multiple cluster analyses on standard risk control data, using different initial cluster centers each time, to determine the optimal initial cluster center; In the process of analyzing standard risk control data using the optimal initial cluster center, the K value is determined using the elbow method or silhouette coefficient method; Determine abnormal clusters formed by clustering through local density estimation and relative distance metrics to identify abnormal behaviors; When abnormal behavior is detected, a risk warning is triggered.
2. The method for processing bank risk control data according to claim 1, characterized in that: Preprocess risk control data, including: Perform data cleaning, including removing duplicate data from risk control data, filling missing values in risk control data, and removing outliers in risk control data; Perform data conversion, including data type conversion, data standardization, and data encoding of risk control data; Conduct data integration, including data merging and data correlation of risk control data; Perform data verification, including data integrity verification, data consistency verification, and data rationality verification of risk control data.
3. The method for processing bank risk control data according to claim 1, characterized in that: Conduct multi-dimensional analysis of standard risk control data, including: Standard risk control data is analyzed in the time, user, business, and data content dimensions. In the time dimension, the number of user visits in the next time period is analyzed based on the distribution of the number of user visits in different time periods. In the user dimension, a user behavior model is constructed and user behavior is analyzed based on the user behavior model. In the business dimension, risk points in the bank's key businesses are identified and data monitoring is performed on these risk points. In the data content dimension, the number of user visits to sensitive content is analyzed. When the analysis results of the time dimension, user dimension, business dimension and data content dimension do not meet the risk indicator threshold, it is determined to be abnormal behavior.
4. The method for processing bank risk control data according to claim 1, characterized in that: Detect abnormal behavior in standard risk control data using machine learning-based anomaly detection methods, including: Increase the normal samples of standard risk control data through data augmentation technology, and increase the abnormal samples of standard risk control data by simulating abnormal generation; Building a machine learning model, wherein the machine learning model includes one of a support vector machine, a random forest, and a neural network; Normal samples and abnormal samples are used to train the machine learning model; the learning method includes active learning or semi-supervised learning; regularization methods, ensemble learning methods or optimization of the hyperparameters of the machine learning model are used during the training process of the machine learning model to prevent overfitting of the machine learning model, and the hyperparameter optimization method includes any one of grid search, random shrinkage and Bayesian optimization.
5. The method for processing bank risk control data according to any one of claims 1 to 4, characterized in that: The bank risk control data processing method further includes: Periodically generate analysis reports based on multi-dimensional analysis results and abnormal behavior detection results in historical time periods.
6. A bank risk control data processing device, characterized in that: include: Acquisition module, used to obtain risk control data; The preprocessing module is used to preprocess the risk control data to form standard risk control data; Anomaly detection module, used to perform multi-dimensional analysis and abnormal behavior detection on standard risk control data to identify abnormal behavior; Detect abnormal behavior on standard risk control data, including: Standard risk control data is tested using statistical, rule-based, and machine learning-based anomaly detection methods to identify abnormal behavior. The standard risk control data is tested according to the statistical anomaly detection method, which is to test the standard risk control data according to the improved mean anomaly detection algorithm, the improved standard difference anomaly detection algorithm, and the improved cluster anomaly detection algorithm in sequence; The improved mean anomaly detection algorithm is used to analyze standard risk control data, including: Use the sliding window method to obtain standard risk control data and calculate the weighted average of the standard risk control data within the sliding window; When the absolute value of the difference between the risk control data and the weighted average value does not meet the first range, the risk control data is determined to be an outlier; wherein, each time the sliding window method is used to determine the weighted average value within the sliding window, the calculation weight of the weighted average value and the size of the first range are dynamically adjusted according to the changing trend of the standard risk control data through an adaptive algorithm; The improved standard deviation detection algorithm is used to analyze standard risk control data, including: Analyze data using kernel density estimation or mixture distribution models to identify abnormal behavior; The improved clustering anomaly detection algorithm is used to analyze standard risk control data, including: Conduct multiple cluster analyses on standard risk control data, using different initial cluster centers each time, to determine the optimal initial cluster center; In the process of analyzing standard risk control data using the optimal initial cluster center, the K value is determined using the elbow method or silhouette coefficient method; Determine abnormal clusters formed by clustering through local density estimation and relative distance metrics to identify abnormal behaviors; The early warning module is used to trigger risk warnings when abnormal behavior is detected.
7. An electronic device, characterized in that: Including the bank risk control data processing device as described in claim 6.
8. A computer-readable storage medium, characterized in that The computer-readable storage medium stores at least one program code, and the program code is executed by a processor to implement the method for processing bank risk control data as described in any one of claims 1 to 5.
Citation Information
Patent Citations
Comprehensive risk management and control method and system based on bank key area multi-dimensional data
CN118297688A