Bank risk control data processing method and device, equipment and storage medium
By pre-processing and multi-dimensional analysis of bank risk control data, combined with multiple anomaly detection methods, the problem of difficulty in identifying abnormal behavior in traditional methods is solved, and the accuracy and satisfaction of risk control data processing is improved.
Patent Information
- Application Number
- CN202510082327.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-20
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2045-01-20
AI Technical Summary
Traditional risk control data processing methods are difficult to identify abnormal behaviors from risk control data of different types and characteristics, and cannot meet the high requirements of banks for risk control data processing.
By acquiring risk control data, pre-processing is performed to form standard data, and performing multi-dimensional analysis and abnormal behavior detection on the standard data, including analysis of time dimensions, user dimensions, business dimensions and data content dimensions, combining anomaly detection methods based on statistics, rules and machine learning.
It improves the detection accuracy of abnormal behavior, can analyze risk control data from multiple angles, identify more abnormal behaviors, and meets the high requirements of banks for risk control data processing.
Smart Images

Figure CN119941372A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of bank risk control data processing, and in particular relates to a bank risk control data processing method and device, equipment and storage medium. Background Art
[0002] With the increasing complexity of banking business and the surge in data volume, data access records and monitoring have become important links in ensuring data security and compliance. Traditional risk control data processing methods have difficulty identifying abnormal behaviors from risk control data of different types and characteristics, and are difficult to meet the bank's needs for risk control data processing. As a result, when processing risk control data, only some abnormal behaviors can be identified, while other abnormal behaviors cannot be identified, which is difficult to meet the bank's high requirements for risk control data processing. Summary of the invention
[0003] The present invention provides a method and device for processing bank risk control data, equipment and storage medium, which can analyze risk control data from multiple angles and more accurately determine abnormal behavior.
[0004] On the one hand, a method for processing bank risk control data is provided, comprising: Obtain risk control data; Pre-process the risk control data to form standard risk control data; Conduct multi-dimensional analysis and abnormal behavior detection on standard risk control data to identify abnormal behaviors; When abnormal behavior is detected, a risk warning is triggered. Optionally, preprocess the risk control data, including: Perform data cleaning, including removing duplicate data in risk control data, filling missing values in risk control data, and removing outliers in risk control data; Perform data conversion, including data type conversion, data standardization and data encoding of risk control data; Conduct data integration, including data merging and data association of risk control data; Perform data verification, including data integrity verification, data consistency verification, and data rationality verification of risk control data.
[0005] Optionally, conduct multi-dimensional analysis of standard risk control data, including: Analyze standard risk control data in the time dimension, user dimension, business dimension and data content dimension respectively; in the time dimension, analyze the number of user visits in the next time period according to the distribution of the number of user visits in different time periods; in the user dimension, build a user behavior model and analyze the user behavior based on the user behavior model; in the business dimension, determine the risk points of the bank's key businesses and conduct data monitoring on the risk points; in the data content dimension, analyze the number of user visits to sensitive content; When the analysis results of the time dimension, user dimension, business dimension and data content dimension do not meet the risk indicator threshold, it is determined to be abnormal behavior.
[0006] Optionally, perform abnormal behavior detection on standard risk control data, including: The standard risk control data is tested in turn according to the statistical-based anomaly detection method, the rule-based anomaly detection method and the machine learning-based anomaly detection method to determine abnormal behavior.
[0007] Optionally, the standard risk control data is detected according to the statistical-based anomaly detection method, which is to detect the standard risk control data according to the improved mean anomaly detection algorithm, the improved standard difference anomaly detection algorithm, and the improved cluster anomaly detection algorithm in sequence; The improved mean anomaly detection algorithm is used to analyze standard risk control data, including: Use the sliding window method to obtain standard risk control data and calculate the weighted average of the standard risk control data within the sliding window; When the absolute value of the difference between the risk control data and the weighted average value does not satisfy the first range, the risk control data is determined to be an abnormal value; wherein, each time the sliding window method is used to determine the weighted average value within the sliding window, the calculation weight of the weighted average value and the size of the first range are adjusted dynamically according to the change trend of the standard risk control data through an adaptive algorithm; The improved standard deviation anomaly detection algorithm is used to analyze standard risk control data, including: Analyze data using kernel density estimation or mixture distribution models to identify abnormal behavior; The improved clustering anomaly detection algorithm is used to analyze standard risk control data, including: Conduct multiple cluster analyses on standard risk control data, using different initial cluster centers each time to determine the optimal initial cluster center; In the process of analyzing the standard risk control data using the optimal initial cluster center, the K value is determined using the elbow method or the silhouette coefficient method; The abnormal clusters formed by clustering are determined by local density estimation and relative distance metrics to identify abnormal behaviors.
[0008] Optionally, abnormal behavior detection is performed on standard risk control data according to an abnormality detection method based on machine learning, including: Increase the normal samples of standard risk control data through data augmentation technology, and increase the abnormal samples of standard risk control data by simulating abnormal generation; Constructing a machine learning model, wherein the machine learning model includes one of a support vector machine, a random forest, and a neural network; Normal samples and abnormal samples are used to train the machine learning model; the learning method includes active learning or semi-supervised learning; regularization methods, ensemble learning methods or hyperparameters of the machine learning model are used in the training process of the machine learning model to prevent the machine learning model from overfitting, and the hyperparameter optimization method includes any one of grid search, random shrinkage and Bayesian optimization.
[0009] Optionally, the bank risk control data processing method further includes: Periodically generate analysis reports based on multi-dimensional analysis results and abnormal behavior detection results in historical time periods.
[0010] On the other hand, a device for processing bank risk control data is provided, comprising: Acquisition module, used to obtain risk control data; The preprocessing module is used to preprocess the risk control data to form standard risk control data; Anomaly detection module, which is used to perform multi-dimensional analysis and abnormal behavior detection on standard risk control data to identify abnormal behavior; The early warning module is used to trigger risk warnings when abnormal behavior is detected.
[0011] On the other hand, an electronic device is provided, which is a device for processing bank risk control data as described above.
[0012] On the other hand, a computer-readable storage medium is provided, in which at least one program code is stored, and the program code is executed by a processor to implement the method for processing bank risk control data as described in any of the above items.
[0013] The beneficial effects brought by the technical solution provided by the present invention are: The present invention provides a method for processing risk control data, which obtains standard risk control data by acquiring risk control data and preprocessing the risk control data. The standard risk control data is subjected to multi-dimensional analysis and abnormal behavior detection. By adopting different methods to analyze the standard risk control data from different angles, the problem that some abnormal behaviors cannot be detected due to the complexity of the types of risk control data and different characteristics can be solved, thereby improving the detection accuracy of abnormal behaviors. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0015] Figure 1 A flow chart of a method for processing bank risk control data provided by the present invention; Figure 2 A structural block diagram of a bank risk control data processing device provided by the present invention; Figure 3 This is a structural block diagram of an electronic device provided by the present invention.
[0016] 11: Acquisition module; 12: Preprocessing module; 13: Anomaly detection module; 14: Early warning module; 21: processor; 22: memory. DETAILED DESCRIPTION
[0017] In order to make the purpose, technical solution and advantages of the present invention clearer, the technical solution of the present invention will be clearly and completely described below in conjunction with the drawings of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0018] Figure 1 This is a flow chart of a method for processing bank risk control data provided by the present invention. Figure 1 ,include: S101. Obtain risk control data.
[0019] In step S101, the risk control data under the business system includes business modules such as deposits, loans, credit cards, transaction information, account status and other data. The risk control data under the channel system includes data from channel systems such as online banking, mobile banking, and self-service terminals. The channel system records the operations performed by customers through different channels, such as login, query, transfer, etc. The risk control data under the external data source includes data provided by credit reporting agencies and anti-fraud service providers to obtain more risk information.
[0020] The collection method of risk control data includes database log collection. For relational databases, the database log function can be used for data collection. The database log records all operations on the database, including insert, update, delete, etc. By parsing the database log, detailed business data changes can be obtained.
[0021] The collection of risk control data also includes system log collection. The various business systems and channel systems of the bank usually generate system logs to record the operating status and operation information of the system. Use log collection tools such as Flume and Logstash to collect and process system logs.
[0022] The collection of risk control data also includes network traffic collection. For some data that cannot be obtained through system logs or database logs, network traffic collection technology can be considered. By deploying traffic collection equipment in the network, network data packets can be captured and the required data can be extracted from them.
[0023] S102: Preprocess the risk control data to form standard risk control data.
[0024] In one example, step S102 includes: Step 1: Perform data cleaning, including removing duplicate data in risk control data, filling missing values in risk control data, and removing outliers in risk control data.
[0025] The collected data may contain duplicate records, which need to be deduplicated to ensure the accuracy and uniqueness of the data. Data deduplication can be performed using hash algorithms, database deduplication functions, etc.
[0026] Identify missing values in the data set and handle them according to the specific situation. For missing values in some important fields, you can fill them by querying other data sources or using default values. For some missing values that cannot be filled, you can consider marking the record as an exception or deleting it, but you need to be careful to avoid affecting the accuracy of data analysis.
[0027] Remove outliers. Detect outliers in the data set, such as abnormally large or small transaction amounts, abnormal transaction times, etc. Statistical methods (such as mean, standard deviation, box plots, etc.) can be used to identify outliers and process them according to business rules, such as deleting outliers, making data corrections, or marking them as abnormal records.
[0028] Step 2: Perform data conversion, including data type conversion, data standardization and data encoding of risk control data.
[0029] Among them, the data formats of different data sources may be different, and format conversion is required to enable unified storage and analysis.
[0030] For example, standardize the date format to a specific standard format, convert string data to numeric data, etc. Standardize the data to make data from different data sources comparable. For example, normalize the transaction amount, standardize the customer rating, etc. Statistical methods (such as Z-score standardization, Min-Max standardization, etc.) can be used to standardize data.
[0031] Step 3: Perform data integration, including data merging and data association of risk control data.
[0032] In this process, data from different data sources are merged to obtain more comprehensive information. For example, basic customer information, transaction records, risk assessment results, etc. can be merged. Data merging can be done using database connection operations (such as inner joins, outer joins, etc.) or data integration tools.
[0033] Data association establishes relationships between data to facilitate multi-dimensional analysis and monitoring. For example, associating a customer's transaction record with the customer's risk rating, associating the time and location of a transaction with a risk event, etc. Data association can be established using database indexes, views, or data mining algorithms.
[0034] Step 4: Perform data verification, including data integrity verification, data consistency verification, and data rationality verification of risk control data.
[0035] Data integrity verification is to check whether the key fields in the data set are complete, such as customer number, transaction time, transaction amount, etc. You can use database constraints (such as non-null constraints, unique constraints, etc.) or data verification tools to perform data integrity verification.
[0036] Data consistency verification is to verify whether the data in the data set is consistent, for example, whether the basic information of customers in different data sources is consistent, whether the sum of transaction records is consistent with the account balance, etc. Data consistency verification can be performed using data comparison tools or data audit tools.
[0037] Data rationality verification is to check whether the data complies with business rules and logic, for example, whether the transaction amount is within a reasonable range, whether the transaction time complies with the business process, etc. You can use a business rule engine or a data rationality check tool to perform data rationality verification.
[0038] Through the implementation of the above data preprocessing, the quality and availability of data can be effectively improved, providing a reliable data foundation for the multi-dimensional data access record and monitoring system of the bank risk control data mart. At the same time, the implementation of data preprocessing also needs to be adjusted and optimized according to specific business needs and data characteristics to ensure the efficiency and accuracy of data processing.
[0039] S103. Perform multi-dimensional analysis and abnormal behavior detection on standard risk control data to determine abnormal behavior.
[0040] In step S103, a data warehouse is first constructed to store data access records from different data sources. The data warehouse should have efficient data storage and query capabilities and be able to support large-scale data processing and analysis.
[0041] Design a reasonable data warehouse architecture, including fact tables and dimension tables. Fact tables are used to store specific data access records, and dimension tables are used to provide different analysis angles, such as time dimension, user dimension, business dimension, etc. Data is extracted from various data sources through data ETL (Extract Transform Load), and then cleaned, transformed, and loaded into the data warehouse. In step S103, a multi-dimensional analysis is performed on the standard risk control data, including: The first step is to analyze the standard risk control data in the time dimension, user dimension, business dimension and data content dimension respectively; in the time dimension, analyze the number of user visits in the next time period according to the distribution of the number of user visits in different time periods; in the user dimension, build a user behavior model and analyze the user behavior based on the user behavior model; in the business dimension, determine the risk points of the bank's key businesses and conduct data monitoring on the risk points; in the data content dimension, analyze the number of users' visits to sensitive content; Step 2: When the analysis results of the time dimension, user dimension, business dimension and data content dimension do not meet the risk indicator threshold, it is determined to be abnormal behavior.
[0042] Exemplarily, performing multi-dimensional analysis includes: Time dimension analysis: Analyze the distribution of data access records at different time points, such as the daily, weekly, and monthly visit trends. Use time series analysis methods to predict future visit trends so that risk prevention measures can be taken in advance. - Detect abnormal time patterns, such as a large number of data accesses during non-working hours or holidays, which may indicate potential risk events. User dimension analysis: Analyze the data access behavior of different user groups. Users can be classified according to factors such as roles, permissions, and geographic locations to study the access habits and risk characteristics of different user groups. Identify abnormal user behaviors, such as frequent access to sensitive data, abnormal access time distribution, etc. Through user profiling technology, establish a user behavior model to detect abnormal users in a timely manner and take corresponding risk control measures. Business dimension analysis: Analyze data access records in different business areas in combination with the bank's business processes and risk points. For example, analyze access to loan business, credit card business, fund transaction business, etc., to identify potential risk links. Monitor changes in business indicators, such as loan approval rate, credit card delinquency rate, etc., and conduct correlation analysis with data access records to identify possible risk factors. Data content dimension analysis: Analyze the specific data content in the data access records. For example, analyze the types of sensitive data accessed, the distribution of data values, etc. Detect abnormal data content, such as large-scale access to specific sensitive data, data values outside the normal range, etc. Through data mining technology, discover potential data leakage risks and abnormal transaction behaviors.
[0043] Risk indicator setting: Based on the results of multi-dimensional analysis, a series of risk indicators are set to monitor the risk level of data access behavior in real time. Risk indicators can include abnormal growth in access volume, abnormal user behavior ratio, and sensitive data access frequency. Determine the threshold of the risk indicator, and trigger a risk warning when the indicator exceeds the threshold. Data analysis technology: You can use a variety of data analysis technologies, such as data mining, machine learning, statistical analysis, etc., to conduct in-depth analysis of data access records. These technologies can help discover potential risk patterns and abnormal behaviors, and improve the accuracy and timeliness of risk warnings. Choose data analysis tools and algorithms that are suitable for bank data characteristics and analysis needs, such as Python, R, SAS, SPSS, etc.
[0044] Data visualization tools: Use data visualization tools to display the results of multi-dimensional analysis in intuitive charts and graphs. Data visualization can help relevant personnel better understand data and risk conditions and improve the efficiency and accuracy of decision-making. Data visualization tools include Tableau, PowerBI, and Echarts. Real-time monitoring technology: In order to achieve real-time risk warning, real-time monitoring technology is needed to monitor data access records in real time. Stream processing technologies such as Apache Kafka and Apache Flink can be used to process and analyze real-time data. A real-time monitoring system is established to ensure that risk events can be discovered and responded to in a timely manner. In short, the multi-dimensional record analysis based on multi-dimensional data access records and monitoring systems designed by the bank's risk control data mart requires the comprehensive use of data storage and management, multi-dimensional analysis methods, risk warning and reporting and other technical means to achieve comprehensive monitoring and risk analysis of data access behaviors. Through the implementation of this analysis layer, it can provide strong support for the bank's risk management and ensure the security and stable operation of the bank's business.
[0045] In step S103, abnormal behavior detection is performed on the standard risk control data, including: The standard risk control data is tested in turn according to the statistical-based anomaly detection method, the rule-based anomaly detection method and the machine learning-based anomaly detection method to determine abnormal behavior.
[0046] In one example, the standard risk control data is detected according to the statistical-based anomaly detection method, which is to detect the standard risk control data in sequence according to the improved mean anomaly detection algorithm, the improved standard difference anomaly detection algorithm, and the improved clustering anomaly detection algorithm.
[0047] Among them, the improved mean anomaly detection algorithm is used to analyze the standard risk control data, including: Use the sliding window method to obtain standard risk control data and calculate the weighted average of the standard risk control data within the sliding window; When the absolute value of the difference between the risk control data and the weighted average value does not satisfy the first range, the risk control data is determined to be an outlier; wherein, each time the sliding window method is used to determine the weighted average value within the sliding window, the calculation weight of the weighted average value and the size of the first range are dynamically adjusted according to the changing trend of the standard risk control data through an adaptive algorithm.
[0048] In one example, the improved mean anomaly detection algorithm is used to analyze standard risk control data, including: First, use robust statistics instead or in combination.
[0049] Median: The median is the value in the middle after sorting the data, and it is not affected by extreme values. The median can be used instead of the mean as an estimate of the central position. For example, when calculating residential housing prices, using the median can avoid the interference of a few luxury housing prices on the central price estimate. When detecting outliers, data points that are too far from the median can be considered abnormal, such as data points that are greater than the median plus three times the median absolute deviation (MAD, MAD is the median of the distance from the data point to the median) are considered abnormal. Weighted mean: Assign a weight to each data point based on the reliability or importance of the data, and then calculate the weighted mean. For example, in sensor network data, more accurate sensor data is given a higher weight, and less accurate sensor data is given a lower weight. This can reduce the impact of inaccurate data on the mean, making the mean more representative of the true data center trend, thereby improving the accuracy of detecting outliers based on the mean. Second, make adjustments based on data distribution characteristics.
[0050] Data preprocessing and distribution fitting: If the data presents an obvious non-normal distribution, the data can be preprocessed first, such as using Box-Cox transformation to convert the data into an approximate normal distribution, and then calculating the mean to detect outliers. At the same time, the data can be fitted with a distribution, such as using a normal distribution, a lognormal distribution, or other suitable distribution model, and a more reasonable range of outliers can be determined based on the fitted distribution parameters and actual data. For example, for some economic data with a positively skewed distribution, after logarithmic transformation to make it close to a normal distribution, the range of outliers is defined based on the mean and standard deviation of the transformed data. Consider quantile information: In addition to the mean, quantiles can also be combined to gain a more comprehensive understanding of the data distribution. For example, calculate the first quartile (Q1) and the third quartile (Q3) of the data, and determine the interquartile range (IQR = Q3 - Q1). Data points that are less than Q1 - 1.5 * IQR or greater than Q3 + 1.5 * IQR can be identified as outliers. This quantile-based method (such as the boxplot method) can also effectively detect anomalies for skewed data, and can be used in conjunction with the mean. When the anomaly detection results based on the mean are inconsistent with the results based on the quantile, further analyze the data to determine whether anomalies really exist. Third, dynamically update the mean and outlier determination criteria.
[0051] Sliding window method: When processing time series or dynamically changing data, the sliding window method is used. For example, for stock price data, a sliding window of fixed length is set, and only the mean of the data in the window is calculated. As the window slides, the mean can be updated in time to reflect the latest data changes. At the same time, the judgment criteria for outliers can be dynamically adjusted according to the fluctuation of the data in the window, such as determining a dynamic outlier range based on the standard deviation of the data in the window. Adaptive algorithm: Use adaptive algorithms to automatically adjust the mean and outlier criteria based on the changing trend of the data. For example, the exponentially weighted moving average (EWMA) algorithm is used, which gives more weight to recent data and can quickly adapt to data changes. When detecting outliers, the data can be judged to be abnormal based on the dynamic mean calculated by EWMA and the standard deviation that is adaptively adjusted according to data fluctuations.
[0052] The improved standard deviation anomaly detection algorithm is used to analyze standard risk control data, including: Analyze data using kernel density estimation or mixture distribution models to identify anomalous behavior.
[0053] In this embodiment, the improved standard deviation abnormal detection algorithm is used to analyze the standard risk control data, including: First, use non-parametric methods or distribution adaptive technology.
[0054] Kernel Density Estimation (KDE): A nonparametric method for estimating the probability density function of data without relying on specific distribution assumptions. Through KDE, data points in low-density areas can be determined as outliers based on the actual distribution of the data, rather than based on the standard deviation under a normal distribution. Mixed distribution model: When the data presents a complex distribution (such as multimodal distribution), a mixed distribution model can be used to fit the data, such as the Gaussian mixture model (GMM). GMM assumes that the data is a mixture of multiple Gaussian distributions. By estimating the parameters of each Gaussian distribution (including mean and standard deviation) and their mixing ratio, the distribution of the data can be described more accurately. When detecting outliers, the probability of each data point belonging to each Gaussian distribution can be used to determine whether it is abnormal, rather than simply relying on a single standard deviation. Second, the local standard deviation and local anomaly factor are used in combination.
[0055] Calculate local standard deviation: In order to reduce the impact of local changes on the overall standard deviation, the local standard deviation can be calculated.
[0056] Combined with Local Outlier Factor (LOF): LOF is a density-based local anomaly detection method. The local standard deviation can be combined with LOF to judge anomalies by comparing the difference in discreteness and density between a data point and its local neighborhood.
[0057] Third, combine semantic information and feature engineering.
[0058] Introducing contextual features: Adding data-related contextual features to the dataset.
[0059] Multivariate analysis and principal component analysis (PCA): If the data contains multiple variables, multivariate analysis can be performed. PCA is a commonly used method that can transform multiple correlated variables into a few uncorrelated principal components. In the new principal component space, the standard deviation or other statistics can be recalculated to determine anomalies.
[0060] The improved clustering anomaly detection algorithm is used to analyze standard risk control data, including: Conduct multiple cluster analyses on standard risk control data, using different initial cluster centers each time to determine the optimal initial cluster center; In the process of analyzing the standard risk control data using the optimal initial cluster center, the K value is determined using the elbow method or the silhouette coefficient method; The abnormal clusters formed by clustering are determined by local density estimation and relative distance metrics to identify abnormal behaviors.
[0061] In this embodiment, an improved clustering anomaly detection algorithm is used to analyze standard risk control data, including: First, optimize clustering algorithm parameters and initialization process Multiple runs and evaluations: For algorithms like K-Means that are sensitive to initial values, you can run the algorithm multiple times, using different initial cluster centers each time. For example, using the K-Means++ initialization method, it can select more appropriate initial centers, or by randomly selecting initial centers multiple times and comparing clustering results, select the optimal clustering scheme to improve the accuracy of outlier detection. Second, automatic parameter selection: Use some methods to automatically determine the parameters of the clustering algorithm. For example, when determining the K value, you can use methods such as the Elbow Method and the Silhouette Coefficient. The Elbow Method observes the curve of the clustering error changing with the K value and finds the "elbow" point of the curve. The K value corresponding to this point is usually a more appropriate number of clusters; the Silhouette Coefficient measures the closeness between each data point and its cluster and adjacent clusters, and selects the appropriate K value by maximizing the Silhouette Coefficient. Third, integrated clustering method: combining the results of multiple different parameters or different clustering algorithms. For example, using K-Means clustering results with multiple different K values, or combining K-Means with density-based clustering (such as DBSCAN) results, to determine outliers through comprehensive judgment. This integrated method can reduce the misjudgment caused by parameter or assumption problems of a single clustering algorithm. Fourth, combine local density and distance information to distinguish abnormal clusters and abnormal points.
[0062] Local density estimation: Calculate the local density around each data point. Data points that are in low-density areas and far away from high-density areas are more likely to be outliers. For example, the local outlier factor (LOF) algorithm is combined with clustering. LOF can measure the local density difference between a data point and its neighbors. After clustering, the LOF value is calculated for the data points in each cluster. Data points with high LOF values may be outliers even within the cluster. For small, low-density clusters, their relationship with other clusters and the distribution of internal data points can be further analyzed to determine whether they are normal small clusters or abnormal clusters. Relative distance metric: In addition to considering the absolute distance from a data point to the cluster center, its relative distance to other data points in the cluster can also be considered. For example, in a cluster, if the distance from a data point to all other data points is significantly greater than the average distance within the cluster, then this data point may be an outlier. Outliers can be detected by calculating the relative distance matrix between data points and combining it with the clustering results. Fifth, adopt efficient clustering algorithms and approximation techniques.
[0063] Sampling-based clustering: For large-scale data sets, you can first perform sampling, cluster the sampled small data sets, and then infer the outliers in the population based on the clustering results and the relationship between the sample and the population. For example, simple random sampling or stratified sampling methods can be used to reduce computational costs while ensuring that the samples are representative. At the same time, incremental clustering algorithms can be used, which can quickly update clustering results when new data is added without re-clustering the entire data set, thereby improving the efficiency of outlier detection. Approximate clustering algorithms: Use some approximate computing clustering algorithms to reduce the consumption of computing resources. For example, locality sensitive hashing (LSH) technology can be used to approximate the distance between data points, thereby accelerating the clustering process. By sacrificing a certain degree of accuracy, clustering and outlier detection of large-scale data sets can be completed within an acceptable time frame.
[0064] In step S103, abnormal behavior detection is performed on the standard risk control data according to an abnormality detection method based on machine learning, including: The first step is to increase the normal samples of standard risk control data through data augmentation technology, and increase the abnormal samples of standard risk control data by simulating abnormal generation; The second step is to build a machine learning model, wherein the machine learning model includes one of a support vector machine, a random forest, and a neural network; The third step is to use normal samples and abnormal samples to train the machine learning model; the learning method includes active learning or semi-supervised learning; in the training process of the machine learning model, regularization methods, ensemble learning methods or optimization of the hyperparameters of the machine learning model are used to prevent the machine learning model from overfitting, and the hyperparameter optimization method includes any one of grid search, random shrinkage and Bayesian optimization.
[0065] In this embodiment, abnormal behavior detection is performed on standard risk control data according to an abnormality detection method based on machine learning, including: First, improve the quality and diversity of training data. Data augmentation technology: For the problem of insufficient data, data augmentation methods can be used. For example, in image anomaly detection, the number of normal samples can be increased by operations such as rotation, flipping, scaling, and adding noise. For abnormal samples, they can be augmented by simulating abnormal generation, such as in financial fraud detection, by modifying certain features of normal transaction data to simulate fraudulent transactions. Active learning and semi-supervised learning: Active learning allows the model to actively select the most valuable samples for labeling, thereby reducing dependence on a large amount of labeled data. Semi-supervised learning uses a large amount of unlabeled data and a small amount of labeled data to train the model. C. Data resampling: Resampling technology is used to address the problem of data distribution bias. For minority classes (abnormal classes), oversampling can be performed, using the SMOTE (Synthetic Minority Over-sampling Technique) algorithm to increase the number of abnormal samples by generating synthetic samples; for majority classes (normal classes), undersampling can be performed to balance the distribution of training data so that the model can better learn abnormal features. Second, optimize the model structure and training process to avoid overfitting and underfitting.
[0066] Regularization methods: Regularization techniques are used during model training, such as L1 and L2 regularization, Dropout, etc. In neural networks, Dropout randomly discards some neurons during training to prevent the model from over-relying on certain features, thereby reducing overfitting.
[0067] Model selection and ensemble learning: Select the appropriate model complexity through methods such as cross-validation. For example, when selecting the depth of a decision tree, use a validation set to evaluate the performance of the model at different depths and select the depth with the best performance. Ensemble learning can combine multiple different models to improve the generalization and robustness of the model. For example, multiple neural networks with different structures or decision trees and SVMs can be combined to judge anomalies through voting or weighted averaging, which can reduce the overfitting or underfitting problems of a single model. Adjust hyperparameters: Reasonably adjust the model's hyperparameters, such as the learning rate, number of layers, number of neurons in each layer, etc. You can use grid search, random search, or more advanced hyperparameter optimization algorithms (such as Bayesian optimization) to find the optimal hyperparameter combination to balance the model's fitting ability and generalization ability. Third, improve model interpretability.
[0068] Explainable artificial intelligence (XAI) methods: Use feature importance analysis methods, such as in decision tree models, you can directly view the importance ranking of features; in neural networks, you can use methods such as LIME (Local Interpretable Model- agnostic Explanations) or SHAP (SHapley Additive exPlanations). These methods can help explain the factors on which the model determines that the data is abnormal. For example, in credit risk assessment anomaly detection, SHAP can explain how the model determines whether the credit risk is abnormal based on factors such as the customer's age, income, and credit record. Build simple and interpretable models to assist in explaining complex models: In addition to using XAI methods, you can also build simple and highly interpretable models (such as linear models and decision trees) to assist in explaining complex machine learning models.
[0069] In another example, abnormal behavior detection is performed on standard risk control data, including: The improved method of detecting outliers based on deep learning is used to detect abnormal behavior of standard risk control data, including the following steps: First, data enhancement and optimization Data augmentation technology: Increase the amount of data by transforming and synthesizing existing data. For example, in image anomaly detection, normal images can be rotated, flipped, scaled, and noise added to generate new samples; for time series data, data can be augmented by sliding windows, resampling, and other methods. In addition, methods such as generative adversarial networks (GANs) can be used to generate virtual data with a distribution similar to that of real data to increase the number of abnormal samples. Active learning and semi-supervised learning: Active learning effectively reduces the reliance on a large amount of labeled data by allowing the model to actively select the most valuable unlabeled samples for labeling. Semi-supervised learning uses a large amount of unlabeled data and a small amount of labeled data to jointly train the model, so that the model can better learn the distribution characteristics of the data and improve the ability to detect anomalies. C. Data resampling: To address the problem of data imbalance, oversampling and undersampling techniques are used. For abnormal samples in the minority class, oversampling algorithms such as SMOTE can be used to generate synthetic samples; for normal samples in the majority class, undersampling can be performed to balance the ratio of normal samples and abnormal samples in the training data, so that the model can learn the characteristics of the two types of samples more evenly. Second, model improvement and regularization.
[0070] Regularization method: Regularization items are added during the training process, such as L1 and L2 regularization, Dropout, etc., to limit the complexity of the model and prevent overfitting. For example, Dropout is used in a neural network to randomly discard some neurons so that the model does not rely too much on certain features, thereby improving generalization ability.
[0071] Model selection and ensemble learning: Select models of appropriate complexity through methods such as cross-validation, and combine multiple different models through ensemble learning. For example, integrate multiple neural networks with different structures or different types of anomaly detection models (such as autoencoders and generative adversarial networks), and use voting, weighted averaging and other strategies to comprehensively judge anomalies, thereby improving the robustness and generalization ability of the model.
[0072] Adjust hyperparameters: Reasonably adjust the model's hyperparameters, such as learning rate, number of layers, number of neurons in each layer, convolution kernel size, etc. You can use hyperparameter optimization algorithms such as grid search, random search, Bayesian optimization, etc. to find the optimal hyperparameter combination to balance the model's fitting ability and generalization ability. Third, computing efficiency is improved.
[0073] Model compression and quantization: Use model compression techniques, such as pruning to remove unimportant connections or parameters, and quantizing to represent parameters as low-precision data types, to reduce the storage space and computational complexity of the model while maintaining the performance of the model as much as possible. Hardware acceleration: Use more powerful computing hardware, such as GPU, FPGA, ASIC and other dedicated chips to accelerate the training and reasoning process of the model. In addition, technologies such as distributed training and parallel computing can also be used to improve computing efficiency and shorten training time. Improvement of optimization algorithms: Choose more efficient optimization algorithms, such as Adagrad, Adadelta, RMSProp, Adam and other adaptive learning rate optimization algorithms, which can automatically adjust the learning rate according to the characteristics of the problem, speed up the convergence of the model and reduce the training time. Fourth, the model’s interpretability is enhanced.
[0074] Explainable artificial intelligence (XAI) methods: Use feature importance analysis methods, such as directly viewing feature importance rankings in decision tree models; in neural networks, methods such as LIME and SHAP can be used to explain the factors on which the model judges data as abnormal, helping people understand the model's decision-making process. Build simple and interpretable models to assist in explaining complex models: In addition to using XAI methods, you can also build simple and highly interpretable models (such as linear models and decision trees) to assist in explaining complex deep learning models. By comparing the decision results and explanations of the two models, you can improve your understanding of anomaly detection results.
[0075] S104. When abnormal behavior is detected, a risk warning is triggered. Risk warning mechanism: Establish a real-time risk warning mechanism. When the risk index reaches the warning threshold, send warning information to relevant personnel in a timely manner. Warning information can be sent via email, SMS, system notification, etc. Classify and prioritize warning information so that relevant personnel can quickly respond to high-risk events. S105. Periodically generate an analysis report based on the multi-dimensional analysis results and abnormal behavior detection results in the historical time period.
[0076] Risk report generation: Generate risk reports regularly to summarize the analysis results and risk status of data access records. Risk reports should include charts and data of multi-dimensional analysis, changing trends of risk indicators, detailed descriptions of risk events, etc. Risk reports can provide decision support for the bank's management and help them formulate effective risk control strategies. Among them, the multi-dimensional data access record and monitoring system report generation layer of the bank risk control data mart is implemented as follows: The report includes: First, the overview part.
[0077] System operation overview: briefly describe the overall operation of the data access logging and monitoring system in a specific time period, including key indicators such as system availability and data collection integrity. For example, you can mention that the data collection success rate reached [X]% and the system maintained stable operation for [specific duration]. Risk Assessment Summary: Give a general description of the risk situation faced by the bank, including the current risk level (such as low, medium, high) and the main risk trend (whether it is rising, falling or stable). This part can be based on the preliminary analysis of multi-dimensional data to allow readers to quickly understand the general outline of the bank's risks. Second, multi-dimensional data analysis report section. Time dimension analysis report: presents data access patterns at different time scales. For example, display the trend chart of access volume changes by day, week, and month, analyze the time points of peak and trough occurrence and their relationship with the operating rules of banking business. At the same time, report data access anomalies at special time points (such as holidays, business peaks, etc.), such as a sudden increase of [X]% in data access volume during a holiday, and analyze the risks that may be caused by this change. User dimension analysis report: Classify and analyze data access behaviors based on factors such as user type (such as individual customers, corporate customers, bank employees, etc.), user role (such as teller, credit manager, etc.), and user geographic location. Detailed description of the access frequency of different user groups, the main types of data accessed, and abnormal access behaviors. For example, the report points out that the frequency of corporate customers in a certain region accessing loan approval data has abnormally increased by [X] times in a specific time period. Business dimension analysis report: Analyze the relationship between data access and business operations for each major business line of the bank (such as savings, loans, credit cards, wealth management, etc.). Display the data access volume, key data points accessed, and possible risk links in each business area. For example, in the loan business, it is found that the number of visits to high-risk loan customer information has increased recently, which may indicate potential credit risk issues. Data content dimension analysis report: In-depth analysis of the specific content of accessed data, including access to sensitive data. List the types of sensitive data that are frequently accessed (such as customer ID numbers, account passwords, etc.) and their access sources, access time, etc. Report data content anomalies, such as large-scale centralized access to a certain type of highly confidential customer information. Third, the anomaly detection report section.
[0078] Abnormal event statistics: summarizes the number of all abnormal events detected during the monitoring period, and classifies and counts them by abnormal type (such as abnormal access frequency, abnormal access path, abnormal data download, etc.). For example, the report shows that in the past month, a total of [X] abnormal access frequency events and [Y] abnormal access path events were detected. Detailed analysis of major abnormal events: Conduct in-depth analysis of major abnormal events with a high risk level, including the time of occurrence, users involved, data content accessed, possible loss assessment, and preliminary cause analysis. For example, describe in detail an event involving abnormal access to a large amount of customer funds information, and analyze whether it may be caused by an external attack or internal illegal operations. Fourth, risk warning and response report section.
[0079] Risk warning report: Lists the indicators and thresholds that trigger risk warnings, and reports the number of warnings triggered and the accuracy of warnings during the monitoring period. For example, when the data access volume exceeds [X] times within an hour, the warning is triggered. In the past week, the warning was triggered [Y] times, and the proportion of accurate warnings was [Z]%. Implementation report of countermeasures: describes the countermeasures taken for different risk warnings (such as restricting user access, strengthening security verification, launching investigation processes, etc.) and the effectiveness evaluation of these measures. For example, the report shows that after restricting the access rights of a suspicious user, the related abnormal access behavior has been effectively curbed. The report generation process is as follows: First, data integration and processing. Data extraction: Extract various data related to the report content from the data warehouse, including multi-dimensional data access records after preprocessing and analysis, anomaly detection results, risk warning information, etc. Use data query languages (such as SQL) or specialized data extraction tools to obtain the required data from the storage system according to predetermined rules and conditions. Data association and integration: Associating and integrating data extracted from different data sources to ensure the consistency and integrity of the report content. For example, associating user dimension data with business dimension data through user IDs so that the report can fully display the user's behavior in different businesses. Use data integration tools or write custom association algorithms to achieve effective data integration. Second, the report generation techniques and tools are as follows: Template engine: Use template engines (such as Jinja2, Velocity, etc.) to define the template structure of the report. The template defines the various parts of the report, the format (such as title, paragraph, chart, table, etc.) and the location of data filling. By filling the extracted and processed data into the template, you can quickly generate a uniformly formatted and beautiful report. Data analysis and visualization tools: Combine data analysis tools (such as Python's Pandas, NumPy, etc.) and visualization tools (such as Matplotlib, Seaborn, Tableau, etc.) to generate charts and data visualization content in reports. These tools can present complex data in intuitive graphics (such as bar charts, line charts, pie charts, heat maps, etc.) and tables to facilitate readers' understanding and analysis. For example, use a bar chart to show the comparison of data access volume of different business lines, and use a line chart to show the access trend changes in the time dimension. Text generation technology: Use natural language generation technology to automatically generate the text description part of the report based on data and analysis results. You can use rule-based text generation methods or machine learning-based text generation models (such as localized applications such as GPT) to generate smooth and accurate text content. At the same time, the generated text is manually reviewed and adjusted to ensure the professionalism and logic of the language expression. Third, report distribution and storage. Distribution channels: Determine the distribution channels of the report, including sending the report to the bank's internal risk management department, senior management, business departments and other relevant personnel. The report can be distributed through email, internal office systems, etc. to ensure that relevant personnel can obtain the report information in a timely manner. For some risk situations that require real-time attention, a real-time push notification function can also be set up within the system. Storage management: Store the generated reports in a dedicated document management system or data repository for subsequent query, audit, and historical comparative analysis. Categorize and manage report storage, and establish a reasonable storage directory structure based on report type (such as daily report, weekly report, special report, etc.), time range, and other factors. At the same time, ensure the security of report storage, set appropriate access rights, and prevent information leakage.
[0080] In this embodiment, a user interface may be further provided. The user interface layer is implemented as follows: First, interface design goals.
[0081] Intuitive: The user interface should be intuitive and easy to understand, so that users from different backgrounds (including risk managers, technical personnel and senior management) can quickly understand and operate it. Through clear layout, clear icons and concise labels, the learning cost of users can be reduced.
[0082] Real-time: Provide real-time data display and updates to ensure that users can understand data access and risk status in a timely manner. The interface should be able to respond quickly to data changes so that users can make decisions quickly.
[0083] Interactivity: Supports interactive operations between users and the system, such as querying specific data, setting parameters, generating reports, etc. The interactive method should be simple and easy to use to improve user work efficiency.
[0084] Visualization: Use visualization technology to display complex data in the form of intuitive charts, graphs and maps to help users better understand data and risk trends.
[0085] Security: Ensure the security of the user interface to prevent unauthorized access and operation. Use strict user authentication and authorization mechanisms to protect sensitive data and system functions.
[0086] Second, interface layout and functional modules The data overview module includes a real-time data dashboard and a time trend chart.
[0087] Real-time data dashboard: Displays real-time data of key indicators, such as data access volume, number of abnormal events, risk level, etc. Using intuitive charts and numbers, users can understand the overall status of the system at a glance.
[0088] Time trend chart: Displays the change trend of data access volume and risk indicators over time. Users can select different time ranges to view in order to analyze long-term and short-term risk trends.
[0089] The multi-dimensional analysis module includes: time dimension analysis, user dimension analysis, business dimension analysis, and data content dimension analysis.
[0090] Time dimension analysis: Provides the function of data access analysis by time dimension. Users can select a specific time period to view the data access pattern, peak and trough times, and abnormal event distribution within that time period.
[0091] User dimension analysis: allows users to analyze data access based on factors such as user type, role, and geographic location. Displays access behavior characteristics and anomalies of different user groups.
[0092] Business dimension analysis: Displays the relationship between data access and business operations for each business area of the bank. Users can gain in-depth understanding of the risk status and data access patterns of specific businesses.
[0093] Data content dimension analysis: Analyze the specific content of accessed data, including access to sensitive data. Users can view the types of sensitive data that are frequently accessed and the sources of access.
[0094] The anomaly detection module includes anomaly event list, anomaly event analysis, Abnormal event list: Lists all abnormal events detected by the system, including detailed information such as the time, type, users involved, and data content of the event. Users can classify, filter, and sort abnormal events to quickly locate and handle high-risk events.
[0095] Abnormal event analysis: Provides in-depth analysis of specific abnormal events. Users can view the details of the event, related data access records, and possible cause analysis. At the same time, the system should provide recommended response measures to help users deal with abnormal events in a timely manner.
[0096] The risk warning module includes warning indicator settings and warning notifications.
[0097] Warning indicator settings: Allow users to set risk warning indicators and thresholds. Users can customize warning rules based on business needs and risk preferences to ensure that the system can detect potential risks in a timely manner.
[0098] Warning notification: When the system detects an event that meets the warning conditions, it will send a notification to the user in a timely manner. The notification method can include email, SMS and system pop-up window, etc., to ensure that users can respond to risk events quickly.
[0099] The report generation module includes custom reports and report export.
[0100] Customized reports: Users can select the content and format of the report according to their needs to generate personalized risk reports. The report can include data overview, multi-dimensional analysis, anomaly detection and risk warning, etc., to provide users with comprehensive risk assessment and decision support.
[0101] Report export: Supports exporting generated reports to PDF, Excel and other formats, making it convenient for users to share and archive with other personnel.
[0102] The user management module includes user authentication and authorization and user rights management.
[0103] User authentication and authorization: Ensure that only authorized users can access the system. Use strict user authentication mechanisms, such as user name and password, digital certificates, etc., to protect the security of the system.
[0104] User rights management: Different rights are assigned to users according to their roles and responsibilities. For example, risk managers can view and process all risk data, while ordinary users can only view some data and reports.
[0105] Third, application of visualization technology.
[0106] Charts and graphs: Use common chart types such as bar charts, line charts, and pie charts to display the distribution and changing trends of indicators such as data access volume, number of abnormal events, and risk level. Use visualization techniques such as heat maps and scatter plots to analyze the hotspots of data access and the distribution of abnormal events. Use interactive charts to allow users to obtain more data information by hovering the mouse, clicking, and other operations.
[0107] Map visualization: For data analysis involving geographic location, map visualization technology can be used to display the distribution of users' visit locations and the geographic distribution of risk events. It supports map zooming, panning, and layer switching functions to facilitate users to query and analyze geographic information.
[0108] Data dashboard: Design a data dashboard that displays key indicators and charts on one page, so that users can quickly understand the overall status of the system. The data dashboard should support custom layout and indicator selection to meet the needs of different users.
[0109] Fourth, interactive design.
[0110] Menu and navigation: Design a concise and clear menu and navigation structure to help users quickly find the required functional modules. Use drop-down menus, sidebars, etc. to organize menus to improve the usability of the interface.
[0111] Search and filter: Provides search function, allowing users to quickly find specific data records and abnormal events. Supports filtering function, users can filter and sort data according to different conditions to better analyze and handle risk events.
[0112] Operation prompts and feedback: Provide clear operation prompts and feedback information when users are operating to ensure that users know whether their operations are successful. For incorrect operations, provide clear error prompts and provide solution suggestions.
[0113] Fifth, safety considerations User authentication and authorization: Adopt strict user authentication and authorization mechanisms to ensure that only authorized users can access the system. Update user passwords regularly and strengthen password strength requirements to prevent password leaks.
[0114] Data encryption: Encrypt sensitive data for storage and transmission to protect user privacy and data security. Use encryption protocols such as SSL / TLS to ensure data security during network transmission.
[0115] Access control: Strictly control the access rights of different users to prevent unauthorized users from accessing sensitive data and system functions. Record user operation logs for auditing and tracing.
[0116] Through the implementation of the above user interface layer, an intuitive, efficient and secure user interaction environment can be provided for the multi-dimensional data access recording and monitoring system of the bank risk control data mart, helping users to better manage and control risks.
[0117] Finally, the beneficial effects of this application include: First, improve data security: through multi-dimensional recording and anomaly detection, timely discover and prevent potential security threats. Second, enhance compliance: regularly conduct compliance assessments to ensure that data access behavior complies with laws, regulations and industry standards. Third, improve recording efficiency: automate recording and report generation, reduce the workload of administrators, and improve recording efficiency. Fourth, enhance traceability: record user access behavior in detail to provide strong evidence for security incident investigations.
[0118] Figure 2 This is a structural block diagram of a bank risk control data processing device provided by the present invention. Figure 2 ,include: An acquisition module 11 is used to acquire risk control data; A preprocessing module 12 is used to preprocess the risk control data to form standard risk control data; Anomaly detection module 13, used to perform multi-dimensional analysis and abnormal behavior detection on standard risk control data to determine abnormal behavior; The early warning module 14 is used to trigger a risk early warning when abnormal behavior is detected.
[0119] Figure 3 This is a structural block diagram of an electronic device provided by the present invention. Figure 3 , electronic devices may include Figure 2 The bank risk control data processing device generally includes a processor 21 and a memory 22 .
[0120] The processor 21 may include one or more processing cores, such as a 4-core processor, an 8-core processor, etc. The processor 21 may be implemented in at least one hardware form of DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), and PLA (Programmable Logic Array). The processor 21 may also include a main processor and a coprocessor. The main processor is a processor for processing data in the awake state, also known as a CPU (Central Processing Unit); the coprocessor is a low-power processor for processing data in the standby state. The memory 22 may include one or more computer-readable storage media, which may be non-transitory. The memory 22 may also include a high-speed random access memory, and a non-volatile memory, such as one or more disk storage devices and flash memory storage devices. In some embodiments, the non-transitory computer-readable storage medium in the memory 22 is used to store at least one instruction, which is used to be executed by the processor 21 to implement the bank risk control data processing method performed by an electronic device provided in the method embodiment of the present application.
[0121] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for processing bank risk control data, characterized in that: include: Obtain risk control data; Pre-process the risk control data to form standard risk control data; Conduct multi-dimensional analysis and abnormal behavior detection on standard risk control data to identify abnormal behaviors; When abnormal behavior is detected, a risk warning is triggered.
2. The method for processing bank risk control data according to claim 1, characterized in that: Preprocess the risk control data, including: Perform data cleaning, including removing duplicate data in risk control data, filling missing values in risk control data, and removing outliers in risk control data; Perform data conversion, including data type conversion, data standardization and data encoding of risk control data; Conduct data integration, including data merging and data association of risk control data; Perform data verification, including data integrity verification, data consistency verification, and data rationality verification of risk control data.
3. The method for processing bank risk control data according to claim 1, characterized in that: Conduct multi-dimensional analysis of standard risk control data, including: Analyze standard risk control data in the time dimension, user dimension, business dimension and data content dimension respectively; in the time dimension, analyze the number of user visits in the next time period according to the distribution of the number of user visits in different time periods; in the user dimension, build a user behavior model and analyze the user behavior based on the user behavior model; in the business dimension, determine the risk points of the bank's key businesses and conduct data monitoring on the risk points; in the data content dimension, analyze the number of user visits to sensitive content; When the analysis results of the time dimension, user dimension, business dimension and data content dimension do not meet the risk indicator threshold, it is determined to be abnormal behavior.
4. The method for processing bank risk control data according to claim 1, characterized in that: Detect abnormal behavior of standard risk control data, including: The standard risk control data is tested in turn according to the statistical-based anomaly detection method, the rule-based anomaly detection method and the machine learning-based anomaly detection method to determine abnormal behavior.
5. The method for processing bank risk control data according to claim 4, characterized in that: The standard risk control data is detected according to the statistical anomaly detection method, that is, the standard risk control data is detected according to the improved mean anomaly detection algorithm, the improved standard difference anomaly detection algorithm, and the improved cluster anomaly detection algorithm in sequence; The improved mean anomaly detection algorithm is used to analyze standard risk control data, including: Use the sliding window method to obtain standard risk control data and calculate the weighted average of the standard risk control data within the sliding window; When the absolute value of the difference between the risk control data and the weighted average value does not satisfy the first range, the risk control data is determined to be an abnormal value; wherein, each time the sliding window method is used to determine the weighted average value within the sliding window, the calculation weight of the weighted average value and the size of the first range are adjusted dynamically according to the change trend of the standard risk control data through an adaptive algorithm; The improved standard deviation anomaly detection algorithm is used to analyze standard risk control data, including: Analyze data using kernel density estimation or mixture distribution models to identify abnormal behavior; The improved clustering anomaly detection algorithm is used to analyze standard risk control data, including: Conduct multiple cluster analyses on standard risk control data, using different initial cluster centers each time to determine the optimal initial cluster center; In the process of analyzing the standard risk control data using the optimal initial cluster center, the K value is determined using the elbow method or the silhouette coefficient method; The abnormal clusters formed by clustering are determined by local density estimation and relative distance metrics to identify abnormal behaviors.
6. The method for processing bank risk control data according to claim 4, characterized in that: Abnormal behavior detection is performed on standard risk control data based on anomaly detection methods based on machine learning, including: Increase the normal samples of standard risk control data through data augmentation technology, and increase the abnormal samples of standard risk control data by simulating abnormal generation; Constructing a machine learning model, wherein the machine learning model includes one of a support vector machine, a random forest, and a neural network; Normal samples and abnormal samples are used to train the machine learning model; the learning method includes active learning or semi-supervised learning; regularization methods, ensemble learning methods or hyperparameters of the machine learning model are used in the training process of the machine learning model to prevent the machine learning model from overfitting, and the hyperparameter optimization method includes any one of grid search, random shrinkage and Bayesian optimization.
7. The method for processing bank risk control data according to any one of claims 1 to 6, characterized in that: The bank risk control data processing method further includes: Periodically generate analysis reports based on multi-dimensional analysis results and abnormal behavior detection results in historical time periods.
8. A bank risk control data processing device, characterized in that: include: Acquisition module, used to obtain risk control data; The preprocessing module is used to preprocess the risk control data to form standard risk control data; Anomaly detection module, which is used to perform multi-dimensional analysis and abnormal behavior detection on standard risk control data to identify abnormal behavior; The early warning module is used to trigger risk warnings when abnormal behavior is detected.
9. An electronic device, characterized in that: The bank risk control data processing device as claimed in claim 8.
10. A computer-readable storage medium, characterized in that: At least one program code is stored in the computer-readable storage medium, and the program code is executed by a processor to implement the method for processing bank risk control data as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Risk management method and device, equipment and storage medium
CN115099680A
Comprehensive risk management and control method and system based on bank key area multi-dimensional data
CN118297688A
Enterprise network sales data early warning management system based on artificial intelligence
CN118840140A
Method for batch construction of risk time sequence features and quantitative information risk adjustment
CN118917935A
Methods and systems for risk mining and for generating entity risk profiles
US20120221485A1
Cited By
Multi-source risk control data cleaning processing method
CN120950837A