Financial risk analysis method based on big data

Through data cleaning, normalization and streaming data processing technology, combined with principal component analysis and random forest algorithm, the analysis problems of multi-source heterogeneity and dynamic nature of financial data are solved, and efficient, accurate and real-time prediction of financial risk analysis is achieved.

CN120147005APending Publication Date: 2025-06-13BAODING RUIXUN INFORMATION TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510204240.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-24
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

In financial risk analysis, the multi-source heterogeneous characteristics and dynamic nature of financial data have led to huge challenges in data integration and real-time analysis. It is difficult for existing technologies to effectively process large-scale real-time data and adjust the model in a timely manner.

Method used

Data cleaning and normalization technology are used to process multi-source heterogeneous data, and the streaming data processing framework is used to collect market conditions and policy changes in real time, and data distribution characteristics are dynamically updated through adaptive algorithms, and credit risk assessment models are constructed using principal component analysis and random forest algorithms.

Benefits of technology

Effectively process the multi-source heterogeneous, real-time dynamic and high-dimensional complex characteristics of financial data, improve the accuracy and timeliness of risk prediction, and reduce the resource consumption and time of model updates.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120147005A_ABST
    Figure CN120147005A_ABST
Patent Text Reader

Abstract

The invention discloses a financial risk analysis method based on big data, and belongs to the field of financial risk analysis, and the method comprises the steps: converting the data of different data sources into a unified standard through normalization processing according to the difference between the data format and the time granularity, and obtaining a standardized data set; in a real-time data stream, the change trend of data distribution is judged through an adaptive algorithm, key features are extracted from multi-dimensional features by adopting a principal component analysis technology aiming at the complexity of high-dimensional data, and the data dimension is reduced; in the sparse features, the importance of the features is judged through a feature selection algorithm, and if the occurrence frequency of the features is lower than a preset threshold value, invalid features are removed; according to the extracted key features and the filtered sparse features, a random forest algorithm is adopted to construct a credit risk assessment model, and the model is trained to predict the risk probability; and according to a final training result, determining an optimal model parameter and feature combination, and generating a risk prediction report.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of financial risk analysis, and particularly relates to a financial risk analysis method based on big data. Background Art

[0002] In financial risk analysis, technical methods based on big data need to deeply mine massive financial data to reveal potential risk points and market rules. However, in actual business scenarios, the multi-source heterogeneous characteristics of financial data pose great challenges to data integration and analysis. For example, bank transaction data, securities market data, and macroeconomic data come from different systems, and there are significant differences in data formats, time granularities, and data quality standards. This heterogeneity leads to a large amount of time and resources being consumed in the data preprocessing stage. Especially in the data cleaning and normalization processes, it is difficult to completely eliminate the influence of data noise and missing values, thereby affecting the accuracy of subsequent analysis.

[0003] In addition, financial data has a high degree of dynamism and timeliness. Market conditions, policy changes, and emergencies will quickly affect the distribution characteristics of financial data. Traditional static analysis methods are difficult to capture such dynamic changes. For example, when predicting market fluctuations, the laws of historical data may fail due to external factors, resulting in the prediction model being unable to adjust in time, thus causing lags or deviations. This dynamism requires analysis methods to have real-time and self-adaptive capabilities. However, existing technologies often face problems such as insufficient computing resources and low model update efficiency when dealing with large-scale real-time data. Summary of the Invention

[0004] To solve the above technical problems, the present invention provides a financial risk analysis method based on big data, including:

[0005] Obtain bank transaction data, securities market data, and macroeconomic data. For the multi-source heterogeneous characteristics of different data sources, use data cleaning technology to remove noise and missing values to obtain data from different data sources;

[0006] According to the differences in data formats and time granularities, convert the data from different data sources into a unified standard through normalization processing to obtain a standardized data set;

[0007] For the dynamism and timeliness of financial data, adopt a stream data processing framework to collect market conditions and rule change data in real time and dynamically update the data distribution characteristics;

[0008] In the real-time data stream, use an adaptive algorithm to judge the change trend of the data distribution. If a significant change is detected, automatically adjust the model parameters;

[0009] Adopt principal component analysis technology to extract key features from multi-dimensional features and reduce the data dimension;

[0010] In sparse features, the importance of features is judged by a feature selection algorithm. If the occurrence frequency of a feature is lower than a preset threshold, the invalid feature is removed;

[0011] According to the extracted key features and the filtered sparse features, a credit risk assessment model is constructed using the random forest algorithm, and the model is trained to predict the risk probability;

[0012] Real-time data is obtained and input into the credit risk assessment model to generate a risk prediction report.

[0013] Preferably, the process of obtaining data from different data sources includes:

[0014] Bank transaction data, securities market data, and macroeconomic data are obtained. For the multi-source heterogeneous characteristics of different data sources, a denoising method is used to remove noise from trading volumes, market values, and economic values;

[0015] A filling method is used to fill in missing values;

[0016] The cleaned data is classified and stored through a storage method;

[0017] An analysis method is used to extract features and perform pattern recognition on the processed data;

[0018] If the data cleaning effect does not meet the preset threshold, the denoising and filling steps are re-executed to generate the data from different data sources.

[0019] Preferably, the process of obtaining the standardized data set includes:

[0020] For the differences in data format and time granularity, the data from different data sources is normalized to obtain a normalized data set;

[0021] Feature values are extracted and pattern values are recognized from the normalized data set. According to the results of pattern value recognition, a clustering algorithm is used to classify the data to obtain the standardized data set.

[0022] Preferably, the process of dynamically updating the data distribution characteristics includes:

[0023] A stream data processing framework is used to obtain real-time data streams from market quotation data sources and policy change data sources. If there are abnormal fluctuations in the data streams, an anomaly detection algorithm is used to identify and filter abnormal data;

[0024] For the dynamics and timeliness in the data streams, the data distribution characteristics are adjusted through a dynamic update mechanism;

[0025] According to the data distribution characteristics after dynamic adjustment, a feature extraction method is used to obtain key feature values. If the key feature values do not meet the preset threshold, the anomaly detection and dynamic update steps are re-executed.

[0026] Preferably, the process of using principal component analysis technology to extract key features from multi-dimensional features and reduce the data dimension includes:

[0027] Use principal component analysis technology to extract key features from high-dimensional data and reduce the data dimension;

[0028] According to the data distribution characteristics after dimensionality reduction, a clustering algorithm is used to classify the key features;

[0029] If there are abnormal categories in the clustering results, an anomaly detection algorithm is used to filter the abnormal categories, and according to the data distribution characteristics after filtering, key feature values are extracted.

[0030] Preferably, the process of using the random forest algorithm to construct a credit risk assessment model based on the extracted key features and filtered sparse features and training the model to predict the risk probability includes:

[0031] According to the data distribution of the sparse features and key features, a random forest algorithm is used to construct a credit risk assessment model;

[0032] Train the model to predict the risk probability. If the risk probability is lower than the preset threshold, it is marked as a low-risk category;

[0033] For high-risk category data, an anomaly detection algorithm is used to filter abnormal samples;

[0034] According to the data distribution after filtering, key feature values are re-extracted;

[0035] Use the random forest algorithm to rank the importance of the key feature values and eliminate low-importance features;

[0036] According to the importance ranking results, update the credit risk assessment model.

[0037] Preferably, the process of updating the credit risk assessment model includes:

[0038] Use cross-validation technology to evaluate the performance of the trained model and obtain the prediction accuracy of the model on the validation set;

[0039] According to the comparison result between the prediction accuracy and the preset standard, judge whether the model performance meets the requirements;

[0040] If the prediction accuracy is lower than the preset standard, enter the feature extraction and parameter adjustment link;

[0041] For the feature extraction step, the random forest algorithm is used to rank the importance of features for the original data, and key features with importance scores higher than a preset threshold are obtained;

[0042] For the parameter adjustment step, the grid search technique is used to search within a preset parameter range to obtain the optimal parameter combination of the model;

[0043] Based on the extracted key features and the optimized parameter combination, the credit risk assessment model is retrained to obtain an updated model;

[0044] The updated model is used for cross-validation to obtain a new prediction accuracy and determine whether it meets the preset criteria;

[0045] If the prediction accuracy meets the preset criteria, the model is determined as the final credit risk assessment model.

[0046] On the other hand, the present invention also provides an electronic device, including a memory, a processor, and a computing program stored in the memory and executable on the processor, and the method is implemented when the processor executes the computer program.

[0047] On the other hand, the present invention also provides a computer-readable storage medium, and the computer-readable storage medium stores a computer program, and the method is implemented when the computer program is executed by a processor.

[0048] Compared with the prior art, the present invention has the following advantages and technical effects:

[0049] The present invention discloses a financial risk analysis method based on big data. This method obtains multi-source heterogeneous bank transaction, securities market, and macroeconomic data, and uses data cleaning and normalization techniques to process noise, missing values, and format differences. In view of the dynamics of financial data, the present invention uses a stream data processing framework to collect market quotes and policy change data in real time, and dynamically updates the data distribution characteristics through an adaptive algorithm. To cope with the complexity of high-dimensional data, the present invention uses principal component analysis and feature selection algorithms to extract key features and reduce the data dimension. Finally, based on the optimized feature set, the present invention uses the random forest algorithm to construct a credit risk assessment model, and evaluates and optimizes the model performance through cross-validation techniques. This method can effectively handle the characteristics of multi-source heterogeneity, real-time dynamics, and high-dimensional complexity of financial data, and improve the accuracy and timeliness of risk prediction. BRIEF DESCRIPTION OF THE DRAWINGS

[0050] The drawings constituting a part of this application are used to provide a further understanding of this application. The schematic embodiments of this application and their descriptions are used to explain this application and do not constitute an improper limitation to this application. In the drawings:

[0051] Figure 1Schematic flowchart of the method according to the embodiments of the present invention. Detailed implementation manners

[0052] It should be noted that, without conflict, the embodiments in the present application and the features in the embodiments may be combined with each other. The present application will be described in detail below with reference to the accompanying drawings and in conjunction with the embodiments.

[0053] It should be noted that the steps shown in the flowchart of the accompanying drawings may be executed in a computer system such as a set of computer-executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than here.

[0054] Embodiment 1

[0055] As Figure 1 shown, a financial risk analysis method based on big data is provided in this embodiment, including:

[0056] S101. Obtain bank transaction data, securities market data, and macroeconomic data, and for the multi-source heterogeneous characteristics of different data sources, adopt data cleaning technology to remove noise and missing values.

[0057] Obtain bank transaction data, securities market data, and macroeconomic data, and for the multi-source heterogeneous characteristics of different data sources, adopt data cleaning technology to remove noise and missing values. Determine the preprocessing method for data cleaning according to the heterogeneity of the data sources. Adopt a denoising method to perform noise removal processing on trading volume, market value, and economic value. Adopt a filling method to fill in the missing values. Classify and store the cleaned data through a storage method. Adopt an analysis method to perform feature extraction and pattern recognition on the processed data. If the data cleaning effect does not meet the preset threshold, re-execute the denoising and filling steps.

[0058] Specifically, bank transaction data involves multiple fields such as savings account transactions, loan account transactions, and credit card consumption. Taking credit card consumption data as an example, the original data may contain problems such as abnormal consumption amount, missing timestamps, and incomplete merchant information. For abnormal consumption amount, the mean method or median method can be used for denoising. For example, if a user spends 10,000 yuan per month and suddenly spends 100,000 yuan, it is necessary to determine whether it is an abnormal value and handle it. For missing timestamps, it can be filled in based on the transaction serial number and adjacent records. Securities market data includes information such as stock prices, trading volume, and market value. Taking stock daily trading data as an example, trading volume data may have noise such as gaps and abnormal volume. The moving average method is used to smooth the trading volume. For example, the five-day average volume line can effectively filter short-term abnormal fluctuations. For price missing caused by suspension, the industry sector average can be used to fill. Market value data needs to be corrected by considering the impact of corporate behaviors such as ex-rights, ex-dividends, and rights issues. Macroeconomic data includes indicators such as gross national product, consumer price index, and employment rate. Taking the consumer price index as an example, the original data may fluctuate due to seasonal factors, emergencies, etc. The seasonal adjustment method is used to eliminate periodic noise. For example, if the price index of a certain month soars due to holidays, seasonal factors need to be eliminated. For missing data in some months, it can be estimated by time series interpolation. The storage of data after cleaning needs to consider the characteristics of the data. The transaction data volume is large and the real-time requirements are high, so it is suitable to use distributed database storage. Securities market data needs to be queried quickly, and column storage can be used to improve query efficiency. Macroeconomic data is updated less frequently, and traditional relational database storage can be used. Feature extraction focuses on the statistical characteristics of the data. For transaction data, extract features such as user consumption habits and credit scores. For securities data, extract features such as technical indicators and capital flows. For macro data, extract features such as trends and cycles. Pattern recognition focuses on discovering the correlation between data, such as the correlation between consumer behavior and economic cycles, the correlation between stock market performance and macro indicators, etc. Multi-dimensional indicators are used to evaluate the effect of data cleaning. For outlier processing, evaluate whether the data distribution before and after processing is reasonable. For missing value filling, evaluate the degree of deviation between the filled value and the true value. If the evaluation indicators do not meet the preset standards, the cleaning method needs to be optimized and re-executed, such as adjusting the outlier judgment threshold, improving the missing value filling algorithm, etc.

[0059] S102. According to the differences in data formats and time granularity, data from different data sources are converted into a unified standard through normalization processing to obtain a standardized data set.

[0060] Obtain bank transaction data, securities market data, and macroeconomic data. For the differences in data formats and time granularities, use normalization techniques to process the data. If there are noise values in the data, use denoising methods to remove noise from trading volumes, market values, and economic values. If there are missing values in the data, use filling methods to fill in the missing values. Classify and store the processed data through storage methods to obtain a standardized data set. Use analysis methods to extract eigenvalue and identify pattern values from the standardized data set. If the eigenvalue extraction results do not meet the preset thresholds, re-execute the denoising and filling steps. According to the results of pattern value identification, use clustering algorithms to classify the data and obtain the final analysis results.

[0061] The normalization of financial data is the basis of data analysis. Taking bank transaction data as an example, the transaction amount units from different channels may vary. Some are in yuan, and some are in ten thousand yuan. Through normalization, all data can be unified to the same dimension. In securities market data, the base periods of the Shanghai Composite Index and the Shenzhen Component Index are different, and they need to be converted to the same base period for comparison. For macroeconomic data, indicators such as the growth rate of gross domestic product and the consumer price index have different time granularities, including monthly data, quarterly data, and annual data, and need to be unified to monthly data through methods such as interpolation. In terms of data denoising, there may be abnormally large transactions in bank transaction data. For example, a customer's one-time transaction amount exceeds ten times the historical transaction average. Use the moving average method to smooth the outliers and make the data more in line with the actual situation. There may be sharp price fluctuations in securities market data due to technical failures, and use the median filtering method to remove these noise points. For missing value processing, customer information in bank transaction data may be incomplete, such as missing fields like age and occupation. It can be filled according to the characteristics of similar customer groups or use multiple imputation methods to handle missing data. In securities market data, data missing due to holiday market closures can be filled with the data of the previous trading day. The conversion of quarterly data to monthly data in macroeconomic data can be achieved using cubic spline interpolation.

[0062] S103. In view of the dynamics and timeliness of financial data, use a stream data processing framework to collect market conditions and policy change data in real time and dynamically update the data distribution characteristics.

[0063] A streaming data processing framework is adopted to obtain real-time data streams from market quotation data sources and policy change data sources. In view of the dynamics and timeliness in the data streams, a dynamic update mechanism is used to adjust the data distribution characteristics. If there are abnormal fluctuations in the data streams, an anomaly detection algorithm is adopted to identify and filter out the abnormal data. According to the data distribution characteristics after dynamic adjustment, a feature extraction method is used to obtain the key feature values. If the key feature values do not meet the preset thresholds, the anomaly detection and dynamic update steps are executed again. A clustering algorithm is used to classify the key feature values to obtain the real-time classification results. According to the real-time classification results, a prediction model is used to predict the market trend to obtain the final prediction results.

[0064] Specifically, market quotation and policy change data usually come from multiple channels such as exchange systems, regulatory announcements, and news media. Using a streaming data processing framework can achieve real-time monitoring and processing of these data sources. For example, new data on indicators such as the price and trading volume of a certain stock will be generated every few seconds during a trading day. This kind of data has obvious characteristics of dynamics and timeliness, and a dynamic update mechanism needs to be adopted to adapt to the changes in data distribution. In practical applications, the latest data characteristics can be dynamically captured through the sliding time window method. For example, a five-minute sliding window is set to continuously observe the stock price fluctuations. When the price of a certain stock fluctuates by more than five percent within five minutes, it may indicate an abnormal fluctuation. At this time, an anomaly detection algorithm needs to be started for analysis, and indicators in multiple dimensions such as historical volatility and trading volume changes are calculated to determine whether it is an abnormal situation. For the detected abnormal data, it cannot be simply and directly eliminated, but needs to be comprehensively judged in combination with the overall market situation and policy environment. For example, during the period of major policy announcements, it is normal for the market to generally show large fluctuations. Although the data deviates from the normal state at this time, it still has important analysis value. Through the feature extraction method, key feature values such as price trends, trading volume change rates, and market sentiment indexes can be obtained from the original data. If the extracted feature values show that the market fluctuates greatly but do not reach the warning threshold, the system will continue to monitor. When it is found that the intraday price volatility of a certain stock exceeds the preset threshold, the system will automatically trigger a new round of anomaly detection and dynamic update processes. Through the clustering algorithm, stocks with similar characteristics can be classified, so as to discover similar behavior patterns in the market. For example, stocks in the same industry, with the same concept, or having similar price trends can be grouped together, which helps to analyze the overall performance of a certain type of stocks.

[0065] S104. In the real-time data stream, an adaptive algorithm is used to judge the change trend of the data distribution. If a significant change is detected, the model parameters are automatically adjusted.

[0066] Obtain real-time data streams using a stream data processing framework, and monitor the changing trend of data distribution through an adaptive algorithm. If a significant change in data distribution is detected, adjust the model parameters according to the change amplitude. Use an anomaly detection algorithm to filter the adjusted data and remove abnormal fluctuations. Extract key feature values based on the characteristics of the filtered data distribution. If the key feature values exceed the preset threshold, re-execute the anomaly detection and parameter adjustment steps. Use a clustering algorithm to classify the extracted key feature values and obtain the classification results. According to the classification results, use a prediction model to predict the market trend and obtain the final prediction result.

[0067] Specifically, the stream data processing framework monitors market changes by collecting and processing data streams in real time. Taking the stock market as an example, when indicators such as stock price and trading volume are continuously generated during a trading day, the stream processing framework can receive and process these data in real time. Through the moving time window method, features such as the price fluctuation range and trading volume change within five minutes can be dynamically calculated. The adaptive algorithm monitors the changing trend by calculating statistical indicators such as the mean and variance of the data distribution. For example, in stock trading, if it is found that the price fluctuation within five minutes exceeds twice the previous volatility, it can be determined as a significant change. At this time, the weight parameters in the prediction model need to be adjusted accordingly to increase the sensitivity to recent data. For the detection of abnormal fluctuations, a statistics-based method can be used to identify outliers. Taking stock price data as an example, if the increase or decrease rate in a certain minute exceeds three times the average fluctuation in the previous ten minutes, it may be an abnormal fluctuation. By setting reasonable detection windows and thresholds, such abnormal data can be effectively filtered. The extraction of key feature values needs to consider multiple dimensions. In the stock market, indicators such as price trend, trading volume change, and turnover rate can be extracted. When these feature values are abnormal, such as when the turnover rate suddenly magnifies to more than five times the daily average level, anomaly detection and parameter adjustment need to be performed again. Cluster analysis can discover different states of the market. By clustering the extracted feature values, the market state can be divided into different categories such as the oscillation period, the rising period, and the falling period. For example, when indicators such as price and trading volume form a certain specific combination, it may indicate that the market is in a strong rising stage. The final prediction model needs to comprehensively consider the market state and historical rules. In a specific market state, the future market trend can be predicted according to the evolution rules of similar states in historical data.

[0068] S105. In view of the complexity of high-dimensional data, use principal component analysis technology to extract key features from multi-dimensional features and reduce the data dimension.

[0069] The principal component analysis technique is used to extract key features from high-dimensional data and reduce the data dimension. According to the data distribution characteristics after dimensionality reduction, a clustering algorithm is used to classify the key features. If there are abnormal categories in the clustering results, an anomaly detection algorithm is used to filter the abnormal categories. According to the data distribution characteristics after filtering, the key feature values are extracted. If the key feature values exceed the preset threshold, the anomaly detection and clustering steps are executed again. A prediction model is used to predict the trend of the classified data. According to the prediction results, the final market trend analysis results are determined.

[0070] Specifically, the principal component analysis technique processes high-dimensional data through dimensionality reduction, which can compress the data from the original multiple dimensions to fewer principal components. Taking stock market data as an example, the original data includes multiple dimensions such as opening price, closing price, highest price, lowest price, trading volume, etc. Through principal component analysis, the most representative features can be extracted, such as price trends and trading volume change trends. Based on the dimensionality-reduced data, clustering analysis is performed. Common methods include distance-based clustering and density-based clustering. Taking real estate market data as an example, after dimensionality reduction of multi-dimensional features such as housing prices, geographical locations, and construction years through principal component analysis, the density clustering method can be used to group properties with similar features into one category, forming different market segmentation groups. When detecting and filtering abnormal categories in the clustering results, statistical methods or machine learning methods can be used. Taking the analysis of consumer behavior in the retail industry as an example, if the purchase amount or frequency of a certain type of consumer far exceeds the normal range, it may be due to data collection errors or abnormal transactions, and filtering needs to be carried out by setting reasonable thresholds.

[0071] S106. Among the sparse features, the importance of features is judged through a feature selection algorithm. If the feature occurrence frequency is lower than the preset threshold, the invalid features are removed.

[0072] A feature selection algorithm is used to calculate the importance of sparse features to obtain the feature importance distribution. If the feature occurrence frequency is lower than the preset threshold, the invalid features are removed and the valid features are retained. According to the data distribution of the valid features, a dimensionality reduction processing technique is used to extract key features. For the key features after dimensionality reduction, a clustering algorithm is used for classification to obtain the category distribution. If there are abnormal categories in the category distribution, an anomaly detection algorithm is used to filter the abnormal data. According to the data distribution after filtering, the key feature values are extracted. If the key feature values exceed the preset threshold, the anomaly detection and clustering steps are executed again. A prediction model is used to predict the trend of the classified data to determine the final analysis results.

[0073] Specifically, during the feature selection process, it is necessary to evaluate the importance of each feature, and metrics such as the coefficient of variation and information gain are calculated to measure it. For example, when analyzing stock data, it may be found that the occurrence frequencies of indicators such as trading volume and price-to-earnings ratio are higher than the preset 80% threshold, while features such as the impact of temporary policies have lower occurrence frequencies. Therefore, the former are retained as effective features. For the dimensionality reduction of high-dimensional data, principal component analysis technology can be used to extract key components from multi-dimensional features. Taking enterprise operation analysis as an example, the original data may contain dozens of financial indicators, and through dimensionality reduction, it can be compressed into several key dimensions such as revenue ability, profitability, and development potential, which not only retains the main features of the data but also reduces the analysis difficulty.

[0074] S107. According to the extracted key features and filtered sparse features, use the random forest algorithm to construct a credit risk assessment model and train the model to predict the risk probability.

[0075] According to the data distributions of the sparse features and key features, use the random forest algorithm to construct a credit risk assessment model. Train the model to predict the risk probability. If the risk probability is lower than the preset threshold, it is marked as a low-risk category. For high-risk category data, use an anomaly detection algorithm to filter out abnormal samples. According to the data distribution after filtering, re-extract the key feature values. Use the random forest algorithm to rank the importance of the key feature values and eliminate low-importance features. According to the importance ranking results, update the credit risk assessment model. Use the updated model to predict the risk of the data and determine the final risk assessment result.

[0076] Specifically, in the field of financial risk control, when building a credit risk assessment model, the data distribution characteristics must be considered first. Taking bank credit business as an example, when customers apply for loans, they will provide multi-dimensional information, including sparse characteristics such as income level, occupation type, and housing situation, as well as key characteristics such as credit card usage frequency and repayment records. The random forest algorithm can effectively process these complex feature combinations to build a risk prediction model. In practical applications, model training requires a large amount of historical data support. A certain bank has collected 100,000 loan records, including 300 feature dimensions. The risk probability threshold is set at 20%. Customers with a probability lower than this threshold are marked as low-risk categories. For high-risk categories exceeding the threshold, it is necessary to further analyze whether there are abnormal situations. Common scenarios for anomaly detection in practice include false information identification. For example, some customers may provide false income certificates or forge asset certificates, and such data will show obvious deviations in the feature distribution. By using density-based anomaly detection algorithms, these abnormal samples can be identified. In a certain case, 5% of the high-risk samples were found to be suspected of data fraud through anomaly detection. After filtering out the abnormal samples, it is necessary to re-extract the key features. Taking the housing loan business as an example, indicators such as property value, down payment ratio, and monthly payment to income ratio are all important features. The random forest algorithm can calculate the importance scores of each feature and sort them. It is found in practice that the importance score of the monthly payment to income ratio is the highest, reaching 35%, followed by the credit record score accounting for 25%. After sorting the feature importance, features with lower scores can be removed. For example, features such as customers' hobbies and whether they have a driver's license, although they are also part of the customer portrait, contribute less to the prediction of credit risk and can be removed from the model. In this way, the feature dimensions of the model can be reduced from 300 to about 50. The updated model performs more stably in risk prediction. In the practice of a certain bank, the model accuracy has been improved from 85% to 92%. This improvement is not only reflected in the overall indicators but also in the recognition rate of high-risk customers. Through dynamic updates and continuous optimization, the model can adapt to the changing market environment.

[0077] S108. After the model training is completed, evaluate the model performance through cross-validation technology. If the prediction accuracy does not reach the preset standard, readjust the feature extraction and model parameters.

[0078] The cross-validation technique is used to evaluate the performance of the trained model, and the prediction accuracy of the model on the validation set is obtained. According to the comparison result between the prediction accuracy and the preset standard, it is judged whether the model performance meets the requirements. If the prediction accuracy is lower than the preset standard, it enters the feature extraction and parameter adjustment link. For the feature extraction link, the random forest algorithm is used to rank the importance of features for the original data, and the key features with importance scores higher than the preset threshold are obtained. For the parameter adjustment link, the grid search technique is used to search within the preset parameter range to obtain the optimal parameter combination of the model. According to the extracted key features and the optimized parameter combination, the credit risk assessment model is retrained to obtain an updated model. The updated model is used for cross-validation to obtain a new prediction accuracy, and it is judged whether it meets the preset standard. If the prediction accuracy meets the preset standard, the model is determined as the final credit risk assessment model; if not, the feature extraction and parameter adjustment process are repeated until the model performance reaches the standard.

[0079] S109. According to the final training result, determine the optimal model parameters and feature combination, and generate a risk prediction report.

[0080] The cross-validation technique is used to evaluate the performance of the trained model, and the prediction accuracy of the model on the validation set is obtained. According to the comparison result between the prediction accuracy and the preset standard, it is judged whether the model performance meets the requirements. If the prediction accuracy is lower than the preset standard, the random forest algorithm is used to rank the importance of features for the original data, and the key features with importance scores higher than the preset threshold are obtained. For the parameter adjustment link, the grid search technique is used to search within the preset parameter range to obtain the optimal parameter combination of the model. According to the extracted key features and the optimized parameter combination, the credit risk assessment model is retrained to obtain an updated model. The updated model is used for cross-validation to obtain a new prediction accuracy, and it is judged whether it meets the preset standard. If the prediction accuracy meets the preset standard, the model is determined as the final credit risk assessment model; if not, the feature extraction and parameter adjustment process are repeated until the model performance reaches the standard. According to the finally determined model parameters and feature combination, a risk prediction report is generated.

[0081] The above is only a preferred specific embodiment of the present application, but the protection scope of the present application is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed in the present application should be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. A financial risk analysis method based on big data, characterized in that: include: Obtain bank transaction data, securities market data and macroeconomic data. Based on the multi-source heterogeneous characteristics of different data sources, use data cleaning technology to remove noise and missing values ​​and obtain data from different data sources; According to the differences in data formats and time granularity, the data from different data sources are converted into a unified standard through normalization processing to obtain a standardized data set; In view of the dynamic and timeliness of financial data, a stream data processing framework is used to collect market conditions and rule change data in real time, and dynamically update data distribution characteristics; In real-time data streams, adaptive algorithms are used to determine the changing trend of data distribution. If significant changes are detected, model parameters are automatically adjusted. Use principal component analysis technology to extract key features from multi-dimensional features and reduce data dimensions; In sparse features, the importance of features is determined by feature selection algorithms. If the frequency of a feature is lower than a preset threshold, invalid features are eliminated. Based on the extracted key features and filtered sparse features, a random forest algorithm is used to build a credit risk assessment model, and the model is trained to predict risk probability; Acquire real-time data, input the real-time data into the credit risk assessment model, and generate a risk prediction report.

2. The method according to claim 1, characterized in that The process of obtaining data from different data sources includes: Obtain bank transaction data, securities market data and macroeconomic data, and use denoising methods to remove noise from transaction volume, market value and economic value based on the multi-source heterogeneity characteristics of different data sources; The missing values ​​were filled using the gap filling method. The cleaned data is classified and stored through the storage method; Use analytical methods to extract features and recognize patterns from processed data; If the data cleaning effect does not meet the preset threshold, the denoising and gap filling steps are re-executed to generate data from the different data sources.

3. The method according to claim 1, characterized in that The process of obtaining the standardized data set includes: According to the differences in data formats and time granularity, the data from the different data sources are normalized to obtain a normalized data set; Feature value extraction and pattern value recognition are performed on the normalized data set, and according to the result of the pattern value recognition, a clustering algorithm is used to classify the data to obtain the standardized data set.

4. The method according to claim 1, characterized in that: The process of dynamically updating data distribution characteristics includes: Use a stream data processing framework to obtain real-time data streams from market data sources and policy change data sources. If there are abnormal fluctuations in the data stream, use anomaly detection algorithms to identify and filter abnormal data. In view of the dynamics and timeliness of data streams, the data distribution characteristics are adjusted through a dynamic update mechanism; According to the dynamically adjusted data distribution characteristics, the key feature values ​​are obtained by using the feature extraction method. If the key feature values ​​do not meet the preset threshold, the anomaly detection and dynamic update steps are re-executed.

5. The method according to claim 1, characterized in that The process of extracting key features from multi-dimensional features using principal component analysis technology and reducing data dimensions includes: Use principal component analysis technology to extract key features from high-dimensional data and reduce data dimensions; According to the data distribution characteristics after dimensionality reduction, clustering algorithm is used to classify key features; If there are abnormal categories in the clustering results, the abnormal categories are filtered using an anomaly detection algorithm, and key feature values ​​are extracted based on the distribution characteristics of the filtered data.

6. The method according to claim 1, characterized in that The process of constructing a credit risk assessment model using a random forest algorithm based on the extracted key features and the filtered sparse features, and training the model to predict risk probability includes: Based on the data distribution of sparse features and key features, a random forest algorithm is used to build a credit risk assessment model; The training model predicts the risk probability. If the risk probability is lower than the preset threshold, it is marked as a low-risk category. For high-risk category data, anomaly detection algorithms are used to filter out abnormal samples; Re-extract key feature values ​​based on the filtered data distribution; The random forest algorithm is used to sort the importance of key feature values ​​and eliminate low-importance features; Update the credit risk assessment model based on the importance ranking results.

7. The method according to claim 6, characterized in that The process of updating the credit risk assessment model includes: Use cross-validation technology to evaluate the performance of the trained model and obtain the prediction accuracy of the model on the validation set; Based on the comparison between the prediction accuracy and the preset standard, determine whether the model performance meets the requirements; If the prediction accuracy is lower than the preset standard, the feature extraction and parameter adjustment phase will be entered; For the feature extraction stage, the random forest algorithm is used to sort the feature importance of the original data and obtain the key features whose importance scores are higher than the preset threshold; For the parameter adjustment phase, grid search technology is used to search within the preset parameter range to obtain the optimal parameter combination of the model; Retrain the credit risk assessment model based on the extracted key features and optimized parameter combinations to obtain an updated model; Use the updated model for cross-validation to obtain new prediction accuracy and determine whether it meets the preset standards; If the prediction accuracy meets the preset standards, the model is determined to be the final credit risk assessment model.

8. An electronic device comprising a memory, a processor, and a computing program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the method described in any one of claims 1 to 7 is implemented.

9. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the method described in any one of claims 1 to 7 is implemented.