Training method of financial risk identification model based on big data
By using technical means such as multimodal data fusion, explanatory enhancement, real-time anomaly detection, NLP sentiment analysis and incremental training in the financial risk identification model, the existing models are solved in capturing market variable risk factors, lack of interpretability and adaptability, and achieving higher risk identification accuracy and model transparency.
Patent Information
- Application Number
- CN202510023645.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-07
- Publication Date
- 2025-05-16
AI Technical Summary
Existing financial risk identification models have flaws in capturing market-changing risk factors, lack interpretability, poor adaptability and real-timeness, and cannot fully utilize potential risk information in unstructured data.
The training method of the financial risk identification model based on big data is adopted, and by collecting multiple data sources, extracting multimodal features, enhancing model interpretability, real-time anomaly detection, combined with NLP sentiment analysis and event-driven prediction, as well as incremental training and hyperparameter optimization.
It improves the accuracy and transparency of risk identification, enhances the adaptability and real-time nature of the model, and can more accurately capture market sentiment and event-driven risk factors, ensuring that the model remains efficient and reliable in complex market environments.
Smart Images

Figure CN120013672A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of financial technology and big data analysis technology, and specifically to a training method for a financial risk identification model based on big data. Background Art
[0002] In the modern financial field, risk identification and prediction are key means to ensure the sound operation of financial institutions. Traditional financial risk identification models are mostly based on structured data, such as financial statements, market transaction data, etc., and their application scope is relatively single. As the financial market becomes increasingly complex, models that rely on a single data source have defects in capturing the changing risk factors of the market. In addition, unstructured data in the financial market, such as news and social media comments, contain potential risk information. How to integrate data and extract effective features from it has become an important technical challenge.
[0003] Existing financial risk identification models mainly rely on traditional machine learning algorithms or statistical models, which have limited interpretability and cannot provide transparency under complex market conditions. For example, deep learning models improve the accuracy of risk prediction in some aspects, but their "black box" characteristics make it difficult for users to understand the decision-making process of the model. Especially in the financial field, the transparency and interpretability of the model are particularly important to financiers. Therefore, the lack of interpretability is one of the main problems in existing technologies.
[0004] In terms of anomaly detection, traditional risk identification methods usually rely on fixed thresholds for detection, which appears rigid when the market fluctuates, resulting in poor adaptability and real-time performance of the model and inability to meet rapidly changing market demands. Risk events occur frequently in the market environment, and traditional anomaly detection algorithms are unable to dynamically respond to sudden changes, affecting the timeliness of risk identification.
[0005] In addition, with the rapid development of natural language processing technology, market sentiment and event-driven risk analysis have received more attention. Emotional information and important events in the financial market can have a direct impact on risk prediction. However, existing technologies do not fully utilize emotional and event-driven data, resulting in the risk identification model's low ability to perceive market sentiment fluctuations, affecting the accuracy and comprehensiveness of risk identification.
[0006] In terms of model training, traditional model training methods use static training data sets and lack a real-time update mechanism. However, data in the financial market is constantly flowing and changing, and static data sets cannot represent real-time market characteristics. Due to the lack of effective incremental training strategies and dynamic hyperparameter optimization mechanisms, existing models show low stability and reduced prediction accuracy when responding to changes in data distribution.
[0007] Therefore, those skilled in the art provide a training method for a financial risk identification model based on big data to solve the problems raised in the above background technology. Summary of the invention
[0008] In view of the deficiencies in the prior art, the present invention provides a training method for a financial risk identification model based on big data to solve the problems raised in the above background technology.
[0009] To achieve the above objectives, the present invention is implemented through the following technical solutions: a training method for a financial risk identification model based on big data, comprising the following steps:
[0010] S1. Collect data sources from structured data, unstructured data and semi-structured data, and pre-process the data;
[0011] S2, extracting multimodal features and fusing them to generate a multimodal feature vector;
[0012] S3. Use model interpretability enhancement methods, including SHAP values, LI ME local explanations, and rule-based interpretability modules to provide transparency of model prediction results;
[0013] S4, real-time anomaly detection through rolling windows and adaptive anomaly detection;
[0014] S5. Sentiment analysis and event-driven forecasting methods based on NLP can be used to obtain market sentiment and event intensity to assist in the training of risk identification models;
[0015] S6. Perform incremental training and hyperparameter optimization on the model, and verify the model performance through multiple evaluation indicators.
[0016] Preferably, the data preprocessing step includes: standardizing the structured data, converting the numerical feature x i Transformed into a distribution with a mean of 0 and a standard deviation of 1, the standardization formula is:
[0017]
[0018] Use the BERT model to embed unstructured text data into vector representation and generate sentence vector X t :
[0019]
[0020] Among them, h i is the embedding vector of the i-th word in the text sequence.
[0021] Preferably, the multimodal feature extraction step comprises:
[0022] Extract time series features from structured data;
[0023] The graph neural network is used to extract the graph structure features of the enterprise relationship network. The embedding update formula of node v is:
[0024]
[0025] Among them, N(v) is the neighbor set of node v, W (k) is a trainable weight matrix and σ is an activation function.
[0026] Preferably, the multimodal feature fusion step adopts the Attention mechanism to combine different modal features X s , X t and X g Perform weighted fusion to obtain the fused feature vector X fusion :
[0027] X fusion =∑ i α i X i ,
[0028] Among them, α i is the weight of each modality feature learned through the self-attention mechanism.
[0029] Preferably, the method for enhancing the interpretability of the model comprises:
[0030] Calculate the SHAP value of each feature to quantify the contribution of each feature to the prediction result. The SHAP value formula is:
[0031]
[0032] Among them, φ i Feature X i SHAP value of feature X i The marginal contribution to the predicted value of the model output; N is the set of all input features; Remove feature X from feature set N i The subset S after the input feature is input; |S| the number of elements in the feature subset S; |N| the number of elements in the feature set N; f(S) is the predicted output value of the model when the input feature subset is S;
[0033] The linear model is fitted in the local area by the LIME method to explain the prediction results of a single sample. The expression of the linear model is:
[0034]
[0035] Among them, g(x) is the local approximate explanation of the linear model at sample x, x i The i-th feature of sample x, w i Features iThe linear weight of b, the bias term of the linear model, and n the total number of features.
[0036] Preferably, the real-time anomaly detection method comprises:
[0037] The detection threshold is dynamically set in the rolling window. The threshold setting formula is:
[0038] |x-μ t |>k·σ t ,
[0039] Among them, μ t is the average value of the data in the window, σ t is the standard deviation, k is the sensitivity parameter;
[0040] Different anomaly detection algorithms are determined according to market volatility, and a density-based local outlier factor algorithm is used in a highly volatile market.
[0041] Preferably, the outlier factor LOF(x) calculation formula of the density-based local outlier factor algorithm is:
[0042]
[0043] Among them, N k (x) represents the k-neighborhood set of sample x, and lrd(x) is the local reachable density of sample x.
[0044] Preferably, the NLP-based sentiment analysis and event-driven prediction steps include:
[0045] Generate text sentiment vector X through BERT model sent And perform sentiment classification;
[0046] Use the event detection module to identify the event intensity in the text and quantify it into event scores:
[0047]
[0048] Among them, Event Score is the event intensity score, n is the total number of trigger words, and Weight is i The weight of the i-th trigger word or event feature, Relevance(T,trigger i ) Text T and the i-th trigger word trigger i The relevance score of .
[0049] Preferably, the incremental training method of the model includes:
[0050] Split the data into training, validation and test sets in chronological order;
[0051] Incremental training is performed within a rolling window, dynamically updating the model parameters θ to minimize the loss function L(θ):
[0052]
[0053] Among them, θ * The optimal parameter represents the optimal solution of the parameter θ obtained by minimizing the loss function L(θ) under the current incremental training data.
[0054] Preferably, the hyperparameter optimization method of the model is implemented by Bayesian optimization to optimize the performance of the model. The Bayesian optimization formula is:
[0055]
[0056] Among them, p(y|θ) is the predicted expectation under the current hyperparameter θ.
[0057] The present invention provides a training method for a financial risk identification model based on big data.
[0058] Beneficial effects:
[0059] 1. The present invention integrates structured data, unstructured data and semi-structured data through multimodal data fusion technology, uses data from different sources to provide the model with more comprehensive feature information, and adopts the Attention mechanism for weighted fusion to ensure that the weight of important features is higher, so that the model can more accurately capture the complex information in the financial market and improve the accuracy of risk identification.
[0060] 2. The present invention introduces SHAP and LIME algorithms, and combines them with the rule base in the financial field to enhance the interpretability of the model. The global contribution of each feature to the model output is calculated through the SHAP value, and local interpretation of individual samples is achieved through LIME, so that the basis of each prediction is clear. Especially in financial risk identification, this interpretability helps financiers understand the decision-making process of the model and enhances the transparency and reliability of the model.
[0061] 3. The present invention utilizes rolling windows and adaptive threshold settings to achieve real-time anomaly detection, enabling the model to quickly respond to changes in the market environment. By dynamically adjusting the threshold, the appropriate anomaly detection algorithm is determined under different market fluctuations, effectively improving the adaptability and real-time detection capabilities of the model, ensuring that potential risks can be quickly identified in a complex and changing market.
[0062] 4. The present invention combines NLP technology and uses a pre-trained language model to perform sentiment analysis and event detection on market text data to generate sentiment scores and event intensity quantification indicators. This information is introduced into the risk identification model, which helps to capture market sentiment fluctuations and important events, and reflect the impact of market sentiment in the forecast, so as to broaden the data dimension of risk identification and enhance the model's responsiveness to event-driven market fluctuations.
[0063] 5. The present invention designs an incremental training strategy based on a rolling window, which enables the model to continuously update parameters in new data streams and dynamically adjust hyperparameters through Bayesian optimization. Incremental training enables the model to continuously adapt to changes in data distribution and maintain the stability of the prediction effect. At the same time, dynamic hyperparameter optimization ensures that the model always runs in the optimal state, thereby improving the real-time and accuracy of the model. BRIEF DESCRIPTION OF THE DRAWINGS
[0064] Figure 1 It is a flow chart of the present invention. DETAILED DESCRIPTION
[0065] In order to make the technical personnel in the technical field understand the scheme of the present invention, the technical scheme in the embodiment of the present invention will be clearly and completely described below in combination with the drawings in the embodiment of the present invention. Obviously, the described embodiment is a partial embodiment of the present invention, not a complete embodiment. Based on the embodiment of the present invention, other embodiments obtained by ordinary technicians in the field without creative work should fall within the scope of protection of the present invention.
[0066] The present invention is described in detail below in conjunction with the accompanying drawings:
[0067] Example:
[0068] Please see attached Figure 1 The present invention provides a method for training a financial risk identification model based on big data, comprising the following steps:
[0069] S1. Collect data sources from structured data, unstructured data and semi-structured data, and pre-process the data;
[0070] S2, extracting multimodal features and fusing them to generate a multimodal feature vector;
[0071] S3. Use model interpretability enhancement methods, including SHAP values, LI ME local explanations, and rule-based interpretability modules to provide transparency of model prediction results;
[0072] S4, real-time anomaly detection through rolling windows and adaptive anomaly detection;
[0073] S5. Sentiment analysis and event-driven forecasting methods based on NLP can be used to obtain market sentiment and event intensity to assist in the training of risk identification models;
[0074] S6. Perform incremental training and hyperparameter optimization on the model, and verify the model performance through multiple evaluation indicators.
[0075] The benefits of S1 are that it can greatly expand data coverage, make full use of multiple data sources, form more comprehensive feature information, and the data preprocessing step standardizes and vectorizes text for different data types, effectively improving data quality and consistency, laying a solid foundation for subsequent feature extraction and model training;
[0076] The benefits of S2 are that it can more comprehensively reflect the multi-dimensional information of financial risks, ensure that important features are given higher weights, enhance the model's ability to capture complex risk factors, and improve the accuracy of risk identification;
[0077] The benefits of S3 are that it helps analyze and understand the contribution of each feature to the prediction results. At the same time, combined with the rule base in the financial field, it can make the model output consistent with the cognition of financiers, enhance the transparency and reliability of the model, and its interpretability enables financiers to intuitively understand the decision-making basis of the model, which is convenient for application and analysis of risk factors in actual business;
[0078] The benefits of S4 are that it can realize real-time risk detection in a dynamic environment. The rolling window technology ensures the dynamic adjustment capability of the threshold. The adaptive detection method flexibly determines different algorithms according to market fluctuations, so that the model can respond quickly to market changes, ensuring that potential risks are detected in time under highly volatile market conditions, and improving the timeliness and adaptability of risk identification.
[0079] The benefit of S5 is that it can effectively capture the market sentiment fluctuations and unexpected event information contained in unstructured text data. Sentiment scores and event intensity indicators are input into the model as additional features to help the model better identify risk factors driven by market sentiment, thereby improving the model's prediction accuracy and sensitivity to market changes.
[0080] The benefits of S6 are that the model can continuously adapt to changes in the market environment and data distribution, ensuring real-time prediction results and model stability. Incremental training ensures the real-time performance of the model. Bayesian optimization achieves the optimal combination of automatic adjustment of hyperparameters. Multi-dimensional evaluation indicators further verify the robustness and predictive performance of the model, providing high-precision and high-stability support for financial risk identification.
[0081] The data preprocessing steps include: standardizing the structured data, converting the numerical features x i Transformed into a distribution with a mean of 0 and a standard deviation of 1, the standardization formula is:
[0082]
[0083] Use the BERT model to embed unstructured text data into vector representation and generate sentence vector X t :
[0084]
[0085] Among them, h i is the embedding vector of the i-th word in the text sequence.
[0086] The benefits of the structured data standardization formula are to avoid the situation where features of different magnitudes produce weight imbalance in the model, accelerate parameter convergence during model training, improve the efficiency of model training, reduce the deviation between features, and help the model maintain a stable learning effect under different features, thereby improving the prediction stability and effect of the model;
[0087] The benefit of the BERT text embedding formula is that it can include the complete sentence meaning and context, and express the semantics of the text more accurately, so that unstructured data can be integrated with structured data, providing the model with more dimensional feature information, capturing the emotions and event information in the text, and enabling the model to better understand market sentiment and emergencies, thereby improving the accuracy of risk identification.
[0088] The multimodal feature extraction steps include:
[0089] Extract time series features from structured data;
[0090] The graph neural network is used to extract the graph structure features of the enterprise relationship network. The embedding update formula of node v is:
[0091]
[0092] Among them, N(v) is the neighbor set of node v, W (k) is a trainable weight matrix and σ is an activation function.
[0093] The benefits of time series feature extraction of structured data are that it effectively reflects the temporal dynamics of structured data, captures the trend and periodic fluctuation of data over time, helps identify potential risk change trends, and provides historical data basis for financial risk prediction, so that the model can provide more accurate predictions based on data changes in risk identification;
[0094] The benefit of the graph neural network feature extraction formula is that it can effectively capture the topological structure and complex relationships between nodes in the enterprise relationship network, give higher weights to important relationships, make the model more accurate in capturing key relationships in the financial network, and comprehensively identify potential risk transmission paths and systemic risks between enterprises, providing support for the real-time and scalability of the model.
[0095] The multimodal feature fusion step uses the Attention mechanism to combine different modal features X s , X t and X g Perform weighted fusion to obtain the fused feature vector X fusion :
[0096] X fusion =∑ i α i X i ,
[0097] Among them, α i is the weight of each modality feature learned through the self-attention mechanism.
[0098] The benefits of the multimodal feature fusion algorithm formula are that it gives text features a higher weight, ensuring that the model can use data from different modalities more accurately, improving the flexibility and effectiveness of feature fusion, allowing the model to focus on the most risk-indicative features when making decisions, enhancing the understanding of complex multi-dimensional financial information, improving the accuracy of risk identification, avoiding excessive attention to low-correlation features, and reducing information redundancy in the model. Its targeted feature fusion strategy ensures the integrity of information, reduces computational complexity, and enables the model to process multimodal data more efficiently. The weight distribution makes the model's decision-making process more interpretable, making it easier for financiers to understand why the model pays more attention to a certain feature in a specific situation, thereby improving the model's transparency and reliability, and better adapting to the dynamically changing market environment.
[0099] Methods to enhance model interpretability include:
[0100] Calculate the SHAP value of each feature to quantify the contribution of each feature to the prediction result. The SHAP value formula is:
[0101]
[0102] Among them, φ i Feature X i SHAP value of feature X i The marginal contribution to the predicted value of the model output; N is the set of all input features; Remove feature X from feature set N i The subset S after the input feature is input; |S| the number of elements in the feature subset S; |N| the number of elements in the feature set N; f(S) is the predicted output value of the model when the input feature subset is S;
[0103] The linear model is fitted in the local area by the LI ME method to explain the prediction results of a single sample. The expression of the linear model is:
[0104]
[0105] Among them, g(x) is the local approximate explanation of the linear model at sample x, x i The i-th feature of sample x, w i Features i The linear weight of b, the bias term of the linear model, and n the total number of features.
[0106] The benefit of the model interpretability enhancement algorithm formula is that the impact of each model feature on the final prediction result can be clearly reflected. The feature importance ranking allows users to intuitively understand which feature is the most critical in risk prediction. It provides global interpretability, helps to reveal the decision logic of the model, improves model transparency, and makes financiers more confident in the decision-making process of the model. Users can detect features that have a smaller impact on the prediction results and perform feature selection or feature engineering optimization to improve model performance, reduce computational complexity, and optimize model training efficiency.
[0107] The benefit of the LI ME local explanation formula is that it can intuitively reveal the specific basis for the sample prediction. Users can clearly understand the degree of dependence of the model on each feature of the sample, enhance the interpretability of the model on a single prediction result, help identify important risk factors in specific cases, and improve the flexibility and applicability of financial risk identification. Financial analysts can clearly understand why the model focuses on a certain feature under specific conditions, which facilitates the evaluation of the rationality of the model's prediction results and supports more reliable decision-making.
[0108] Real-time anomaly detection methods include:
[0109] The detection threshold is dynamically set in the rolling window. The threshold setting formula is:
[0110] |x-μ t |>k·σ t ,
[0111] Among them, μ t is the average value of the data in the window, σ t is the standard deviation, k is the sensitivity parameter;
[0112] Determine different anomaly detection algorithms based on market volatility, and use a density-based local outlier factor algorithm in a highly volatile market;
[0113] The calculation formula of the outlier factor LOF(x) of the density-based local outlier factor algorithm is:
[0114]
[0115] Among them, N k (x) represents the k-neighborhood set of sample x, and lrd(x) is the local reachable density of sample x.
[0116] The benefits of the real-time anomaly detection algorithm formula are that the detection threshold automatically adapts to the market environment. Especially in the financial market, where prices and trading volumes fluctuate frequently, dynamic thresholds can more accurately identify anomalies, ensuring real-time and adaptability. The model can effectively distinguish between normal fluctuations and true anomalies, while reducing false positives and missed negatives, and improving the accuracy of anomaly detection;
[0117] The advantage of the density-based local outlier factor algorithm is that it can identify local anomalies more accurately. The outlier factor algorithm can distinguish between normal points in high-density areas and abnormal points in low-density areas by calculating the local density of samples. Therefore, the outlier factor algorithm shows higher detection accuracy and adaptability under high volatility conditions, effectively avoiding the misdetection of normal points in dense areas as abnormal points, reducing false positives, and improving detection accuracy. It is particularly effective for high-frequency transactions and dense volatility situations that are common in financial data, ensuring real-time responsiveness to market changes.
[0118] The steps of NLP-based sentiment analysis and event-driven prediction include:
[0119] Generate text sentiment vector X through BERT model sent And perform sentiment classification;
[0120] Use the event detection module to identify the event intensity in the text and quantify it into event scores:
[0121]
[0122] Among them, Event Score is the event intensity score, n is the total number of trigger words, and Weight is i The weight of the i-th trigger word or event feature, Relevance(T,trigger i ) Text T and the i-th trigger word trigger i The relevance score of .
[0123] The benefits of NLP-based sentiment analysis and event-driven prediction algorithm formulas are to accurately capture market sentiment fluctuations and improve the model's risk perception ability. Sentiment analysis generates text sentiment vectors and classifications through the BERT model, which can effectively capture market sentiment, such as optimism, pessimism and other emotional changes, and provide key sentiment input features for risk identification models, so that the model can more accurately perceive the impact of market sentiment fluctuations on financial risks. It can quantify the influence of events mentioned in the text and help the model understand the potential risks of the event. Its quantification method makes it easier for the model to identify events with high-risk indicators during training, improve the model's sensitivity to emergencies and prediction accuracy, and the adaptability of weights ensures that the model is It can be flexibly adjusted in different contexts to enhance the flexibility of event detection. The BERT-based text vectorization and trigger word relevance score improve the semantic understanding of the text, enabling the model to deeply explore the hidden information in unstructured data such as news and reports, and provide more diverse information sources for risk identification. Broaden the data sources for risk identification and improve the ability to respond to event-driven market fluctuations. Sentiment analysis and event-driven prediction provide the risk identification model with two layers of information dimensions, namely, emotions and events, so that the model can focus on financial indicators and respond to real-time information such as news and social media, so that the model can respond quickly to event-driven market fluctuations, improving real-time and prediction capabilities.
[0124] The incremental training methods of the model include:
[0125] Split the data into training, validation and test sets in chronological order;
[0126] Incremental training is performed within a rolling window, dynamically updating the model parameters θ to minimize the loss function L(θ):
[0127]
[0128] Among them, θ * The optimal parameter represents the optimal solution of the parameter θ obtained by minimizing the loss function L(θ) under the current incremental training data.
[0129] The benefits of the incremental training method algorithm formula are that it ensures that the model adapts to real-time changes in the market in a timely manner. Compared with the traditional static training method, incremental training can adjust the model parameters when each round of new data is input, thereby ensuring that the model remains real-time and updated in a changing market environment, and improving the real-time response capability of risk identification. Real-time adjustment of parameters can help the model overcome changes in data distribution, maintain the stability and high accuracy of the prediction effect, prevent the model from being "outdated" in long-term operation, avoid the computational overhead of retraining the entire model, reduce the demand for computing resources, and ensure that the model is updated in a short time, which is conducive to efficient and real-time risk identification in an environment with limited computing resources. The rolling window requires the model to process part of the new data without loading all the data, thereby improving the scalability of the model on large-scale data, so that the model can run for a long time without being limited by the size of the data, and helping to identify and adapt to new risk characteristics or market trends. For emergencies or emerging risk factors in the financial market, the incremental training method allows the model to quickly capture and reflect changes, improve the model's ability to identify new risks in a dynamic market, and enhance the accuracy and stability of the model, making the risk identification results more reliable.
[0130] The hyperparameter optimization method of the model is implemented through Bayesian optimization to optimize the performance of the model. The Bayesian optimization formula is:
[0131]
[0132] Among them, p(y|θ) is the predicted expectation under the current hyperparameter θ.
[0133] The benefits of the Bayesian optimization algorithm formula are that it can intelligently determine the area in the hyperparameter space that is most likely to improve model performance for evaluation. Bayesian optimization determines the optimal hyperparameter combination with a smaller number of trials, greatly saving computing resources. The Bayesian optimization algorithm automatically determines and updates the hyperparameter combination, so that the hyperparameter optimization process does not require human intervention. The automated process reduces the complexity of tuning, and is particularly suitable for dealing with situations where there are a large number of hyperparameters and complex influencing relationships, providing a convenient means for the model to achieve optimal performance. Bayesian optimization uses the expectation maximization strategy to gradually converge to the optimal hyperparameter combination by constructing a priori distribution and updating the posterior distribution, thereby improving optimization efficiency and ensuring that a better hyperparameter combination is determined under limited resources. By optimizing the determined hyperparameter combination, the model can improve generalization performance and make its performance on the test set or new data more robust. It handles discrete parameters or mixed parameter combinations and is widely used in the hyperparameter tuning process of various models to adapt to different types of model structures and hyperparameter requirements. Bayesian optimization can determine the hyperparameter combination that optimizes model performance, and improves the prediction accuracy of the model, which is particularly critical for high-demand tasks such as financial risk identification, ensuring that the model has higher accuracy and reliability in practical applications.
[0134] Although embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions and variations may be made to the embodiments without departing from the principles and spirit of the present invention, and that the scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A training method for a financial risk identification model based on big data, characterized in that: The following steps are involved: S1. Collect data sources from structured data, unstructured data and semi-structured data, and pre-process the data; S2, extracting multimodal features and fusing them to generate a multimodal feature vector; S3. Use model interpretability enhancement methods, including SHAP values, LIME local explanations, and rule-based interpretability modules to provide transparency of model prediction results. S4, real-time anomaly detection through rolling windows and adaptive anomaly detection; S5. Sentiment analysis and event-driven forecasting methods based on NLP can be used to obtain market sentiment and event intensity to assist in the training of risk identification models; S6. Perform incremental training and hyperparameter optimization on the model, and verify the model performance through multiple evaluation indicators.
2. According to the training method of a financial risk identification model based on big data according to claim 1, it is characterized in that: The data preprocessing step includes: standardizing the structured data, converting the numerical feature x i Transformed into a distribution with a mean of 0 and a standard deviation of 1, the standardization formula is: Use the BERT model to embed unstructured text data into vector representation and generate sentence vector X t : Among them, h i is the embedding vector of the i-th word in the text sequence.
3. According to the training method of a financial risk identification model based on big data in claim 1, it is characterized in that: The multimodal feature extraction step comprises: Extract time series features from structured data; The graph neural network is used to extract the graph structure features of the enterprise relationship network. The embedding update formula of node v is: Among them, N(v) is the neighbor set of node v, W (k) is a trainable weight matrix and σ is an activation function.
4. The training method of a financial risk identification model based on big data according to claim 1 is characterized in that: The multimodal feature fusion step adopts the Attention mechanism to combine different modal features X s , X t and X g Perform weighted fusion to obtain the fused feature vector X fusion : X fusion =∑ i a i X i , Among them, α i is the weight of each modality feature learned through the self-attention mechanism.
5. The training method of a financial risk identification model based on big data according to claim 1 is characterized in that: Methods to enhance the interpretability of the model include: Calculate the SHAP value of each feature to quantify the contribution of each feature to the prediction result. The SHAP value formula is: Among them, φ i Feature X i SHAP value of feature X i The marginal contribution to the predicted value of the model output; N is the set of all input features; Remove feature X from feature set N i The subset S after the input feature is input; |S| the number of elements in the feature subset S; |N| the number of elements in the feature set N; f(S) is the predicted output value of the model when the input feature subset is S; The linear model is fitted in the local area by the LIME method to explain the prediction results of a single sample. The expression of the linear model is: Among them, g(x) is the local approximate explanation of the linear model at sample x, x i The i-th feature of sample x, w i Features i The linear weight of b, the bias term of the linear model, and n the total number of features.
6. The training method of a financial risk identification model based on big data according to claim 1 is characterized in that: The real-time anomaly detection method comprises: The detection threshold is dynamically set in the rolling window. The threshold setting formula is: |x-m t |>k·s t , Among them, μ t is the average value of the data in the window, σ t is the standard deviation, k is the sensitivity parameter; Different anomaly detection algorithms are determined according to market volatility, and a density-based local outlier factor algorithm is used in a highly volatile market.
7. The training method of a financial risk identification model based on big data according to claim 6 is characterized in that: The outlier factor LOF(x) calculation formula of the density-based local outlier factor algorithm is: Among them, N k (x) represents the k-neighborhood set of sample x, and lrd(x) is the local reachable density of sample x.
8. The training method of a financial risk identification model based on big data according to claim 1 is characterized in that: The NLP-based sentiment analysis and event-driven prediction steps include: Generate text sentiment vector X through BERT model sent And perform sentiment classification; Use the event detection module to identify the event intensity in the text and quantify it into event scores: Among them, Event Score is the event intensity score, n is the total number of trigger words, and Weight is i The weight of the i-th trigger word or event feature, Relevance(T,trigger i ) Text T and the i-th trigger word trigger i The relevance score of .
9. The training method of a financial risk identification model based on big data according to claim 1 is characterized in that: The incremental training method of the model includes: Split the data into training, validation and test sets in chronological order; Incremental training is performed within a rolling window, dynamically updating the model parameters θ to minimize the loss function L(θ): Among them, θ * The optimal parameter represents the optimal solution of the parameter θ obtained by minimizing the loss function L(θ) under the current incremental training data.
10. The training method of a financial risk identification model based on big data according to claim 1, characterized in that: The hyperparameter optimization method of the model is implemented through Bayesian optimization to optimize the performance of the model. The Bayesian optimization formula is: Among them, p(y|θ) is the predicted expectation under the current hyperparameter θ.
Citation Information
Cited By
Cross-modal transaction feature self-generation and anomaly detection system and method fused with large language model
CN120258990A