Intelligent financial risk monitoring method based on AI

By collecting multi-source heterogeneous data in real time and using GNN and Transformer models for feature extraction and fusion, combined with Monte Carlo simulation, the data silos and lack of dynamics of traditional financial risk monitoring methods are solved, achieving more accurate and flexible risk monitoring and response.

CN120725784AInactive Publication Date: 2025-09-30上海御胜信息科技股份有限公司
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510846400.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-24
Publication Date
2025-09-30
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing financial risk monitoring methods rely on a single data source, resulting in data silos, single data types, insufficient dynamism, unclear risk propagation mechanisms, and high demand for manual intervention. These methods make it difficult to fully reflect market dynamics and rapid changes, and lack scientific risk response strategies.

Method used

An AI-based intelligent financial risk monitoring method is adopted to collect multi-source heterogeneous data in real time, and feature extraction and fusion are performed through the GNN model and time series Transformer model. Combined with the Monte Carlo simulation method, a risk diffusion probability matrix is ​​constructed to determine systemic risks and initiate response mechanisms.

Benefits of technology

It improves the accuracy and flexibility of risk monitoring, can respond to market changes in a timely manner, clarify the risk transmission path, provide scientific risk response strategies, reduce manual intervention, and reduce potential losses.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120725784A_ABST
    Figure CN120725784A_ABST
Patent Text Reader

Abstract

The invention provides an intelligent financial risk monitoring method based on AI, relates to the technical field of financial science and technology, and aims to collect multi-source heterogeneous data including structured, semi-structured and unstructured data in the financial field in real time, overcome the data island phenomenon and promote information sharing. A multi-dimensional data set is obtained through data cleaning, feature extraction and fusion are performed by using a GNN model and a time sequence Transform model, and multi-dimensional feature space and time sequence features are formed. And quantifying the characteristics to obtain risk quantitative indexes, determining a risk diffusion probability matrix based on a Monte Carlo simulation method, and further calculating a probability value of systematic risk outbreak. And when the probability value exceeds a set critical value, determining that a systematic risk exists and starting a risk response mechanism. According to the method, the accuracy and timeliness of risk monitoring are improved, manual intervention is reduced, the ability of financial institutions to deal with sudden risk events is enhanced, and the potential loss risk is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of financial technology, and in particular to an AI-based intelligent financial risk monitoring method. Background Art

[0002] In modern financial markets, risk management has become a core component of financial institutions' operations. As market complexity and uncertainty continue to grow, traditional approaches to financial risk monitoring face numerous challenges. Existing risk monitoring systems typically rely on a single data source, such as historical transaction data or financial statements. This single-source approach to data processing leads to the following issues:

[0003] (1) Data silos: The lack of an effective data sharing mechanism among different financial institutions has led to the formation of information silos. Each institution can only rely on its own data for risk assessment and is unable to fully understand market dynamics and potential risks.

[0004] (2) Single data type: Traditional systems focus primarily on structured data (such as transaction flows and financial statements), while ignoring the importance of semi-structured and unstructured data (such as social media sentiment and financial news). This makes risk monitoring unable to fully reflect changes in market sentiment and the external environment.

[0005] (3) Lack of dynamism: Existing risk assessment models are often based on static data and are unable to capture the rapid changes in the market and the dynamic risk characteristics. This lag makes financial institutions less responsive to sudden risk events, increasing the risk of potential losses.

[0006] (4) Unclear risk transmission mechanisms: Traditional methods lack in-depth analysis of risk transmission pathways and are unable to effectively predict the risk contagion effects across industries. This leaves financial institutions without a scientific basis for formulating risk response strategies, making it difficult to effectively prevent the outbreak of systemic risks.

[0007] (5) High demand for manual intervention: Existing systems often rely on manual setting of thresholds and frequencies for risk warning and monitoring, resulting in insufficient flexibility of the warning mechanism and difficulty in adapting to rapid market changes. Summary of the Invention

[0008] In order to overcome the shortcomings of the existing technology, the purpose of the present invention is to provide an AI-based intelligent financial risk monitoring method, which not only improves the prediction accuracy of the probability of systemic risk outbreak, but also provides data support for financial institutions to formulate scientific risk response strategies.

[0009] To achieve the above object, the present invention provides the following solutions:

[0010] An AI-based intelligent financial risk monitoring method, comprising:

[0011] Real-time collection of multi-source heterogeneous data in the financial field; the multi-source heterogeneous data includes structured data, semi-structured data and unstructured data;

[0012] Performing data cleaning on the multi-source heterogeneous data to obtain a cleaned multidimensional data set;

[0013] Based on the GNN model and the time series Transformer model, feature extraction and feature fusion are performed on the multidimensional data set to obtain multidimensional feature space and time series features;

[0014] quantifying the multidimensional feature space and the time series features to obtain a risk quantification index;

[0015] Based on the Monte Carlo simulation method, a risk diffusion probability matrix is ​​determined according to the risk quantification indicators;

[0016] Determining the probability value of a systemic risk outbreak based on the risk diffusion probability matrix;

[0017] When the probability value is greater than the set critical value, it is determined that there is a systemic risk and the risk response mechanism is activated.

[0018] Preferably, the structured data includes transaction flows and balance sheets; the semi-structured data includes supply chain relationship networks; and the unstructured data includes financial news and social media sentiment.

[0019] Preferably, performing data cleaning on the multi-source heterogeneous data to obtain a cleaned multidimensional dataset includes:

[0020] Grouping the structured data according to a preset collection period to obtain multiple data groups;

[0021] Calculate the difference coefficient between the current data group and the previous data group in sequence;

[0022] Determine whether the value of the coefficient of difference is within a preset range;

[0023] If the value of the coefficient of difference is not within the preset range, the corresponding data group will be removed;

[0024] If the value of the difference coefficient is within the preset range, the corresponding data group will be retained until all data groups are traversed to obtain structured data after data cleaning;

[0025] Performing data standardization and anomaly detection on the semi-structured data to obtain cleaned semi-structured data;

[0026] Performing text preprocessing and data standardization on the unstructured data to obtain cleaned unstructured data;

[0027] The cleaned multidimensional data set is constructed according to the cleaned structured data, the cleaned semi-structured data and the cleaned unstructured data.

[0028] Preferably, the coefficient of difference calculation formula is:

[0029]

[0030] Among them, p X,Y is the coefficient of difference, cov(X,Y) represents the covariance between the current data set X and the previous data set Y, α X Represents the mean of the current data set X, β Y Represents the mean of the previous data set Y.

[0031] Preferably, the abnormality detection step includes:

[0032] Based on the semi-structured data after data standardization, the relationship strength between each enterprise is calculated;

[0033] It is determined whether the relationship strength is greater than a preset reasonable strength range. If so, the relationship between the corresponding enterprises is removed to obtain semi-structured data after data cleaning.

[0034] Preferably, the calculation formula for the relationship strength is:

[0035] R ij =α·T ij +β·V ij +γ·C ij +δ·S ij

[0036] Among them, R ij is the strength of the relationship between enterprise i and enterprise j. The higher the value, the closer the relationship between the two. ij V is the transaction amount, which represents the total transaction amount between enterprise i and enterprise j. The larger the transaction amount, the closer the relationship. ij is the transaction frequency, which indicates the number of transactions between enterprise i and enterprise j within a certain period of time. The higher the frequency, the more frequent the interaction between the two. ij is the number of cooperative projects, indicating the number of projects jointly participated by enterprise i and enterprise j. The more cooperative projects there are, the closer the relationship between the two is. ij is social media interaction, which indicates the degree of interaction between enterprise i and enterprise j on social media. The more interaction, the more active the relationship between the two. α, β, γ, and δ are all weight coefficients used to adjust the influence of various factors on the relationship strength.

[0037] Preferably, based on the GNN model and the time series Transformer model, feature extraction and feature fusion are performed on the multidimensional dataset to obtain multidimensional feature space and time series features, including:

[0038] Extracting asset features from the transaction flow and the balance sheet to obtain a numerical feature vector; the numerical vector features include a current ratio and an asset-liability ratio;

[0039] Analyzing the supply chain relationship network through a GNN model to extract network feature vectors of nodes and edges in the graph structure of the supply chain relationship network; the network feature vectors include equity chains and guarantee chains;

[0040] Analyzing the financial news and social media sentiment using natural language processing technology to extract sentiment scores and keywords to obtain text feature vectors;

[0041] Performing feature splicing on the numerical vector feature, the network feature vector, and the text feature vector to obtain a fusion feature;

[0042] Input the fused features into the GNN model to capture the relationship and topological structure between features;

[0043] The relationship and topological structure between the features output by the GNN model are input into the time series Transformer model to process the time series data to obtain the multidimensional feature space and the time series features; the multidimensional feature space includes the company's financial indicators, market sentiment, and supply chain relationships; the time series features include historical data of the company and the industry.

[0044] Preferably, the multidimensional feature space and the time series features are quantified to obtain risk quantification indicators, including:

[0045] Performing feature selection on the multidimensional feature space and the time series features to select risk-related features and obtain financial indicators, market sentiment indicators, and historical data indicators;

[0046] Performing feature standardization processing on the financial indicator, the market sentiment indicator, and the historical data indicator to obtain standardized features;

[0047] Performing feature extraction and feature dimensionality reduction on the standardized features based on a principal component analysis method to obtain a feature matrix after dimensionality reduction;

[0048] Defining risk quantification indicators based on the feature matrix after dimensionality reduction; the risk quantification indicators include financial risk indicators and market risk indicators;

[0049] A comprehensive risk score is performed based on the financial risk index and the market risk index to obtain the risk quantification index; the calculation formula of the risk quantification index is: R = w1 × R f +w2×R m ; Wherein, R is the value of the risk quantification indicator, R f and R m are the values ​​of the financial risk indicator and the market risk indicator respectively; w1 and w2 are the preset weight coefficients of the financial risk indicator and the market risk indicator respectively.

[0050] Preferably, based on the Monte Carlo simulation method, determining the risk diffusion probability matrix according to the risk quantification index includes:

[0051] Based on the normal distribution method, random samples are generated according to the risk quantification indicators;

[0052] The risk diffusion probability between enterprises is calculated based on random samples; the calculation formula of the risk diffusion probability is: Among them, P ij represents the risk diffusion probability of enterprise i to enterprise j, R i and R j are the risk quantitative indicators of enterprise i and enterprise j respectively; max(R) is the maximum value of all enterprise risk quantitative indicators;

[0053] The risk diffusion probability of each simulation is accumulated to construct the final risk diffusion probability matrix; the calculation formula of the risk diffusion probability matrix is: Where N is the number of simulations, is the risk diffusion probability of the kth simulation.

[0054] Preferably, the calculation formula of the probability value is:

[0055]

[0056] Among them, P systemic is the probability value of systemic risk outbreak, which means the probability of at least one enterprise experiencing a risk event under a given risk diffusion probability matrix; N' is the total number of enterprises, which means the number of enterprises included in the risk diffusion probability matrix.

[0057] According to the specific embodiments provided by the present invention, the present invention discloses the following technical effects:

[0058] The present invention provides an AI-based intelligent financial risk monitoring method, comprising: real-time collection of multi-source heterogeneous data in the financial field; the multi-source heterogeneous data including structured data, semi-structured data, and unstructured data; data cleaning of the multi-source heterogeneous data to obtain a cleaned multidimensional dataset; feature extraction and feature fusion of the multidimensional dataset based on a GNN model and a time series Transformer model to obtain a multidimensional feature space and time series features; feature quantification of the multidimensional feature space and the time series features to obtain a risk quantification index; determining a risk diffusion probability matrix based on the risk quantification index based on a Monte Carlo simulation method; determining a probability value of a systemic risk outbreak based on the risk diffusion probability matrix; and determining the presence of a systemic risk when the probability value is greater than a set critical value, and initiating a risk response mechanism. By collecting multi-source heterogeneous data in real time, including structured, semi-structured, and unstructured data, the present invention overcomes the data silo phenomenon and promotes information sharing among different financial institutions. This diversified data processing approach enables risk monitoring to comprehensively reflect market dynamics and potential risks, avoiding the reliance on a single data source in traditional methods. In addition, the use of GNN models and time-series Transformer models for feature extraction and fusion enhances the ability to respond to dynamic market changes, making risk assessment more timely and accurate. The present invention calculates the risk diffusion probability matrix through the Monte Carlo simulation method, deeply analyzes the risk transmission mechanism, and clarifies the risk contagion effect between industries. This innovative method not only improves the accuracy of predicting the probability of systemic risk outbreaks, but also provides data support for financial institutions to formulate scientific risk response strategies. At the same time, the automated risk monitoring and response mechanism reduces dependence on manual intervention, improves the flexibility and adaptability of the early warning mechanism, and enables financial institutions to respond to sudden risk events more quickly, thereby reducing the risk of potential losses. BRIEF DESCRIPTION OF THE DRAWINGS

[0059] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0060] Figure 1 A flow chart of a method provided by an embodiment of the present invention;

[0061] Figure 2 A schematic diagram of the system structure provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0062] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0063] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the present invention is further described in detail below with reference to the accompanying drawings and specific embodiments.

[0064] Figure 1 A flow chart of the method provided in the embodiment of the present invention is shown in FIG. Figure 1 As shown, the present invention provides an AI-based intelligent financial risk monitoring method, comprising:

[0065] Step 100: Real-time collection of multi-source heterogeneous data in the financial field; multi-source heterogeneous data includes structured data, semi-structured data, and unstructured data;

[0066] Step 200: performing data cleaning on multi-source heterogeneous data to obtain a cleaned multidimensional dataset;

[0067] Step 300: Based on the GNN model and the time series Transformer model, feature extraction and feature fusion are performed on the multidimensional data set to obtain multidimensional feature space and time series features;

[0068] Step 400: quantify the multidimensional feature space and time series features to obtain risk quantification indicators;

[0069] Step 500: Determine a risk diffusion probability matrix based on risk quantification indicators using a Monte Carlo simulation method;

[0070] Step 600: Determine the probability value of systemic risk outbreak according to the risk diffusion probability matrix;

[0071] Step 700: When the probability value is greater than the set critical value, it is determined that there is a systemic risk and the risk response mechanism is activated.

[0072] Preferably, the structured data includes transaction flows and balance sheets; the semi-structured data includes supply chain relationship networks; and the unstructured data includes financial news and social media sentiment.

[0073] Specifically, step 100 of this embodiment includes:

[0074] This embodiment utilizes API interfaces, web crawlers, and data stream processing technologies to connect to various financial data sources, such as exchanges, banks, social media platforms, and news websites. Through these technologies, the system can automatically capture and collect structured data (such as transaction flows and balance sheets), semi-structured data (such as supply chain relationship networks), and unstructured data (such as financial news and social media sentiment), ensuring real-time and comprehensive data. To handle different types of data, this embodiment allows structured data to be stored directly in a relational database, while semi-structured data (such as supply chain relationship networks in JSON or XML format) is converted through a data parser to facilitate subsequent analysis. For unstructured data, natural language processing (NLP) technology is used to perform text analysis on financial news and social media sentiment, extracting valuable information and sentiment indicators. This process not only improves data availability but also lays the foundation for subsequent risk monitoring and analysis. Finally, by setting up data validation rules and real-time monitoring mechanisms, this embodiment enables the system to promptly identify and process outliers or missing values ​​in the data, preventing data quality issues from affecting subsequent analysis results. In addition, data collection strategies are regularly updated and maintained to adapt to changes in the market environment, ensuring that the system can continuously and effectively obtain the latest multi-source heterogeneous data, thereby providing a solid data foundation for financial risk monitoring.

[0075] Preferably, performing data cleaning on the multi-source heterogeneous data to obtain a cleaned multidimensional dataset includes:

[0076] Grouping the structured data according to a preset collection period to obtain multiple data groups;

[0077] Calculate the difference coefficient between the current data group and the previous data group in sequence;

[0078] Determine whether the value of the coefficient of difference is within a preset range;

[0079] If the value of the coefficient of difference is not within the preset range, the corresponding data group will be removed;

[0080] If the value of the difference coefficient is within the preset range, the corresponding data group will be retained until all data groups are traversed to obtain structured data after data cleaning;

[0081] Performing data standardization and anomaly detection on the semi-structured data to obtain cleaned semi-structured data;

[0082] Performing text preprocessing and data standardization on the unstructured data to obtain cleaned unstructured data;

[0083] The cleaned multidimensional data set is constructed according to the cleaned structured data, the cleaned semi-structured data and the cleaned unstructured data.

[0084] Preferably, the coefficient of difference calculation formula is:

[0085]

[0086] Among them, p X,Y is the coefficient of difference, cov(X,Y) represents the covariance between the current data set X and the previous data set Y, α X Represents the mean of the current data set X, β Y Represents the mean of the previous data set Y.

[0087] Specifically, step 200 of this embodiment includes:

[0088] Step 201: Group structured data according to a preset collection period. This process involves dividing the collected structured data (such as transaction flows and balance sheets) into time periods, such as by day, week, or month. In this way, this embodiment can organize the data into multiple data groups, facilitating subsequent difference analysis and processing.

[0089] Step 202: The coefficient of variation (CDR) between the current data set and the previous data set is calculated. The CDR is an important statistical metric used to measure the degree of change between two data sets. By calculating the covariance between the current and previous data sets and combining their respective means, the system can quantify the difference between the two data sets. This calculation process requires ensuring data integrity and accuracy to obtain a reliable CDR value.

[0090] Step 203: After calculating the coefficient of variation, this embodiment determines whether the value is within a preset range. This preset range can be set based on historical data analysis and business requirements, typically including reasonable upper and lower limits. If the coefficient of variation value exceeds this range, this embodiment will deem the data set to be abnormal, possibly due to data collection errors or market fluctuations. Therefore, this embodiment will remove the corresponding data set to ensure the accuracy of subsequent analysis.

[0091] Step 204: This embodiment retains data sets whose coefficient of variation falls within a preset range until all data sets have been traversed. In this way, the system effectively selects structured data that meets the standards, ensuring the high quality of the resulting dataset. This process not only improves data reliability but also provides a solid foundation for subsequent risk monitoring and analysis.

[0092] Step 205: When processing semi-structured data, the system performs data standardization and anomaly detection. Semi-structured data (such as supply chain relationship networks) is typically stored in JSON or XML format. This embodiment requires parsing and converting this data. Data standardization involves unifying data from different sources into a unified format for subsequent analysis. This embodiment also performs anomaly detection to identify potential data errors or inconsistencies, ensuring the accuracy and consistency of the cleaned semi-structured data.

[0093] Step 206: For unstructured data (such as financial news and social media sentiment), this embodiment performs text preprocessing and data normalization. Text preprocessing includes removing stop words, punctuation, and stemming to extract meaningful keywords and sentiment information. Data normalization converts the processed text data into a unified format to facilitate subsequent analysis and modeling. This process ensures that unstructured data can be effectively integrated into the overall data analysis framework.

[0094] Step 207: A cleansed multidimensional dataset is constructed based on the cleansed structured, semi-structured, and unstructured data. This multidimensional dataset integrates information from various data sources to form a comprehensive view, facilitating subsequent feature extraction and risk monitoring analysis. In this way, this embodiment fully leverages the advantages of multi-source heterogeneous data to improve the accuracy and effectiveness of financial risk monitoring.

[0095] In summary, this embodiment systematically cleanses heterogeneous data from multiple sources, enabling the financial risk monitoring system to ensure high-quality and reliable data, thereby providing a solid foundation for subsequent risk assessment and decision-making. This process not only improves data availability but also enhances the system's responsiveness to market dynamics.

[0096] Preferably, the abnormality detection step includes:

[0097] Based on the semi-structured data after data standardization, the relationship strength between each enterprise is calculated;

[0098] It is determined whether the relationship strength is greater than a preset reasonable strength range. If so, the relationship between the corresponding enterprises is removed to obtain semi-structured data after data cleaning.

[0099] Preferably, the calculation formula for the relationship strength is:

[0100] R ij =α·T ij +β·V ij +γ·C ij +δ·S ij

[0101] Among them, R ij is the strength of the relationship between enterprise i and enterprise j. The higher the value, the closer the relationship between the two. ij V is the transaction amount, which represents the total transaction amount between enterprise i and enterprise j. The larger the transaction amount, the closer the relationship. ij is the transaction frequency, which indicates the number of transactions between enterprise i and enterprise j within a certain period of time. The higher the frequency, the more frequent the interaction between the two. ij is the number of cooperative projects, indicating the number of projects jointly participated by enterprise i and enterprise j. The more cooperative projects there are, the closer the relationship between the two is. ij is social media interaction, which indicates the degree of interaction between enterprise i and enterprise j on social media. The more interaction, the more active the relationship between the two. α, β, γ, and δ are all weight coefficients used to adjust the influence of various factors on the relationship strength.

[0102] Specifically, this embodiment first conducts factor analysis to determine weights to identify key factors influencing the strength of relationships between enterprises. These factors include transaction amount, transaction frequency, number of collaborative projects, and social media interactions. By statistically analyzing historical data, the importance of each factor in enterprise relationships can be assessed. For example, correlation analysis or regression analysis can be used to quantify the relationship between each factor and the strength of enterprise relationships, thereby assigning a preliminary weight coefficient to each factor. This embodiment utilizes data mining techniques and machine learning algorithms to ensure the scientific and rationality of the weight coefficients. To further optimize the weight coefficients, this embodiment also utilizes expert evaluation and multivariate decision-making methods. Industry experts and data analysts are invited to review and adjust the initially determined weight coefficients to ensure they align with actual business scenarios and market dynamics. Furthermore, this embodiment can also utilize multivariate decision-making tools such as the Analytic Hierarchy Process (AHP), combined with expert opinion and historical data, to comprehensively evaluate and optimize the weights. Ultimately, the determined weight coefficients are used to calculate the strength of relationships between enterprises, effectively identifying and removing unreasonable enterprise relationships during anomaly detection, thereby improving the quality and reliability of the cleaned semi-structured data.

[0103] For example, the optimal weight distribution determined in this embodiment is as follows:

[0104] Transaction amount: 40%

[0105] The transaction amount is usually an important indicator to measure the strength of the relationship between enterprises. The larger the amount, the closer the economic connection between the two.

[0106] Transaction frequency: 30%

[0107] Transaction frequency reflects the degree of interaction between enterprises. The higher the frequency, the more active the cooperative relationship between the two.

[0108] Number of cooperative projects: 20%

[0109] The number of jointly participated projects can reflect the long-term cooperative relationship between enterprises. The more projects, the closer the relationship.

[0110] Social media interactions: 10%

[0111] While the level of interaction on social media is important, it may have less influence than economic transactions and collaborative projects, and therefore has a relatively low weight.

[0112] Preferably, based on the GNN model and the time series Transformer model, feature extraction and feature fusion are performed on the multidimensional dataset to obtain multidimensional feature space and time series features, including:

[0113] Extracting asset features from the transaction flow and the balance sheet to obtain a numerical feature vector; the numerical vector features include a current ratio and an asset-liability ratio;

[0114] Analyzing the supply chain relationship network through a GNN model to extract network feature vectors of nodes and edges in the graph structure of the supply chain relationship network; the network feature vectors include equity chains and guarantee chains;

[0115] Analyzing the financial news and social media sentiment using natural language processing technology to extract sentiment scores and keywords to obtain text feature vectors;

[0116] Performing feature splicing on the numerical vector feature, the network feature vector, and the text feature vector to obtain a fusion feature;

[0117] Input the fused features into the GNN model to capture the relationship and topological structure between features;

[0118] The relationship and topological structure between the features output by the GNN model are input into the time series Transformer model to process the time series data to obtain the multidimensional feature space and the time series features; the multidimensional feature space includes the company's financial indicators, market sentiment, and supply chain relationships; the time series features include historical data of the company and the industry.

[0119] Specifically, step 300 of this embodiment includes:

[0120] Step 301: Extract asset features from the transaction flow and balance sheet to generate a numerical feature vector. This process in this embodiment involves calculating key financial indicators such as the current ratio and debt-to-asset ratio. The current ratio reflects a company's short-term debt repayment ability, while the debt-to-asset ratio measures its level of financial leverage. By calculating and standardizing these indicators, a structured numerical feature vector is generated, which serves as the basis for subsequent analysis.

[0121] Step 302: Graph Neural Network (GNN) models are used to analyze the supply chain relationship network to extract feature vectors for nodes and edges within the network. GNNs effectively capture relationships and patterns within graph-structured data. By analyzing equity and guarantee chains, they can identify potential risks and partnerships between companies. These network feature vectors provide crucial information for subsequent feature fusion, helping to understand the position and influence of companies within the supply chain.

[0122] Step 303: When processing unstructured data, this embodiment applies natural language processing (NLP) technology to analyze financial news and social media sentiment to extract sentiment scores and keywords. Sentiment scores reflect market sentiment toward a company, while keywords provide important information related to the company. By cleaning and processing text data, text feature vectors are generated. These vectors are used in feature fusion along with numerical and network feature vectors.

[0123] Step 304: The numerical feature vector, network feature vector, and text feature vector are concatenated to form a comprehensive fused feature vector. This fused feature vector not only contains information about the company's financial indicators, market sentiment, and supply chain relationships, but also reflects the relationships between different features, providing rich contextual information for subsequent model input.

[0124] Step 305: This embodiment inputs the fused features into the GNN model to capture the relationships and topological structure between features. Leveraging its graph structure, the GNN model effectively learns relationships between nodes and extracts deeper feature representations. This process helps the system understand enterprise behavior patterns and potential risks within complex networks, providing more accurate evidence for risk monitoring.

[0125] Step 306: The feature relationships and topological structures output by the GNN model are input into the Time Series Transformer model to process the time series data. The Time Series Transformer can capture dynamic changes and trends in time series and, combined with historical data, generate a multidimensional feature space and time series features. These features will include the dynamic changes in a company's financial indicators, market sentiment, and supply chain relationships, providing comprehensive data support for subsequent risk assessment and decision-making. Through this series of steps, this embodiment can achieve in-depth analysis and feature extraction of multidimensional datasets, providing strong technical support for financial risk monitoring.

[0126] Preferably, the multidimensional feature space and the time series features are quantified to obtain risk quantification indicators, including:

[0127] Performing feature selection on the multidimensional feature space and the time series features to select risk-related features and obtain financial indicators, market sentiment indicators, and historical data indicators;

[0128] Performing feature standardization processing on the financial indicator, the market sentiment indicator, and the historical data indicator to obtain standardized features;

[0129] Performing feature extraction and feature dimensionality reduction on the standardized features based on a principal component analysis method to obtain a feature matrix after dimensionality reduction;

[0130] Defining risk quantification indicators based on the feature matrix after dimensionality reduction; the risk quantification indicators include financial risk indicators and market risk indicators;

[0131] A comprehensive risk score is performed based on the financial risk index and the market risk index to obtain the risk quantification index; the calculation formula of the risk quantification index is: R = w1 × R f +w2×R m ; Wherein, R is the value of the risk quantification indicator, R f and R m are the values ​​of the financial risk indicator and the market risk indicator respectively; w1 and w2 are the preset weight coefficients of the financial risk indicator and the market risk indicator respectively.

[0132] Specifically, step 400 of this embodiment includes:

[0133] Step 401: This embodiment first performs feature selection to identify key risk-related features. This process involves analyzing financial indicators, market sentiment indicators, and historical data indicators to select features that have a significant impact on risk assessment. For example, indicators such as current ratio, debt-to-asset ratio, sentiment score, and transaction frequency may be important risk-related features. This approach allows us to focus on the most representative features, laying the foundation for subsequent quantitative analysis.

[0134] Step 402: Standardize the selected financial indicators, market sentiment indicators, and historical data indicators to eliminate dimensional differences between different features. In this embodiment, the standardization process involves converting the value of each feature into a distribution with a mean of 0 and a standard deviation of 1. This process ensures the comparability of each feature in subsequent analysis, allowing data from different sources to be effectively compared on the same scale, thereby improving the accuracy of risk quantification.

[0135] Step 403: After standardization, this embodiment applies principal component analysis (PCA) to extract and reduce the dimensionality of the standardized features. PCA can map high-dimensional data to a low-dimensional space through linear transformation while preserving the data's primary information as much as possible. This process not only reduces the dimensionality of the features and computational complexity, but also eliminates multicollinearity between features, thereby improving the model's stability and predictive power. Ultimately, this embodiment produces a reduced-dimensional feature matrix, which serves as the basis for risk quantification indicators.

[0136] Step 404: Based on the reduced feature matrix, financial risk indicators and market risk indicators are defined. These indicators reflect the risk level of the enterprise in terms of its financial health and market performance. Financial risk indicators include the current ratio and debt-to-asset ratio, while market risk indicators include market sentiment scores and stock price volatility. Through comprehensive analysis of these indicators, a comprehensive risk quantification indicator can be generated for each enterprise.

[0137] For example, when determining the weight coefficients of financial risk indicators and market risk indicators, this embodiment adopts an expert evaluation and data-driven approach. First, through interviews with industry experts and questionnaires, subjective evaluations of the importance of each indicator are collected to form a preliminary weight distribution. Secondly, combined with historical data analysis, statistical methods (such as regression analysis) are used to verify and adjust these weights. Finally, assuming that the weight of the financial risk indicator is 0.6 and the weight of the market risk indicator is 0.4, such a distribution reflects the importance of financial health in the overall risk assessment, while also taking into account the impact of market sentiment. In this way, the system can ensure the scientificity and rationality of risk quantification indicators and provide a reliable basis for subsequent risk management.

[0138] Preferably, based on the Monte Carlo simulation method, determining the risk diffusion probability matrix according to the risk quantification index includes:

[0139] Based on the normal distribution method, random samples are generated according to the risk quantification indicators;

[0140] The risk diffusion probability between enterprises is calculated based on random samples; the calculation formula of the risk diffusion probability is: Among them, P ij represents the risk diffusion probability of enterprise i to enterprise j, R i and R j are the risk quantitative indicators of enterprise i and enterprise j respectively; max(R) is the maximum value of all enterprise risk quantitative indicators;

[0141] The risk diffusion probability of each simulation is accumulated to construct the final risk diffusion probability matrix; the calculation formula of the risk diffusion probability matrix is: Where N is the number of simulations, is the risk diffusion probability of the kth simulation.

[0142] Specifically, step 500 of this embodiment includes:

[0143] Step 501: Generate a set of random samples based on the mean and standard deviation of the risk quantification indicators. These random samples will reflect the risk performance of enterprises under different market conditions, ensuring the diversity and authenticity of the simulation results. In this way, this embodiment can provide basic data for subsequent risk diffusion probability calculations and simulate risk transfer between enterprises under different scenarios.

[0144] Step 502: The generated random samples are used to calculate the risk diffusion probability between enterprises. Specifically, this embodiment calculates the enterprise-to-enterprise risk diffusion probability based on the quantitative risk indicators of each pair of enterprises and the random samples. This calculation takes into account the mutual influence and risk transmission mechanisms between enterprises. After each simulation, the system accumulates the calculated risk diffusion probabilities to ultimately construct a complete risk diffusion probability matrix. This matrix provides an important basis for subsequent risk assessment and management, helping financial institutions identify potential systemic risks and inter-industry risk contagion effects. Through multiple simulations, this embodiment can obtain more accurate and reliable risk diffusion probabilities, thereby providing scientific support for decision-making.

[0145] Preferably, the calculation formula of the probability value is:

[0146]

[0147] Among them, P systemicis the probability value of systemic risk outbreak, which means the probability of at least one enterprise experiencing a risk event under a given risk diffusion probability matrix; N' is the total number of enterprises, which means the number of enterprises included in the risk diffusion probability matrix.

[0148] Specifically, in step 600 of this embodiment, a risk diffusion probability matrix is ​​analyzed. This matrix contains the risk diffusion probabilities between enterprises and reflects the likelihood of risk transmission between different enterprises. To calculate the probability of a systemic risk outbreak, this embodiment first assesses the independent probability of a risk event occurring at each enterprise. This process involves analyzing each element in the risk diffusion probability matrix to identify which enterprises are more susceptible to risk events under specific conditions. This embodiment then uses these independent probabilities to calculate the overall probability of at least one enterprise experiencing a risk event. Specifically, this embodiment considers the risk diffusion probabilities of all enterprises and, using logical reasoning and basic principles of probability theory, calculates the probability of at least one enterprise experiencing a risk event given the given risk diffusion probability matrix. This calculation process may involve accumulating the risk event probabilities for each enterprise, while also taking into account the mutual influence and risk transmission effects between enterprises to ensure the accuracy of the final result. After completing these calculations, this embodiment will obtain a comprehensive systemic risk outbreak probability value, which provides important decision-making basis for financial institutions. By analyzing this probability value, financial institutions can identify potential systemic risks and implement appropriate risk management measures to reduce the overall risk level. Furthermore, the calculation of the probability of systemic risk outbreak can provide a reference for policymakers, helping them formulate more effective regulatory policies to maintain financial market stability. Through this series of steps, this embodiment can effectively assess and monitor systemic risk, providing scientific support for financial risk management.

[0149] Optionally, step 700 of this embodiment includes:

[0150] Step 701: When the calculated probability of a systemic risk outbreak exceeds a set threshold, this embodiment automatically determines that a systemic risk exists. This threshold is set based on historical data analysis, industry standards, and expert opinion, aiming to identify potential risk levels. By reviewing and analyzing historical risk events, financial institutions can determine a reasonable threshold so they can take timely action when the risk probability reaches that threshold. This process ensures the foresight and effectiveness of risk management, enabling financial institutions to prepare for risks before they materialize.

[0151] Step 702: Once a systemic risk is determined, this embodiment immediately activates the risk response mechanism. This mechanism involves multiple aspects, starting with the rapid assessment and response to risk events. Financial institutions will establish a dedicated risk management team responsible for conducting in-depth analysis of current market conditions and assessing the potential impact of risk events. This team will utilize real-time data and analytical tools to rapidly capture market dynamics and ensure timely understanding of risk trends and changes.

[0152] Step 703: After the risk response mechanism is activated, financial institutions will develop corresponding response strategies. These strategies may include adjusting investment portfolios, increasing liquidity reserves, and strengthening risk monitoring and information disclosure. Through these strategies, financial institutions can effectively mitigate potential losses and protect the interests of investors. Furthermore, institutions may communicate with regulators and other market participants to share risk information to jointly address systemic risks. This collaboration not only helps reduce risks for individual institutions but also enhances the stability of the entire financial system.

[0153] Step 704: Finally, this embodiment conducts subsequent monitoring and evaluation of the risk response mechanism. Financial institutions should continuously track market dynamics and risk indicators to ensure timely adjustments to response strategies as risk events develop. Furthermore, institutions should evaluate the effectiveness of their risk response mechanisms, analyzing lessons learned to optimize and improve future risk management. This cyclical process will help enhance financial institutions' risk management capabilities and strengthen their resilience to systemic risks, thereby maintaining stable and sustainable development in a complex and volatile market environment.

[0154] Corresponding to the above method, such as Figure 2 As shown, this embodiment also provides an AI-based intelligent financial risk monitoring method, including:

[0155] Data collection unit, used to collect multi-source heterogeneous data in the financial field in real time; multi-source heterogeneous data includes structured data, semi-structured data and unstructured data;

[0156] A data cleaning unit is used to clean multi-source heterogeneous data to obtain a cleaned multidimensional data set;

[0157] The feature processing unit is used to extract and fuse features of multidimensional data sets based on the GNN model and the time series Transformer model to obtain multidimensional feature space and time series features;

[0158] Feature quantification unit, used to quantify multi-dimensional feature space and time series features to obtain risk quantification indicators;

[0159] A matrix construction unit is used to determine the risk diffusion probability matrix based on risk quantification indicators based on the Monte Carlo simulation method;

[0160] A probability determination unit, used to determine the probability value of a systemic risk outbreak based on a risk diffusion probability matrix;

[0161] The risk determination unit is used to determine that there is a systemic risk when the probability value is greater than the set critical value and to activate the risk response mechanism.

[0162] The beneficial effects of the present invention are as follows:

[0163] (1) The present invention effectively solves several technical problems faced by traditional financial risk monitoring systems. First, by collecting multi-source heterogeneous data in real time, including structured, semi-structured, and unstructured data, the data island phenomenon is overcome and information sharing between different financial institutions is promoted. This diversified data processing method enables risk monitoring to fully reflect market dynamics and potential risks, avoiding the reliance on a single data source in traditional methods. In addition, the use of GNN models and time series Transformer models for feature extraction and fusion enhances the ability to respond to dynamic market changes, making risk assessment more timely and accurate.

[0164] (2) This invention uses the Monte Carlo simulation method to calculate the risk diffusion probability matrix, deeply analyzes the risk transmission mechanism, and clarifies the risk contagion effect between industries. This innovative method not only improves the accuracy of predicting the probability of systemic risk outbreaks, but also provides data support for financial institutions to formulate scientific risk response strategies. At the same time, the automated risk monitoring and response mechanism reduces reliance on manual intervention, improves the flexibility and adaptability of the early warning mechanism, and enables financial institutions to respond to sudden risk events more quickly, thereby reducing the risk of potential losses.

[0165] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. Reference can be made to the common and similar parts between the various embodiments. For the systems disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple, and the relevant parts can be referred to the method description.

[0166] This document uses specific examples to illustrate the principles and implementation methods of the present invention. The above examples are only intended to help understand the method and core concept of the present invention. At the same time, those skilled in the art will find that the specific implementation methods and application scopes may vary based on the concept of the present invention. In summary, the contents of this specification should not be construed as limiting the present invention.

Claims

1. An AI-based intelligent financial risk monitoring method, characterized in that: include: Real-time collection of multi-source heterogeneous data in the financial field; The multi-source heterogeneous data includes structured data, semi-structured data and unstructured data; Performing data cleaning on the multi-source heterogeneous data to obtain a cleaned multidimensional data set; Based on the GNN model and the time series Transformer model, feature extraction and feature fusion are performed on the multidimensional data set to obtain multidimensional feature space and time series features; quantifying the multidimensional feature space and the time series features to obtain a risk quantification index; Based on the Monte Carlo simulation method, a risk diffusion probability matrix is ​​determined according to the risk quantification indicators; Determining the probability value of a systemic risk outbreak based on the risk diffusion probability matrix; When the probability value is greater than the set critical value, it is determined that there is a systemic risk and the risk response mechanism is activated.

2. The AI-based intelligent financial risk monitoring method according to claim 1, characterized in that: The structured data includes transaction flows and balance sheets; the semi-structured data includes supply chain relationship networks; and the unstructured data includes financial news and social media sentiment.

3. The AI-based intelligent financial risk monitoring method according to claim 2, characterized in that: Performing data cleaning on the multi-source heterogeneous data to obtain a cleaned multidimensional dataset includes: Grouping the structured data according to a preset collection period to obtain multiple data groups; Calculate the difference coefficient between the current data group and the previous data group in sequence; Determine whether the value of the coefficient of difference is within a preset range; If the value of the coefficient of difference is not within the preset range, the corresponding data group will be removed; If the value of the difference coefficient is within the preset range, the corresponding data group will be retained until all data groups are traversed to obtain structured data after data cleaning; Performing data standardization and anomaly detection on the semi-structured data to obtain cleaned semi-structured data; Performing text preprocessing and data standardization on the unstructured data to obtain cleaned unstructured data; The cleaned multidimensional data set is constructed according to the cleaned structured data, the cleaned semi-structured data and the cleaned unstructured data.

4. The AI-based intelligent financial risk monitoring method according to claim 3, characterized in that: The coefficient of difference calculation formula is: Among them, p X,Y is the coefficient of difference, cov(X,Y) represents the covariance between the current data set X and the previous data set Y, α X Represents the mean of the current data set X, β Y Represents the mean of the previous data set Y.

5. The AI-based intelligent financial risk monitoring method according to claim 3, characterized in that: The steps of anomaly detection include: Based on the semi-structured data after data standardization, the relationship strength between each enterprise is calculated; It is determined whether the relationship strength is greater than a preset reasonable strength range. If so, the relationship between the corresponding enterprises is removed to obtain semi-structured data after data cleaning.

6. The AI-based intelligent financial risk monitoring method according to claim 5, characterized in that: The calculation formula for the relationship strength is: R ij =α·T ij +β·V ij +γ·C ij +δ·S ij Among them, R ij is the strength of the relationship between enterprise i and enterprise j. The higher the value, the closer the relationship between the two. ij V is the transaction amount, which represents the total transaction amount between enterprise i and enterprise j. The larger the transaction amount, the closer the relationship. ij is the transaction frequency, which indicates the number of transactions between enterprise i and enterprise j within a certain period of time. The higher the frequency, the more frequent the interaction between the two. ij is the number of cooperative projects, indicating the number of projects jointly participated by enterprise i and enterprise j. The more cooperative projects there are, the closer the relationship between the two is. ij is social media interaction, which indicates the degree of interaction between enterprise i and enterprise j on social media. The more interaction, the more active the relationship between the two. α, β, γ, and δ are all weight coefficients used to adjust the influence of various factors on the relationship strength.

7. The AI-based intelligent financial risk monitoring method according to claim 2, characterized in that: Based on the GNN model and the time series Transformer model, feature extraction and feature fusion are performed on the multidimensional dataset to obtain multidimensional feature space and time series features, including: Extracting asset features from the transaction flow and the balance sheet to obtain a numerical feature vector; the numerical vector features include a current ratio and an asset-liability ratio; Analyzing the supply chain relationship network through a GNN model to extract network feature vectors of nodes and edges in the graph structure of the supply chain relationship network; the network feature vectors include equity chains and guarantee chains; Analyzing the financial news and social media sentiment using natural language processing technology to extract sentiment scores and keywords to obtain text feature vectors; Performing feature splicing on the numerical vector feature, the network feature vector, and the text feature vector to obtain a fusion feature; Input the fused features into the GNN model to capture the relationship and topological structure between features; The relationship and topological structure between the features output by the GNN model are input into the time series Transformer model to process the time series data to obtain the multidimensional feature space and the time series features; the multidimensional feature space includes the company's financial indicators, market sentiment, and supply chain relationships; the time series features include historical data of the company and the industry.

8. The AI-based intelligent financial risk monitoring method according to claim 1, characterized in that: The multidimensional feature space and the time series features are quantified to obtain risk quantification indicators, including: Performing feature selection on the multidimensional feature space and the time series features to select risk-related features and obtain financial indicators, market sentiment indicators, and historical data indicators; Performing feature standardization processing on the financial indicator, the market sentiment indicator, and the historical data indicator to obtain standardized features; Performing feature extraction and feature dimensionality reduction on the standardized features based on a principal component analysis method to obtain a feature matrix after dimensionality reduction; Defining risk quantification indicators based on the feature matrix after dimensionality reduction; the risk quantification indicators include financial risk indicators and market risk indicators; A comprehensive risk score is performed based on the financial risk index and the market risk index to obtain the risk quantification index; the calculation formula of the risk quantification index is: R = w1 × R f +w2×R m ; Wherein, R is the value of the risk quantification indicator, R f and R m are the values ​​of the financial risk indicator and the market risk indicator respectively; w1 and w2 are the preset weight coefficients of the financial risk indicator and the market risk indicator respectively.

9. The AI-based intelligent financial risk monitoring method according to claim 1, characterized in that: Based on the Monte Carlo simulation method, the risk diffusion probability matrix is ​​determined according to the risk quantification indicators, including: Based on the normal distribution method, random samples are generated according to the risk quantification indicators; The risk diffusion probability between enterprises is calculated based on random samples; the calculation formula of the risk diffusion probability is: Among them, P ij represents the risk diffusion probability of enterprise i to enterprise j, R i and R j are the risk quantitative indicators of enterprise i and enterprise j respectively; max(R) is the maximum value of all enterprise risk quantitative indicators; The risk diffusion probability of each simulation is accumulated to construct the final risk diffusion probability matrix; the calculation formula of the risk diffusion probability matrix is: Where N is the number of simulations, is the risk diffusion probability of the kth simulation.

10. The AI-based intelligent financial risk monitoring method according to claim 9, characterized in that: The calculation formula of the probability value is: Among them, P systemic is the probability value of systemic risk outbreak, which means the probability of at least one enterprise experiencing a risk event under a given risk diffusion probability matrix; N' is the total number of enterprises, which means the number of enterprises included in the risk diffusion probability matrix.