Financial data analysis and statistical model construction method

By building a financial data analysis model, integrating multi-source data and designing an event-driven update mechanism, the problem of untimely multi-modal data fusion and knowledge graph update in financial data analysis is solved, and the accuracy of risk assessment and decision transparency are improved.

CN120598682APending Publication Date: 2025-09-05ZHONGNAN UNIVERSITY OF ECONOMICS AND LAW

Patent Information

Application Number
CN202510729406.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-03
Publication Date
2025-09-05

AI Technical Summary

Technical Problem

Existing financial data analysis methods are difficult to effectively integrate multimodal data sources, and the knowledge graph is not updated in time, resulting in the analysis results lagging behind the actual situation and lack of interpretability.

Method used

Through multiple channels, financial transaction flows, public opinion texts and industrial chain data are collected, initial knowledge graphs are built, and event-driven update mechanisms are designed, data changes are continuously monitored, map reorganization is automatically triggered, and model interpretability is enhanced by combining feature importance analysis and decision-making path tracking.

Benefits of technology

It achieves dynamic integration and timely updating of multi-source data, improves the accuracy of risk assessment and the transparency of decision-making, reduces risk misjudgment, and enhances the decision-making support capabilities of financial institutions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120598682A_ABST
    Figure CN120598682A_ABST
Patent Text Reader

Abstract

The invention discloses a financial data analysis and statistical model construction method, and relates to the technical field of financial big data analysis, and the method comprises the steps: collecting a customer transaction flow, a public opinion text and industrial chain data from a transaction system, a social media platform and an industrial chain database of a financial institution; performing cleaning, noise data removal and standardization processing on the collected data, and extracting key features; and constructing an initial knowledge graph according to the extracted information, continuously monitoring the time sequence change of the data, and continuously updating the knowledge graph. Through comprehensive fusion of multi-source data and deep mining of potential information, more accurate risk assessment and timely discovery of potential associations and risks can be realized, risk misjudgment and event-driven adaptive updating caused by a single data source are reduced, and the timeliness of the atlas can closely follow the market change, so that the accuracy and timeliness of decision making are improved, and the user experience is improved. And a three-dimensional attribution analysis system is constructed, so that the decision transparency can be improved, and the credibility and fairness of financial data analysis are enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field related to financial data analysis, and in particular to a method for financial data analysis and statistical model construction. Background Art

[0002] In the financial sector, particularly in the banking and insurance industries, data sources are extensive and diverse. Customer transaction flow data contains a wealth of information, such as transaction amounts, times, and recipients. This data reflects customers' spending behavior, capital flows, and potential risk appetite. Public opinion text data, obtained from social media, news, and other channels, reflects market participants' views and sentiment on financial institutions, policy changes, and the macroeconomic environment. Industry chain data, encompassing information on multiple links, including suppliers, manufacturers, and distributors, is crucial for assessing industry trends, corporate competitiveness, and supply chain risks.

[0003] Traditional financial data analysis methods mainly focus on statistical analysis of structured data. However, it is difficult to effectively integrate different types of data sources and cannot fully reflect the true situation of the market. The lack of explainability of the decision-making process makes the analysis results difficult to understand and trust.

[0004] A knowledge graph is a graphical structure used to represent and store knowledge, consisting of nodes and edges. Nodes represent entities, such as customers, financial institutions, and products, while edges represent relationships between entities, such as transaction relationships, ownership relationships, and partnerships. By building a knowledge graph, dispersed data can be presented in a more intuitive and clear manner, facilitating knowledge reasoning and analysis.

[0005] However, most existing financial knowledge graphs are static and cannot reflect market changes and data updates in a timely manner, resulting in analysis results lagging behind actual conditions.

[0006] Therefore, it is necessary to propose a financial data analysis and statistical model construction method to solve the above problems. Summary of the Invention

[0007] In response to the shortcomings of the existing technology, the present invention provides a method for financial data analysis and statistical model construction, which solves the problems of difficulty in multimodal data fusion and untimely dynamic updating of knowledge graphs in existing financial data analysis methods.

[0008] To achieve the above objectives, the present invention is implemented through the following technical solutions:

[0009] A method for financial data analysis and statistical model construction, comprising the following steps:

[0010] Step 1: Data Collection: Collect customer transaction flows, public opinion texts, and industry chain data from multiple channels, including financial institutions’ transaction systems, social media platforms, and industry chain databases.

[0011] Step 2: Data Processing and Feature Extraction: Clean the collected data to remove noise, including outliers in transaction flows, advertisements and irrelevant information in public opinion texts, and standardize it for subsequent analysis. Extract key features from transaction flows, public opinion texts, and industry chain data from the processed data. These features are used to reflect the temporal relationships, trend changes, and anomalies of the data.

[0012] Step 3: Construct an initial knowledge graph: Based on the extracted information, construct an initial knowledge graph. Taking the customer transaction network of a bank as an example, customers are used as nodes, and transaction relationships between customers are used as edges. The weight of the edges is determined by the transaction amount and frequency. At the same time, the evaluation information related to customers in the public opinion text is associated with the corresponding customer nodes, and the upstream and downstream cooperation relationships related to enterprises in the industrial chain data are associated with the enterprise nodes.

[0013] Step 4: Knowledge graph update: Continuously monitor the time series changes of the data and continuously update the knowledge graph. When it is found that a customer's recent transaction frequency has increased significantly and there are frequent financial transactions with certain high-risk accounts, update the customer's risk management attributes in the knowledge graph and adjust the associated edge relationships.

[0014] Optionally, the data collection described in step one includes a data input layer for obtaining customer transaction flows, public opinion texts and industry chain data; the transaction flow data: includes transaction time, transaction amount, and transaction object information; the public opinion text data: collects public opinion texts within a specific time period from various data sources (such as social media platforms, news websites, forums, etc.) to ensure the integrity and accuracy of the data, and the texts include comments, posts, and articles; the industry chain data: includes upstream and downstream relationships in the industry and changes in market share.

[0015] Optionally, the data processing and feature extraction described in step 2 include: transaction flow data processing, public opinion text data processing and industrial chain data processing; transaction flow data processing: discretize the transaction time, and perform segmented statistics on transaction frequency and average transaction amount according to time units of hours and days to calculate the transaction amount ratio between transaction objects; public opinion text data processing: preprocess the collected text, remove irrelevant information (such as HTML tags, special characters, stop words), and perform word segmentation operations to divide the text into meaningful words or phrases for subsequent sentiment analysis. For the sentiment tendency score, calculate the average sentiment tendency over a period of time, add up the sentiment tendency scores of all public opinion texts in the time period, and obtain the total sentiment tendency score; divide the total sentiment tendency score by the number of public opinion texts in the time period to obtain the average sentiment tendency score; industrial chain data processing: for the upstream and downstream relationships of the industry, construct an adjacency matrix.

[0016] Optionally, the construction of the initial knowledge graph described in step 3 includes the following steps:

[0017] S1: Node generation: Based on the processed data, different types of nodes are generated, including customer nodes, financial product / service nodes, enterprise nodes, and hot topic nodes. The attributes of the nodes are set according to the corresponding data characteristics. The attributes of customer nodes include age and asset size, and the attributes of enterprise nodes include industry type and registered capital.

[0018] S2: Edge generation and weight calculation: Generate relationship edges between customers and financial products / services, and between customers based on transaction flow data;

[0019] S3: Public opinion association edges: Generate relationship edges between customers and hot topics, and between companies and hot topics from public opinion text data;

[0020] S4. Industry chain relationship edges: Generate upstream and downstream relationship edges between enterprises based on industry chain data. The weight of the edges is determined based on market share changes and transaction amount factors.

[0021] Optionally, the knowledge graph described in step four is updated as follows: set an update threshold. When the market volatility exceeds the threshold (defined by the change in the market volatility index), rerun the data processing and feature extraction process to update the nodes, edges and their attributes in the knowledge graph. The updated node attributes are modified according to the new data processing results, and the existence and weight of the edges are updated based on the new relationship calculation.

[0022] Optionally, the knowledge graph update described in step 4 also includes an event-driven update mechanism for monitoring financial market fluctuations. When market fluctuations exceed a set threshold, the graph structure reorganization is automatically triggered. The event-driven update mechanism includes market data collection, event monitoring and identification, event impact analysis, graph structure reorganization, and topology optimization.

[0023] Market data collection: Acquire real-time market data from multiple financial data sources, including stock trading markets, futures trading markets, and foreign exchange markets, including key indicators such as price and trading volume. Clean, organize, and standardize the data collected from different sources for subsequent analysis and use. Fill in or correct missing or abnormal data to ensure data quality and consistency.

[0024] Event monitoring and identification: Set monitoring indicators for market fluctuations, such as stock price volatility (measured by calculating the standard deviation of prices over a certain period of time), calculate market volatility indicators (such as volatility), and compare them with preset thresholds to determine whether a market volatility event is triggered.

[0025] Optionally, the event impact analysis: for different types of events (market fluctuations), impact assessment models are constructed separately. For market fluctuation events, a regression analysis model based on historical data is considered to analyze the impact of market fluctuations on different financial entities (such as individual stocks and industry indices);

[0026] The graph structure reorganization: for entities affected by the event, their attributes are updated according to the results of the event impact analysis, and the relationships between entities are adjusted based on the new entity attributes and the event impact analysis results. For newly emerging entity relationships, they are added according to the event information and entity attribute changes;

[0027] The topology structure optimization is as follows: after completing the entity attribute update and relationship adjustment, the topology structure of the knowledge graph is optimized, and the entities in the knowledge graph are divided into communities using a community discovery algorithm, so that entities in the same community have closer relationships, which is convenient for subsequent analysis and query. The modularity-optimized community discovery algorithm is used, and its goal is to maximize the difference between the edge density within the community and the edge density of the entire graph.

[0028] Optionally, the model building method further includes an interpretability enhancement module for a three-dimensional attribution analysis system, which analyzes the financial model from three dimensions: feature importance (global), decision path (local), and counterfactual explanation (causal); including feature importance analysis, decision path tracking, and counterfactual explanation application;

[0029] The feature importance analysis includes the following steps:

[0030] Data preprocessing: Collect and organize data related to financial models, including various characteristic variables such as the insured's age, health status, and occupation, as well as the corresponding insurance premium target variables. Clean and standardize the data to ensure data quality and consistency.

[0031] Select evaluation metrics: Based on the characteristics and objectives of the problem, select information gain as the metric to measure feature importance. The importance of a feature is determined by calculating the change in information entropy before and after the feature is partitioned into the dataset.

[0032] Calculate feature importance: Use the selected evaluation metric to calculate the contribution of each feature to the prediction result. For the insurance pricing model, the decision tree algorithm is used to calculate the information gain.

[0033] Result visualization and analysis: The calculated feature importance results are visualized and a bar chart or heat map is drawn to intuitively observe the importance ranking of each feature, analyze key features and their impact on pricing decisions, and provide a basis for the interpretation and optimization of insurance pricing models.

[0034] Optionally, the decision path tracking includes the following steps:

[0035] Sample selection and data preparation: Select sample data of individual loan applicants, including information on their income level, credit history, and collateral value, to ensure data accuracy and completeness;

[0036] Application of model interpretation tools: Use the LIME (Local Interpretable Model-Sensitive Explanations) tool to track the model's decision-making process. By generating a large number of perturbation samples near the sample point, the difference between the prediction results of the perturbed samples and the original samples is calculated, and the impact of each feature on the decision is determined based on the size of the difference.

[0037] Decision path visualization and analysis: Visualize the model's decision-making process for a single sample by drawing a decision tree or flowchart. By analyzing the decision path, we can understand how the model gradually makes approval or rejection decisions based on the values ​​of different features.

[0038] Optionally, the counterfactual explanation application includes the following steps:

[0039] Case selection and hypothesis setting: Select a case of insurance claim rejection as the research object, and set different hypothetical situations for this case, such as changing the insured's disease type or treatment method;

[0040] Model recalculation and comparative analysis: Input the hypothetical data into the insurance claims model for recalculation to obtain the model's claims decision results under the hypothetical circumstances, and then compare and analyze the actual situation with the decision results under the hypothetical circumstances.

[0041] Causal relationship interpretation and summary: Through comparative analysis, we gain an in-depth understanding of the causal relationships behind model decisions. We use causal inference methods, such as matching and inverse probability weighting, to quantify the impact of assumptions on claims decisions. Ultimately, we summarize the causal relationships between different assumptions and claims decisions, providing a reference for optimizing insurance claims models and controlling risks.

[0042] The present invention provides a method for financial data analysis and statistical model construction, which has the following beneficial effects:

[0043] 1. The present invention converts customer transaction flows, public opinion texts, and industrial chain data into dynamic graphs through a temporal relationship reasoning algorithm, which can integrate information from different dimensions. In loan risk assessment, it not only considers the repayment ability reflected by the customer's transaction flow, but also combines information about the development trends of the customer's industry in the public opinion text and the operating conditions of upstream and downstream companies in the industrial chain data to more comprehensively assess loan risks and reduce risk misjudgments caused by a single data source.

[0044] 2. The present invention designs an event-driven update mechanism, which enables the knowledge graph to automatically trigger the reorganization of the graph structure when the market fluctuates or policies are released. When macroeconomic policy adjustments lead to interest rate changes, the knowledge graph can be updated in a timely manner to reflect the changes in the risk-return relationship of different financial products under different interest rate environments, providing financial institutions with the latest decision-making basis so that they can quickly adjust their investment strategies.

[0045] 3. Decision support based on feature importance analysis of the present invention Through feature importance analysis, the key factors and their weights that affect financial decisions are determined from a global perspective. In insurance pricing, insurance companies can clearly understand the impact of factors such as customer age, gender, health status, and occupation on insurance premiums, thereby formulating more reasonable and fair insurance premiums. At the same time, the pricing basis can be explained to customers to improve their understanding and acceptance of pricing.

[0046] 4. The decision path analysis of the present invention can record the decision-making process of a single sample in detail. During the bank loan approval process, if a loan application is rejected, the decision path analysis can show the customer the specific links and factors that led to the application being rejected, such as insufficient credit score, poor income stability, etc., so that the customer can understand the reasons for the decision and reduce unnecessary disputes. It also helps banks to discover possible problems in the approval process and optimize them. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] Figure 1 Schematic diagram of the process of financial data analysis and statistical model construction of the present invention. DETAILED DESCRIPTION

[0048] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0049] Example 1

[0050] See also Figure 1 , a financial data analysis and statistical model construction method, comprising the following steps:

[0051] Step 1: Data Collection: Collect customer transaction flows, public opinion texts, and industry chain data from multiple channels, including financial institutions’ transaction systems, social media platforms, and industry chain databases.

[0052] Step 2: Data Processing and Feature Extraction: Clean the collected data to remove noise, including outliers in transaction flows, advertisements and irrelevant information in public opinion texts, and standardize it for subsequent analysis. Extract key features from transaction flows, public opinion texts, and industry chain data from the processed data. These features are used to reflect the temporal relationships, trend changes, and anomalies of the data.

[0053] Step 3: Construct an initial knowledge graph: Based on the extracted information, construct an initial knowledge graph. Taking the customer transaction network of a bank as an example, customers are used as nodes, and transaction relationships between customers are used as edges. The weight of the edges is determined by the transaction amount and frequency. At the same time, the evaluation information related to customers in the public opinion text is associated with the corresponding customer nodes, and the upstream and downstream cooperation relationships related to enterprises in the industrial chain data are associated with the enterprise nodes.

[0054] Step 4: Knowledge Graph Update: Continuously monitor changes in the data's time series and continuously update the knowledge graph. When a customer's recent trading frequency increases significantly and they frequently have financial transactions with certain high-risk accounts, update the customer's risk management attributes in the knowledge graph and adjust the associated edge relationships. Set an update threshold. When market volatility exceeds the threshold (defined by the change in the market volatility index), rerun the data processing and feature extraction process, update the nodes, edges, and their attributes in the knowledge graph, modify the updated node attributes according to the new data processing results, and update the existence and weight of the edges based on the new relationship calculations.

[0055] The data collection in step one includes the data input layer, which is used to obtain customer transaction flow, public opinion text and industry chain data;

[0056] Transaction flow data: including transaction time T transaction Transaction Amount transactionTransaction Object O transaction Information, assuming that the transaction flow data set within a certain period of time is:

[0057] D transaction ={(t1,a1,o1,)(t2,a2,o2,)…(t n ,a n ,o n ,)}, where n is the number of transaction records;

[0058] Public opinion text data: Collect public opinion texts within a specific time period from various data sources (such as social media platforms, news websites, forums, etc.) to ensure the integrity and accuracy of the data. The texts include comments, posts, and articles. For each public opinion text P i , with release time Tp i , sentiment tendency score Sp i ∈[-1,1] (obtained from the sentiment analysis model, 1 represents positive, 1 represents negative), hot topic category Hp i (obtained through the text classification model), the public opinion text dataset is recorded as:

[0059] D opinion ={(p1,tp1,sp1,p1)(p2,tp2,sp2,p2)…(p m ,tp m ,sp m ,p m )},

[0060] Among them, m is the number of public opinion texts;

[0061] Industry chain data: including upstream and downstream industry relationships updown (For example, if enterprise E1 is the upstream supplier of enterprise E2, it is represented by R updown (E1, E2), market share changes MS change , the industry chain dataset is represented as: D industry ={(r1,ms1)(r2,ms2)…(r k ,ms k )}, where k is the number of industry chain relationship records;

[0062] The data processing and feature extraction in step 2 include: transaction flow data processing, public opinion text data processing, and industry chain data processing;

[0063] Transaction flow data processing: Discretize transaction time and perform segmented statistics on transaction frequency by hour or day. and average transaction amount Calculate the transaction amount ratio between transaction objects:

[0064]

[0065] Among them, a ij Indicates that from the transaction object o i to o j The transaction amount, ∑a ik Indicates that i The sum of all relevant transaction amounts;

[0066] Public opinion text data processing: Preprocess the collected text to remove irrelevant information (such as HTML tags, special characters, and stop words), and perform word segmentation to divide the text into meaningful words or phrases for subsequent sentiment analysis. For sentiment tendency scores, calculate the average sentiment tendency over a period of time:

[0067]

[0068] Where m′ is the number of public opinion texts in the time period. The sentiment tendency score can be a discrete category label (such as positive, negative, neutral) or a continuous numerical score. The sentiment tendency scores of all public opinion texts in the time period are added together to obtain the total sentiment tendency score. If the sentiment tendency score is a discrete category label, it is converted into a corresponding numerical value for accumulation. For example, positive is recorded as +1, negative is recorded as -1, and neutral is recorded as 0. The total sentiment tendency score is divided by the number of public opinion texts in the time period to obtain the average sentiment tendency score.

[0069] Industrial chain data processing: For upstream and downstream relationships in the industry, construct the adjacency matrix A updown , if R updown (E i ,E j2 ) exists, then A updown [i,j]=1, otherwise 0. The market share change is normalized, and the mathematical expression is:

[0070]

[0071] Among them, min(MS change ) and max(MS change ) are the minimum and maximum values ​​of market share changes, respectively, mapping the data to the [0,1] interval;

[0072] Step 3 of building the initial knowledge graph includes the following steps:

[0073] S1: Node generation: Generate different types of nodes based on the processed data, including client node N customer , Financial product / service node N product , enterprise node N enterprise , hot topic node Ntopic ,The attributes of the nodes are set according to the corresponding data ,features. The attributes of the customer node include age and asset size (if ,there are relevant information), and the attributes of the enterprise node include ,industry type and registered capital;

[0074] S2: Edge generation and weight calculation: Generate edge relationships between customers and financial products / services, and between customers based on transaction flow data. For example, if customer C i Purchased financial product P j , then generate edge E transaction (C i ,P j ), whose weight is determined by the transaction amount and transaction frequency factors, and is mathematically expressed as: Among them, f is a comprehensive calculation function, which is a linear weighted function;

[0075] S3: Public opinion related edges: Generate relationship edges between customers and hot topics, and between companies and hot topics from public opinion text data. For example, if customer C i Posted about hot topic H k The public opinion text generates edge E opioion (C i ,H k ), whose weight is determined according to the sentiment tendency score, and is mathematically expressed as: W opioion (C i ,H k )=g(S pi ), where g is a monotonically increasing or decreasing function (determined according to business requirements, such as an absolute value function);

[0076] S4. Industry chain relationship edge: Generate upstream and downstream relationship edges between enterprises based on industry chain data. The weight of the edge is determined by market share changes and transaction amount factors. The mathematical expression is: W industry (E i ,E j )=h(ms pi ,a′ ij ), where h is the calculation function, a′ ij It is a value related to the transaction amount between enterprises obtained through relevant transaction data of the industrial chain.

[0077] In this embodiment, data from different sources and structures, such as customer transaction flows, public opinion texts, and industrial chain data, can be integrated together. For example, customer transaction flows can reflect the customer's financial behavior and preferences, public opinion texts reflect market sentiment and public opinion, and industrial chain data shows the connections between upstream and downstream industries and the market structure. This fusion of multi-source data makes the analysis more comprehensive and avoids the limitations brought by a single data source. In financial risk management, accurate risk assessment is key. Using a dynamic knowledge graph containing multimodal data, credit risk, market risk, etc. can be more comprehensively assessed. For example, for loan business, in addition to considering the customer's financial status (transaction flow), public opinion texts can also be combined to judge their social reputation and potential risk factors. At the same time, reference can be made to industrial chain data to assess the risk trends of their industry, thereby making more accurate loan risk assessment decisions.

[0078] By constructing a knowledge graph, rich semantic relationships are endowed to these data. For example, the graph can clearly identify the transaction relationship between customers and financial products, the public opinion correlation between customers and hot topics, and the upstream and downstream relationships between enterprises in the industrial chain. These semantic relationships help to deeply explore the potential information behind the data. Financial institutions can provide more personalized financial services based on the comprehensive information of customers in the knowledge graph. For example, banks can tailor financial products, loan plans or insurance plans for customers based on their consumption habits (transaction flow), interests (public opinion text) and the position of their industry in the industrial chain, thereby improving customer satisfaction and loyalty.

[0079] Financial business processes can be optimized based on knowledge graphs. For example, in the credit approval process, knowledge graphs can be used to quickly integrate and evaluate multi-source customer data, reduce manual review steps, and improve approval efficiency. In marketing, the characteristics of potential customer groups analyzed by knowledge graphs can be used to accurately locate marketing goals, reduce marketing costs, and improve marketing effectiveness.

[0080] Example 2

[0081] This embodiment is a further optimization based on the first embodiment. Specifically, the knowledge graph update in step 4 also includes an event-driven update mechanism to monitor financial market fluctuations. When market fluctuations exceed a set threshold (such as a stock index rise or fall by more than 5% in a short period of time), the graph structure reorganization is automatically triggered. The event-driven update mechanism includes market data collection, event monitoring and identification, event impact analysis, graph structure reorganization, and topology optimization.

[0082] Market data collection: Acquire real-time market data from multiple financial data sources, including stock trading markets, futures trading markets, and foreign exchange markets, including key indicators such as price and trading volume. Clean, organize, and standardize the data collected from different sources for subsequent analysis and use. Fill in or correct missing or abnormal data to ensure data quality and consistency.

[0083] Event monitoring and identification: Set market volatility monitoring indicators, such as stock price volatility (measured by calculating the standard deviation of prices over a certain period of time), calculate market volatility indicators (such as volatility), and compare them with preset thresholds to determine whether a market volatility event is triggered. Market volatility calculation (taking standard deviation as an example) assumes P t is the stock price on the tth trading day, n is the number of trading days in the observation period, and the calculation formula for market volatility σ is:

[0084]

[0085] in, For example, the stock price is considered to be the mean of the stock price. For example, the stock price fluctuation range within a certain period of time is set to exceed a certain percentage as the judgment standard for a volatility event. When the market data exceeds the set threshold, it is determined that a market volatility event has occurred.

[0086] Event impact analysis: For different types of events (market fluctuations), impact assessment models are constructed separately. For market volatility events, a regression analysis model based on historical data is considered to analyze the impact of market fluctuations on different financial entities (such as individual stocks and industry indices). For example, with stock prices as the dependent variable and market volatility as the independent variable, a linear regression model is established (for market volatility impact assessment): Y = β0 + β1X + ∈, where Y is the dependent variable (such as stock price changes), X is the independent variable (such as market volatility), β0 and β1 are regression coefficients, ∈ is the error term, and the regression coefficient β1 is estimated by the least squares method:

[0087]

[0088] in, and where β1 is the sample mean of X and Y, respectively. By estimating β1, we can understand the average impact of market volatility on stock prices. Based on the entity relationships in the knowledge graph, we can analyze the scope of entities affected by the event. For example, when the central bank adjusts interest rates (monetary policy), it will not only affect the credit business of banking and financial institutions (directly related entities), but also downstream entities related to the real estate market, consumer companies, and bank credit funds (indirectly related entities). By traversing and reasoning about entity relationships in the knowledge graph, we can determine the relevant entities and their relationships that need to be updated.

[0089] Graph structure reorganization: For entities affected by events, their attributes are updated based on the results of the event impact analysis. For example, after market fluctuations cause a stock price to drop sharply, the market value attribute (market value = stock price × total share capital) and risk rating attribute of the stock entity are updated (risk rating is reassessed based on price fluctuations). Based on the new entity attributes and the results of the event impact analysis, the relationship between entities is adjusted. For example, when a company gains a competitive advantage due to policy support (such as tax incentives), the association between the company and the policy-issuing department is strengthened in the knowledge graph (such as increasing the relationship weight), while the competitive relationship between the company and its competitors may be weakened (such as reducing the relationship weight). For newly emerging entity relationships, they are added based on event information and changes in entity attributes. For example, after the release of emerging industry policies, investment and financing relationships between emerging companies and investment institutions may emerge, and this new relationship is added to the knowledge graph.

[0090] Topology optimization: After completing entity attribute updates and relationship adjustments, the topology of the knowledge graph is optimized. A community discovery algorithm is used to divide entities in the knowledge graph into communities, so that entities within the same community have closer relationships, facilitating subsequent analysis and querying. The modularity-optimized community discovery algorithm aims to maximize the difference between the edge density within a community and the edge density of the entire graph.

[0091] The objective function in the community discovery algorithm (taking modularity optimization as an example) is to assume that A is the adjacency matrix of the knowledge graph, e ij Represents node i R and node j R Are they connected? ij =1 R Indicates connection, e ij =0 R Indicates not connected), Represents node i R The degree, Represents the total number of edges in the graph, and modularity is expressed as:

[0092]

[0093] Among them, δ(c i ,c i ) is an indicator function that takes the value of 1 when nodes i and j belong to the same community, and takes the value of 0 otherwise. It optimizes the community division by maximizing the modularity Q.

[0094] In this embodiment, when the market fluctuates, the event-driven update mechanism can capture and reflect the latest market dynamics in real time, ensuring that the information in the knowledge graph is synchronized with the actual market situation. Timely data updates help companies respond quickly and adjust strategies to cope with the rapidly changing market environment. Through the real-time updated knowledge graph, decision makers can obtain the most accurate and comprehensive data support, thereby making more informed decisions.

[0095] The dynamic information provided by the event-driven update mechanism helps companies better predict market trends and develop forward-looking strategic plans. Real-time data support enables companies to more accurately identify market demand and resource shortages, thereby optimizing resource allocation and improving operational efficiency.

[0096] The event-driven update mechanism enables the knowledge graph to flexibly adapt to the ever-changing market environment and maintain the freshness and relevance of the data. This mechanism enhances the robustness of the system, enabling it to maintain stable operation in the face of uncertainty and emergencies.

[0097] Example 3

[0098] This embodiment is based on Example 1 or Example 2 and is optimized as follows. Specifically, the model building method also includes an interpretability enhancement module for a three-dimensional attribution analysis system. This module analyzes financial models from three dimensions: feature importance (global), decision path (local), and counterfactual explanation (causal). This includes feature importance analysis, decision path tracking, and counterfactual explanation application.

[0099] Feature importance analysis involves the following steps:

[0100] Data preprocessing: Collect and organize data related to financial models, including various characteristic variables such as the insured's age, health status, and occupation, as well as the corresponding insurance premium target variables. Clean and standardize the data to ensure data quality and consistency.

[0101] Select evaluation metrics: Based on the characteristics and objectives of the problem, select information gain as the metric to measure feature importance. The importance of a feature is determined by calculating the change in information entropy before and after the feature is partitioned into the dataset.

[0102] Calculate feature importance: Use the selected evaluation index to calculate the contribution of each feature to the prediction result. For the insurance pricing model, the decision tree algorithm is used to calculate the information gain. First, calculate the information entropy H(D) of the entire data set, and then for each feature A i , Feature A i The value set is {a i 1,a i 2,…,a i n}, calculate the conditional entropy of the data set under different values ​​of this feature (D|Ai ), then the information gain IG is calculated as:

[0103] IG(D,A i )=H(D)H(D|A i ), calculate feature A i The greater the information gain, the greater the impact of the feature on the insurance premium rate, that is, the higher the feature importance; where H(D) is the information entropy of the data set D, and the calculation formula is:

[0104] Among them, p(y k ) represents the category y k The probability of appearing in the data set D; H(D|A i ) is the dataset D in feature A i Under the condition of entropy, the calculation formula is:

[0105]

[0106] Among them, p(a ij ) represents feature A i The value is a ij The probability of H(D|A i =a ij ) indicates that in feature A i The value is a ij The conditional entropy of the data set under the condition of;

[0107] Result visualization and analysis: Visualize the calculated feature importance results and draw bar charts or heat maps to intuitively observe the importance ranking of each feature, analyze key features and their impact on pricing decisions, and provide a basis for explaining and optimizing insurance pricing models.

[0108] Decision path tracing includes the following steps:

[0109] Sample selection and data preparation: Select sample data of individual loan applicants, including information on their income level, credit history, and collateral value, to ensure data accuracy and completeness;

[0110] Application of model interpretation tools: Use the LIME (Local Interpretable Model-Sensitive Explanations) tool to track the model's decision-making process. By generating a large number of perturbation samples near the sample point, the difference between the prediction results of the perturbed samples and the original samples is calculated, and the impact of each feature on the decision is determined based on the size of the difference.

[0111] Decision path visualization and analysis: Visualize the model's decision-making process for a single sample by drawing a decision tree or flowchart. By analyzing the decision path, we can understand how the model makes approval or rejection decisions based on the values ​​of different features. For example, if the applicant has a high income and a good credit history, the model may be inclined to approve the loan application; however, if the collateral value is low and there is a history of overdue payments, the model may reject the loan application.

[0112] Perturbation sample generation and explanation score calculation in LIME: Let the sample be x, the model be f, and the perturbation sample set be: {x1,x2,…,x m}, the prediction result set corresponding to the perturbation sample is: {f(x1),f(x2),…,f(x m )}, then interpret the fraction e i The calculation formula is:

[0113]

[0114] in, Represents the perturbation sample x j The value of the i-th feature, x i Represents the value of sample x on the i-th feature, explaining the score e i It reflects the influence of the i-th feature on the model prediction results. The larger the absolute value, the greater the influence.

[0115] The application of counterfactual explanations involves the following steps:

[0116] Case selection and hypothesis setting: Select a case of insurance claim rejection as the research object, and set different hypothetical situations for this case, such as changing the insured's disease type or treatment method;

[0117] Model recalculation and comparative analysis: Input hypothetical data into the insurance claims model for recalculation to obtain the model's claim decision results under the hypothetical circumstances. The model then compares the actual situation with the hypothetical decision results. For example, if the policyholder's original disease type is a rare disease with a complex treatment method, resulting in a claim rejection, assume that the disease type is changed to a common disease with a more standardized treatment method, and recalculate the model's claim decision. If the model approves the claim under the hypothetical circumstances, it means that the change in disease type and treatment method has affected the claim decision.

[0118] Causal interpretation and summary: Through comparative analysis, we gain a deeper understanding of the causal relationships behind model decisions. We use causal inference methods, such as matching and inverse probability weighting, to quantify the impact of hypothetical factors on claims decisions. Ultimately, we summarize the causal relationships between different hypothetical factors and claims decisions, providing a reference for optimizing insurance claims models and controlling risks.

[0119] Calculation of matching weight in matching method: Let the actual case be C, the hypothetical case be C′, and the matching weight be w(C, C′). Then the calculation formula of matching weight is:

[0120]

[0121] Among them, d i (C, C′) represents the difference between the actual case C and the hypothetical case C′ in the i-th feature, which is calculated using the Euclidean distance method. The matching weight is used to adjust the impact of the hypothetical case on the actual case, so that the actual situation can be more accurately reflected when calculating the causal effect.

[0122] In this embodiment, the feature importance dimension can accurately determine which factors play a key role in pricing decisions by calculating the influence coefficient of each factor on the prediction result. For example, in an insurance pricing model, it can be determined which factors, among factors such as the insured's age, health status, and occupation, have a more significant impact on insurance premiums. This helps insurance companies formulate pricing strategies more accurately, helps insurance companies optimize the pricing mechanism of insurance products, improves the accuracy and rationality of pricing, avoids pricing errors caused by incomplete consideration of factors, and enhances the company's competitiveness in the market. At the same time, decision makers can adjust the design or marketing strategy of insurance products according to actual conditions, such as conducting risk assessment and management on certain key factors to improve the company's risk control capabilities.

[0123] Decision path tracking records the model's decision-making process for a single sample in detail, making the loan approval process more transparent. Applicants can clearly understand how factors such as their income level, credit history, and collateral value affect the loan approval result, increasing applicants' understanding and trust in the approval result, reducing unnecessary disputes and misunderstandings, and improving customer satisfaction. It also helps financial institutions analyze the specific circumstances of individual samples and identify factors that may lead to poor decisions. For example, if a specific problem is found in an applicant's credit record that led to a loan rejection, but the problem can be improved in some way, targeted suggestions can be provided to the applicant, which can not only improve the service quality of the financial institution, but also help applicants improve their own credit status and increase their chances of obtaining a loan in the future.

[0124] The application of counterfactual interpretation can deeply explore the causal relationship behind model decisions by comparing actual situations with hypothetical situations. Taking the case of rejected insurance claims as an example, it can clearly understand how changes in factors such as the insured's disease type or treatment method affect the claims decision. This will help insurance companies further optimize claims policies, improve the fairness and rationality of claims, and reduce the occurrence of erroneous claims rejections.

[0125] In summary: By comprehensively integrating multi-source data and deeply mining potential information, the present invention can more accurately assess risks and promptly discover potential connections and risks. By using a temporal relationship reasoning algorithm, customer transaction flows, public opinion texts, and industry chain data are converted into dynamic graphs. This can integrate information from different dimensions, more comprehensively assess loan risks, and reduce risk misjudgments caused by a single data source. The implicit relationships between different data sources are revealed in the dynamic graph, and it may be possible to discover that corporate customers in certain specific industry links have higher insurance claim risks under specific transaction patterns.

[0126] Event-driven adaptive updates keep the graph up-to-date and keep pace with market changes, thereby improving the accuracy and timeliness of decision-making. An event-driven update mechanism is designed to enable the knowledge graph to automatically trigger graph restructuring when the market fluctuates or policies are released. The knowledge graph can be updated in a timely manner to reflect the changes in the risk-return relationship of different financial products under different interest rate environments, providing financial institutions with the latest decision-making basis and enabling them to quickly adjust their investment strategies. The automatically updated graph ensures the timeliness of the data and avoids erroneous decisions caused by the use of outdated data.

[0127] Building a three-dimensional attribution analysis system can improve decision-making transparency and enhance the credibility and fairness of financial data analysis. Through feature importance analysis, the key factors and their weights that affect financial decisions can be determined from a global perspective, thereby formulating more reasonable and fair insurance premiums. At the same time, the pricing basis can be explained to customers to improve their understanding and acceptance of pricing. Decision path analysis can record the decision-making process of a single sample in detail. Through decision path analysis, customers can be shown the specific links and factors that led to the failure of the application, reducing unnecessary disputes. By comparing the actual risk events with the situation where the event was assumed not to have occurred for causal analysis, the effectiveness of risk control measures can be more accurately evaluated, and the impact of the occurrence or non-occurrence of a risk event on the value of the financial institution's asset portfolio can be analyzed, which helps to optimize risk management strategies and improve the ability of financial institutions to cope with risks.

[0128] Finally, it should be noted that the above is only a preferred embodiment of the present invention and is not intended to limit the present invention. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art can still modify the technical solutions described in the aforementioned embodiments or make equivalent substitutions for some of the technical features therein. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.

Claims

1. A method for financial data analysis and statistical model construction, characterized in that: The following steps are involved: Step 1: Data Collection: Collect customer transaction flows, public opinion texts, and industry chain data from multiple channels, including financial institutions’ transaction systems, social media platforms, and industry chain databases. Step 2: Data Processing and Feature Extraction: Clean the collected data to remove noise, including outliers in transaction flows, advertisements and irrelevant information in public opinion texts, and perform standardization. Key features of transaction flows, public opinion texts, and industry chain data are extracted from the processed data. Step 3: Construct an initial knowledge graph: Based on the extracted information, construct an initial knowledge graph. Taking the customer transaction network of a bank as an example, customers are used as nodes, and transaction relationships between customers are used as edges. The weight of the edges is determined by the transaction amount and frequency. At the same time, the evaluation information related to customers in the public opinion text is associated with the corresponding customer nodes, and the upstream and downstream cooperation relationships related to enterprises in the industrial chain data are associated with the enterprise nodes. Step 4: Knowledge graph update: Continuously monitor the time series changes of data and continuously update the knowledge graph.

2. A financial data analysis and statistical model building method according to claim 1, characterized in that: The data collection described in step one includes a data input layer, which is used to obtain customer transaction flows, public opinion texts and industry chain data; the transaction flow data: includes transaction time, transaction amount, and transaction object information; the public opinion text data: collects public opinion texts within a specific time period from various data sources, and the texts include comments, posts, and articles; the industry chain data: includes upstream and downstream relationships in the industry and changes in market share.

3. The method for financial data analysis and statistical model construction according to claim 2, characterized in that: The data processing and feature extraction described in step 2 include: transaction flow data processing, public opinion text data processing, and industry chain data processing; Transaction flow data processing: Discretize transaction time, perform segmented statistics on transaction frequency and average transaction amount by hour or day, and calculate the transaction amount ratio between transaction parties; Public opinion text data processing: pre-process the collected text, remove irrelevant information, and perform word segmentation operations to divide the text into meaningful words or phrases. For the sentiment tendency score, calculate the average sentiment tendency over a period of time, add up the sentiment tendency scores of all public opinion texts in the time period, and obtain the total sentiment tendency score. Divide the total sentiment tendency score by the number of public opinion texts in the time period to obtain the average sentiment tendency score; industrial chain data processing: for upstream and downstream industry relationships, construct an adjacency matrix.

4. The method for financial data analysis and statistical model construction according to claim 1, characterized in that: Step 3 of constructing the initial knowledge graph includes the following steps: S1: Node generation: Based on the processed data, different types of nodes are generated, including customer nodes, financial product / service nodes, enterprise nodes, and hot topic nodes. The attributes of the nodes are set according to the corresponding data characteristics. The attributes of customer nodes include age and asset size, and the attributes of enterprise nodes include industry type and registered capital. S2: Edge generation and weight calculation: Generate relationship edges between customers and financial products / services, and between customers based on transaction flow data; S3: Public opinion association edges: Generate relationship edges between customers and hot topics, and between companies and hot topics from public opinion text data; S4. Industry chain relationship edges: Generate upstream and downstream relationship edges between enterprises based on industry chain data.

5. The method for financial data analysis and statistical model construction according to claim 1, characterized in that: The method of updating the knowledge graph described in step 4 is: setting an update threshold. When the market fluctuation exceeds the threshold, re-run the data processing and feature extraction process to update the nodes, edges and their attributes in the knowledge graph. The updated node attributes are modified according to the new data processing results, and the existence and weight of the edges are updated based on the new relationship calculation.

6. The method for financial data analysis and statistical model construction according to claim 1, characterized in that: The knowledge graph update described in step 4 also includes an event-driven update mechanism to monitor financial market fluctuations. When market fluctuations exceed a set threshold, the graph structure is automatically triggered to restructure. The event-driven update mechanism includes market data collection, event monitoring and identification, event impact analysis, graph structure restructure, and topology optimization. Market data collection: Acquire real-time market data from multiple financial data sources, including key indicators such as price and trading volume. Clean, organize, and standardize the data collected from different sources, and fill in or correct missing or abnormal data. Event monitoring and identification: Set market volatility monitoring indicators, calculate market volatility indicators, which are volatility rates, and compare them with preset thresholds to determine whether a market volatility event is triggered. When market data exceeds the set threshold, it is determined that a market volatility event has occurred.

7. The method for financial data analysis and statistical model construction according to claim 6, characterized in that: The event impact analysis: for different types of events, impact assessment models are constructed respectively. For market volatility events, a regression analysis model based on historical data is used to analyze the impact of market fluctuations on different financial entities; The graph structure reorganization: for entities affected by the event, their attributes are updated according to the results of the event impact analysis, and the relationships between entities are adjusted based on the new entity attributes and the event impact analysis results. For newly emerging entity relationships, they are added according to the event information and entity attribute changes; The topology structure optimization is as follows: after completing the entity attribute update and relationship adjustment, the topology structure of the knowledge graph is optimized, and the entities in the knowledge graph are divided into communities using a community discovery algorithm. The community discovery algorithm with modularity optimization is used, and its goal is to maximize the difference between the edge density within the community and the edge density of the entire graph.

8. The method for financial data analysis and statistical model building according to claim 1, characterized in that: The model building method also includes an explainability enhancement module for a three-dimensional attribution analysis system, which analyzes financial models from three dimensions: feature importance, decision path, and counterfactual explanation. include Feature importance analysis, decision path tracking, and counterfactual interpretation applications; The feature importance analysis includes the following steps: Data preprocessing: Collect and organize data related to financial models, including various characteristic variables such as the insured's age, health status, and occupation, as well as the corresponding insurance premium target variables, and perform data cleaning and standardization preprocessing operations; Select evaluation metrics: Based on the characteristics and objectives of the problem, select information gain as the metric to measure feature importance. The importance of a feature is determined by calculating the change in information entropy before and after the feature is partitioned into the dataset. Calculate feature importance: Use the selected evaluation metrics to calculate the contribution of each feature to the prediction results. For insurance pricing models, use the decision tree algorithm to calculate information gain. Visualize and analyze results: Visualize the calculated feature importance results, draw bar charts or heat maps, and analyze key features and their impact on pricing decisions.

9. A financial data analysis and statistical model building method according to claim 8, characterized in that: The decision path tracking includes the following steps: Sample selection and data preparation: Select sample data of individual loan applicants, including information about their income level, credit history, and collateral value; Application of model interpretation tools: Use the LIME tool to track the model's decision-making process. By generating a large number of perturbation samples near the sample point, the difference between the prediction results of the perturbed samples and the original samples is calculated. The impact of each feature on the decision is determined based on the size of the difference. Decision path visualization and analysis: Visualize the model's decision-making process for a single sample by drawing a decision tree or flowchart. By analyzing the decision path, we can understand how the model gradually makes approval or rejection decisions based on the values ​​of different features.

10. The method for financial data analysis and statistical model building according to claim 8, characterized in that: The counterfactual explanation application includes the following steps: Case selection and hypothesis setting: Select a case of insurance claim rejection as the research object, and set different hypothetical situations for this case; Model recalculation and comparative analysis: Input hypothetical data into the insurance claims model for recalculation to obtain the model's claim decision results under the hypothetical circumstances. The actual situation is then compared and analyzed with the hypothetical decision results. Causal relationship explanation and summary: Through comparative analysis, we gain an in-depth understanding of the causal relationship behind model decisions. We use causal inference methods to quantify the impact of hypothetical factors on claims decisions, and ultimately summarize the causal relationship between different hypothetical factors and claims decisions.

Citation Information

Patent Citations

  • Method and device for evaluating model interpretation tool

    CN111325344A

  • Method and system for mining risk propagation path based on knowledge graph

    CN112214614A

  • Automatic feature mining-based interpretable credit default rate prediction method and system

    CN115936159A

  • Relationship graph construction method and device, relation graph query method and device, electronic equipment and storage medium

    CN117764695A

  • Financial risk monitoring system and monitoring method

    CN118674554A

Cited By

  • Intelligent training system with multi-branch interactive plot and related equipment

    CN121788296A