Financial risk assessment system based on emotion perception and dynamic knowledge graph
Through the combination of multi-source data collection, emotional perception and dynamic knowledge graph, the financial risk assessment system has solved the problem of single data dimension, insufficient real-timeness and poor model interpretability in the model, and achieved high-precision, interpretable and dynamically optimized risk assessment, improving the comprehensiveness and accuracy of risk identification.
Patent Information
- Application Number
- CN202510452911.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-11
- Publication Date
- 2025-08-15
AI Technical Summary
The existing financial risk assessment system has limitations in dealing with complex market correlation, dynamic risk transmission and unstructured emotional signals. It is difficult to capture the real-time risk transmission path of implicit guarantee chains and supply chain financial networks among enterprises, and lacks the ability to dynamically adjust market sentiment fluctuations, resulting in a lag in risk warning and a one-sided evaluation result.
The multi-source data acquisition module is used to integrate structured and unstructured data, quantify customer emotional characteristics through the emotion perception engine, combine dynamic knowledge graph construction module to update the knowledge graph in real time, use multi-modal risk assessment model to integrate emotional features and knowledge graph embedding, and iteratively update the risk assessment logic through dynamic feedback and optimization modules to realize dynamic tracking and risk score of the emotion communication model.
It improves the comprehensiveness and accuracy of risk identification, improves the robustness and adaptability of risk prediction through multimodal data fusion, enhances the interpretability and real-time response capabilities of the model, and supports high-precision risk prevention and control and decision-making of financial institutions.
Smart Images

Figure CN120494975A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of financial risk assessment, and more specifically, to a financial risk assessment system based on emotion perception and dynamic knowledge graph. Background Art
[0002] With the in-depth application of large-scale models, knowledge graph technology, and multimodal sentiment analysis in the banking industry, traditional financial risk assessment methods are increasingly limited in addressing complex market correlations, dynamic risk transmission, and processing unstructured sentiment signals. Traditional models rely on static financial indicators and historical data, making it difficult to capture the real-time risk transmission paths within implicit inter-enterprise guarantee chains and supply chain finance networks. They also fail to effectively integrate the potential impact of public sentiment fluctuations on market confidence.
[0003] For example, machine learning-based credit scoring models lack the ability to model risk transmission and rely on static data. Traditional models (such as logistic regression and decision trees) primarily rely on historical credit records and structured financial data, failing to capture the dynamic transmission of market risk across industries and institutions. In supply chain finance scenarios, the diffusion path of credit risk from core enterprises to upstream and downstream small and medium-sized enterprises is difficult to quantify.
[0004] Limitations of complex relationship modeling: Although machine learning (such as random forests and neural networks) can handle nonlinear relationships, it lacks the ability to explicitly model risk transmission links (such as multi-hop risk propagation path analysis), resulting in delayed systemic risk warnings.
[0005] Lack of sentiment impact analysis and insufficient coverage of unstructured data: Existing models rarely integrate unstructured information such as social media sentiment data and news opinions, making it difficult to quantify the amplifying effect of market sentiment fluctuations on credit risk, such as liquidity crises caused by panic selling.
[0006] Weak correlation between sentiment and risk: The model does not establish a causal relationship between sentiment polarity, such as the intensity of negative news, and the probability of default, making it impossible to dynamically adjust risk weights. For example, the real-time impact of market sentiment during the release of major policies is not incorporated into the scoring system.
[0007] Taking the risk assessment system based on expert rules as an example, its dynamic adaptability of risk transmission is poor and rule updates are delayed: expert rules rely on manual experience summary and are difficult to respond quickly to changes in the market environment, such as chain reactions caused by emergencies. The rule base update cycle usually takes weeks to months.
[0008] Weak cross-scenario generalization capabilities: Rule systems are mostly designed for specific scenarios, such as corporate bond assessment, and lack a unified modeling framework for cross-market risk transmission, such as the linkage between the stock and bond markets.
[0009] Limitations of sentiment impact analysis: Subjectivity and insufficient coverage: The sentiment judgment criteria in the rule base rely on manual definition, such as keyword matching, and cannot cover diverse forms of emotional expression, such as sarcasm and metaphor. It also lacks fine-grained sentiment classification, such as the difference between "strongly negative" and "mildly negative."
[0010] Real-time defects: Emotional data needs to be manually labeled and matched with rules, resulting in high processing delays, such as hours to days, making it difficult to capture high-frequency emotional fluctuations in a timely manner, such as the rapid fermentation of hot events on social media.
[0011] In summary, the existing financial assessment systems and methods have the following shortcomings:
[0012] Single data dimension: Traditional systems rely on structured financial data, such as debt-to-asset ratios and credit scores, and ignore unstructured data, such as customer behavior, social media sentiment, and customer service conversation records, resulting in one-sided risk assessment.
[0013] Static knowledge graphs: Existing knowledge graphs are mostly statically constructed based on historical data. The machine learning models lack real-time performance and cannot integrate market dynamics in real time, such as policy changes, industry trends, and public opinion events, resulting in delayed risk assessment.
[0014] Lack of sentiment analysis: There is a lack of quantitative modeling of customer emotional tendencies, such as complaint sentiment and investment preferences, resulting in a high false alarm rate and difficulty in predicting sudden risks such as bank runs and mass defaults.
[0015] Poor model interpretability: Black box algorithms, such as deep neural networks, have opaque decision logic and are difficult to meet the compliance requirements of financial regulators. Summary of the Invention
[0016] The purpose of this invention is to provide a financial risk assessment system based on emotion perception and dynamic knowledge graph, which can quantify the impact of user emotions on risk, use knowledge graph to capture the relationship between entities, and improve the comprehensiveness and accuracy of risk identification.
[0017] The above technical objectives of the present invention are achieved through the following technical solutions: A financial risk assessment system based on emotion perception and dynamic knowledge graph, comprising:
[0018] The multi-source data acquisition module is used to connect to multiple data sources, collect structured data and unstructured data, and standardize the collected data to obtain standard data;
[0019] Emotion perception engine, including text sentiment analysis module and speech sentiment recognition module, used to quantify customer emotional characteristics;
[0020] The dynamic knowledge graph construction module is used to extract entities and corresponding relationships through the entity extraction model. It is also used to attenuate the weight of historical data according to the time dimension when incrementally updating the knowledge graph, and update the knowledge graph in real time.
[0021] Multimodal risk assessment model, which is used to integrate standard data, sentiment features and knowledge graph embedding to output risk scores;
[0022] The dynamic feedback and optimization module iteratively updates the risk assessment logic based on risk scores and market feedback, and performs incremental training on the multimodal risk assessment model through misjudgment data.
[0023] As a preferred technical solution of the present invention, the standardization process includes:
[0024] Duplicate data filtering: All collected data is converted into hash values through a hash algorithm to obtain a unique hash value for the data. The hash values of all data are compared. If data with the same hash value exists, it is considered duplicate and the duplicate data is then removed.
[0025] Missing value filling: For all collected data, detect missing fields and fill in the missing fields;
[0026] Data normalization: For the collected numerical data, the Min-Max normalization formula is used to obtain normalized data.
[0027] As a preferred technical solution of the present invention, the Min-Max normalization formula is:
[0028]
[0029] As a preferred technical solution of the present invention, the text sentiment analysis module outputs fine-grained sentiment labels and confidence scores by training a pre-trained model; models the sentiment propagation path to obtain a sentiment propagation model to simulate the negative sentiment diffusion path;
[0030] The speech emotion recognition module extracts acoustic features from speech data, inputs them into a classification model, and outputs an emotion classification label.
[0031] As a preferred technical solution of the present invention, the dynamic knowledge graph construction module includes an event extraction and entity linking module and a graph incremental update module. The event extraction and entity linking module is used to extract entities and corresponding relationships through an entity extraction model, calculate the similarity of all entities, and determine multiple entities with the same meaning and the different meanings of one entity based on the similarity.
[0032] The graph incremental update module is used to incrementally update the knowledge graph and decay the weight of historical data according to the time dimension. The decay formula is: λ(t) = e -kt , where k is the decay rate and t is the time interval.
[0033] As a preferred technical solution of the present invention, the architecture of the multimodal risk assessment model integrates a learning framework for processing structured data; integrates knowledge graph embedding for capturing relationships between entities; and also uses emotional features as independent input branches.
[0034] As a preferred technical solution of the present invention, the dynamic feedback and optimization module is provided with a feedback closed-loop mechanism, including backtracking of risk misjudgment cases and adjustment of knowledge graph weights.
[0035] As a preferred technical solution of the present invention, the dynamic feedback and optimization module is also used to generate visual reports, including risk transmission path diagrams and emotional heat maps.
[0036] In summary, the present invention has the following beneficial effects: cross-domain technology integration: combining sentiment analysis with dynamic knowledge graphs, quantifying the impact of user emotions on financial risks through sentiment propagation models, and using knowledge graphs to capture dynamic associations between entities, thereby improving the comprehensiveness and accuracy of risk identification.
[0037] Multimodal data fusion capability: Integrate multi-dimensional data such as emotions, transactions, and social interactions, adopt weighted average or neural network fusion strategies, and build a multimodal risk assessment model to significantly improve the robustness and adaptability of risk prediction.
[0038] Explainability and dynamic optimization: The Shapley value algorithm is used to quantify feature contributions, enhancing model transparency. In combination with the rule engine, the model and business rules are collaboratively verified to ensure the credibility and traceability of the evaluation results.
[0039] Real-time response and self-iteration capabilities: The attenuation factor and real-time data update mechanism based on the dynamic knowledge graph support dynamic tracking of risk transmission paths, and optimize model parameters through user feedback to achieve continuous system iteration.
[0040] Technical barriers and market value: Build multi-layered technical barriers covering emotional perception, dynamic graphs, and multimodal fusion to form differentiated competitive advantages, enabling rapid commercialization and market capture.
[0041] In summary, this patent combines technological innovation, application flexibility and commercial potential, providing financial institutions with a high-precision, explainable and dynamically optimized risk assessment tool to help improve risk prevention and control and decision-making efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] Figure 1 It is a schematic diagram of the system architecture of the present invention;
[0043] Figure 2 It is an application flow chart of the present invention. DETAILED DESCRIPTION
[0044] The present invention will be further described in detail below with reference to the accompanying drawings.
[0045] like Figure 1 and 2 As shown, the present invention provides an intelligent financial risk assessment system based on emotion perception and dynamic knowledge graph, including:
[0046] Multi-source data acquisition module:
[0047] It is used to connect to multiple data sources and collect structured and unstructured data; structured data includes bank transaction records, credit reports, financial statements, etc.; real-time access is achieved through API interfaces such as RESTful; unstructured data includes social media texts, customer service call recordings, news and public opinion, etc., and is dynamically collected using distributed crawlers such as Apache Nutch.
[0048] And standardize the collected structured data and unstructured data to obtain standard data;
[0049] Specifically, standardization processing includes:
[0050] Duplicate data filtering: All collected data is hashed using a hash algorithm, such as SHA-256, to obtain a unique hash value. This identifies the unique data entry and compares the hash values of all data. If data with the same hash value exists, it is considered duplicate and the duplicate data is removed.
[0051] Missing value filling: For all collected data, the XGBoost model is used to detect and predict missing fields, such as customer income levels, and fill in the missing fields. Through the actual application of this invention, the accuracy rate of missing value filling exceeds 92%;
[0052] Data normalization: For the collected numerical data, such as debt ratio, the Min-Max normalization formula is used to obtain normalized data to ensure the comparability of data of different dimensions. The Min-Max normalization formula is:
[0053]
[0054] The multi-source data acquisition module integrates structured and unstructured data sources, completes data cleaning and feature extraction, collects structured and unstructured data, and standardizes the data. It ensures the comparability of data of different dimensions by filtering duplicate data, filling missing values, and normalizing data.
[0055] Emotion Perception Engine:
[0056] Used to quantify customer sentiment characteristics and predict irrational financial behavior risks;
[0057] Emotion perception engine, including text emotion analysis module and speech emotion recognition module,
[0058] The text sentiment analysis module trains pre-trained models, such as BERT-Fin, based on financial forecasts, to output fine-grained sentiment labels and confidence scores. Specific fine-grained sentiment labels include strongly negative, moderately negative, neutral, and so on.
[0059] The output of the model is: Where Wi is the event weight. emo The higher the value, the more positive the overall emotion, and the lower the value, the more negative it is.
[0060] The text sentiment analysis module is also used to model the path of sentiment transmission. Based on the SEIR (Susceptible-Exposed-Infected-Recovered) infectious disease dynamics model, a sentiment transmission model is obtained to simulate the diffusion path of negative emotions. The spread of negative emotions is compared to the spread process of infectious diseases. The SEIR model in epidemiology is borrowed to quantify the dynamics of sentiment transmission in social networks.
[0061] Susceptible (S): Has not been exposed to emotions, but may be infected (unaffected users).
[0062] Lurkers (E): They are exposed to emotions but do not express them (e.g., seeing negative news but not forwarding it).
[0063] Infector (I): Actively spreads emotions (such as posting negative comments and forwarding).
[0064] Recoverers (R): Emotions subside or become immune (e.g., they stop spreading after rational users or platform intervention).
[0065] (P)(sentimenti): The sentiment probability of the i-th sentiment word, which indicates the possibility or confidence that the sentiment word has a certain specific emotion (such as positive, negative, etc.).
[0066] (n): The total number of sentiment words in the text, that is, the number of sentiment words involved in the calculation.
[0067] (t): represents time.
[0068] The following differential equations describe how the proportions of the four groups of people change over time:
[0069] Fewer susceptible people: β is the transmission rate (social interaction intensity), SI is the emotional transmission caused by the interaction between susceptible and infected people, dS and dt represent the calculus, the small change in variable S and time t, respectively, which are used in differential equations to describe the rate of change of variables over time.
[0070] Lurker change rate: SI: represents the number of susceptible and infected individuals. σ is the rate at which a latent state transitions to an infected state (e.g., the delay between a user seeing a negative message and forwarding it).
[0071] Infected changes: γ is the recovery rate (the rate at which users stop spreading emotions, such as when the platform deletes posts or users calm down).
[0072] Restorer Additions: This is the ratio of infected people to recovered people.
[0073] The meanings of the parameters in the present invention are as follows:
[0074] β represents the activity of the social platform and the infectiousness of information.
[0075] σ is the hesitation time between user acceptance and dissemination (such as the immediacy of Weibo hot searches and the in-depth thinking of long articles).
[0076] γ is the platform control efficiency or user self-regulation ability.
[0077] The speech emotion recognition module uses the Open SMILE tool to extract acoustic features from speech data, such as fundamental frequency (F0) and energy (RMS), and inputs these features into a classification model to output emotion classification labels such as anxiety / excitement. In practical applications, the technical solution of the present invention has a classification accuracy rate of over 88%.
[0078] The entire emotion perception engine uses the BERT-Fin model to perform sentiment classification on text, extract speech features, and uses the SEIR model for emotion recognition and propagation. It loads the pre-trained BERT-Fin model and SEIR model, pre-processes the input text and speech data, and inputs the pre-processed data into the corresponding model to obtain the sentiment analysis results.
[0079] Dynamic knowledge graph building blocks:
[0080] Used to integrate market dynamics in real time and build a knowledge graph of industry risks that can be used for fireworks, specifically including:
[0081] Through the entity extraction model, entities and corresponding relationships are extracted:
[0082] The dynamic knowledge graph construction module includes event extraction and entity linking modules. The entity extraction model used in this module is the BiLSTM-CRF model, which is used to extract entities and their relationships from news text or other target text. For example, entities include companies and policies, and relationships between entities include acquisitions and defaults.
[0083] After extracting entities and corresponding relationships through the entity extraction model, similarity is calculated for all entities. This similarity is used to identify multiple entities with the same meaning and the different meanings of a single entity. This is done by calculating entity similarity based on a knowledge base, such as Wikipedia, to eliminate homonyms.
[0084] The dynamic knowledge graph construction module also includes a graph incremental update module, which is used to attenuate the weight of historical data according to the time dimension when incrementally updating the knowledge graph, and to update the knowledge graph in real time to ensure that the knowledge graph can dynamically reflect market changes.
[0085] When the weight of historical data is attenuated according to the time dimension, the attenuation formula is:
[0086] λ(t)=e -kt , where k is the decay rate and t is the time interval.
[0087] When the dynamic knowledge graph construction module is actually applied, its storage architecture can use the Neo4j graph database to store entity relationships and use Elasticsearch as the search engine to support full-text retrieval.
[0088] Multimodal risk assessment model:
[0089] Used to fuse structured data, sentiment features and knowledge graph embedding to output risk scores;
[0090] The architecture of the multimodal risk assessment model is as follows: it integrates a learning framework to process structured data, integrates a knowledge graph embedding tool to capture complex relationships between entities, and also uses sentiment features as an independent input branch;
[0091] Among them, structured data can be processed by XGBoost, including traditional risk factors such as financial indicators and transaction records;
[0092] Emotional features, as an independent input branch, come from text analysis or speech emotion recognition;
[0093] Knowledge graph embeddings, generated using GraphSAGE, can capture complex relationships between entities.
[0094] The integrated learning framework of the multimodal risk assessment model adopts a weighted loss function, whose formula is: L = α·L XGB +β·L Graph +γ· / / S emo -Y risk / / 2, where α, β, and γ are all hyperparameters, which are optimized and adjusted based on actual network search results.
[0095] L is the loss function, which is the objective function that the entire framework needs to minimize, and is achieved by weighting the losses of different components.
[0096] α is a weighting coefficient used to adjust the weight of the XGBoost model loss function LXGBL. By changing the value of α, you can control the contribution of the XGBoost model to the final loss function.
[0097] β is a weighted coefficient that adjusts the weight of the graph neural network model loss function LGraphLGraph. Similar to α, β controls the influence of the graph neural network model in the final loss function.
[0098] γ is a weighting coefficient that adjusts the weight of the difference between the sentiment score Semo and the risk label Yrisk. This difference reflects the consistency between sentiment analysis and risk assessment, and its importance in the final loss function can be controlled by γ.
[0099] LXGB: The XGBoost model loss function is based on the loss calculated by the XGBoost algorithm. It is an efficient gradient boosting decision tree algorithm commonly used for classification and regression tasks.
[0100] LGraph: The graph neural network model loss function is based on the loss calculated by the graph neural network. The graph neural network can process graph structured data and is suitable for data modeling with complex relationships.
[0101] Semo: Sentiment score is obtained through sentiment analysis model, which reflects the emotional tendency in text or data.
[0102] Yrisk: The target output of the risk assessment model, usually a numerical value or category label, used to measure the risk level.
[0103] The technical highlights of the multimodal risk assessment model are:
[0104] Model interpretability design: Shapley value analysis is used to quantify the marginal contribution of each feature to the final risk, explaining the high-risk customer scores. A rule engine is designed to provide an auditable business rule layer, allowing for predefined if-then rule bases, such as "If the proportion of negative sentiment is greater than 30%, then the risk level is increased by 1," enhancing decision transparency.
[0105] Dynamic feedback and optimization module:
[0106] Based on the output of the multimodal risk assessment model and market feedback, the risk assessment logic of the multimodal risk assessment model is iteratively updated and optimized. The misjudged data is added to the training set of the multimodal risk assessment model to perform incremental training on the multimodal risk assessment model, thereby dynamically adjusting the edge weights according to the frequency of new entity relationships.
[0107] The dynamic feedback and optimization module incorporates a closed-loop feedback mechanism, including backtracking of risk misjudgment cases. This includes incorporating misjudged data, such as actual defaults but low-risk data from the model, into the training set, triggering incremental model training. It also includes knowledge graph weight adjustment, dynamically adjusting edge weights based on the frequency of new entity relationships. For example, the weight of the "policy risk" entity, which has appeared frequently recently, was increased.
[0108] The dynamic feedback and optimization module also generates visual reports, including risk transmission path diagrams. Based on the knowledge graph, these diagrams show the source of risk, such as negative public opinion about a company, and the transmission links to affected entities. These diagrams also include sentiment heat maps, which map regional sentiment distribution and assist banks in developing differentiated risk control strategies.
[0109] The advantages of the technical solution of the present invention are:
[0110] Cross-domain technology integration: Combining sentiment analysis with dynamic knowledge graphs, quantifying the impact of user emotions on financial risks through sentiment propagation models, and using knowledge graphs to capture dynamic associations between entities to improve the comprehensiveness and accuracy of risk identification.
[0111] Multimodal data fusion capability: Integrate multi-dimensional data such as emotions, transactions, and social interactions, adopt weighted average or neural network fusion strategies, and build a multimodal risk assessment model to significantly improve the robustness and adaptability of risk prediction.
[0112] Explainability and dynamic optimization: The Shapley value algorithm is used to quantify feature contributions, enhancing model transparency. In combination with the rule engine, the model and business rules are collaboratively verified to ensure the credibility and traceability of the evaluation results.
[0113] Real-time response and self-iteration capabilities: The attenuation factor and real-time data update mechanism based on the dynamic knowledge graph support dynamic tracking of risk transmission paths, and optimize model parameters through user feedback to achieve continuous system iteration.
[0114] Technical barriers and market value: Build multi-layered technical barriers covering emotional perception, dynamic graphs, and multimodal fusion to form differentiated competitive advantages, enabling rapid commercialization and market capture.
[0115] In summary, this patent combines technological innovation, application flexibility and commercial potential, providing financial institutions with a high-precision, explainable and dynamically optimized risk assessment tool to help improve risk prevention and control and decision-making efficiency.
[0116] The following is an example of the actual application of the core formula in the technical solution based on banking business scenarios.
[0117] Data normalization formula:
[0118] To standardize customer debt ratios, banks need to compare customer debt data of different dimensions, such as a debt ratio range of 0% to 500%, for use in credit assessment models. The formula used for this standardization is the Min-Max normalization formula. For example, if the original data contains Customer A's debt ratio = 380%, Customer B's debt ratio = 120%, the minimum debt ratio in the entire data set is 0% (Xmin), and the maximum debt ratio is 500% (Xmax).
[0119] The standardized calculation formula is:
[0120]
[0121] The reason for this calculation is to eliminate dimensional differences, so that the value of high-debt customers (such as A) is close to 1 and the value of low-debt customers (such as B) is close to 0, which facilitates unified processing of the model.
[0122] SEIR sentiment propagation model:
[0123] To predict bank run risks, when rumors of a "bank liquidity crisis" break out in a region, it is necessary to quantify the impact of the spread of negative emotions on bank runs.
[0124] Parameter settings:
[0125] Initial state: S = 0.95, that is, 95% of users have not been exposed to rumors, E = 0.04, I = 0.01, R = 0.
[0126] The transmission coefficient is β = 0.3, i.e. the transmission power of social media, the incubation period is σ = 0.2, and the recovery rate is γ = 0.1.
[0127] Deduction process (taking t=7t=7 days as an example):
[0128] Calculate daily variation (Euler discretization):
[0129] ΔS=-βS / Δt=-0.3×0.95×0.01×1=-0.00285;
[0130] ΔI=(σE-γ / )Δt=(0.2×0.04-0.1×0.01)×1=0.007;
[0131] Update the status value to:
[0132] S t+1 =0.95-0.00285=0.94715;
[0133] l t+1 =0.01+0.007=0.017;
[0134] As a result, the infection ratio II increased from 1% to 1.7%, triggering the risk control system to issue an early warning, and the bank needs to activate liquidity reserves.
[0135] Knowledge graph time decay factor:
[0136] Used to dynamically adjust industry weights. For example, in the early stages of the virus spread, the weight of historical retail data needs to be reduced to reflect the impact of policy changes in real time. Initial weight: The historical correlation weight of a company is: W0 = 0.8;
[0137] The attenuation calculation process is:
[0138] λ(30)=e -0.1×30 =e -3 ≈0.0498;
[0139] w new =w0×λ(t)=0.8×0.0498≈0.0398;
[0140] After 30 days, the historical weight is reduced to 4%, forcing the model to rely on real-time policy data, such as news on "suspension of offline business", to avoid misjudging loan risks.
[0141] Multimodal joint loss function:
[0142] This optimizes risk assessment models by integrating structured customer credit data with unstructured social media sentiment data to improve default prediction accuracy. Parameter settings: α = 0.6 (credit data weighting), β = 0.3 (knowledge graph weighting), γ = 0.1 (sentiment bias penalty). Single-sample loss: LXGB = 0.4, LGraph = 0.2, Semo = 0.8 (negative sentiment score), and Yrisk = 1 (actual default).
[0143] Deduction calculation:
[0144] L=0.6×0.4+0.3×0.2+0.1× / / 0.8-1 / / 2=0.24+0.06+0.1×0.2=0.32;
[0145] The sentiment score of 0.8 is close to the actual default label 1. The model reduces the loss of this sample and strengthens the importance of sentiment features.
[0146] Shapley Value Analysis: Explaining high-risk customer scores. Model interpretability is crucial in financial risk assessment. Customer A of a certain bank has a risk score of 0.85, indicating high risk. The model's decision-making basis needs to be explained. Input features: debt ratio (0.65), monthly income (40k), negative sentiment score (0.72), and industry risk index (0.55).
[0147] Defining the contribution of feature subsets: Debt ratio alone contributes to an increase in risk of 0.15, with a baseline risk of 0.5; income alone contributes to a decrease in risk of 0.12; emotion and debt ratio contribute jointly to an increase in risk of 0.28 (nonlinear interaction effect);
[0148]
[0149] φi is the main part of the formula, representing a quantity or function value related to element ii. It means to sum all subsets S of the set N except element i.
[0150] It is a variation of the combinatorial number, which represents the ratio of the number of combinations of |S| elements selected from n elements to the number of combinations of elements selected from n elements other than S and i. Here n is the number of elements in the set N.
[0151] [ν(S∪{i})-ν(S)] represents the change in the value of the function ν when element i is added to the set S.
[0152] Where N = {debt ratio, income, emotion, industry risk}.
[0153] Feature contribution decomposition:
[0154] Debt ratio: Calculate the marginal contribution in all subset combinations, and the final Shapley value is φ debt ratio = 0.25, φ debt ratio = 0.2534.
[0155] Negative emotion score: Due to the emotion contagion effect, its contribution is φemotion=0.18, φemotion=0.1823.
[0156] Industry risk: Influenced by the dynamic knowledge graph attenuation factor λ(t), the contribution φindustry = 0.12 and φindustry = 0.1212. Interpretation: Debt ratio is the largest risk driver, accounting for 29.4%, consistent with the bank's risk strategy rule of "highly indebted customers require manual review." The contribution of sentiment score exceeded expectations, indicating that negative comments made by customer A on social media, such as "broken funding chain," significantly impacted the score.
[0157] The rule engine and Shapley values are collaborating to verify the rule trigger logic: IF negative sentiment > 30% AND Shapley value > 0.15 → THEN risk level +1. Actual match: Customer A's negative sentiment ratio was 37% and the Shapley value was 0.18, triggering the rule to upgrade the risk level. Data verification: A review of 1,000 high-risk cases showed 89% consistency between the Shapley values and the rule engine. The discrepancies were primarily due to the rapid decay of industry risk factors, indicating that the λ(t) parameter requires optimization.
[0158] The dynamic feedback and optimization module's closed-loop feedback mechanism allows for risk misjudgment case retrospectives: adding misjudgment data (e.g., actual defaults but low risk in the model) to the training set to trigger incremental model training. Knowledge graph weight adjustment: Dynamically adjust edge weights based on the frequency of new entity relationships (e.g., increasing the weight of the recently frequently appearing "policy risk" entity).
[0159] Visual report generation: Risk transmission path map: Based on the knowledge graph, it shows the transmission chain from the source of risk (such as negative public opinion about a company) to the affected entity. Sentiment heat map: It marks the regional sentiment distribution on the map to assist banks in formulating differentiated risk control strategies.
[0160] Shapley value-driven feature contribution analysis, contribution calculation: For each feature xi, calculate its Shapley value φi to quantify its marginal contribution to the risk score f(x):
[0161]
[0162] Where N is the full set of features and S is the subset of features. The rule engine and model collaborate on decision-making. Rule types include: Hard rules: Regulatory compliance constraints, such as "If the capital adequacy ratio is less than 8%, the risk rating must be ≥ A," which directly overrule the model score; Soft rules: Empirical rules, such as "If short-term liabilities / current assets is greater than 1.2, the risk weight must be +0.1," which are weighted and integrated with the model score.
[0163] Self-optimization mechanism: The rule base dynamically adjusts rule weights through reinforcement learning, such as the PPO algorithm. For example, if a rule conflicts with the model decision 10 times in a row, the weight is triggered to decrease, such as the learning rate η = 0.01.
[0164] Case study: In the XX Bank corporate loan dataset, Shapley value analysis showed that the contribution of "supply chain risk transmission path length" was 0.18, significantly higher than the explanatory power of the "asset-liability ratio" in the traditional model, for example, with a contribution of 0.07; False alarm rate reduction: By intercepting abnormal decisions of the black box model, such as emotional misjudgment, the false alarm rate of mass defaults was further reduced from 9% to 5%.
[0165] Multimodal feature fusion based on the attention mechanism: data source classification and embedding, structured data such as corporate financial indicators are encoded into vectors vstructt through a fully connected layer; unstructured sentiment data such as news texts are extracted into sentiment feature vectors vsentiment through the domain pre-training model BERT-Fin; knowledge graph entity relationships are encoded into vectors vkg through graph embedding algorithms (such as TransE). Dynamic weight allocation, using the multi-head attention mechanism (Multi-headAttention) to calculate the weights of each modal feature,
[0166] αi = Softmax(Wq[vstruct∥vsentiment∥vkg]TWkT), where Wq and Wk are learnable parameters, and αi represents the weight coefficient of the i-th modality. Data support: Validated on the XX Bank dataset, the dynamic weighted fusion model achieved an AUC of 0.91, a 7% improvement over the fixed weight fusion model (AUC = 0.85).
[0167] Adaptive weighting driven by data quality and timeliness: Weight adjustment rules: For high-noise data, such as short social media posts, the confidence score (ConfidenceScore) is used to reduce its weight. For high-timeliness events, such as sudden policy adjustments, the weight is dynamically increased using the time decay factor λ(t) = e-kt, where k is the event's urgency factor. Case study: During a regional bank run in 2023, the confidence score of social media sentiment data was 0.32, indicating low quality, and the weight was automatically reduced to 0.15. However, the central bank's reserve requirement ratio cut had a time efficiency factor of k = 0.8, and the weight was increased to 0.63.
[0168] Real-time knowledge graph update mechanism: Event-driven update mechanism, event detection and classification, using event extraction models such as BERT-CRF to identify event types such as "policy release" and "corporate bankruptcy" from unstructured data such as news and announcements; event urgency classification: Level 1 events, such as policy adjustments, trigger real-time updates, while Level 2 events, such as changes in industry trends, trigger scheduled batch updates. Graph adjustment strategy, time decay factor: For historical relationship edges, such as corporate guarantee relationships, the weight is w(t) = w0·e -λtAttenuation, λ = 0.1; Dynamic relationship weight: Based on the intensity of the event impact, such as the quantitative score of policy intensity, the relationship weight between entities is adjusted. For example, the central bank's reserve requirement ratio cut event increases the correlation weight between "banking" and "real estate" from 0.5 to 0.8.
[0169] Incremental Update Algorithm: Utilizing the dynamic graph representation learning algorithm DyRep, this algorithm updates only the local subgraph affected by an event, rather than reconstructing the entire graph. This improves efficiency by 83%, reducing the update time from 6.2 hours to 1.1 hours. Real-time performance was verified using a test scenario simulating a debt default event involving a real estate company in 2023. The system updated the knowledge graph within 15 minutes of the event trigger, adding a new "debt default" node and associated edges, compared to 4.2 hours using traditional methods. Performance metrics: Event response latency ≤ 30 minutes for a Level 1 event, based on manual review sampling, and graph update accuracy ≥ 95%.
[0170] The following is a table showing the effects of the present invention when it is actually implemented:
[0171] Table 1:
[0172]
[0173] The above description is merely a preferred embodiment of the present invention. The scope of protection of the present invention is not limited to the above embodiment. All technical solutions based on the concept of the present invention are within the scope of protection of the present invention. It should be noted that for those skilled in the art, various improvements and modifications that do not depart from the principles of the present invention should also be considered within the scope of protection of the present invention.
Claims
1. A financial risk assessment system based on emotion perception and dynamic knowledge graph, characterized by: include: Multi-source data acquisition module, used to connect to multiple data sources and collect structured and unstructured data; And standardize the collected data to obtain standard data; Emotion perception engine, including text emotion analysis module and speech emotion recognition module, used to analyze standard data and quantify customer emotional characteristics; The dynamic knowledge graph construction module is used to extract entities and relationships between entities from standard data through the entity extraction model. It is also used to attenuate the weight of historical data according to the time dimension when incrementally updating the knowledge graph, and update the knowledge graph in real time. Multimodal risk assessment model, which is used to integrate standard data, sentiment features and knowledge graph embedding to output risk scores; The dynamic feedback and optimization module iteratively updates the risk assessment logic based on risk scores and market feedback, and performs incremental training on the multimodal risk assessment model through misjudgment data.
2. The financial risk assessment system based on emotion perception and dynamic knowledge graph according to claim 1 is characterized by: The standardization process includes: Duplicate data filtering: All collected data is converted into hash values through a hash algorithm to obtain a unique hash value for the data. The hash values of all data are compared. If data with the same hash value exists, it is considered duplicate and the duplicate data is then removed. Missing value filling: For all collected data, detect missing fields and fill in the missing fields; Data normalization: For the collected numerical data, the Min-Max normalization formula is used to obtain normalized data.
3. The financial risk assessment system based on emotion perception and dynamic knowledge graph according to claim 2 is characterized by: The Min-Max normalization formula is:
4. The financial risk assessment system based on emotion perception and dynamic knowledge graph according to claim 3 is characterized by: The text sentiment analysis module trains the pre-trained model to output fine-grained sentiment labels and confidence scores; models the sentiment propagation path to obtain a sentiment propagation model to simulate the negative sentiment diffusion path; The speech emotion recognition module extracts acoustic features from speech data, inputs them into a classification model, and outputs an emotion classification label.
5. The financial risk assessment system based on emotion perception and dynamic knowledge graph according to claim 4 is characterized by: The dynamic knowledge graph construction module includes an event extraction and entity linking module and a graph incremental update module. The event extraction and entity linking module is used to extract entities and corresponding relationships through the entity extraction model, calculate the similarity of all entities, and determine the different meanings of multiple entities with the same meaning and a single entity based on the similarity; The graph incremental update module is used to incrementally update the knowledge graph and decay the weight of historical data according to the time dimension. The decay formula is: λ(t) = e -kt , where k is the decay rate and t is the time interval.
6. The financial risk assessment system based on emotion perception and dynamic knowledge graph according to claim 5 is characterized by: The architecture of the multimodal risk assessment model integrates a learning framework to process structured data; integrates knowledge graph embedding to capture the relationship between entities; and also uses sentiment features as an independent input branch.
7. The financial risk assessment system based on emotion perception and dynamic knowledge graph according to claim 6 is characterized by: The dynamic feedback and optimization module is equipped with a closed-loop feedback mechanism, including backtracking of risk misjudgment cases and adjustment of knowledge graph weights.
8. The financial risk assessment system based on emotion perception and dynamic knowledge graph according to claim 7 is characterized by: The dynamic feedback and optimization module is also used to generate visual reports, including risk transmission path diagrams and sentiment heat maps.
Citation Information
Cited By
Stock price collapse risk analysis reference system based on stock market interaction data identification
CN121052933A
Credit rating report generation method and device based on multi-modal data and storable medium
CN121257479A
Data analysis method and system fusing knowledge graph and deep learning
CN121327780A
Digital human interaction method, device, equipment and program product
CN121999777A
Insurance policy customer allocation method, system and equipment based on relation graph, and medium
CN122089491A