Data structure adaptive visualization method based on large language model

By building a dynamic financial knowledge graph and a large language model to process data, the problems of low generation efficiency and insufficient accuracy in financial risk analysis are solved, efficient and accurate visualization results and automated analysis are achieved, and the efficiency and accuracy of financial risk management are improved.

CN120744145AActive Publication Date: 2025-10-03ZHEJIANG FULIN TECH CO LTD

Patent Information

Application Number
CN202511270675.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-08
Publication Date
2025-10-03
Estimated Expiration
2045-09-08

AI Technical Summary

Technical Problem

Existing technologies generate inefficiencies and inaccuracies in financial risk analysis, requiring analysts to spend a significant amount of time on manual code review and verification to ensure the accuracy of visualization results, limiting the technology's usefulness in fast-paced, high-precision financial risk management scenarios.

Method used

By building a dynamic financial knowledge graph, using a large language model to process quantitative data and text data, updating risk transmission weights in real time, generating multiple synthetic risk data paths, and providing interactive risk simulation and automatic attribution analysis of visualization results, including counterfactual analysis and attribution reports.

Benefits of technology

It improves the accuracy and efficiency of visualization results, enables the rapid generation of visualization results of multiple future possibilities, automates the analysis process, reduces the time for manual verification, and improves the work efficiency of analysts.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120744145A_ABST
    Figure CN120744145A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of financial knowledge maps, in particular to a large language model-based data structure adaptive visualization method, which comprises the following steps of: acquiring internal quantitative data and external unstructured text data of an investment portfolio; processing the quantized data and the text data by using a large language model, constructing a dynamic financial knowledge graph, and updating a risk conduction weight between entities in the dynamic knowledge graph in real time by using a causal event extracted from the text data; receiving a risk impact hypothesis input by a user, and performing prospective simulation based on the dynamic financial knowledge graph and the risk impact hypothesis to generate a plurality of synthetic risk data paths; generating a visual result based on the plurality of synthetic risk data paths; and in response to the selection of the user on the specific simulation result on the visual result, executing anti-fact analysis, and generating an attribution report containing the contribution degree, so that the interactive risk simulation and automatic attribution analysis of the visual result are realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of financial knowledge graphs, and specifically to a data structure adaptive visualization method based on a large language model. Background Art

[0002] In the field of financial risk analysis, analysts often need to conduct ad hoc, exploratory data analysis on complex data such as market volatility and credit risk. To improve the efficiency of this process, the industry has begun exploring technologies that leverage large language models to directly convert natural language queries into visual code. The goal is to quickly generate visual charts for testing hypotheses or communicating findings.

[0003] However, in practical applications of financial risk analysis, this approach suffers from low generation efficiency. Large language models can generate inaccuracies and ambiguities when interpreting analyst intent. Because financial decision-making requires extremely accurate data visualization, any error can lead to serious risk misjudgments. Therefore, these automatically generated charts cannot be directly trusted and used. They must invest significant time in manual code review, verification, and repeated debugging to ensure the final presentation is accurate and reliable. This necessary verification and correction process constitutes a major bottleneck, limiting the technology's usefulness in fast-paced, high-precision financial risk management scenarios.

[0004] To this end, a data structure adaptive visualization method based on large language model is proposed. Summary of the Invention

[0005] The present invention aims to provide a data structure adaptive visualization method based on a large language model. This method implements interactive risk simulation and automatic attribution analysis of visualization results through a dynamic knowledge graph and interactive simulation attribution. First, internal quantitative data and external unstructured text data of an investment portfolio are acquired. The large language model is used to process the quantitative and textual data to construct a dynamic financial knowledge graph. Causal events extracted from the text data are used to update the risk transmission weights between entities in the dynamic knowledge graph in real time. Risk shock hypotheses input by the user are received, and forward-looking simulations are performed based on the dynamic financial knowledge graph and the risk shock hypotheses to generate multiple synthetic risk data paths. Visualization results are generated based on the multiple synthetic risk data paths. In response to the user's selection of a specific simulation result in the visualization, counterfactual analysis is performed to generate an attribution report containing contribution, thus implementing interactive risk simulation and automatic attribution analysis of visualization results.

[0006] To achieve the above object, the present invention provides the following technical solutions: A data structure adaptive visualization method based on a large language model, comprising: Acquire internal quantitative data and external unstructured text data of the investment portfolio; process the external unstructured text data using a large language model to extract structured causal events; Constructing a dynamic financial knowledge graph using internal quantitative data and the structured causal events, and updating risk transmission weights between entities in the dynamic financial knowledge graph in real time using the structured causal events; receiving a risk impact hypothesis input by a user, performing a forward-looking simulation based on the dynamic financial knowledge graph and the risk impact hypothesis to generate a plurality of synthetic risk data paths representing a variety of future possibilities; generating a visualization result based on the plurality of synthetic risk data paths, wherein the visualization result is bound to a traceability link; and in response to the user selecting a specific simulation result on the visualization result, performing a path sensitivity analysis to identify a set of driving factors; A modified simulation method is executed based on the driving factors to perform counterfactual analysis, determine the quantitative contribution of the driving factors to the specific simulation results, and generate an attribution report.

[0007] Preferably, the step of obtaining internal quantitative data and external unstructured text data of the investment portfolio, processing the external unstructured text data using a large language model, and extracting structured causal events includes: The internal quantitative data includes daily portfolio holdings details, historical trading records, and daily profit and loss time series; the external unstructured text data includes regulatory agency announcements, listed company financial reports, and financial news articles; A large language model is used to perform named entity recognition and relationship extraction on the external unstructured text data, and the text data is converted into structured causal events including event types, participating entities and influencing parameters.

[0008] Preferably, the process of constructing a dynamic financial knowledge graph includes: Initializing a base knowledge graph containing a predefined financial ontology, wherein the financial ontology defines entity types, attribute types, and relationship types; The structured causal events are fused with the base knowledge graph. The fusion process includes creating and updating nodes corresponding to entities in the structured causal events in the knowledge graph, creating and updating relationship edges representing the structured causal events between nodes, and adjusting the risk transmission weights associated with the relationship edges and nodes according to the impact parameters of the structured causal events; adding a time-effect decay function to the updated risk transmission weights, and assigning a confidence score to the relationship edges and weights according to the information source of the text data.

[0009] Preferably, the step of adjusting the risk transmission weights associated with the relationship edges and nodes according to the impact parameters of the structured causal events includes: Based on the impact scope of the structured causal event, a group of virtual economic agents related to the structured causal event are instantiated from a predefined agent library, each agent being assigned a specific role and behavior pattern; Presenting the structured causal event as input information to the instantiated group of agents, and using the large language model to drive each agent to generate a decision and behavioral response to the structured causal event based on its role and behavior pattern; Summarize the decisions and behavioral responses generated by the agent group, quantify the summary results into adjustment values, and use the adjustment values ​​to update the risk transmission weights.

[0010] Preferably, the step of performing forward-looking simulation based on the dynamic financial knowledge graph and the risk shock hypothesis includes: Parsing risk impact hypotheses inputted in natural language into structured intervention parameter objects using the large language model, and generating clarifying questions based on the dynamic financial knowledge graph for user confirmation when there is ambiguity in the hypothesis parsing; Applying the structured intervention parameter object to the dynamic financial knowledge graph to adjust the quantitative attributes of graph elements in the dynamic financial knowledge graph that are associated with the intervention parameter object; Perform Monte Carlo simulation, use the quantitative properties of the adjusted graph elements in the dynamic financial knowledge graph to define the random process and related parameters of the Monte Carlo simulation, and generate the multiple synthetic risk data paths.

[0011] Preferably, the step of generating a visualization result includes: For each of the multiple synthetic risk data paths and statistically derived elements, a traceability link pointing to the corresponding driving factor in the dynamic financial knowledge graph is established; on the visualization result, the traceability link is bound to the corresponding visualization element, and the user triggers a counterfactual analysis of the driving factor by interacting with the visualization element.

[0012] Preferably, the process of establishing a traceability link to a corresponding driving factor in the dynamic financial knowledge graph includes: performing a path sensitivity analysis on a specified path in the synthetic risk data path to obtain a quantitative impact score of each driving factor, calculating the sum of the absolute values ​​of the quantitative impact scores of all driving factors, sorting the driving factors in descending order, and selecting, from the top, the driving factor whose cumulative sum of the absolute values ​​of the quantitative impact scores reaches a preset percentage of the total as the main driving risk factor; Based on the analysis results, a data structure including ranked driver identifications and quantified impact scores is generated as the traceability link.

[0013] Preferably, the step of performing counterfactual analysis includes: performing a modified simulation on the main driving risk factors by temporarily suppressing their risk transmission weights in the dynamic financial knowledge graph; comparing the specific simulation results with the modified simulation results to calculate the quantitative contribution of the main driving risk factors.

[0014] Compared with the prior art, the present invention has the following beneficial effects: 1. By constructing a dynamic financial knowledge graph that integrates real-time text events, this approach provides deep business context for analysis. Compared to existing techniques where large language models generate a large number of ambiguous and erroneous visualization codes due to a lack of understanding of isolated data, this approach builds analysis on a structured, connected knowledge network that reflects the latest market dynamics. This ensures the accuracy and relevance of the analysis from the very beginning, resulting in higher quality initial visualization results and significantly improved efficiency.

[0015] 2. By introducing forward-looking simulation capabilities based on natural language hypotheses, this method transforms a single visualization request into a one-time, multi-path exploration of future scenarios. By describing a complex risk hypothesis in natural language, thousands of simulations can be automatically executed, generating a single visualization of the probability distribution of multiple future possibilities. This "batch" exploration model compresses the repetitive chart generation process, which previously took hours or even days, into a single, efficient simulation and visualization process, significantly improving the overall efficiency of exploratory risk analysis.

[0016] 3. By providing interactive counterfactual analysis and instant attribution reporting functions, this method solves the core efficiency bottleneck in the background technology caused by the need for manual review and debugging of results. In the previous workflow, even if the chart was generated correctly, analysts still had to spend a lot of time tracing the data and checking the logic to understand the reasons behind a certain abnormal result. The present invention fully automates this process: users can select specific simulation results on the visualization results to perform counterfactual analysis immediately and generate an attribution report containing the contribution of each risk factor. This replaces the time-consuming and labor-intensive manual "error correction" with a one-click automatic "explanation", effectively shifting the focus of the analysis work from tedious technical verification to core business insights and decision-making, thereby significantly optimizing the workflow and improving analysis efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] Figure 1 A flow chart of a data structure adaptive visualization method based on a large language model provided by an embodiment of the present invention; Figure 2 A schematic diagram of the process of constructing a dynamic financial knowledge graph provided by an embodiment of the present invention; Figure 3A schematic diagram of a process for generating a visual interface is provided for an embodiment of the present invention. DETAILED DESCRIPTION

[0018] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0019] See also Figures 1 to 3 The present invention provides a data structure adaptive visualization method based on a large language model, and the technical solution is as follows: Example 1: In order to improve the efficiency and accuracy of visualization code generation in the practical application of financial risk analysis, a data structure adaptive visualization method based on a large language model is introduced. The specific process is as follows: Figure 1 shown.

[0020] Acquire internal quantitative data and external unstructured text data of an investment portfolio; process the external unstructured text data using a large language model to extract structured causal events.

[0021] The initial data processing phase involves an automated data aggregation process. Through a secure API, the system calls the institution's back-office portfolio management system, executing pre-defined database queries to retrieve internal quantitative data, including daily portfolio holdings, historical trading records, and daily profit and loss time series. All retrieved internal data is timestamped to a unified snapshot and stored in a database. Furthermore, the system continuously polls APIs from multiple external data sources. Based on the list of financial instrument IDs in the holdings, it then initiates batch requests to market data providers' real-time data APIs to retrieve market time series for these instruments, as well as key market benchmarks such as the S&P 500, VIX, and US Treasury bond yields. A keyword list related to the company holdings is maintained and used to filter queries against the News API and the US Securities and Exchange Commission's EDGAR database, obtaining raw content from regulatory announcements, public company earnings reports, and financial news articles published within the past 24 hours. After all raw data is aggregated in a temporary database, the retrieved external unstructured text data is fed into a large language model fine-tuned in the financial domain to perform causal event extraction. This processing process is a natural language processing task that uses named entity recognition technology to read through all texts to identify key financial entities, and uses relationship extraction technology to identify the actions and relationships between these key financial entities.

[0022] The final result is standardized structured causal event data objects, each of which contains a set of predefined fields to clearly define the event type, participating entities, and specific impact parameters. The original text information is converted into machine-readable structured event data and, together with the quantitative data obtained at the same timestamp, forms a complete input for subsequent steps.

[0023] By accurately converting external unstructured text into standardized structured causal event data, it provides high-quality, directly usable input for the subsequent construction of dynamic financial knowledge graphs and real-time weight updates. The aggregation and structuring of multimodal data provide the necessary data support for subsequent realistic forward-looking simulations and precise attribution analysis.

[0024] A dynamic financial knowledge graph is constructed using internal quantitative data and the structured causal events, and the risk transmission weights between entities in the dynamic financial knowledge graph are updated in real time using the structured causal events.

[0025] First, load a predefined base financial knowledge graph. The base financial knowledge graph contains a financial ontology that defines core entity types in financial markets and typical relationship types between these core entity types. These core entity types include, but are not limited to, organizations (subtypes of which include listed companies, financial institutions, and regulatory agencies); financial instruments (subtypes of which include stocks, bonds, and derivatives); geopolitical entities (representing countries or economic regions); individuals (specifically, executives or politicians with significant market influence); macroeconomic indicators (such as the Consumer Price Index (CPI) or unemployment rate); and industries (used to categorize organizations). Typical relationship types include affiliation and ownership relationships, such as "employed by," "is a subsidiary of," and "issued by"; economic connections (such as "has a supplier," "competes with," and "belongs to the industry of"); regulatory and rating relationships (such as "regulated by" and "rated by"); and causal influence relationships (such as "influences" and "cause of"). The base financial knowledge graph provides an initial, structured financial framework. Subsequently, a data fusion operation is performed to map internal quantitative data, including portfolio holdings details, into the knowledge graph, establishing relationships between the current portfolio and its financial assets. Simultaneously, structured causal event data is integrated. For each causal event record, entity nodes related to the event are identified and created in the knowledge graph. Edges are then established and modified between these nodes, representing the causal relationships described by the event.

[0026] The dynamic adjustment mechanism for risk transmission weights is implemented through a subroutine based on multi-agent simulation. This subroutine is initiated when the impact of a causal event extracted from text on the risk network needs to be quantified. First, a group of virtual economic agents is instantiated from a pre-defined agent library based on the nature of the event. These agents are assigned different market roles, such as traders, analysts, and managers, each with specific behavioral patterns and decision-making logic. Next, the structured causal event is presented as input to all instantiated agents. Driven by a large language model, each agent independently generates a decision or behavioral response to the event based on its assigned role. These responses are output as natural language text, simulating the diverse reactions that different market participants might have to the same information in the real world. The decision responses of all agents are then collected and aggregated. To accurately quantify these responses, a deep learning model based on the BERT architecture is employed. This BERT model has been pre-trained using domain adaptation on a large corpus of financial text and fine-tuned for two downstream tasks: sentiment analysis and intent classification. Specifically, each agent's decision response text is input into the fine-tuned BERT model. The model outputs two results: a sentiment score ranging from -1 to 1, representing the response's negative or positive bias; and the behavioral intent category that the response most closely matches. Preset categories include "buy," "sell," "hold," and "wait and see." After analyzing all agent responses, this method performs a multi-step aggregation and quantification process to calculate the final weight adjustment. The first step is to calculate the base sentiment impact value. First, the arithmetic mean of all sentiment scores is calculated. This average is then multiplied by a base sentiment impact coefficient to generate a base adjustment component that reflects overall market sentiment. The base sentiment impact coefficient is determined through backtesting analysis of historical event data and market reactions. The second step is to calculate the comprehensive behavioral impact value. This aims to comprehensively quantify the market behavior structure composed of four different behavioral intents. First, the proportion of responses identified as "buy," "sell," "hold," and "wait and see" is calculated. These proportions are then weighted and summed to calculate a comprehensive behavioral impact component. Specifically, the "buy" percentage is multiplied by a positive behavior coefficient, the "sell" percentage is multiplied by a negative behavior coefficient, the "hold" percentage is multiplied by a stable behavior coefficient, and the "wait-and-see" percentage is multiplied by an uncertainty coefficient. The algebraic sum of these four products constitutes an additional adjustment component that reflects the complex market behavior structure. The four behavioral coefficients—positive, negative, stable, and uncertain—serve as model parameters in this method. Their specific values ​​are determined through backtesting analysis based on historical financial event data and their corresponding market risk fluctuations. The third step is to combine these factors to generate the final adjustment value.The weighted sum of the basic emotional impact value and the comprehensive behavioral impact value calculated in the first two steps is used to obtain the final weight adjustment value. The weight adjustment value is ultimately used to update the risk transmission weight of the relevant relationship edge in the knowledge graph, thereby converting macro-level, qualitative event information into a quantitative impact on the strength of the micro-level risk transmission path.

[0027] By simulating a micro-market composed of multiple actors, we can capture the diverse reactions and interactive effects of different real-world market participants when faced with the same event. By leveraging an intelligent agent driven by a large language model, its decision-making logic is no longer constrained by rigid, pre-set rules, but rather possesses human-like reasoning capabilities for processing complex and ambiguous information. Ultimately, by aggregating "group consensus" to quantify impact, this not only makes the weight adjustment process more transparent and traceable, but also brings the results closer to the real market dynamics emerging from the actions of a large number of heterogeneous actors, significantly improving the accuracy of the entire dynamic risk analysis model and its simulation of market complexity.

[0028] Finally, to ensure the timeliness and reliability of the knowledge graph model, all risk transmission weights updated due to events are assigned two attributes: a time-attenuation function, which simulates the natural fading of information's influence over time; and a confidence score, whose value depends on the authority of the original source of the event information. Through these operations, this method constructs and maintains a dynamic financial knowledge graph that can adaptively reflect market changes, providing a solid foundation for subsequent forward-looking analysis.

[0029] By introducing a weight adjustment mechanism based on multi-agent simulation, this method can transform sudden, qualitative news events into quantitative adjustments to risk transmission path weights in real time. This process, utilizing agent-based simulation driven by a large language model, imbues the complex connotations of news events with a dynamic and repeatable quantitative framework, aiming to improve the timeliness and market adaptability of risk models. Furthermore, this method provides greater transparency for the quantification of risk transmission weights. Rather than relying on fixed mapping rules, it simulates a micro-market with multiple actors and extracts "group consensus," providing an interpretable generative logic based on complex behavior for the origin of weight adjustments. This enhances the inherent rationality of risk analysis results and the credibility of decision-making.

[0030] Receive risk impact assumptions input by the user, perform forward-looking simulations based on the dynamic financial knowledge graph and the risk impact assumptions, and generate multiple synthetic risk data paths representing multiple future possibilities; generate visualization results based on the multiple synthetic risk data paths, and bind the visualization results to traceability links.

[0031] First, user input is processed by an intelligent natural language hypothesis parsing subroutine. This subroutine receives free-text user input and invokes a large language model configured with a specific tool set. This model operates in proxy mode. Its core task is not to generate free text, but rather to select the most appropriate internal function from a predefined tool library based on the semantics of the user input and generate formatted parameters for it. This process decomposes a complex, free-form user input into multiple standardized, machine-readable intervention parameter objects. The subroutine also incorporates a knowledge graph-based ambiguity clarification mechanism. During the parsing process, if the large language model identifies ambiguity in the user's statement, such as "If a conflict occurs among major oil-producing countries, oil prices will fluctuate dramatically," the clarification mechanism is triggered. It proactively queries the dynamic financial knowledge graph for specific information related to these ambiguous terms and generates clarification questions accordingly. For example, the user will be presented with: "Your question has been identified. Please confirm or modify the following details: "Does 'major oil-producing countries' refer to the Organization of Petroleum Exporting Countries and its allies (OPEC+)?", "Does 'oil' refer to 'West Texas Intermediate (WTI)' or 'Brent Crude Oil'?", and "Does 'severe volatility' mean 'a 20% price increase in one day, or an increase in realized volatility to 80% in the next month'?" This closed-loop interaction ensures the accuracy of simulation input.

[0032] Furthermore, each intervention parameter object is traversed, and the node corresponding to its "target entity" is located in the knowledge graph through index query. Then, based on the "target attributes" and "impact description" defined in the object, the numerical attributes of the node or its associated edges are temporarily modified programmatically. For example, an intervention object related to a crude oil price shock will directly adjust the price attribute baseline value of the "WTI Crude Oil" node in the knowledge graph. This operation creates an in-memory knowledge graph snapshot representing the initial state of the market after a specific external shock, which serves as the starting point for subsequent simulations.

[0033] Based on the modified knowledge graph snapshot, a high-dimensional Monte Carlo simulation is performed, systematically converting the graph's topology and weights into the mathematical parameters required for the simulation. The specific process is as follows: First, a subgraph containing all key entities directly or indirectly affected is extracted from the graph snapshot. Then, a weighted adjacency matrix is ​​calculated for this subgraph, where the elements represent the risk transmission weights between entities. To capture the networked, multi-step risk transmission paths, graph theory algorithms are further applied. First, the corresponding graph Laplacian matrix is ​​calculated based on this weighted adjacency matrix. In graphical model theory, the graph Laplacian matrix is ​​considered the precision matrix of a Gaussian Markov random field. By taking the pseudo-inverse of this graph Laplacian matrix, a covariance matrix is ​​generated that comprehensively reflects all direct and indirect connections in the graph. This covariance matrix is ​​then used as the core driving parameter of a multivariate geometric Brownian motion random process, which is used to perform thousands of independent simulations. Each simulation generates a synthetic risk data path representing the portfolio's profit and loss over a predetermined period of time. Together, these paths form the basis for a probabilistic analysis of future possibilities.

[0034] By introducing a natural language parsing function driven by a large language model and equipped with an ambiguity clarification mechanism, the complex operation of traditional risk simulation, which requires users to manually enter precise parameters, is transformed into an intuitive natural language dialogue. This greatly reduces the threshold for using professional risk analysis tools, enabling decision makers without a quantitative background to conduct in-depth exploration. Secondly, this method systematically transforms the topological structure and weights of the dynamic financial knowledge graph into the core parameters of the Monte Carlo simulation through a graph algorithm. This ensures that the initial state and evolution path of the simulation not only reflect historical statistical laws, but also can incorporate the dynamic impact of unexpected events and user-defined assumptions in real time, making the simulation results more closely resemble the complexity of the real world.

[0035] In response to a user's selection of a specific simulation result on the visualization result, a path sensitivity analysis is performed to identify a set of driving factors; a modified simulation method is performed based on the driving factors to perform a counterfactual analysis, determine the quantitative contribution of the driving factors to the specific simulation result, and generate an attribution report.

[0036] First, all synthetic paths are statistically processed to generate a set of basic visualization charts. These charts include a probability density plot showing the final profit and loss distribution, a risk path chart showing the evolution of risk over time, and a dashboard containing key risk indicators such as value at risk and expected loss.

[0037] Furthermore, an efficient sensitivity analysis is automatically performed for a specified extreme loss path. This utilizes a gradient calculation technique based on the adjoint state method. The core principle of this gradient calculation technique stems from the chain rule, which achieves efficient gradient calculation by constructing a reverse calculation graph accompanying the original simulation process. Specifically, it starts with the final output (i.e., the path's profit and loss value) and propagates the gradient backwards, layer by layer, from the output end back to the input end. The significant advantage of this reverse calculation is that, regardless of the number of input parameters, the partial derivatives of the final output with respect to all input parameters can be obtained simultaneously through a single forward calculation and a single reverse calculation. The calculated partial derivative values ​​are directly used as the quantitative impact score of each driver factor on the specified path. All driver factors are sorted in descending order by the absolute value of their impact scores, and the top-ranked group is selected as the primary driver risk factors. Specifically, the cumulative contribution method is adopted. First, the sum of the absolute values ​​of the impact scores of all driving factors is calculated, and then they are added up one by one in descending order. When the cumulative value reaches a preset proportion of the total, all driving factors in this accumulation process are identified as major driving risk factors. The preset proportion is found through statistical methods to find the proportion that performs best in historical data.

[0038] During gradient backpropagation using the adjoint state method, the computational graph not only transmits gradient values ​​but also maintains a "path backpointer" for each node. As the gradient propagates from a node to its adjacent predecessor, the predecessor node records the adjacent node with the largest contribution from the gradient source as its backpointer. After the backpropagation process is complete, starting from the final portfolio P / L node, the backpointers are traced back in reverse until the node with the highest impact score is reached. This sequence of nodes linked by backpointers is identified as the key risk transmission path. This path, in a data-driven manner, reveals the primary chain of transmission that risks traverse from their source to their final point of impact. Finally, the results of these two tasks are encapsulated into an enhanced traceability link data structure using a universal data exchange format. The data structure contains the following core fields: "Main Driving Risk Factors" list: each entry records in detail the unique identifier, type, name and quantitative impact score of a driving factor; "Key Conduction Path" list: is an ordered array, each element in the array is an entity object representing a node on the path, which fully reproduces the step-by-step process of risk conduction.

[0039] By applying a gradient calculation technique based on the adjoint state method, a single reverse calculation can simultaneously obtain the quantitative contribution of all input factors to a specific output, significantly improving the computational efficiency of attribution analysis in high-dimensional parameter spaces. Furthermore, this method goes beyond identifying the primary risk drivers. By introducing a path backtracking pointer mechanism, it can also accurately track and extract the key paths of risk transmission from complex knowledge graph networks. This provides users with in-depth information on how risks form and spread, thereby improving the interpretability of the analysis from identifying a single cause to understanding the entire chain of influence.

[0040] When an analysis is completed, a structured data object containing the analysis results is generated, such as a traceability-linked data structure containing a list of "key risk drivers." This triggers the chart recommendation step, which first parses the metadata and structure of the data object, identifying key features such as data dimensions, field types, and basic statistical properties. It then determines the underlying analytical intent, including "comparative ranking" and "trend analysis." The parsed data structure features and the determined analytical intent are then combined to form a structured request, which is sent to the large language model. The large language model receives the request, executes the decision, and returns a standardized object containing the optimal chart type recommendation and its logical justification. Finally, the visualization rendering engine receives this recommendation, automatically selects and configures the appropriate chart library components, and, using the original structured data object as input, generates and presents expert-level visualizations that are most suitable for interpreting the current analysis results. This approach enables customized visualization recommendations for different analyses, improving user viewing efficiency.

[0041] In front-end applications, these linked data structures are deeply bound to corresponding visualization elements. For example, when rendering using a graphics library, they are stored as metadata in the data attributes of the visualization curve element representing the extreme loss path. In this way, the visualization results constructed by this method are no longer static images, but rather an intelligent interactive interface rich with back-end model information. The front-end application is configured to listen for user click events on these elements. Once an interaction is detected, it immediately reads the linked information stored in its metadata and uses the list of key driving risk factors encapsulated therein as a parameter to automatically initiate a targeted counterfactual analysis request to the back-end. This completes the entire process from data visualization to triggering deep attribution analysis through user interaction, providing clear input and initiation signals for the final counterfactual analysis step.

[0042] By building and binding "traceability links" containing deep model information, the transparency and explainability of risk analysis are significantly enhanced. Compared to traditional visualizations that focus on presenting simulation results, the core of this method is to empower users to explore the underlying causes of the results. By introducing efficient sensitivity analysis and gradient path tracing technology based on the adjoint state method, this method can instantly and quantitatively identify the main driving risk factors for a given simulation result and further trace the key paths of risk transmission in the knowledge graph. It provides users with clear insights from macro results to micro causes, thereby improving the basis and quality of decision-making.

[0043] When a user selects a visualization element associated with a specific simulation result on the generated visualization interface—for example, a path representing extreme losses—the front-end application reads the "traceability link" data structure bound to the visualization element. The list of "primary risk drivers" contained in the link is then sent to the back-end via an API request, triggering a counterfactual analysis. The analysis process uses a "one-by-one suppression" correction simulation approach. For each primary risk driver received from the front-end, the back-end performs a separate correction simulation based on a post-impact knowledge graph snapshot of the path that led to the extreme losses.

[0044] In practice, the analysis engine temporarily suppresses the impact of a single driver. For example, if the driver is a node attribute, its value is reset to its original pre-impact value. If the driver is a relationship weight, its risk transmission weight is reset to zero. If the driver is an initial shock introduced by the user, the shock is temporarily removed. After suppressing the impact of the single driver, the analysis engine reruns the forward-looking simulation using this locally modified knowledge graph as a starting point. This modified simulation generates a new set of synthetic risk data paths. The analysis engine then calculates the average final profit and loss of this modified set of paths and compares it with the profit and loss of the original extreme loss path selected by the user. This method compares the average final profit and loss of the newly generated set of synthetic risk data paths with the final profit and loss of the original extreme loss path, quantifying the contribution of the current single driver to the current extreme loss. The modified simulation and contribution calculation are then repeated for each major driver risk factor in the list. Finally, the quantified contributions of all major driver risk factors are aggregated and normalized to generate a clear attribution report. The attribution report is presented to the user in the form of charts and text summaries, intuitively showing the extent of each driver's responsibility for causing the specific simulation results.

[0045] By employing a "suppress one-by-one" correction simulation approach, users can instantly and interactively verify the true impact of each key driver after selecting a simulation result of interest. By re-simulating a control scenario after suppressing a specific risk factor, the absolute contribution of that factor to the original result can be precisely calculated. This clear and intuitive attribution approach provides users with a data-driven attribution report, greatly enhancing confidence in the entire risk model and providing the most direct and powerful data support for developing precise hedging or management strategies for specific risk sources.

[0046] By constructing a dynamic financial knowledge graph that integrates real-time text events, this invention provides business context for analysis, improving the accuracy and relevance of analysis from the source. To improve efficiency, its forward-looking simulation function transforms repetitive chart generation work into a single multi-path scenario exploration. In addition, the method of the present invention provides interactive counterfactual analysis and attribution reporting functions, replacing manual review and debugging with automated attribution. Users can select a specific simulation result in the visualization to obtain a quantitative contribution report of its driving factors, shifting the focus of work from technical verification to business decision-making, further improving the efficiency of the workflow.

[0047] Example 2: In order to improve the visualization clarity of financial data structures, a data structure adaptive visualization method based on a large language model is introduced.

[0048] This method uses an API from a mainstream financial data provider to retrieve daily portfolio details, historical trading records, and daily profit and loss time series for the current investment portfolio as of the previous trading day. A continuously running online information collection module, through real-time monitoring of multiple mainstream financial media outlets, captured a key news item: "A national regulatory department has issued a draft for comment on adjustments to new energy vehicle subsidy policies."

[0049] The original content of the news item was fed into a large language model fine-tuned with a large amount of Chinese financial corpus. The model performed causal event extraction, generating a structured causal event data object. The core content of the object was parsed into the event type "industry regulatory policy changes," the participating entities "national competent authorities" and "new energy vehicle industry," and the influencing parameters "subsidy policy adjustments" and "potential negative impacts."

[0050] This causal event triggers a dynamic update of the knowledge graph. A virtual economic agent analysis group, comprised of multiple roles, is instantiated. Upon receiving this event information, the agent group generates a variety of decision responses. By summarizing and analyzing these responses, a short-term negative consensus on the new energy vehicle industry is identified and quantified as a specific adjustment value. This adjustment value is used to increase the weight of the negative risk transmission from the "national competent authority" node to the "new energy vehicle industry" node in the knowledge graph, while also increasing the weight of the competitive relationship between the nodes of leading enterprises within the industry.

[0051] A user enters the query "If the final draft of the subsidy policy next month confirms a significant reduction, what will be the impact on my holdings?" into the natural language interface. The large language model's hypothesis parsing subroutine converts this query into structured intervention parameters. After a clarification conversation with the user and confirmation of the specific percentage of "significant reduction," the impact is applied to the knowledge graph snapshot, specifically by temporarily and significantly lowering the "policy support" attribute of nodes related to the "new energy vehicle industry" in the knowledge graph. Based on this modified knowledge graph snapshot, a Monte Carlo simulation consisting of thousands of iterations is initiated, generating thousands of synthetic risk data paths for the portfolio over the next quarter.

[0052] The simulation results were generated in real time into an interactive dashboard. The probability density plot in the dashboard showed that the portfolio's expected profit and loss distribution exhibited a significant negative skewness. The worst-case scenario (the fifth percentile path) in the risk path fan chart indicated the potential for significant future losses. An automated sensitivity analysis of this worst-case scenario, using the adjoint state approach, identified its key drivers and transmission pathways. This information was then encapsulated as a traceable link to the corresponding loss path curve on the fan chart.

[0053] On the dashboard, the user clicks on the worst-case loss path. This interactive action immediately triggers the counterfactual analysis process. The backend analytical engine reads the links associated with this path and performs a "one-by-one" correction simulation on the key risk drivers contained within it.

[0054] The resulting attribution report, presented as a waterfall chart, clearly reveals the factors contributing to the extreme losses: the user's hypothesis of a "significant subsidy cut" was the largest negative contributor; the updated knowledge graph revealed that the increased weight of competition among industry leaders, amplifying the negative impact within the industry, was the second largest contributor; and the high stock price volatility of the power battery manufacturer at the heart of the supply chain was the third largest contributor. Through this comprehensive process, fund managers not only quantified policy risks but also gained a deeper understanding of the risk transmission mechanism within the industry, enabling them to develop more refined risk response strategies.

[0055] Example 3: To enhance the intelligence, personalization, and risk interpretability of individual users' financial planning, this paper introduces a data structure adaptive visualization method based on a large language model, which is applied to the intelligent wealth management advisor scenario.

[0056] First, the system comprehensively captures the user's financial status and personal goals. Through authorized APIs, the system automatically aggregates the user's internal quantitative data. For example, a user with total assets of 500,000 yuan might have a portfolio consisting of 300,000 yuan in stocks, 100,000 yuan in funds, and 100,000 yuan in cash. The user also enters their specific financial goals, such as "planning to save 1 million yuan for a down payment on a home within five years," and completes a risk appetite assessment, which results in a "conservative" risk profile. Furthermore, the system continuously monitors external unstructured text data related to the user's holdings.

[0057] Next, the user's personal profile is mapped and integrated into a macro, dynamic financial knowledge graph. The user's asset portfolio, specific holdings, and financial goals are all created as independent entity nodes in the knowledge graph. When a new causal event is extracted from external text data, such as a new regulatory policy affecting a specific industry, a multi-agent simulation driven by a large language model is used to perform inferences. The resulting "consensus" is quantified into a specific adjustment value, such as increasing the risk transmission weight of the policy for a specific technology company from 0.3 to 0.6, and the knowledge graph is updated in real time.

[0058] Based on a dynamic knowledge graph, users can interact through natural language to test and explore their financial plans. Users can pose hypothetical questions, such as "Based on my current plan, considering a possible 20% market correction, what is the probability of achieving my financial goals?" A large language model parses this query, converting it into structured intervention parameters. It then performs a Monte Carlo simulation on the knowledge graph, spanning 5,000 iterations, to generate multiple synthetic risk data paths for the user's future asset growth. These simulation results are presented in a user-friendly, interactive financial health dashboard, displayed as a bar chart, visually illustrating the probability of achieving the goal (e.g., a 72% probability), along with the distribution of various possible outcomes.

[0059] On the dashboard, if users are concerned about an unsatisfactory simulation result (such as a path that failed to achieve a goal), they can directly click on the visualization element to interact. This action immediately triggers the back-end counterfactual analysis process. Using a "one-by-one suppression" correction simulation method, the contribution of each driver to the result is accurately calculated and a clear and easy-to-understand attribution report is generated. This report not only explains "why" the plan failed (for example, market systemic risk contributed 50%, specific asset volatility contributed 30%, and insufficient savings rate contributed 20%), but also automatically generates specific, actionable optimization strategies based on the attribution results. For example, "To increase your probability of achieving your goal to 90%, we recommend adding 2,000 yuan to your monthly fixed investment or shifting 15% of your high-risk positions to more stable assets. Would you like to simulate your adjusted plan?"

[0060] While embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions, and variations may be made to these embodiments without departing from the principles and spirit of the invention, and that the scope of the invention is defined by the appended claims and their equivalents.

Claims

1. A data structure adaptive visualization method based on a large language model, characterized in that: include: Access internal quantitative data and external unstructured text data of the investment portfolio; Using a large language model to process the external unstructured text data and extract structured causal events; Constructing a dynamic financial knowledge graph using internal quantitative data and the structured causal events, and updating risk transmission weights between entities in the dynamic financial knowledge graph in real time using the structured causal events; receiving a risk impact hypothesis input by a user, and performing a forward-looking simulation based on the dynamic financial knowledge graph and the risk impact hypothesis to generate a plurality of synthetic risk data paths representing a plurality of future possibilities; generating a visualization result based on the plurality of synthetic risk data paths, wherein the visualization result is bound to a traceability link; In response to a user selecting a specific simulation result on the visualization result, performing a path sensitivity analysis to identify a set of driving factors; A modified simulation method is executed based on the driving factors to perform counterfactual analysis, determine the quantitative contribution of the driving factors to the specific simulation results, and generate an attribution report.

2. The data structure adaptive visualization method based on a large language model according to claim 1, characterized in that: The steps of obtaining internal quantitative data and external unstructured text data of an investment portfolio, processing the external unstructured text data using a large language model, and extracting structured causal events include: The internal quantitative data includes daily portfolio holdings details, historical trading records, and daily profit and loss time series; the external unstructured text data includes regulatory agency announcements, listed company financial reports, and financial news articles; A large language model is used to perform named entity recognition and relationship extraction on the external unstructured text data, and the text data is converted into structured causal events including event types, participating entities and influencing parameters.

3. The data structure adaptive visualization method based on a large language model according to claim 1, characterized in that: The process of constructing a dynamic financial knowledge graph includes: Initializing a base knowledge graph containing a predefined financial ontology, wherein the financial ontology defines entity types, attribute types, and relationship types; The structured causal events are fused with the base knowledge graph. The fusion process includes creating and updating nodes corresponding to entities in the structured causal events in the knowledge graph, creating and updating relationship edges representing the structured causal events between nodes, and adjusting the risk transmission weights associated with the relationship edges and nodes according to the impact parameters of the structured causal events; adding a time-effect decay function to the updated risk transmission weights, and assigning a confidence score to the relationship edges and weights according to the information source of the text data.

4. The data structure adaptive visualization method based on a large language model according to claim 3, characterized in that: The step of adjusting the risk transmission weights associated with the relationship edges and nodes according to the impact parameters of the structured causal events includes: Based on the impact scope of the structured causal event, a group of virtual economic agents related to the structured causal event are instantiated from a predefined agent library, each agent being assigned a specific role and behavior pattern; Presenting the structured causal event as input information to the instantiated group of agents, and using the large language model to drive each agent to generate a decision and behavioral response to the structured causal event based on its role and behavior pattern; Summarize the decisions and behavioral responses generated by the agent group, quantify the summary results into adjustment values, and use the adjustment values ​​to update the risk transmission weights.

5. The data structure adaptive visualization method based on a large language model according to claim 1, characterized in that: The step of performing forward-looking simulation based on the dynamic financial knowledge graph and the risk shock hypothesis includes: Parsing risk impact hypotheses inputted in natural language into structured intervention parameter objects using the large language model, and generating clarifying questions based on the dynamic financial knowledge graph for user confirmation when there is ambiguity in the hypothesis parsing; Applying the structured intervention parameter object to the dynamic financial knowledge graph to adjust the quantitative attributes of graph elements in the dynamic financial knowledge graph that are associated with the intervention parameter object; Perform Monte Carlo simulation, use the quantitative properties of the adjusted graph elements in the dynamic financial knowledge graph to define the random process and related parameters of the Monte Carlo simulation, and generate the multiple synthetic risk data paths.

6. The data structure adaptive visualization method based on a large language model according to claim 1, characterized in that: The step of generating a visualization result includes: For each of the multiple synthetic risk data paths and statistically derived elements, a traceability link pointing to the corresponding driving factor in the dynamic financial knowledge graph is established; on the visualization result, the traceability link is bound to the corresponding visualization element, and the user triggers a counterfactual analysis of the driving factor by interacting with the visualization element.

7. The data structure adaptive visualization method based on a large language model according to claim 6, characterized in that: The process of establishing a traceability link to a corresponding driving factor in the dynamic financial knowledge graph includes: performing a path sensitivity analysis on a specified path in the synthetic risk data path to obtain a quantitative impact score of each driving factor, calculating the sum of the absolute values ​​of the quantitative impact scores of all driving factors, sorting the driving factors in descending order, and selecting, from the top, the driving factor whose cumulative absolute value of the quantitative impact score reaches a preset percentage of the total as the main driving risk factor; Based on the analysis results, a data structure including ranked driver identifications and quantified impact scores is generated as the traceability link.

8. The data structure adaptive visualization method based on a large language model according to claim 1, characterized in that: The step of performing counterfactual analysis includes: performing a modified simulation on the main driving risk factors by temporarily suppressing their risk transmission weights in the dynamic financial knowledge graph; comparing the specific simulation results with the modified simulation results to calculate the quantitative contribution of the main driving risk factors.

Citation Information

Patent Citations

  • Dynamic financial credit risk model construction method based on rule engine

    CN117709446A

  • Pig farm disease risk attribution method and system based on knowledge graph

    CN119207779A

  • Diagnostic dose key information extraction system based on multi-agent architecture

    CN119311789A

  • Financial transaction anomaly detection and risk assessment method and device based on artificial intelligence

    CN119693111A

  • Knowledge graph-based intelligent quantitative research platform system and method thereof

    CN119693155A

Cited By

  • Sustainable consumption behavior simulation system fusing large language model and multiple agents

    CN121615533A

  • Sustainable consumption behavior simulation system fusing large language model and multi-agent

    CN121615533B

  • Intelligent training system with multi-branch interactive plot and related equipment

    CN121788296A

  • Knowledge base construction method based on multi-modal fusion

    CN121935854A