A Multi-Scenario Intelligent Financial Analysis System Based on Large Models and Composite Agents
By using a multi-scenario intelligent financial analysis system based on large models and composite intelligent agents, the shortcomings of existing financial systems in multi-scenario integration, vertical domain knowledge fusion, and intelligent application are solved, achieving efficient and accurate financial data analysis and improving the scientific nature and efficiency of corporate financial decision-making.
Patent Information
- Application Number
- CN202511615249.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-06
- Publication Date
- 2026-01-30
- Estimated Expiration
- 2045-11-06
AI Technical Summary
Existing financial systems are inadequate in terms of multi-scenario integration, vertical domain knowledge fusion, flexible and accurate indicator calculation, and intelligent application. They are unable to meet the complex and diverse financial analysis needs of enterprises, resulting in fragmented and inefficient analysis processes, as well as a lack of accurate understanding of financial terminology and complex indicators.
We employ a multi-scenario intelligent financial analysis system based on large models and composite intelligent agents, including data sources, a foundation layer, a technology stack, and an access terminal. Through deep embedding of agent intelligent agents and LLM large model technology, we achieve multi-intent recognition, intelligent task planning, multimodal perception, and automated processing of financial data. Combined with NL2SQL, RAG search tools, and other technologies, we open up internal and external data channels, build an industry terminology library and indicator formula library, and support natural language queries and report generation.
It enhances the intelligence and accuracy of financial data analysis, lowers the barrier to entry, improves query efficiency and data analysis flexibility, adapts to complex scenarios, supports multi-table joins and nested queries, covers most business needs, and enhances the scientific nature and efficiency of corporate financial decision-making.
Smart Images

Figure CN121073689B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of financial analysis technology, specifically to a multi-scenario intelligent financial analysis system based on a large model and a composite intelligent agent. Background Technology
[0002] Currently, most financial systems remain at the "simulation" stage of traditional accounting work (such as basic functions like automatic voucher generation and report preparation), while showing significant shortcomings in the two dimensions of "extension" (refined management across business units) and "expansion" (decision support based on big data). They tend to focus on single financial scenarios, only able to achieve basic accounting and simple report generation, lacking the ability to integrate multiple scenarios. This makes it difficult to cope with the complex and diverse "finance+" business needs of enterprises, and it is unable to efficiently integrate scenarios such as financial document Q&A, knowledge Q&A, and intelligent data analysis, resulting in a fragmented and inefficient financial analysis process.
[0003] While some AI-powered financial analysis solutions incorporate basic algorithmic models, they are largely based on general-domain data analysis and lack deep integration of expert experience and industry knowledge within the vertical financial field. These general models often lack sufficient understanding of financial terminology and complex indicator formulas. For example, when dealing with scenarios involving nested calculations of debt-to-equity ratio derivatives, the lack of industry knowledge base support can easily lead to analytical biases, compromising professionalism and accuracy.
[0004] In terms of indicator calculation, existing technologies have obvious shortcomings. Financial data analysis relies on a large number of basic and derived indicators. Existing tools are weak in real-time and batch calculation capabilities in scenarios involving nested indicators and multi-time granularity calculations. Often, due to rigid calculation logic, they cannot flexibly adapt to complex financial logic, resulting in lagging or erroneous indicator calculation results, which affects financial analysis decisions.
[0005] Meanwhile, existing financial analysis technologies need improvement in terms of intelligence. Most solutions do not fully utilize Large Language Modeling (LLM) and multi-agent technologies, and lack intelligent empowerment for the entire financial data analysis process.
[0006] The core reasons for the shortcomings of existing technologies are as follows: First, there is an insufficient understanding of the complexity and diversity of financial business scenarios, and an inability to build an effective architecture for multi-scenario integration, making it difficult to break down data and functional barriers between scenarios. Second, the accumulation of knowledge in the vertical financial field is neglected, and general data analysis models cannot accurately adapt to the needs of financial professionals. The integration of industry knowledge is difficult and requires a lot of manpower to sort out industry terminology, indicator formulas and other knowledge systems. Third, the design of indicator calculation logic does not fully consider the complexity of financial indicators, and lacks a flexible and efficient calculation framework. To balance real-time performance and accuracy, it is necessary to overcome technical challenges such as computing resource scheduling and complex logic decomposition. Fourth, the breadth and depth of intelligent technology application are insufficient. The implementation of LLM and multi-agent technologies in the entire financial analysis process requires solving problems such as model adaptation and multi-module collaboration, which is difficult to integrate.
[0007] These intertwined issues have resulted in a significant gap in existing technologies to meet enterprises' needs for efficient, professional, and intelligent financial data analysis. The "Intelligent Financial Analyst" of this invention addresses these pain points by integrating multiple scenarios, fusing vertical domain knowledge, performing flexible and accurate indicator calculations, and applying deep intelligence to achieve a comprehensive breakthrough and innovation in financial data analysis. Summary of the Invention
[0008] The purpose of this invention is to provide a multi-scenario intelligent financial analysis system based on a large model and a composite intelligent agent to solve the problems mentioned above.
[0009] The objective of this invention can be achieved through the following technical solutions:
[0010] A multi-scenario intelligent financial analysis system based on a large model and composite intelligent agents includes: data source, basic layer, technology stack, application layer and access terminal;
[0011] Data sources are used to provide data to the foundational layer, including internal business data and external publicly available data; internal business data includes financial data and operational data; external publicly available data includes knowledge data and real-time data.
[0012] The base layer is used for data processing, and analyzes and processes the data through the data integration module, data governance module, data query module, and data security module;
[0013] The technology stack analyzes the processed data, specifically by deeply embedding LLM large model technology based on Agent / composite agent technology and capability toolkit. The capability toolkit includes: NL2SQL tool, mathematical calculator tool, chart visualization tool, RAG retrieval tool, professional financial search tool, and general domain search tool.
[0014] The application layer will be used to transform the analyzed data into financial business scenario functions;
[0015] The access point is used for interactive entry, including visual interface interaction and API interface calls.
[0016] As a further aspect of the present invention: the data integration module is used to enable user uploads, internet queries, direct database connections, and to establish a data channel between internal business data and external public data;
[0017] The data governance module includes capabilities such as indicator merging, user query standardization, formula extraction, and template extraction, and it also accumulates an industry terminology library, indicator formula library, and report template library.
[0018] Data query module: constructed with wide tables, alias mapping, and SQL standard constraints;
[0019] Data security module: Strictly controls data access through tenant isolation and row-level permissions.
[0020] As a further aspect of the present invention: the Agent's specific structure includes:
[0021] Sensing module: Acquires information about the external environment through sensors;
[0022] Decision module: Based on built-in rules, LLM models or knowledge bases, it analyzes the perceived information and generates action strategies, including calling tools, answering questions, and executing code;
[0023] Execution module: Transforms action strategies into specific actions, including calling functions, sending requests, operating hardware, and feeding back the results to the environment or itself;
[0024] Memory module: Stores historical interaction data, intermediate results, or experience.
[0025] As a further aspect of this invention: the technology stack relies on two major systems—Agent intelligent agents and capability toolkits—and deeply embeds LLM large model technology to reshape the financial analysis process, specifically including:
[0026] Multi-intent recognition: parsing natural language requirements;
[0027] LLM large model: Based on financial knowledge and business data training, it provides capabilities such as intelligent question answering, logical reasoning, and text generation;
[0028] Intelligent task planning: Decompose complex financial analysis tasks and schedule multiple modules to work together to execute sub-tasks, including data query, indicator writing, and text writing.
[0029] Voice conversion: Supports voice input and output;
[0030] Multi-turn dialogue memory: Retains dialogue context and supports continuous, related question interactions;
[0031] Multimodal perception: Integrating multimodal data processing such as text, speech, and tables.
[0032] As a further aspect of the present invention: the composite intelligent agent includes sqlagent, searchagent, and reportagent;
[0033] The sqlagent is an intelligent system that automatically converts natural language into executable SQL and returns results. It combines the LLM model and database operation capabilities to achieve end-to-end semantic querying.
[0034] The sqlagent uses the NL2SQL tool, mathematical calculator tool, and chart visualization tool from the toolbox;
[0035] The searchagent is an intelligent system that performs search tasks, searches for information within a specific range according to rules set by the user, and returns the results to the user.
[0036] The searchagent uses the RAG search tool, professional financial search tool, and general domain search tool from the toolbox;
[0037] The Reportagent is an intelligent agent used to automate the entire process of report generation, optimization, and updating. By integrating data collection, analysis, content generation, and formatting capabilities, it transforms complex reporting tasks into an automated process.
[0038] The Reportagent uses all the tools in the toolbox.
[0039] As a further aspect of the present invention, the technology of the sqlagent specifically includes:
[0040] Natural Language Understanding: Parsing users' natural language questions and extracting key information through entity recognition, intent understanding, and relational reasoning;
[0041] Database knowledge integration: Understanding the structure of the target database;
[0042] SQL generation and validation: Based on LLM's text generation capabilities, natural language is converted into SQL statements, and the generated SQL statements are fully self-validated through validation mechanisms such as syntax validation and semantic validation.
[0043] Execution and Result Processing: Execute the generated SQL through the database connection, obtain the result set, and convert the query results into a user-friendly format;
[0044] Iterative optimization: If the generated SQL fails to execute, the Agent can optimize it through error feedback.
[0045] As a further aspect of the present invention, the technology of the searchagent specifically includes:
[0046] Query understanding: Using natural language processing technology to parse the query statement entered by the user, identify keywords, query intent, etc., and transform the user's needs into executable search instructions;
[0047] Information retrieval: Based on the query instructions, search for relevant information by using the provided toolkit capabilities, indexing the database, or by interfaceing with external search engines;
[0048] Results processing: The retrieved results are filtered, sorted, and integrated; based on relevance, authority, and update time, the results that meet the user's needs are ranked first.
[0049] As a further aspect of the present invention, the Reportagent technology includes sqlagent technology and searchagent technology.
[0050] As a further aspect of the present invention: the RAG retrieval tool combines information retrieval and generative AI technology, specifically including the following steps:
[0051] (1) Data preparation and preprocessing;
[0052] Data collection: Collect professional knowledge documents in the field of finance;
[0053] Cleaning and Segmentation: a. Cleaning: Remove irrelevant information and retain the core text;
[0054] b. Chunking: Breaking a long document into short text blocks;
[0055] Supplementary metadata: Add source, timestamp, keywords, etc. to each block for easy and quick filtering and subsequent tracing;
[0056] (2) Construct a vector database;
[0057] Text embedding: Converting text blocks into high-dimensional vectors using an embedding model;
[0058] Storing vectors: The generated vectors are stored in a vector database and indexed to accelerate subsequent retrieval;
[0059] (3) Retrieval stage: Matching user queries with relevant text;
[0060] (4) Generation stage: Generate answers based on search results.
[0061] As a further aspect of the present invention, the methods for using the LLM large model include: direct invocation, fine-tuning and customization, and combining with tools to expand the capability boundaries.
[0062] The beneficial effects of this invention are:
[0063] This invention is based on the Large Language Model (LLM), which has powerful semantic understanding and generation capabilities. It can understand complex contexts and generate fluent text that conforms to human expression habits. It has strong generalization ability and adapts to multiple tasks. The general language knowledge learned in the pre-training stage means that it does not need to be trained separately for each task. It can directly support dozens of tasks through prompt words. It also has strong context understanding ability, supports the processing of long contexts, and can remember previous information and generate coherent content based on it. Compared with traditional models, it has low data dependence. Traditional machine learning models require a lot of labeled data to complete specific tasks, while LLM can complete tasks through prompt words.
[0064] This system employs NL2SQL (Natural Language to SQL), a technology that automatically converts natural language questions into structured SQL query statements. This allows non-technical users to directly query the database using everyday language, without needing to master SQL syntax, thus lowering the barrier to entry. Non-technical users do not need to learn SQL; they can directly query the database using natural language, significantly reducing database usage costs. It also improves query efficiency by automating SQL generation, reducing the time and errors associated with manual coding. It is particularly suitable for scenarios with frequent queries and adapts to complex scenarios, supporting the generation of complex SQL statements such as multi-table joins and nested queries. It covers most business needs, has strong scalability, can adapt to different types of databases, and can be quickly migrated to new business scenarios through fine-tuning the model. Attached Figure Description
[0065] The invention will now be further described with reference to the accompanying drawings.
[0066] Figure 1 This is a schematic diagram of the system structure of the present invention. Detailed Implementation
[0067] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0068] The purpose of this invention is to comprehensively solve the problems in the existing technologies mentioned above by leveraging large-scale models and agent technology. By developing conversational financial data AI, it organically integrates five major "finance+" scenarios to achieve comprehensive financial data analysis, broaden data sources, and merge internal business and financial data with externally available data, providing rich data support for comprehensive and accurate analysis of corporate financial conditions.
[0069] This invention also deeply integrates expert experience and industry knowledge to construct an industry terminology library, indicator formula library, industry knowledge base, and report template library, enhancing professionalism in vertical fields and improving the accuracy and professionalism of financial data analysis. Simultaneously, through real-time and batch computing technologies, it optimizes indicator calculation logic, flexibly and accurately handling complex calculation problems such as indicator nesting and multiple time granularities, ensuring the efficiency and reliability of financial indicator calculation. Furthermore, it deeply applies LLM and multi-agent technologies to multiple modules of the entire architecture, comprehensively improving the intelligence level of financial data analysis, providing enterprises with more intelligent, efficient, and accurate financial analysis services, helping them make scientific and rational financial decisions in a complex and ever-changing market environment, and enhancing their competitiveness.
[0070] Example 1
[0071] Please see Figure 1 As shown, the present invention is a multi-scenario intelligent financial analysis system based on a large model and a composite intelligent agent, including: data source, basic layer, technology stack, application layer and access terminal;
[0072] Among them: the data source provides data to the basic layer, that is, the "raw materials";
[0073] The base layer is used to process data, making it "usable and easy to use";
[0074] The technology stack analyzes the processed data to enhance the technology's ability to analyze "thinking and decision-making".
[0075] The application layer will be used to transform the analyzed data into financial business scenario functions;
[0076] The access point is used as the entry point for use;
[0077] By constructing a complete value chain of "data-governance-intelligence-application-interaction" between each layer, a comprehensive breakthrough in the efficiency, quality, and intelligence of financial data analysis can be achieved.
[0078] Specifically:
[0079] As the data foundation of the entire architecture, the data source integrates two types of data: including:
[0080] Internal business data includes financial data, such as core accounting information like revenue, costs, and expenses; and business data, such as process data like sales, operations, procurement, and customer data. It deeply covers the entire chain of internal business management, providing a business-related foundation for financial analysis.
[0081] External publicly available data includes knowledge data, such as professional knowledge in accounting, economic law, and financial cost management; and real-time data, including market dynamics data such as news, stock prices, and industry research, expanding the analytical perspective to the industry environment and macro market.
[0082] The data source integrates data from multiple sources, breaking down "data silos" for enterprises and providing rich material for comprehensive financial analysis, supporting full-dimensional insights from micro-finance to macro-business development;
[0083] The foundation layer processes data through four modules: data integration, data governance, data query, and data security, thereby accumulating industry knowledge assets and ensuring data quality and analytical accuracy.
[0084] Specifically:
[0085] Data integration module: used to enable user uploads, internet queries, direct database connections, etc., to open up internal and external data channels and achieve the convergence of financial data and market data;
[0086] Data governance module: Based on expert experience and industry knowledge, it builds capabilities such as indicator merging, user query standardization, formula extraction, and template extraction, and accumulates industry terminology library, indicator formula library, and report template library; for example, by merging indicators to standardize the definition of financial indicators, and by extracting formulas to solidify professional calculation logic, it enables data to be adapted to the needs of financial professionals, solving the problem of traditional data analysis "not understanding financial business".
[0087] Data query module: Optimizes query experience with wide table construction, alias mapping, and SQL standard constraints, adapts to the habits of finance personnel, and improves data retrieval efficiency and accuracy;
[0088] Data security module: Strictly controls data access through tenant isolation and row-level permissions to ensure the security of sensitive corporate financial data;
[0089] The foundational layer deeply integrates industry knowledge with data governance, enabling data analysis to "understand finance and business," laying the cornerstone for professional analysis.
[0090] Technology Stack: It is the core engine of intelligent analysis, relying on two major systems: Agent intelligent agent / composite intelligent agent technology and capability toolbox, and deeply embedding LLM large model technology to reshape the financial analysis process;
[0091] Specifically: Through agent-based intelligent technology, the financial analysis process can be automated and intelligently collaborative.
[0092] An agent is an intelligent entity capable of autonomously perceiving its environment, making decisions, and executing actions. Its core characteristics can be summarized as a closed loop of "perception-decision-execution," with the following specific structure:
[0093] Perception module: Acquires external environmental information (such as text, data, images, user commands, etc.) through sensors (such as data interfaces, APIs, sensor hardware);
[0094] Decision module: Based on built-in rules, models (such as LLM, reinforcement learning models) or knowledge base, analyze the perceived information and generate action strategies (such as "call tools", "answer questions", "execute code").
[0095] Execution module: Transforms action strategies into specific actions (such as calling functions, sending requests, and operating hardware), and feeds back the results to the environment or itself;
[0096] Memory module (optional): Stores historical interaction data, intermediate results or experience to optimize subsequent decisions (e.g., short-term memory caches dialogue, long-term memory stores knowledge).
[0097] The essence of an agent is an "autonomous intelligent execution unit" that can autonomously complete its goals in a dynamic environment without continuous human intervention;
[0098] Specifically, the steps include:
[0099] Multi-intent recognition: Accurately analyzes the natural language needs of finance personnel, such as "analyze quarterly cash flow trends", and identifies multiple intents such as query, analysis, and reporting, enabling the tool to "understand" the language of financial business.
[0100] LLM Large Model: Based on financial knowledge and business data training, it provides capabilities such as intelligent question answering, logical reasoning, and text generation; for example, it can answer complex financial standard questions, deduce the calculation logic of derived indicators, and generate the first draft of analysis reports, breaking through the limitations of traditional tools that are "intelligent in calculation but cannot think".
[0101] Intelligent task planning: Decompose complex financial analysis tasks, such as "generating a viscous financial diagnostic report", and schedule multiple modules to work together to execute sub-tasks, so as to realize the automatic decomposition and advancement of tasks. Sub-tasks include data query, indicator writing, and text writing.
[0102] Voice conversion: Supports voice input and output, adapting to mobile office and hands-free scenarios for finance personnel, improving the convenience of interaction;
[0103] Multi-turn dialogue memory: Retains dialogue context and supports continuous and related question interaction, such as first asking "the reasons for the decline in profits" and then asking "the proportion of the impact of procurement costs", simulating the in-depth communication experience of a real consultant;
[0104] Multimodal perception: Integrates multimodal data processing such as text, voice, and tables, supports uploading and parsing financial statements, and voice command-driven data analysis, expanding the boundaries of interaction;
[0105] The technology stack leverages LLM and multi-agent deep collaboration to upgrade financial analysis from "tool execution" to "intelligent decision support," covering the entire process of perception, thinking, and execution, and reshaping the intelligent analysis experience.
[0106] Application Layer: Focusing on the implementation of financial business scenarios, it integrates five core functions to achieve multi-scenario coverage and comprehensive analysis.
[0107] Financial Document Q&A: Supports uploading company financial statements, financial reimbursement policies and other documents, and uses natural language Q&A to parse the content (such as "reasons for abnormal inventory turnover days in financial statements") to quickly extract key information from the documents;
[0108] Financial knowledge Q&A: Based on knowledge graphs and LLM, it answers professional knowledge questions on financial standards, tax policies and other related topics (such as "the applicable conditions for the new regulations on additional deduction of R&D expenses"), becoming a "portable knowledge base" for finance personnel;
[0109] Intelligent financial data retrieval: Through NL2SQL and multi-intent recognition, it enables natural language-driven data queries (such as "gross profit margin changes over the past three years by region"), replacing traditional SQL writing and improving data retrieval efficiency;
[0110] Financial Theme Analysis: Based on a scenario model library and intelligent task planning, in-depth analysis is conducted on specific themes (such as "cost control optimization" and "revenue growth strategy"), and diagnostic conclusions are output by integrating multi-source data.
[0111] Financial report generation: Leveraging LLM text generation, BI visualization, and template extraction capabilities, the system automatically integrates analysis results to generate standardized or customized financial reports (such as annual reports and management dashboard reports), significantly shortening the report preparation cycle.
[0112] Five major scenarios cover the entire financial analysis process, from basic queries to in-depth diagnosis and report output, achieving "one-stop" comprehensive analysis and solving the pain points of traditional tools being scattered and repetitive.
[0113] Access point: Flexible interactive entry point;
[0114] It offers two access methods: a visual interface and an API call.
[0115] The visual interface is designed for finance professionals, using low-code and interactive design to achieve "what you see is what you get" operation, adapting to daily analysis needs;
[0116] The API interface supports integration with existing enterprise systems (such as ERP and BI platforms), allowing intelligent analysis capabilities to be embedded in business processes and expanding application boundaries.
[0117] In summary, this invention, based on the Large Language Model (LLM), possesses powerful semantic understanding and generation capabilities. It can understand complex contexts and generate fluent text that conforms to human expression habits. It has strong generalization ability, adapts to multiple tasks, and the general language knowledge learned in the pre-training stage eliminates the need for separate training for each task. It can directly support dozens of tasks through prompt words, and has strong contextual understanding capabilities, supporting the processing of long contexts. It can remember previous information and generate coherent content based on it. Compared with traditional models, it has low data dependence. Traditional machine learning models require a large amount of labeled data to complete specific tasks, while LLM can complete tasks through prompt words.
[0118] This system employs NL2SQL (Natural Language to SQL), a technology that automatically converts natural language questions into structured SQL query statements. This allows non-technical users to directly query the database using everyday language, without needing to master SQL syntax, thus lowering the barrier to entry. Non-technical users do not need to learn SQL; they can directly query the database using natural language, significantly reducing database usage costs. It also improves query efficiency by automating SQL generation, reducing the time and errors associated with manual coding. It is particularly suitable for scenarios with frequent queries and adapts to complex scenarios, supporting the generation of complex SQL statements such as multi-table joins and nested queries. It covers most business needs, has strong scalability, can adapt to different types of databases, and can be quickly migrated to new business scenarios through fine-tuning the model.
[0119] Example 2
[0120] Based on the above embodiments, this embodiment describes a technical solution for handling unsolvable problems encountered by the Agent during financial analysis using a multi-agent approach:
[0121] Specifically: Multi-Agent is a system composed of multiple agents. Its core lies in the collaboration and interaction between agents. Through division of labor, negotiation, and cooperation, it solves complex tasks that a single agent cannot complete. The essence of Multi-Agent is a "distributed intelligent collaborative network" that makes up for the limitations of a single agent through the interaction of multiple agents.
[0122] This invention includes "sqlagent" (SQL agent), "searchagent" (search agent), and "reportagent" (report agent):
[0123] sqlagent: An intelligent system that automatically converts natural language into executable SQL and returns results. Its core is to combine Large Language Model (LLM) and database operation capabilities to achieve end-to-end semantic querying.
[0124] The technical principles include:
[0125] Natural Language Understanding (NLU): SQLagent first parses the user's natural language question and extracts key information through entity recognition, intent understanding, relational reasoning, and other methods;
[0126] Database knowledge integration: SQLagent needs to understand the structure (Schema) of the target database;
[0127] SQL generation and validation: Based on LLM's text generation capabilities, natural language is converted into SQL statements, and the generated SQL statements are fully self-validated through validation mechanisms such as syntax validation and semantic validation.
[0128] Execution and Result Processing: Execute the generated SQL through the database connection, obtain the result set, and convert the query results into a user-friendly format;
[0129] Iterative optimization: If the generated SQL fails to execute, the Agent can optimize it through error feedback;
[0130] SQLagent uses the "NL2SQL tool", "mathematical calculator tool", and "chart visualization tool" from the toolbox;
[0131] Searchagent: An intelligent system that can automatically perform search tasks. It can search for relevant information within a specific scope (such as web pages, databases, etc.) based on user needs or set rules, and return the results to the user.
[0132] The technical principles include:
[0133] Query understanding: Using natural language processing technology to parse the query statement entered by the user, identify keywords, query intent, etc., and transform the user's needs into executable search instructions;
[0134] Information retrieval: Based on the query instructions, search for relevant information by using the provided toolkit capabilities, indexing the database, or by interfaceing with external search engines;
[0135] Results processing: The system filters, sorts, and integrates a large number of retrieved results; based on factors such as relevance, authority, and update time, the results that best meet the user's needs are ranked first, and the system can also perform processing such as extracting summaries from the results so that users can obtain key information more quickly.
[0136] Searchagent uses the "RAG Search Tool", "Professional Financial Search Tool", and "General Domain Search Tool" from the toolbox;
[0137] Reportagent: An intelligent agent that can automate the entire process of report generation, optimization, and updating. By integrating data collection, analysis, content generation, and formatting capabilities, it transforms complex reporting tasks into an automated process.
[0138] The technical principles include the related technologies of sqlagent and searchagent;
[0139] Specifically, the following tools were used: "NL2SQL Tool", "Mathematical Calculator Tool", "Chart Visualization Tool", "RAG Search Tool", "Professional Financial Search Tool", and "General Domain Search Tool" from the toolbox.
[0140] Example 3
[0141] Based on the various technical tools used in the Multi-Agent (composite intelligent agent) in Embodiment 2, this embodiment will provide a technical description based on the tools used by "sqlagent" (SQL intelligent agent), "searchagent" (search intelligent agent), and "reportagent" (report intelligent agent):
[0142] RAG retrieval tools combine information retrieval and generative AI technologies. The core goal is to enable generative models to generate answers based on "real data retrieved" rather than relying solely on the model's own training data.
[0143] RAG knowledge retrieval: Combining retrieval and generation technologies, it quickly locates professional knowledge and integrates it into analysis, solving the problems of outdated and unprofessional LLM knowledge and making the analysis conclusions more reliable; including:
[0144] BI Visualization: Automatically generates visual charts of financial data, presenting analysis results intuitively and assisting in decision-making;
[0145] Scenario Model Library: Pre-built industry-standard and customized analysis models, enabling one-click access to complete complex scenario analysis and improve analysis efficiency;
[0146] Its principle can be broken down into four key steps:
[0147] (1) Data preparation and preprocessing;
[0148] Data collection: Collect professional knowledge documents in the field of finance;
[0149] Cleaning and Chunking: a. Cleaning: Remove irrelevant information (such as formatting errors and duplicate content) and retain the core text. b. Chunking: Break down long documents into short text blocks (usually 100-500 words).
[0150] Supplementary metadata: Add source, timestamp, keywords, etc. to each block for easy and quick filtering and subsequent tracing;
[0151] (2) Construct a vector database;
[0152] Text embedding: Converting text blocks into high-dimensional vectors using an embedding model;
[0153] Storing vectors: The generated vectors are stored in a vector database and indexed to accelerate subsequent retrieval;
[0154] (3) Retrieval stage: Matching user queries with relevant text;
[0155] When a user enters a query, it is first converted into a vector using the same embedding model. Then, the top N text blocks that are closest to the query vector are found in the vector database using "similarity retrieval" (using cosine similarity). These text blocks contain information most relevant to the user's question.
[0156] (4) Generation stage: Generate answers based on search results;
[0157] The retrieved Top N text blocks are input into the generative model along with the user query, and the model is guided by prompt words to combine the retrieved real data to output accurate and evidence-based answers.
[0158] NL2SQL (Natural Language to SQL) is a technology that automatically converts natural language questions into structured SQL query statements. Its core goal is to allow non-technical users to query the database directly using everyday language without needing to master SQL syntax. This technology converts user queries (i.e., natural language form) into corresponding SQL statements, executes the converted SQL statements, retrieves relevant data from the database, integrates the retrieved data, and returns it to the user.
[0159] NL2SQL: Transforms financial natural language requirements (such as "query departments with sales expenses exceeding 1 million in Q1 2025") into SQL statements, allowing non-technical finance personnel to directly access database analysis, thus lowering the technical barrier; including:
[0160] Knowledge graph: Construct a network of financial knowledge connections, accumulate classic financial themes, and use them for the analysis of complex financial relationships and financial content to uncover the deep value of data;
[0161] Mathematical Calculator: Adapted to specific financial calculation needs, it supports complex indicator nesting and batch calculations, ensuring calculation accuracy;
[0162] Its technical principle can be broken down into three key aspects:
[0163] (1) Natural Language Understanding (NLU):
[0164] First, the input natural language question is parsed to extract core information:
[0165] Entity recognition: Identify key entities in the question (such as "sales volume", "XX city", "2023") and their corresponding fields (such as sales), table names (such as region), or specific values (such as 2023) in the database.
[0166] Intent and Relationship Analysis: Analyze the semantic intent of the question (such as "query", "sum", "sort") and the relationships between entities (such as "sales volume of XX city" corresponding to the filter condition "region = XX city").
[0167] Context modeling: For complex questions (such as the multi-turn dialogue "What about XX city?"), it is necessary to understand the referential relationship by combining the historical dialogue context ("that" refers to the "sales volume query" mentioned earlier);
[0168] (2) Database schema mapping:
[0169] Associating the parsed natural language information with the database structure (schema): clarifying which table and field in the database the entity in the question corresponds to (e.g., "sales amount" maps to the sales_amount field in the orders table). Handling ambiguity (e.g., "product name" may correspond to the name or product_title field in the products table, which needs to be eliminated through context or database metadata);
[0170] (3) SQL generation and validation:
[0171] Based on the results of the first two steps, generate SQL statements that conform to the syntax rules and perform validation:
[0172] Generation: Select SQL (Structured Query Language) operations (such as SELECT "query", SUM() "sum", ORDERBY "sort") based on semantic intent, and construct WHERE conditions, GROUP BY grouping, etc. by combining entity mapping.
[0173] Validation: Ensure that the generated SQL syntax is correct (e.g., field names exist, table join logic is reasonable) and can be adapted to the target database type (e.g., syntax differences between MySQL and PostgreSQL).
[0174] Technological evolution: Early versions relied on rule templates (e.g., "query A's B" corresponds to SELECT B FROM table WHERE entity = A), but had poor generalization ability; the current mainstream is based on pre-trained language models (e.g., T5, BART, LLaMA), which take "natural language question + database schema" as input, generate SQL end-to-end, and combine Prompt engineering or fine-tuning to adapt to specific database scenarios.
[0175] Its core advantage lies in
[0176] (1) Lowering the barrier to entry: Non-technical users (such as business personnel and customer service) do not need to learn SQL and can directly query the database using natural language (such as "showing the number of users in each quarter of 2024"), which greatly reduces the cost of using the database;
[0177] (2) Improve query efficiency: Automatically generate SQL, reduce the time and errors of manual writing (such as field name spelling errors, logical errors), especially suitable for scenarios with frequent queries (such as enterprise report generation, customer service data query);
[0178] (3) Adapt to complex scenarios: Supports the generation of complex SQL such as multi-table joins and nested queries (e.g., "query the names of employees in each department whose salary is higher than the average salary of that department"), covering most business needs;
[0179] (4) High scalability:
[0180] It can be adapted to different types of databases (relational, data warehouse, etc.), and can be quickly migrated to new business scenarios by fine-tuning the model (such as expanding from e-commerce order query to logistics delivery query).
[0181] Example 4
[0182] This embodiment describes the technical solution based on the LL large model in the technology stack, specifically:
[0183] Large Language Model (LLM) is a deep learning model trained on massive amounts of text data. Its core principle is to capture the statistical patterns and semantic logic of human language through "self-supervised learning" and then output text that conforms to the context through "autoregressive generation".
[0184] Its underlying architecture is based on the Transformer (attention mechanism), which uses a "self-attention mechanism" to allow the model to dynamically focus on the relationships between different words when processing text. This mechanism enables the model to capture long-distance semantic dependencies (such as the logical relationships between sentences in a paragraph), far exceeding the capabilities of traditional RNNs and LSTMs.
[0185] The core advantages of this LLM model lie in its "generality" and "depth of language understanding," specifically manifested as follows:
[0186] (1) Powerful semantic understanding and generation capabilities: It can understand complex contexts (such as ambiguity, metaphor, and long sentence logic) and generate fluent text that conforms to human expression habits;
[0187] (2) Strong generalization ability and adaptability to multiple tasks: The general language knowledge learned in the pre-training stage makes it possible to support dozens of tasks directly through prompt words without separate training for each task.
[0188] (3) Contextual understanding ability (long text processing): Supports processing long contexts (such as GPT-4 supporting tens of thousands of tokens), can remember previous information and generate coherent content based on it;
[0189] (4) Low data dependence (compared to traditional models): Traditional machine learning models require a large amount of labeled data to complete a specific task, while LLM can complete the task with cue words (zero samples / few samples);
[0190] Its usage methods include:
[0191] 1. Direct call: Direct call requires no or minimal code. This invention adopts the direct call method, which guides the model to output results through prompts, quickly solving general tasks. The core technique is prompt engineering.
[0192] 2. Fine-tuning and customization: This method is adapted to specific domains or tasks. When direct calls are not effective, such as when there are many domain terms or the task is special, fine-tuning can be used to make the model "more professional".
[0193] 3. Integrating tools to expand capabilities: This invention uses a tool-integrated approach to create a custom mathematical calculator; LLM itself lacks real-time information (such as current news in 2025) and computational capabilities (such as complex mathematical problems), so it is necessary to call tools to make up for these deficiencies.
[0194] Example 5
[0195] Based on the data processing process of the base layer on the data source provided by the data source in the above embodiments, this embodiment provides the following technical solution:
[0196] Data integration module: used to connect internal and external data channels and aggregate data from data sources;
[0197] The process includes the following steps: acquiring data from the data source, denoting the data source in the batch window time as j, matching and merging the data records in the data source, and then denoting the matched and merged data record as i, where j and i are both positive integers; i is 1, 2, 3...N; j is 1, 2, 3...n;
[0198] It should be noted that the data source is the data provided by the basic layer, such as financial data and business data; the data record is the record data at the time of collection, such as income and expenses in financial data and sales and operations in business data;
[0199] Retrieve the data record value D of data record i in the data source. i,j Simultaneously, obtain the confidence score ω of the data records in data source j. j Then, based on the completeness and validity of data record i, obtain its quality score Q in field j of data source i. i,j ;
[0200] Record confidence is calculated by comprehensively evaluating the reliability of the data source, data quality, and historical records. Quality score is calculated by comprehensively evaluating the completeness, correctness, consistency, and timeliness of the data source. Both record confidence and quality score are conventional data calculation methods.
[0201] Then through Calculate the merged value R of the obtained data records. i ;
[0202] Then through Calculate the overall confidence level S of data record i. i ;
[0203] A single data record confidence score may contain errors or be outdated during the calculation process. Similarly, while a single quality score can guarantee the accuracy of the data field itself, it may be based on low confidence scores, leading to inaccuracies in the entire data source. By comprehensively processing data record confidence scores and quality scores, the resulting data records are more accurate.
[0204] In this step, data records from multiple sources are merged and calculated to ensure the high credibility of the data source; at the same time, the overall confidence level of each data record is calculated to provide a basis for subsequent governance and calculation.
[0205] Data governance module: Constructs indicator merging to adapt data to financial needs;
[0206] The data records are labeled with fields, and the fields of each data record i are denoted as k, where m is the total number of fields;
[0207] Similarly, through Calculate the combined field value R i,k ;
[0208] Among them, D i,j,k For data record i, specify the field value in data source j; ω j,k The confidence score of field k for data source j;
[0209] Then, retrieve the fields from the data records and use the VIF function to determine the validity of the fields;
[0210] If field k is a valid field, the value of the VIF function is 1; if field k is an invalid field, the value of the VIF function is 0.
[0211] Then through Calculate the integrity score Cpf(i) of the obtained data records;
[0212] Then through Calculate the corrected confidence level SX i ;
[0213] Based on the valid values and confidence levels of the fields in the data records, the completeness of the records is calculated, and the overall confidence level is further corrected to effectively improve the credibility of the overall data records and reduce the error of subsequent queries.
[0214] Finally, merge the data records into value R. i The data is processed based on a function calculated from the indicator data, and the confidence level SX is adjusted. i ,pass The weight value M of indicator data e is calculated. e ;
[0215] Among them, G e ( ) is the index calculation function;
[0216] Among them, the indicator data is the analysis result calculated and summarized based on the data records using formulas, such as the total revenue obtained after summation, the net profit obtained after calculation, etc. Therefore, the indicator calculation function is a function that performs calculations and processing on different data records, and is a function adapted to the current data record.
[0217] Data query module: Optimizes query data metrics to improve data retrieval efficiency and accuracy;
[0218] Based on the query conditions, the confidence level SX is adjusted. i Indicator weight value M eand index calculation function G e ( The query is processed and the final query result summary Rq is obtained;
[0219] Specifically: ;
[0220] Where α is the confidence coefficient, which can be calculated by predicting probability and error distribution. It is set by technicians based on data records to control the impact of different quality data on query results; qest is the set of records that meet all query conditions, which refers to the data entries available in all data sources.
[0221] Then the indicator function is calculated and extracted from the data source to filter out the record set that meets the conditions;
[0222] The index function is used to calculate the adjusted confidence level for each record.
[0223] The final query results are obtained;
[0224] Data security module: Controls access to sensitive data to ensure the security of sensitive corporate financial data;
[0225] Based on the query result summary Rq, determine whether there are sensitive fields in the query results. If there are sensitive fields, mark the sensitive fields Ls as 1; if there are no sensitive fields, mark the sensitive fields Ls as 0.
[0226] Then, based on the summary of query results and the marking of sensitive fields, the results are processed to obtain the output query result RL;
[0227] Specifically: ;
[0228] This module performs security control on the information in the query results. When there are sensitive data records or sensitive fields in the query results, the sensitive fields of the query record are marked as 0. At this time, it means that the query information is sensitive information, and the output of this query result is canceled.
[0229] This data processing workflow begins with multi-source data integration, merging internal business data with externally available data according to weights and quality scores to obtain preliminary integration results and ensure data reliability. Subsequently, the data governance module standardizes the integrated data, merges indicators, refines formulas, and establishes templates to improve data consistency and usability. Furthermore, it delves deeper into the data records from the digital sources, processing data down to the field level, resulting in a high-quality dataset directly usable for analysis. The data query module constructs wide tables based on the governed data, maps aliases, and applies SQL standards, weighting the query results to make the output more accurate and aligned with business realities. Finally, the data security module isolates and anonymizes sensitive fields, generating secure and compliant final data for application layer use.
[0230] The foregoing has provided a detailed description of one embodiment of the present invention, but this description is merely a preferred embodiment and should not be construed as limiting the scope of the invention. All equivalent variations and modifications made within the scope of the claims of this invention should still fall within the patent coverage of this invention.
Claims
1. A multi-scene intelligent financial analysis system based on a large model and a composite agent, characterized in that, Comprise; Data sources for providing data for the base layer, including internal business data and external public data; internal business data includes financial data and business data; external public data includes knowledge data, real-time data; The base layer is used for processing the data provided by the data source; The process of data processing by the base layer includes: Data integration module: for opening internal and external data channels, and gathering data in the data source; Including the following steps: obtaining data source data, matching and merging data records in the data source; calculating the data record merging value and comprehensive confidence of the data record; Data governance module: building index merging through data adaptation financial needs; Mark the data record by field, calculate the field merging value; judge the field validity by VIF function; then calculate the integrity score of the data record; calculate the correction confidence and index proportion value of the index data; Data query module: optimize query data index, improve data retrieval efficiency and accuracy; Data security module: control access to sensitive data; Technology stack, based on Agent intelligent agent / complex intelligent agent technology and ability toolbox, deeply embedded LLM large model technology to reshape and analyze the data processed by the base layer; Application layer, which is used to convert the data after reshaping and analysis into financial business scenarios.
2. The multi-scene intelligent financial analysis system based on a large model and a composite agent according to claim 1, characterized in that, The specific structure of the Agent includes: Perception module: obtain external environment information through sensors; Decision module: analyze the perceived information to generate action strategies; Execution module: convert action strategies into specific actions and feed back results to the environment or itself; Memory module: store historical interaction data, intermediate results or experience.
3. The multi-scene intelligent financial analysis system based on a large model and a composite agent according to claim 2, characterized in that, The technology stack based on Agent intelligent agent and ability toolbox two systems, deeply embedded LLM large model technology in the process of reshaping and analyzing the data processed by the base layer specifically includes: Multi-intention recognition: analyze natural language requirements; LLM large model: based on financial knowledge and business data training, providing intelligent question and answer, logical reasoning, text generation; Intelligent task planning: decompose complex financial analysis tasks, and schedule multiple modules for collaborative execution of subtasks; Multi-round dialogue memory: retain dialogue context to support continuous and related problem interaction; Multi-modal perception: fusion of text, speech, table multi-modal data processing.
4. The multi-scene intelligent financial analysis system based on a large model and a composite agent according to claim 1, characterized in that, The complex intelligent agent includes sqlagent, searchagent and reportagent; The ability toolbox includes: NL2SQL tool, mathematical calculator tool, chart visualization tool, RAG retrieval tool, professional financial search tool, general domain search tool; The sqlagent is an intelligent system for automatically converting natural language into executable SQL and returning results; The searchagent is an intelligent system for performing search tasks; The Reportagent is an intelligent agent that can automatically complete the whole process of report generation, optimization and update.
5. The multi-scene intelligent financial analysis system based on a large model and a composite agent according to claim 4, characterized in that, The technology of sqlagent specifically includes: Natural language understanding: analyze users' natural language problems and extract key information; Database knowledge integration: understand the structure of the target database; SQL generation and verification: Based on the text generation capability of LLM large model, natural language is converted into SQL statements, and the generated SQL statements are comprehensively self-checked through syntax verification mechanism and semantic verification mechanism; Execution and result processing: Execute the generated SQL through database connection, get the result set, and convert the query result format; Iterative optimization: If the generated SQL execution fails, the Agent optimizes through error feedback.
6. The multi-scene intelligent financial analysis system based on a large model and a composite agent according to claim 4, characterized in that, The technology of searchagent includes: Query understanding: Analyze the user's input query statement through natural language processing technology; Information retrieval: Index the database according to the query instruction and search for related information; Result processing: Filter, sort and integrate the retrieved results.
7. The multi-scene intelligent financial analysis system based on a large model and a composite agent according to claim 4, characterized in that, The technology of Reportagent includes sqlagent technology and searchagent technology.
8. The multi-scene intelligent financial analysis system based on a large model and a composite agent according to claim 1, characterized in that, The RAG retrieval tool combines information retrieval and generative AI technology, which includes the following steps: (1) Data preparation and preprocessing; (2) Constructing a vector database; (3) Retrieval stage: matching user queries with related texts; (4) Generation stage: generating answers based on retrieval results.
9. The multi-scene intelligent financial analysis system based on a large model and a composite agent according to claim 1, characterized in that, The use method of LLM large model includes: direct calling, fine-tuning customization, and combining tool to expand the capability boundary.
Citation Information
Patent Citations
Financial field data intelligent dialogue type analysis system based on AI large model
CN120653739A