A financial intelligence question precision method based on dynamic calculation optimization
By combining dynamic computation optimization and large language model (LLM), the problems of computational flexibility and adaptation to complex scenarios in the financial vertical field are solved, and the accuracy, security and efficiency of multi-entity and multi-dimensional financial queries are achieved.
Patent Information
- Application Number
- CN202511604643.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-05
- Publication Date
- 2026-02-06
- Estimated Expiration
- 2045-11-05
AI Technical Summary
Existing intelligent query technology suffers from insufficient computational flexibility and poor adaptability to complex scenarios in the financial vertical field. It is difficult to handle financial query needs with multiple entities and multiple time dimensions, and traditional tools cannot achieve parallel processing of multiple subtasks and dynamic adaptation.
By employing a dynamic computation optimization approach, this method utilizes the Large Language Model (LLM) and the Semantic Representation Model (BERT) to decompose and compute financial data. Combined with federated learning and homomorphic encryption techniques, it enables multi-entity, multi-dimensional financial analysis and constructs a dynamic rule base and a dual-trigger update mechanism to support secure access and efficient querying of multi-source financial data.
It achieves greater accuracy and flexibility in financial indicator calculation, improves the accuracy of cross-system data matching, lowers the operational threshold for non-professional users, and ensures the security and query efficiency of financial data.
Smart Images

Figure CN121071105B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of question precision, and in particular to a financial intelligent question precision method based on dynamic calculation optimization. BACKGROUND
[0002] Under the wave of enterprise digital transformation, intelligent question technology has become an important tool for financial data analysis, but the existing general field intelligent question scheme has exposed significant limitations in the application of the financial vertical field. The current technology lacks innovation breakthroughs in the flexibility of financial indicator calculation and the adaptability of complex scenarios.
[0003] In the prior art, the prior art lacks an effective processing mechanism for the calculation differences of financial indicators at different time granularities and data sections. The complex nested scenarios of financial analysis place high demands on the logical analysis capabilities of the computing system. Existing tools mostly use single-layer computing frameworks, which are difficult to cope with. Enterprise financial queries often involve multiple entities, multiple time dimensions, and multiple indicators. The existing technology lacks precise multi-entity extraction and correlation analysis capabilities, making it difficult to decompose complex problems into executable subtasks. The computing scheduling capability is insufficient to enable parallel processing of multiple subtasks. Traditional query solutions based on physical data warehouses cannot dynamically adapt to multi-entity correlation logic, resulting in insufficient accuracy of cross-organizational data matching.
[0004] Therefore, there is a need to provide a financial intelligent question precision method based on dynamic calculation optimization. SUMMARY
[0005] The present application aims to provide a financial intelligent question precision method based on dynamic calculation optimization. To solve the above-mentioned problems of the prior art, the present application achieves the following technical solutions:
[0006] In a first aspect, the present application provides a financial intelligent question precision method based on dynamic calculation optimization, which specifically includes the following:
[0007] The user inputs the financial inquiry requirement in natural language, generates financial question information, and takes it as the starting point for system processing.
[0008] Relying on the accumulated industry knowledge base and enterprise knowledge base in the vertical field, the user's question is processed with industry knowledge fusion and standardization, and the user's question is preprocessed.
[0009] For complex user questions, a complex inquiry decomposition process is started to decompose complex requirements into simple logical subtasks, supporting multi-entity and multi-dimensional financial analysis.
[0010] For the disassembled sub-tasks, the processing is circularly carried out, real-time dynamic calculation is sequentially executed, formula flexible expansion and caliber adaptation are carried out, and the granularity difference and index nesting problems are solved.
[0011] The calculation results of each sub-task are collected, and integrated processing is carried out relying on a large language model LLM.
[0012] In the second aspect, the application embodiment provides a financial intelligent question and answer precision method based on dynamic calculation optimization, which specifically comprises the following steps:
[0013] Step one: obtaining financial question and answer information, dividing input form, rewriting financial semantics through double knowledge base, standardizing time for multi-dimensional rules, and obtaining standardized text through text clarification;
[0014] Step two: constructing a data access gateway based on the standardized text, deploying a local data node, transmitting data metadata, decomposing complex requirements into logically simple sub-tasks, supporting multi-entity and multi-dimensional financial analysis, partitioning according to business domain and time granularity, and performing data quality optimization processing;
[0015] Step three: after the data quality optimization processing is completed, a financial index dynamic rule base is constructed, intelligent adaptation is carried out, and visual rules are generated; based on the visual rules, a double-trigger update mechanism is set;
[0016] Step four: based on the standardized text, entity annotation processing is carried out, entity probability is calculated, and the association relationship between entities is established; complex inquiries are decomposed into independent sub-tasks and sorted;
[0017] Step five: based on the obtained sub-tasks, a calculation engine is constructed, the processing is circularly carried out, real-time NL2SQL data retrieval calculation is executed, the index results are calculated by assigning to hierarchical calculation nodes according to priority, the total execution time is analyzed, and threefold verification is carried out by comparing with the preset financial reasonable interval; based on the large language model LLM, comprehensive intelligent understanding is carried out according to the user's question content and SQL query results, the effectiveness of the index results is judged, core information is summarized, and the conclusion is structured and displayed.
[0018] The application has the following advantages:
[0019] 1. Combine federated learning, homomorphic encryption and Apache-Doris partitioning to achieve secure access to SAP-ERP, CRM multi-source financial data, avoid raw data leakage, and optimize query efficiency through time / business domain partitioning; design dynamic rule base and double-trigger update mechanism to integrate accounting standards parameters, industry adjustment coefficients, and time granularity correction coefficients into index calculation logic, and automatically update through RSS-feed subscription and business mapping, breaking the limitations of traditional fixed calculation logic; use BIO labeling, BERT entity recognition and semantic dependency parsing to improve query understanding accuracy from text matching to entity association and sub-task structuring, solving the problem of complex query disassembly ambiguity;
[0020] 2. Establish three types of input forms of natural language, optimized voice and visual selection, optimize voice recognition model for financial terminology, reduce the threshold for non-professional users, cover multiple scenarios of professional financial personnel fast input and business personnel convenient operation, which is better than the existing single input mode; build a data processing closed loop, where the correction Z-score avoids extreme value interference, the keyword weight and cosine similarity are fused to realize cross-system field standardization, solve the problems of multi-source heterogeneous, quality uneven and privacy sensitive of financial data, dynamic rule base defines the complete index life cycle management of basic attributes, calculation logic, adaptation parameters and update conditions, combined with multi-parameter dynamic calculation and triple verification to ensure the reliability of index calculation, which is different from the existing mode which only focuses on numerical calculation. BRIEF DESCRIPTION OF DRAWINGS
[0021] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings needed to be used in the embodiment description. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can also be obtained by those skilled in the art without creative labor.
[0022] Figure 1 is a step flow chart of a financial intelligent question number accurate method based on dynamic calculation optimization provided by embodiments 1, 2, 3 and 4 of the present application;
[0023] Figure 2 is a step flow chart of a financial intelligent question number accurate method based on dynamic calculation optimization provided by embodiment 5 of the present application. DETAILED DESCRIPTION
[0024] In order to enable personnel in the technical field to better understand the present application, the technical solutions in the embodiments of the present application will be described clearly and completely below in combination with the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, but not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by a person of ordinary skill in the art without creative labor should belong to the protection scope of the present application.
[0025] Embodiment 1: As shown in the figure, the financial intelligent question precision method based on dynamic calculation optimization provided by the embodiment of the present application specifically includes the following: Figure 1
[0026] The user inputs the financial inquiry requirement in natural language, generates the financial question inquiry information, and takes it as the starting point of system processing;
[0027] Relying on the accumulated vertical field industry knowledge base and enterprise knowledge base, the user question is subjected to industry knowledge fusion and standardized processing, and the user question is preprocessed;
[0028] It should be noted that the local knowledge base is in the form of a dictionary, i.e. "tag + value", and the tag and value are in a one-to-one mapping relationship. The dictionary type includes two types: industry knowledge base and enterprise knowledge base. If the user question hits the tag, i.e. the keyword, the rewriting is triggered, the tag is replaced and interpreted as the value corresponding to the tag, the user question is supplemented, the fuzzy expression and the abbreviation term in the user question are identified based on the financial expert experience and the industry standard, the expression ambiguity is corrected, the user intention in the financial context and the enterprise context is accurately understood; the time standardization rewrites the time content in the user question according to the regular matching and keyword hitting mode, accumulates the regular rules and keywords of different granularities of years, quarters, months and days from the time node and time period, and when the user question hits the regular rules or keywords, the corresponding content is mapped and replaced; for example, "in the past 2 years" is explicitly "2023 and 2024", and "from the beginning of the year to now" is rewritten as "from January to the current month";
[0029] Embodiment 2: As shown in the figure, the financial intelligent question precision method based on dynamic calculation optimization provided by the embodiment of the present application specifically includes the following steps: Figure 1
[0030] For complex user questions, a composite inquiry disassembly process is started, the complex requirements are disassembled into logically simple sub-tasks, and multi-entity and multi-dimensional financial analysis is supported;
[0031] Entity Extraction: Identify key entities in user queries, including: time entities, company entities, indicator entities, and analysis entities, accurately locating the analysis object and dimension; specifically, the Named Entity Recognition (NER) algorithm of the semantic representation model BERT is used to extract various types of entities. When the extraction confidence exceeds a preset threshold, the algorithm extraction result is used; otherwise, it is discarded.
[0032] Relationship matching: Constructing the association and pointing relationships between entities, specifically using the entity relationship extraction algorithm of the semantic representation model BERT, filtering by confidence threshold, traversing the filtered reliable relationship pairs, and adding the relationship pairs to the relationship structure in sequence according to the association direction, finally forming a relationship tree diagram;
[0033] Problem decomposition: The complex question is broken down into simple subtasks to reduce computational complexity. A relation tree diagram is constructed by dependency matching. Along the direction of relation flow, the original complex question is decomposed into simple subtasks in sequence to form a list of subtasks.
[0034] Example 3: As Figure 1 As shown in the figure, the present invention provides a method for improving the accuracy of financial intelligent data collection based on dynamic calculation optimization, which specifically includes the following steps:
[0035] The decomposed batch of subtasks are processed in a loop, and real-time dynamic calculations are performed sequentially. By flexibly expanding the formulas and adapting the scope, the problems of granularity differences and nested indicators are solved, ensuring the accuracy and flexibility of complex financial indicator calculations.
[0036] Formula matching: Based on the indicator entity, relevant indicator calculation formulas are obtained by matching and associating them from the accumulated indicator formula library through keyword matching. The indicator formula library contains the calculation of derived indicators and is stored in the form of a dictionary of "label + value". When using it, the problem of multiple matching is solved by using the inverted dictionary label + keyword matching priority.
[0037] Formula expansion: Handles nested indicators and nested analyses, including: formula drill-down, which flattens out the formula to obtain the final expanded formula for nested indicators; and analysis nesting, which overlays year-on-year / month-on-month analysis logic on top of the indicator calculation formula.
[0038] Formula rewriting: The calculation formulas for indicators are refined and adjusted to accommodate different calculation methods, distinguishing between monthly / annual and single-point / cross-sectional data requirements. The expert rule model, incorporating financial experience, mainly includes two parts: main logic processing and basic indicator processing.
[0039] Main logic processing: if the query does not contain "ring ratio" and does not contain "same ratio", call the base_index_deal function to process the financial question information, return the updated financial question information, formula replacement dictionary, formula replacement key list; otherwise, initialize an empty formula replacement dictionary and replacement key list;
[0040] If the query contains "ring ratio" and the current time is the latest date, the current time = the first element of the time combination, the comparison time = the previous month of the current time, generate the ring ratio formula, update the financial question information, call the base_index_deal function to process the indicators of the current time and the comparison time respectively, merge the formula replacement dictionary processed twice, merge the replacement key list processed twice, add the ring ratio formula to the formula replacement dictionary;
[0041] If the query contains "same ratio" and the current time is the latest date, the current time = the first element of the time combination, the comparison time = the previous year of the current time, generate the same ratio formula, update the financial question information, call the base_index_deal function to process the indicators of the current time and the comparison time respectively, merge the formula replacement dictionary processed twice, merge the replacement key list processed twice, add the same ratio formula to the formula replacement dictionary;
[0042] Basic index processing:
[0043] Initialize an empty formula replacement dictionary and replacement key list, extract time and indicators, determine the type of indicators and the source of corresponding data tables, and process according to time format classification:
[0044] Time is year: if the type of indicator is "basic index" type; if the data table is not a specific table and the query does not contain the "monthly / this month" keyword, add the sum suffix to the financial question information, otherwise, replace the time "2024" with "2024 December"; if the type of indicator is other, generate the time of last year, get the indicator formula from the formula dictionary, replace "this year" and "last year" with the current year and last year respectively, update the formula replacement dictionary, and add the indicator to the replacement key list;
[0045] Time is year-month: if the type of indicator is "basic index" type, if the data table is not a specific table and the query does not contain the "monthly / this month" keyword, replace the time "2024 February" with "2024 January to February" and add the sum suffix;
[0046] If the type of indicator is other, generate the time of last year, get the indicator formula from the formula dictionary, replace "this year", "last year" and "this month" with the current year, last year and current month respectively, update the formula replacement dictionary, and add the indicator to the replacement key list;
[0047] Time is quarterly / first and second half: Convert quarterly / first and second half to month range, if the index type is "basic index", if the data table is not a specific table, add the sum suffix to the financial question information; otherwise, replace it with the end month of the month range; if the index type is other, generate the last year time from the formula dictionary to get the index formula, replace "this year", "last year" and "this month" with the current year, last year and the end month of the month range respectively; Update the formula replacement dictionary and add the index to the replacement key list;
[0048] Time is month range: Split the range into single months, update the financial question information, list the index of each month, and for each single month, repeat the time year-month processing logic in the "year-month" format;
[0049] Return the processed financial question information, formula replacement dictionary and replacement key list;
[0050] Data extraction calculation:
[0051] NL2SQL data extraction converts user query content into SQL statements intelligently, extracts the required related data from the financial database, and the implementation is to combine the associated metadata and sqlagent NL2SQL technology, by matching the relevant tables and fields in advance, and as context injection, further improve the accuracy of NL2SQL data extraction;
[0052] Associated metadata:
[0053] It is divided into table matching and field mapping; Table matching: use the text classification algorithm of semantic representation model BERT to understand the content of the rewritten user query, intelligently match the user query with the corresponding data table in the database according to the semantics, improve the accuracy of data table association, and continue operation on the corresponding table matched and associated;
[0054] Field mapping: The field information of all tables in the database is uniformly stored in the form of a dictionary, that is, "label + value", the label is the standardized Chinese name of the field, and the value is the English name of the field in the data table; Based on the standardized Chinese name, all indicators involved in the user query are identified through keyword matching to complete the data table field mapping of the indicators, which improves the accuracy of data table extraction for subsequent NL2SQL;
[0055] Sqlagent:
[0056] Combine sqlagent to intelligently convert user query content into accurate SQL statements to extract the required related data from the financial database;
[0057] Specifically, demand understanding: based on the large language model LLM, the user's question content is intelligently understood, and the core elements in the user's demand are identified: entity, index, intention and constraint;
[0058] SQL generation: based on the injected associated metadata and the understood user demand, the natural language is automatically converted into executable SQL statements conforming to the database structure;
[0059] SQL verification and optimization: the generated SQL is verified in various aspects, if the verification fails, corresponding optimization is performed until the verification passes, syntax verification: check whether the SQL conforms to the syntax rules;
[0060] Semantic verification: check whether the fields and table names exist in the database; efficiency verification: adjust inefficient queries to avoid full table scan operations;
[0061] SQL execution: after the verification passes, the agent executes the SQL through the database connection tool and obtains the query result;
[0062] Result processing: the agent intelligently thinks whether the current SQL query result meets the expectation, if not, how much difference is there, and what operations need to be done, re-perform the demand understanding; if it meets, end the thinking and give the SQL query result;
[0063] Mathematical formula calculation:
[0064] Call the custom mathematical calculator, according to the index calculation formula, perform dynamic real-time calculation to obtain the derived index result, combine the large language model LLM, according to the user's question content and the SQL query result, perform comprehensive intelligent understanding, extract the mathematical formula to be executed, and pass the parameter in the form of a list to the custom mathematical calculator, and return the calculation result in the form of a list after the calculation is completed inside the calculator;
[0065] Embodiment 4: as Figure 1 shown, the financial intelligent question precision method based on dynamic calculation optimization provided by the embodiment of the application specifically includes the following steps:
[0066] Collect the calculation results of each subtask, and integrate and process them relying on the large language model LLM;
[0067] Based on the large language model LLM, all subtask results are sorted out, summarized and refined to extract key conclusions so that the user can better understand;
[0068] Chart visualization: convert result data into intuitive charts to assist financial analysis and decision-making. The chart visualization module mainly includes two steps: factor extraction: combine large language model LLM and text conclusions obtained from content induction link to intelligently extract key elements needed for charting and pass them to the charting tool in the form of a dictionary, i.e. "label + value", where label is the element name and value is the corresponding numerical list of the element;
[0069] Visualization: according to the elements given by the factor extraction and the chart type selected by the user on the interactive interface, the chart is drawn to realize the result visualization;
[0070] It should be noted that BERT is a pre-trained language model based on the Transformer Encoder structure. It achieves strong semantic representation ability through bidirectional context learning. BERT is pre-trained through the Masked Language Model (MLM) task, allowing the model to focus on the left and right context of a word during training. Compared with unidirectional models, it more accurately captures the meaning of ambiguous and polysemous words. After pre-training on large-scale unlabeled text, BERT quickly adapts to downstream tasks through fine-tuning, eliminating the need to train the model from scratch for each task, significantly reducing data requirements and training costs. BERT avoids the complexity of designing specialized models for different tasks in traditional NLP by making minor structural adjustments to handle all tasks uniformly. BERT supports text input up to 512 words, capturing long-range dependencies. Based on the semantic representation model BERT, the downstream text task extracts entities and classifies them from the text. The input text sequence is directly input into BERT, and each word corresponds to an output vector. A Conditional Random Field (CRF) or linear layer is added after each word output to predict entity labels. The text category is determined by converting the text to a word vector through BERT's Tokenizer and adding special symbols. The output vector of the special symbol position is connected to a fully connected layer and a softmax activation function, and the classifier is trained through labeled data. The entity relationship extraction algorithm simultaneously captures text semantics, entity position, and potential relationships to identify specific semantic relationships between entity pairs in the text. The entity recognition loss and relationship classification loss are weighted and summed as the total loss function. The parameters of BERT and the two branches are updated simultaneously during training to achieve collaborative optimization of the two tasks. Expert rule model is a model that processes tasks based on artificial pre-set logical rules. The core principle is to convert domain knowledge into executable conditional-action rules. By matching features in text or data to satisfy rule conditions, the corresponding conclusions are triggered.
[0071] When the model runs, it scans the input data and matches the conditions in the rule library one by one. If the conditions are met, the corresponding results of the rules are output; if not, "no match" or the default result is returned. The rules are defined by humans, and the basis for each decision is clear and visible, making it easy to debug and trust. For simple scenarios or clearly defined domains, you can quickly implement the function by writing rules, without the need for a large amount of labeled data. In the scenario covered by the rules, as long as the rules are designed accurately and the output results are accurate, the rules can be modified or increased or decreased without the need to retrain the model, which is suitable for dynamic adjustment.
[0072] It should be noted that the large language model LLM is a deep learning model trained on massive text data. The core principle is to capture the statistical rules and semantic logic of human language through self-supervised learning, and then generate output text that conforms to the context through autoregressive generation. The large language model is based on the Transformer architecture, which uses self-attention mechanisms to dynamically focus on the relationships between different words in the text. It can understand complex contexts and generate fluent text that conforms to human expression habits. The general language knowledge learned during the pre-training phase allows the model to support dozens of tasks directly through prompt words without the need for separate training for each task. It can handle long contexts, remember previous information, and generate coherent content based on it.
[0073] It should be noted that NL2SQL is a technology that automatically converts natural language questions into structured SQL query statements. The core goal is to allow non-technical users to query databases directly through everyday language without needing to master SQL syntax.
[0074] It should be noted that SQLAgent is an intelligent system that automatically converts natural language into executable SQL and returns the results. Its core is the combination of large language models LLM and database operation capabilities to achieve end-to-end semantic queries.
[0075] It should be noted that LLMMathChain is a component in the LangChain framework specifically designed to handle natural language mathematical problems. The core goal is to address the shortcomings of large language models LLM in precise calculations. The principle is a collaborative mode of natural language understanding + tool-assisted calculation. Large language models are good at understanding natural language and breaking down problem logic, but they are prone to errors in complex mathematical operations. LLMMathChain divides the work between the LLM, which is responsible for understanding the problem and planning the steps, and external computing tools, which are responsible for performing precise calculations, to solve precise calculations.
[0076] Embodiment 5: As shown in the Figure 2 The financial intelligent question and number accurate method based on dynamic calculation optimization provided by the embodiment of the application specifically includes the following steps:
[0077] Step one: Obtain financial question asking information, divide input forms, rewrite financial semantics through double knowledge base, standardize time for multi-dimensional rules, and obtain standardized text through text clarity;
[0078] In specific embodiments, financial question asking information is obtained, and the input method of the asking information is divided into three input forms, covering multi-scenario financial question needs, including:
[0079] Natural language input, exemplary: "Query the gross profit margin and net profit margin of company A in 2023Q1 and company B in 2024Q2" "Asset-liability ratio of ProGroup in the past two years year-on-year";
[0080] Voice input, integrating Aliyun speech recognition API, optimizing speech recognition model for financial terms;
[0081] Visual point selection input, providing "company list, time granularity and index classification" drop-down menu to reduce the operation threshold of non-professional users;
[0082] Rewrite financial semantics through double knowledge base, double knowledge base structure is stored in the form of "label + value" dictionary, and double knowledge base includes two types of core libraries:
[0083] Industry knowledge base: covering financial index aliases and accounting standard terms, exemplary: "revenue" represents "operating income", "ROE" represents "return on equity", "IFRS16" represents "lease standard";
[0084] Enterprise knowledge base: storing enterprise full name and abbreviation mapping and business line exclusive indicators, exemplary: "ProGroup" represents "ProGroup Co., Ltd.", "product revenue" represents "core product sales revenue";
[0085] Rewriting rules: if the input text hits the "label" in the dictionary and the keyword matching degree is greater than or equal to the preset matching degree threshold, then it is automatically replaced with "value", and the complete semantics is supplemented, exemplary: the original question "ProGroup's profitability in the past two years" is rewritten as "ProGroup Co., Ltd. 2023-2024 gross profit margin, net profit margin, return on equity, total asset return";
[0086] Standardize time for multi-dimensional rules, rule system: integrate regular expressions and time keywords for scene-by-scene processing:
[0087] Time period class, exemplary: "the past two years" represents "the current year minus one year to the current year";
[0088] Relative time class, exemplary: "from the beginning of the year to now" represents "from January of the current year to the current month";
[0089] Fuzzy Period Category: "Same Period" means "Last Year Same Quarter or Same Month";
[0090] Based on time standardization, the text is clear to get the standardized text, Jieba segmentation and financial dictionary are used for segmentation to avoid splitting "gross profit rate" into "gross profit" and "rate"; stop words are removed to filter meaningless words; part-of-speech tagging is used to tag four types of entities, including "company", "time", "indicator" and "analysis dimension", and provides the basis for disassembly;
[0091] Step two: Based on the standardized text, build a data access gateway, deploy a local data node, transmit data metadata, and partition according to business domain and time granularity, and perform data quality optimization processing;
[0092] In specific embodiments, the data access gateway is built through the FATE (Federated AITechnology Enabler) federated learning framework:
[0093] Each data source, such as SAP-ERP, Jinsan Third Phase, and Salesforce CRM, deploys a local data node to transmit data metadata, including field name, data type, and timestamp;
[0094] Homomorphic encryption technology is used to encrypt metadata to avoid leakage of original financial data;
[0095] Based on the data access gateway, Apache Doris columnar database is used to partition according to business domain and time granularity:
[0096] Business domains include but are not limited to revenue domain, cost domain, R&D domain, and liability domain;
[0097] According to the time granularity, for example, "revenue domain-2023Q1" and "R&D domain-202305";
[0098] Query optimization: support fast slicing by time granularity, for example: query "2023 monthly R&D expenses", scan "R&D domain-202301" to "R&D domain-202312" partitions;
[0099] Design data cleaning, data standardization, and data alignment to perform data quality optimization processing, the specific rules are as follows:
[0100] Data cleaning: correct the standard score, identify potential outliers, and perform secondary verification according to the preset financial reasonable interval;
[0101] Specifically, for any financial data sequence, the correction standard score calculation formula is:
[0102]
[0103] The modified standard score of the financial data sequence is calculated , wherein, is the i-th data sample, is the median of the data sequence X, avoiding the influence of extreme values on the mean, is the data sequence, is the median absolute deviation of the data sequence X, measuring the degree of dispersion, and is not disturbed by extreme values;
[0104] Based on the obtained modified standard score, secondary verification is performed according to a preset reasonable financial interval. If the absolute value of the modified standard score is greater than a preset standard score threshold, secondary verification is performed. If the data sample satisfies any one of the preset reasonable financial intervals in the secondary verification, it is determined that the business is normally fluctuating, and the abnormal mark is cancelled;
[0105] For example, if the data sample ∈ [historical same period data x 0.7, historical same period data x 1.3], it is determined that the business is normally fluctuating, and the abnormal mark is cancelled; if the data sample ∈ [industry mean x 0.8, industry mean x 1.2], it is determined that the business is normally fluctuating, and the abnormal mark is cancelled;
[0106] Based on the completion of data cleaning, data standardization is performed through financial ontology mapping. In view of the differences between different system fields, the combination of preset weights of keywords and cosine similarity is adopted to realize standardization. The similarity between the to-be-mapped field f and the standard financial field F is calculated by the formula: , is the preset weight of the keyword, is the keyword of the to-be-mapped field f, is the standard financial field, is the total number of keywords, is the preset weight index, is the cosine similarity between the keyword and the standard field;
[0107] For example, the to-be-mapped field f is “main income”, the keywords of the to-be-mapped field f are “main” and “income”, the preset weight of the keyword “main” is 0.6, and the preset weight of the keyword “income” is 0.4. The cosine similarity between the keyword “main” and the standard field “main business” is 0.92.
[0108] Based on the completion of data standardization, data alignment is performed through multi-time granularity adaptation;
[0109] Specifically, the data is aligned according to the minimum time granularity through the aggregation rule:
[0110]
[0111] The target granularity data after the analysis is obtained , wherein, is a coefficient of aggregation, is original granularity data, is end-of-period granularity data, is an index of the original granularity data;
[0112] For example, the annual revenue in 2023 is 120 million yuan, and when the monthly data is split, if there is no detailed allocation, it is allocated according to the historical monthly proportion. If the historical monthly proportion of January is 8%, the revenue in January 2023 is 960 million yuan, and if the historical monthly proportion of January is 7.5%, the revenue in January 2023 is 900 million yuan, ensuring the consistency of multi-granularity calculation;
[0113] Step three: based on the data quality optimization processing, a financial indicator dynamic rule library is constructed, and intelligent adaptation is performed to generate visual rules; based on the visual rules, a double-trigger update mechanism is set;
[0114] Specifically, the structure of the financial indicator dynamic rule library is designed, and each indicator rule includes five parts: basic attributes, calculation logic, granularity adaptation, industry parameters, and update conditions;
[0115] For example, taking the R&D expense capitalization rate as an example, the financial indicator dynamic rule library structure of the R&D expense capitalization rate is designed through Table 1 shown below:
[0116] Table 1 Financial indicator dynamic rule library structure statistical table of R&D expense capitalization rate
[0117] Rule field R&D capitalization rate Data type Explanation Indicator ID IND-FIN-001 String Unique identifier for associating calculation logic Indicator name R&D capitalization rate String Financial standard name Base indicator dependency Capitalized R&D expenses (IND-FIN-0011), Total R&D expenses (IND-FIN-0012) List Associated base indicator IDs to ensure calculation link integrity Calculation logic template Capitalization rate = Capitalized R&D expenses / Total R&D expenses × 100% Formula string Supports variable substitution, adapts to different standards Granularity adaptation coefficient Year: p = 1; Quarter: p = 0.95; Month: p = 0.9 Key-value pair Corrects calculation bias caused by short-term data fluctuations Accounting standards parameters IFRS: α = 1 (includes development stage expenses); China GAAP: α = 0 (does not meet 5 conditions) Key-value pair Adapts to differences in capitalization scope under different standards Industry adjustment parameters High-tech industry: γ = 1.75 (addition and deduction coefficient); Traditional industry: γ = 1.0 Key-value pair Covers industry-specific accounting requirements Update trigger conditions Revision of accounting standards (e.g. IASB releases new R&D standards), addition of new R&D lines List Defines the trigger scenarios for rule updates
[0118] Based on the financial indicator dynamic rule library structure, configuration is completed through a Web interface to generate visual rules;
[0119] Specifically, the basic indicators are selected: the dependent items are selected from the basic indicator library;
[0120] Edit the calculation logic: drag the formula component to generate a logic template;
[0121] Set the adaptation parameters: input the granularity coefficient p, the criterion parameter a, and the industry parameter g;
[0122] Rule verification: the system verifies through a rule verification engine, and writes into the Neo4j database after verification;
[0123] Dynamically calculate the financial indicators in combination with the input granularity coefficient p, the criterion parameter a, and the industry parameter g;
[0124] Exemplary, for example, with R&D capitalization rate, integrate standards, granularity, industry parameters, dynamic calculation formula:
[0125]
[0126] The R&D capitalization rate is calculated , wherein, is the capitalization of R&D expenses, is the accounting standards parameter, is the industry adjustment parameter, is the cost of R&D expenses, is the time granularity coefficient;
[0127] Based on the visualization rules, set up a double trigger update mechanism;
[0128] Specifically, the standard change trigger: the system subscribes to the RSS-feed of the Ministry of Finance and IASB website, and monitors the accounting standards revision in real time; If the revision involves index calculation logic, automatically push the update reminder to financial experts; After the expert confirms, the system adjusts the calculation logic template of the related indicators in batches;
[0129] Business adjustment trigger: if the enterprise adds a new business line, the system automatically adds exclusive indicators through the business-index pre-configuration mapping relationship; The associated basic indicators depend on and inherit the general calculation logic;
[0130] Step four: based on the standardized text, perform entity annotation processing, calculate entity probability, establish the association relationship between entities, and decompose the complex inquiry into independent sub-tasks and sort them;
[0131] In specific embodiments, based on the obtained standardized text, further entity annotation processing is performed;
[0132] The BIO annotation system is adopted, wherein B is the entity start, I is the entity middle, O is the non-entity, and four types of core entities are annotated, including: company entity, time entity, index entity and analysis dimension entity;
[0133] For the text sequence after entity annotation, the entity probability of each token is calculated through the BERT model, through the formula:
[0134]
[0135] The entity probability is calculated , wherein, is the i-th token in the text sequence, and is the entity type, is the weight matrix of the entity type , and is the entity type The weight matrix, For the output of the BERT model Hidden layer vectors For entity type The bias term, For entity type The bias term;
[0136] Based on semantic dependency analysis, relationships between entities are established, and complex queries are broken down into independent subtasks. Each subtask includes four elements: computational object, time granularity, financial indicators, and visualization rules, expressed through the formula:
[0137]
[0138] Get subtask ,in, Index for subtasks For the calculation object, For time granularity, For financial indicators, For visualizing rules, For analysis dimensions;
[0139] The data acquisition difficulty and computational complexity are weighted and sorted to prioritize the execution of time-consuming subtasks.
[0140] Step 5: Based on the obtained subtasks, build a computing engine, allocate the results to the hierarchical computing nodes according to priority, analyze the total execution time, compare it with the preset financial reasonable range for triple verification, determine the validity of the indicator results, summarize the core information, and present the conclusions in a structured manner.
[0141] In a specific embodiment, a computing engine is constructed, and subtasks are assigned to computing nodes according to priority;
[0142] Layered calculations are performed using basic indicators, single-layer derived indicators, and multi-layer nested indicators to cover requirements of varying complexity.
[0143] Specifically, basic indicators are extracted based on data quality optimization; single-layer derived indicators are obtained by combining time granularity coefficient correction; and multi-layer nested indicators are calculated through recursive decomposition and layered execution.
[0144] Based on hierarchical computing, using the formula:
[0145]
[0146] The total execution time was calculated. ,in, This represents the total number of subtasks. Index for subtasks is the computation consumption of the subtask j, is the number of computing nodes, is the processing capacity per second of a single node;
[0147] The triple verification is performed on the preset financial reasonable interval to determine the effectiveness of the index result, and the formula is:
[0148]
[0149] The analysis result is effective , wherein, is the index result, is the minimum value of the preset financial reasonable interval, is the maximum value of the preset financial reasonable interval;
[0150] The trend similarity of the index result and the historical data is calculated to avoid sudden abnormalities, and the formula is:
[0151]
[0152] The trend similarity is calculated , wherein, is the index result, is the last period index result, is the last but one period index result;
[0153] The result is determined whether it meets the business needs in combination with the business goals of the enterprise, and the formula is:
[0154]
[0155] The demand effectiveness is calculated , wherein, is the business goal value of the enterprise, is the index result;
[0156] If or appears in the triple verification, the calculation parameters are automatically adjusted, and the calculation parameters include granularity coefficient and industry parameters;
[0157] Based on all the subtask results, the core information is summarized, and the conclusion is structured and displayed according to the structure of "company, time, index, calculation result, benchmark value and verification result". LLM extracts the drawing elements from the summarized conclusion and passes them to the visualization tool in the form of "label + value" dictionary.
[0158] It should be noted that LLM represents a large language model, which is a deep learning model trained based on massive text data, and its core capabilities are understanding, generating and reasoning human language.
[0159] The above describes one embodiment of the present application in detail, but the content is only the preferred embodiment of the present application and cannot be considered to limit the scope of the present application; the above formulas are all dimensionless values, and the formulas are obtained by collecting a large amount of data to simulate a formula of the most recent real situation, and the preset parameters in the formula are set by the person skilled in the art according to the actual situation and historical experience, and can be adjusted according to the actual situation; the above is only the preferred embodiment of the present application and cannot be used to limit the present application, and all equivalent changes and improvements made according to the scope of the present application should still belong to the patent coverage range of the present application.
Claims
1. A method for improving the accuracy of financial intelligent data collection based on dynamic calculation optimization, characterized in that, Includes the following steps: The system acquires financial question information and classifies the input format. It then rewrites the financial semantics using a dual knowledge base, standardizes multi-dimensional rules over time, and clarifies the text to obtain standardized text. A data access gateway is built based on standardized text, local data nodes are deployed, data metadata is transmitted, complex requirements are broken down into logically simple sub-tasks, supporting multi-entity and multi-dimensional financial analysis; data is partitioned according to business domain and time granularity, and data quality optimization processing is performed. Based on the completion of data quality optimization processing, a dynamic rule base for financial indicators is constructed and intelligently adapted to generate visualized rules; Based on visual rules, a dual-trigger update mechanism is set up; The method for setting up the dual-trigger update mechanism is as follows: Design a dynamic rule base structure for financial indicators. Based on this structure, configuration is completed through a web interface to generate visual rules. Combine the input granularity coefficient p, the standard parameter α, and the industry parameter γ to dynamically calculate financial indicators. Based on the visual rules, a dual-trigger update mechanism is set up. Standard change trigger: real-time monitoring of accounting standard revisions. If the revision involves the calculation logic of indicators, an update reminder will be automatically pushed to the financial experts; after the experts confirm, the system will adjust the calculation logic template of the relevant indicators in batches; business adjustment trigger: if the enterprise adds a new business line, the system will automatically add exclusive indicators through the business-indicator pre-configuration mapping relationship; associate basic indicator dependencies and inherit the general calculation logic; Entity annotation is performed based on standardized text, entity probabilities are calculated, relationships between entities are established, and complex queries are broken down into independent subtasks and sorted. Based on the obtained subtasks, a computing engine is built to process them in a loop, execute real-time NL2SQL data retrieval and calculation, allocate the results to hierarchical computing nodes according to priority to obtain the indicator results, analyze the total execution time, perform triple verification by comparing with the preset financial reasonable range, and based on the large language model LLM, perform comprehensive intelligent understanding based on user questions and SQL query results to determine the validity of the indicator results, summarize the core information, and present the conclusions in a structured manner. The method for performing triple verification is as follows: The validity of the indicator results is determined by comparing them with a preset reasonable financial range using a triple verification method, and by formula: Analysis of the validity of the results ,in, For the indicator results, To preset the minimum value of the reasonable financial range, To set the maximum value of the pre-defined reasonable financial range; To calculate the trend similarity between the indicator results and historical data and avoid abrupt changes, the formula is: Calculate the trend similarity ,in, For the indicator results, This refers to the results of the previous period's indicators. This refers to the indicator results from the period before last; Based on the company's business objectives, determine whether the results meet operational needs using the formula: Calculate the validity of the demand ,in, For the enterprise's business target value, For indicator results; If the triple check fails or Then the calculation parameters will be adjusted automatically.
2. The method for improving the accuracy of financial intelligent data collection based on dynamic calculation optimization according to claim 1, characterized in that, The method for text clearing is as follows: We acquire financial questions and categorize the input methods into three types, then rewrite the financial semantics using a dual knowledge base. Time standardization is performed on multi-dimensional rules, and regular expressions and time keywords are integrated for scenario-based processing; Based on the time standardization results, four types of entities are labeled to provide a basis for decomposition, handle the nesting of indicators and analysis, distinguish the granularity of monthly / annual and single-point / cross-sectional data, and adapt to different calculation methods; the main logic is processed in conjunction with an expert rule model.
3. The method for improving the accuracy of financial intelligent data collection based on dynamic calculation optimization according to claim 1, characterized in that, The method for NL2SQL data retrieval and calculation is as follows: NL2SQL intelligently transforms user queries into precise SQL statements, extracting the required data from the financial database. This is achieved by combining NL2SQL technology with associated metadata and SQLagent, pre-matching relevant tables and fields, and injecting them as context. By combining SQLAgent with the intelligent conversion of user queries into SQL statements, the SQLAgent process extracts the necessary data from the financial database. The SQLAgent process includes: requirement understanding, SQL generation, SQL verification and optimization, semantic verification, SQL execution, and result processing.
4. The method for improving the accuracy of financial intelligent data collection based on dynamic calculation optimization according to claim 3, characterized in that, The method for SQL validation and optimization is as follows: Needs Understanding: Based on the Large Language Model (LLM), intelligent understanding of user queries is performed to identify the core elements of user needs: entities, metrics, intents, and constraints; SQL generation: Based on injected relational metadata and understood user requirements, it automatically converts natural language into executable SQL statements that conform to the database structure; SQL validation and optimization: The generated SQL is validated in various aspects. If the validation fails, the corresponding optimization is performed until the validation passes. Syntax validation: Checks whether the SQL conforms to the syntax rules.
5. The method for improving the accuracy of financial intelligent data collection based on dynamic calculation optimization according to claim 3, characterized in that, The method for processing the results is as follows: Check if the fields and table names exist in the database; Efficiency verification: Adjust inefficient queries to avoid full table scans; After successful verification, the agent executes SQL through a database connection tool and retrieves the query results; The agent uses intelligent thinking to determine whether the current SQL query result meets expectations. If it does not, it determines the extent of the difference, what further operations are needed, and re-executes the requirement analysis. If it does meet the expectations, the thinking process ends, and the SQL query result is provided.
6. The method for improving the accuracy of financial intelligent data collection based on dynamic calculation optimization according to claim 1, characterized in that, The method for establishing relationships between entities is as follows: Obtain standardized text for further entity annotation processing; The BIO annotation system is adopted to annotate four types of core entities; For the text sequence after entity annotation, the entity probability of each token is calculated using the BERT model, using the formula: Calculate the entity probability ,in, For the i-th token in the text sequence, and For entity type, For entity type The weight matrix, For entity type The weight matrix, For the output of the BERT model Hidden layer vectors For entity type The bias term, For entity type The bias term; Based on semantic dependency analysis, relationships between entities are established, and complex queries are decomposed into independent subtasks, using the formula: Get subtask ,in, Index for subtasks For the calculation object, For time granularity, For financial indicators, For visualizing rules, For analysis dimensions; The data acquisition difficulty and computational complexity are weighted and sorted to prioritize the execution of time-consuming subtasks.
7. The method for improving the accuracy of financial intelligent data collection based on dynamic calculation optimization according to claim 1, characterized in that, The method for obtaining the total execution time in the analysis is as follows: Build a computing engine and assign subtasks to computing nodes according to priority; Layered calculations are performed using basic indicators, single-layer derived indicators, and multi-layer nested indicators to cover requirements of varying complexity. Basic indicators are extracted and corrected using time granularity coefficients to obtain single-layer derived indicators; multi-layer nested indicators are calculated through recursive decomposition and layered execution. Based on hierarchical computing, using the formula: The total execution time was calculated. ,in, This represents the total number of subtasks. Index for subtasks The computational cost of subtask j. To calculate the number of nodes, The processing capacity per second for a single node.
8. The method for improving the accuracy of financial intelligent data collection based on dynamic calculation optimization according to claim 1, characterized in that, The method for setting up integrated intelligent understanding is as follows: Call the custom mathematical calculator to perform dynamic real-time calculations based on the indicator calculation formulas to obtain the derived indicator results. Combined with the Large Language Model (LLM), it performs comprehensive intelligent understanding based on the user's question content and SQL query results, extracts the mathematical formulas to be executed, and passes them to the custom mathematical calculator in list form. After the calculator completes the calculation internally, it returns the calculation results in list form. Based on the Large Language Model (LLM), the results of all subtasks are analyzed, summarized, and key conclusions are extracted. The results data are transformed into intuitive charts to assist financial analysis and decision-making. The chart visualization module mainly includes two steps: Element extraction: Combining the Large Language Model (LLM), the key elements needed for charting are intelligently extracted from the text conclusions obtained in the content summarization stage, and passed to the charting tool in dictionary form, namely "label + value", where the label is the element name and the value is the list of numerical values corresponding to the element.
Citation Information
Patent Citations
Financial field intelligent data analysis platform based on large language model
CN119025544A
Automated Prompt Augmentation And Engineering Using ML Automation In SQL Query Engine
US20250284688A1