Financial question and answer numerical reasoning system and method based on deep learning method
By adopting heterogeneous graph structure and Gaussian process in the financial numerical inference model, the shortcomings in table structure information processing, diversity of generation process and context dependence are solved, and efficient financial question-and-answer numerical inference is achieved, improving the performance and generation quality of the model.
Patent Information
- Application Number
- CN202510254001.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-05
- Publication Date
- 2025-06-17
AI Technical Summary
The existing financial numerical inference model has shortcomings in processing table structure information, generation process diversity and context dependence, resulting in poor performance of the model in complex financial scenarios.
A heterogeneous graph structure is used to build a graph retriever, which enhances information interaction and structural information supplement through graph structure, and introduces meaningful diversity through Gaussian processes to improve the quality of generated text.
It realizes efficient retrieval and expression generation of document and table information, improving the accuracy of the model in complex financial scenarios and the quality of generated results.
Smart Images

Figure CN120163249A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of financial Q&A numerical reasoning, and particularly to a deep model for financial Q&A numerical reasoning based on a deep learning method. Background Art
[0002] Financial numerical reasoning is a core task in the financial field and plays a very important role in financial report analysis. Financial analysis is the core method for evaluating a company's performance, and a slight error in the analysis may lead to huge losses. [1] . In order to make high-quality and timely decisions, professionals need to perform complex quantitative analysis on financial reports to extract important insights. These analyses not only rely on the accurate processing of data but also require strong numerical reasoning abilities, such as comparing the profitability or growth rate of enterprises. However, with the rapid increase in corporate financial documents [2] , the analysis difficulty has also increased accordingly, which has triggered the thinking about whether deep financial analysis can be carried out through automated means. Therefore, improving the numerical reasoning ability of artificial intelligence in financial analysis has become particularly urgent and crucial.
[0003] Most traditional financial numerical reasoning methods are end-to-end models, which use sequence-to-sequence models [3] or syntactic parse trees [4] to generate mathematical expressions for answering questions. By inputting the text into the model for encoding, the corresponding expression parameters are generated directionally through a decoder or a tree structure. However, this method has certain limitations. Financial reports in the financial field are already very large, and their text length far exceeds the scope that common pre-trained models can handle. At the same time, only a few sentences may be required to answer the question, and a large number of sentences are irrelevant and will cause interference. Inputting all sentences into the model will invisibly increase the difficulty of the model to solve the problem, making this task enter a bottleneck period.
[0004] The proposal of the method based on the retriever-generator architecture [5] , has put forward new ideas for this task. This type of method first uses a retriever to retrieve the sentences in the document that are most relevant to the question, so as to reduce the impact of overly long text and a large number of irrelevant sentences on the model. Then, the retrieved sentences are input into the generator model, and the generator is used to generate the final expression sequence.
[0005] However, the existing methods based on the retriever-generator architecture also have some deficiencies:
[0006] (1) Lack of association in table structure: Many existing models directly flatten table data into a one-dimensional sequence and convert it into sentence expressions in a fixed form, ignoring the two-dimensional row-column structure of the table. The structured information in the table is crucial for understanding the relationships between data and the reasoning process, which can assist the model in quickly locating the correct table data. Losing this information will make it difficult to obtain key facts during the retrieval process, thereby affecting the model's reasoning ability. Therefore, how to capture the internal associations of the table while unifying the table and document representations has become an urgent problem to be solved.
[0007] (2) Insufficient diversity in the generation process: When generating arithmetic expressions, due to the limitation of the training objective, existing models tend to generate a pre-designed standard expression, lacking sufficient diversity. In some complex problems, there may be multiple different solutions to the same problem. For example, both 2×4 and 4 + 4 can yield the same result of 8. Letting the model learn both of these expressions can better improve the model's reasoning ability. Therefore, increasing the generation diversity can provide more reasoning paths for the model and improve the quality of the generated results. However, how to introduce meaningful diversity while maintaining the quality of the generated text is a key challenge.
[0008] (3) Insufficient context-dependent information: Numerical reasoning often depends on context information, such as the scenario where the problem is located, the data relationships in the table, etc. Existing methods have limitations in modeling complex context dependencies and fail to interact information from multiple perspectives. Especially in scenarios in the financial field involving long texts and multi-table data, this also leads to the problem of insufficient information in the reasoning process and needs to be corrected.
[0009] To solve the above problems, the present invention introduces a heterogeneous graph structure to construct a graph retriever, enhances information interaction and structural information complement through the graph structure, and constructs a controllable perturbation through Gaussian processes. Without affecting semantic information, meaningful diversity is introduced, thereby improving the quality of the generated text.
[0010] [References]
[0011] [1] Jerven M. Poor numbers: how we are misled by African development statistics and what to do about it [M]. Cornell University Press, 2013.
[0012] [2]MacKenzie D,Beunza D,Millo Y,et al.Drilling through the AlleghenyMountains:Liquidity,materiality and high-frequency trading[J].Journal ofcultural economy,2012,5(3):279–296.
[0013] [3]Zhu F,Lei W,Huang Y,et al.TAT-QA:A question answering benchmark onahybrid of tabular and textual content in finance[J].arXiv preprint arXiv:2105.07624,2021.
[0014] [4]Xie Z,Sun S.A goal-driven tree-structured neural model for mathword problems.[C].In Ijcai.,2019:5299–5305.
[0015] [5]Chen Z,Chen W,Smi ley C,et al.Finqa:A dataset of numericalreasoning over financial data[J].arXiv preprint arXiv:2109.00122,2021. Summary of the Invention
[0016] The present invention aims to solve the problems in existing financial numerical reasoning models, such as insufficient processing of tabular structure information, lack of diversity in the generation process, and insufficient context-dependent modeling, and provides an efficient solution for financial report data. Specifically, the present invention proposes a financial question-answering numerical deep reasoning system based on deep learning methods, which retains multi-dimensional information of tabular and text data through a heterogeneous graph structure, and enhances the diversity of generated expressions through a Gaussian process random function, thereby improving the performance of the model in complex financial scenarios.
[0017] The objectives of the present invention are achieved through the following technical solutions:
[0018] A financial Q&A numerical reasoning system based on deep learning methods, the deep reasoning system includes a retriever and a generator; the retriever consists of a first encoder, a graph attention module and a semantic aggregation module; the generator consists of a second encoder, a semantic optimization module and a decoder; where:
[0019] The graph attention module establishes numerical associations of question, text, and table-related information through a heterogeneous graph; where:
[0020] The semantic aggregation module is trained according to the numerical associations to obtain a fact-based feature representation;
[0021] The generator optimizes and generates multiple groups of numerical semantic expressions based on the fact-based feature representation through a random function.
[0022] Furthermore, the process by which the graph attention module constructs numerical associations of question, text, and table-related information through a heterogeneous graph according to the input question includes:
[0023] 101. Create node representations of the table, document, and question:
[0024] For each cell in the table, create a table node; each table node represents a data point in the table;
[0025] For each sentence in the document, create a document sentence node; each document sentence node represents a semantic unit in the document;
[0026] For each financial question to be answered, create a question node; the question node represents the query or question raised by the user;
[0027] 201. Create edge representations of the table, document, and question:
[0028] Create internal table relationship edges according to the row and column relationships between table nodes;
[0029] Create internal document relationship edges according to the semantic relationships between document sentence nodes;
[0030] Create associated edges between the table and the document according to the relationship between table data and document content between table nodes and document sentence nodes;
[0031] Create associated edges between the question and table data or document content according to the relationship between the question node and the table node and the document sentence node;
[0032] 301. Construct a heterogeneous graph representing the two-dimensional structure information of the table according to the node representations in step 101 and the edge representations in step 201;
[0033] 401. The graph attention module aggregates nodes according to the type of edge.
[0034] 501. The semantic aggregation module calculates and outputs a fact feature representation according to the aggregation nodes based on similarity.
[0035] Furthermore, the semantic aggregation module obtains the process of fact feature representation through training based on numerical association, including:
[0036] Extract the numerical association information of each node in the heterogeneous graph for weighted aggregation;
[0037] Perform cross-node information exchange on each weighted node according to the following formula:
[0038]
[0039] e uv = LeakyReLU(a T [e u ||e v )
[0040]
[0041] where, represents the feature representation of node u at the l-th layer, N(u) is the neighborhood of node u, e uv is the interaction feature representation between node u and node v, is the attention correlation coefficient between node u and node v when the edge type is k.
[0042] Furthermore, the generator optimizes and generates multiple groups of numerical semantic expressions based on the fact feature representation through a random function, including:
[0043] The second encoder performs latent variable representation based on the fact feature representation according to the following formula:
[0044] h = Encoder(x)
[0045] The optimization semantic module optimizes the fact feature representation according to the following random parameters:
[0046] z = g(h)+ε
[0047] The decoder calculates and outputs multiple groups of numerical semantic representations for the optimized numerical semantics according to the following formula:
[0048] y t = Decoder(y t-1 ,z)
[0049] where z is the input of the decoder, sampled from a probability distribution, ε is a trainable parameter, g() is a random function, x is the input of the encoder, h is the output of the encoder, yt is the output at time t;
[0050] Furthermore, the generator trains the random function through the following steps, including:
[0051] Obtain the random function g() through variational inference:
[0052] logp(y|h)≥E q [log p(y|z)] - KL(q(z|h,y)||p(z|h))
[0053] where q(z|h,y) is the approximate result;
[0054] Simplify the verification of multiple groups of numerical semantics through the following formula:
[0055]
[0056] where μ and σ 2 are the mean and covariance of the network.
[0057] Beneficial effects
[0058] Compared with the prior art, the beneficial effects brought by the technical solution of the present invention are:
[0059] 1. The novel financial Q&A numerical reasoning deep model based on the heterogeneous graph retriever and Gaussian process-inspired generator architecture proposed by the present invention realizes the retrieval of document and table information and the generation of corresponding expressions. On the validation set of the FinQA dataset, it reaches a correct rate of 70.96%, and on the test set of the FinQA dataset, it reaches a correct rate of 69.04%.
[0060] 2. The present invention adopts the collaborative optimization of the generator and the retriever. The heterogeneous graph and Gaussian process random function of the present invention cooperate with each other, significantly improving the performance of the model in both the retrieval and generation stages. The retriever enhances the model's ability to identify key facts through the heterogeneous graph, while the generator increases the flexibility of the generation process through the Gaussian process. Through this collaborative optimization, the present invention can process complex financial data and generate high-quality answers.
[0061] 3. The present invention solves numerical reasoning problems based on the sequence-to-sequence (Seq2Seq) model. After the input data is encoded by the BERT encoder, the generator generates a mathematical expression for the problem by combining the problem and the retrieved key facts.
[0062] Different from traditional generative models, the present invention introduces a Gaussian process random function in the encoder to add perturbations containing context information to the latent variables. These randomized context perturbations enable the generator to flexibly generate different expressions, which is very helpful for complex numerical reasoning problems. BRIEF DESCRIPTION OF THE DRAWINGS
[0063] Figure 1 FIG. is a structural diagram of a financial Q&A numerical reasoning system based on a deep learning method of the present invention;
[0064] Figure 2 FIG. is a flowchart of a financial Q&A numerical reasoning method based on a deep learning method of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0065] The following further describes in detail the technical solution of the present invention in conjunction with the attached Figures 1 to 2 drawings, but the protection scope of the present invention is not limited to the following.
[0066] As Figure 1 shown, the present invention provides a financial Q&A numerical reasoning system based on a deep learning method. The deep reasoning model is based on a graph neural network and includes a retriever and a generator; the retriever consists of a first encoder, a graph attention module, and a semantic aggregation module; the generator consists of a second encoder, a semantic optimization module, and a decoder; wherein:
[0067] In the retriever, the graph attention module constructs numerical associations of question, text, and table-related information through a heterogeneous graph; including:
[0068] In the data preprocessing stage of the present invention, the input financial table and document data are modeled into a heterogeneous graph structure. For each cell in the table, a table node is created; for each sentence in the document, a document sentence node is created; for each financial question to be answered, a question node is created. The nodes are connected by different types of edges, such as the edges between table rows and columns, the edges between sentences within the document, and the edges between the table, document, and question nodes. Then the graph neural network aggregates information according to different types of edges.
[0069] The modeling process of the heterogeneous graph preserves the two-dimensional structure information of the table, enabling the model to better capture the hierarchical structure and row-column relationship of the table data during the retrieval process, and avoiding the information loss problem caused by flattening the table data in traditional methods.
[0070] Graph Attention Network Module
[0071] After the heterogeneous graph is constructed, the model uses the Graph Attention Network (GAT) to process the nodes in the graph. GAT performs information-weighted aggregation on each node to find the most beneficial information from the information of different nodes, thus achieving efficient cross-node information exchange.
[0072] In the retrieval stage of the system, the question node interacts with the table and document nodes through GAT, thereby enhancing the model's ability to identify key facts. GAT can capture the fine-grained relationships between nodes through multiple levels, ensuring that the model can extract key information related to the question from complex financial data. Among them: The graph attention module numerically associates the information related to the question, text, and table according to the input question through the heterogeneous graph, including:
[0073] 101. Create node representations of the table, document, and question:
[0074] Create a table node for each cell in the table; each table node represents a data point in the table;
[0075] Create a document sentence node for each sentence in the document; each document sentence node represents a semantic unit in the document;
[0076] Create a question node for each financial question to be answered; the question node represents the query or question raised by the user;
[0077] 201. Create edge representations of the table, document, and question:
[0078] Create internal table relationship edges according to the row and column relationships between table nodes;
[0079] Create internal document relationship edges according to the semantic relationships between document sentence nodes;
[0080] Create associated edges between the table and the document according to the relationship between the table data and the document content between the table node and the document sentence node;
[0081] Create associated edges between the question and the table data or the document content according to the relationship between the question node and the table node and the document sentence node;
[0082] 301. Construct a heterogeneous graph representing the two-dimensional structure information of the table according to the node representation in step 101 and the edge representation in step 201;
[0083] 401. The graph attention module aggregates the nodes according to the type of edge;
[0084] 501. The semantic aggregation module calculates the output based on the fact feature representation according to the similarity of the aggregated nodes.
[0085] The semantic aggregation module is trained according to numerical associations to obtain a factual feature representation process, including:
[0086] Use a graph attention network (GAT) to process the nodes in the graph. GAT realizes efficient information exchange across nodes by performing information weighted aggregation on each node and finding the most beneficial information for itself from the information of different nodes. Specifically, its formula is as follows:
[0087]
[0088] e uv = LeakyReLU(a T [e u ||e v )
[0089]
[0090] Among them, represents the feature representation of node u in the l-th layer, N(u) is the neighborhood of node u, and e uv is the interaction feature representation between node u and node v, is the attention correlation coefficient between node u and node v when the edge type is k.
[0091] The generator optimizes and generates multiple groups of numerical semantic expressions based on the factual feature representation through a random function;
[0092] In the generation stage, the present invention introduces a Gaussian process stochastic function (GPSF) to enhance the diversity of the generated expressions. Specifically, GPSF generates multiple possible context representations by perturbing the hidden states in the encoder. GPSF introduces Gaussian noise into the encoder, making the hidden states have appropriate variability while retaining the original information. This mechanism enables the generator to generate multiple different arithmetic expressions, which is particularly suitable for problems with multiple solutions in the financial field. Specifically:
[0093] The second encoder performs latent variable representation based on the factual feature representation according to the following formula:
[0094] h = Encoder(x)
[0095] The optimization semantic module optimizes the factual feature representation according to the following random parameters:
[0096] z = g(h)+ε
[0097] The decoder calculates and outputs multiple groups of numerical semantic representations for the optimized numerical semantics according to the following formula:
[0098] y t= Decoder(y t-1 , z)
[0099] where z is the input of the decoder, sampled from a probability distribution, ε is a trainable parameter, g() is a random function, x is the input of the encoder, h is the output of the encoder, and y t is the output at time t. Meanwhile, the following steps are adopted for the training process of the random function, including:
[0100] And obtain the random function g() through variational inference:
[0101] logp(y|h) ≥ E q [log p(y|z)] - KL(q(z|h,y) || p(z|h))
[0102] where q(z|h,y) is the approximate result;
[0103] Simplify the verification of multiple groups of numerical semantics through the following formula:
[0104]
[0105] where μ and σ 2 are the mean and covariance of the network.
[0106] Finally, the system is tested and trained on two question-and-answer datasets. The hyperparameter settings in the experiment are shown in Table 1, and the detailed descriptions of the two question-and-answer datasets are shown in Table 1.
[0107] Table 1 Experimental hyperparameter settings
[0108]
[0109] Table 2 Statistical results of each classification dataset
[0110]
[0111] Table 3 Experimental results of the question-and-answer matching dataset It can be seen from Table 3 that compared with the previous baseline model, our work has achieved the best performance on the FinQA dataset and performed well on the ConvFinQA dataset. This shows that improving the retrieval performance is of great significance for this framework model, and the graph structure has a good performance in information modeling of tables. At the same time, enhancing the diversity of the generation process also brings better performance for the model without reducing the quality of text generation.
[0112] The present invention is not limited to the embodiments described above. The above description of the specific embodiments is intended to describe and illustrate the technical solutions of the present invention. The above specific embodiments are merely illustrative and not restrictive. Without departing from the spirit of the present invention and the scope protected by the claims, those of ordinary skill in the art can make many specific transformations in various forms under the inspiration of the present invention, and these all fall within the protection scope of the present invention.
Claims
1. A financial question-answering numerical reasoning system based on deep learning method, characterized in that: The deep reasoning system includes a retriever and a generator; the retriever is composed of a first encoder, a graph attention module and a semantic aggregation module; the generator is composed of a second encoder, a semantic optimization module and a decoder; wherein: The graph attention module establishes numerical associations among questions, texts, and table-related information through heterogeneous graphs; wherein: The semantic aggregation module is trained according to numerical association to obtain fact-based feature representation; The generator generates multiple groups of numerical semantic expressions by optimizing the representation of fact features through random functions.
2. A financial question-answering numerical reasoning system based on deep learning method according to claim 1, characterized in that: The graph attention module constructs the numerical association process of question, text and table related information through heterogeneous graphs according to the input question, including:
101. Create node representations of tables, documents, and issues: According to each cell in the table, create a table node; each table node represents a data point in the table; According to each sentence in the document, a document sentence node is created; each document sentence node represents a semantic unit in the document; A question node is created for each financial question that needs to be answered; the question node represents a query or question raised by a user; 201. Create edge representations of tables, documents, and questions: Create internal relationship edges in the table based on the row and column relationships between table nodes; Create document internal relationship edges based on the semantic relationship between document sentence nodes; Create association edges between tables and documents based on the relationship between table data and document content between table nodes and document sentence nodes; Create association edges between question nodes and table nodes or document sentence nodes; 301. Construct a heterogeneous graph representing two-dimensional structural information of the table according to the node representation of step 101 and the edge representation of step 201; 401. The graph attention module aggregates nodes according to the type of edges; 501. The semantic aggregation module calculates and outputs the fact feature representation based on the aggregation nodes according to the similarity.
3. A financial question-answering numerical reasoning system based on deep learning method according to claim 2, characterized in that: The semantic aggregation module is trained according to the numerical association to obtain the fact-based feature representation process; including: Extract the numerical association information of each node in the heterogeneous graph and perform weighted aggregation; The information exchange across nodes is performed for each weighted node according to the following formula: have been uv =LeakyReLU(a T [have been u ||e v ]) in, represents the feature representation of node u in the lth layer, N(u) is the neighborhood of node u, e uv is the interactive feature representation of node u and node v, It is the attention correlation coefficient between node u and node v when the edge type is k.
4. A financial question-answering numerical reasoning system based on deep learning method according to claim 1; characterized in that: The generator generates multiple sets of numerical semantic expressions by optimizing the fact feature representation through random functions, including: The second encoder performs latent variable representation based on the fact feature representation according to the following formula: h=Encoder(x) The optimized semantic module optimizes the fact-based feature representation according to the following random parameters: z=g(h)+ε The decoder calculates and outputs multiple groups of numerical semantic representations for the optimized numerical semantics according to the following formula: and t =Decoder(and t-1 ,z) Among them, z is the input of the decoder, which is sampled from the probability distribution, ε is a trainable parameter, g() is a random function, x is the input of the encoder, h is the output of the encoder, and y t is the output at time t.
5. A financial question-answering numerical system based on deep learning method according to claim 1; characterized in that: The generator trains the random function through the following steps, including: Obtain the g() random function through variational inference: logp(y|h)≥E q [log p(y|z)]-KL(q(z|h,y)||p(z|h)) Among them, q(z|h,y) is the approximate result; The following formula is used to simplify the semantics of checking multiple sets of numerical values: Among them, μ and σ 2 are the mean and covariance of the network.
6. A method for numerical reasoning in financial question answering based on deep learning method, characterized in that: The steps include executing the deep reasoning model of claims 1-5 to realize the retrieval of documents and table information based on the raised financial questions and the generation of corresponding expressions.