Financial document automatic analysis and question-answering system based on large model
Through multi-grained segmentation and semantic embedding coding technology, the problems of insufficient semantic correlation across paragraphs and model illusions in the financial document question and answer system are solved, and efficient and accurate intelligent analysis of financial documents is achieved.
Patent Information
- Application Number
- CN202510433555.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-08
- Publication Date
- 2025-07-29
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
When dealing with complex financial problems, the existing financial document Q&A system has problems such as insufficient semantic correlation across paragraphs, serious model illusions and poor domain adaptability, resulting in inaccurate information extraction and lack of credible reasoning paths.
Multi-grained segmentation and semantic embedding coding technology are used to map financial documents to vector space, dynamically match user questions and candidate document fragments through multi-layer search mechanisms, cross-paragraph semantic association modeling is carried out, and a traceable inference path is constructed based on the thinking chain inference generation model.
It improves the analysis depth and accuracy of complex financial problems, provides intelligent analysis services that combine accuracy and credibility, and reduces the risk of large-scale models' hallucinations.
Smart Images

Figure CN120387437A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of data intelligent analysis, and more specifically, to a large model-based automated financial document parsing and question answering system. Background Art
[0002] In the current financial market, documents such as financial bond prospectuses, as the core carriers of information disclosure, carry a large amount of structured and unstructured data such as company profiles, operating conditions, financial indicators, and risk factors. Such documents are not only important bases for investors' decisions but also key objects for regulatory compliance reviews. However, financial documents generally have characteristics such as large text scale, dense professional terms, and complex logical levels. For example, a single bond prospectus often contains tens of thousands of words of cross-chapter related content, involving multi-modal information such as industry analysis and financial models. Traditional manual parsing methods rely on experts to read and annotate paragraph by paragraph, which is not only time-consuming and laborious but also difficult to ensure the accuracy and consistency of information extraction.
[0003] With the development of large language model technology, more and more document question answering systems have started to adopt such technology to improve the efficiency and accuracy of information processing. Existing technologies such as the knowledge graph enhancement solutions proposed in patents such as CN116775847B and CN117056493B construct structured knowledge bases through entity recognition and relationship extraction, showing high question answering accuracy in the general field. However, in the face of the unique technical challenges of financial documents, existing systems still have significant defects: First, most systems adopt a paragraph-level semantic similarity matching strategy, which can extract local text fragments but lacks global modeling of cross-paragraph semantic associations, making it difficult to accurately respond to problems that require comprehensive multi-chapter information such as company business analysis and risk transmission paths; Second, the inherent "hallucination" phenomenon of large models is amplified in the financial scenario. When dealing with ambiguous terms or contradictory statements, the system may generate conclusions without factual basis, and the existing black-box architecture cannot provide a credible reasoning path for tracing; Third, in the existing retrieval-augmented generation (RAG) framework, the document encoding and question parsing modules are often optimized independently, without fully considering the domain adaptability under the financial term system, resulting in insufficient matching of professional concepts. For example, the general retrieval disclosed in CN117033608B is prone to semantic confusion with ordinary bond terms when dealing with composite concepts such as "perpetual bond interest deferral clauses".
[0004] Therefore, there is a need for a large model-based automated financial document parsing and question answering solution. Summary of the Invention
[0005] To solve the above technical problems, this application is proposed. Embodiments of this application provide a large model-based automated financial document parsing and question answering system.
[0006] According to one aspect of the present application, there is provided a large model-based automated financial document parsing and question-answering system, which includes:
[0007] A financial document acquisition module for acquiring financial documents uploaded by users;
[0008] A financial document embedding and encoding module for performing semantic embedding and encoding on the financial document at the segment granularity to obtain a sequence distribution of semantic embedding and encoding features describing the financial document at the segment granularity;
[0009] A user question acquisition module for acquiring user questions;
[0010] A user question embedding and encoding module for performing semantic embedding and encoding on the user question to obtain user question semantic embedding and encoding features;
[0011] A retrieval and generation module for performing retrieval and generation on the sequence distribution of semantic embedding and encoding features describing the financial document at the segment granularity and the user question semantic embedding and encoding features to obtain an answer reply text, wherein the retrieval and generation module includes: a retrieval encoding unit for performing spatial constraint context semantic encoding based on multi-layer retrieval on the user question semantic embedding and encoding features and the sequence distribution of semantic embedding and encoding features describing the financial document at the segment granularity to obtain candidate answer context semantic encoding features; a generation unit for obtaining the answer reply text based on the candidate answer context semantic encoding features.
[0012] Compared with the prior art, the large model-based automated financial document parsing and question-answering system provided by the present application first performs multi-granularity segmentation and semantic embedding and encoding on the financial document to map the document data into a vector space, then adopts a multi-layer retrieval mechanism to dynamically match the user question with candidate document segments in the feature space to accurately locate relevant contexts, then conducts cross-paragraph semantic association modeling on the candidate segments by introducing dynamic propagation of feature space constraints to break through the information fragmentation limitation of traditional paragraph matching, realize global reasoning of complex questions, and finally combine the thought chain reasoning generation model to synchronously construct a traceable reasoning path when outputting answers, significantly reducing the large model hallucination risk. In this way, the parsing depth of complex financial questions can be improved, and thus an intelligent analysis service with both accuracy and credibility can be provided to users. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained according to these drawings without creative efforts. In the drawings:
[0014] Figure 1 It is a system block diagram of a large model-based financial document automated parsing and question answering system according to an embodiment of the present application.
[0015] Figure 2 It is a block diagram of a financial document embedding and encoding module in a large model-based financial document automated parsing and question answering system according to an embodiment of the present application.
[0016] Figure 3 It is a block diagram of a retrieval and generation module in a large model-based financial document automated parsing and question answering system according to an embodiment of the present application.
[0017] Figure 4 It is a block diagram of a retrieval and encoding unit in a large model-based financial document automated parsing and question answering system according to an embodiment of the present application. Detailed implementation manners
[0018] Next, exemplary embodiments according to the present application will be described in detail with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments of the present application. It should be understood that the present application is not limited by the exemplary embodiments described herein.
[0019] In the current financial market, documents such as financial bond prospectuses, as the core carriers of information disclosure, integrate multi-dimensional information such as enterprise operation data, financial indicators, and risk elements. They are not only the basis for investment decisions but also the key points for regulatory compliance reviews. Such documents have characteristics such as large text volume, strong professionalism, and complex logical structures. A single prospectus often contains tens of thousands of words of cross-chapter related content, covering multi-modal information such as industry analysis and financial models. Traditional manual parsing methods have limitations such as low efficiency and poor consistency, and it is difficult to meet the actual needs.
[0020] Although the document question answering system driven by large language model technology has improved the information processing ability through the knowledge graph enhancement solutions proposed in CN116775847B and CN117033608B, it still faces three major bottlenecks when dealing with the special needs of the financial field: First, existing systems mostly adopt paragraph-level semantic matching strategies and lack the ability to analyze cross-chapter information associations, resulting in insufficient responses to complex problems involving multi-dimensional cross-reasoning; Second, the phenomenon of model "hallucination" is likely to produce conclusions with insufficient factual basis when parsing financial ambiguous terms, and the existing black box architecture lacks a credible reasoning path tracing mechanism; Third, the general retrieval augmented generation (RAG) framework is not fully adapted to the financial professional term system, and it is easy to cause semantic confusion when dealing with compound concepts such as "perpetual bond interest deferral clauses", reflecting the domain adaptability defects of the document encoding and question parsing modules.
[0021] Based on this, the technical concept of this application is to effectively address the three core challenges in intelligent financial document processing by constructing an end-to-end parsing and question-answering process. Specifically, first, the financial document is segmented at multiple granularities, and the document paragraphs are mapped to the vector space through semantic embedding encoding; then, a multi-layer retrieval mechanism is adopted to dynamically match the user's question with the candidate document fragments in the feature space to accurately locate the relevant context; by introducing the dynamic propagation of feature space constraints, semantic association modeling across paragraphs is performed on the candidate fragments, breaking through the information fragmentation limitation of traditional paragraph matching and realizing global reasoning for complex questions; finally, combined with the chain-of-thought reasoning generation model, a traceable reasoning path is constructed synchronously when outputting the answer, significantly reducing the risk of large model hallucinations. This technical solution significantly improves the parsing depth of complex financial questions while maintaining a high retrieval recall rate. At the same time, the recognition accuracy of professional concepts such as perpetual bonds and cross-default clauses is enhanced through a domain-adapted semantic encoding strategy, providing intelligent analysis services with both accuracy and credibility for investors and regulators.
[0022] Figure 1 FIG. is a system block diagram of a large model-based automated financial document parsing and question-answering system according to an embodiment of the present application. As Figure 1 shown, in the large model-based automated financial document parsing and question-answering system 100, it includes: a financial document acquisition module 110 for acquiring the financial document uploaded by the user; a financial document embedding and encoding module 120 for performing semantic embedding encoding on the financial document based on segment granularity to obtain a sequence distribution of semantic embedding encoding features describing the financial document at the segment granularity; a user question acquisition module 130 for acquiring the user's question; a user question embedding and encoding module 140 for performing semantic embedding encoding on the user's question to obtain user question semantic embedding encoding features; a retrieval and generation module 150 for performing retrieval and generation on the sequence distribution of semantic embedding encoding features describing the financial document at the segment granularity and the user question semantic embedding encoding features to obtain an answer reply text.
[0023] In the embodiment of the present application, the financial document acquisition module 110 is used to acquire financial documents uploaded by users. It should be understood that financial documents such as financial bond prospectuses contain a large amount of key information such as company profiles, operating conditions, financial indicators, and risk factors. Only by acquiring the financial documents uploaded by users can the data therein be extracted and analyzed, so as to provide support for subsequent question answering and business analysis. In particular, the financial documents uploaded by users are in PDF format, and their content includes structured text (such as tables, paragraph headings) and unstructured text (such as natural language descriptions), and may also include complex layouts such as charts, headers, and footers. Although the PDF format is convenient for human reading, its binary encoding and paging characteristics make it difficult for computers to directly parse semantic content. After acquiring the financial documents, it is necessary to use PDF parsing technology to convert their content into processable text data, that is, to perform a preliminary parsing of the PDF file using PyMuPDF to extract page text.
[0024] In the embodiment of the present application, the financial document embedding encoding module 120 is used to perform semantic embedding encoding on the financial document based on segment granularity to obtain a sequence distribution of semantic embedding encoding features of the financial document segment granularity description. Specifically, Figure 2 It is a block diagram of the financial document embedding encoding module in the financial document automatic parsing and question answering system based on a large model according to the embodiment of the present application. As Figure 2 shown, the financial document embedding encoding module 120 includes: a financial text segment preprocessing unit 121, which is used to perform segmentation processing on the financial document to obtain a sequence distribution of the financial document segment granularity description; a financial document segment granularity semantic embedding unit 122, which is used to perform semantic embedding encoding on each financial document segment granularity description in the sequence distribution of the financial document segment granularity description to obtain a sequence distribution of semantic embedding encoding vectors of the financial document segment granularity description as the sequence distribution of semantic embedding encoding features of the financial document segment granularity description.
[0025] In the embodiment of the present application, the financial text preprocessing unit 121 is configured to segment the financial document to obtain a sequence distribution described at the granularity of financial document segments. Accordingly, considering that professional documents such as financial documents usually adopt a multi-level nested structure, for example, there is a logical progression relationship between chapters such as "issuance terms - guarantee situation - risk factors", and heterogeneous content such as data tables and financial formula derivations may be included within a single chapter. For example, when analyzing the "completeness of the issuer's related party transaction disclosure", the relevant descriptions may be scattered in sub-paragraphs of multiple chapters such as "corporate governance", "major events", and "notes to the financial statements". If the entire document is directly processed, it is not only difficult but also difficult to capture the internal relevance of semantic units across paragraphs. Based on this, the present application segments the financial document to segment the entire document into text segments, obtaining a sequence distribution described at the granularity of financial document segments. In this way, through the semantic coherent paragraph division, the semantic embedding encoding can focus on the complete business logic unit (such as the description of all rights and obligations of a single guarantee clause), so as to better capture the logical connection relationship between clauses. To ensure the integrity of the text content, a text block splitting and splicing strategy is introduced, and the paragraph and page information of the PDF are represented in the form of a sequential matrix T, where:
[0026] T = {T i,j | i ∈ Pages, j ∈ Paragraphs}
[0027] Each T i,j represents the text block of the j-th paragraph on the i-th page. By merging adjacent blocks within the matrix, a complete financial document text string, i.e., a sequence distribution described at the granularity of financial document segments, is generated, which lays a foundation for subsequent information extraction and analysis.
[0028] In the embodiment of the present application, the financial document segment granularity semantic embedding unit 122 is used to perform semantic embedding encoding on each financial document segment granularity description in the sequence distribution of the financial document segment granularity description to obtain a sequence distribution of financial document segment granularity description semantic embedding encoding vectors as the sequence distribution of the financial document segment granularity description semantic embedding encoding features. Specifically, in the embodiment of the present application, the financial document segment granularity semantic embedding unit is used to: use a semantic embedding encoder based on sentence-BERT to perform semantic embedding encoding on each financial document segment granularity description in the sequence distribution of the financial document segment granularity description to obtain the sequence distribution of the financial document segment granularity description semantic embedding encoding vectors. Correspondingly, considering that the highly specialized term system in financial documents (such as "subordinated debt", "cross-default trigger conditions") often has strict domain definitions, and the semantics of the same vocabulary may change dynamically according to the context in different clauses. For example, "liquidity support" specifically refers to the capital injection commitment of a third-party institution in the guarantee clause, while in the financial analysis section, it may refer to the enterprise's cash flow management ability. This phenomenon of polysemy is likely to cause conceptual confusion if only relying on traditional word frequency statistics or general semantic models. At the same time, the implicit logic of financial clauses is often constructed through long-distance context associations. For example, the conditional statement "if the coverage requirement is not met for two consecutive quarters" in a certain risk disclosure clause needs to be associated with the subsequent "default event handling process" section to fully understand its semantics, which poses extremely high requirements for the context capture ability of semantic representation. Therefore, in the technical solution of the present application, semantic embedding encoding is performed on each financial document segment granularity description in the sequence distribution of the financial document segment granularity description to obtain a sequence distribution of financial document segment granularity description semantic embedding encoding vectors. In particular, in a specific example of the present application, semantic embedding encoding can be performed on each financial document segment granularity description in the sequence distribution of the financial document segment granularity description by using a semantic embedding encoder based on sentence-BERT to obtain the sequence distribution of the financial document segment granularity description semantic embedding encoding vectors. That is, the semantic embedding encoder optimized based on the sentence-BERT architecture, the core design of which is to achieve accurate semantic mapping of the financial segment granularity description through a domain-adapted dual-tower encoding strategy. Specifically, the encoder introduces financial domain corpus (such as bond clauses, regulatory documents) in the pre-training stage for continuous fine-tuning, so that the model deeply internalizes the industry term system and expression paradigm. When processing the segmented financial document segment granularity description, the encoder not only extracts the local semantic features of the word sequence within the segment, but also dynamically captures the cross-sentence logical relationship through the attention mechanism (such as the conditional causal chain of "if the issuer fails to meet... then trigger..."), and finally generates a high-dimensional vector representation with both local semantic integrity and global relevance.
[0029] In the embodiment of the present application, the user question acquisition module 130 is used to acquire user questions. It should be understood that the content of user questions usually focuses on various key information in financial documents. When facing a financial bond prospectus, questions may be asked about company-level information, such as the company's financial situation, including specific data such as revenue, profit, and assets and liabilities, and the overall health. Attention will also be paid to the company's business opportunities, such as the issuance of different bond products and market share improvement strategies. Risk factors are also common question directions, such as what potential risks the company faces, the transmission path and impact degree of the risks. In short, user questions are the target instructions for the operation of the Q&A system. Without questions, the system cannot clarify the content to be parsed (such as "whether to analyze financial data or risk terms"), and subsequent data processing (retrieval, reasoning, generation) will lose direction.
[0030] In the embodiment of the present application, the user question embedding and encoding module 140 is used to perform semantic embedding and encoding on the user question to obtain user question semantic embedding and encoding features. Specifically, in the embodiment of the present application, the user question embedding and encoding module is used to: use a BERT-based semantic embedding encoder to perform semantic embedding and encoding on the user question to obtain a user question semantic embedding and encoding vector as the user question semantic embedding and encoding feature. Correspondingly, considering that questions in the financial field often involve professional terms and complex semantic relationships. For example, "the company's future business opportunities" involves an understanding of various aspects such as the company's strategy, market trends, and industry development; "financial risk factors" includes considerations of factors such as financial statements, macroeconomic environment, and financial policies. Therefore, in order to capture these professional and complex semantic information and convert it into a vector form that can be processed by a computer, in the technical solution of the present application, semantic embedding and encoding are performed on the user question to obtain user question semantic embedding and encoding features. In particular, in a specific example of the present application, a BERT-based semantic embedding encoder is used to perform semantic embedding and encoding on the user question to obtain a user question semantic embedding and encoding vector as the user question semantic embedding and encoding feature. It can be understood that the BERT model has learned rich language knowledge and semantic information during the pre-training process. It adopts a bidirectional Transformer architecture, can consider context information at the same time, and has strong semantic understanding ability for text. In the financial field, user questions may contain complex professional terms and logical relationships, and the BERT-based semantic embedding encoder can better capture this information, accurately understand the semantics of the questions, and convert them into appropriate vector representations. For example, for a complex question such as "to what extent will the company's future business opportunities be affected by macroeconomic policies", BERT can comprehensively analyze each word and the relationship between them, and extract accurate semantic features.
[0031] In the embodiment of the present application, the retrieval and generation module 150 is configured to perform retrieval and generation on the sequence distribution of the semantic embedding encoding features of the financial document segment granularity description and the semantic embedding encoding features of the user question to obtain an answer reply text. Specifically, Figure 3 It is a block diagram of the retrieval and generation module in the financial document automated parsing and question answering system based on a large model according to the embodiment of the present application. As Figure 3 shown, the retrieval and generation module 150 includes: a retrieval encoding unit 151, configured to perform spatial constraint context semantic encoding based on multi-layer retrieval on the sequence distribution of the semantic embedding encoding features of the user question and the semantic embedding encoding features of the financial document segment granularity description to obtain candidate answer context semantic encoding features; a generation unit 152, configured to obtain the answer reply text based on the candidate answer context semantic encoding features.
[0032] In the embodiment of the present application, the retrieval encoding unit 151 is configured to perform spatial constraint context semantic encoding based on multi-layer retrieval on the sequence distribution of the semantic embedding encoding features of the user question and the semantic embedding encoding features of the financial document segment granularity description to obtain candidate answer context semantic encoding features. Specifically, Figure 4 It is a block diagram of the retrieval encoding unit in the financial document automated parsing and question answering system based on a large model according to the embodiment of the present application. As Figure 4 shown, the retrieval encoding unit 151 includes: a multi-layer semantic retrieval and screening sub-unit 1511, configured to input the sequence distribution of the semantic embedding encoding vectors of the user question and the semantic embedding encoding vectors of the financial document segment granularity description into a multi-layer retrieval module to obtain the sequence distribution of the semantic embedding encoding vectors of the candidate answer financial document segment granularity description; a candidate answer context global encoding sub-unit 1512, configured to perform context semantic global encoding based on spatial constraints on the sequence distribution of the semantic embedding encoding vectors of the candidate answer financial document segment granularity description to obtain a candidate answer context semantic encoding vector as the candidate answer context semantic encoding feature.
[0033] In the embodiment of the present application, the multi-layer semantic retrieval and screening sub-unit 1511 is configured to input the sequence distributions of the user question semantic embedding encoding vector and the financial document segment granularity description semantic embedding encoding vector into a multi-layer retrieval module to obtain the sequence distribution of the candidate answer financial document segment granularity description semantic embedding encoding vector. It should be understood that although the vector sequence of the financial document segment granularity description already contains semantic information, there are significant defects in directly performing global similarity matching: First, financial documents usually contain tens of thousands of words of cross-chapter content. Directly calculating the similarity between the user question and all paragraph vectors in full volume will lead to an explosion in computational complexity. Second, user questions (such as "the impact of cross-default triggering on debt-servicing ability") often involve the associated reasoning of multiple types of information, and a single similarity metric is difficult to distinguish the relevance weights of different semantic dimensions such as legal terms, financial data, and risk warnings. For example, when retrieving "perpetual bond interest deferral clauses", it is necessary to simultaneously match the clause definitions in the legal chapter, the accounting treatment descriptions in the financial chapter, and the trigger condition analyses in the risk section. Traditional single-layer retrieval is prone to missing key paragraphs or confusing the relevance ranking. Therefore, in the technical solution of the present application, the sequence distributions of the user question semantic embedding encoding vector and the financial document segment granularity description semantic embedding encoding vector are input into a multi-layer retrieval module to obtain the sequence distribution of the candidate answer financial document segment granularity description semantic embedding encoding vector. In particular, the multi-layer retrieval module combines multiple retrieval strategies, such as vector retrieval, BM25 algorithm, and semantic similarity calculation. The user question semantic embedding encoding vector and the financial document segment granularity description semantic embedding encoding vector contain rich semantic information, which can be combined with these retrieval strategies to give full play to the advantages of various strategies, so as to more comprehensively and accurately find the segment granularity description content related to the user question.
[0034] Specifically, first, a coarse-grained retrieval operation is performed: Calculate the BM25 similarity between each financial document segment granularity description semantic embedding encoding vector in the sequence distribution of the financial document segment granularity description semantic embedding encoding vector and the user question semantic embedding encoding vector (denoted as ). The BM25 formula is as follows:
[0035]
[0036] where D is the financial document segment granularity description semantic embedding encoding vector, f(q i , D) is the occurrence frequency of each user question semantic embedding encoding feature in the financial document segment granularity description semantic embedding encoding vector, IDF(q i)The inverse document frequency of the semantic embedding encoding features of the user's question in the semantic embedding encoding vector at the financial document segment granularity, |D| is the length of the semantic embedding encoding vector at the financial document segment granularity, avgdl is the average length of the semantic embedding encoding vector at the financial document segment granularity, k1 and b are adjustment parameters, and finally, several semantic embedding encoding vectors at the financial document segment granularity with the highest BM25 scores are selected for fine-grained ranking.
[0037] Then perform the fine-grained ranking operation: For the input semantic embedding encoding vector of the user's question and the semantic embedding encoding vectors of the financial document segments retrieved roughly, first use the encoder model of the DistilBERT two-tower structure to process them again to obtain the semantic feature vector v of the user's question Q and the semantic feature vector v of the financial document segment granularity D . The embedding representation of the dual-encoder model is calculated as follows:
[0038] v Q = f encoder (Q), v D = f encoder (D)
[0039] where f encoder is the pre-trained encoder. Then calculate the cosine similarity between the semantic feature vector of the user's question and the semantic feature vector of the financial document segment granularity to determine their semantic relevance. The calculation formula for semantic similarity is:
[0040]
[0041] The semantic embedding encoding vectors of the financial document segments can be freely selected according to the similarity scores and arranged in a certain order before the semantic feature vectors of the financial document segments as candidate answer semantic embedding encoding vectors of the financial document segments to obtain the sequence distribution of the final candidate answer semantic embedding encoding vectors of the financial document segments.
[0042] In the embodiment of the present application, the candidate answer context global encoding subunit 1512 is configured to perform context semantic global encoding based on spatial constraints on the sequence distribution of the candidate answer financial document segment granularity description semantic embedding encoding vectors to obtain candidate answer context semantic encoding vectors as the candidate answer context semantic encoding features. Specifically, in the embodiment of the present application, the candidate answer context global encoding subunit includes: an end axial information propagation constraint factor calculation secondary subunit, configured to calculate the end information propagation constraint factor and the axial information propagation constraint factor of each candidate answer financial document segment granularity description semantic embedding encoding vector in the sequence distribution of the candidate answer financial document segment granularity description semantic embedding encoding vectors to obtain a candidate answer financial document semantic end information propagation constraint factor and a candidate answer financial document semantic axial information propagation constraint factor; a semantic information propagation constraint factor calculation secondary subunit, configured to calculate the candidate answer financial document semantic information propagation constraint factor of each candidate answer financial document segment granularity description semantic embedding encoding vector based on the candidate answer financial document semantic end information propagation constraint factor and the candidate answer financial document semantic axial information propagation constraint factor of each candidate answer financial document segment granularity description semantic embedding encoding vector in the sequence distribution of the candidate answer financial document segment granularity description semantic embedding encoding vectors; a candidate answer context semantic feature generation secondary subunit, configured to perform semantic information dynamic constraint propagation on the sequence distribution of the candidate answer financial document segment granularity description semantic embedding encoding vectors based on the candidate answer financial document semantic information propagation constraint factor of each candidate answer financial document segment granularity description semantic embedding encoding vector to obtain the candidate answer context semantic encoding vectors.
[0043] It should be understood that key information in financial documents (such as risk conduction paths and legal clause linkage mechanisms) often is distributed in different chapters in a discontinuous and multi-level manner. For example, the triggering conditions of the "cross-default clause" of a certain bond may be scattered in the main body of the contract, supplementary agreements, and financial appendices, and its legal consequences need to be comprehensively analyzed in combination with related paragraphs such as default disposal processes and collateral liquidation clauses. Traditional encoding methods based on recurrent neural networks or self-attention mechanisms can capture local context dependencies, but in the face of the long-range dependencies (such as clause associations separated by dozens of pages) and structural hierarchies (such as the nesting of main clauses and exception clauses) unique to financial documents, semantic focus drift or key logical breakpoints are likely to occur. For example, when analyzing the "perpetual bond interest deferral triggering conditions", if only relying on local associations in adjacent paragraphs, special agreements in the attachments of the issuance agreement may be missed, resulting in a one-sided conclusion. Therefore, in the technical solution of the present application, context semantic global encoding based on spatial constraints is performed on the sequence distribution of the candidate answer financial document segment granularity description semantic embedding encoding vectors to obtain candidate answer context semantic encoding vectors.
[0044] Specifically, first, extract the end-constrained anchoring encoding vector of the candidate answer paragraph to capture the end semantic features of the current retrieval result (such as the semantic expression of the latest financial data), which serves as the stability boundary for information transmission; at the same time, generate the axial-constrained anchoring encoding vector through clustering analysis to characterize the global distribution pattern of the candidate paragraph set (such as the common features of risk factor paragraphs). On this basis, by dynamically calculating the end and axial information propagation constraint factors, the association strength between each paragraph vector and the local semantic features and global structural features is quantified respectively - for example, for paragraphs related to "accounting treatment of perpetual bonds", the end factor strengthens its association with the latest accounting standards provisions, and the axial factor enhances its clustering characteristics in the classification of equity instruments. Finally, through the horizontal compactification covariance strategy, the two types of constraint factors are fused into dynamic weights to guide the directional propagation of semantic information between paragraphs, ensuring the semantic integrity of the key logical chain (such as "trigger conditions - warning indicators - disposal measures") during the information transmission process. In this way, after obtaining the candidate answer context semantic encoding vector, it can provide richer and more accurate semantic features for subsequent answer generation, reasoning, and matching with the user's question. These features can better reflect the information related to the question in the financial document and improve the ability of the model to generate accurate answers.
[0045] Specifically, in the embodiment of the present application, the secondary subunit for calculating the end-axial information propagation constraint factor is used to: extract the last candidate answer financial document segment granularity description semantic embedding encoding vector from the sequence distribution of the candidate answer financial document segment granularity description semantic embedding encoding vector as the end-constrained anchoring encoding vector of the candidate answer financial document semantic propagation space. This process can be expressed by the formula:
[0046] O = {v1, v2,..., v i ,..., v t}
[0047] v tail = v t
[0048] where O is the sequence distribution of the candidate answer financial document segment granularity description semantic embedding encoding vector, v1, v2, v i and v i are the 1st, 2nd, i-th, and t-th candidate answer financial document segment granularity description semantic embedding encoding vectors in the sequence distribution of the candidate answer financial document segment granularity description semantic embedding encoding vector respectively, and v tail is the end-constrained anchoring encoding vector of the candidate answer financial document semantic propagation space;
[0049] Perform clustering analysis on the sequence distribution of the granularity description semantic embedding coding vectors of the candidate answer financial document segments to obtain the candidate answer financial document semantic propagation space axial constraint anchoring coding vectors. This process can be expressed by the formula:
[0050]
[0051] where Cluster is the clustering analysis operation, max(v i ) and min(v i ) are respectively the maximum and minimum values of v i , η is the adjustment parameter, a i is the i-th candidate answer financial document semantic constraint anchoring value in the sequence distribution of the candidate answer financial document semantic constraint anchoring values, softmax is the normalization function, e i is the i-th candidate answer financial document semantic constraint anchoring weight value in the sequence distribution of the candidate answer financial document semantic constraint anchoring weight values, t is the number of vectors in O, and v axis is the candidate answer financial document semantic propagation space axial constraint anchoring coding vector;
[0052] Calculate the candidate answer financial document semantic end information propagation constraint factor of each candidate answer financial document segment granularity description semantic embedding coding vector in the sequence distribution of the candidate answer financial document segment granularity description semantic embedding coding vectors relative to the candidate answer financial document semantic propagation space end constraint anchoring coding vector. This process can be expressed by the formula:
[0053]
[0054] where f(v i , v tail ) is to calculate the end information propagation constraint factor between v i and v tail , v ij is the eigenvalue at the j-th position of v i , is the eigenvalue at the j-th position of v tail , log2 is the logarithmic function value with base 2, n is the number of eigenvalues in v i , exp is the exponential function value with base e (the natural constant), α i is the candidate answer financial document semantic end information propagation constraint factor corresponding to v i ;
[0055] Calculating the candidate answer financial document semantic axial information propagation constraint factor of each candidate answer financial document segment granularity description semantic embedding coding vector in the sequence distribution of the candidate answer financial document segment granularity description semantic embedding coding vector relative to the candidate answer financial document semantic propagation space axial constraint anchoring coding vector, this process can be expressed by the formula:
[0056]
[0057] where f(v i , v axis ) is to calculate the axial information propagation constraint factor between v i and v axis , ||·|| 2 is to calculate the square of the Euclidean norm of the vector, arccosh is the inverse hyperbolic cosine function, and β i is the candidate answer financial document semantic axial information propagation constraint factor corresponding to v i .
[0058] It should be understood that the logical chain of key information in financial documents (such as "default trigger conditions - early warning indicators - disposal measures") usually has a clear time sequence or hierarchical end point (such as the latest financial data or the final interpretation of legal terms). Due to the lack of explicit constraints on the sequence end point, traditional coding methods are prone to losing key semantic focuses in long-distance dependence scenarios (such as missing supplementary agreement content across pages). By using the last candidate answer financial document segment granularity description semantic embedding coding vector in the sequence distribution of the candidate answer financial document segment granularity description semantic embedding coding vector as the candidate answer financial document semantic propagation space end constraint anchoring coding vector (such as the embedding representation of the latest accounting standards terms), it can provide a stability boundary for information transmission. This processing is similar to setting the boundary conditions of a dynamic system, ensuring that the semantic propagation direction always points to the logical end point (such as the final judgment basis of "interest deferral trigger conditions"), and avoiding semantic drift caused by local noise or paragraph jumps. For example, when analyzing the terms of perpetual bonds, the candidate answer financial document semantic propagation space end constraint anchoring coding vector can strengthen the association with the latest accounting treatment rules, ensuring that the model always converges to the final terms during reasoning.
[0059] Accordingly, similar information in financial documents (such as risk factors, financial indicators) is often scattered across multiple chapters, but their semantic patterns share commonalities (such as the "cross-default clause" is often associated with paragraphs related to the realization of collateral, legal consequences, etc.). Although the traditional self-attention mechanism can capture local relationships, it is difficult to identify global distribution characteristics (such as the clustering characteristics of risk-related paragraphs). By performing clustering analysis on the sequence distribution of the semantic embedding coding vectors of the candidate answer financial document segment granularity descriptions, the global distribution pattern of the candidate paragraphs (such as the common vector of risk factor paragraphs) can be extracted, and the candidate answer financial document semantic propagation space axial constraint anchoring coding vector can be generated. This vector serves as the global semantic skeleton (such as the clustering center of "equity instrument classification"), constraining the information transmission to proceed along the main structural direction of the document, preventing the model from over-focusing on irrelevant paragraphs. For example, when dealing with the "accounting treatment of perpetual bonds", the axial constraint can enhance the clustering association between paragraphs and the equity instrument classification, ensuring that the model preferentially integrates similar clauses (such as special agreements in the appendix of the issuance agreement) rather than analyzing individual paragraphs in isolation.
[0060] It should be understood that the local dynamics of financial documents (such as accounting standard updates) require the model to distinguish the association strength between different paragraphs and the current latest paragraph. The candidate answer financial document semantic end information propagation constraint factor dynamically assigns local weights by quantifying the similarity between the semantic embedding coding vectors of each candidate answer financial document segment granularity description and the candidate answer financial document semantic propagation space end constraint anchoring coding vector (such as the semantic distance between the "accounting treatment" paragraph and the latest standard). For example, when analyzing the equity instrument classification, the end information propagation constraint factor will give paragraphs involving new standards higher weights, ensuring that information transmission preferentially associates with clauses with strong timeliness rather than historical version content. That is, through conditional probability-based weighting, the model focuses on local features strongly related to the current semantic endpoint, reducing the interference of outdated information.
[0061] Accordingly, cross-paragraph semantic consistency needs to match the document structure hierarchy (such as the equity instrument classification needs to conform to the global accounting framework). By calculating the projection relationship of each candidate answer financial document segment granularity description semantic embedding coding vector relative to the candidate answer financial document semantic propagation space axial constraint anchoring coding vector, the candidate answer financial document semantic axial information propagation constraint factor is generated, which can quantify the degree of fit between each paragraph and the overall distribution pattern. That is, this factor ensures that information transmission conforms to the internal structure of the document through regularization constraints. For example, mapping the default definition in the main contract and the exemption conditions in the supplementary agreement to the same legal logic axis, preventing incorrect parsing of the clause linkage mechanism caused by overfitting of local features.
[0062] Specifically, in the embodiments of the present application, the semantic information propagation constraint factor calculation secondary subunit is configured to: perform horizontal compact covariance on the candidate answer financial document semantic end information propagation constraint factor and the candidate answer financial document semantic axial information propagation constraint factor corresponding to the candidate answer financial document segment granularity description semantic embedding encoding vector by using a canonical space constraint matrix to obtain a candidate answer financial document semantic end information propagation compact covariance constraint factor and a candidate answer financial document semantic axial information propagation compact covariance constraint factor; based on the candidate answer financial document semantic end information propagation compact covariance constraint factor and the candidate answer financial document semantic axial information propagation compact covariance constraint factor, perform propagation constraint integration processing on the candidate answer financial document semantic end information propagation constraint factor and the candidate answer financial document semantic axial information propagation constraint factor to obtain an optimized candidate answer financial document semantic end information propagation constraint factor and an optimized candidate answer financial document semantic axial information propagation constraint factor; perform non-linear activation processing on the optimized candidate answer financial document semantic end information propagation constraint factor and the optimized candidate answer financial document semantic axial information propagation constraint factor to obtain the candidate answer financial document semantic information propagation constraint factor corresponding to the candidate answer financial document segment granularity description semantic embedding encoding vector. The above process can be expressed by the formula:
[0063]
[0064] y i = Sigmoid(ω1·α i ′ + ω2·β i ′)
[0065] where δ i and ε i are respectively the candidate answer financial document semantic end information propagation compact covariance constraint factor and the candidate answer financial document semantic axial information propagation compact covariance constraint factor after horizontal compact covariance of α i and β i , T represents the transpose operation, |·| is the absolute value operation, α i ′ and β i ′ are respectively the optimized candidate answer financial document semantic end information propagation constraint factor and the optimized candidate answer financial document semantic axial information propagation constraint factor corresponding to v i , ω1 and ω2 are respectively weighting parameters, Sigmoid is a weight mapping function, and y i is the candidate answer financial document semantic information propagation constraint factor corresponding to v i .
[0066] It should be understood that the complex semantics of financial documents need to satisfy both local precise alignment and global structural consistency simultaneously. By using dynamic weighted fusion to combine the candidate answer financial document semantic end information propagation constraint factor and the candidate answer financial document semantic axial information propagation constraint factor, the candidate answer financial document semantic information propagation constraint factor can be generated. For example, when analyzing cross-default clauses, it is possible to strengthen the local association between the triggering conditions of the master contract and the tail financial data (end information propagation constraint factor), and at the same time ensure its global risk classification consistency with the guarantee clause (axial information propagation constraint factor). That is to say, by fusing to generate the candidate answer financial document semantic information propagation constraint factor, the model can maintain the integrity of the logical chain while retaining the sensitivity to details, and avoid one-sided conclusions caused by single constraints.
[0067] Specifically, in the deep scenario of financial document semantic parsing, the interaction mechanism between end constraints and axial constraints needs to face the problem of dynamic balance of feature propagation. Specifically, when the candidate answer financial document segment granularity description semantic embedding coding vector performs semantic diffusion in the horizontal transmission field, it may generate uncontrolled semantic radiation due to the enhancement of local dynamic sensitivity (for example, when analyzing debt restructuring clauses, the model over-focuses on the detailed description of individual historical cases and deviates from the core framework of the latest accounting standards). This diffusion effect and the global structural stability forced by the candidate answer financial document semantic axial information propagation constraint factor through regularization are prone to form a tension, which is specifically manifested as the adversarial contradiction between the local dynamic adaptation and global pattern convergence of the feature flow. Therefore, based on the principle of horizontal compactification, the covariance integration strategy constructs a canonical space constraint matrix, and places the candidate answer financial document semantic end information propagation constraint factor (such as the correlation strength between the perpetual debt interest deferral clause and the latest financial data) and the candidate answer financial document semantic axial information propagation constraint factor (such as the clustering center directivity of the equity instrument classification) in a unified covariant tensor space, and uses the compactification operation of the matrix to realize the orthogonal projection reconstruction of the two types of constraint factors, so as to eliminate the interference of redundant semantic dimensions on the feature propagation path. Further, through the basis vector rotation transformation of the invertible space metric, a spinor finite-dimensional compact constraint is imposed on the reconstructed semantic representation, so that the timeliness characteristics of the end constraint (such as the update of the triggering conditions in the supplementary agreement) can be adjusted in a precessional manner along the main logical direction determined by the axial constraint (such as the hierarchical framework of legal clauses), and finally realize the coupling optimization of the dynamic propagation field and the static structure field. For example, when the model is processing multi-layer guarantee clauses, the horizontal compact covariance operation can accurately strip the noise features brought by regional regulatory differences, while the spinor constraint ensures that the core guarantee logic always converges along the main axis of the merger and acquisition risk distribution, which not only maintains the parsing sensitivity to special clauses, but also avoids clause linkage misjudgment caused by structural defocus.
[0068] Specifically, in the embodiments of the present application, the candidate answer context semantic feature generation secondary subunit is configured to: based on the candidate answer financial document semantic information propagation constraint factor of the semantic embedding encoding vector described by each candidate answer financial document segment granularity, perform semantic information dynamic constraint propagation on the sequence distribution of the semantic embedding encoding vector described by the candidate answer financial document segment granularity to obtain the candidate answer context semantic encoding vector, which can be expressed by the following formula:
[0069]
[0070] where z is the candidate answer context semantic encoding vector.
[0071] It should be understood that traditional encoders are prone to losing key logical chains because they treat the features of all paragraphs equally. By performing semantic information dynamic constraint propagation on the corresponding candidate answer financial document segment granularity description semantic embedding encoding vector through the candidate answer financial document semantic information propagation constraint factor (such as imposing a high weight on the "trigger condition" segment and suppressing irrelevant appendices), the generated context encoding features can explicitly retain cross-level semantic relationships. For example, when analyzing the disposal of collateral, paragraphs scattered in the main text (disposal process), appendix (liquidation rules), and attachment (time limit) can be spliced into a complete logical unit according to the weight, avoiding problems such as gradient decay of RNN or attention dilution of Transformer. That is, the finally obtained candidate answer context semantic encoding vector has both local sensitivity and global structure, and can provide a high-fidelity semantic representation for downstream tasks.
[0072] In an embodiment of the present application, the generating unit 152 is configured to obtain the answer reply text based on the semantic encoding features of the candidate answer context. Specifically, in an embodiment of the present application, the generating unit is configured to: input the semantic encoding vector of the candidate answer context into a chain-of-thought reasoning generation model to obtain the answer reply text. It should be understood that the semantic encoding vector of the candidate answer context contains the semantic information of the document paragraphs related to the question, but this information is in a scattered encoding form. The chain-of-thought reasoning generation model can integrate this semantic information, construct an inference path based on the knowledge graph, and gradually associate the candidate answer with the question step by step. Through this model, the potential logical relationships between these semantic information can be mined, so as to conduct in-depth reasoning and analysis on the question. For example, in a financial document, for the semantic encoding vector of a candidate answer regarding a company's financial status, the chain-of-thought reasoning generation model can associate it with relevant information such as the company's business model and market environment through the inference path for comprehensive reasoning. That is to say, through the processing of the semantic encoding vector of the candidate answer context by the chain-of-thought reasoning generation model, relevant semantic information can be integrated, and reasoning can be carried out along the inference path to generate an accurate, comprehensive, and logically coherent answer reply text. The model will comprehensively consider the results of multiple inference paths to ensure that the answer covers all aspects of the question and provides detailed and valuable information for users. For example, when answering a question about financial market trends, the answer will include analyses of multiple aspects such as macroeconomic factors, industry dynamics, and policy changes.
[0073] Specifically, first, a knowledge graph is constructed. The system first calculates the similarity of the sequence distribution of the semantic embedding encoding vectors of the financial document segment granularity descriptions obtained after processing to establish the connections between entities in different financial document segment granularity descriptions. Pairwise similarity calculations are performed on all entities {e1, e2,..., e n}, and the similarity matrix S is defined as:
[0074]
[0075] where S i,j > σ entity pairs are regarded as having semantic associations and are connected in the graph to generate a preliminary knowledge graph. Subsequently, the constructed knowledge graph is stored in Neo4j, and further multi-dimensional relationship optimization is performed using GraphSAGE. In the knowledge graph, nodes represent entities and edges represent relationships. The hidden representation h v of each node is defined as:
[0076]
[0077] where, is the neighbor node set of node v, and c v,uis the normalization coefficient, and σ is the activation function. Through graph optimization, the semantic associations of entities in the granular descriptions of different financial document segments can be enhanced. Then, using the optimized knowledge graph to generate inference paths, the entities closely related to the semantic encoding vectors of the question and candidate answer contexts are gradually associated using the chain-of-thought model. The system first selects relevant entity nodes and generates multiple inference paths along the relationship edges in the knowledge graph. Let the set of inference paths be where each path p i represents a logical association chain from the question node to the entity closely related to the semantic encoding vector of the candidate answer context. Then, using the pre-trained large language model Llama3 and combined with the chain-of-thought prompt, a multi-path inference answer to the question is generated. Let the user's question be Q and the generated inference path be p i , then the answer generated by the model is:
[0078] A i = Llama3 CoT (Q, p i )
[0079] The answer A for each path i represents the inference result obtained along path p i . To ensure the accuracy of the answer, a self-consistency verification method is adopted, that is, the answers of multiple inference paths are verified for consistency, and the answer with the highest frequency of occurrence is selected. If more than half of the answers are consistent on all inference paths, it is used as the final answer; otherwise, a weighted voting method is used, with the answer with the highest probability as the main one. The verification formula is as follows:
[0080]
[0081] where, δ(A i , A j ) = 1 when A i = A j , otherwise 0; Weight(p j ) is the path weight. After the generated answer is post-processed, the system displays the final answer reply text to the user through the Web interface. In the answer, the system provides the traceability link of the knowledge graph node cited, enabling the user to trace the specific source of the question answer and improving the transparency of the reply and user trust.
[0082] In summary, the large model-based financial document automated parsing and question answering system 100 according to the embodiments of the present application is elucidated. It first performs multi-granularity segmentation and semantic embedding encoding on financial documents to map document data into a vector space. Subsequently, a multi-layer retrieval mechanism is adopted to dynamically match user questions with candidate document fragments in the feature space to accurately locate relevant contexts. Then, by introducing dynamic propagation with feature space constraints, semantic association modeling across paragraphs is performed on candidate fragments to break through the information fragmentation limitation of traditional paragraph matching and achieve global reasoning for complex questions. Finally, combined with the thought chain reasoning generation model, a traceable reasoning path is constructed synchronously when outputting answers, significantly reducing the large model hallucination risk. In this way, the parsing depth of complex financial questions can be improved, and then an intelligent analysis service with both accuracy and credibility can be provided for users.
[0083] In other embodiments of the present application, a company profile and financial information extraction system based on financial prospectuses is provided. It includes a PDF parsing module, an information extraction module, and a result display module. It can parse financial prospectus PDF files to extract key contents such as company profiles, financial conditions, and risk factors, and generate a company profile report. Specifically, the PDF parsing module initially parses the PDF file using PyMuPDF to extract text, and introduces a text block segmentation and splicing strategy to represent the paragraph and page information of the PDF in the form of an ordered matrix. By merging adjacent blocks within the matrix, a complete prospectus text string is generated. In this embodiment, the PDF parsing module limits the uploaded file size to 50MB to ensure parsing speed and stability. The threshold for text block merging is set to 50 characters to ensure the integrity of long paragraphs. The information extraction module then loads the pre-trained weights in the financial field using the FinBERT model to identify category entities such as "company name", "financial data", "risk factors", etc., and sets the semantic similarity threshold τ = 0.85. When the similarity is higher than this threshold, semantic relationships are considered to exist between entities and are connected in the knowledge graph. During syntactic parsing, sentences with a syntactic dependency depth not exceeding 3 layers are selected to reduce noise. Finally, the result display module fills the extracted company profile information including company name, registered address, incorporation date, main business, and financial information into the company profile template for display. For financial data (such as Revenue, Profit, Assets, Liabilities), the system calculates the financial health score through the health score formula and classifies it as "high risk", "medium health", or "healthy". The company profile template is shown as follows:
[0084]
[0085] where, e i is the specifically extracted entity information.
[0086] In other embodiments of the present application, a user Q&A generation method is also provided, which can respond to specific questions of users based on the content in the financial prospectus. Specifically, in step 1, the user's question is first preprocessed to limit the length of the question characters within 300 characters, and no more than 10 keywords are extracted through the TF-IDF algorithm to ensure that the retrieved content focuses on relevant information. In step 2, a coarse-grained and fine-grained retrieval method is used to retrieve the keywords in the financial prospectus. In the coarse-grained retrieval, the parameters of the BM25 algorithm are configured as k1 = 1.5, b = 0.75, and the average document length is set to 200 words. It should be noted that although the BM25 algorithm formula is the same in the coarse-grained retrieval process, the meanings represented by each letter are different.
[0087]
[0088] Among them, D is the document, f(q i , D) is the frequency of occurrence of the keyword in the document, IDF(q i ) is the inverse document frequency of the keyword, |D| is the document length, avgdl is the average document length, and k1 and b are adjustment parameters. Through the relevant content after coarse-grained retrieval, the retrieval results are further optimized in the fine-grained ranking to ensure the acquisition of the most relevant text fragments. Specifically, for the user input question and the paragraphs retrieved by the coarse retrieval, the DistilBERT model is used to calculate the semantic similarity between the user question and the paragraphs, and the threshold is set to 0.8 to ensure the selection of the most relevant paragraphs. In step 3, based on the chain of thought reasoning model, multi-path answers are generated, and the number of chain of thought reasoning paths is set to 5 in the answer generation. Each path contains at most 3 reasoning steps, and a consistency verification algorithm is adopted. When the result consistency of multiple reasoning paths exceeds 60%, the answer is output; otherwise, the system votes according to the path weights to generate the final answer. In step 4, regular expression character matching is performed on the generated answer to ensure that the financial data is formatted into a currency format with two decimal places, and redundancy detection is performed on the content. The cosine similarity threshold is set to 0.75 to ensure that no duplicate information is displayed.
[0089] In other embodiments of the present application, there is also provided a method for updating and optimizing a knowledge graph. By regularly updating the knowledge graph, the adaptability to new financial data is enhanced, and the timeliness of financial terms and financial information is ensured. In step 1, the embedding vectors are updated every 7 days. The sentence-BERT is used to perform embedding calculations on the newly added financial terms and data. The number of newly added nodes per update does not exceed 100 to control the scale of the graph. In step 2, in graph optimization, the number of layers of the GraphSAGE model is set to 2 layers to avoid information loss caused by excessive convolution. The initial learning rate of the Adam optimizer is 0.001 and gradually decays after update, with a decay factor of 0.95 to ensure the stability of embedding updates. In step 3, LoRA fine-tuning is performed on the low-rank adapter of the Llama3 model. The rank of the adapter matrix is set to 16 to ensure that the model can quickly adapt to new financial data, and a dataset is generated for fine-tuning every two weeks based on 5% of the high-frequency problems feedback by users, thereby improving the answering ability.
Claims
1. A financial document automated parsing and question-answering system based on a large model, characterized in that, Including: A financial document acquisition module for acquiring financial documents uploaded by users; A financial document embedding and encoding module for performing semantic embedding encoding on the financial document at the segment granularity to obtain a sequence distribution of semantic embedding encoding features describing the financial document at the segment granularity; A user question acquisition module for acquiring user questions; A user question embedding and encoding module for performing semantic embedding encoding on the user question to obtain user question semantic embedding encoding features; A retrieval and generation module for performing retrieval and generation on the sequence distribution of semantic embedding encoding features describing the financial document at the segment granularity and the user question semantic embedding encoding features to obtain an answer reply text. Among them, the retrieval and generation module includes: a retrieval encoding unit for performing space-constrained context semantic encoding based on multi-layer retrieval on the user question semantic embedding encoding features and the sequence distribution of semantic embedding encoding features describing the financial document at the segment granularity to obtain candidate answer context semantic encoding features; a generation unit for obtaining the answer reply text based on the candidate answer context semantic encoding features.
2. The financial document automated parsing and question answering system based on a large model according to claim 1, characterized in that The financial document embedding and encoding module includes: A financial text preprocessing unit for segmenting the financial document to obtain a sequence distribution of descriptions of the financial document at the segment granularity; A financial document segment granularity semantic embedding unit for performing semantic embedding encoding on each description of the financial document at the segment granularity in the sequence distribution of descriptions of the financial document at the segment granularity to obtain a sequence distribution of semantic embedding encoding vectors of the descriptions of the financial document at the segment granularity as the sequence distribution of semantic embedding encoding features of the descriptions of the financial document at the segment granularity.
3. The financial document automated parsing and Q&A system based on a large model according to claim 2, wherein The financial document segment granularity semantic embedding unit is used to: use a semantic embedding encoder based on sentence-BERT to perform semantic embedding encoding on each description of the financial document at the segment granularity in the sequence distribution of descriptions of the financial document at the segment granularity to obtain the sequence distribution of semantic embedding encoding vectors of the descriptions of the financial document at the segment granularity.
4. The financial document automated parsing and question answering system based on a large model according to claim 3, wherein The user question embedding and encoding module is used to: use a semantic embedding encoder based on BERT to perform semantic embedding encoding on the user question to obtain a user question semantic embedding encoding vector as the user question semantic embedding encoding features.
5. The financial document automatic parsing and Q&A system based on a large model according to claim 4, wherein The retrieval encoding unit includes: A multi-layer semantic retrieval and screening sub-unit for inputting the user question semantic embedding encoding vector and the sequence distribution of semantic embedding encoding vectors of the descriptions of the financial document at the segment granularity into a multi-layer retrieval module to obtain a sequence distribution of semantic embedding encoding vectors of candidate answer financial document segments at the granularity; A candidate answer context global encoding sub-unit for performing context semantic global encoding based on space constraints on the sequence distribution of semantic embedding encoding vectors of candidate answer financial document segments at the granularity to obtain a candidate answer context semantic encoding vector as the candidate answer context semantic encoding features.
6. The financial document automated parsing and question answering system based on a large model according to claim 5, wherein The candidate answer context global encoding sub-unit includes: The end - axial information propagation constraint factor calculation secondary subunit is used to calculate the end - information propagation constraint factor and the axial - information propagation constraint factor of each candidate answer financial document segment granularity description semantic embedding coding vector in the sequence distribution of the candidate answer financial document segment granularity description semantic embedding coding vectors, so as to obtain the candidate answer financial document semantic end - information propagation constraint factor and the candidate answer financial document semantic axial - information propagation constraint factor; The semantic information propagation constraint factor calculation secondary subunit is used to calculate the candidate answer financial document semantic information propagation constraint factor of each candidate answer financial document segment granularity description semantic embedding coding vector based on the candidate answer financial document semantic end - information propagation constraint factor and the candidate answer financial document semantic axial - information propagation constraint factor of each candidate answer financial document segment granularity description semantic embedding coding vector in the sequence distribution of the candidate answer financial document segment granularity description semantic embedding coding vectors; The candidate answer context semantic feature generation secondary subunit is used to perform semantic information dynamic constraint propagation on the sequence distribution of the candidate answer financial document segment granularity description semantic embedding coding vectors based on the candidate answer financial document semantic information propagation constraint factor of each candidate answer financial document segment granularity description semantic embedding coding vector to obtain the candidate answer context semantic coding vector.
7. The financial document automated parsing and question answering system based on a large model according to claim 6, characterized in that, The end - axial information propagation constraint factor calculation secondary subunit is used for: Extracting the last candidate answer financial document segment granularity description semantic embedding coding vector from the sequence distribution of the candidate answer financial document segment granularity description semantic embedding coding vectors as the candidate answer financial document semantic propagation space end - constraint anchoring coding vector; Performing clustering analysis on the sequence distribution of the candidate answer financial document segment granularity description semantic embedding coding vectors to obtain the candidate answer financial document semantic propagation space axial - constraint anchoring coding vector; Calculating the candidate answer financial document semantic end - information propagation constraint factor of each candidate answer financial document segment granularity description semantic embedding coding vector in the sequence distribution of the candidate answer financial document segment granularity description semantic embedding coding vectors relative to the candidate answer financial document semantic propagation space end - constraint anchoring coding vector; Calculating the candidate answer financial document semantic axial - information propagation constraint factor of each candidate answer financial document segment granularity description semantic embedding coding vector in the sequence distribution of the candidate answer financial document segment granularity description semantic embedding coding vectors relative to the candidate answer financial document semantic propagation space axial - constraint anchoring coding vector.
8. The financial document automated parsing and question answering system based on a large model according to claim 7, wherein The semantic information propagation constraint factor calculation secondary subunit is used for: Using a canonical space constraint matrix to perform horizontal compact covariance on the candidate answer financial document semantic end - information propagation constraint factor and the candidate answer financial document semantic axial - information propagation constraint factor corresponding to the candidate answer financial document segment granularity description semantic embedding coding vector to obtain the candidate answer financial document semantic end - information propagation compact covariance constraint factor and the candidate answer financial document semantic axial - information propagation compact covariance constraint factor; Based on the candidate answer financial document semantic end - information propagation compact covariant constraint factor and the candidate answer financial document semantic axial - information propagation compact covariant constraint factor, perform propagation constraint integration processing on the candidate answer financial document semantic end - information propagation constraint factor and the candidate answer financial document semantic axial - information propagation constraint factor to obtain an optimized candidate answer financial document semantic end - information propagation constraint factor and an optimized candidate answer financial document semantic axial - information propagation constraint factor; Perform non - linear activation processing on the optimized candidate answer financial document semantic end - information propagation constraint factor and the optimized candidate answer financial document semantic axial - information propagation constraint factor to obtain the candidate answer financial document semantic information propagation constraint factor corresponding to the candidate answer financial document segment - granularity description semantic embedding coding vector.
9. The financial document automated parsing and question answering system based on a large model according to claim 8, wherein The generating unit is configured to: input the candidate answer context semantic coding vector into the thought - chain reasoning generation model to obtain the answer reply text.
Citation Information
Patent Citations
A question-answering method and system based on knowledge graphs and large language models
CN116775847B
A Generative Question Answering Method and System Based on a Knowledge Graph of a Large Language Model
CN117033608B
Large language model medical question answering system based on medical record knowledge graph
CN117056493B
Cited By
Preparation method and system of ultrathin printed circuit board
CN120166642A
High-temperature-resistant and high-pressure-resistant lubricating oil and preparation method thereof
CN120340644A
Financial document editing method and system based on large language model, medium and electronic equipment
CN121031610A
Document hierarchical index construction method based on layout visual features and path constraints
CN121092651A
AI knowledge service system and method based on financial intelligent context protocol
CN121213242A