Low-resource financial annual report large model question answering system construction method
By building structured and semi-structured Q&A modules in the financial annual report big model question and answer system, the problem of insufficient semi-structured data processing capabilities in low-resource scenarios is solved, and high-quality data generation and generalization capabilities of complex tables are improved.
Patent Information
- Application Number
- CN202510358722.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-25
- Publication Date
- 2025-06-24
AI Technical Summary
The existing technology is difficult to effectively process semi-structured data in low-resource scenarios, and cannot guarantee data quality and diversity, resulting in insufficient financial annual report big model Q&A system in understanding and processing financial data.
A low-resource financial annual report big model question and answer system construction method is proposed, including calling module, structured question and answer module and semi-structured question and answer module. Through data preprocessing, the generation of structured and semi-structured data sets, and training based on large language models, a question-and-answer system with cross-table computing, in-depth text understanding and data comparison capabilities are built.
It realizes the generation of high-quality financial field instruction data in low-resource scenarios, reduces the dependence of model training on labeled resources, improves the generalization ability of complex tables, and ensures data quality and diversity.
Smart Images

Figure CN120196723A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of automatic question answering, and particularly to a method for constructing a large model question answering system for financial annual reports with low resources. Background Art
[0002] In the financial field, financial report data, as an important basis for analysts to evaluate the operating conditions of enterprises, involves multiple aspects such as financial indicators, corporate governance, business development, and market environment. First, the formats of financial statements are diverse and the structures are complex, posing high challenges to table parsing technology; second, these data are often presented in a combination of tables and texts, and the semi-structured form of mixed text tables poses significant challenges to model understanding. Finally, due to the characteristics of vertical fields, high-quality instruction data for fine-tuning large models is scarce. Therefore, how to construct and train a large model question answering system with cross-table calculation, in-depth text understanding, and data comparison capabilities in a low-resource scenario has become an important research topic for improving the accuracy and effectiveness of information extraction.
[0003] The challenges in constructing a financial report question answering system mainly lie in the understanding of semi-structured data and the generation of high-quality instruction fine-tuning data in low-resource scenarios. Existing methods mainly use financial report data to generate the data required for pre-training and instruction fine-tuning, and iterate through user feedback data. They require high-quality financial datasets for cold start, and overly rely on the subjective scoring of users for answers, lacking a screening mechanism for the quality and diversity of training data, and unable to guarantee the quality and diversity of the final user feedback data, or mainly focus on the processing of text data and have limited capabilities in processing semi-structured data containing tables. Summary of the Invention
[0004] Object of the Invention: The object of the present invention is to overcome the defects of insufficient semi-structured data processing capabilities and uncontrollable data quality in low-resource scenarios in the prior art, and provide a method for constructing a large model question answering system for financial annual reports with low resources.
[0005] To solve the above technical problems, the present invention discloses a method for constructing a large model question answering system for financial annual reports with low resources. The question answering system includes a calling module, a structured question answering module, and a semi-structured question answering module;
[0006] The calling module is constructed based on existing large language models;
[0007] The construction steps of the structured question answering module and the semi-structured question answering module are as follows:
[0008] Step 1: Input the financial annual report, and the file name of the financial annual report consists of the full company name, short name, and year;
[0009] Step 2: Perform data preprocessing, extract financial data from the annual financial reports to generate a financial database, and construct a vector database through the vector set obtained after chunking;
[0010] Step 3: According to a small amount of manually annotated template Q&A data, the financial database, and the manually annotated financial indicator synonym mapping dictionary, use the structured data synthesis strategy to generate a structured data set, and train a structured Q&A model from the base model; according to the content of the annual financial report, use the semi-structured data synthesis strategy to generate a semi-structured data set, and train a semi-structured Q&A model from the base model;
[0011] Step 4: Build a structured Q&A module based on the structured Q&A model; build a semi-structured Q&A module based on the semi-structured Q&A model.
[0012] The Q&A process of the annual financial report large model Q&A system is as follows:
[0013] Step 2.1: Input the user's query statement;
[0014] Step 2.2: Call the module to determine whether the user's query statement is a financial-related question, and extract the company name and year included in the query statement to construct a binary list P of company name - year; if it is a financial-related question, call the structured Q&A module for processing; if it is not a financial-related question, call the semi-structured Q&A module for processing;
[0015] Step 2.3: The corresponding semi-structured Q&A module or structured Q&A module gives a reply.
[0016] In Step 2, the data preprocessing process is specifically as follows:
[0017] Step 2-1: Use a text recognition tool to identify the text area and table area in the annual financial report, and parse the text area and table area into a document tree. The document tree of the i-th annual report can be represented in the following form;
[0018] T i =(V i ,E i )
[0019] where V i represents the collection of chapter titles and chapter contents, and E i represents the collection of hierarchical relationships. The chapter titles include sub-chapter titles;
[0020] Step 2-2: Find the financial data included in the document through the matching of the chapter titles of the document tree, and generate a financial database;
[0021] Step 2-3: For each annual report document tree T i, the subtree with the first-level chapter as the root node is texturized to obtain the corresponding sub-chapter text;
[0022] Step 2-4: Divide the obtained sub-chapter text into chunks. The text chunks include table areas and text areas, and are divided according to delimiters (including full stops, line breaks, etc.). The maximum length of the chunks is determined by manual experience, generally between 1024 tokens and 4096 tokens, not exceeding the upper limit of the context length of the model. When the length of the table exceeds the set maximum length of the text chunk, truncate the table area, and take the first m rows (the value of m is determined by manual experience, usually between 1 row and 5 rows, and the length does not exceed 64 tokens) of the text area before the table area and the table header as the metadata of the table (the metadata is a string including the full company name, company abbreviation, and annual report year). There is a maximum limit on the number of characters of the metadata, and add the metadata text of the table before the truncated table area; when the length of the ordinary text exceeds the set maximum length of the text chunk, do not process the truncated ordinary text, and obtain the set of text chunks {d1, d2,.., d n} where n represents the total number of chunks, and use the embedding model f e to encode the text chunks vector i = f e (d i ), and obtain the vector set {vector1, vector2,…, vector n};
[0023] Step 2-5: Build a vector database based on the text chunks and text chunk vectors. The annual report title and first-level chapter title of the document tree are stored in the vector library in the form of metadata, and the corresponding text chunks and text chunk vectors can be indexed through the annual report title and first-level chapter title.
[0024] In Step 3, the specific process of generating the structured data set is as follows:
[0025] Step 3-1-1: Manually label a number of seed data. A single piece of seed data can be expressed as (q user , s sql ), where q user represents the question template data, and s sql represents the SQL (Structured Query Language) template data corresponding to the question template data;
[0026] Step 3-1-2: According to the seed data and the financial database, the teacher model f teacher generates new user question template data q' user and the corresponding SQL template data s' sql :
[0027] (q′ user , s′ sql) = f teacher (q user , s sql )
[0028] Repeat several times to obtain an augmented template dataset S struct , the template dataset S struct Each question-and-answer pair sample in is composed of user question template data and corresponding SQL template data;
[0029] Step 3-1-3, for the template dataset S struct Use the greedy subset selection algorithm to obtain a subset with higher diversity than the template dataset S struct ;
[0030] Step 3-1-4, for each question template and SQL template in the subset obtained in Step 3-1-3, perform field filling, and use the synonym dictionary to perturb the filling of the question template to obtain a structured dataset;
[0031] In Step 3, the specific process of generating the semi-structured dataset is as follows:
[0032] Step 3-2-1, randomly extract chapter fragments other than financial data from the financial annual report, and the teacher model f teacher Generate documents and question-and-answer pairs according to the chapter fragments, repeat the generation process several times, and the obtained documents and question-and-answer pairs form an initial dataset;
[0033] Step 3-2-2, define several perturbation operators, and perform perturbation enhancement on the initial dataset to obtain a perturbation-enhanced dataset;
[0034] Step 3-2-3, use the greedy subset selection algorithm for the perturbation-enhanced dataset to obtain a subset with higher diversity relative to the perturbation-enhanced dataset as the semi-structured dataset.
[0035] The specific greedy subset selection algorithm in Step 3-1-3 or Step 3-2-3 is as follows:
[0036] Step 3-1, set the target subset as S greedy , and set its target size as k;
[0037] Step 3-2, randomly select an initial sample s in the input dataset S i , and initialize the target subset S greedy ;
[0038] Step 3-3, calculate the vector center c of S greedy , from the input dataset S excluding the target subset S struct greedy Select the sample with the lowest cosine similarity in the part of s j And add it to the target subset:
[0039]
[0040] Step 3-4: Repeat Step 3-3 until the size of the target subset reaches k;
[0041] The input data set S is equal to the template data set S in Step 3-1-2 struct Or the perturbed enhanced data in Step 3-2-2.
[0042] The specific steps of the perturbation enhancement by the perturbation operator in Step 3-2-2 are as follows:
[0043] Step 3-2-2-1: Define the perturbation operation, which includes attribute deletion, content shuffling, and structure adjustment;
[0044] Step 3-2-2-2: Preset a prompt word for each perturbation operation;
[0045] Step 3-2-2-3: The teacher model f teacher Rewrites the table content according to the requirements of the prompt word to obtain the perturbed enhanced data set;
[0046] Specifically expressed as follows:
[0047] Perturbation operation function set:
[0048]
[0049] Each function among them Represents a specific table perturbation method; for the table area table in each text block, randomly select an operator To perturb the table block and obtain the perturbed document Thus, the perturbed enhanced data set is obtained.
[0050] In Step 4, the specific process of the structured Q&A module is as follows:
[0051] Step 4-1-1: Use the structured Q&A model to generate the SQL statement for querying by the company name and target year in the user's question;
[0052] Step 4-1-2: Execute the generated SQL statement in the financial database to obtain the table of the query result;
[0053] Step 4-1-3: Convert the obtained table into Markdown format text and use it as the context for the structured Q&A model to generate an answer.
[0054] In step 4, the specific process of the semi-structured question-answering module is as follows:
[0055] Step 4-2-1: For each company name-year pair p in the binary list P of company name-year i (p i ∈P), form a string sp i , and calculate the BM25 similarity between the string sp i and the metadata in the document tree of the vector database;
[0056] Step 4-2-2: According to the BM25 similarity, recall the list L of the top x annual report document trees with the highest similarity, L = {T1, T2,..., T x} where x is greater than or equal to 1 and less than or equal to the number of document trees in the vector database;
[0057] Step 4-2-3: Select the most suitable document tree T from the document tree list L by the base large model f top ;
[0058] Step 4-2-4: Use the base large model f to filter out the most relevant chapter subset from the subset of sub-chapters under the first-level chapter of the document tree T top :
[0059] Step 4-2-5: For each text block in the most relevant chapter subset, retrieve all corresponding text block vectors from the vector database;
[0060] Step 4-2-6: Encode the user question q user using the embedding model f e to obtain f e (q user ), and calculate the similarity with each text block vector vector i corresponding to the text block di:
[0062]
[0063] Take the text blocks corresponding to the top y text block vectors with the highest vector similarity (y usually ranges from 5 to 20 and does not exceed the total number of document blocks) as the recalled document set D r ;
[0064] Step 4-2-7: Use the semi-structured question-answering model f semi-struct to generate an answer a;
[0065] a = f semi-struct (D r , q user )
[0066] In the present invention, preferably, the teacher model f teacher is the existing GPT-4 model.
[0067] In the present invention, preferably, the base large model f is the existing CHATGLM2-6B model.
[0068] In the present invention, preferably, the embedding model f e is the existing BGE-M3 model.
[0069] Advantageous effects:
[0070] Low-resource data generation: Design an automatic generation framework for large-scale training data that only requires a small amount of manual annotation. Through question-and-answer pair generation and data diversification rewriting, synthesize high-quality instruction data in the financial field, and reduce the dependence of model training on annotation resources.
[0071] Improve the generalization ability of complex tables: In view of the characteristics of the mixed structured and semi-structured tables in financial reports, implement different training strategies, and respectively optimize the NL2SQL (Natural Language to SQL) ability of the model for regular tables (such as balance sheets) and the semi-structured understanding ability of the model for irregular tables (such as company information tables). Description of the drawings
[0072] Figure 1 is the system question-and-answer flowchart.
[0073] Figure 2 is the visualization diagram of the partial structure of the document tree.
[0074] Figure 3 is the structured data generation flowchart.
[0075] Figure 4 is the semi-structured data generation flowchart. Specific implementation manners
[0076] The present invention proposes a construction and training method for a financial report question-and-answer system based on a large model. According to the complexity of the table, the user's questions are divided into two categories: structured table question-and-answer and semi-structured table question-and-answer for processing. For structured tables, the NL2SQL paradigm is used to efficiently extract the target value, and for the semi-structured table text mixing scenario, the retrieval-enhanced paradigm is used for question-and-answer. In addition, we also designed a set of data generation and training processes in low-resource scenarios. It can realize the generation of large-scale high-quality data and large model training in low-resource scenarios.
[0077] The embodiment of the present application discloses a method for constructing a low-resource financial annual report large model question-and-answer system. The question-and-answer system includes a calling module, a structured question-and-answer module, and a semi-structured question-and-answer module;
[0078] The calling module is built based on existing large language models;
[0079] The steps for building the structured Q&A module and the semi-structured Q&A module are as follows:
[0080] Step 1: Input the PDF file of the financial annual report. The file name of the financial annual report consists of the full company name, short name, and year, such as "Jiangsu Anchor Intelligent Transmission Engineering Technology Co., Ltd.__Anchor Smart Electric__2022__Annual Report";
[0081] Step 2: Perform data preprocessing. Extract financial data from the financial annual report to generate a financial database, store it using a relational database, and construct a vector database from the vector set obtained after chunking and store it using the vector database;
[0082] Step 3: According to a small amount of manually annotated template Q&A data, the financial database, and the manually annotated financial indicator synonym mapping dictionary, use the structured data synthesis strategy to generate a structured data set and train ChatGLM2-6B to obtain a structured Q&A model; according to the content of the financial annual report, use the semi-structured data synthesis strategy to generate a semi-structured data set and train ChatGLM2-6B to obtain a semi-structured Q&A model;
[0083] Step 4: Build a structured Q&A module based on the structured Q&A model; build a semi-structured Q&A module based on the semi-structured Q&A model.
[0084] As Figure 1 shown, the process of specific Q&A examples of the financial annual report large model Q&A system is as follows:
[0085] Step 2.1: The user inputs a question, such as "What are the net assets of Anchor Smart Electric in 2019 and 2020?";
[0086] Step 2.2: The calling module determines whether the user's query statement is a finance-related question, extracts the company name and year contained in the query statement, and constructs a binary list P of company name-year. The corresponding value in this example is [('Anchor Smart Electric', 2019), ('Anchor Smart Electric', 2020)]; if it is a finance-related question, call the structured Q&A module for processing; if it is not a finance-related question, call the semi-structured Q&A module for processing. The corresponding value in this example is a finance-related question, and the structured Q&A module should be called;
[0087] Step 2.3: The corresponding semi-structured question-and-answer module or structured question-and-answer module gives an answer. The corresponding answer in the example here is "The net assets of Anke Smart Electric in 2019 and 2020 were 828,521,233.86 yuan and 929,391,576.81 yuan respectively."
[0088] In step 2, the specific example of the data preprocessing process is as follows:
[0089] Step 2-1, use Python's pdfplumber library to identify the text area and table area in the financial annual report PDF file, use the candidate regular expression list (such as "^第[一二三四五六七八九]+.*") to match the chapter title and its subtitle, and parse the text area and table area into a document tree according to the hierarchical structure of the chapter title and its subtitle. The document tree of the i-th annual report can be expressed as follows:
[0090] T i =(V i ,E i )
[0091] Among them, V i Indicates a collection of chapter titles and chapter contents. Chapter titles include sub-chapter titles. i Represents a collection of hierarchical relationships. The specific structure visualization is as follows Figure 2 As shown;
[0092] Step 2-2, match the chapter title containing "Financial Report" through the chapter title of the document tree T, and search for the target table (consolidated balance sheet, consolidated cash flow statement and consolidated income statement) according to the following rules: match the row containing the target table name in the chapter content, and search down to the first table. The table must meet the following conditions: 1) The row name contains the specified field (consolidated balance sheet is "monetary funds", consolidated cash flow statement is "operating income and interest income", and consolidated cash flow statement is "cash received"), and 2) The table header does not contain the word "adjustment". After finding the corresponding table, write it into the relational database to form a financial database.
[0093] Step 2-3: For each annual report document tree T i , the subtree with the first-level chapter as the root node is textualized to obtain the corresponding sub-chapter text;
[0094] Step 2-4: Chunk the obtained sub-chapter text. The maximum number of tokens for each text chunk is 1024. The text chunk includes a table area and a text area. When the table exceeds 1024 tokens, truncate the table area, take the first 3 lines of the text area before the table area and the table header as the metadata of the table. The maximum number of tokens for the metadata is 64. Add the metadata text of the table before the truncated table area, and do not process the truncated ordinary text, to obtain a set of chunked text chunks {d1, d2,.., d n},where n represents the total number of chunks. Use the BGE-M3 model (denoted as f e ) to encode the text chunks to obtain a vector collection {f e (d1), …, f e (d n )};
[0095] Step 2-5: Build a vector database based on the text chunks and text chunk vectors. The annual report title and the first-level chapter titles of the document tree are stored in the vector library in the form of metadata, and the corresponding text chunks and text chunk vectors can be indexed through the annual report title and the first-level chapter titles.
[0096] In Step 3, as Figure 3 shown, the specific process of generating the structured data set is as follows:
[0097] Step 3-1-1: Manually annotate 30-100 pieces of seed data. A single piece of seed data can be expressed as (q user , s sql ), where q user represents the problem template data, such as "Tell me how many companies' {attribute} are greater than {threshold} in the {year} year?", and s sql represents the SQL template data corresponding to the problem template data, such as "SELECT count(company's Chinese name) FROM finance WHERE year = {year} and {attribute} > {threshold}";
[0098] Step 3-1-2: According to the seed data and the financial database, use GPT-4 as the teacher model f teacher to generate new user problem template data q' user and the corresponding SQL template data s' sql :
[0099] (q user , s sql ) = f teacher (q user , s sql )
[0100] Given prompts to the teacher model, such as "Please generate a new Q&A pair template with the following requirements: The only variables allowed are {attribute}, {company_name}, {year}, and {threshold}. The question template needs to describe a query scenario based on one or more of these four variables. For example: Example question: Tell me how many companies have {attribute} greater than {threshold} in the year {year}? SQL example: SELECT count(company's Chinese name) FROM finance WHERE year = {year} and {attribute} > {threshold} The fields in the finance table are: company's Chinese name, year, {attribute} (multiple financial indicator names) Please generate a template as different as possible from the above template. The answer to your question doesn't have to be a company name. It can be a year, a ratio, an indicator value, or other calculated values"
[0101] The generated example q' user is "What is the ratio of the highest value to the lowest value of the {attribute} indicator among all companies in the year {year}?", and the corresponding s' sql is "SELECT MAX({attribute}) / MIN({attribute}) AS ratio FROM finance WHERE year = {year};"
[0102] Repeat several times to obtain an augmented template dataset S struct , the template dataset S struct Each Q&A pair sample of which consists of user question template data and corresponding SQL template data;
[0103] Step 3-1-3. For the template dataset S struct Use the greedy subset selection algorithm to obtain a subset with higher diversity than the template dataset S struct ;
[0104] Step 3-1-4: For each question template and SQL template in the subset obtained in Step 3-1-3, perform field filling. For example, "Tell me how many companies have {attribute} greater than {threshold} in the {year} year?" After filling, it becomes "Tell me how many companies have total owner's equity greater than 10,000,000 in the 2019 year?", and the corresponding SQL is "SELECT count(Chinese name of the company) FROM finance WHERE year = 2019 and owner's equity > 10000000". Use the synonym dictionary to perturb the filling of the question template. For example, synonyms for "total owner's equity" include "total owner's equity", "total shareholder's equity", "net assets", etc. Only change the values in the question, and keep the SQL template unchanged (consistent with the financial database). For example, "Tell me how many companies have net assets greater than 10,000,000 in the 2019 year?". After completing the random filling of all templates, a structured dataset is obtained;
[0105] In Step 3, as Figure 4 shown, the specific process of generating the semi-structured dataset is as follows:
[0106] Step 3-2-1: Randomly extract chapter fragments other than the consolidated balance sheet, consolidated cash flow statement, and consolidated income statement from the financial annual report. Use GPT-4 as the teacher model f teacher to generate documents and question-and-answer pairs based on the chapter fragments,
[0107] Given prompts to the teacher model, such as "Please ask three questions based on the following annual report excerpt. The questions must include the company name and the year.\n\nCurrent annual report: Annual Report of China Gezhouba Group Co., Ltd. 2020\n\nCurrent table of contents hierarchy: Section VIII Directors, Supervisors, Senior Management and Employees >> VI. Employees of the Parent Company and Major Subsidiaries >> (1) Employee Information\n| Number of employees in the parent company | 614 |\n| Number of employees in the major subsidiaries | 38,457 |\n| Total number of employees | 39,071 |\n| Number of retired employees for whom the parent company and major subsidiaries bear expenses | 29,912 |\n\n| Category of professional composition | Number of people in professional composition |\n| Skilled workers | 10,294 |\n| Engineering and technical personnel | 10,514 |\n| Management personnel | 16,585 |\n| Service personnel | 1,678 |\n| Total | 39,071 |\n\n| Category of education level | Number of people |\n| Doctoral students | 38 |\n| Master's students | 2,307 |\n| Bachelor's degree | 16,550 |\n| Associate degree | 9,058 |\n| Below associate degree | 11,118 |\n| Total | 39,071 |\n\nFinal output format\n[Q1]\n\nAnalysis:\n[Summarize the relevant data or content here to explain the basis for formulating the question.]\n\nQuestion:\n[Pose a question-and-answer question, indicating the company name and the year of the annual report.]\n\nAnswer:\n[The content of the answer should be able to obtain relevant information from the original annual report.]\n\nAnalysis:\n[Explain the basis and logic of the answer.]\n\n[Q2]\n... same format...\n\n[Q3]\n... same format...”
[0108] One corresponding question generated is: In the 2020 annual report of China Gezhouba Group Co., Ltd., what is the number of employees in the parent company?
[0109] The answer corresponding to the generated question is: In the 2020 annual report of China Gezhouba Group Co., Ltd., the number of employees in the parent company is 614 people.
[0110] Repeat the generation process several times, and the resulting documents and question-and-answer pairs form the initial dataset;
[0111] Step 3-2-2, Define several perturbation operators, and perform perturbation enhancement on the initial dataset to obtain a perturbed and enhanced dataset;
[0112] Step 3-2-3, Use the greedy subset selection algorithm on the perturbed and enhanced dataset to obtain a subset with higher diversity relative to the perturbed and enhanced dataset as the semi-structured dataset.
[0113] The specific greedy subset selection algorithm described in step 3-1-3 or step 3-2-3 is as follows:
[0114] Step 3-1: Given the existing dataset S, set the target subset as S greedy , and set its target size as i.e., half of the size of dataset S;
[0115] Step 3-2: Randomly select an initial sample s from the input dataset S i , and initialize the target subset S greedy .
[0116] Step 3-3: Calculate the vector center c of S greedy , and select the sample with the lowest cosine similarity cosine_similarity from the part of the input dataset S that does not include the target subset S struct , and add s greedy to the target subset, : j
[0117]
[0118] Step 3-4: Repeat step 3-3 until the size of the target subset reaches
[0119] The input dataset S is equal to the template dataset S in step 3-1-2 struct or the perturbed enhanced data in step 3-2-2
[0120] The specific steps of perturbation enhancement using the perturbation operator described in step 3-2-2 are as follows:
[0121] Step 3-2-2-1: Define the perturbation operations, which include attribute deletion, content shuffling, and structure adjustment;
[0122] Step 3-2-2-2: Preset a prompt word for each perturbation operation;
[0123] Prompt word for attribute deletion, such as "Please randomly delete the element values in the following table in the document, set them to empty or delete the corresponding rows, without affecting the answer to the user's question. Just reply with the modified document, without including other content: {Fragment of financial annual report}\nQuestion: {User's question}"
[0124] Prompt word for content shuffling, such as "Please randomly shuffle the rows or columns in the following table in the document, you can split or merge the table, without affecting the answer to the user's question. Just reply with the modified document, without including other content: {Fragment of financial annual report}\nQuestion: {User's question}"
[0125] Structure adjustment prompt, e.g., "Please rewrite the following table structure without adding or deleting table content to make it more complex, such as adding a hierarchical structure. You cannot add content and do not need to follow Markdown syntax. Convert it to HTML / CSV / Latex or a custom format without affecting the answer to the user's question. Just reply with the modified document without including other content: {Fragment of financial annual report}\nQuestion: {User's question}"
[0126] Step 3-2-2-3, teacher model f teacher Rewrite the table content according to the requirements of the prompt to obtain a perturbation-enhanced dataset;
[0127] Specifically expressed as follows:
[0128] Perturbation operation function set:
[0129]
[0130] Each function among them represents a specific table perturbation method; for the table area table in each text block, randomly select an operator to perturb the table block and obtain the perturbed table area Thus, a perturbation-enhanced dataset is obtained.
[0131] In step 4, the specific process of the structured question answering module is as follows:
[0132] Step 4-1-1, use the structured question answering model to generate a query SQL statement according to the user's question, such as "What were the net assets of Anchor Smart Grid in 2019 and 2020?", and a company name-year binary list, such as [('Anchor Smart Grid', 2019), ('Anchor Smart Grid', 2020)], like "SELECT the Chinese name of the company, year, total owner's equity FROM finance WHERE year in (2019, 2020) and the Chinese name of the company = 'Anchor Smart Grid'. Here, due to the effect of synonym enhancement training, 'net assets' is replaced by the field 'total owner's equity' in the financial database by the model;
[0133] Step 4-1-2, execute the generated SQL statement in the financial database to obtain a table of the query results, such as
[0134] Company's Chinese name Year Total owners' equity ANKO Power 2019 828,521,233.86 ANKO Power 2020 929,391,576.81
[0135] Step 4-1-3: Convert the table into a text in Markdown format through Python code, and use it as the context for the structured Q&A model to generate an answer, that is, "The net assets of Anchor Smart Grid in 2019 were 828,521,233.86 yuan, and in 2020 they were 929,391,576.81 yuan."
[0136] In step 4, the specific process of the semi-structured Q&A module is as follows:
[0137] Step 4-2-1: For each company name-year pair p in the binary list P of company name-year i (p i ∈P) to form a string sp i , such as [(Anchor Smart Grid, 2019), (Anchor Smart Grid, 2020)], calculate the bm25 similarity between the string sp i and the annual report T j 's metadata M j (The annual report metadata is a string including the full company name, company abbreviation, and annual report year, such as "Jiangsu Anchor Smart Grid Transmission Engineering Technology Co., Ltd. Anchor Smart Grid 2019");
[0138] Step 4-2-2: According to the bm25 similarity, recall the list L of the top five annual reports with the highest similarity = {T1, T2,..., T5};
[0139] The example of the list is as follows:
[0140] 1. Company name: Jiangsu Anchor Smart Grid Transmission Engineering Technology Co., Ltd., Company abbreviation: Anchor Smart Grid, Year: 2019
[0141] 2. Company name: Jiangsu Anchor Smart Grid Transmission Engineering Technology Co., Ltd., Company abbreviation: Anchor Smart Grid, Year: 2020
[0142] 3. Company name: Jiangsu Anchor Smart Grid Transmission Engineering Technology Co., Ltd., Company abbreviation: Anchor Smart Grid, Year: 2021
[0143] 4. Company name: Hangzhou Reliable Nursing Products Co., Ltd., Company abbreviation: Reliable Co., Year: 2019
[0144] 5. Company name: San'an Optoelectronics Co., Ltd., Company abbreviation: San'an Optoelectronics, Year: 2019
[0145] Step 4-2-3: Let CHATGLM2-6B be the base large model f to select the most suitable document tree T from the document tree list L top ;
[0146] The following prompt words are given for the first element in the list of pairs (AnKao Smart Electric Co., Ltd., 2019):
[0147] Please determine which of the following 5 annual report names is the same annual report as the queried annual report "AnKao Smart Electric 2019".
[0148] The candidate companies are as follows:
[0149] 1. Company name: Jiangsu AnKao Smart Transmission Engineering Technology Co., Ltd., Company abbreviation: AnKao Smart Electric, Year: 2019
[0150] 2. Company name: Jiangsu AnKao Smart Transmission Engineering Technology Co., Ltd., Company abbreviation: AnKao Smart Electric, Year: 2020
[0151] 3. Company name: Jiangsu AnKao Smart Transmission Engineering Technology Co., Ltd., Company abbreviation: AnKao Smart Electric, Year: 2021
[0152] 4. Company name: Hangzhou Reliable Nursing Products Co., Ltd., Company abbreviation: Reliable Co., Ltd., Year: 2019
[0153] 5. Company name: Sanan Optoelectronics Co., Ltd., Company abbreviation: Sanan Optoelectronics, Year: 2019
[0154] Please only return the matching company number (a number between 1 and 5), if none match, return 0. Only return the number, no other explanations are required.
[0155] Here the model returns 1, for the "Jiangsu AnKao Smart Transmission Engineering Technology Co., Ltd.__300617__AnKao Smart Electric__2019__Annual Report".
[0156] This example involves two pairs [(AnKao Smart Electric, 2019), (AnKao Smart Electric, 2020)]. After selection according to the above process respectively, two annual reports are obtained, namely the "Jiangsu AnKao Smart Transmission Engineering Technology Co., Ltd.__300617__AnKao Smart Electric__2019__Annual Report" and the "Jiangsu AnKao Smart Transmission Engineering Technology Co., Ltd.__300617__AnKao Smart Electric__2020__Annual Report"
[0157] Step 4-2-4. Use the base large model f(CHATGLM2-6B) to filter out the most relevant chapter subset from the subset of chapters under the first-level chapters of the document tree T top The following prompt words are given:
[0158] "The following is the table of contents structure of the {year} annual report of {company_name}. Please select **all** the chapters that may be relevant to the question. You cannot return an empty list.\nTable of contents structure:\n{content_structure}\n\nQuestion:\n{question}\n\nYour output format should be as follows:\n[Analysis]\nxxx\n[Chapters]\n```json\n[\n\"Chapter Name 1\",\n\"Chapter Name 2\",\n...\n]\n```\n"
[0159] In this example, the corresponding relevant chapter is ["Section II Company Profile and Main Financial Indicators"]
[0160] Step 4-2-5: For each text block in the most relevant chapter subset, retrieve the corresponding text block vector from the vector database;
[0161] Step 4-2-6: For the user question q user (e.g., "Where is the office address of Anchor Smart Grid in 2019? What about in 2020?") Use the embedding model f e to encode and get f e (q user ), calculate the vector similarity with all candidate text block vectors:
[0162]
[0163] Take the text blocks corresponding to the top 10 text block vectors with the highest vector similarity as the recall document set D r ;
[0164] Step 4-2-7: Use the semi-structured question answering model f semi-struct to generate the answer a;
[0165] a = f semi-struct (D r , q user )
[0166] In this example, the answer a here is "The office addresses of Anchor Smart Grid in 2019 and 2020 are both at No. 100, Tianmu Lake Avenue, Liyang City, Jiangsu Province."
[0167] The source of the financial annual report data involved in this embodiment is:
[0168] Regular reports disclosed by the Shanghai Stock Exchange: https: / / www.sse.com.cn / disclosure / listedinfo / regular / ;
[0169] Regular reports disclosed by the Shenzhen Stock Exchange: https: / / www.szse.cn / disclosure / listed / fixed / index.html。
[0170] The present invention provides a method for constructing a low-resource financial annual report large model Q&A system. There are many methods and ways to specifically implement this technical solution. The above description is only the preferred embodiment of the present invention. It should be noted that for those of ordinary skill in the art in this technical field, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention. Each component not clearly defined in this embodiment can be implemented using existing technologies.
Claims
1. A method for constructing a large model question-answering system for financial annual reports with low resources, characterized in that: The question-answering system includes a calling module, a structured question-answering module, and a semi-structured question-answering module; The calling module is constructed based on the existing large language model; The steps for constructing the structured question-answering module and the semi-structured question-answering module are as follows: Step 1: Enter the annual financial report. The file name of the annual financial report consists of the company's full name, abbreviation and year; Step 2: Perform data preprocessing, extract financial data from the financial annual report to generate a financial database, and construct a vector database through the vector set obtained after segmentation; Step 3: Based on a small amount of manually annotated template question-answering data, the financial database, and the manually annotated synonym mapping dictionary of financial indicators, a structured data synthesis strategy is used to generate a structured data set, and a structured question-answering model is trained. Based on the content of the financial annual report, a semi-structured data synthesis strategy is used to generate a semi-structured data set, and a semi-structured question-answering model is trained. Step 4: Build a structured question-answering module based on the structured question-answering model; Construct a semi-structured question and answer module based on the semi-structured question and answer model.
2. According to claim 1, a method for constructing a low-resource financial annual report large model question-answering system is characterized in that: The question-answering process of the financial annual report large model question-answering system is as follows: Step 2.1: Enter the user's query statement; Step 2.2: The calling module determines whether the user's query statement is a financial-related question, and extracts the company name and year contained in the query statement to construct a binary list P of company name-year; If the question is financial, the structured question-answering module is used for processing; if the question is not financial, the semi-structured question-answering module is used for processing; Step 2.3: The corresponding semi-structured question-answering module or structured question-answering module gives a response.
3. The method for constructing a low-resource financial annual report large model question-answering system according to claim 1, characterized in that: In step 2, the data preprocessing process is as follows: Step 2-1: Use the text recognition tool to identify the text area and table area in the financial annual report, and parse the text area and table area into a document tree. The document tree of the i-th annual report is represented in the following form: T i =(V i ,E i ) Among them, V i Indicates a collection of chapter titles and chapter contents, E i Represents a collection of hierarchical relationships; Step 2-2, find the financial data contained in the document by matching the chapter titles of the document tree, and generate a financial database; Step 2-3: For each annual report document tree T i , the subtree with the first-level chapter as the root node is textualized to obtain the corresponding sub-chapter text; Step 2-4: Divide the sub-chapter text obtained in step 2-3 into blocks and use the embedding model f e Encode the text blocks obtained by segmentation to obtain text block vectors; Step 2-5: construct a vector database based on text blocks and text block vectors. The annual report title and first-level chapter title in the document tree are stored in the vector database in the form of metadata. The corresponding text blocks and text block vectors can be indexed through the annual report title and first-level chapter title.
4. The method for constructing a low-resource financial annual report large model question-answering system according to claim 1, characterized in that: In step 3, the specific process of generating a structured data set is as follows: Step 3-1-1: Manually annotate seed data. A single seed data is represented as (q user ,s sql ), q user Represents the question template data, s sql represents the SQL template data corresponding to the question template data; Step 3-1-2: Based on the seed data and financial database, the teacher model f teacher Generate new user question template data q' user and the corresponding SQL template data s' sql : (q’ user ,s’ sql )=f teacher (q user ,s sql ) Repeat several times to obtain the augmented template dataset S struct , template datasets struct Each question-answer pair sample consists of user question template data and corresponding SQL template data; Step 3-1-3: For the template data set S struct Use the greedy subset selection algorithm to obtain the data subset; Step 3-1-4: Fill in the fields of each question template and SQL template of the data subset obtained in step 3-1-3, and use a synonym dictionary to perturb the filling of the question template to obtain a structured data set.
5. The method for constructing a low-resource financial annual report large model question-answering system according to claim 1, characterized in that: In step 3, the specific process of generating a semi-structured data set is as follows: Step 3-2-1: Randomly select chapter segments other than financial data in the financial annual report, and let the teacher model f teacher Generate documents and question-answer pairs according to the chapter fragments, repeat the generation process several times, and the obtained documents and question-answer pairs constitute an initial data set; Step 3-2-2, define a disturbance operator, and perform disturbance enhancement on the initial data set to obtain a disturbance enhanced data set; Step 3-2-3: Use a greedy subset selection algorithm on the perturbation-enhanced dataset to obtain a subset with higher diversity than the perturbation-enhanced dataset as a semi-structured dataset.
6. A method for constructing a low-resource financial annual report large model question-answering system according to claim 4 or 5, characterized in that: The greedy subset selection algorithm described in step 3-1-3 or step 3-2-3 is specifically: Step 3-1: Set the target subset to S greedy , set its target size to k; Step 3-2: Randomly select an initial sample s from the input data set S i , initialize the target subset S greedy ; Step 3-3, calculate S greedy The vector center c struct , from the input data set S excluding the target subset S greedy Select the sample with the lowest cosine similarity cosine_sunukarity from the part of s j Add to target subset: Step 3-4: Repeat step 3-3 until the size of the target subset reaches k; The input data set S is equal to the template data set S in step 3-1-2 struct Or the perturbation-enhanced data in step 3-2-2.
7. A method for constructing a low-resource financial annual report large model question-answering system according to claim 6, characterized in that: The specific steps of performing disturbance enhancement by the disturbance operator described in step 3-2-2 are as follows: Step 3-2-2-1, define the disturbance operation; Step 3-2-2-2, preset a prompt word for each disturbance operation; Step 3-2-2-3, Teacher Model f teacher The table content is rewritten according to the requirements of the prompt words to obtain a perturbation-enhanced dataset.
8. A method for constructing a low-resource financial annual report large model question-answering system according to claim 1 or 2, characterized in that: In step 4, the specific process of the structured question-answering module is as follows: Step 4-1-1: Use the structured question-answering model to generate SQL statements for queries based on user questions using NL2SQL; Step 4-1-2, execute the generated SQL statement in the financial database to obtain a table of query results; Step 4-1-3: Convert the obtained table into text and use it as context for the structured question-answering model to generate answers.
9. A method for constructing a low-resource financial annual report large model question-answering system according to claim 1 or 2, characterized in that: In step 4, the specific process of the semi-structured question-answering module is as follows: Step 4-2-1: For each company name-year pair p in the company name-year binary list P i (p i ∈P) to form a string sp i , calculate the string sp i bm25 similarity with metadata in the document tree in the vector database; Step 4-2-2: Based on the bm25 similarity, recall the top x annual report document tree lists with the highest similarity L = {T1, T2, ..., T x }, x is greater than or equal to 1 and less than or equal to the number of document trees in the vector database; Step 4-2-3: The base model f selects the document tree T that best meets the target from the document tree list L top ; Step 4-2-4, use the base model f from the document tree T top Filter out the most relevant chapter subsets from the sub-chapter collection under the first-level chapter: Step 4-2-5: for each text block in the most relevant chapter subset, retrieve all corresponding text block vectors from the vector database; Step 4-2-6: Answer user questions user Using the embedding model f e Encode to get f e (q user ), calculate and each text block d i The corresponding text block vector vector i Similarity: Take the text blocks corresponding to the text block vectors of the top y names in vector similarity as the recalled document set D r , y takes a value between 5 and 20, not exceeding the total number of document blocks; Step 4-2-7: Use the semi-structured question answering model semi-struct Generate answer a; a=f semi-struct (D r ,q user )。 10. The method for constructing a low-resource financial annual report large model question-answering system according to claim 1, characterized in that: The financial data described in step 2 include the consolidated balance sheet, consolidated cash flow statement and consolidated income statement.