Financial statement balance verification automation method and device based on large language model

By adopting a new table representation format and two-stage numerical verification method in financial statement calibration, combined with a single-sample learning method of multi-level retrieval, the problem of inaccurate numerical calculation and complex table structure understanding in financial statement data calibration is solved, and efficient and accurate financial statement verification sum is achieved to reduce the risk of missed inspection.

CN119940299APending Publication Date: 2025-05-06UNIV OF SCI & TECH OF CHINA
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510061476.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-15
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

The prior art has problems in the verification of financial statement data such as inaccurate numerical calculations, inability to understand complex table structures, lack of field-specific adaptability and relying on large-scale training data, resulting in low verification efficiency and high risk of missed inspection.

Method used

A financial statement balanced verification automation method based on a large language model is adopted, and the model can better understand the table structure and semantics by designing a new tabular representation format. The two-stage numerical verification method is used to ensure accuracy, and the single-sample learning method of multi-level search is reduced to the demand for large-scale annotation data sets, and the knowledge base is dynamically updated to identify unseen abnormal situations.

Benefits of technology

It improves the accuracy of table semantic understanding and numerical verification, realizes efficient field adaptability and automated report generation, reduces the risk of missed inspection, and improves the efficiency and accuracy of financial statement audits.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119940299A_ABST
    Figure CN119940299A_ABST
Patent Text Reader

Abstract

The invention provides a financial statement balance verification automation method based on a large language model. The method comprises the steps that S1, a table custom format conversion module is constructed; s2, constructing a typical sample library; s3, constructing a table verification information generation module; s4, connecting character calculators in series; step S5, generating a verification report; and S6, carrying out confidence coefficient test and dynamically updating the typical sample library. According to the method, the semantic understanding and numerical value verification accuracy of the table can be improved, efficient field adaptability is achieved, automatic report generation is achieved, user friendliness is achieved, and the enterprise management efficiency can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of artificial intelligence technology, and in particular relates to an automated method and device for financial statement balance verification based on a large language model. Background Art

[0002] Enterprises provide accounting information related to the financial status, operating results and cash flow of enterprises to users of financial accounting reports through financial reports. It reflects the overall operating conditions of enterprises and is also an important basis for supporting enterprises to conduct financial analysis and management decisions. Traditional financial reports rely on manual completion, which is prone to problems such as long preparation time, untimely information reflection, insufficient financial data and artificial cover-up, resulting in hidden dangers in the authenticity, integrity and timeliness of financial reports.

[0003] At present, the financial management of enterprises has gone through the development stage from computerization to informatization. With the emergence and gradual maturity of information technologies such as big data, cloud computing, and artificial intelligence, financial reporting has also ushered in the opportunity for automation and intelligent transformation. Among the emerging tools, RPA (Robotic Process Automation) has been widely used in the financial field to realize the automation of multiple financial activities such as cash expenditure, accounts receivable management, financial control, financial reporting, taxation, financial planning, risk management, and auditing.

[0004] In the fields of business management, financial statements, audit reports, etc., tabular data is the main carrier of information management and analysis. As the amount of tabular data processed by enterprises and organizations continues to increase, how to efficiently process this data has become a key challenge in modern information processing systems, especially in the financial field, where a large number of tabular values ​​need to be calculated and verified through accounting rules.

[0005] The existing schemes for financial statement data verification mainly include three categories: manual operation, rule-based verification and LLM-based verification methods. At present, the numerical verification in accounting tables mostly relies on traditional rule-based systems or manual reconciliation. The common practice is to use set verification rules, such as numerical range restrictions, formula verification (such as debit and credit balance), document verification, etc., combined with written scripts or financial management software to automate some verification work. Traditional financial data processing and auditing work mostly rely on manual operations, which are inefficient and prone to errors. Before the emergence of Large Language Models (LLMs), researchers have tried to combine tabular data with neural networks for tasks such as NLP and data management. With the development of intelligence, LLM-based methods have gradually replaced inefficient manual verification and traditional rigid rule verification and become a new type of table verification method. LLM automatically extracts key indicators from financial statements and verifies them according to the logical relationship of values ​​in the table. This method can not only significantly improve the efficiency of financial data processing, but also reduce human errors, thereby reducing operational risks, saving a lot of labor costs, improving the accuracy and efficiency of data analysis, and promoting the intelligent process of financial work.

[0006] Traditional rule-based methods have weak intelligence and adaptability when dealing with sudden data errors or outliers, and can often only identify errors within a preset range. They are ineffective for verification that requires contextual semantics (such as correlation verification between tables) or cross-validation of highly complex data. These methods require financial personnel to clearly define verification rules, and the rules have different applicability in different tables and scenarios. For situations with high diversity in table structures, traditional rule-based verification systems may not be able to flexibly deal with complex data structures and implicit errors, resulting in increased manual verification workload and low efficiency. The large language model (LLM)-based method shows great potential due to the excellent capabilities of LLM in all aspects, but it still has the following defects:

[0007] (1) Inaccurate numerical calculation: Although the existing LLM technology has a certain effect in processing tabular data, it is essentially a probabilistic model and cannot perform accurate numerical calculations, which affects the final verification results.

[0008] (2) Unable to understand complex table structures: Complex financial tables often contain hierarchical nesting, cross-row and cross-column data structures and strict numerical calculation logic. Large language models naturally have excellent understanding capabilities for unstructured data, but cannot understand complex table structures such as structured data such as tables.

[0009] (3) Lack of domain-specific adaptability: Existing LLM models are usually generic and lack expertise in specific financial fields, which makes them difficult to understand when processing relevant tabular data involving specific names and rules, thus limiting their effectiveness.

[0010] (4) Reliance on large-scale training data: The current LLM model requires a large amount of training data for pre-training and fine-tuning. In the financial field, the types of tables are diverse, and the labeling cost of the data set is high and it is difficult to cover tables of various categories and styles. It is also unable to handle unseen abnormal situations and it is difficult to effectively identify those abnormal situations that have not been seen in the training stage, which may lead to the omission of potential financial errors. Summary of the invention

[0011] Deeper semantic understanding of tables: This paper designs a new table representation format that enables LLMs to better understand the structure and semantic relationship of the table, which is better than the existing HTML format. It enables the model to accurately extract and analyze key indicators in financial tables, which is the basis for table value verification and error detection.

[0012] Multi-precision numerical verification method: This paper proposes a two-stage method: in the first stage, a large model is used to analyze key indicators and extract relevant values; in the second stage, an external calculator is introduced in the numerical calculation stage to ensure the accuracy of the verification results and avoid the shortcomings of the large model in calculation. Multi-precision verification is designed to meet the needs of different scenarios.

[0013] Efficient and versatile domain adaptation method: This paper designs a single-sample learning method based on multi-level retrieval to build a typical knowledge base, reduce the need for large-scale annotated data sets, and avoid long-term fine-tuning processes. This makes large models more effective in processing diverse financial tables. At the same time, this method has better generalization performance, can identify unseen anomalies during use, dynamically add new samples to the knowledge base when the confidence level is lower than the threshold, and continuously update, thereby improving the ability to control missed detection risks.

[0014] In summary, the present invention adopts the following technical solution: an automated method for balance verification of financial statements based on a large language model, comprising the following steps:

[0015] Step S1: construct a table custom format conversion module; used to convert the table format into a clearly structured format that can be understood by a large language model;

[0016] Step S2: construct a typical sample library; used to provide table samples that are most similar to the input table features as references, and use a large language model to generate verification information for the table samples;

[0017] Step S3: construct a table verification information generation module; used to generate verification information for the input table, including verification rules and calculation formulas for each checkpoint;

[0018] Step S4, a serial character calculator; used to execute the verification rules and formula calculations for each checkpoint involved in the table verification process;

[0019] Step S5, generating a verification report; used to generate a verification report, the verification report includes the inspection results of each inspection point;

[0020] Step S6: Confidence test and dynamic update of typical sample library.

[0021] The present invention also proposes an automated device for financial statement balance verification based on a large language model, comprising the following modules:

[0022] The table custom format conversion module is used to convert the table format into a clearly structured format that can be understood by the large language model;

[0023] The typical sample library is used to provide table samples that are most similar to the input table features as references, and use the large language model to generate verification information for the table samples;

[0024] A table verification information generation module is used to generate verification information for the input table, including verification rules and calculation formulas for each checkpoint;

[0025] A concatenated character calculator for executing the verification rules and formula calculations for each checkpoint involved in the form verification process;

[0026] A verification report generation module is used to generate a verification report, which contains the inspection results of each inspection point;

[0027] Confidence testing and updating module, used for confidence testing and dynamic updating of typical sample library.

[0028] The present invention has the following beneficial effects:

[0029] 1. Improve the semantic understanding and numerical verification accuracy of tables: By adopting a newly designed table representation format, the large language model (LLM) can deeply understand the structure of the table and its internal semantic relationships. This method is superior to using the traditional HTML format, allowing the model to accurately extract and analyze key indicators in financial tables, thus laying a solid foundation for numerical verification and error detection. At the same time, the two-stage numerical verification method of the present invention ensures high accuracy: in the first stage, the large model is used to analyze key indicators and extract relevant numerical values, and in the second stage, an external Python calculator is introduced to perform specific numerical calculations, avoiding the limitations of the large language model in mathematical calculations.

[0030] 2. Efficient domain adaptability and automated report generation: The single-sample learning method for multi-level retrieval proposed in this invention effectively constructs a typical knowledge base, reduces the need for large-scale annotated data sets, and makes large models more efficient when processing diverse financial tables. At the same time, this method has good generalization performance, can identify unseen anomalies and dynamically update the knowledge base, and improve the ability to control the risk of missed detection. In addition, the system can automatically generate a detailed verification report, clearly indicating the location of potential errors and their exact values, greatly improving the efficiency of financial personnel in discovering and solving problems.

[0031] 3. User-friendliness and improved enterprise management efficiency: The intelligent process design of the present invention allows users to provide only table data, and the system can automatically analyze and verify, which reduces the operation threshold. Combined with the above advantages, this solution will effectively improve the efficiency of enterprises in financial statement audits, reduce errors and delays caused by manual audits, and ultimately achieve more efficient financial management and decision support. BRIEF DESCRIPTION OF THE DRAWINGS

[0032] Figure 1 It is a flowchart of the steps of the automated method for balance verification of financial statements based on a large language model;

[0033] Figure 2 It is a flowchart of the steps to build a typical sample library;

[0034] Figure 3 is a schematic diagram of a verification information generation module;

[0035] Figure 4 It is a flow chart of confidence test. DETAILED DESCRIPTION

[0036] In order to make the purpose, technical scheme and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention. In addition, the technical features involved in each embodiment of the present invention described below can be combined with each other as long as they do not conflict with each other. To achieve the above-mentioned purpose, the present invention adopts the following technical scheme.

[0037] See also Figure 1 As shown, the present invention is an automated method for balance verification of financial statements based on a large language model, comprising the following steps:

[0038] Step S1, constructing a table custom format conversion module;

[0039] It is used to convert a table in HTML text format into a custom JSON format. The custom JSON format is described as follows, and includes two main fields:

[0040] data: stores the specific content of each row of data, the format is data = [{column_1: value_1,column_2: value_2}, ...].

[0041] mergeCells: describes the starting and ending positions of the merged cells and their values. The format is mergeCells = [{start: cell_1, end: cell_2, value: merged_value}].

[0042] Step S2, constructing a typical sample library;

[0043] Specifically, Figure 2 As shown, including:

[0044] Step S201, extracting typical table samples; including: sorting out 15 common table categories, and selecting a table from each category;

[0045] Step S202, converting the selected table into a custom table format to obtain basic table information;

[0046] Step S203, using the large language model to generate table verification information;

[0047] Step S204, manual verification, proofreading and correction;

[0048] Step S205: adding a typical sample library.

[0049] This embodiment sorts out 15 common categories from the large number of tables that need to be verified in the existing business processes. Select a table from each category and add it to the sample library. Each table data in the sample library contains two tags, including basic information and verification information, as follows:

[0050] (1) Basic information;

[0051] File Information: Table file name.

[0052] Field structure:

[0053] Row Names: A list of specific names or categories representing each row in the table.

[0054] Column Names: A list representing the specific names or categories of each column in the table.

[0055] Data content: All data in the table is stored in JSON format. The content structure is a table of rows and columns, and each item corresponds to a specific value.

[0056] HTML format: HTML representation of the table for further visualization.

[0057] (2) Verification information;

[0058] There are two types of checkpoints: horizontal checkpoints (i.e., indicators that need to be verified by the relationship between columns and fields) and vertical checkpoints (i.e., indicators that need to be verified by the relationship between rows and fields). Each checkpoint contains the following information:

[0059] Name: row name + column name where the value point is located;

[0060] Position: row and column number;

[0061] Values: Raw values ​​disclosed in the table;

[0062] Sub-item value: used to calculate the sub-items and their values ​​required for this checkpoint;

[0063] Calculation formula: A calculation formula in string format containing specific values.

[0064] The verification information of the tables in the sample library is generated by the large language model (LLM) and manually proofread for correct labels.

[0065] Step S3, constructing a table verification information generation module;

[0066] Specifically, Figure 3 As shown in this embodiment, the specific construction process is as follows:

[0067] Step S301, input table format conversion;

[0068] Use the table custom format conversion module constructed in step S1 to convert the input HTML table into a custom JSON format;

[0069] Step S302, using a hierarchical retrieval algorithm to retrieve learning samples;

[0070] Given an input table, by calculating the similarity score between the basic information of the input table and each table sample in the typical sample library, and sorting the similarity scores from high to low, find the table sample in the typical sample library that is most similar to the input table in structure and content as a learning sample. The learning sample is used to embed the prompt word together with the input table in step S303, and input the large language model to generate verification information of the input table. The designed hierarchical retrieval algorithm extracts and uses the features of three levels of the table to calculate the similarity score between the input table and similar tables. Including a. Calculate the cosine similarity of the row field name ; b. Calculate the cosine similarity between column field names ; c. Cosine similarity of descriptive sentences containing both row fields and column fields . Define formula (1), cosine similarity The calculation formula is:

[0071] (1)

[0072] in, and Respectively represent the i-th value represented by the vector of the corresponding feature in the two tables, Represents the feature vector and the eigenvector The cosine similarity between .

[0073] Define formula (2), similarity score The weighted calculation formula is as follows:

[0074] (2)

[0075] ;

[0076] in A custom weight.

[0077] right The calculation steps are as follows:

[0078] Score calculation: Use the Chinese embedding model bge-m3 to convert the row field names of each table sample in the input table and the typical sample library into numerical vector representations, where the row field name vector representation of the input table is recorded as ; Typical sample library The row field name vector representation of a table sample is recorded as ; Use formula (1) to calculate the input table and the typical sample library The similarity scores of table samples on this feature dimension ;

[0079] Score calculation: Use the Chinese embedding model bge-m3 to convert the column field names of each table sample in the input table and the typical sample library into numerical vector representations, where the column field name vector representation of the input table is recorded as ; Typical sample library The column field name vector representation of the table sample is recorded as ; Use formula (1) to calculate the input table and the typical sample library The similarity scores of table samples on this feature dimension ;

[0080] Score calculation: Use the Chinese embedding model bge-m3 to convert the descriptive statements of each table sample in the input table and the typical sample library into numerical vector representations, where the descriptive statement vector representation of the input table is recorded as ; Typical sample library The descriptive sentence vector representation of the table sample is recorded as ; Use formula (1) to calculate the input table and the typical sample library The similarity scores of table samples on this feature dimension ;

[0081] Use formula (2) to calculate the final similarity score of each table sample in the typical sample library. The final similarity score of the table samples is calculated as:

[0082] ;

[0083] By combining the metrics of the three levels of features, the table sample with the highest final similarity score is obtained as the retrieved learning sample, and its similarity score is recorded as .

[0084] Step S303, prompt word design and embedding;

[0085] Prompt refers to the prompt input to the large language model, which is intended to guide the model to generate the target output. The goal of prompt design is to help the model understand how to process the input data and generate output that meets the requirements. The designed prompt consists of four parts: format description, input table, output requirements, and learning sample. The prompt is obtained by splicing these four parts, as shown below:

[0086] ;

[0087] Specifically, the first component of the prompt word, Format discription, is a specific explanation of the custom table format, which is intended to describe the structure, type, and format of the input table. It helps the model understand the layout of the table data and how to parse and process the data.

[0088] Input table is the input table after format conversion in step S301; it is embedded in the second part of the prompt word as input.

[0089] The third part of the prompt word Requirements is the specific requirements for the table inspection information that needs to be generated, including the format and value type, to help the model clarify the task objectives, generated content, and output specifications.

[0090] Example is a learning sample retrieved in step S302. It contains manually verified table verification information, shows an actual table data and the corresponding expected output, and helps the model understand the specific content and format of the task. The purpose is to guide the model to understand the task by providing a complete example.

[0091] Step S304, use the large language model to generate verification information. Input the designed prompt word (prompt) into the large language model (LLM), and use the large language model to generate verification information of the table that is consistent with the verification information label format in the sample library. Specifically, the process includes position identification and formula extraction. Position identification is to determine the position (row name + column name) of the value to be verified. Formula extraction is to extract the relevant calculation formula based on the identified checkpoints. The formula contains the fields used to calculate the value and the logical relationship between them, and the formula is converted into an executable string form.

[0092] Step S4, serial character calculator;

[0093] The serial character calculator performs specific numerical calculations and result comparisons for each checkpoint extracted from the table through the embedded Python calculator. First, a Python serial character calculator is constructed, and then the Python serial character calculator is used to calculate the calculation formula of each checkpoint in the verification information obtained in step S3 to obtain the precise numerical result after calculation, and the calculation result is added to the verification information; the calculation result is compared with the original numerical value in the table. If the numerical difference is less than the error threshold margin_for_error, it is marked as "correct", otherwise it is marked as "wrong", and the comparison result is also added to the verification information. Users can adjust the verification accuracy by modifying the error threshold parameter margin_for_error (the default value is 0.01) according to the business needs of different scenarios. The table data that records the calculation results and comparison results in step S4 is recorded as Checked Data. Compared with the verification information defined in step S2, the verification information structure of Checked Data is as follows:

[0094] Verify information;

[0095] There are two types of checkpoints: horizontal checkpoints (i.e., indicators that need to be verified by the relationship between columns and fields) and vertical checkpoints (i.e., indicators that need to be verified by the relationship between rows and fields). Each checkpoint contains the following information:

[0096] Name: row name + column name where the value point is located;

[0097] Position: row and column number;

[0098] Values: Raw values ​​disclosed in the table;

[0099] Sub-item value: used to calculate the sub-items and their values ​​required for this checkpoint;

[0100] Calculation formula: a calculation formula containing specific values ​​in string format;

[0101] Calculation result: the numerical result after the calculation formula is calculated;

[0102] Comparison result: The result of comparing "Calculation result" and "Value", the value is "Correct" or "Incorrect".

[0103] Step S5, generating a verification report;

[0104] The numerical comparison results are input into the large language model to generate a descriptive table verification report. The report indicates the numerical values ​​and their exact values ​​of the locations with potential errors in all verified tables. That is:

[0105] ;

[0106] The formula parameters and operators are defined as follows:

[0107] Report: The final generated table verification report contains all verified table information, potential error values ​​and their correction values. The report is a structured text designed to help users quickly locate and correct errors.

[0108] LLM: refers to the Large Language Model (here we use GPT-4o from the GPT series), which is used to generate verification reports described in natural language. The input content of the large language model is in brackets. The model accepts verified data as input and automatically generates business-related, easy-to-understand error descriptions and correction suggestions;

[0109] Checked Data: refers to the table data obtained in step S4 that records the numerical calculation results and comparison results.

[0110] Step S6, confidence test and dynamic update of typical sample library;

[0111] Specifically, Figure 4 As shown. Includes:

[0112] Step S601, setting the search confidence threshold to p; wherein the value of p ranges from 0 to 1;

[0113] Step S602, checking the similarity score obtained in step S302 Is it lower than the threshold value p? If it is lower than the threshold value p, then execute step S603; otherwise, end the entire step S6;

[0114] Step S603, asking whether to add the new sample to the sample library; if the user chooses yes, then step S604 is executed; otherwise, the entire step S6 is terminated;

[0115] Step S604, the user inputs the manually corrected form verification information;

[0116] Step S605, checking whether the user's input format is correct; if the format is correct, executing step S606; if the format is incorrect, returning to step S604, re-entering the form verification information;

[0117] Step S606: Add the new sample including the basic information and the verification information to the typical sample library.

[0118] The search confidence threshold is set to p, and the similarity score obtained in step S302, ie, SimilarityScore, is checked. If it is lower than the threshold, the confidence in the result is low, and it is inquired whether to add the table to the sample library to dynamically update the typical sample library.

Claims

1. An automated method for financial statement balance verification based on a large language model, characterized in that: The steps include: Step S1: construct a table custom format conversion module; used to convert the format of the table into a clearly structured format that can be understood by the large language model; Step S2: construct a typical sample library; used to provide table samples that are most similar to the input table features as references, and use a large language model to generate verification information for the table samples; Step S3: construct a table verification information generation module; used to generate verification information for the input table, including verification rules and calculation formulas for each checkpoint; Step S4, a serial character calculator; used to execute the verification rules and formula calculations for each checkpoint involved in the table verification process; Step S5, generating a verification report; used to generate a verification report, the verification report includes the inspection results of each inspection point; Step S6: Confidence test and dynamic update of typical sample library.

2. The method for automating financial statement balance verification based on a large language model according to claim 1, characterized in that: Step S1 includes: The table custom format conversion module is used to convert the table in HTML text format into a custom JSON format. The custom JSON format is described as follows and includes two fields: data: stores the specific content of each row of data in the format of data = [{column_1: value_1, column_2: value_2}, ...]; mergeCells: describes the starting and ending positions of the merged cells and their values. The format is mergeCells = [{start: cell_1, end: cell_2, value: merged_value}].

3. The method for automating financial statement balance verification based on a large language model according to claim 1, characterized in that: Step S2 includes: Step S201, extracting typical table samples; Step S202, converting the selected table into a custom table format to obtain basic table information; Step S203, using the large language model to generate table verification information; Step S204, manual verification, proofreading and correction; Step S205: adding a typical sample library.

4. The method for automating financial statement balance verification based on a large language model according to claim 3, characterized in that: Step S201 specifically includes: sorting out typical category tables from a large number of tables, selecting a table from each category as a table sample, and adding it to a typical sample library.

5. The method for automating financial statement balance verification based on a large language model according to claim 4, characterized in that: Basic information and verification information are as follows: Basic information includes: file information, field structure, data content and HTML format; file information includes: table file name; field structure includes: row name: a list of specific names or categories for each row in the table; column name: a list of specific names or categories for each column in the table; data content includes: all data in the table is stored in JSON format, and the content structure is a table form of rows and columns, and each item corresponds to a specific value; HTML format includes: HTML representation of the table for further visualization; Verification information includes two types: horizontal checkpoints, which are indicators verified by the relationship between columns and vertical checkpoints, which are indicators verified by the relationship between rows. Each checkpoint contains the following information: Name: row name + column name where the value point is located; Position: row and column number; Values: Raw values ​​disclosed in the table; Sub-item value: used to calculate the sub-items and their values ​​required for this checkpoint; Calculation formula: a calculation formula containing specific values ​​in string format; The verification information of the tables in the typical sample library is generated by LLM and manually proofread for correct labels.

6. The method for automating financial statement balance verification based on a large language model according to claim 5, characterized in that: Step 3 The specific construction process is as follows: Step S301, input table format conversion; the table custom format conversion module converts the input HTML table into a custom JSON format; Step S302, using a hierarchical retrieval algorithm to retrieve learning samples; given an input table, by calculating the similarity score between the basic information of the input table and each table sample in the typical sample library, and sorting the similarity scores from high to low, find the table sample in the typical sample library that is most similar to the input table in structure and content as the learning sample, and use the hierarchical retrieval algorithm to extract and use the features of the three levels of the table to calculate the similarity score between the input table and similar tables; Including a. Calculate the cosine similarity of row field names ; b. Calculate the cosine similarity between column field names ; c. Cosine similarity of descriptive sentences containing both row and column fields ; Definition formula (1), cosine similarity The calculation formula is: (1) in, and Respectively represent the i-th value represented by the vector of the corresponding feature in the two tables, Represents the feature vector and the eigenvector The cosine similarity between ; Define formula (2), similarity score The weighted calculation formula is as follows: (2) ; in, For custom weights; right The calculation steps are as follows: Score calculation: Use the Chinese embedding model bge-m3 to convert the row field names of each table sample in the input table and the typical sample library into numerical vector representations, where the row field name vector representation of the input table is recorded as ; Typical sample library The row field name vector representation of a table sample is recorded as ; Use formula (1) to calculate the input table and the typical sample library The similarity scores of the tables on this feature dimension ; Score calculation: Use the Chinese embedding model bge-m3 to convert the column field names of each table sample in the input table and the typical sample library into numerical vector representations, where the column field name vector representation of the input table is recorded as ; Typical sample library The column field name vector representation of the table sample is recorded as ; Use formula (1) to calculate the input table and the typical sample library The similarity scores of table samples on this feature dimension ; Score calculation: Use the Chinese embedding model bge-m3 to convert the descriptive statements of each table sample in the input table and the typical sample library into numerical vector representations, where the descriptive statement vector representation of the input table is recorded as ; Typical sample library The descriptive sentence vector representation of the table sample is recorded as ; Use formula (1) to calculate the input table and the typical sample library The similarity scores of table samples on this feature dimension ; Use formula (2) to calculate the final similarity score of each table sample in the typical sample library. The final similarity score of the table samples is calculated as: ; By combining the metrics of the three levels of features, the table sample with the highest final similarity score is obtained as the retrieved learning sample, and its similarity score is recorded as ; Step S303, prompt word design and embedding; The prompt word Prompt refers to the prompt input to the large language model, which is intended to guide the large language model to generate the target output. The prompt word is used to help the large language model understand how to process the input data and generate output that meets the requirements. The prompt word consists of four parts: format description Format discription, input table Input table, output requirements Requirements and learning sample Example. These four parts are concatenated to obtain the prompt word, which is expressed as follows: ; Format description is a specific explanation of the custom table format, which aims to describe the structure, type and format of the input table. It helps the large language model understand the layout of the table data and how to parse and process the data. Input table is the input table after format conversion; Output requirements are the specific requirements of the table inspection information to be generated, including the format and value type, to help the big prediction model clarify the task objectives, generated content, and output specifications; The learning sample Example is the learning sample retrieved in step S302. It contains manually verified table verification information, displays an actual table data and the corresponding expected output, helps the large language model understand the specific content and format of the task, and guides the large language model to understand the task by providing a complete example; Step S304: Use the large language model to generate verification information, input the designed prompt word prompt into the large language model, and use the large language model to generate verification information of a table that is consistent with the verification information label format in the sample library; specifically, the process includes position identification and formula extraction, position identification is to determine the position of the value to be verified, that is, the row name + column name, formula extraction is to extract the relevant calculation formula according to the identified checkpoint, the formula includes the fields used to calculate the value and the logical relationship between them, and the formula is converted into an executable string form.

7. The method for automating financial statement balance verification based on a large language model according to claim 6, characterized in that: Step S4 includes: The serial character calculator performs specific numerical calculations and result comparisons on each checkpoint extracted from the table through the embedded Python calculator. First, a Python serial character calculator is constructed, and then the Python serial character calculator is used to calculate the calculation formula of each checkpoint in the verification information obtained in step S3 to obtain the precise numerical result after calculation, and the calculation result is added to the verification information; the calculation result is compared with the original numerical value in the table. If the numerical difference is less than the error threshold margin_for_error, it is marked as "correct", otherwise it is marked as "wrong", and the comparison result is also added to the verification information. The table data that records the calculation results and comparison results in step S4 is recorded as CheckedData. The verification information structure of Checked Data is as follows: Verification information; includes two types, horizontal checkpoints and vertical checkpoints, each of which contains the following information: Name: row name + column name where the value point is located; Position: row and column number; Values: Raw values ​​disclosed in the table; Sub-item value: used to calculate the sub-items and their values ​​required for this checkpoint; Calculation formula: a calculation formula containing specific values ​​in string format; Calculation result: the numerical result after the calculation formula is calculated; Comparison result: The result of comparing "Calculation result" and "Value", the value is "Correct" or "Incorrect".

8. The method for automating financial statement balance verification based on a large language model according to claim 7, characterized in that: Step S5 includes: Input the Checked Data into the large language model to generate a descriptive table verification report, which indicates the numerical values ​​and accurate values ​​of the locations with potential errors in all verified tables, namely: ; The formula parameters and operators are defined as follows: Report: The final generated table verification report contains all verified table information, potential error values ​​and their correction values. The report is a structured text that helps users quickly locate and correct errors. LLM: refers to the large language model used to generate verification reports described in natural language. The large language model accepts the verified data as input and automatically generates business-related, easy-to-understand error descriptions and correction suggestions; Checked Data: refers to the table data recording the numerical calculation results and comparison results obtained in step S4.

9. The method for automating financial statement balance verification based on a large language model according to claim 7, characterized in that: Step S6 includes: Step S601, setting the search confidence threshold to p; wherein the value of p ranges from 0 to 1; Step S602, checking the similarity score obtained in step S302 Is it lower than the threshold value p? If it is lower than the threshold value p, then execute step S603; otherwise, end the entire step S6; Step S603, asking whether to add the new sample to the sample library; if the user chooses yes, then step S604 is executed; otherwise, the entire step S6 is terminated; Step S604, the user inputs the manually corrected form verification information; Step S605, checking whether the user's input format is correct; if the format is correct, executing step S606; if the format is incorrect, returning to step S604, re-entering the form verification information; Step S606, adding the new sample containing basic information and verification information to the typical sample library; The search confidence threshold is set to p, and the similarity score obtained in step S302 is checked. If it is lower than the threshold, the confidence in the result is low, and it is inquired whether to add the table to the sample library to dynamically update the typical sample library.

10. An automated device for financial statement balance verification based on a large language model, characterized in that: Includes the following modules: The table custom format conversion module is used to convert the table format into a clearly structured format that can be understood by the large language model; The typical sample library is used to provide table samples that are most similar to the input table features as references, and use the large language model to generate verification information for the table samples; A table verification information generation module is used to generate verification information for the input table, including verification rules and calculation formulas for each checkpoint; A concatenated character calculator for executing the verification rules and formula calculations for each checkpoint involved in the form verification process; A verification report generation module is used to generate a verification report, which contains the inspection results of each inspection point; Confidence testing and updating module, used for confidence testing and dynamic updating of typical sample library.

Citation Information

Cited By

  • Commission settlement and reconciliation method and device for social e-commerce, computer equipment and storage medium

    CN120707206A

  • Financial data cleaning and verification system and method

    CN121116959A