Financial report element extraction method and device, equipment and storage medium
By screening key pages of financial reports and automatically extracting financial report elements using a multimodal large model, the problems of low efficiency, poor accuracy, and high cost in existing technologies have been solved, achieving efficient and accurate extraction of financial report elements.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-27
- Publication Date
- 2026-03-17
AI Technical Summary
Existing technologies for extracting financial reporting elements are inefficient, inaccurate, and costly, mainly due to their reliance on manual processing.
By selecting key pages of financial reports, multimodal large models and preset prompt word templates are used to generate element extraction instructions, and character recognition and semantic understanding are combined to automatically extract financial report elements.
It improves the efficiency and accuracy of extracting financial report elements, reduces labor costs, and achieves automated processing without human intervention.
Smart Images

Figure CN121686500A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of financial technology, and in particular to a financial report element extraction method, device, equipment and storage medium. BACKGROUND
[0002] The enterprise financial report contains corresponding financial report elements. The bank can analyze the financial report elements in the enterprise financial report to assess the credit risk of the enterprise (solvency analysis, profitability verification, cash flow health monitoring, etc.), loan approval and pricing (credit limit calculation, interest rate pricing basis, etc.), post-loan management and risk warning (financial indicator monitoring, industry comparison analysis) and other core business activities. Therefore, after obtaining the enterprise financial report, the financial report elements in the enterprise financial report need to be extracted first.
[0003] Currently, the bank extracts financial report elements in the following way: the enterprise submits the financial report through offline (paper file) or online (email, bank system upload), usually in the form of pictures (JPG, JPEG, PNG, etc.) or PDF format. The bank's customer manager or credit officer checks whether the financial report is complete and audited, and preliminarily judges the authenticity of the financial report. Then, the key elements are extracted from the financial report manually, such as: obtaining cash, accounts receivable, inventory, fixed assets, etc. from the balance sheet; obtaining operating income, operating cost, net profit, etc. from the profit statement.
[0004] However, first of all, the efficiency of manually extracting financial report elements is low. It may take 1-2 days to complete the processing of a financial report, and multiple posts are designed to collaborate. At the same time, a large amount of manpower is occupied by repetitive work (such as data entry); secondly, the accuracy of manually extracting financial report elements is poor. Manual entry is prone to errors (such as number of digits, unit confusion), and the identification of financial report abnormalities relies on experience and is highly subjective; finally, the cost of manually extracting financial report elements is high. It is necessary to equip customer managers or outsourcing teams with professional financial backgrounds, and the proportion of human cost is high. SUMMARY
[0005] In order to improve the efficiency and accuracy of financial report element extraction and reduce the cost, the present application provides a financial report element extraction method, device, equipment and storage medium.
[0006] In a first aspect, the present application provides a financial report element extraction method, comprising:
[0007] Screening key pages from the original financial report, the key pages at least including one of balance sheet and profit statement;
[0008] determine an element extraction instruction based on the financial statement header of the key page and a preset prompt word template; wherein the prompt word template comprises a to-be-filled area for filling a key element of the financial statement corresponding to the financial statement header;
[0009] determine a character recognition sequence based on the key page, the element extraction instruction, a preset model parameter, and a multi-modal large model; wherein the multi-modal large model is a model obtained by using a preset historical financial report as training data for training;
[0010] perform semantic understanding on the character recognition sequence to determine a target financial report element.
[0011] Through the above implementation, based on screening the key page, the corresponding prompt word (element extraction instruction) can be generated from the key page, and then the multi-modal large model is controlled using the element extraction instruction to recognize the characters (character recognition sequence) in the key page. After that, the character recognition sequence is directly subjected to semantic understanding, and the financial report element can be extracted. This process does not require human intervention, and can effectively improve the efficiency and accuracy of financial report element extraction and reduce costs.
[0012] Preferably, the original financial report at least includes one of an original financial report picture and an original financial report PDF file;
[0013] Correspondingly, the key page is screened from the original financial report, comprising:
[0014] adjusting the resolution of the original financial report picture to a preset resolution, and / or paginating the original financial report PDF file to obtain a target financial report;
[0015] processing the target financial report based on a preset visual model to obtain a key page.
[0016] Through the above implementation, by pre-processing (resolution adjustment and / or pagination) of the original financial report, the original financial report can meet the processing requirements of the visual model, and the accuracy of generating the key page can be improved.
[0017] Preferably, the element extraction instruction is determined based on the financial statement header of the key page and a preset prompt word template, comprising:
[0018] determining a key element of a financial statement corresponding to the financial statement header of the key page;
[0019] determining an element extraction instruction based on the key element of the financial statement and a preset prompt word template.
[0020] By the above implementation, the report type is combined with the prompt word template as the element extraction instruction, and the accuracy of characters generated by the model can be conveniently guided by the element extraction instruction subsequently.
[0021] Preferably, the determining the financial report key element corresponding to the financial report header of the key page comprises:
[0022] Performing header identification on the key page to obtain a financial report header;
[0023] Performing semantic identification on the financial report header to obtain a report type;
[0024] Determining a financial report key element corresponding to the report type.
[0025] By the above implementation, the report type is determined through the header of the key page, and since the header generally contains the report type, the accuracy of the report type determination can be improved.
[0026] Preferably, the semantic understanding of the character recognition sequence to determine a target financial report element comprises:
[0027] Performing semantic understanding on the character recognition sequence to obtain an initial financial report element;
[0028] Correcting the initial financial report element to obtain a target financial report element; wherein the correction at least includes one of table layout recovery and error correction.
[0029] By the above implementation, the initial financial report element generated initially is corrected, and the accuracy of the target financial report element output subsequently can be improved.
[0030] Preferably, the financial report element extraction method provided by the application further comprises:
[0031] Packaging the target financial report element to obtain an element packaging result;
[0032] Sending the element packaging result to a downstream classification layer in a distributed stream processing platform.
[0033] By the above implementation, the target financial report element is packaged and sent to the downstream classification layer in the distributed stream processing platform, and the use convenience of the target financial report element can be improved.
[0034] In a second aspect, the application provides a financial report element extraction device, comprising:
[0035] A data screening module is configured to screen a key page from an original financial report, wherein the key page at least includes one of a balance sheet and a profit statement;
[0036] The instruction generation module is configured to determine an element extraction instruction based on the financial statement header of the key page and a preset prompt word template. The prompt word template includes a to-be-filled area for filling a key element of the financial statement corresponding to the financial statement header.
[0037] The character recognition module is configured to determine a character recognition sequence based on the key page, the element extraction instruction, preset model parameters, and a multi-modal large model. The multi-modal large model is a model obtained by training using preset historical financial reports as training data.
[0038] The semantic understanding module is configured to perform semantic understanding on the character recognition sequence to determine a target financial report element.
[0039] Through the above implementation, on the basis of screening the key page, the corresponding prompt word (element extraction instruction) can be generated from the key page, and then the multi-modal large model is controlled using the element extraction instruction to recognize the characters (character recognition sequence) in the key page. After that, the character recognition sequence is directly subjected to semantic understanding, and the financial report element can be extracted. This process does not require human intervention, and can effectively improve the efficiency and accuracy of the financial report element extraction and reduce the cost.
[0040] In a third aspect, the present application provides a computer device, which comprises a memory and a processor. The memory stores a computer program, and the processor implements the steps of the above method when executing the computer program.
[0041] In a fourth aspect, the present application provides a computer readable storage medium, which stores a computer program. The computer program is executed by a processor to implement the steps of the above method.
[0042] In a fifth aspect, the present application further provides a computer program product. The computer program product comprises a computer program, which is executed by a processor to implement the steps of any of the above method embodiments.
[0043] The financial report element extraction method, device, equipment and storage medium extract the key page from the original financial report, the key page at least includes one of the balance sheet and the profit table; determine the element extraction instruction based on the financial statement header of the key page and the preset prompt word template; wherein the prompt word template includes a to-be-filled area for filling the key elements of the financial statement corresponding to the financial statement header; determine the character recognition sequence based on the key page, the element extraction instruction, the preset model parameter and the multi-modal large model; wherein the multi-modal large model is a model obtained by using the preset historical financial report as training data for training; perform semantic understanding on the character recognition sequence to determine the target financial report element. Through the above implementation, based on the screening of the key page, the corresponding prompt word (element extraction instruction) can be generated through the key page, and then the multi-modal large model is controlled by using the element extraction instruction to recognize the characters (character recognition sequence) in the key page, and then the semantic understanding is directly performed on the character recognition sequence, so that the financial report element can be extracted. The process does not require human participation, can effectively improve the efficiency and accuracy of the financial report element extraction and reduce the cost.
[0044] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present application, nor is it used to limit the scope of the present application. Other features of the present application will become apparent from the following description. BRIEF DESCRIPTION OF DRAWINGS
[0045] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the drawings needed to be used in the embodiments. It should be understood that the following drawings only show some embodiments of the present application, and therefore should not be regarded as limiting the scope. For those skilled in the art, other related drawings can also be obtained without creative labor.
[0046] Figure 1 A financial report element extraction method flow chart provided in an embodiment of the present application;
[0047] Figure 2 A structural schematic diagram of a financial report element extraction device provided in an embodiment of the present application;
[0048] Figure 3 A structural schematic diagram of a computer device provided in an embodiment of the present application;
[0049] Figure 4 An internal structure diagram of a computer readable storage medium provided in an embodiment of the present application. DETAILED DESCRIPTION
[0050] To make the objectives, technical solutions, and advantages of this disclosure clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this disclosure.
[0051] It should be noted that the terms "first," "second," etc., used in the specification, claims, and accompanying drawings herein are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, apparatus, product, or device that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or devices.
[0052] In this article, the term "and / or" is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, and B alone. Additionally, the character " / " in this article generally indicates that the preceding and following related objects have an "or" relationship.
[0053] Example 1
[0054] Figure 1 A flowchart of a financial report element extraction method provided in Embodiment 1 of this application is shown below. Figure 1 The method can be executed by a device that performs the method, which can be implemented in software and / or hardware, and the method includes:
[0055] S110. Select key pages from the original financial reports. The key pages must include at least one of the balance sheet or income statement.
[0056] It should be noted that, in order to facilitate banks and other financial institutions in conducting core business activities such as credit risk assessment, loan approval and pricing, post-loan management, and risk warning based on clients' corporate financial reports, business personnel can upload the client's authorized corporate financial reports to a pre-set cloud platform (OAAS) and encapsulate the corresponding metadata of the corporate financial reports (application name, domain identifier, storage data center, etc.) into Kafka producer messages and write them into a designated Topic (category layer). If the back-office staff of the financial institution needs to use the corporate financial reports, they can download the corporate financial reports from OAAS by subscribing to Topic data and according to the corresponding parameter information. The downloaded corporate financial reports are recorded as the original financial reports. It should be noted that the data format of the original financial reports can be image format and / or PDF format, without any specific restrictions.
[0057] It should also be noted that several key elements can be extracted from the original financial report. These key elements are the main data that financial institutions rely on to provide core business activities to enterprises. For example, the key elements contained in the balance sheet of the original financial report include: cash and cash equivalents, accounts receivable, inventory, fixed assets, etc.; the key elements contained in the income statement of the original financial report include: operating revenue, operating costs, operating profit, etc.
[0058] The original financial report consists of multiple pages, and the aforementioned key elements are generally found in the core financial statements such as the balance sheet and income statement in the original financial report; the pages containing the core financial statements are also known as the key pages.
[0059] Specifically, taking one page from the original financial report, we determine whether the page is a key page by checking if it contains a preset core financial report statement. If the page contains a core financial report statement, it is determined to be a key page. In this embodiment, the core financial report statement includes at least one of the balance sheet and the income statement. In other embodiments, the specific details are not limited.
[0060] S120. Based on the financial statement header of the key page and the preset prompt word template, determine the element extraction instruction. The prompt word template includes a field to be filled for the key elements of the financial statement corresponding to the financial statement header.
[0061] It should be noted that, in order to improve the accuracy and efficiency of extracting key elements from corporate financial reports, this embodiment aims to use a large language model to extract key elements from key pages; in order to guide the large language model to extract the required key elements from key pages, corresponding prompt words need to be generated first, and these prompt words are recorded as element extraction instructions.
[0062] In order to generate the element extraction instruction, this embodiment has a preset prompt word template. The area to be filled in the prompt word template can be filled by the attribute information corresponding to the key page. For example, the attribute information is the key element of the financial statement corresponding to the header of the financial statement of the key page.
[0063] For example, the key elements of the financial statements corresponding to the header of the balance sheet include: cash and cash equivalents, accounts receivable, inventory, fixed assets, etc.; the key elements of the financial statements corresponding to the header of the income statement include: operating revenue, operating costs, net profit, etc.; no specific restrictions are imposed.
[0064] Specifically, the attribute information corresponding to the key page is filled into the area to be filled in the prompt word template, thereby obtaining the corresponding prompt words, which are the feature extraction instructions.
[0065] S130. Based on the key page, the element extraction instruction, the preset model parameters, and the multimodal large model, determine the character recognition sequence; wherein, the multimodal large model is a model obtained by training using preset historical financial reports as training data.
[0066] In this embodiment, a pre-trained large language model is used. To ensure the accuracy of the large language model in the scenario of extracting financial report elements, the embodiment uses pre-set historical financial reports as training data to train the initial large language model. The historical financial reports can be internal historical financial reports of financial institutions and / or historical financial reports authorized by users of financial institutions, and there is no specific limitation.
[0067] This large language model processes key pages from the input to output corresponding character recognition sequences, which contain key elements of a company's financial report. To enable the large language model to process these key pages, preset model parameters must be transmitted to it. For example, these parameters include a temperature coefficient (Temperature) and a kernel sampling parameter (Top_p). The temperature coefficient controls the randomness of the generated text; a lower temperature coefficient tends to select words with the highest probability, while a higher temperature coefficient selects lower-probability words more randomly. The kernel sampling parameter (Top_p) limits the range of words selected by the model, sorting candidate words by probability from high to low, accumulating probabilities until a threshold p is reached, and then retaining only these words. In other embodiments, the specific model parameters are not limited. This large language model is referred to as a multimodal large model.
[0068] Specifically, the preset model parameters are first transmitted to the multimodal large model, and then the key page and element extraction instructions are output to the multimodal large model. The multimodal large model performs target detection (locating text regions) and OCR recognition (recognizing the text in the text regions as character sequences) on the key page according to the element extraction instructions, thereby obtaining a character recognition sequence; among which, the character recognition sequence contains the key elements of the enterprise financial report.
[0069] S140. Perform semantic understanding on the character recognition sequence to determine the target financial reporting elements.
[0070] It should be noted that character recognition sequences are usually a large, disorganized text or simply located text blocks, which cannot directly identify key elements. In order to identify key elements from character recognition sequences, this embodiment adopts a semantic understanding method to identify key elements from character recognition sequences.
[0071] Specifically, the character recognition sequence can be processed through a preset semantic model to obtain the key elements in the character recognition sequence, and these key elements can be recorded as the target financial report elements.
[0072] It should be noted that this embodiment selects key pages from the original financial reports, including at least one of the balance sheet or income statement; based on the financial statement headers of the key pages and a preset prompt word template, an element extraction instruction is determined; wherein, the prompt word template includes a region to be filled for the key elements of the financial statements corresponding to the financial statement headers; based on the key pages, the element extraction instruction, preset model parameters, and a multimodal large model, a character recognition sequence is determined; wherein, the multimodal large model is a model trained using preset historical financial reports as training data; semantic understanding is performed on the character recognition sequence to determine the target financial report elements. Through the above implementation, based on the selection of key pages, corresponding prompt words (element extraction instructions) can be generated from the key pages, and then the element extraction instruction can be used to control the multimodal large model to recognize the characters (character recognition sequences) in the key pages. After that, semantic understanding can be directly performed on the character recognition sequence to extract the financial report elements. This process does not require manual intervention, which can effectively improve the efficiency and accuracy of financial report element extraction and reduce costs.
[0073] Example 2
[0074] This application provides a method for extracting financial report elements in Embodiment 2, which optimizes the "screening key pages from the original financial report" method in Embodiment 1. It should be noted that for parts not detailed in this embodiment, please refer to the descriptions in other embodiments. The method includes:
[0075] S211. The original financial report includes at least one of an original financial report image and an original financial report PDF file; the resolution of the original financial report image is adjusted to a preset resolution, and / or the original financial report PDF file is paginated to obtain the target financial report.
[0076] It should be noted that the original financial report is multimodal data, which may include financial reports in image format and / or financial reports in PDF text format. The financial reports in image format will be referred to as the original financial report image, and the financial reports in PDF text format will be referred to as the original financial report PDF file.
[0077] It should also be noted that the large language model has input requirements for the original financial reports it inputs. For example, if the original financial report is an image, the resolution of the image needs to be adjusted to a preset resolution; if the original financial report is a PDF file (composed of multiple pages), the PDF file needs to be paginated. In other words, before inputting the original financial report into the large language model, it is necessary to perform model input preprocessing on the original financial report; and the original financial report that has completed model input preprocessing (resolution adjustment or pagination, etc.) is denoted as the target financial report.
[0078] It should be noted that by preprocessing the original financial report, not only can the preprocessed financial report meet the input requirements of the large language model, but the preprocessed financial report can also focus more on its key content (core financial statements), thereby improving the accuracy of the large language model in processing the preprocessed financial report.
[0079] S212. Process the target financial report based on a preset visual model to obtain key pages; the key pages include at least one of the balance sheet and the income statement.
[0080] It should be noted that, in order to select key pages from the target financial report, this embodiment has a preset visual model. This visual model is used to visually inspect each page of the target financial report to determine whether the page contains core financial statements (balance sheet, income statement, etc.). If it is detected, the corresponding page is marked as a key page.
[0081] By implementing the above methods, preprocessing the original financial report (resolution adjustment and / or pagination) can make the original financial report meet the processing requirements of the visual model and improve the accuracy of generating key pages.
[0082] S220. Based on the financial statement header of the key page and the preset prompt word template, determine the element extraction instruction; wherein, the prompt word template includes the area to be filled for filling the key elements of the financial statement corresponding to the financial statement header.
[0083] S230. Based on the key page, the element extraction instruction, the preset model parameters, and the multimodal large model, determine the character recognition sequence; wherein, the multimodal large model is a model obtained by training using preset historical financial reports as training data.
[0084] S240. Perform semantic understanding on the character recognition sequence to determine the target financial reporting elements.
[0085] Example 3
[0086] This application provides a method for extracting financial report elements in Embodiment 3. This method optimizes the "determining element extraction instructions based on the financial statement header of the key page and the preset prompt word template" in Embodiment 1. It should be noted that for parts not described in detail in this embodiment, please refer to the descriptions in other embodiments. The method includes:
[0087] S310. Select key pages from the original financial reports. The key pages must include at least one of the balance sheet or income statement.
[0088] S321. Determine the key elements of the financial statements corresponding to the header of the financial statements on the key page.
[0089] It should be noted that the key pages contain core financial statements, which, for example, include a balance sheet and an income statement. Each core financial statement has its corresponding header, which is designated as the financial statement header. To facilitate the subsequent construction of relevant prompts and enable the large model to accurately extract key elements from the key pages, this embodiment pre-sets corresponding key elements for each core financial statement in key-value pairs. The key of each key-value pair is the financial statement header of the core financial statement, and the value is the key element of the financial statement.
[0090] For example, when the core financial statement is the balance sheet, the key of the corresponding key-value pair is "balance sheet", and the value of the corresponding key-value pair is "cash and cash equivalents, accounts receivable, inventory, fixed assets"; when the core financial statement is the income statement, the key of the corresponding key-value pair is "income statement", and the value of the corresponding key-value pair is "revenue, cost of goods sold, net profit".
[0091] S322. Based on the key elements of the financial statements and the preset prompt word template, determine the element extraction instruction.
[0092] Among them, by filling multiple key financial statement elements of the core financial statements into the area to be filled in the preset prompt word template, prompt words can be obtained, and the prompt words are recorded as element extraction instructions.
[0093] Through the above implementation, key elements of financial statements and prompt word templates are combined into element extraction instructions, which can then be used to guide the accuracy of characters generated by the model.
[0094] S330. Based on the key page, the element extraction instruction, the preset model parameters, and the multimodal large model, determine the character recognition sequence; wherein, the multimodal large model is a model obtained by training using preset historical financial reports as training data.
[0095] S340. Perform semantic understanding on the character recognition sequence to determine the target financial reporting elements.
[0096] Example 4
[0097] This application provides a method for extracting financial report elements in Embodiment 4, which optimizes the "determining the key financial report elements corresponding to the financial report header of the key page" in Embodiment 3. It should be noted that for parts not detailed in this embodiment, please refer to the descriptions in other embodiments. The method includes:
[0098] S410. Select key pages from the original financial reports. The key pages must include at least one of the balance sheet or income statement.
[0099] S421A. Perform header recognition on the key page to obtain the financial statement header.
[0100] The key pages contain core financial statements, and the header of the core financial statements is recorded as the financial statement header. The financial statement header is text and is used to indicate the type of the core financial statement.
[0101] Specifically, in this embodiment, the financial statement header of the core financial statements contained in the key page is identified using an OCR recognition tool. In other embodiments, the specific implementation is not limited.
[0102] S421B. Perform semantic recognition on the header of the financial statement to obtain the report type.
[0103] The core financial statements have corresponding report types, and these report types are subsequently used to determine the key elements of the financial statements to be filled into the prompt word template. Therefore, it is necessary to first determine the report types of the core financial statements contained in the key page; for example, the report types include: balance sheet, income statement, and others.
[0104] Specifically, if the metadata of the core financial statements on the key page contains the corresponding report type, the report type of the corresponding core financial statements can be directly extracted from the metadata; if the metadata of the core financial statements on the key page does not contain the corresponding report type, the report type of the core financial statements can be determined by semantic recognition of the core financial statement headers.
[0105] Since the header of a financial statement can indicate the type of the core financial statement, semantic recognition of the header can accurately determine the type of the core financial statement, which is then recorded as the statement type.
[0106] Specifically, the header of the financial statement is semantically recognized using a pre-defined semantic recognition model to obtain the report type of the core financial statement; for example, the report types include balance sheet, income statement, and others.
[0107] S421C. Determine the key elements of the financial statements corresponding to the aforementioned report type.
[0108] In this embodiment, the report type and the corresponding key elements of the financial statement are stored in key-value pairs, where the report type is the key and the corresponding key elements of the financial statement are the value.
[0109] S422. Based on the key elements of the financial statements and the preset prompt word template, determine the element extraction instruction.
[0110] S430. Based on the key page, the element extraction instruction, the preset model parameters, and the multimodal large model, determine the character recognition sequence; wherein, the multimodal large model is a model obtained by training using preset historical financial reports as training data.
[0111] S440. Perform semantic understanding on the character recognition sequence to determine the target financial reporting elements.
[0112] Example 5
[0113] This application provides a method for extracting financial report elements in Embodiment 5. This method optimizes the "semantic understanding of the character recognition sequence to determine the target financial report elements" in Embodiment 1. It should be noted that for parts not described in detail in this embodiment, please refer to the descriptions in other embodiments. The method includes:
[0114] S510. Select key pages from the original financial reports. The key pages must include at least one of the balance sheet or income statement.
[0115] S520. Based on the financial statement header of the key page and the preset prompt word template, determine the element extraction instruction; wherein, the prompt word template includes the area to be filled for filling the key elements of the financial statement corresponding to the financial statement header.
[0116] S530. Based on the key page, the element extraction instruction, the preset model parameters, and the multimodal large model, determine the character recognition sequence; wherein, the multimodal large model is a model obtained by training using preset historical financial reports as training data.
[0117] S541. Perform semantic understanding on the character recognition sequence to obtain initial financial report elements.
[0118] In this process, by using a pre-defined semantic recognition model to perform semantic understanding on the character recognition sequence, key elements in the financial report (cash, accounts receivable, inventory, fixed assets, operating revenue, operating costs, operating profit, etc.) can be initially obtained, and these initially obtained key elements are recorded as initial financial report elements.
[0119] S542. Correct the initial financial reporting elements to obtain the target financial reporting elements; wherein the correction includes at least one of table layout restoration and error correction.
[0120] It should be noted that, firstly, the initial financial report elements are a set of key elements directly output and are not arranged according to the layout in the corresponding financial statements, which is not conducive to the subsequent direct analysis of the initial financial report elements; secondly, some elements in the initial financial report elements may contain errors, such as typos or overall recognition errors, which will reduce the accuracy of the subsequent use of the initial financial report elements; to solve the above two problems, after obtaining the initial financial report elements, it is necessary to correct them first.
[0121] In this embodiment, the correction package uses two methods: table layout restoration and error correction. Table layout restoration involves formatting the initial financial report elements according to the layout in the financial statements. Error correction involves correcting the initial financial report elements using a spell checker, domain dictionary, or regular expressions. The corrected initial financial report elements are then denoted as the target financial report elements.
[0122] By implementing the above methods, the accuracy of the initially generated initial financial report elements can be improved, thereby enhancing the accuracy of the subsequent output target financial report elements.
[0123] Example 6
[0124] This application provides a method for extracting financial report elements in Embodiment Six, which supplements the method shown in Embodiment One. It should be noted that for parts not detailed in this embodiment, please refer to the descriptions in other embodiments. The method includes:
[0125] S610. Select key pages from the original financial reports. The key pages must include at least one of the balance sheet or the income statement.
[0126] S620. Based on the financial statement header of the key page and the preset prompt word template, determine the element extraction instruction; wherein, the prompt word template includes the area to be filled for filling the key elements of the financial statement corresponding to the financial statement header.
[0127] S630. Based on the key page, the element extraction instruction, the preset model parameters, and the multimodal large model, determine the character recognition sequence; wherein, the multimodal large model is a model obtained by training using preset historical financial reports as training data.
[0128] S640. Perform semantic understanding on the character recognition sequence to determine the target financial reporting elements.
[0129] S650. Encapsulate the target financial report elements to obtain the element encapsulation result.
[0130] It should be noted that, in order to facilitate the use of the extracted target financial reporting elements in subsequent processes, the target financial reporting elements can be packaged in a predetermined format first, and then the packaged result can be sent to the publish / subscribe platform for subsequent subscription and use.
[0131] Specifically, in this embodiment, the target financial report elements are encapsulated in JSON format as a unit of financial report files to obtain the element encapsulation result.
[0132] S660. The element encapsulation result is sent to the downstream classification layer in the distributed stream processing platform.
[0133] This embodiment provides a publish-subscribe platform, referred to as a distributed stream processing platform, exemplified by Kafka. Kafka producers can write the element encapsulation results into downstream classification layers (Topics) within the distributed stream processing platform. Subsequently, business personnel can obtain the element encapsulation results from the distributed stream processing platform in real time through their business systems, facilitating applications such as customer credit assessment and risk warning.
[0134] By encapsulating the target financial reporting elements and sending them to the downstream classification layer in the distributed stream processing platform, the ease of use of the target financial reporting elements can be improved.
[0135] It should be understood that although the steps in the flowcharts of the embodiments described above are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the embodiments described above may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages of other steps.
[0136] Example 7
[0137] Based on the same inventive concept, this embodiment also provides a financial reporting element extraction device for implementing the financial reporting element extraction method described above. The solution provided by this device is similar to the solution described in the above method; therefore, the specific limitations in one or more embodiments of the financial reporting element extraction device provided below can be found in the limitations of the financial reporting element extraction method described above, and will not be repeated here.
[0138] In this embodiment, as Figure 2 As shown, a financial report element extraction device is provided, comprising:
[0139] The data filtering module is used to filter out key pages from the original financial report. The key pages must include at least one of the balance sheet and income statement.
[0140] The instruction generation module is used to determine the element extraction instruction based on the financial statement header of the key page and the preset prompt word template; wherein, the prompt word template includes a field to be filled for filling the key elements of the financial statement corresponding to the financial statement header;
[0141] The character recognition module is used to determine the character recognition sequence based on the key page, the element extraction instruction, preset model parameters, and a multimodal large model; wherein, the multimodal large model is a model obtained by training using preset historical financial reports as training data;
[0142] The semantic understanding module is used to perform semantic understanding on the character recognition sequence to determine the target financial reporting elements.
[0143] Each module in the aforementioned financial report element extraction device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device, or stored in the memory of a computer device as software, so that the processor can call and execute the operations corresponding to each module.
[0144] It should be noted that this embodiment selects key pages from the original financial reports, including at least one of the balance sheet or income statement; based on the financial statement headers of the key pages and a preset prompt word template, an element extraction instruction is determined; wherein, the prompt word template includes a region to be filled for the key elements of the financial statements corresponding to the financial statement headers; based on the key pages, the element extraction instruction, preset model parameters, and a multimodal large model, a character recognition sequence is determined; wherein, the multimodal large model is a model trained using preset historical financial reports as training data; semantic understanding is performed on the character recognition sequence to determine the target financial report elements. Through the above implementation, based on the selection of key pages, corresponding prompt words (element extraction instructions) can be generated from the key pages, and then the element extraction instruction can be used to control the multimodal large model to recognize the characters (character recognition sequences) in the key pages. Then, semantic understanding is directly performed on the character recognition sequence to extract the financial report elements. This process requires no manual intervention, effectively improving the efficiency and accuracy of financial report element extraction and reducing costs.
[0145] In an optional embodiment, the original financial report includes at least one of an image of the original financial report and an original financial report PDF file;
[0146] Accordingly, the process of selecting key pages from the original financial report includes:
[0147] Adjust the resolution of the original financial report image to a preset resolution, and / or paginate the original financial report PDF file to obtain the target financial report;
[0148] The target financial report is processed based on a preset visual model to obtain key pages.
[0149] In an optional embodiment, determining the element extraction instruction based on the financial statement header of the key page and a preset prompt word template includes:
[0150] Determine the key elements of the financial statements corresponding to the header of the financial statements on the key page;
[0151] Based on the key elements of the financial statements and the preset prompt template, the element extraction instruction is determined.
[0152] In an optional embodiment, determining the key financial statement elements corresponding to the financial statement header of the key page includes:
[0153] The header of the key page is identified to obtain the financial statement header;
[0154] Semantic recognition is performed on the header of the financial statements to obtain the report type;
[0155] Determine the key elements of the financial statements corresponding to the aforementioned report type.
[0156] In an optional embodiment, the step of semantically understanding the character recognition sequence to determine the target financial reporting element includes:
[0157] Semantic understanding is performed on the character recognition sequence to obtain initial financial report elements;
[0158] The initial financial reporting elements are corrected to obtain the target financial reporting elements; wherein the correction includes at least one of table layout restoration and error correction.
[0159] In an optional embodiment, the financial reporting element extraction method performed by the financial reporting element extraction device provided in this embodiment further includes:
[0160] The target financial reporting elements are encapsulated to obtain the element encapsulation results;
[0161] The encapsulated result of the elements is sent to the downstream classification layer in the distributed stream processing platform.
[0162] Example 8
[0163] In this embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be as follows. Figure 3 As shown, the computer device includes a processor, memory, and a network interface connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and a database. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage media. The database stores data. The network interface communicates with external terminals via a network connection. When executed by the processor, the computer program implements a method for extracting financial reporting elements.
[0164] Those skilled in the art will understand that Figure 3The structure shown is merely a block diagram of a portion of the structure related to the present disclosure and does not constitute a limitation on the computer device to which the present disclosure is applied. A specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0165] Example 9
[0166] In this embodiment, a computer-readable storage medium is provided, such as... Figure 4 As shown, a computer program is stored thereon, and when the computer program is executed by the processor, it implements the steps in the above-described method embodiments.
[0167] Example 10
[0168] In this embodiment, a computer program product is provided, including a computer program that, when executed by a processor, implements the steps in the above-described method embodiments.
[0169] It should be noted that the information collected is information and data authorized by the user or fully authorized by all parties, and the collection, storage, use, processing, transmission, provision, disclosure and application of the relevant data all comply with the relevant laws, regulations and standards of the relevant countries and regions, necessary confidentiality measures have been taken, and it does not violate public order and good morals. Corresponding operation portals are provided for users to choose to authorize or refuse.
[0170] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, databases, or other media used in the embodiments provided in this disclosure can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM). The databases involved in the embodiments provided in this disclosure may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this disclosure may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, etc., and are not limited to these.
[0171] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0172] The embodiments described above are merely illustrative of several implementations of this disclosure, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of this patent disclosure. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this disclosure, and these all fall within the protection scope of this disclosure. Therefore, the protection scope of this disclosure should be determined by the appended claims.
Claims
1. A financial reporting element extraction method, characterized by, The method comprises the following steps: Screening key pages from the original financial report, wherein the key pages at least include one of balance sheet and profit table; Based on the financial statement header of the key page and the preset prompt word template, determine the element extraction instruction; wherein the prompt word template includes a to-be-filled area for filling the key elements of the financial statement corresponding to the financial statement header; Based on the key page, the element extraction instruction, the preset model parameter and the multi-modal large model, determine the character recognition sequence; wherein the multi-modal large model is a model obtained by using the preset historical financial report as training data for training; Perform semantic understanding on the character recognition sequence to determine the target financial report element.
2. The method of claim 1, wherein, The original financial report at least includes one of an original financial report picture and an original financial report PDF file; Correspondingly, the step of screening key pages from the original financial report comprises the following steps: Adjust the resolution of the original financial report picture to a preset resolution, and / or page the original financial report PDF file to obtain a target financial report; Process the target financial report based on a preset visual model to obtain the key page.
3. The method of claim 1, wherein, The step of determining the element extraction instruction based on the financial statement header of the key page and the preset prompt word template comprises the following steps: Determine the financial statement key elements corresponding to the financial statement header of the key page; Based on the financial statement key elements and the preset prompt word template, determine the element extraction instruction.
4. The method of claim 3, wherein, The step of determining the financial statement key elements corresponding to the financial statement header of the key page comprises the following steps: Perform header recognition on the key page to obtain a financial statement header; Perform semantic recognition on the financial statement header to obtain a report type; Determine the financial statement key elements corresponding to the report type.
5. The method of claim 1, wherein, The step of performing semantic understanding on the character recognition sequence to determine the target financial report element comprises the following steps: Perform semantic understanding on the character recognition sequence to obtain an initial financial report element; Correct the initial financial report element to obtain a target financial report element; wherein the correction at least includes one of table layout recovery and error correction.
6. The method of claim 1, wherein, Further comprising the following steps: Package the target financial report element to obtain an element packaging result; Send the element packaging result to a downstream classification layer in a distributed stream processing platform.
7. A financial reporting element extraction apparatus characterized by comprising: The device comprises: A data screening module for screening key pages from the original financial report, wherein the key pages at least include one of balance sheet and profit table; An instruction generation module for determining the element extraction instruction based on the financial statement header of the key page and the preset prompt word template; wherein the prompt word template includes a to-be-filled area for filling the financial statement key elements corresponding to the financial statement header; A character recognition module for determining the character recognition sequence based on the key page, the element extraction instruction, the preset model parameter and the multi-modal large model; wherein the multi-modal large model is a model obtained by using the preset historical financial report as training data for training; A semantic understanding module for performing semantic understanding on the character recognition sequence to determine the target financial report element.
8. A computer device comprising a memory and a processor, the memory storing a computer program, characterized in that, The computer program is executed by the processor to implement the steps of the method of any one of claims 1 to 6.
9. A computer readable storage medium having stored thereon a computer program, characterized in that, The computer program is executed by the processor to implement the steps of the method of any one of claims 1 to 6.
10. A computer program product comprising a computer program, characterized in that, The computer program is executed by the processor to implement the steps of the method of any one of claims 1 to 6.