Rule-driven financial statement automatic generation method and system

CN122414142BActive Publication Date: 2026-09-11NANTONG VOCATIONAL COLLEGE
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202610874391.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-06-17
Publication Date
2026-09-11
Estimated Expiration
2046-06-17

AI Technical Summary

Technical Problem

[0003]本申请提供基于规则驱动的财务报表自动生成方法及系统,用于针对解决现有技术中业务单据与会计科目难以准确匹配,导致财务报表生成效率低的技术问题

Benefits of technology

本申请获取目标企业的企业基础核算配置信息进行数字化建模,构建会计科目层级树,并调取预置通用财务词典对所述会计科目层级树中的科目与报表行次进行合规义项补充和向量化编码,获得多个科目层级向量;实时捕获前端业务系统输入的业务交互单据,将业务交互单据中非结构化摘要文本和半结构化辅助核算字段组合转化为目标业务文本向量;基于所述多个科目层级向量和所述目标业务文本向量进行跨层级交叉注意力矩阵计算,获得多个科目相关性得分,根据所述多个科目相关性得分过滤不相关文本并锁定目标财务候选实体集合;遍历所述目标财务候选实体集合,根据财务借贷先后勾稽逻辑构建要素关联有向无环图,利用路径扩展算法在所述要素关联有向无环图中搜索并解耦出多条独立的结构化要素事件链;将所述多条独立的结构化要素事件链转化为分布式记账凭证数据记入科目余额表,并调用所述会计科目层级树进行金额向上汇总归集,自动生成并渲染目标财务报表。本发明解决现有技术中业务单据与会计科目难以准确匹配,导致财务报表生成效率低的技术问题,通过构建会计科目层级树,并结合向量化编码、跨层级交叉注意力计算和要素事件链解耦处理,达到提高财务报表自动生成准确性和效率的技术效果。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122414142B_ABST
    Figure CN122414142B_ABST
Patent Text Reader

Abstract

The application discloses a rule-driven-based financial statement automatic generation method and system, relates to the technical field of data management, and comprises the following steps: constructing an accounting subject hierarchical tree, combining compliance item supplement and vectorization coding, and forming a subject hierarchical vector; converting abstract text and auxiliary accounting fields in a business interactive document into target business text vectors, determining subject relevance through cross-level cross-attention calculation, screening out irrelevant information and locking financial candidate entities. Subsequently, a directed acyclic graph of element correlation is constructed according to financial debit-credit reconciliation logic, a path expansion algorithm is used to decouple a structured element event chain, and finally, a bookkeeping voucher is generated, subject balances are summarized, and a financial statement is automatically rendered. The application solves the technical problem that business documents and accounting subjects are difficult to accurately match in the prior art, leading to low efficiency of financial statement generation, and achieves the technical effect of improving the accuracy and efficiency of financial statement automatic generation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data management technology, specifically to a rule-driven method and system for automatically generating financial statements. Background Technology

[0002] As the data volume of enterprise business systems, financial systems, and reporting systems continues to expand, business interaction documents typically contain a large amount of unstructured summary text, auxiliary accounting fields, and multi-dimensional transaction elements. Financial statement generation requires first identifying the document content and then mapping the business information to the corresponding accounting subjects and report rows. Due to differences in the subject systems, accounting standards, and business expression methods among different enterprises, there is a lack of stable and accurate association rules between document text and accounting subjects. This easily leads to subject matching deviations, omissions of business elements, or aggregation errors. Consequently, the subsequent generation of accounting vouchers and report summarization processes require significant manual verification, impacting the accuracy and processing efficiency of financial statement generation. Summary of the Invention

[0003] This application provides a rule-driven method and system for automatically generating financial statements, which addresses the technical problem of low efficiency in financial statement generation due to the difficulty in accurately matching business documents with accounting subjects in the prior art.

[0004] In view of the above problems, this application provides a rule-driven method and system for automatically generating financial statements.

[0005] The first aspect of this application provides a rule-driven method for automatically generating financial statements, the method comprising: The system acquires the target enterprise's basic accounting configuration information for digital modeling, constructs an accounting subject hierarchy tree, and uses a pre-built general financial dictionary to supplement and vectorize the subjects and report rows in the accounting subject hierarchy tree, obtaining multiple subject hierarchy vectors. It captures business interaction documents input from the front-end business system in real time, combining unstructured summary text and semi-structured auxiliary accounting fields in the business interaction documents into target business text vectors. Based on the multiple subject hierarchy vectors and the target business text vectors, it performs cross-level attention matrix calculations to obtain multiple subject relevance scores. Irrelevant text is filtered based on these relevance scores, and a target financial candidate entity set is locked. The system traverses the target financial candidate entity set, constructs a directed acyclic graph (DAG) of element associations based on the sequential reconciliation logic of financial debits and credits, and uses a path expansion algorithm to search and decouple multiple independent structured element event chains from the DAG. These independent structured element event chains are converted into distributed accounting voucher data and recorded in the account balance sheet. The system then calls the accounting subject hierarchy tree to aggregate and summarize amounts upwards, automatically generating and rendering the target financial statements.

[0006] A second aspect of this application provides a rule-driven automatic financial statement generation system, the system comprising: The encoding module is used to acquire the target enterprise's basic accounting configuration information for digital modeling, construct an accounting subject hierarchy tree, and retrieve a pre-built general financial dictionary to supplement the subjects and report rows in the accounting subject hierarchy tree with compliant definitions and vectorize them, obtaining multiple subject hierarchy vectors. The conversion module is used to capture business interaction documents input from the front-end business system in real time, and convert the unstructured summary text and semi-structured auxiliary accounting fields in the business interaction documents into target business text vectors. The filtering module is used to perform cross-level cross-attention matrix calculation based on the multiple subject hierarchy vectors and the target business text vectors to obtain... The system employs a multi-subject relevance score mechanism to filter irrelevant text and lock in a target set of financial candidate entities. A traversal module iterates through this set, constructing a directed acyclic graph (DAG) of element relationships based on the sequential reconciliation logic of financial debits and credits. A path expansion algorithm is then used to search and decouple multiple independent structured element event chains within this DAG. Finally, a summary and aggregation module converts these independent structured element event chains into distributed accounting voucher data, records it in the account balance sheet, and calls the accounting subject hierarchy tree to perform upward summarization and aggregation of amounts, automatically generating and rendering the target financial statements.

[0007] One or more technical solutions provided in this application have at least the following technical effects or advantages: This application acquires the target enterprise's basic accounting configuration information for digital modeling, constructs an accounting subject hierarchy tree, and retrieves a pre-built general financial dictionary to supplement the subjects and report rows in the accounting subject hierarchy tree with compliant definitions and vectorize them, obtaining multiple subject hierarchy vectors; it captures business interaction documents input from the front-end business system in real time, and combines unstructured summary text and semi-structured auxiliary accounting fields in the business interaction documents into target business text vectors; it performs cross-level cross-attention matrix calculation based on the multiple subject hierarchy vectors and the target business text vectors to obtain multiple subject relevance scores, filters irrelevant text based on the multiple subject relevance scores, and locks the target financial candidate entity set; it traverses the target financial candidate entity set, constructs a directed acyclic graph of element associations based on the financial debit and credit sequence reconciliation logic, and uses a path expansion algorithm to search and decouple multiple independent structured element event chains in the directed acyclic graph of element associations; it converts the multiple independent structured element event chains into distributed accounting voucher data and records them into the account balance sheet, and calls the accounting subject hierarchy tree to aggregate the amounts upwards, automatically generating and rendering the target financial statements. This invention addresses the technical problem of low efficiency in financial statement generation due to the difficulty in accurately matching business documents with accounting subjects in existing technologies. By constructing a hierarchical tree of accounting subjects and combining vectorized coding, cross-level cross-attention calculation, and decoupling of element event chains, it achieves the technical effect of improving the accuracy and efficiency of automatic financial statement generation. Attached Figure Description

[0008] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0009] Figure 1 A schematic diagram of the rule-driven automatic financial statement generation method provided in this application embodiment; Figure 2 This is a schematic diagram of the structure of a rule-driven automatic financial statement generation system provided in an embodiment of this application.

[0010] Figure labeling: Encoding module 11, Conversion module 12, Filtering module 13, Traversal module 14, Summarization module 15. Detailed Implementation

[0011] This application provides a rule-driven method and system for automatically generating financial statements. It addresses the technical problem of low efficiency in financial statement generation due to the difficulty in accurately matching business documents with accounting subjects in existing technologies. By constructing a hierarchical tree of accounting subjects and combining vectorized coding, cross-level cross-attention calculation, and decoupling processing of element event chains, it achieves the technical effect of improving the accuracy and efficiency of automatic financial statement generation.

[0012] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. All other embodiments obtained by those skilled in the art based on the embodiments of this application without creative effort are within the scope of protection of this application.

[0013] It should be noted that any variation of the terms "comprising" and "having" is intended to cover non-exclusive inclusion, for example, a process, method, system, product, or server that includes a series of steps or units is not necessarily limited to those steps or units that are explicitly listed, but may include other steps or modules that are not explicitly listed or that are inherent to such processes, methods, products, or devices.

[0014] Example 1, as Figure 1 As shown, this application provides a rule-driven method for automatically generating financial statements, the method comprising: Step S100: Obtain the basic accounting configuration information of the target enterprise for digital modeling, construct an accounting subject hierarchy tree, and retrieve a pre-set general financial dictionary to supplement the subjects and report rows in the accounting subject hierarchy tree with compliant definitions and vectorize them to obtain multiple subject hierarchy vectors.

[0015] In this embodiment, the target enterprise's basic accounting configuration information is obtained by reading the enterprise's financial software, ERP platform, and accounting configuration database, including a standard chart of accounts and standard financial statement item rows. The standard chart of accounts is the basic data used to describe various accounting accounts of the enterprise, including account codes, account names, account levels, parent accounts, account attributes, and balance directions. The standard financial statement item rows are data used to describe the position and data retrieval rules of each item in the financial statements, including the report item name, row number, corresponding account range, and summary scope. When digitally modeling the above information, data fields from different sources are first identified, cleaned, and formatted to transform account codes, account names, hierarchical relationships, and report row correspondences into computable structured data. Then, each accounting account is used as a node, the hierarchical relationship between accounts is used as a connection relationship, and a mapping relationship between the final-level account, the parent account, and the report item rows is established, thereby constructing an accounting account hierarchy tree.

[0016] Next, when using a pre-built general financial dictionary to supplement and vectorize the compliance definitions of accounts and report rows in the accounting subject hierarchy tree, the subject objects and report objects to be supplemented and encoded are determined by extracting the category name corresponding to each last-level subject code and the text title corresponding to the standard financial statement item row from the accounting subject hierarchy tree. Then, based on the pre-built general financial dictionary, the standard financial standard definition and accounting attribute description that match the category name and text title are retrieved and concatenated with the original subject and report row text as compliance definitions to form subject extended text containing subject meaning, standard basis, and accounting attributes. Subsequently, a dense vector encoder is used to perform high-dimensional feature space mapping on the subject extended text, so that the accounting subjects and report rows are transformed from text semantics into computable vector expressions, and finally, multiple subject-level vectors corresponding one-to-one with each subject and report row are output.

[0017] Furthermore, the method provided in the application embodiment, which involves retrieving a pre-set general financial dictionary to supplement the accounts and report rows in the accounting subject hierarchy tree with compliant definitions and perform vectorized encoding to obtain multiple subject hierarchy vectors, also includes: Extract the category name of each final-level account code in the accounting subject hierarchy tree and the text title of the standard financial statement item row number; retrieve the standard financial standard definition and accounting attribute description corresponding to the category name and the text title from the pre-set general financial dictionary, and use them as compliance terms to concatenate the text to obtain the subject extended text; use a dense vectorized encoder to perform high-dimensional feature space mapping on the subject extended text, and output multiple subject-level vectors that correspond one-to-one with each subject and report row number.

[0018] In this embodiment, the accounting subject hierarchy tree is traversed node by node, and the subject code, subject name, parent subject code, and child node record are read one by one from each subject node. If a subject node has no lower-level child nodes, or if there is no lower-level record in the parent-child relationship record that uses the subject code as the parent subject code, then the subject code is determined as the last-level subject code. For each last-level subject code, its subject name field is read, and the code prefix, hierarchy number, spaces, parentheses, remarks, and auxiliary accounting identifier are deleted in sequence, retaining the text content indicating the accounting category, to obtain the category name of the last-level subject code. The row number, report name, project name, project title, and data retrieval description fields are read line by line for each standard financial statement item. When the project title field contains content, the project title field is extracted as the text title; when the project title field is empty, the project name field is extracted as the text title, and the serial number, spaces, newline characters, and blank placeholders in the text title are cleared to obtain the text title corresponding to the standard financial statement item row.

[0019] Next, the category name and text title are used as search keywords to perform term retrieval in a pre-built general financial dictionary. During the search, the search keywords are first matched character-by-character with the term names in the pre-built general financial dictionary. If no matching term is found, the thesaurus, near-synonym list, and financial standard term mapping table in the pre-built general financial dictionary are read to generate a candidate term set. For each candidate term name, the same cleaning and word segmentation processes are performed on both the search keywords and the candidate term name, resulting in two word element sequences. All words appearing in the two word element sequences are summarized into a word element dictionary, and the frequency of each word element in the search keywords and candidate term names is counted to form the keyword frequency vector and the candidate term frequency vector. When calculating text similarity, the values ​​of the two word frequency vectors at the same word positions are multiplied one by one and summed to obtain the word overlap degree. Then, the sum of the squares of each value in the two word frequency vectors is calculated, and the square root of the sum is taken to obtain the length of the two vectors. Finally, the word overlap degree is divided by the product of the two vector lengths to obtain the text similarity between the search keyword and the candidate term name. If there are multiple candidate terms, the candidate term with the highest text similarity and reaching the preset matching threshold is selected as the hit term. The standard financial standard definition field and the accounting attribute description field in the hit term are read and used as compliance terms. Then, the category name, text title, standard financial standard definition, and accounting attribute description are concatenated in a fixed field order. During the concatenation process, duplicate content is deleted, missing fields are marked with null values, and uniform separators are added between adjacent fields to obtain the subject extended text.

[0020] Finally, the extended subject text is input into a dense vectorized encoder for high-dimensional feature space mapping. The dense vectorized encoder is an encoding model used to convert text content into continuous numerical vectors, including a token embedding layer, a context encoding layer, and a vector output layer. During processing, the extended subject text is first segmented into words, resulting in multiple tokens arranged in text order. Then, the token embedding layer queries the initial numerical vector corresponding to each token and assembles a token vector matrix according to the token order. When calculating the semantic association weights between tokens, the context encoding layer first converts the token vector matrix into a query vector set, a key vector set, and a value vector set, respectively. Then, each query vector is multiplied dimension-wise by each key vector and summed to obtain the association score between that token and other tokens. Finally, the association score is divided by the square root of the key vector dimension to perform scaling, avoiding excessively large scores. Subsequently, softmax processing is performed on a set of scaled association scores corresponding to the same word unit. First, the maximum value in the set of association scores is taken, and this maximum value is subtracted from each other's association scores to obtain a stable score. Then, an exponential operation is performed on each stable score with the natural constant e as the base to obtain the corresponding exponential value. Then, each exponential value is divided by the sum of all exponential values ​​in the set to obtain the semantic association weight of the word unit to other words. Based on the semantic association weight, the set of value vectors is weighted and summed to obtain the word unit encoding result containing the context semantics. Then, all word unit encoding results are averaged along the same dimension to obtain the text representation vector of a subject extended text. Finally, the text representation vector is normalized. First, the square root of the sum of the squares of the values ​​in each dimension of the text representation vector is calculated to obtain the vector length. Then, the value of each dimension is divided by the vector length to make the output vector uniform in numerical scale, forming a subject-level vector. Each subject extension text generates a subject-level vector according to the above process, and stores it in association according to the corresponding subject code or report row number, outputting multiple subject-level vectors that correspond one-to-one with each subject and report row.

[0021] Step S200: Capture business interaction documents input from the front-end business system in real time, and combine the unstructured summary text and semi-structured auxiliary accounting fields in the business interaction documents into the target business text vector.

[0022] In this embodiment, the submission, saving, or approval of business interaction documents serves as the trigger condition to read document data input from the front-end business system in real time. This document data includes document number, document type, business occurrence time, unstructured summary text, and semi-structured auxiliary accounting fields. The unstructured summary text includes continuous text content such as summary, remarks, and usage descriptions. The semi-structured auxiliary accounting fields include data with field names and values, such as counterparty entity, department, project, amount, tax, date, and business category. After reading, the unstructured summary text is cleaned, specifically by removing spaces, duplicate punctuation, and invalid symbols, and standardizing the date to a year-month-day format, the amount to a numeric format, and the tax to a numeric format. The semi-structured auxiliary accounting fields are then text-based, converting each field into field text in the form of field name plus field value; for example, counterparty entity plus corresponding entity name, amount plus corresponding amount value, and tax plus corresponding tax value.

[0023] Next, the cleaned unstructured summary text and the text-based semi-structured auxiliary accounting fields are concatenated in a fixed order: document type, business occurrence time, counterparty, department, project, business category, amount, tax amount, and unstructured summary text, thus forming the business combination text. The business combination text is then segmented into multiple business terms. Each business term is then queried against a pre-set word vector table. If a business term is not found in the pre-set word vector table, its corresponding vector is set to all zeros to ensure that each business term has a numerical representation of the same dimension.

[0024] Next, the word vectors of all business terms are averaged to obtain the average business text vector. The calculation involves first summing the word vectors of all business terms along the same dimension to obtain a vector sum; then, the number of business terms is counted, and the value of each dimension in the vector sum is divided by the number of business terms to obtain the average business text vector. Subsequently, the average business text vector is normalized by squaring the values ​​of each dimension and summing them, then taking the square root of the sum to obtain the vector length; finally, the value of each dimension in the average business text vector is divided by this vector length to obtain a target business text vector with a uniform scale. Thus, the unstructured summary text and semi-structured auxiliary accounting fields in the business interaction documents are combined and transformed into the target business text vector.

[0025] Step S300: Calculate the cross-level attention matrix based on the multiple subject-level vectors and the target business text vector to obtain multiple subject-related scores. Filter irrelevant texts and lock the target financial candidate entity set based on the multiple subject-related scores.

[0026] In this embodiment, when calculating the cross-level attention matrix based on multiple subject-level vectors and the target business text vector, the multiple subject-level vectors are aggregated according to the subject-level relationship and the correspondence between report rows to form a query retrieval matrix. The target business text vector is used as the key-value matrix, so that the vector representation of accounting subjects and report rows and the vector representation of business interaction documents are included in the same calculation process. Then, the dot product operation is performed between the query retrieval matrix and the transpose of the key-value matrix to calculate the matching degree between each subject-level vector and the target business text vector, and an initial similarity matrix is ​​obtained. The initial similarity matrix is ​​then scaled to reduce the influence of vector dimension on the similarity value, and the similarity value is converted into a weight distribution through normalized activation to output the cross-level attention weight matrix. Finally, the weight distribution values ​​corresponding to each subject and report row in the cross-level attention weight matrix are used as the relevance scores of multiple subjects.

[0027] Next, irrelevant text is filtered and the target financial candidate entity set is locked based on the relevance scores of multiple subjects. In this process, this step iterates through multiple subject relevance scores, compares the subject relevance score of each text paragraph with a preset relevance filtering threshold, and removes non-core business text paragraphs with scores below the threshold, retaining those with scores above or equal to the threshold, thus assembling a set of key text sentences. Subsequently, a named entity recognition component is used to scan the key text sentence set, extracting multiple attribute entities according to financial business semantics, and associating and encapsulating multiple attribute entities belonging to the same text sentence to form multiple attribute entity combinations. Finally, the target financial candidate entity set is composed of these multiple attribute entity combinations.

[0028] Furthermore, the method provided in the application embodiment, which calculates a cross-level attention matrix based on the multiple subject-level vectors and the target business text vector to obtain multiple subject relevance scores, also includes: The multiple subject-level vectors are aggregated to construct a query retrieval matrix, with the target business text vector as the key-value matrix. The dot product of the query retrieval matrix and the transpose of the key-value matrix is ​​calculated to obtain an initial similarity matrix. The initial similarity matrix is ​​scaled and normalized to output a cross-level cross-attention weight matrix, and the corresponding weight distribution values ​​in the cross-level cross-attention weight matrix are used as the relevance scores of multiple subjects.

[0029] In this embodiment, when aggregating multiple account-level vectors, the account code, report row number, and vector dimension corresponding to each account-level vector are first read, and the vector arrangement sequence number is determined according to the hierarchical order in the accounting account hierarchy tree and the arrangement order of the standard financial statement item rows. Let n be the number of accounts and report rows involved in the calculation, and each account-level vector contain d dimensions. The account-level vector with arrangement sequence number 1 is written to the first row, the account-level vector with arrangement sequence number 2 is written to the second row, and so on up to the nth row, forming an n-row, d-column query retrieval matrix. Then, the values ​​of the first to d dimensions in the target business text vector are read and arranged in the original dimension order into a 1-row, d-column matrix, which is used as the key-value matrix.

[0030] Next, the dot product of the query matrix and the transpose of the key matrix is ​​calculated. First, the 1xd key matrix is ​​transposed into a dx1 transpose matrix, making it multiplyable with the nxd query matrix. During the calculation, each row of the query matrix's subject-level vector is selected sequentially. The first dimension value of that row is multiplied by the first dimension value of the target business text vector, the second dimension value is multiplied by the second dimension value of the target business text vector, and so on, until the corresponding multiplication of the dth dimension is completed. Then, the d products are summed to obtain an initial similarity value corresponding to that subject or report row. Following the same calculation process, rows 1 to n of the query matrix are calculated row by row to obtain n initial similarity values, which are then arranged into an nx1 initial similarity matrix according to the original subject and report row order.

[0031] Next, the initial similarity matrix is ​​scaled. First, the dimension *d* of the subject-level vector is obtained, and the square root of *d* is taken to obtain the scaling factor. Then, each initial similarity value in the initial similarity matrix is ​​divided by this scaling factor to obtain the scaled similarity matrix. When normalizing the activation of the scaled similarity matrix, softmax is used. First, the maximum value is selected from all scaled similarity values, and this maximum value is subtracted from each other to obtain the stabilized similarity value. Then, an exponential operation is performed on each stabilized similarity value with the natural constant *e* as the base to obtain the corresponding exponential value. All exponential values ​​are then summed to obtain the total exponential value. Finally, each exponential value is divided by the total exponential value to obtain the corresponding weight distribution value. All weight distribution values ​​are arranged in order of subject and report row to form a cross-level cross-attention weight matrix. The weight distribution values ​​corresponding to each subject and report row in the cross-level cross-attention weight matrix are used as the relevance scores for multiple subjects.

[0032] Furthermore, the method provided in the application embodiments, which filters irrelevant text and locks the target financial candidate entity set based on the relevance scores of the multiple subjects, further includes: The relevance scores of the multiple subjects are iterated through, and non-core business paragraphs with scores below the preset relevance filtering threshold are removed. Text paragraphs with scores above or equal to the relevance filtering threshold are extracted and assembled into a set of golden text sentences. The golden text sentence set is then scanned using a named entity recognition component to extract multiple attribute entities. Multiple attribute entities belonging to the same text sentence are associated and encapsulated to obtain multiple attribute entity combinations. These multiple attribute entity combinations constitute the target financial candidate entity set.

[0033] Furthermore, the method provided in the application embodiments also includes: Each of the multiple attribute entities is one of the following: counterparty entity, core transaction amount, associated input and output tax amount, and transaction occurrence time.

[0034] In this embodiment, when iterating through multiple subject relevance scores, the text content in the business interaction document is first divided into multiple business paragraph texts according to periods, semicolons, line breaks, field separators, and document field boundaries, and each business paragraph text is associated with a corresponding subject relevance score. When a business paragraph text corresponds to multiple subject relevance scores, all scores corresponding to that business paragraph text are read, and the maximum value is taken as the paragraph relevance score of that business paragraph text. Subsequently, the paragraph relevance score of each business paragraph text is compared item by item with a preset relevance filtering threshold. If the paragraph relevance score is less than the preset relevance filtering threshold, it indicates that the relevance of the business paragraph text to the accounting subject or report row is insufficient, and it is removed from the text to be processed as a non-core business paragraph text. If the paragraph relevance score is greater than or equal to the relevance filtering threshold, the text paragraph is retained and arranged according to its original appearance order in the business interaction document to form a set of golden text sentences.

[0035] After obtaining the set of golden text sentences, a named entity recognition component is used to scan each sentence. Specifically, each sentence is first cleaned by removing extra spaces, duplicate punctuation, and invalid characters, and converting full-width characters in amounts, taxes, and dates to half-width characters. Then, the sentences are segmented according to a financial field dictionary, a subject suffix dictionary, amount regularization rules, date regularization rules, and a general word segmentation dictionary. During segmentation, consecutive numbers, decimal points, and currency units are identified to form amount words; consecutive date content connected by years, months, days, slashes, and hyphens is identified to form time words; and subject suffixes such as company, unit, supplier, customer, and school, along with their preceding consecutive names, are identified to form subject words. The remaining text is matched from left to right using the general word segmentation dictionary to the maximum length. Single characters or consecutive characters that cannot be matched are retained as ordinary words, thus obtaining a word sequence arranged in the original text order. Subsequently, entities were identified according to four attribute categories: counterparty, core transaction amount, associated input and output VAT, and transaction occurrence time. The counterparty was identified using entity name suffixes such as company, unit, supplier, customer, and school. The core transaction amount was identified using amount field names, numbers, decimal points, currency units, and semantic terms such as price, receipt, payment, and income. Associated input and output VAT was identified using tax field names such as input VAT, output VAT, tax amount, and tax rate, along with their corresponding values. The transaction occurrence time was identified using time field names such as transaction date, invoice date, payment / receipt date, and accounting date, as well as year-month-day, date separators, and consecutive time values. For consecutively occurring suffixes belonging to the same entity type, they were merged into a single attribute entity, thus extracting multiple attribute entities from the golden text sentence set.

[0036] When encapsulating multiple extracted attribute entities, the same text sentence is used as the encapsulation boundary. The counterparty, core transaction amount, associated input and output tax, and business occurrence time identified in the text sentence are read, and multiple attribute entities belonging to the same text sentence are written into the same data unit. If multiple attribute entities of the same type exist in the same text sentence, they are recorded sequentially according to their appearance order in the text sentence, and the entity type is determined by combining the field name; if a certain type of attribute entity does not appear in the text sentence, the corresponding position is recorded as null. Subsequently, the attribute entities in the same text sentence are encapsulated according to the order of business occurrence time, counterparty, core transaction amount, and associated input and output tax to obtain an attribute entity combination. The above process is repeated for each text sentence in the gold text sentence set to obtain multiple attribute entity combinations, which together constitute the target financial candidate entity set.

[0037] Step S400: Traverse the target financial candidate entity set, construct a directed acyclic graph of element associations based on the sequential reconciliation logic of financial lending, and use the path expansion algorithm to search and decouple multiple independent structured element event chains in the directed acyclic graph of element associations.

[0038] In this embodiment, by traversing the target financial candidate entity set, each attribute entity in the multiple attribute entity combinations is read one by one, and a preset financial reconciliation topology rule is invoked to divide each attribute entity into subject nodes, behavior nodes, value nodes, or time nodes according to financial semantics. Then, according to the evolution order in the financial loan reconciliation logic, directed connection edges are established between subject nodes, time nodes, behavior nodes, and value nodes. The subject node is used as the root node of the graph, the time node and behavior node are used as intermediate logical routing nodes, and the value node is used as the leaf tail node. This makes the financial elements in the target financial candidate entity set form a topology structure with directional relationships and no loops, thereby constructing a directed acyclic graph of element association.

[0039] Next, a path expansion algorithm is used to search the directed acyclic graph of element associations. Starting from the root node and performing a depth-first search along the direction of the directed connecting edges, multiple complete topological paths from the root node to the leaf tail node are extracted, forming multiple initial independent event sequences. Then, topological features are extracted from the multiple initial independent event sequences to identify the attribute intersection type and the number of similar attributes, and multiple event complexity indices are calculated to characterize the information mixing density. Then, the multiple event complexity indices are filtered according to a preset document segmentation threshold to map and extract multiple leading complex associated event chains. Then, based on the attribute entity combinations corresponding to the multiple leading complex associated event chains, the adjacent independent event sequences under the same root node are associated and arranged, and the adjacency influence is corrected with the goal of optimal alignment between capital flow and business flow, resulting in multiple sets of corrected element event chains. Finally, based on the multiple sets of corrected element event chains, the multiple initial independent event sequences are merged and pruned to obtain multiple independent structured element event chains.

[0040] Furthermore, the method provided in the application embodiment, which involves traversing the target financial candidate entity set and constructing a directed acyclic graph of element relationships based on the sequential reconciliation logic of financial lending, also includes: Retrieve preset financial reconciliation topology rules, and divide each attribute entity in the target financial candidate entity set into subject nodes, behavior nodes, numerical nodes, or time nodes respectively; establish directed connection edges according to the evolution order in the financial loan reconciliation logic, wherein the subject node is used as the root node of the graph, the time node and the behavior node are used as intermediate logical routing nodes, and the numerical node is used as the leaf tail node, thus constructing a directed acyclic graph of element association.

[0041] In this embodiment, when retrieving the preset financial reconciliation topology rules, multiple attribute entity combinations in the target financial candidate entity set are first read, and the attribute entities within each attribute entity combination are parsed one by one. For each attribute entity, its entity type, entity text, source text sentence, and field source are read; when the entity type is a counterparty entity, or the entity text contains entity name features such as customer, supplier, unit, company, school, etc., the attribute entity is classified as a subject node; when the entity type is the core transaction amount or associated input and output tax amount, or the entity text contains numerical features such as amount value, tax amount value, currency unit, input tax amount, output tax amount, etc., the attribute entity is classified as a value node; when the entity type is the business occurrence time, or the entity text contains time features such as year, month, day, date separator, invoice date, transaction date, payment date, receipt and payment date, the attribute entity is classified as a time node; when the entity text or its source text sentence contains business action words such as purchase, sales, receipt, payment, reimbursement, invoicing, settlement, etc., the corresponding business action is classified as a behavior node. Therefore, each attribute entity in the target financial candidate entity set is converted into a subject node, behavior node, numerical node, or time node.

[0042] When establishing directed edges based on the evolutionary order of financial lending and borrowing logic, the processing scope is based on combinations of entities with the same attribute. First, it checks if a principal node exists within the combination and designates it as the root node of the graph. Next, it searches for time nodes and behavior nodes within the combination and designates them as intermediate logical routing nodes. Finally, it finds the numerical nodes corresponding to the core transaction amount and associated input / output tax and designates them as leaf nodes. When establishing edges, the direction is determined according to the order of principal node, time node, behavior node, and numerical node. If both principal and time nodes exist, a directed edge is established from the principal node to the time node; if both time and behavior nodes exist, a directed edge is established from the time node to the behavior node; if both behavior and numerical nodes exist, a directed edge is established from the behavior node to the numerical node. If a combination of entities lacks a time node, the principal node directly points to the behavior node; if a behavior node is missing, the time node directly points to the numerical node; if both time and behavior nodes are missing, the principal node directly points to the numerical node, ensuring continuous connection of financial elements within the same combination of entities.

[0043] After all attribute entity combinations have been partitioned into nodes and directed connections established, the nodes are merged and deduplicated. For counterparty entities originating from different attribute entity combinations but with identical entity text, they are merged into a single entity node; for similar time nodes occurring at the same time, they are merged into a single time node; for behavior nodes corresponding to the same business action words, they are merged along the same entity node and time node path; for numerical nodes corresponding to core transaction amounts and associated input / output tax amounts, they are retained separately according to the amount value, tax value, and source text sentence. During the construction process, connections are only established in the directions from entity nodes to time nodes, time nodes to behavior nodes, and behavior nodes to numerical nodes. Connections returning from numerical nodes to behavior nodes, time nodes, or entity nodes are not established, nor are back-pointing connections between nodes at the same level established, ensuring the graph structure is free of loops. Thus, with the entity node as the root node, the time node and behavior node as intermediate logical routing nodes, and the numerical node as the leaf tail node, a directed acyclic graph (DAG) of element associations is constructed.

[0044] Furthermore, the method provided in the application embodiment, which uses a path expansion algorithm to search and decouple multiple independent structured feature event chains in the directed acyclic graph of feature association, also includes: Starting from the root node of the directed acyclic graph (DAG) associated with the elements, a depth-first search is performed along the direction of the directed connecting edges to extract multiple complete topological paths from the root node to the leaf tail node, obtaining multiple initial independent event sequences. Topological features are extracted from these initial independent event sequences to identify attribute intersection types and the number of similar attributes, calculating multiple event complexity indices representing information mixing density. These event complexity indices are then filtered according to a preset document segmentation threshold to map and extract multiple leading complex associated event chains. Based on the attribute entity combinations corresponding to these leading complex associated event chains, adjacent independent event sequences under the same root node are associated and categorized. Adjacency impact correction is performed with the goal of optimal alignment between capital flow and business flow, obtaining multiple sets of corrected element event chains. Based on these sets of corrected element event chains, the initial independent event sequences are merged and pruned to obtain multiple independent structured element event chains.

[0045] In this embodiment, when performing a depth-first search starting from the root node of a directed acyclic graph (DAG), the system first reads the main nodes with zero incoming edges and uses each main node as the starting point for the search. Then, it sequentially visits time nodes, behavior nodes, and value nodes along the direction of the directed edges. During the search, a current path record is established. For each node visited, the system records the node type, node text, source attribute entity combination, and direction of the connecting edge. When a value node without any lower-level directed edges is visited, the current path is defined as a complete topological path from the root node to the leaf tail node. If the current node has multiple lower-level nodes, the system enters the next branch sequentially according to the order in which the lower-level nodes appear in the original text. The current path record is copied each time a branch is entered, allowing different branches to form independent paths. This search is repeated for each root node until all leaf tail nodes under that root node have been visited, resulting in multiple initial independent event sequences.

[0046] When extracting topological features from multiple initial independent event sequences, the main nodes, time nodes, behavior nodes, and numerical nodes in each initial independent event sequence are read one by one, and the adjacent connection relationships of each node in the sequence within the directed acyclic graph of element association are checked. If the same main node connects to multiple time nodes, it is identified as a time attribute intersection; if the same time node connects to multiple behavior nodes, it is identified as a behavior attribute intersection; if the same behavior node connects to multiple core transaction amount nodes or associated input and output tax amount nodes, it is identified as a numerical attribute intersection. For each type of attribute intersection, the number of branches connected downwards to the corresponding node is counted, and the value after subtracting one from the number of branches is taken as the intersection increment of the attribute intersection; at the same time, the number of similar attributes of core transaction amount, associated input and output tax amount, business occurrence time, and behavior nodes within the same root node is counted. When the number of a certain type of attribute is greater than one, the number exceeding one is taken as the similar attribute increment. The various intersection increments and the various similar attribute increments are added together to obtain the information mixing density value corresponding to the initial independent event sequence, and this information mixing density value is used as the event complexity index.

[0047] Subsequently, when filtering multiple event complexity indicators according to a preset document segmentation threshold, the event complexity indicator corresponding to each initial independent event sequence is read one by one and compared with the preset document segmentation threshold. If the event complexity indicator is less than the preset document segmentation threshold, the original independent state of the initial independent event sequence is retained; if the event complexity indicator is greater than or equal to the preset document segmentation threshold, the complete topological path, attribute cross-type, number of similar attributes, and source attribute entity combination corresponding to the initial independent event sequence are read, and the initial independent event sequence is marked as a leading complex associated event chain. In this way, event sequences that meet the segmentation conditions among multiple event complexity indicators are mapped and extracted into multiple leading complex associated event chains.

[0048] Next, based on the attribute entity combinations corresponding to multiple leading complex accompanying event chains, adjacent independent event sequences under the same root node are associated and categorized. First, the main node of each leading complex accompanying event chain is read, and the initial independent event sequences adjacent to it are searched within the root node range corresponding to the main node. Adjacent includes sharing the same main node, connecting the same time node, connecting the same behavior node, or the numerical nodes originating from the same text sentence. Subsequently, the business occurrence time, counterparty entity, core transaction amount, and accompanying input and output tax amount in the adjacent independent event sequence are compared item by item with the attribute entity combinations corresponding to the leading complex accompanying event chain. When the two have the same counterparty entity, the same or adjacent business occurrence time, and the core transaction amount and accompanying input and output tax amount have the same sentence source, the same behavior node source, or a corresponding relationship between amount and tax, the adjacent independent event sequence is categorized into the neighborhood independent event sequence set corresponding to the leading complex accompanying event chain, thus obtaining the neighborhood independent event sequence set used for adjacency impact correction.

[0049] Subsequently, when correcting adjacency impact with the goal of optimizing the synergy alignment between capital flow and business flow, the first leading complex associated event chain and its corresponding first attribute entity combination are extracted. Then, based on the topological position of the first leading complex associated event chain in the directed acyclic graph of element associations, the synergy alignment of the adjacent neighborhood independent event sequence set is identified under an accrual basis, yielding the first synergy stability score. Next, the neighborhood independent event sequence set is subjected to multiple random perturbations according to a preset adjustment span, forming multiple adjusted neighborhood event sequence sets. The adjustment synergy stability score between each adjusted neighborhood event sequence set and the first attribute entity combination is calculated. When there is an item among the multiple adjustment synergy stability scores that is greater than or equal to the first synergy stability score, the adjusted neighborhood event sequence set corresponding to the maximum adjustment synergy stability score is selected as the first corrected element event chain set, and this first corrected element event chain set is added to the multiple corrected element event chain sets, thus obtaining multiple corrected element event chain sets.

[0050] Finally, when merging and pruning multiple initial independent event sequences based on multiple sets of modified element event chains, the main node, time node, behavior node, core transaction amount node, and associated input and output tax node contained in each set of modified element event chains are read first, and the corresponding initial independent event sequence is determined according to the combination of the node source attribute entities. For multiple initial independent event sequences from the same set of modified element event chains that have the same source main node, the same business occurrence time, or the same behavior node, they are merged in the order of main node, time node, behavior node, and value node, so that the core transaction amount and associated input and output tax scattered in different initial independent event sequences are included in the same event chain. For main nodes, time nodes, and behavior nodes that appear repeatedly during the merging process, only one node record is retained. Duplicate amount nodes, duplicate tax nodes, or isolated nodes that do not form an effective connection with main nodes, time nodes, or behavior nodes are pruned and deleted. For initial independent event sequences that are not included in any set of modified element event chains and have a complete connection from the root node to the leaf tail node, their independent state is maintained. After merging and pruning, the node order is rearranged according to the main node, business occurrence time, business behavior, core transaction amount and associated input and output tax of each event chain, and the directed connection relationship between the nodes is updated to obtain multiple independent structured element event chains.

[0051] Furthermore, the method provided in the application embodiment, which aims to optimize the alignment between cash flow and business flow by correcting adjacency effects and obtaining a set of multiple corrected element event chains, also includes: Extract the first leading complex associated event chain and its corresponding first attribute entity combination. Combine its topological position in the graph to identify the cooperative alignment degree of the adjacent neighborhood independent event sequence set under the accrual basis, and obtain the first cooperative stability score. For the neighborhood independent event sequence set, perform multiple random perturbation adjustments according to a preset adjustment span to obtain multiple adjusted neighborhood event sequence sets. Again, identify the cooperative alignment degree of the multiple adjusted neighborhood event sequence sets with the first attribute entity combination to obtain multiple adjusted cooperative stability scores. If there is an item in the multiple adjusted cooperative stability scores that is greater than or equal to the first cooperative stability score, then the adjusted neighborhood event sequence set corresponding to the maximum value is taken as the first corrected element event chain set, and the first corrected element event chain set is added to the multiple corrected element event chain sets.

[0052] In this embodiment, when extracting the first leading complex associated event chain and its corresponding first attribute entity combination, one currently to be processed is selected from multiple leading complex associated event chains as the first leading complex associated event chain, and its corresponding first attribute entity combination is read. The first attribute entity combination includes the counterparty entity, the core transaction amount, the associated input and output tax amount, and the business occurrence time. Based on the topological position of the first leading complex associated event chain in the element association directed acyclic graph, its main node, time node, behavior node, numerical node, and directed connection edges between nodes are read, and the set of neighboring independent event sequences adjacent to the first leading complex associated event chain obtained by the aforementioned association arrangement is called. When identifying the collaborative alignment degree of a set of independent event sequences in a neighborhood under the accrual basis, the counterparty entity, business occurrence time, business behavior, core transaction amount, and associated input and output tax amount in each independent event sequence are read one by one and compared with the first attribute entity combination. Specifically, the entity alignment item is recorded as one if the counterparty entity is the same, otherwise it is recorded as zero; the time alignment item is recorded as one if the business occurrence time is the same or belongs to the same accounting period, otherwise it is recorded as zero; the cash flow alignment item is recorded as one if there is a price-tax correspondence between the core transaction amount and the associated input and output tax amount, or if the amount direction is consistent with the lending direction corresponding to the business behavior, otherwise it is recorded as zero; the business flow alignment item is recorded as one if the business behavior is consistent with the business semantics in the first attribute entity combination, otherwise it is recorded as zero. The entity alignment item, time alignment item, cash flow alignment item, and business flow alignment item are added together and divided by four to obtain the collaborative alignment degree value of a single independent event sequence in the neighborhood; then, the average of all collaborative alignment degree values ​​in the set of independent event sequences in the neighborhood is calculated to obtain the first collaborative stability score.

[0053] Next, the set of independent event sequences in the neighborhood is subjected to multiple random perturbations according to the adjustment step size. First, the preset adjustment step size is read, which limits the number of independent event sequences in the neighborhood that can be adjusted in each perturbation adjustment. In each perturbation adjustment, no more than the number of independent event sequences in the neighborhood are randomly selected from the set of independent event sequences in the neighborhood, and adjusted according to their correspondence with the subject, time, capital flow, and business flow of the first attribute entity combination. For independent event sequences in the neighborhood with the same subject node but adjacent time nodes, they are adjusted to the adjacent time path of the first leading complex associated event chain. For independent event sequences in the neighborhood with the same subject node and time node but different behavior nodes, their arrangement position in the behavior node path is adjusted according to the order of business occurrence. For independent event sequences in the neighborhood where the amount node or tax node does not have a correspondence with the first attribute entity combination, they are removed from the neighborhood in this round of perturbation. After each perturbation adjustment is completed, the adjusted arrangement of independent event sequences is retained as an adjusted neighborhood event sequence set. Multiple perturbation adjustments are repeated to obtain multiple adjusted neighborhood event sequence sets.

[0054] Next, the coordination alignment degree of multiple adjusted neighborhood event sequence sets was identified again in combination with the first attribute entity. For each adjusted neighborhood event sequence set, the adjusted event sequence was read one by one, and recalculated according to the subject alignment item, time alignment item, capital flow alignment item, and business flow alignment item. Each item was still marked as one if the condition was met and zero if the condition was not met. The four results were added together and divided by four to obtain the coordination alignment degree value of the adjusted single event sequence. Subsequently, the coordination alignment degree values ​​of all event sequences in the same adjusted neighborhood event sequence set were averaged to obtain the adjustment coordination stability score corresponding to the adjusted neighborhood event sequence set. The above calculation was performed on multiple adjusted neighborhood event sequence sets one by one to obtain multiple adjustment coordination stability scores.

[0055] If any of the multiple adjustment coordination stability scores contains an item greater than or equal to the first coordination stability score, then the adjustment coordination stability score with the largest value is selected from the satisfying adjustment coordination stability scores, and the adjustment neighborhood event sequence set corresponding to this maximum value is read. This adjustment neighborhood event sequence set is determined as the first corrective element event chain set, indicating that, compared to the neighborhood independent event sequence set before the disturbance adjustment, it has a cash flow and business flow coordination alignment degree no lower than the original result under the accrual basis. Subsequently, the first corrective element event chain set is added to multiple corrective element event chain sets; the above identification, disturbance, scoring, and selection process is repeated for other leading complex accompanying event chains to obtain multiple corrective element event chain sets.

[0056] Furthermore, the method provided in the application embodiments also includes: If the cooperative stability scores of the multiple adjustments are all less than the first cooperative stability score, the current perturbation adjustment method is marked with a "disabled times" label, and the invocation of the specific perturbation parameters corresponding to the current perturbation adjustment method is restricted within a preset cooling period; the adjustment step size is changed again, and a new round of random perturbation adjustment is performed on the neighborhood independent event sequence set according to the updated preset adjustment range, and the multiple adjustment neighborhood event sequence sets are iteratively updated according to the adjustment results of the new round.

[0057] In this embodiment, if multiple adjustment coordination stability scores are all lower than the first coordination stability score, each adjustment coordination stability score is read one by one and compared with the first coordination stability score. If all comparison results are lower than the first coordination stability score, it indicates that the current disturbance adjustment method has not improved the coordination alignment between cash flow and business flow. Subsequently, specific disturbance parameters corresponding to the current disturbance adjustment method are read. These specific disturbance parameters include the disturbance object, disturbance direction, disturbance count, and adjustment step size. The adjustment step size is the number of neighboring independent event sequences that can be moved, replaced, removed, or rearranged in a single random disturbance adjustment. Based on the specific disturbance parameters, a disable count label is added to the current disturbance adjustment method, and the multiple adjustment coordination stability scores generated in one round of random disturbance adjustment are used as a judgment object. When the multiple adjustment coordination stability scores generated in this round are all lower than the first coordination stability score, the disable count label corresponding to the current disturbance adjustment method is incremented by one, and the completion time of this round of judgment is recorded as the starting point of the disable cycle. Within the preset cooling period, subsequent random perturbation adjustments will no longer call the current perturbation adjustment method and its corresponding specific perturbation parameters that have been marked with a "disabled times" tag, thereby avoiding the same inefficient perturbation parameter from repeatedly participating in the adjustment of the neighborhood independent event sequence set.

[0058] After completing the restrictions of the current perturbation adjustment method, the current adjustment step size is read and changed in conjunction with the number of times the perturbation was disabled. If the number of times the perturbation was disabled increases, the current adjustment step size is reduced by one unit. If the reduced adjustment step size is less than one unit, the adjustment step size is maintained at one unit, and the perturbation object or direction is changed. An updated preset adjustment magnitude is formed based on the changed adjustment step size. This updated preset adjustment magnitude limits the number of neighborhood independent event sequences that can be adjusted in the new round of random perturbation adjustment. Then, a new round of random perturbation adjustment is performed on the set of neighborhood independent event sequences according to the updated preset adjustment magnitude, ensuring that each round of adjustment only changes the number of neighborhood independent event sequences within the limit set by the updated preset adjustment magnitude, resulting in a new round of adjustment results. Subsequently, the set of adjusted neighborhood event sequences generated by the restricted perturbation adjustment method is deleted from multiple sets of adjusted neighborhood event sequences, and the new round of adjustment results is added to multiple sets of adjusted neighborhood event sequences. Thus, multiple sets of adjusted neighborhood event sequences are iteratively updated based on the new round of adjustment results.

[0059] Step S500: Convert the multiple independent structured element event chains into distributed accounting voucher data and record them into the account balance sheet. Then, call the accounting account hierarchy tree to aggregate the amounts upwards and automatically generate and render the target financial statement.

[0060] In this embodiment, when converting multiple independent structured event chains into distributed ledger voucher data, the counterparty entity, business occurrence time, business behavior, core transaction amount, and associated input and output tax amount in each structured event chain are read one by one, and the corresponding debit and credit accounts are matched according to the business behavior. Subsequently, the core transaction amount and associated input and output tax amount are written into the corresponding journal entries, generating distributed ledger voucher data containing voucher number, event chain identifier, business occurrence time, account code, debit / credit direction, journal entry amount, tax amount, and counterparty entity. The system performs a balance check on the total debit and credit amounts under the same structured element event chain. If the total debit amount matches the total credit amount, the distributed ledger voucher data is recorded in the account balance sheet according to the account code, accounting period, and debit / credit direction, and the current period debit amount, current period credit amount, and ending balance of the corresponding account are updated. If the total debit amount does not match the total credit amount, the distributed ledger voucher data is marked as voucher data to be checked, and the difference amount, corresponding event chain identifier, and abnormal account code are recorded. At the same time, an abnormal prompt message is generated and sent to the financial audit end, where financial personnel verify the information and decide whether to correct it and record it in the account balance sheet.

[0061] Next, the accounting subject hierarchy tree is invoked to aggregate and summarize the amounts upwards. The current period's debit and credit amounts and ending balances of each bottom-level subject in the account balance sheet are read. Then, according to the parent-child node relationship in the accounting subject hierarchy tree, the amounts of the bottom-level subjects are accumulated to the upper-level subjects level by level to obtain the total amount for each level of subjects. Subsequently, based on the mapping relationship between subjects and report rows in the accounting subject hierarchy tree, the corresponding subject amounts are written into the corresponding rows of the target financial statements. For total rows, the amounts of their corresponding lower-level rows are added to obtain the total amount; for net rows, the difference between the debit and credit balances is filled in. When rendering the target financial statements, the row number, item name, and amount fields in the report template are read. The generated report row amounts are filled into the corresponding cells, and the amounts are processed by unifying decimal places, displaying thousands of percent, and padding zeros for empty values, forming a displayable target financial statement.

[0062] In summary, the embodiments of this application have at least the following technical effects: This application acquires the target enterprise's basic accounting configuration information for digital modeling, constructs an accounting subject hierarchy tree, and retrieves a pre-built general financial dictionary to supplement the subjects and report rows in the accounting subject hierarchy tree with compliant definitions and vectorize them, obtaining multiple subject hierarchy vectors; it captures business interaction documents input from the front-end business system in real time, and combines unstructured summary text and semi-structured auxiliary accounting fields in the business interaction documents into target business text vectors; it performs cross-level cross-attention matrix calculation based on the multiple subject hierarchy vectors and the target business text vectors to obtain multiple subject relevance scores, filters irrelevant text based on the multiple subject relevance scores, and locks the target financial candidate entity set; it traverses the target financial candidate entity set, constructs a directed acyclic graph of element associations based on the financial debit and credit sequence reconciliation logic, and uses a path expansion algorithm to search and decouple multiple independent structured element event chains in the directed acyclic graph of element associations; it converts the multiple independent structured element event chains into distributed accounting voucher data and records them into the account balance sheet, and calls the accounting subject hierarchy tree to aggregate the amounts upwards, automatically generating and rendering the target financial statements. This invention addresses the technical problem of low efficiency in financial statement generation due to the difficulty in accurately matching business documents with accounting subjects in existing technologies. By constructing a hierarchical tree of accounting subjects and combining vectorized coding, cross-level cross-attention calculation, and decoupling of element event chains, it achieves the technical effect of improving the accuracy and efficiency of automatic financial statement generation.

[0063] Example 2 is based on the same inventive concept as the rule-driven automatic financial statement generation method in the previous examples, such as... Figure 2 As shown, this application provides a rule-driven automatic financial statement generation system. The system and method embodiments in this application are based on the same inventive concept. The system includes: Encoding module 11 is used to acquire the target enterprise's basic accounting configuration information for digital modeling, construct an accounting subject hierarchy tree, and retrieve a pre-set general financial dictionary to supplement the subjects and report rows in the accounting subject hierarchy tree with compliant definitions and vectorize them, obtaining multiple subject hierarchy vectors; Conversion module 12 is used to capture business interaction documents input from the front-end business system in real time, and convert the unstructured summary text and semi-structured auxiliary accounting fields in the business interaction documents into target business text vectors; Filtering module 13 is used to perform cross-level cross-attention matrix calculation based on the multiple subject hierarchy vectors and the target business text vectors, obtaining The system obtains relevance scores for multiple subjects, filters irrelevant text based on these scores, and locks in a target set of financial candidate entities. A traversal module 14 traverses the target set of financial candidate entities, constructs a directed acyclic graph (DAG) of element associations based on the sequential reconciliation logic of financial loans, and uses a path expansion algorithm to search for and decouple multiple independent structured element event chains within the DAG. A summary and aggregation module 15 converts these independent structured element event chains into distributed accounting voucher data, records it in the account balance sheet, and calls the accounting subject hierarchy tree to perform upward summarization and aggregation of amounts, automatically generating and rendering the target financial statements.

[0064] Furthermore, the system is also used to implement the following functions: Extract the category name of each final-level account code in the accounting subject hierarchy tree and the text title of the standard financial statement item row number; retrieve the standard financial standard definition and accounting attribute description corresponding to the category name and the text title from the pre-set general financial dictionary, and use them as compliance terms to concatenate the text to obtain the subject extended text; use a dense vectorized encoder to perform high-dimensional feature space mapping on the subject extended text, and output multiple subject-level vectors that correspond one-to-one with each subject and report row number.

[0065] Furthermore, the system is also used to implement the following functions: The multiple subject-level vectors are aggregated to construct a query retrieval matrix, with the target business text vector as the key-value matrix. The dot product of the query retrieval matrix and the transpose of the key-value matrix is ​​calculated to obtain an initial similarity matrix. The initial similarity matrix is ​​scaled and normalized to output a cross-level cross-attention weight matrix, and the corresponding weight distribution values ​​in the cross-level cross-attention weight matrix are used as the relevance scores of multiple subjects.

[0066] Furthermore, the system is also used to implement the following functions: The relevance scores of the multiple subjects are iterated through, and non-core business paragraphs with scores below the preset relevance filtering threshold are removed. Text paragraphs with scores above or equal to the relevance filtering threshold are extracted and assembled into a set of golden text sentences. The golden text sentence set is then scanned using a named entity recognition component to extract multiple attribute entities. Multiple attribute entities belonging to the same text sentence are associated and encapsulated to obtain multiple attribute entity combinations. These multiple attribute entity combinations constitute the target financial candidate entity set.

[0067] Furthermore, the system is also used to implement the following functions: Each of the multiple attribute entities is one of the following: counterparty entity, core transaction amount, associated input and output tax amount, and transaction occurrence time.

[0068] Furthermore, the system is also used to implement the following functions: Retrieve preset financial reconciliation topology rules, and divide each attribute entity in the target financial candidate entity set into subject nodes, behavior nodes, numerical nodes, or time nodes respectively; establish directed connection edges according to the evolution order in the financial loan reconciliation logic, wherein the subject node is used as the root node of the graph, the time node and the behavior node are used as intermediate logical routing nodes, and the numerical node is used as the leaf tail node, thus constructing a directed acyclic graph of element association.

[0069] Furthermore, the system is also used to implement the following functions: Starting from the root node of the directed acyclic graph (DAG) associated with the elements, a depth-first search is performed along the direction of the directed connecting edges to extract multiple complete topological paths from the root node to the leaf tail node, obtaining multiple initial independent event sequences. Topological features are extracted from these initial independent event sequences to identify attribute intersection types and the number of similar attributes, calculating multiple event complexity indices representing information mixing density. These event complexity indices are then filtered according to a preset document segmentation threshold to map and extract multiple leading complex associated event chains. Based on the attribute entity combinations corresponding to these leading complex associated event chains, adjacent independent event sequences under the same root node are associated and categorized. Adjacency impact correction is performed with the goal of optimal alignment between capital flow and business flow, obtaining multiple sets of corrected element event chains. Based on these sets of corrected element event chains, the initial independent event sequences are merged and pruned to obtain multiple independent structured element event chains.

[0070] Furthermore, the system is also used to implement the following functions: Extract the first leading complex associated event chain and its corresponding first attribute entity combination. Combine its topological position in the graph to identify the cooperative alignment degree of the adjacent neighborhood independent event sequence set under the accrual basis, and obtain the first cooperative stability score. For the neighborhood independent event sequence set, perform multiple random perturbation adjustments according to a preset adjustment span to obtain multiple adjusted neighborhood event sequence sets. Again, identify the cooperative alignment degree of the multiple adjusted neighborhood event sequence sets with the first attribute entity combination to obtain multiple adjusted cooperative stability scores. If there is an item in the multiple adjusted cooperative stability scores that is greater than or equal to the first cooperative stability score, then the adjusted neighborhood event sequence set corresponding to the maximum value is taken as the first corrected element event chain set, and the first corrected element event chain set is added to the multiple corrected element event chain sets.

[0071] Furthermore, the system is also used to implement the following functions: If the cooperative stability scores of the multiple adjustments are all less than the first cooperative stability score, the current perturbation adjustment method is marked with a "disabled times" label, and the invocation of the specific perturbation parameters corresponding to the current perturbation adjustment method is restricted within a preset cooling period; the adjustment step size is changed again, and a new round of random perturbation adjustment is performed on the neighborhood independent event sequence set according to the updated preset adjustment range, and the multiple adjustment neighborhood event sequence sets are iteratively updated according to the adjustment results of the new round.

[0072] It should be noted that the order of the embodiments described above is merely for descriptive purposes and does not represent the superiority or inferiority of the embodiments. Furthermore, the above description focuses on specific embodiments of this specification. The processes depicted in the accompanying drawings do not necessarily require a specific or sequential order to achieve the desired results. In some implementations, multitasking and parallel processing are possible or may be advantageous.

[0073] The above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention in any way. Although the present invention has been disclosed above with reference to preferred embodiments, it is not intended to limit the present invention. Any person skilled in the art can make some modifications or alterations to the above-disclosed technical content to create equivalent embodiments without departing from the scope of the present invention. Any modifications, equivalent changes, and alterations made to the above embodiments based on the technical essence of the present invention without departing from the scope of the present invention shall still fall within the scope of the present invention.

Claims

1. A method for automatic generation of financial statements based on rule driving, characterized in that, The method includes: The basic accounting configuration information of the target enterprise is obtained for digital modeling, an accounting subject hierarchy tree is constructed, and a pre-set general financial dictionary is retrieved to supplement the subjects and report rows in the accounting subject hierarchy tree with compliant definitions and vectorized encoding, so as to obtain multiple subject hierarchy vectors. Real-time capture of business interaction documents input from the front-end business system; conversion of unstructured summary text and semi-structured auxiliary accounting fields in the business interaction documents into target business text vectors. Based on the multiple subject-level vectors and the target business text vector, a cross-level attention matrix is ​​calculated to obtain multiple subject relevance scores. Irrelevant text is filtered and a target financial candidate entity set is locked based on the multiple subject relevance scores. Traverse the target set of financial candidate entities, construct a directed acyclic graph of element associations based on the sequential reconciliation logic of financial lending, and use the path expansion algorithm to search and decouple multiple independent structured element event chains in the directed acyclic graph of element associations. The multiple independent structured element event chains are transformed into distributed accounting voucher data and recorded into the account balance sheet. The accounting account hierarchy tree is then called to aggregate the amounts upwards, automatically generating and rendering the target financial statements. This involves retrieving a pre-built general financial dictionary to supplement the accounts and report rows in the accounting subject hierarchy tree with compliant definitions and perform vectorized encoding, obtaining multiple subject hierarchy vectors, including: Extract the category name of each last-level account code in the accounting subject hierarchy tree, and the text title of the standard financial statement item row number; Retrieve the standard financial standard definitions and accounting attribute descriptions corresponding to the category name and the text title from the pre-set general financial dictionary, and use them as compliance terms to concatenate the text to obtain the account extended text; Using a dense vectorized encoder, the subject extended text is mapped to a high-dimensional feature space, and multiple subject-level vectors are output that correspond one-to-one with each subject and report row.

2. The method of claim 1, wherein the rules are defined in a rule definition file. Based on the multiple subject-level vectors and the target business text vector, a cross-level cross-attention matrix is ​​calculated to obtain multiple subject relevance scores, including: The multiple subject-level vectors are aggregated to construct a query retrieval matrix, and the target business text vector is used as the key-value matrix. Calculate the dot product of the query retrieval matrix and the transpose of the key value matrix to obtain the initial similarity matrix; The initial similarity matrix is ​​scaled and normalized for activation, and a cross-level cross-attention weight matrix is ​​output. The corresponding weight distribution values ​​in the cross-level cross-attention weight matrix are used as the relevance scores of multiple subjects.

3. The method of claim 2, wherein the rules are defined by a user. Based on the relevance scores of the multiple subjects, irrelevant text is filtered out and a target set of financial candidate entities is identified, including: Iterate through the relevance scores of the multiple subjects, remove non-core business paragraph texts that are less than the preset relevance filtering threshold, extract text paragraphs that are greater than or equal to the relevance filtering threshold, and assemble them into a set of golden text sentences. The named entity recognition component is used to scan the set of golden text sentences, extract multiple attribute entities, and associate and encapsulate multiple attribute entities belonging to the same text sentence to obtain multiple attribute entity combinations. The multiple attribute entity combinations constitute the target financial candidate entity set.

4. The method of claim 3, wherein the rules are defined in a rule definition file. Each of the multiple attribute entities is one of the following: counterparty entity, core transaction amount, associated input and output tax amount, and transaction occurrence time.

5. The method of claim 1, wherein the rules are defined by a user.

5. The method of claim 1, wherein the rules are defined by a user. Traverse the target set of financial candidate entities and construct a directed acyclic graph of element relationships based on the logical consistency of financial lending and borrowing, including: Retrieve preset financial reconciliation topology rules and divide each attribute entity in the target financial candidate entity set into subject nodes, behavior nodes, value nodes or time nodes respectively. Directed connections are established based on the evolutionary order of the financial lending sequence logic. The main node is used as the root node of the graph, the time node and the behavior node are used as intermediate logical routing nodes, and the numerical node is used as the leaf tail node, thus constructing a directed acyclic graph of element association.

6. The method of claim 5, wherein the rules are defined in a rule definition file. The path expansion algorithm is used to search and decouple multiple independent structured feature event chains in the directed acyclic graph of feature associations, including: Starting from the root node of the directed acyclic graph associated with the elements, a depth-first search is performed along the direction of the directed connecting edges to extract multiple complete topological paths from the root node to the leaf tail node, thereby obtaining multiple initial independent event sequences. Topological features are extracted from the multiple initial independent event sequences to identify attribute cross-types and the number of similar attributes, and multiple event complexity indices representing information mixing density are calculated. The multiple event complexity indicators are filtered according to the preset document segmentation threshold, and multiple chains of events leading to complex co-occurrence are mapped and extracted. Based on the attribute entity combination corresponding to the multiple leading complex associated event chains, the adjacent independent event sequences under the same root node are associated and arranged, and the adjacent influence is corrected with the goal of optimal alignment between capital flow and business flow, so as to obtain a set of multiple correction element event chains. Based on the multiple sets of modified feature event chains, the multiple initial independent event sequences are merged and pruned to obtain multiple independent structured feature event chains.

7. The method of claim 6, wherein the rules are defined in a rules file, and the rules file is updated by a user. With the goal of achieving optimal alignment between cash flow and business flow, adjacency impact correction is performed, resulting in a set of event chains for multiple correction elements, including: Extract the first leading complex associated event chain and its corresponding first attribute entity combination, and combine it with its topological position in the graph to identify the cooperative alignment degree of the neighboring independent event sequence set under the accrual basis, and obtain the first cooperative stability score. The set of independent event sequences in the neighborhood is subjected to multiple random perturbations according to a preset adjustment span to obtain multiple adjusted neighborhood event sequence sets; The multiple sets of adjustment neighborhood event sequences are then combined with the first attribute entity to identify the degree of coordination alignment, and multiple adjustment coordination stability scores are obtained. If there is an item in the plurality of adjustment collaborative stability scores that is greater than or equal to the first collaborative stability score, then the set of adjustment neighborhood event sequences corresponding to the maximum value is taken as the first correction element event chain set, and the first correction element event chain set is added to the plurality of correction element event chain sets.

8. The method of claim 7, wherein the rules are defined in a rules file, and the rules file is updated by a user. The method further includes: If the multiple adjustment coordination stability scores are all less than the first coordination stability score, the current disturbance adjustment method is marked with a number of times it is disabled, and the call of the specific disturbance parameters corresponding to the current disturbance adjustment method is restricted within a preset cooling period. The specific disturbance parameters include the disturbance object, disturbance direction, number of disturbances, and adjustment step size. The adjustment step size is changed again, and a new round of random perturbation adjustment is performed on the set of independent event sequences in the neighborhood according to the preset adjustment range. The multiple adjusted neighborhood event sequence sets are then iteratively updated based on the results of the new round of adjustment.

9. A rule-driven based financial statement auto-generation system characterized in that, The system is used to execute the rule-driven automatic financial statement generation method as described in any one of claims 1-8, and the system includes: The encoding module is used to obtain the basic accounting configuration information of the target enterprise for digital modeling, construct an accounting subject hierarchy tree, and call up a pre-set general financial dictionary to supplement the subjects and report rows in the accounting subject hierarchy tree with compliant definitions and vectorize them to obtain multiple subject hierarchy vectors. The conversion module is used to capture business interaction documents input from the front-end business system in real time and convert the unstructured summary text and semi-structured auxiliary accounting fields in the business interaction documents into target business text vectors. The filtering module is used to perform cross-level cross-attention matrix calculation based on the multiple subject-level vectors and the target business text vector to obtain multiple subject-related scores, filter irrelevant text based on the multiple subject-related scores, and lock the target financial candidate entity set. The traversal module is used to traverse the target financial candidate entity set, construct a directed acyclic graph of element associations based on the sequential reconciliation logic of financial lending, and use the path expansion algorithm to search and decouple multiple independent structured element event chains in the directed acyclic graph of element associations. The aggregation module is used to convert the multiple independent structured element event chains into distributed accounting voucher data and record it into the account balance sheet, and call the accounting account hierarchy tree to aggregate the amount upwards, automatically generating and rendering the target financial statement.

Citation Information

Patent Citations

  • Visual display method for multi-dimensional incidence relation of financial statements

    CN120336603A

  • Financial statement generation method and system based on cloud computing

    CN120764505A