Automatic analysis method and system based on financial statement data processing

By constructing a financial statement data network and employing reinforcement learning strategies, core financial events are identified and accompanying rules are generated. This addresses the issue of synchronized data in financial statement analysis, improving the accuracy of anomaly identification and data integrity.

CN122387973APending Publication Date: 2026-07-14湖南省地质灾害调查监测所(湖南省地质灾害应急救援技术中心)
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
湖南省地质灾害调查监测所(湖南省地质灾害应急救援技术中心)
Filing Date
2026-06-12
Publication Date
2026-07-14

AI Technical Summary

Technical Problem

Existing financial statement analysis systems struggle to identify whether data should appear synchronously after a core financial event occurs, leading to omissions in anomaly detection.

Method used

Construct a financial statement data network, identify core financial event nodes, generate a set of accompanying rules, adjust the rule weights through reinforcement learning strategies, map the accompanying data nodes, form a void constraint surface, and perform integrity, delay, splitting, and volatility compression verification to generate void backtracking paths and anomaly chains.

Benefits of technology

It improves the data integrity and traceability of financial statement analysis, identifies data gaps that are difficult to find in traditional analysis, reduces the probability of invalid alarms and false alarms, and can quickly locate the source and scope of anomalies.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122387973A_ABST
    Figure CN122387973A_ABST
Patent Text Reader

Abstract

The application provides an automatic analysis method and system based on financial statement data processing, which comprises the following steps: obtaining financial statement data to be analyzed, standardizing the report data, the text of notes and the data source records, extracting a set of financial data elements, and constructing an actual financial statement data network; identifying core financial event nodes in the actual financial statement data network, generating a set of accompanying rules, adjusting the rule weights in combination with a reinforcement learning strategy model, generating a set of accompanying data node sets and event attribution relationships; mapping the set of accompanying data node sets to the actual financial statement data network, generating accompanying hit records for the hit data, generating negative evidence placeholder nodes for the unmatched positions, and forming a hollow constraint surface; based on the hollow constraint surface, performing completeness, delay, splitting and fluctuation compression verification on the negative evidence placeholder nodes to generate negative evidence hollow records.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of financial data processing technology, and in particular to an automatic analysis method and system based on financial statement data processing. Background Technology

[0002] As businesses expand, the number of account levels, period data, amount reconciliation, cross-report relationships, and notes in financial statements increases. It becomes difficult to detect anomalies hidden among multiple reports by relying solely on manual verification or analysis of a single indicator.

[0003] Most existing financial statement analysis methods identify anomalies through data standardization, account classification, indicator calculation, trend comparison, and reconciliation checks. They can detect issues such as significant fluctuations in account amounts, inconsistencies in totals between different levels of accounts, and financial indicators exceeding thresholds. However, their analysis primarily focuses on data already present in the statements, emphasizing the matching of existing data. In actual financial statement processing, a core financial event is usually accompanied by changes in multiple data points.

[0004] For example, revenue recognition is usually accompanied by changes in accounts receivable, contract liabilities, taxes payable, operating cash inflows, or cost of goods sold; asset impairment is usually accompanied by changes in the asset's carrying amount, impairment losses, and notes to the financial statements. If the system only analyzes existing data without determining whether accompanying data that should appear after a core financial event has occurred has appeared simultaneously, it is prone to omissions and anomalies. Therefore, this invention proposes an automated analysis method and system based on financial statement data processing. Summary of the Invention

[0005] The purpose of this invention is to address the shortcomings of existing technologies by providing an automated analysis method and system based on financial statement data processing, thereby resolving the technical problems mentioned in the background section.

[0006] To achieve the above objectives, the present invention provides the following technical solution: An automated analysis method based on financial statement data processing includes the following steps: S1. Obtain the financial statement data to be analyzed, standardize the statement data, notes text and data source records, extract the set of financial data elements, and construct the actual financial statement data network based on the hierarchical relationship of accounts, the corresponding relationship of periods, the reconciliation relationship of amounts and the cross-statement relationship; S2. Identify core financial event nodes in the actual financial statement data network, generate a set of accompanying rules based on historical period data, industry report templates, enterprise business types, accounting reconciliation rules and semantic notes, and adjust the rule weights by combining reinforcement learning strategy models to generate a set of accompanying data nodes and event attribution relationships. S3. Map the set of accompanying data nodes to the actual financial statement data network, generate accompanying hit records for the hit data, generate negative evidence placeholder nodes for the unmatched positions, and form a void constraint surface based on period constraints, account constraints and amount constraints. S4. Based on the void constraint, perform integrity, delay, splitting and fluctuation compression checks on the negative evidence occupant node to generate a negative evidence void record. S5. Generate a void backtracking path based on the void records of negative evidence, connect the core financial event nodes, upstream source nodes and downstream impact nodes to form a void-type anomaly chain, and output the automatic analysis results of financial statement anomalies.

[0007] S1 specifically includes: acquiring the financial statement data to be analyzed; aligning the header names, account columns, period columns, amount columns, and notes paragraph titles according to a preset report field table; unifying the period granularity, amount unit, and account names; deleting duplicate rows, blank rows, and invalid identifier rows; and generating a standardized report dataset. It also involves extracting financial data elements from the standardized report dataset, binding element identifiers, report categories, period information, account names, amount values, debit / credit direction, notes paragraph positions, and data source positions to generate a set of financial data elements; and establishing account hierarchy, period correspondence, amount reconciliation, and cross-report relationships based on the set of financial data elements to construct the actual financial statement data network.

[0008] S2 specifically includes: reading the account name, period information, amount value, debit / credit direction, and related edges of financial data elements in the actual financial statement data network; calculating the event trigger score based on the preset event identification table; identifying the financial data elements that reach the preset trigger threshold as core financial event nodes; using the event type, triggering account, triggering period, and related edge list of the core financial event nodes as the retrieval entry point; calling historical period data, industry report templates, accounting reconciliation rules, and semantic notes to generate a set of accompanying rules; and adjusting the rule weights using a reinforcement learning strategy model; and generating a set of accompanying data nodes and event attribution relationships based on the core financial event nodes and the set of accompanying rules.

[0009] S3 specifically includes: mapping the set of accompanying data nodes to the actual financial statement data network according to accompanying accounts, corresponding periods, report categories, and trigger sources; generating accompanying hit records for matching results that reach the preset hit threshold; and generating negative evidence placeholder nodes for non-hit locations; establishing period constraints, account constraints, and amount constraints based on core financial event nodes, the set of accompanying data nodes, accompanying hit records, and negative evidence placeholder nodes, and generating a set of constraint relationships; and classifying the accompanying data nodes, accompanying hit records, negative evidence placeholder nodes, and constraint relationship sets corresponding to the same core financial event node into the event constraint region according to the event attribution relationship, forming a void constraint surface.

[0010] S4 specifically includes: performing integrity verification on negative evidence placeholder nodes based on void constraints; marking non-occurring negative evidence nodes when the review search still fails to find the node, the period constraint, account constraint, and amount constraint are valid, and the corresponding rule weight reaches the preset verification threshold; performing delayed verification and split verification on negative evidence placeholder nodes that are not directly found, determining whether the accompanying data appears late or is dispersed to multiple similar accounts, footnote paragraphs, or report positions, and marking delayed negative evidence nodes or split negative evidence nodes; performing fluctuation compression verification on accompanying accounts, marking nodes with amount changes lower than the expected fluctuation range as compressed negative evidence nodes, and summarizing to generate negative evidence void records.

[0011] S5 specifically includes: using a void record of negative evidence as the trigger input, and the corresponding negative evidence placeholder node as the starting point of the path, tracing back upstream source nodes and downstream impact nodes along the data source edge, account level edge, period corresponding edge, amount reconciliation edge, and cross-report correlation edge in the actual financial statement data network to generate a void backtracking path; connecting core financial event nodes, void records of negative evidence, negative evidence placeholder nodes, upstream source nodes, and downstream impact nodes according to the void backtracking path to generate a void-type anomaly chain and determining the chain strength; generating automatic financial statement anomaly analysis results based on the void-type anomaly chain, and outputting anomaly cause explanations and evidence source locations according to review priority scoring and sorting.

[0012] Automated analysis systems based on financial statement data processing include: The data network construction module is used to acquire financial statement data to be analyzed, standardize the statement data, notes text and data source records, and extract the set of financial data elements. The accompanying node generation module is used to identify core financial event nodes in the actual financial statement data network and generate a set of accompanying rules based on historical period data, industry report templates, enterprise business types, accounting reconciliation rules and semantic notes. The void constraint forming module is used to map the set of accompanying data nodes to the actual financial statement data network and generate accompanying hit records for the hit data; The void verification record module is used to perform integrity, delay, splitting, and fluctuation compression verification on negative evidence occupant nodes based on void constraints, and generate negative evidence void records. The anomaly chain output module is used to generate a void backtracking path based on the void records of negative evidence, and connect the core financial event nodes, upstream source nodes and downstream impact nodes to form a void-type anomaly chain.

[0013] The beneficial effects of this invention are as follows: This invention constructs a network of actual financial statement data, unifying and linking statement data, footnote text, and data source records. This enables the identification and tracking of core financial events, accompanying accounts, period relationships, and monetary reconciliation within the same data structure, improving the data integrity and traceability of automated financial statement analysis. By generating a set of accompanying data nodes around core financial event nodes, it not only determines whether existing data is abnormal but also whether the accompanying data should have appeared effectively after the core financial event occurred. This allows it to identify data gaps that are difficult to detect with traditional indicator analysis, such as "data that should have appeared but did not."

[0014] This invention distinguishes and verifies the absence, delayed appearance, split appearance, and fluctuation compression of accompanying data by using negative evidence placeholder nodes and void constraint surfaces. This avoids simply assuming all missing data is anomaly, thus improving the accuracy of anomaly identification. A reinforcement learning strategy model is introduced to adjust the weights of accompanying rules based on historical void verification results, manual review results, and anomaly hit results. This allows the set of accompanying data nodes to be continuously optimized with review feedback, reducing the probability of invalid alarms and false alarms.

[0015] This invention generates void backtracking paths and void-type anomaly chains based on void records of negative evidence. It can simultaneously output abnormal financial events, missing accompanying accounts, the location of evidence sources, and the location of affected financial statements, facilitating reviewers to quickly locate the source and scope of the anomaly. It can identify anomalies where the overall financial statements appear balanced and the monetary relationships are not obviously incorrect, but there are gaps in the accompanying data. This improves the ability to detect hidden financial anomalies such as premature revenue recognition, delayed expense recognition, splitting of receivables into accounts, and cost transfer presentation. Attached Figure Description

[0016] Figure 1 This is a schematic diagram of the automatic analysis method based on financial statement data processing according to the present invention; Figure 2 This is a schematic diagram of the framework of the automatic analysis system based on financial statement data processing according to the present invention. Detailed Implementation

[0017] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0018] Example 1: As Figure 1 As shown, this embodiment provides an automated analysis method based on financial statement data processing, including the following steps: S1. Obtain the financial statement data to be analyzed, standardize the statement data, notes text and data source records, extract the set of financial data elements, and construct the actual financial statement data network based on the hierarchical relationship of accounts, the corresponding relationship of periods, the reconciliation relationship of amounts and the cross-statement relationship; S2. Identify core financial event nodes in the actual financial statement data network, generate a set of accompanying rules based on historical period data, industry report templates, enterprise business types, accounting reconciliation rules and semantic notes, and adjust the rule weights by combining reinforcement learning strategy models to generate a set of accompanying data nodes and event attribution relationships. S3. Map the set of accompanying data nodes to the actual financial statement data network, generate accompanying hit records for the hit data, generate negative evidence placeholder nodes for the unmatched positions, and form a void constraint surface based on period constraints, account constraints and amount constraints. S4. Based on the void constraint, perform integrity, delay, splitting and fluctuation compression checks on the negative evidence occupant node to generate a negative evidence void record. S5. Generate a void backtracking path based on the void records of negative evidence, connect the core financial event nodes, upstream source nodes and downstream impact nodes to form a void-type anomaly chain, and output the automatic analysis results of financial statement anomalies.

[0019] S1 specifically includes the following sub-steps: S110. Obtain the financial statement data to be analyzed. The financial statement data to be analyzed includes balance sheet data, income statement data, cash flow statement data, statement of changes in equity data, notes to the financial statements, and data source records. Data source records refer to the record information that can trace the source of financial data, including the original file name, worksheet location, row and column position, export batch, and export time.

[0020] After reading the financial statement data to be analyzed, the system first aligns the header name, account column, period column, amount column, and notes paragraph title according to the preset report field table; then it unifies the period information to one period granularity of year, quarter, or month, and unifies the amount unit to the same unit of measurement, such as "yuan" or "ten thousand yuan".

[0021] For account names, the system performs normalization based on the preset account mapping table, uniformly mapping expressions such as "accounts receivable" and "accounts receivable balance" to the same standard account name "accounts receivable". For the notes text, the system retains its chapter title, paragraph number, page number and corresponding report item name, so as to establish cross-report relationships in the future.

[0022] After completing field alignment and normalization, the system deletes duplicate rows, blank rows, and invalid identifier rows. Duplicate rows are those where the original file name, worksheet position, row and column position, account name, period information, and amount value are all identical. Blank rows are those where both the account name and amount value are empty. Invalid identifier rows are those in headers and footers, unit descriptions, annotations, and formatting prompts that are not involved in amount calculations.

[0023] After the above processing, a standardized report dataset is generated. A standardized report dataset refers to a collection of financial statement data with consistent field formats, period granularity, account names, and monetary units, while retaining records of the data source. It serves as the sole input for S120 to extract financial data elements, and the original financial statement data to be analyzed will no longer be directly used in network construction.

[0024] S120. Extract financial data elements based on standardized report datasets. A financial data element is the smallest data object that can be included in the actual financial statement data network. Each financial data element corresponds to a specific report item, period, amount, and source location.

[0025] The system reads the standardized report dataset line by line, generates a financial data element for each valid financial record, and binds an element identifier, account name, report type, period information, amount, debit / credit direction, footnote paragraph position, and data source position to this financial data element. The element identifier is used to uniquely identify the financial data element in subsequent identification, mapping, verification, and backtracking processes, and it is generated by the report type, period information, account name, data source position, and row / column position.

[0026] To ensure that element identifiers are reproducible, they can be generated in the following way:

[0027] in, Let H be the element identifier of the i-th financial data element, and let H be the hash mapping function. For the report category of the i-th financial data element, For the period information of the i-th financial data element, Let i be the account name of the i-th financial data element. The data source location for the i-th financial data element. This represents the row and column position of the i-th financial data element.

[0028] The debit / credit direction is preferentially determined by the existing debit / credit direction field in the standardized report dataset. If the field does not exist, it is determined based on the account attribute and whether the amount is positive or negative. Asset and expense accounts default to debit increases, while liability, owner's equity, and revenue accounts default to credit increases. Negative amounts are recorded in the opposite direction. For amounts or risk items in the notes text, the system binds the note paragraph position to its corresponding main table item name, enabling cross-report relationships in subsequent processing. After the above processing, a set of financial data elements is generated.

[0029] Each financial data element in the set of financial data elements has a traceable data source location and a monetary value that can be used for calculation. This set serves as the basis for S130 to establish hierarchical relationships between accounts, corresponding relationships between periods, reconciliation relationships between amounts, and cross-report relationships.

[0030] S130. Construct an actual financial statement data network based on the set of financial data elements. The actual financial statement data network refers to a data network consisting of financial data elements as nodes and defined relationships between financial data elements as edges. It is used to uniformly carry out the identification of subsequent core financial event nodes, the mapping of accompanying data node sets, and the generation of negative evidence placeholder nodes.

[0031] The system first establishes a hierarchical relationship between accounts based on a preset account system. If the account name corresponding to one financial data element is a subordinate account of the account name corresponding to another financial data element, an account hierarchy edge is established between them. For example, a subordinate-to-superior account hierarchy edge is established between "Accounts Receivable" and "Current Assets". The system secondly establishes a period correspondence relationship based on period information. If two financial data elements have the same account name and their periods are adjacent, or belong to corresponding items in different reports within the same period, a period correspondence edge is established between them.

[0032] The system then establishes a reconciliation relationship based on the monetary value and the debit / credit direction. This reconciliation relationship indicates the consistency between the sum of amounts in the higher-level account and the amounts in the lower-level account. The determination method is as follows:

[0033] Where D is the difference in monetary reconciliation. This refers to the monetary value of the superior account. Let be the amount value of the j-th sub-item, and n be the number of sub-items. An amount reconciliation boundary is established when the amount discrepancy D is not greater than the preset tolerance threshold. The preset tolerance threshold can be set according to the amount unit and report precision; for example, it can be set to 0.01 million yuan when the amount unit is ten thousand yuan.

[0034] The system also establishes cross-report relationships based on report categories and note paragraph positions. If a financial data element corresponds to a certain account in the main table, and there are descriptions of the same account name, the same period, or the same data source location in the note text, then a cross-report relationship edge is established between the main table node and the note node.

[0035] In a real financial statement data network, each node must include at least a node identifier, financial data element, report type, period information, account name, amount, and data source location; each edge must include at least an edge type, starting node, ending node, relationship source, and relationship weight. Relationship weights are determined based on the degree of account matching, period matching, and report correlation, and are used to prioritize the identification of strongly related financial events in subsequent analyses.

[0036] After establishing the aforementioned nodes and edges, the system outputs the actual financial statement data network. This actual financial statement data network serves as the input for S210 to identify core financial event nodes, and as the basis for S310 to map the accompanying data node set and generate negative evidence placeholder nodes.

[0037] S2 specifically includes the following sub-steps: S210. Identify core financial event nodes in the actual financial statement data network. Core financial event nodes are financial data nodes that can cause one or more accompanying accounts to change synchronously, and are used to represent financial events such as revenue recognition, cost transfer, cash inflow, tax accrual, asset impairment, accounts receivable formation, inventory transfer out, or contract liabilities to revenue.

[0038] The system takes each financial data element in the actual financial statement data network as the object to be judged, reads its account name, report type, period information, amount value, debit / credit direction, account level edge, amount reconciliation edge and cross-report association edge, and determines whether it has financial event triggering characteristics according to the preset event identification table.

[0039] The preset event identification table is used to record the triggering accounts, triggering reports, debit / credit directions, and necessary related edges corresponding to various financial events. For example, the triggering accounts for revenue recognition events include operating revenue and main business revenue, the triggering report is the income statement, the debit / credit direction is an increase on the credit side, and it should be directly or indirectly related to accounts receivable, contract liabilities, cash received from sales of goods, or tax accrual items; the triggering accounts for asset impairment events include asset impairment losses, credit impairment losses, and impairment explanations in the notes, and it should be related to a decrease in the amount of the relevant asset account.

[0040] In one specific embodiment, the preset event identification table is stored using a relational data structure, and its fields include at least: event type ID, set of main triggering accounts (e.g., ["Main Business Revenue", "Other Business Revenue"]), related report category (e.g., "Profit and Loss Statement"), debit / credit direction characteristics (e.g., "Credit Amount"), and set of mandatory related paths (e.g., ["Accounts Receivable_Debit", "Contract Liabilities_Debit", "Taxes Payable_Credit"]). The system extracts matching features by traversing the actual financial statement data network and comparing the node attributes with the above fields.

[0041] The system calculates an event trigger score for each object to be judged, and the calculation method is as follows:

[0042] in, Score the event trigger for the i-th financial data element. For subject path matching values, For the direction of lending, Matching values ​​for changes in amount. To associate and match values ​​across reports, , , , These are the preset weight coefficients for the corresponding matching values.

[0043] To establish the feasibility of this scoring mechanism, in a specific application scenario, the preset weight coefficients can be configured as follows: =0.4 (focusing on subject accuracy), =0.2, =0.3, =0.1. The range of each matching value is normalized to [0,1]. The system's preset trigger threshold is, for example, 0.75, when the event triggers a score. When the value is ≥0.75, core financial event nodes can be accurately identified, and routine data changes can be filtered out.

[0044] The subject path matching value indicates whether the subject name of the financial data element falls within the trigger subject range of the preset event identification table. The debit / credit direction matching value indicates whether the debit / credit direction of the financial data element conforms to the increase / decrease direction of the corresponding event. The amount change matching value indicates whether the change of the financial data element relative to the same subject in the historical period reaches the preset change threshold. The cross-report association matching value indicates whether the financial data element has an association edge with other report items or note paragraphs.

[0045] When the event trigger score reaches the preset trigger threshold, the system identifies the financial data element as a core financial event node and writes the event node identifier, event type, triggering account, triggering period, triggering amount, trigger score, triggering basis, source financial data element, and related edge list to it.

[0046] The list of associated edges includes account-level edges, amount reconciliation edges, and cross-report associated edges that are directly connected to the core financial event node, which serve as input for S220 to generate the candidate set of accompanying rules.

[0047] S220. Generate a set of accompanying rules based on core financial event nodes. Accompanying rules refer to rules used to describe the accompanying accounts, amount directions, and fluctuation requirements that should occur in the corresponding period or adjacent period after a core financial event node occurs.

[0048] The system uses the event type, triggering account, triggering period, triggering amount, and related edge list of core financial event nodes as search entry points, and calls historical period data, industry report templates, enterprise business types, accounting reconciliation rules, and semantic notes to generate a candidate set of accompanying rules.

[0049] Historical period data is used to extract the combination of accounts that actually accompanied similar events in the past period; industry report templates are used to supplement the report items that usually accompany the industry; enterprise business type is used to exclude candidate accounts that are inconsistent with the enterprise's business model; accounting reconciliation rules are used to determine the amount direction and period relationship of accompanying accounts; and the semantics of the notes text are used to supplement accompanying matters that are not directly listed in the main table but have been explained in the notes.

[0050] For example, for revenue recognition events, the candidate set of accompanying rules may include increases in accounts receivable, decreases in contract liabilities, increases in taxes payable, increases in cash received from sales of goods, decreases in inventory, and increases in operating costs; for asset impairment events, the candidate set of accompanying rules may include decreases in the carrying amount of related assets, increases in impairment losses, changes in deferred tax assets, and the appearance of impairment statements in the notes. The system further introduces a reinforcement learning strategy model to filter and weight the candidate set of accompanying rules.

[0051] The reinforcement learning strategy model refers to a strategy model that adjusts the weights of accompanying rules based on historical feedback results. Its state includes the event type, trigger amount, trigger period, enterprise business type, and current candidate set of accompanying rules for core financial event nodes; its actions include retaining candidate rules, deleting candidate rules, and adjusting the weights of candidate rules; and its reward value is jointly determined by historical void verification results, manual review results, and abnormal hit results.

[0052] For each candidate rule, the system updates the rule weight as follows:

[0053] in, Let be the current rule weight of the accompanying rule in round t. For the updated rule weights, To learn step length, The reward value for round t is formed by the historical void verification results, manual review results, and abnormal hit results.

[0054] If a candidate rule corresponds to a real anomaly multiple times in historical hole verification and is confirmed by manual review, its reward value increases; if a candidate rule causes invalid holes or false alarms multiple times, its reward value decreases. The system normalizes the updated rule weights and adds candidate rules whose weights reach a preset retention threshold to the accompanying rule set, while removing candidate rules whose weights are below the preset retention threshold.

[0055] The accompanying rule set includes rule identifier, event type, accompanying account, applicable period, expected amount direction, candidate source, rule weight, and rule generation basis, which serve as the basis for S230 to generate the accompanying data node set.

[0056] To achieve the transformation from qualitative feedback to quantitative parameters, reward value The quantization mapping formula is configured as follows:

[0057] in, The historical hole verification feedback value (for example, if the rule successfully constrained a valid hole, it is recorded as +1; if it caused a false constraint, it is recorded as -1; and if there is no constraint, it is recorded as 0). This is the value provided by manual review (for example, +2 is recorded if the manual review confirms it as a genuine anomaly, and -2 is recorded if it is determined to be a reasonable fluctuation in business). The abnormal hit feedback value (e.g., if the accompanying rule ultimately points to a material financial misstatement, it is recorded as +1, otherwise as 0). , , The corresponding preset reward coefficients are set to 0.3, 0.5, and 0.2 in one embodiment, respectively. Through the above mapping, the model can convert expert experience into updatable rule weights. The mathematical basis.

[0058] S230. Generate a set of accompanying data nodes based on the core financial event node and the set of accompanying rules. The accompanying data nodes refer to the data nodes that should appear or fluctuate accordingly in the actual financial statement data network after the occurrence of the core financial event node, according to the accompanying rules. They are used to match the financial data elements in the actual financial statement data network in the subsequent S310.

[0059] The system centers on each core financial event node, reading its event node identifier, event type, trigger period, trigger amount, and trigger account, and then calling the accompanying rules corresponding to its event type for each event. For each accompanying rule, the system generates a corresponding data node and writes the data node identifier, the corresponding core financial event node identifier, the accompanying account, the corresponding period, the expected amount direction, the expected fluctuation range, the trigger source, and the rule weight.

[0060] The corresponding period is determined based on the triggering period of the core financial event and the applicable period in the accompanying rules. If the rules require it to occur simultaneously in the current period, the corresponding period is the same as the triggering period. If the rules require it to be reflected in a subsequent period, the corresponding period is extended according to the lag period recorded in the rules. The expected amount direction is determined based on the increase or decrease direction of the account in the accompanying rules. For example, after a revenue recognition event is triggered, the expected amount direction of accounts receivable can be an increase, and the expected amount direction of contract liabilities can be a decrease.

[0061] The expected fluctuation range is determined based on the mean and standard deviation of the amount changes of similar accompanying items during the historical period, calculated as follows:

[0062] in, This represents the lower bound of the expected amount for the i-th accompanying data node. The upper limit of the expected amount for the i-th data node that should accompany it. This represents the average change in amount for similar accompanying items over a historical period. This represents the standard deviation of the amount changes for similar accompanying items during the historical period. A preset range coefficient is used to control the tolerance for reasonable business fluctuations. In one embodiment, based on the statistical principle of normal distribution, The possible values ​​are 1.96 (corresponding to a 95% confidence interval) or 2.58 (corresponding to a 99% confidence interval), which covers the vast majority of normal business amount changes and locks data changes outside this range as unexpected fluctuations.

[0063] If historical data is insufficient to calculate the mean and standard deviation of change, the system uses the default range from the industry report template as the expected fluctuation range and marks the source of the default range in the corresponding accompanying data node. The system also establishes event attribution relationships, including the core financial event node identifier, the accompanying data node identifier, the source of the accompanying rule, and the rule weight, to ensure that each accompanying data node belongs only to its corresponding core financial event node. After the above processing, the system outputs the set of accompanying data nodes and the event attribution relationships.

[0064] The accompanying data node set and event attribution relationship are used as inputs to the S310 mapping to the actual financial statement data network to determine whether the accompanying data node can match the corresponding financial data element in the actual financial statement data network, and generate negative evidence placeholder nodes at the unmatched positions.

[0065] S3 specifically includes the following sub-steps: S310. Map the set of accompanying data nodes to the actual financial statement data network to generate accompanying hit records or negative evidence placeholder nodes. Negative evidence placeholder nodes are placeholder data nodes generated when a core financial event node has occurred and, according to the accompanying rules, corresponding accompanying data should exist, but no valid financial data element is found in the actual financial statement data network. Candidate financial data elements are financial data elements in the actual financial statement data network that can be matched.

[0066] The system reads the set of accompanying data nodes and event attribution relationships output by S230, and obtains the corresponding core financial event node identifier, accompanying account, corresponding period, expected amount direction, expected fluctuation range, trigger source and rule weight for each accompanying data node. Then, it searches for candidate financial data elements in the actual financial statement data network using accompanying account, corresponding period, report type and trigger source as search conditions.

[0067] During the retrieval, the system first performs an exact match. If the account name, period information, report type, and amount direction of the candidate financial data element are consistent with the corresponding accompanying data node, an accompanying hit record is generated. If the exact match fails, an approximate match is performed. Approximate matching allows the account name of the candidate financial data element to belong to the same parent account as the accompanying account, or the period information of the candidate financial data element to be within the lag period allowed by the accompanying rules, but the amount direction must not be opposite to the expected amount direction.

[0068] To ensure the matching process is reproducible, the system calculates a matching score:

[0069] in, The matching score between the i-th accompanying data node and the candidate financial data element is given. For subject matching values, For the period matching value, For the direction of the amount, For source matching values, , , , These are preset matching weight coefficients for subject, period, amount direction, and source, respectively.

[0070] In a preferred embodiment, to highlight the core matching dimension, the preset matching weight coefficient is set as follows: =0.5 (subject matching has the highest weight), =0.3, =0.15, =0.05. Meanwhile, the preset hit threshold is set to 0.85. When the match score... When the value is ≥0.85, the system determines that the candidate financial data element is valid accompanying data and generates an accompanying hit record.

[0071] If the matching score reaches the preset hit threshold, the system generates an accompanying hit record. The accompanying hit record includes the identifier of the accompanying data node, the identifier of the hit financial data element, the matching score, the hitting method, the hitting period, the hitting account, and the hitting amount direction. If the matching score does not reach the preset hit threshold, or no candidate financial data element is found, the system generates a negative evidence placeholder node at the corresponding account position, period position, and report position.

[0072] Negative evidence placeholder nodes should include at least the placeholder node identifier, the corresponding core financial event node identifier, the accompanying data node identifier, the accompanying account, the corresponding period, the expected amount direction, the expected fluctuation range, the reason for not hitting, the rule weight, and the placeholder position.

[0073] After the above processing, the system outputs the accompanying hit record and negative evidence placeholder node, which serve as the input for S320 to establish the constraint relationship set.

[0074] S320. Based on the core financial event nodes, the set of accompanying data nodes, the accompanying hit records, and the negative evidence placeholder nodes, establish period constraints, account constraints, and amount constraints, and generate a set of constraint relationships.

[0075] The constraint set refers to the data set used to record whether each accompanying data node meets the time, account, and amount requirements. Its function is to provide a definite judgment basis for S330 to form a void constraint surface. The system first establishes period constraints, which are used to determine whether the corresponding period of the accompanying data node meets the triggering period of the core financial event node and the occurrence time limited by the accompanying rules.

[0076] The offset during the period is determined as follows:

[0077] in, This is the period offset. To correspond to the period of the accompanying data nodes, This refers to the triggering period for core financial event nodes. When the accompanying rule requires the event to occur simultaneously in the current period, the period offset should be zero; when the accompanying rule allows the event to occur in a subsequent period, the period offset should fall within the preset allowed period range.

[0078] The system then establishes account constraints, which are jointly determined by the accompanying accounts in the accompanying rule set, the account hierarchy edges in the actual financial statement data network, and the preset account system. If the accompanying account name is consistent with the account name of the hit financial data element, or if they belong to the same parent account and do not violate the accompanying rule limitations, then the account constraint is valid. The system also establishes amount constraints, which are used to determine whether the direction and range of amount changes of the hit financial data element conform to the expected amount direction and expected fluctuation range of the accompanying data node.

[0079] The change in amount is determined based on the difference between the candidate financial data element in the corresponding period and the base period. The method for determining the amount constraint is as follows:

[0080] in, Let be the change in amount of the i-th candidate financial data element relative to the base period. To correspond to the lower limit of the expected amount for the accompanying data nodes, The expected upper limit of the accompanying data node that should be matched for the i-th candidate financial data element.

[0081] When the change in amount is between the lower and upper limits of the expected amount, and the direction of the change is consistent with the direction of the expected amount, the amount constraint is valid. If the direction of the change is consistent but the change in amount is lower than the lower limit of the expected amount, the data node that should accompany it is marked as a subsequent fluctuation compression verification object. The system writes the period constraint results, account constraint results, and amount constraint results into the constraint relationship set.

[0082] The constraint relationship set includes at least the core financial event node identifier, the accompanying data node identifier, the accompanying hit record identifier or the negative evidence placeholder node identifier, the period constraint result, the account constraint result, the amount constraint result, the constraint source and the constraint status, as inputs for S330 to form a void constraint surface.

[0083] S330. Forming a void constraint surface based on a set of constraint relationships. A void constraint surface refers to a constraint-type data structure formed around the same core financial event node, not an abstract mathematical surface; it is composed of period constraints, account constraints, amount constraints, rule weights, accompanying hit records, and negative evidence placeholder nodes under the same core financial event node, used to limit whether the negative evidence placeholder node has the conditions to enter the subsequent void verification.

[0084] The system categorizes data nodes, associated hit records, negative evidence placeholder nodes, and constraint relationship sets corresponding to the same core financial event node into the same event constraint region based on event attribution. An event constraint region refers to the data range corresponding to only one core financial event node, used to prevent the mixing of associated data nodes and negative evidence placeholder nodes for different core financial event nodes.

[0085] For each event constraint region, the system takes the core financial event node as the center, connects the corresponding accompanying hit record or negative evidence placeholder node according to the accompanying data node identifier, and writes the period constraint result, account constraint result, amount constraint result and rule weight into the connection edge to form a void constraint surface that can be read later.

[0086] A void constraint surface must include at least the constraint surface identifier, the core financial event node identifier, a list of associated accompanying data nodes, a list of associated negative evidence placeholder nodes, a list of associated accompanying hit records, period constraint results, account constraint results, amount constraint results, rule weights, and constraint surface status. If the period constraint, account constraint, and amount constraint corresponding to a negative evidence placeholder node all have valid judgment results, the system marks the constraint surface status of the void constraint surface it belongs to as verifiable; if the necessary constraint results are missing, it is marked as pending supplementation, and it will not proceed to subsequent void verification for the time being.

[0087] After the above processing, the system outputs a void constraint surface. The void constraint surface serves as the common basis for integrity verification in S410, delay and splitting verification in S420, and fluctuation compression verification in S430.

[0088] S4 specifically includes the following sub-steps: S410. Perform integrity verification on negative evidence placeholder nodes based on void constraints. Integrity verification refers to determining whether the corresponding accompanying data does not actually exist in the actual financial statement data network, provided that the core financial event node has occurred, rather than being missed due to differences in account names, period offsets, or differences in source location.

[0089] The system reads the void constraint surface output by S330, obtains the core financial event node identifier, negative evidence placeholder node, accompanying data node, period constraint result, account constraint result, amount constraint result and rule weight, and uses the accompanying account, corresponding period, report type, expected amount direction and trigger source in the negative evidence placeholder node as review conditions to search for financial data elements again in the actual financial statement data network.

[0090] The review retrieval process first searches for financial data elements that are consistent in account name, period information, report type, and amount direction; if no such elements are found, it then searches for similar financial data elements that belong to the same superior account as the accompanying account and whose period is within the range allowed by the accompanying rules.

[0091] If neither of the above two types of searches yields a match, and the period constraints, account constraints, and amount constraints are all valid, and the rule weight reaches the preset verification threshold, then the negative evidence placeholder node is marked as a non-occurring negative evidence node. A non-occurring negative evidence node refers to a negative evidence node that should appear according to the core financial event node and accompanying rules, but for which no valid corresponding data can be found in the actual financial statement data network.

[0092] To ensure the reproducibility of the decision-making process, the system can adopt the following decision-making method:

[0093] in, The non-occurrence judgment value for the i-th negative evidence placeholder node. To match the hit status, use 1 for a miss and 0 for a hit. To ensure the constraints are valid, a value of 1 is assigned when all three constraints—period constraint, account constraint, and amount constraint—are valid; otherwise, a value of 0 is assigned. The rule is considered valid when its weight reaches a preset verification threshold; otherwise, it is set to 0. When the value equals 1, it is marked as a negative evidence node that has not appeared. If the enterprise's business type indicates that the accompanying subject is not applicable, or the rule weight is lower than the preset verification threshold, the system does not directly create a negative evidence gap, but instead marks the negative evidence placeholder node as a rule-pending verification node.

[0094] Negative evidence nodes that have not appeared and rule nodes that need to be reviewed are both retained in the void constraint surface, serving as input for S420 to continue judging delayed appearance and split appearance.

[0095] S420. Based on the void constraint, perform delayed verification and split verification on the negative evidence placeholder node. Delayed verification refers to determining whether the accompanying data did not appear in the corresponding period, but appeared in a subsequent period allowed by the accompanying rules; split verification refers to determining whether the accompanying data did not appear in a single account or a single report position, but was dispersed across multiple approximate accounts, multiple footnote paragraphs, or multiple report positions.

[0096] For negative evidence placeholder nodes that are not directly hit by S410, the system reads the corresponding period, accompanying subject, expected amount direction, expected fluctuation range and rule weight, and determines the delayed search range based on the allowed lag period in the accompanying rules.

[0097] If a financial data element appears after the corresponding period, the system calculates the delay period offset, which is the difference between the actual occurrence period and the corresponding period of the accompanying data node. When the delay period offset is greater than zero and does not exceed the allowable lag period, it is recorded as a delayed hit within the allowable range, and a delayed hit record is generated. When the delay period offset exceeds the allowable lag period, or the delayed occurrence results in a lack of valid accompanying data in the original corresponding period, it is marked as a delayed negative evidence node. Delayed negative evidence nodes are used to indicate that although the accompanying data appears in a subsequent period, it does not meet the synchronization requirements of the corresponding period of the core financial event node.

[0098] For split verification, under the same core financial event node, the system retrieves multiple candidate financial data elements that belong to the same superior account, the same corresponding period or within the allowed lag period, and whose amount direction is consistent with the expected amount direction, and merges their amount change values. The merged amount is determined as follows:

[0099] in, The amount to be split and merged for the i-th data node should be included. Let be the change in amount of the k-th candidate financial data element, where k is the step count variable and m is the number of candidate financial data elements being merged.

[0100] If the direction of the merged amount is consistent with the expected direction and the merged amount falls within the expected fluctuation range, a split hit record is generated; if the merged amount still does not fall within the expected fluctuation range, or if the candidate financial data element does not belong to the category allowed by the accompanying rules, the corresponding negative evidence placeholder node is marked as a split-type negative evidence node.

[0101] Delayed negative evidence nodes, split negative evidence nodes, delayed hit records, and split hit records are all written back to the void constraint surface, serving as the basis for S430 to perform fluctuation compression verification and generate negative evidence void records.

[0102] S430. Based on void constraints, perform fluctuation compression verification on negative evidence placeholder nodes and corresponding accompanying accounts, and generate negative evidence void records. Fluctuation compression verification refers to the situation where, although the financial data elements exist and the direction of the amount is consistent with the expected direction of the amount, the magnitude of their amount change is lower than the expected fluctuation range, which cannot effectively support the verification process of core financial event nodes.

[0103] The system reads accompanying hit records, delayed hit records, split hit records, and negative evidence placeholder nodes that are still in a hit state from the void constraint surface to determine the base period amount and corresponding period amount for each accompanying account. The base period can be the previous accounting period, the same period in history, or a period specified by the accompanying rules. The system subtracts the base period amount from the corresponding period amount to obtain the actual amount change value and compares it with the expected lower limit of the amount in the accompanying data node.

[0104] To facilitate the assessment of compression levels, the system calculates the fluctuation ratio:

[0105] in, Let i be the fluctuation realization ratio of the i-th accompanying subject. Let be the actual change in amount of the financial data element corresponding to the i-th accompanying account. This represents the lower limit of the expected amount for the accompanying data nodes associated with the i-th accompanying item. When the amount direction is consistent with the expected amount direction and the fluctuation realization ratio is lower than the preset compression threshold, the system marks the node corresponding to the accompanying item as a compressed negative evidence node.

[0106] Compressed negative evidence nodes are used to represent accompanying data that exists on the surface, but whose monetary fluctuations are suppressed and cannot form a valid accompanying relationship matching the core financial event node. The system summarizes non-occurring negative evidence nodes, delayed negative evidence nodes, split negative evidence nodes, and compressed negative evidence nodes to generate negative evidence void records.

[0107] Negative evidence gap records refer to data records used to describe the data gaps accompanying core financial event nodes. They must include at least the core financial event node identifier, missing accompanying account, abnormal period, gap formation method, gap impact scope, rule weight, corresponding negative evidence placeholder node position, verification basis, and evidence source location. Evidence source location refers to the original file name, worksheet position, row / column position, footnote paragraph position, or data source location used to support negative evidence gap records or gap-type anomaly chains.

[0108] After the manual review results are returned, the system feeds back the verification results of this step to the reinforcement learning strategy model. The feedback includes the type of negative evidence node, the corresponding rule weight, the manual review result, the anomaly hit result, and the false alarm identifier. This feedback is used to update the rule weights of the accompanying rule set when executing S220 in the next round or subsequent batches. The negative evidence void record serves as the input for backtracking the void-type anomaly chain in S510.

[0109] S5 specifically includes the following sub-steps: S510. Based on the negative evidence void record, generate a void backtracking path in the actual financial statement data network. The void backtracking path refers to a directed path formed by using the negative evidence void record as the trigger input, the corresponding negative evidence placeholder node as the path starting point, and tracing the data source and scope of influence along the defined relationship edges in the actual financial statement data network. It is used to explain where the negative evidence void is triggered and which report positions may be affected.

[0110] The system reads the negative evidence gap records output by S430, obtains the core financial event node identifier, missing accompanying account, abnormal period, gap formation method, rule weight, corresponding negative evidence placeholder node position and evidence source position, and uses the corresponding negative evidence placeholder node as the path starting point.

[0111] The system first traces back to the source nodes upstream. The upstream source nodes refer to the data source nodes that lead to the formation of core financial event nodes and negative evidence placeholder nodes, including the original data source location, the note paragraph location, the trigger account source node, similar nodes in the historical period, and the superior account node that has a monetary correlation with the trigger account.

[0112] During backtracking, the system searches layer by layer along the data source edge, subject level edge, period corresponding edge, amount reconciliation edge, and cross-report association edge; when backtracking to the original file name, worksheet position, row and column position, or note paragraph position, the node is recorded as a traceable source node.

[0113] The system then traces downstream impact nodes, which are report nodes that may be affected by negative evidence gaps. These include affected report items, affected periods, affected amount reconciliation edges, affected cross-report relationship edges, and notes to the financial statements. For example, if accounts receivable corresponding to a revenue recognition event are missing, the system continues to track operating revenue, operating cash inflows, taxes payable, contract liabilities, and notes to the revenue recognition statement.

[0114] To prevent the path from expanding indefinitely, the system sets a preset maximum backtracking level and a list of visited nodes. Backtracking in that direction stops when the path reaches the preset maximum backtracking level, encounters a visited node, or has no valid associated edges. The system calculates a path association score for each candidate path.

[0115] in, Score the path association. For source-related values, For the period-related value, For monetary value For report-related values, , , , These are the preset weighting coefficients corresponding to the source, period, amount, and report association, respectively.

[0116] If the path association score reaches the preset path retention threshold, the system writes it into the holed backtracking path; if it does not reach the preset path retention threshold, it will not be included in the subsequent anomaly chain generation. The holed backtracking path serves as the input for S520 to generate holed anomaly chains.

[0117] S520. Generate a void-type anomaly chain based on the void backtracking path. A void-type anomaly chain refers to an anomaly description chain formed by connecting negative evidence void records, upstream source nodes, and downstream impact nodes around the same core financial event node in the order of event occurrence and report impact. It is used to express the location and direction of the formation of the accompanying data gap.

[0118] The system reads the void backtracking path output by S510, first connects the core financial event node with the void record of negative evidence, then connects the negative evidence placeholder node corresponding to the void record of negative evidence, and then connects the upstream source node and the downstream impact node.

[0119] Nodes within the same period are sorted according to the hierarchical relationship of accounts and the reconciliation of amounts, with lower-level detailed nodes preceding higher-level summary nodes, and higher-level summary nodes preceding the affected report positions; nodes across periods are sorted according to the order of the periods, with the triggering period node preceding the subsequent period node; nodes across reports are connected according to the order of influence of the income statement, balance sheet, cash flow statement, and notes.

[0120] If two void backtracking paths have the same core financial event node, the same missing accompanying account, and the same abnormal period, the system will merge them into the same void-type abnormal chain; if the downstream impact nodes are different, the different downstream impact nodes will be treated as branch records of the same void-type abnormal chain, and independent abnormal chains will not be generated repeatedly.

[0121] The system assigns an anomaly chain identifier, a core financial event node identifier, a missing accompanying account, the period of the anomaly, the method of hole formation, a list of upstream source nodes, a list of downstream affected nodes, a list of affected report locations, and a chain length for each void-type anomaly chain. Chain length refers to the number of valid nodes traversed from the core financial event node to the final downstream affected node, and is used for subsequent calculation of review priority. The system also determines the chain strength based on the rule weights, the number of negative evidence nodes, the number of affected report locations, and the number of valid constraints in the negative evidence void record.

[0122] Chain strength is used to indicate the degree of correlation between the void-type anomalous chain and the financial statement anomalies, and is calculated as follows:

[0123] in, For the chain strength of a void-type anomalous chain, For the rule weights in the void record of negative evidence, The number of nodes containing negative evidence. The number of affected report locations. This represents the number of effective constraints in the void constraint surface. , , , These are the preset weight coefficients corresponding to the rule weight, the number of negative evidence nodes, the number of affected report locations, and the number of valid constraints, respectively.

[0124] The higher the chain strength, the higher the priority the abnormal chain should be for review. After the above processing, the system outputs a hollow abnormal chain, which serves as the basis for the S530 to generate automatic analysis results for financial statement anomalies.

[0125] S530. Generate automatic financial statement anomaly analysis results based on void-type anomaly chains. Automatic financial statement anomaly analysis results refer to the structured analysis results output by the system to reviewers, used to explain abnormal financial events, accompanying data gaps, sources of evidence, and review order.

[0126] The system reads the void-type anomaly chain output by S520, generates analysis results one by one according to the anomaly chain identifier, and writes the core financial event node identifier, abnormal financial event, missing accompanying account, abnormal period, void formation method, related report location, upstream source summary, downstream impact summary, anomaly cause explanation, evidence source location, chain strength, and suggested review entry.

[0127] The explanation of the cause of the anomaly is automatically generated based on the formation method of the void, the type of negative evidence node, the upstream source node, and the downstream impact node. For example, if the formation method of the void is "not present", the missing accompanying item is accounts receivable, and the core financial event node is the revenue recognition node, then the explanation of the cause of the anomaly is recorded as follows: The revenue recognition node exists, but no accompanying account receivable node that meets the rule weight, period constraint, account constraint, and amount constraint was found in the corresponding period.

[0128] The system further calculates the review priority score:

[0129] in, Scoring based on review priority For the rule weights in the void record of negative evidence, The number of affected report locations. The length of the void-type anomaly chain. The severity of the cavity formation method, , , , These are preset weight coefficients corresponding to rule weight, number of affected report locations, anomaly chain length, and severity of hole formation method.

[0130] The severity of void formation can be set according to four types: non-occurring, delayed, split, and compressed. Higher values ​​indicate higher review priority. The system automatically analyzes financial statement anomalies based on review priority scores from highest to lowest, and outputs anomaly chain identifiers, explanations of anomaly causes, and the location of evidence sources, enabling reviewers to directly locate the original documents, report locations, and footnote paragraphs.

[0131] The results of the automatic analysis of financial statement anomalies are the final output of this method, used to indicate financial statement anomalies that have undergone surface balancing but have accompanying data gaps.

[0132] Example 2: Figure 2 As shown, this embodiment provides an automated analysis system based on financial statement data processing, including: The data network construction module is used to acquire the financial statement data to be analyzed, standardize the statement data, notes text and data source records, extract the set of financial data elements, and construct the actual financial statement data network according to the account hierarchy, period correspondence, amount reconciliation and cross-statement relationship. The accompanying node generation module is used to identify core financial event nodes in the actual financial statement data network. It generates a set of accompanying rules based on historical period data, industry report templates, enterprise business types, accounting reconciliation rules, and semantic notes. It also adjusts the rule weights by combining reinforcement learning strategy models to generate a set of accompanying data nodes and event attribution relationships. The void constraint forming module is used to map the set of accompanying data nodes to the actual financial statement data network, generate accompanying hit records for hit data, generate negative evidence placeholder nodes for unmatched positions, and form void constraint surfaces based on period constraints, account constraints, and amount constraints. The void verification record module is used to perform integrity, delay, splitting, and fluctuation compression verification on negative evidence occupant nodes based on void constraints, and generate negative evidence void records. The anomaly chain output module is used to generate a void backtracking path based on the void records of negative evidence, connect the core financial event nodes, upstream source nodes and downstream impact nodes to form a void-type anomaly chain, and output the automatic analysis results of financial statement anomalies.

[0133] All the above formulas are performed using dimensionless numerical calculations; the relevant formulas are based on empirical models that approximate the real situation, obtained through extensive data collection and software simulation fitting. The preset parameters and thresholds involved in the formulas can be conventionally set and adjusted by those skilled in the art according to the physical constraints of the actual application scenario.

[0134] Those skilled in the art will recognize that the modules and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0135] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

[0136] In conclusion, the above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. An automatic analysis method based on financial statement data processing, characterized in that, Includes the following steps: S1. Obtain the financial statement data to be analyzed, standardize the statement data, notes text and data source records, extract the set of financial data elements, and construct the actual financial statement data network based on the hierarchical relationship of accounts, the corresponding relationship of periods, the reconciliation relationship of amounts and the cross-statement relationship; S2. Identify core financial event nodes in the actual financial statement data network, generate a set of accompanying rules based on historical period data, industry report templates, enterprise business types, accounting reconciliation rules and semantic notes, and adjust the rule weights by combining reinforcement learning strategy models to generate a set of accompanying data nodes and event attribution relationships. S3. Map the set of accompanying data nodes to the actual financial statement data network, generate accompanying hit records for the hit data, generate negative evidence placeholder nodes for the unmatched positions, and form a void constraint surface based on period constraints, account constraints and amount constraints. S4. Based on the void constraint, perform integrity, delay, splitting and fluctuation compression checks on the negative evidence occupant node to generate a negative evidence void record.

2. The automatic analysis method based on financial statement data processing according to claim 1, characterized in that, Also includes: S5. Generate a void backtracking path based on the void records of negative evidence, connect the core financial event nodes, upstream source nodes and downstream impact nodes to form a void-type anomaly chain, and output the automatic analysis results of financial statement anomalies.

3. The automatic analysis method based on financial statement data processing according to claim 1, characterized in that, S1 specifically includes: Obtain the financial statement data to be analyzed, align the header name, account column, period column, amount column and note paragraph title according to the preset report field table, unify the period granularity, amount unit and account name, delete duplicate rows, blank rows and invalid identifier rows, and generate a standardized report dataset; Extract financial data elements from standardized report datasets, bind element identifiers, report categories, period information, account names, amount values, debit / credit directions, note paragraph positions, and data source positions to generate a set of financial data elements; Based on the set of financial data elements, establish account hierarchy, period correspondence, amount reconciliation and cross-report relationship to construct the actual financial statement data network.

4. The automatic analysis method based on financial statement data processing according to claim 1, characterized in that, S2 specifically includes: In the actual financial statement data network, the account name, period information, amount value, debit / credit direction and related edge of the financial data elements are read. The event trigger score is calculated according to the preset event identification table, and the financial data elements that reach the preset trigger threshold are identified as core financial event nodes. Using the event type, triggering account, triggering period, and related edge list of core financial event nodes as the retrieval entry point, historical period data, industry report templates, accounting reconciliation rules, and semantic annotation text are called to generate an accompanying rule set, and the rule weights are adjusted using a reinforcement learning strategy model; Generate a set of accompanying data nodes and event attribution relationships based on the core financial event nodes and the set of accompanying rules.

5. The automatic analysis method based on financial statement data processing according to claim 1, characterized in that, S3 specifically includes: The set of accompanying data nodes is mapped to the actual financial statement data network according to the accompanying account, corresponding period, report type and trigger source. Accompanying hit records are generated for matching results that reach the preset hit threshold, and negative evidence placeholder nodes are generated for non-hit positions. Based on the core financial event nodes, the set of accompanying data nodes, the accompanying hit records, and the negative evidence placeholder nodes, establish period constraints, account constraints, and amount constraints to generate a set of constraint relationships.

6. The automatic analysis method based on financial statement data processing according to claim 5, characterized in that, Also includes: Based on the event attribution relationship, the accompanying data nodes, accompanying hit records, negative evidence placeholder nodes, and constraint relationship sets corresponding to the same core financial event node are grouped into the event constraint area to form a void constraint surface.

7. The automatic analysis method based on financial statement data processing according to claim 1, characterized in that, S4 specifically includes: Based on the void constraint, the integrity of the negative evidence placeholder node is checked. When the review and search still fail to find the node, the period constraint, account constraint and amount constraint are valid, and the corresponding rule weight reaches the preset verification threshold, the non-occurring negative evidence node is marked. For negative evidence placeholder nodes that are not directly hit, perform delayed verification and split verification to determine whether the accompanying data appears late or is scattered to multiple similar subjects, footnote paragraphs or report positions, and mark delayed negative evidence nodes or split negative evidence nodes.

8. The automatic analysis method based on financial statement data processing according to claim 7, characterized in that, Also includes: The accompanying items are subjected to fluctuation compression verification. Nodes with changes in amount below the expected fluctuation range are marked as compressed negative evidence nodes, and negative evidence gap records are generated by summarizing them.

9. The automatic analysis method based on financial statement data processing according to claim 2, characterized in that, S5 specifically includes: Using the void record of negative evidence as the trigger input and the corresponding negative evidence placeholder node as the starting point of the path, the upstream source node is traced back and the downstream impact node is tracked along the data source edge, account level edge, period corresponding edge, amount reconciliation edge and cross-report association edge in the actual financial statement data network to generate a void backtracking path. Based on the void backtracking path, core financial event nodes, void records of negative evidence, void nodes of negative evidence, upstream source nodes, and downstream impact nodes are connected to generate void-type anomaly chains and determine the chain strength. The system generates automatic analysis results of financial statement anomalies based on void-type anomaly chains, and outputs explanations of the causes of anomalies and the location of evidence sources according to the review priority scoring.

10. An automatic analysis system based on financial statement data processing, employing the automatic analysis method based on financial statement data processing as described in any one of claims 1 to 9, characterized in that, include: The data network construction module is used to acquire financial statement data to be analyzed, standardize the statement data, notes text and data source records, and extract the set of financial data elements. The accompanying node generation module is used to identify core financial event nodes in the actual financial statement data network and generate a set of accompanying rules based on historical period data, industry report templates, enterprise business types, accounting reconciliation rules and semantic notes. The void constraint forming module is used to map the set of accompanying data nodes to the actual financial statement data network and generate accompanying hit records for the hit data; The void verification record module is used to perform integrity, delay, splitting, and fluctuation compression verification on negative evidence occupant nodes based on void constraints, and generate negative evidence void records. The anomaly chain output module is used to generate a void backtracking path based on the void records of negative evidence, and connect the core financial event nodes, upstream source nodes and downstream impact nodes to form a void-type anomaly chain.