Financial statement intelligent standardization processing method and system

By building a hierarchical standard knowledge base and a weighted sequence alignment algorithm, combined with the structural characteristics of financial statements and the relationships between items, the problems of format differences and typos in financial statement item names are solved, achieving standardized processing with high accuracy and efficiency.

CN120765411APending Publication Date: 2025-10-10SHANGHAI JIAOTONG UNIV
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202511157661.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-19
Publication Date
2025-10-10

AI Technical Summary

Technical Problem

The existing financial reporting system has low matching accuracy when dealing with format differences, typos, and OCR recognition errors in project names, and lacks fault tolerance and error correction mechanisms, resulting in high manual processing costs and low efficiency.

Method used

A hierarchical standard knowledge base is constructed, and weighted sequence alignment algorithm and context-aware rule matching are adopted. The structural characteristics of financial statements and the relationship between projects are combined to achieve the standardized processing of project names.

Benefits of technology

The matching accuracy of financial statement item names has been improved to over 95%, processing efficiency has been increased by 20 times, and labor costs have been significantly reduced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120765411A_ABST
    Figure CN120765411A_ABST
Patent Text Reader

Abstract

The invention provides a financial statement intelligent standardization processing method and system, and the method comprises the steps: constructing a hierarchical standard knowledge base which comprises a standard item name reference system and a feature vector; by taking the standard knowledge base as a reference, matching an original financial statement item name with the standard item name by adopting a weighted sequence comparison algorithm to realize preliminary matching; a matching result is optimized by adopting rule matching of context awareness; and constructing a mapping relation table of the original items and the standard items, and converting the original financial statements according to the mapping relation table to finish standardized processing of the financial statements. The problems of format differences, wrongly written characters, OCR identification errors and the like of financial statement item names can be effectively solved, the matching accuracy rate reaches 95% or above, the processing efficiency is improved by 20 times, the labor cost is remarkably reduced, and reliable technical support is provided for automatic processing and analysis of financial data.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of data intelligence, in particular, to a financial statement intelligent standardization processing method and system. BACKGROUND

[0002] Financial statements are important manifestations of the financial position and operating results of enterprises, mainly including balance sheets, profit statements, cash flow statements, etc. Although accounting standards have some specifications for the format of financial statements, in actual application, there are significant differences in the project name expressions of financial statements of different enterprises. At present, financial statements have gradually become digital and intelligent, and many inconsistent expressions of project names cause inconvenience in the specific application process of digital data, and also waste a lot of manpower for checking, affecting the quality and efficiency of financial statement preparation.

[0003] With the development of digital technology, there have been many research results on automatic processing of financial statements. However, the traditional method mainly has the following problems / one of them:

[0004] 1. Limitation of accurate matching: existing systems mostly use accurate string matching, which cannot handle project names containing format differences. For example, "1. Currency funds" and "currency funds" cannot be matched.

[0005] 2. Insufficient ability to correct errors: traditional methods cannot recognize the similarity between "accounts receivable" and "accounts receivable".

[0006] 3. Lack of error handling for OCR recognition: existing technologies lack a fault-tolerant mechanism for OCR recognition errors, such as "assets" being recognized as "assets".

[0007] 4. Lack of context-related analysis: existing methods mostly match independent projects, without considering the structural characteristics of financial statements and the positional relationship between projects.

[0008] 5. High cost of manual processing: a large number of non-standard financial statements need to be checked and corrected one by one by manual, which is inefficient and prone to errors.

[0009] During the search, it was found that there are also public methods and systems for automatically generating corporate financial statements in the prior art. For example, one of the methods includes: obtaining corporate financial data and financial reporting items in corporate financial statements; performing preliminary screening of corporate financial data to remove non-financial data; performing belief propagation-based associated sentence extraction on the corporate financial data after preliminary screening, retaining multiple initial financial sentences with financial quadruple; mapping bill predicates to financial reporting items, and merging the initial financial sentences according to the mapping results to obtain first-level financial data with the same bill predicate; performing consistency verification on the first-level financial data and corporate financial data; if the verification passes, merging the first-level financial data according to the bill predicate and outputting second-level financial data; filling the second-level financial data into the corresponding financial reporting items in the corporate financial statements, and outputting the obtained target corporate financial statements, thereby improving the efficiency and accuracy of corporate financial statement generation. However, this method mainly solves the problem of generating financial statements from raw financial data, rather than the problem of standardizing existing statements. It assumes that the input is standardized corporate financial data, lacks consideration of the format differences of financial statements in actual business (it does not consider actual data quality issues such as OCR recognition errors and manual entry typos), lacks fault tolerance and error correction mechanisms, and does not consider the processing of project name format differences.

[0010] Based on this, there is an urgent need to study a new method that can use multidimensional engine technology to achieve accurate matching of project names, improve the quality and preparation efficiency of financial statements, and meet the needs of different users. Summary of the Invention

[0011] In response to one of the defects in the existing technology, the purpose of this application is to provide a method and system for intelligent standardization processing of financial statements to solve the technical problems in the existing technology of low accuracy in matching financial statement item names and inability to handle complex format differences.

[0012] In a first aspect of the present application, a method for intelligent standardization processing of financial statements is provided, comprising:

[0013] Constructing a hierarchical standard knowledge base, wherein the standard knowledge base includes a reference system of standard project names and feature vectors;

[0014] Using the standard knowledge base as a reference, a weighted sequence alignment algorithm is used to match the original financial statement item names with the standard item names to achieve a preliminary match;

[0015] Based on the preliminary matching, context-aware rule matching is used to optimize the matching results, taking into account the structural characteristics of the financial statements and the relationships between items;

[0016] Based on the optimization matching results, a mapping relationship table between original items and standard items is constructed, and the original financial statements are converted accordingly to complete the normalization processing of the financial statements.

[0017] Optionally, the standard knowledge base includes establishing a three-layer standard project name system and defining project feature vectors, wherein:

[0018] The three-tier standard project name system is specifically:

[0019] First level: standard project name;

[0020] Second layer: common variant collection;

[0021] The third layer: potential error set;

[0022] The project feature vector is specifically:

[0023] Keyword features: extract the core words of the project name;

[0024] Position characteristics: the typical position of the item in the report;

[0025] Numerical characteristics: the numerical range and characteristics of the item;

[0026] Association characteristics: association relationship with other items.

[0027] Optionally, the use of a weighted sequence alignment algorithm to match the original financial statement item names with the standard item names to achieve a preliminary match includes:

[0028] Pre-process the original financial statement item names to unify the format;

[0029] The pre-processed financial statements are matched using a weighted sequence alignment algorithm to obtain the matching score of each item's feature vector;

[0030] A comprehensive matching score is calculated based on the matching scores of the feature vectors of each item to obtain a preliminary matching result.

[0031] Optionally, the pre-processing of the original financial statement item names includes:

[0032] Remove all serial numbers from project names;

[0033] Remove the brackets and the contents within the brackets from the project name;

[0034] Normalize characters.

[0035] Optionally, the pre-processed financial statements are matched using a weighted sequence alignment algorithm to obtain a matching score for each item's feature vector, including:

[0036] Create a score matrix with the initial value set to 0;

[0037] The matrix is ​​filled using a dynamic programming method, wherein for each character position in the two strings to be compared, the matching score of the current position is calculated and the maximum value is taken according to the following three conditions:

[0038] The score of the current character match between the two strings plus the score of the previous position;

[0039] The score of the character at the current position in the first string and the gap in the second string plus the score of the previous position;

[0040] The score of the missing character in the first string and the character at the current position in the second string plus the score of the character at the left position;

[0041] Finally, the final score of the matrix is ​​divided by the length of the longer of the two strings to obtain a normalized similarity score.

[0042] Optionally, the comprehensive matching score is calculated as follows:

[0043] Score=α×SeqScore+β×KeywordScore+γ×PositionScore

[0044] Among them, αβ and γ are weight coefficients, and α+β+γ=1; SeqScore is the sequence alignment score, KeywordScore is the keyword matching score, and PositionScore is the position matching score.

[0045] Optionally, the context-aware rule matching includes sequentially executing a single matching enhancement rule, multiple intelligent matching rules, and matching verification based on machine learning, wherein:

[0046] The multiple intelligent matching rules use the Hungarian algorithm to solve the optimal matching solution and obtain matching pairs;

[0047] The machine learning validation uses a pre-trained classification model to assess the reliability of the matching pairs.

[0048] Optionally, completing the normalization processing of financial statements includes:

[0049] Construct a mapping table between original project names and standard project names;

[0050] Convert the original financial statements according to the mapping relationship table and unify the project name expressions.

[0051] A second aspect of the present application provides a system for intelligent and standardized processing of financial statements, comprising:

[0052] A standard knowledge base construction module, wherein the constructed standard knowledge base includes a standard project name reference system and feature vectors;

[0053] A preliminary matching module, using the standard knowledge base as a reference, uses a weighted sequence alignment algorithm to match the original financial statement item names with the standard item names to achieve preliminary matching;

[0054] A matching optimization module, based on the preliminary matching, combines the structural characteristics of the financial statements and the relationship between items, and uses context-aware rule matching to optimize the matching results;

[0055] The normalized output module constructs a mapping relationship table between original items and standard items based on the optimized matching results, converts the original financial statements accordingly, and completes the normalized processing of the financial statements.

[0056] A third aspect of the present application provides an application system for intelligent and standardized processing of financial statements, comprising:

[0057] Knowledge base management module, used to maintain the hierarchical standard project name system and feature vector library;

[0058] An intelligent matching engine for executing sequence alignment rules, matching, and machine learning verification; the intelligent matching engine employs the aforementioned intelligent normalization processing method for financial statements or the aforementioned intelligent normalization processing system for financial statements;

[0059] Data preprocessing module, used to format and extract features of input data;

[0060] A quality control module to assess matching quality and generate audit reports;

[0061] The interface service module is used to provide API interface and batch processing functions.

[0062] Optionally, the application system further includes a cache module for storing matched project name mapping relationships.

[0063] The intelligent normalization processing method and system for financial statements provided in this application establish a three-layer standard project name system and project feature vectors by constructing a hierarchical standard knowledge base; perform intelligent matching based on a sequence alignment algorithm, and further implement context-aware rule matching, and finally generate normalized output. It can effectively handle format differences, typos, OCR recognition errors and other problems in financial statement project names, with a matching accuracy rate of over 95%, a 20-fold increase in processing efficiency, and a significant reduction in labor costs, providing reliable technical support for the automated processing and analysis of financial data.

[0064] Other technical effects brought about by the additional features will be further explained in the corresponding embodiments. BRIEF DESCRIPTION OF THE DRAWINGS

[0065] Other features, objects and advantages of the present application will become more apparent upon reading the detailed description of non-limiting embodiments with reference to the following drawings:

[0066] Figure 1 The present invention is a flowchart of a method for intelligent normalization processing of financial statements according to an exemplary embodiment;

[0067] Figure 2 is a flow chart of a sequence alignment algorithm according to an exemplary preferred embodiment;

[0068] Figure 3 is a schematic diagram showing multiple matching rules according to an exemplary preferred embodiment;

[0069] Figure 4 A diagram showing a comparison of matching effects according to an exemplary preferred embodiment;

[0070] Figure 5 The diagram is a diagram of an application system deployment architecture according to an exemplary preferred embodiment. DETAILED DESCRIPTION

[0071] The present application is described in detail below with reference to specific embodiments. The following examples will help those skilled in the art to further understand the present application, but are not intended to limit the present application in any form. It should be noted that, without departing from the concept of the present application, a number of variations and improvements may be made by those skilled in the art, and these all fall within the scope of protection of the present application. Parts not described in detail in the following examples may be implemented using existing technologies.

[0072] Currently, there are still some problems with existing financial statements, such as not considering actual data quality issues such as OCR recognition errors and manual typos, lacking fault tolerance and error correction mechanisms, not considering the processing of project name format differences, and insufficient matching depth and accuracy. To address these issues, the following embodiments of this application provide solutions.

[0073] Reference Figure 1 As shown, the embodiment of the present application provides a method for intelligent normalization processing of financial statements, including the following steps:

[0074] S100, building a hierarchical standard knowledge base, which includes a reference system of standard project names and feature vectors;

[0075] In this step, in order to achieve unified and standardized subsequent processing and identification of financial statements, a standard reference system and feature vectors of three-level standard project names were established to provide a benchmark basis for subsequent matching.

[0076] S200, using the standard knowledge base as a reference, uses a weighted sequence alignment algorithm to match the original financial statement item names with the standard item names to achieve a preliminary match;

[0077] In this step, based on the intelligent matching of the sequence alignment algorithm, the standard knowledge base built by S100 is used as a reference. After preprocessing the original financial statement item names, the similarity score between them and the standard item names is calculated through the weighted sequence alignment algorithm, and the comprehensive matching score is calculated in combination with multi-dimensional features to achieve preliminary matching.

[0078] S300, based on preliminary matching, combines the structural characteristics of financial statements and the relationship between items, uses context-aware rule matching to optimize matching results;

[0079] In this step, context-aware rule matching is used to optimize the matching results and improve matching reliability based on the preliminary matching results of S200, combined with the structural characteristics of the financial statements and the correlation between items, through single matching enhancement rules, multiple intelligent matching rules and matching verification based on machine learning.

[0080] S400: Based on the optimized matching results, a mapping relationship table between the original items and the standard items is constructed, and the original financial statements are converted accordingly to complete the normalization processing of the financial statements.

[0081] In this step, a standardized output is generated. Based on the matching results optimized in S300, a mapping table between the original items and the standard items is constructed. The original financial statements are converted accordingly. Furthermore, a report containing a matching quality assessment can be generated to complete the standardized processing of the financial statements.

[0082] The above embodiment of the present application can effectively handle the format differences, typos, OCR recognition errors and other issues of financial statement item names through steps S100-S400, with a matching accuracy of over 95%, a processing efficiency improvement of 20 times, and a significant reduction in labor costs, thus enabling automated processing and analysis of financial data. Figure 3 shown.

[0083] In some embodiments of the present application, S100 constructs a hierarchical standard knowledge base, wherein:

[0084] S101, establish a three-tier standard project name system, including:

[0085] First level: standard item name (such as "monetary funds");

[0086] Second level: a collection of common variants (e.g., "1. Monetary Funds", "I. Monetary Funds");

[0087] The third layer: potential error sets (such as "currency market funds" and "currency consultation funds");

[0088] In this application, the standard project names adopt the national standardized terminology for financial data to comply with the basic financial reporting standards required by law. The potential error set is mainly a set of typos, which can accurately solve problems such as typos and OCR errors. The common variant set is an extension and supplement to the standard project name, which is conducive to the recognition and unification of the same project name and avoids matching limitations. In addition to being able to handle project names with format differences, it can also match project names with different terms, thereby improving the accuracy of matching.

[0089] The above-mentioned three-tier standard project name system can effectively handle issues such as format differences, typos, and OCR recognition errors in financial statement project names, thereby improving subsequent recognition rates and error tolerance.

[0090] S102: Based on the three-tier standard project name system, a feature vector is defined for each standard project, including:

[0091] Keyword features: extract the core words of the project name;

[0092] Position characteristics: the typical position of the item in the report;

[0093] Numerical characteristics: the numerical range and characteristics of the item;

[0094] Association characteristics: association with other items;

[0095] In the above feature vector, the keyword feature extracts the core vocabulary of the project name, which provides key semantic clues for matching the project name. When faced with complex and diverse financial statement item expressions, keywords can quickly locate the essence of the project. For example, whether it is "cash and cash equivalents" or "monetary assets available at any time for the company", by extracting the two keywords "currency" and "funds", a close connection can be established with the standard item name "cash and cash equivalents". For example, for "accounts receivable (net amount)", the keywords "receivable" and "account" are extracted. Even if there is an error in the word "account", it can be preliminarily associated with the standard item "accounts receivable" based on the keyword, and then the accurate match can be confirmed through the subsequent matching process. The keyword feature greatly improves the pertinence and accuracy of the matching, and reduces the matching difficulties caused by differences in expression.

[0096] Positional features define the typical position of items in financial statements, providing important structural information for matching and standardizing financial statement items. In financial statements, different items often have a relatively fixed positional order. Taking the balance sheet as an example, "cash and cash equivalents" is typically at the forefront of asset items, and "inventory" also has a common position within current assets. By identifying the typical first position of "cash and cash equivalents" in the balance sheet and combining it with other matching factors for standard items, the reliability of the matching is further enhanced. When dealing with items such as "cash outflows from investing activities," which may have incomplete descriptions, the positional feature of the item, typically located after the "subtotal of cash inflows from operating activities" in the cash flow statement, combined with other matching scores, can more accurately confirm its correspondence with the standard items. Positional features help optimize matching results at the report structure level during subsequent sequence alignment and rule matching, improving matching accuracy and stability. They are particularly effective in resolving matching issues for some ambiguous items.

[0097] Numerical features include the numerical range and characteristics of an item, providing a key basis for matching and detecting anomalies in financial statement items. Different financial statement items have specific numerical characteristics. For example, the value of "cash and cash equivalents" generally reflects the company's cash reserves and has a relatively reasonable range within a certain industry and company size. By studying a large amount of corporate income statement data, a normal range model for the values ​​of different items can be established. When encountering a new income statement, if the value of the "operating income" item deviates significantly from the normal range, the system can not only identify that there may be a problem with the item, but also combine other features during the matching process to further confirm its matching relationship with the standard item. Numerical features help to discover anomalies in the data and avoid incorrect matches caused by erroneous data. At the same time, they assist in judgment from the perspective of numerical rationality during the matching process, improving the accuracy of matching and the ability to control data quality.

[0098] The association feature reflects the relationship between an item and other items, providing comprehensive information support for matching financial statement items and overall logical verification. In financial statements, there is a close logical connection between various items. For example, "operating income" and "operating costs" have a matching relationship, and "total assets" is equal to the sum of various asset items. By identifying the correlations between items, the system can use these correlations to cross-validate the matching results of individual items during the matching process. If the matching score of the "cash outflow from investing activities" item is low, but the matching of its related items such as "cash inflow from investing activities" is accurate and overall conforms to the logical relationship of investing cash flows, the association feature can be used to increase the confidence level of the matching of the "cash outflow from investing activities" item. The association feature helps optimize the matching results from the overall logical level of the financial statements, improve the accuracy and reliability of the matching, and can also detect potential logical errors in the reports, which is beneficial to ensuring the logical consistency and data quality of the financial statements.

[0099] Therefore, the feature vector system composed of keyword features, position features, numerical features, and association features in the above-mentioned embodiments of this application provides comprehensive and in-depth support for the intelligent standardization of financial statement items from multiple dimensions, including semantics, structure, numerical rationality, and logical association. Their mutual collaboration can jointly improve the accuracy and stability of subsequent matching and the ability to control data quality, greatly improving matching accuracy and processing efficiency.

[0100] In the above-mentioned embodiments of this application, a special processing mechanism for typos and OCR recognition errors is designed to address various practical application scenarios such as format differences and expression changes in financial statements. Fuzzy matching is achieved through a sequence alignment algorithm, which has strong fault tolerance. A complete alias library is established to handle various expression variations, solving the actual business pain point of inconsistent expression of project names in different companies' reports. It can systematically solve the problem of various format differences in project names, achieve comprehensive format unification through standardized project name libraries and alias libraries, and has the ability to handle complex format variations.

[0101] Reference Figure 2 As shown, in some embodiments of the present application, in S200, the intelligent matching based on the sequence alignment algorithm specifically includes the following steps:

[0102] S201, pre-process the original project name, and perform the following operations in sequence:

[0103] Remove all kinds of serial numbers in project names, such as the numerical serial number "1.", the Chinese serial number "一、", etc.;

[0104] Remove the brackets and the content in the brackets from the item name, for example, change "Accounts Receivable (Net)" to "Accounts Receivable";

[0105] Normalize characters, such as full-width and half-width conversion, and unify special symbols.

[0106] In a specific embodiment, the pre-processing of the original project name in S201 can be implemented using the following computer programming language:

[0107]

[0108] In this application, by preprocessing the above-mentioned project names, the error tolerance capability can be improved, and typos, OCR recognition errors and other problems can be effectively handled, with an error tolerance rate of 85%.

[0109] S202, executing a weighted sequence alignment algorithm:

[0110] Based on S201, follow the steps below:

[0111] First, create a score matrix with the initial value set to 0;

[0112] Then, the matrix is ​​filled using a dynamic programming method: for each character position in the two strings to be compared, the matching score of the current position is calculated and the maximum value is taken according to the following three cases:

[0113] The score of the current character match between the two strings plus the score of the previous position;

[0114] The score of the character at the current position in the first string and the gap in the second string plus the score of the previous position;

[0115] The score of the missing character in the first string and the character at the current position in the second string plus the score of the character at the left position;

[0116] Finally, the final matrix score is divided by the length of the longer of the two strings to obtain a normalized similarity score.

[0117] The present application can significantly improve the matching accuracy. In some embodiments, by combining a sequence alignment algorithm with contextual rules, the matching accuracy is increased from 65% of traditional methods to over 95%.

[0118] In a specific embodiment, the execution of the weighted sequence alignment algorithm in S202 can be implemented using the following computer programming language:

[0119]

[0120] In a specific embodiment, the matching score can be calculated using a character similarity matrix:

[0121] Exact match: +2.0;

[0122] Similar characters: +1.5 (e.g. "账" and "账");

[0123] Homophone matching: +1.0;

[0124] Mismatch: -1.0.

[0125] Of course, in other embodiments, other matching scores may be used and are not limited to the above limitations.

[0126] S203, calculate the comprehensive matching score Score:

[0127] Score=α×SeqScore+β×KeywordScore+γ×PositionScore

[0128] Among them, α, β, and γ are weight coefficients, and α+β+γ=1; SeqScore is the sequence alignment score, KeywordScore is the keyword matching score, and PositionScore is the position matching score.

[0129] In a specific embodiment, α, β, and γ are weight coefficients, which can be selected as 0.5, 0.3, and 0.2, respectively, i.e., the comprehensive matching score Score = 0.5 × sequence alignment score + 0.3 × keyword matching score + 0.2 × position matching score. Of course, in other embodiments, other weight coefficients can also be used, and are not limited to the above limitations.

[0130] In some embodiments of the present application, the context-aware rule matching in S300 specifically includes:

[0131] S301, single matching enhancement rule, includes the following operations:

[0132] When the items at positions i and i+2 in the standard report have been successfully matched, but the item at position i+1 has not been matched:

[0133] S3011, extract the unmatched items at the corresponding position of the original report;

[0134] S3012, calculating the edit distance and semantic similarity between the matching item and the standard item;

[0135] S3013: If the comprehensive score exceeds the set threshold (eg, 0.75), the project is confirmed to be successfully matched.

[0136] In this embodiment, the single item matching enhancement rule is mainly used to optimize the matching of items at specific positions in the financial statements. This rule comes into play when the items at adjacent positions (such as positions i and i+2) in the standard report have been successfully matched, but the item at the middle position (i+1) has not yet been matched. It calculates the edit distance and semantic similarity between the item and the standard item by extracting the unmatched items at the corresponding position of the original report. The edit distance is used to measure the minimum number of operations required for two strings to be converted into the same string after inserting, deleting, and replacing characters, and the semantic similarity considers the degree of similarity between the two from a semantic level. If the score calculated by combining these two dimensions exceeds the set threshold (for example, 0.75), it can be confirmed that the unmatched item has been successfully matched with the corresponding standard item. By adopting this rule, the structural correlation between the adjacent positions of the financial statement items can be utilized, which can effectively solve the problem of missing matches caused by subtle differences in the item descriptions or some interference information. For example, in balance sheet processing, "cash and cash equivalents" and "inventory" have been matched. The "accounts receivable" in the middle may have scored low during the initial matching due to errors in the word "account". However, through the single-item matching enhancement rules, combined with its editing distance and semantic similarity with the standard item "accounts receivable", the match was finally confirmed, which greatly improved the accuracy and completeness of the single-item matching, thereby improving the success rate of the overall report item matching.

[0137] S302, multiple intelligent matching rules, refer to Figure 3 As shown, the following steps are included:

[0138] First, a similarity matrix is ​​constructed for the two sets of item lists, and the similarity between each original item and each standard item is calculated;

[0139] Then, the Hungarian algorithm is used to solve the optimal matching solution and find the project matching combination with the highest overall similarity;

[0140] Finally, the matching quality is verified and only matching pairs (item matching combinations) with similarity exceeding a threshold are retained.

[0141] In a specific embodiment, the multiple intelligent matching rules of S302 can be implemented using the following computer programming languages:

[0142]

[0143] The multiple intelligent matching rules in the above embodiments of the present application can optimize the matching effect of multiple items in the financial statements from a holistic level. It first constructs a similarity matrix for the original report item list and the standard item list. In this matrix, the similarity between each original item and each standard item is calculated in detail, and the various characteristic dimensions of the project name are comprehensively considered. Then, the Hungarian algorithm is used to solve the optimal matching solution. As a classic algorithm for solving the maximum matching problem of bipartite graphs, the Hungarian algorithm can find the project matching combination that makes the overall similarity reach the highest in this similarity matrix, ensuring the best matching effect from the overall perspective of multiple projects. Finally, the matching quality is strictly verified, and only matching pairs whose similarity exceeds a pre-set threshold are retained. This rule fully takes into account the mutual relationship and overall structural characteristics between financial statement items, and avoids isolated errors that may occur when matching single items. For example, in the batch processing of income statements, faced with the task of matching items for a large number of income statements in different formats, a number of intelligent matching rules have been optimized overall to effectively solve the matching difficulties of some items caused by format differences, ambiguous expressions and other issues, greatly improving the accuracy and stability of item matching in batch report processing, enabling the system to efficiently and accurately process large-scale financial statement data, and ensuring the efficiency and reliability of standardized processing of financial statements.

[0144] S303, matching verification based on machine learning;

[0145] After the multiple intelligent matching rules of S302 are executed, the matching verification based on machine learning includes the following steps:

[0146] S3031, extracting feature vectors of matched item pairs

[0147] S3032, use pre-trained classification models to evaluate the reliability of matching results

[0148] S3033: Mark low-confidence matching results as requiring manual confirmation

[0149] In the aforementioned S303 of this application, machine learning-based match verification plays a key role in quality control and intelligent optimization throughout the entire intelligent normalization process for financial statements. After preliminary matching results are obtained through sequence alignment and rule matching, this step extracts feature vectors for the matched item pairs. These feature vectors encompass multi-dimensional information such as the item name's keywords, location, value, and association with other items. A pre-trained classification model is then used to assess the reliability of the matching results. The classification model has been trained on a large amount of historical financial statement data and has learned the pattern characteristics of accurate and incorrect matches. For low-confidence matches, the system will mark them and prompt for manual confirmation. This approach, on the one hand, leverages the powerful learning and prediction capabilities of the machine learning model to automatically detect potential incorrect matches, reducing the likelihood of incorrect matches entering the final normalized report and improving the accuracy and reliability of the matching results. On the other hand, for low-confidence matches that are difficult to accurately determine through the model, a manual confirmation step is introduced, achieving human-machine collaboration, leveraging the efficiency of machine processing while leveraging the flexibility and experience of human judgment. For example, in the processing of special items in the cash flow statement, for some item matching results that have slight differences in expression but are difficult to fully determine through conventional matching rules, machine learning-based matching verification can effectively identify potential risks, ensure the accuracy of cash flow statement item matching, and ultimately provide a solid guarantee for the generation of high-quality standardized financial statements, ensuring that the entire financial statement intelligent standardization processing system can operate stably and reliably, and meet the strict requirements for the accuracy of financial statement data in actual business.

[0150] In the above embodiments of the present application, matching verification based on machine learning is adopted, and the matching strategy is continuously optimized through the machine learning model, which can improve the adaptive learning ability and the system performance continues to improve with the use time.

[0151] The existing intelligent matching algorithms in financial statement processing lack technical depth. For example, the adoption of semantic-based belief propagation and quadruple extraction lacks a precise character-level matching algorithm, making it impossible to match project names. This leaves a lack of effective means for processing similar but not identical project names. To address these insufficient matching issues, the above-mentioned embodiments of the present application employ a sequence alignment algorithm based on dynamic programming to achieve precise character-level matching, design a multi-level matching mechanism (main matching + rule matching), cover various matching scenarios, and solve multiple matching problems through an optimization algorithm. This can significantly improve matching accuracy while increasing processing efficiency.

[0152] In some embodiments of the present application, generating a normalized output in S400 specifically includes:

[0153] S401, constructing a mapping relationship table between original project names and standard project names;

[0154] S402, converting the original financial statement according to the mapping relationship table, and unifying the project name expression;

[0155] S403, generating a quality evaluation report, including: matching success rate, contribution degree of each matching method, low confidence matching list, and possible manual review project suggestions.

[0156] The above embodiments of the application can significantly improve the matching accuracy, have strong error tolerance, and improve the adaptive learning ability. At the same time, the processing efficiency is improved, and compared with manual processing, the efficiency is improved by more than 20 times, and the processing time of a single report is reduced from 30 minutes to 1.5 minutes. It can greatly reduce the operating cost and reduce the demand for 90% of manual intervention, which can greatly save the labor cost.

[0157] Based on the same technical concept, in another embodiment of the application, a financial statement intelligent standardization processing system is provided, comprising: a standard knowledge base construction module, a preliminary matching module, a matching optimization module, and a standardization output module, wherein: the standard knowledge base constructed by the standard knowledge base construction module comprises a standard project name reference system and a characteristic vector; the preliminary matching module matches the project name of the original financial statement with the standard project name by using a weighted sequence comparison algorithm with reference to the standard knowledge base, to realize preliminary matching; the matching optimization module optimizes the matching result by using context-aware rule matching based on the preliminary matching, in combination with the structural characteristics of the financial statement and the correlation between projects; and the standardization output module constructs a mapping relationship table between the original project and the standard project based on the optimized matching result, and converts the original financial statement accordingly, to complete the standardization processing of the financial statement.

[0158] In the above embodiment of the financial statement intelligent standardization processing system, the specific implementation techniques of each module can refer to the corresponding steps in the financial statement intelligent standardization processing method, which will not be described here.

[0159] Based on the same technical concept, in another embodiment, the application provides a financial statement intelligent standardization processing application system, comprising:

[0160] The knowledge base management module: maintains a hierarchical standard project name system and a characteristic vector library;

[0161] The intelligent matching engine: realizes sequence comparison, rule matching and machine learning verification; the intelligent matching engine adopts the financial statement intelligent standardization processing method of any one of the above embodiments or runs the financial statement intelligent standardization processing system of the above embodiments;

[0162] The data preprocessing module: formats and extracts features from the input data;

[0163] Quality control module: evaluates matching quality and generates audit reports;

[0164] Interface service module: provides API interface and batch processing functions.

[0165] In specific applications, the application system for intelligent normalization processing of financial statements also includes a cache module for storing matched project name mapping relationships to improve the processing efficiency of duplicate data.

[0166] In this embodiment, the knowledge base management module is responsible for maintaining a hierarchical system of standard item names and a library of feature vectors, providing a benchmark reference for the entire system's matching process. By building a comprehensive standard reference system that covers various scenarios, such as formatting differences, typos, and OCR errors in financial statement items, it provides a unified and detailed basis for subsequent matching, ensuring matching accuracy from the source and reducing matching deviations caused by missing standards.

[0167] In this embodiment, the intelligent matching engine serves as the core processing module of the system, integrating three technical means: sequence alignment, rule matching, and machine learning verification. Based on the standard data of the knowledge base, the weighted sequence alignment algorithm is executed on the pre-processed original project names to calculate the similarity, and the results are optimized by combining context rules (such as single-item enhancement, multiple Hungarian algorithm matching), and the matching reliability is verified by the machine learning model. It can achieve accurate recognition of complex name differences, effectively handle format differences, typos, OCR errors and other problems, and increase the matching accuracy from 65% of traditional methods to more than 95%.

[0168] In this embodiment, the data preprocessing module formats and extracts features from the input raw financial statement data, specifically including operations such as removing serial numbers, bracketed content, and character normalization, eliminating format interference and extracting key features of the project (such as keywords and location information), thereby reducing noise in the raw data, unifying the data format, and providing clean and standardized input data for the subsequent intelligent matching engine, reducing the processing difficulty of the matching algorithm, and indirectly improving matching efficiency and accuracy.

[0169] In this embodiment, the quality control module assesses match quality and generates an audit report that includes the match success rate, the contribution of various matching methods, a list of low-confidence matches, and recommended items for manual review. This enables full-process quality monitoring of standardized processing results, identifying strengths and weaknesses in the matching process. By flagging low-confidence items, manual review focuses on key issues, reducing the need for manual intervention by 90% while ensuring the reliability of the final output and preventing erroneous data from entering downstream analysis.

[0170] In this embodiment, the interface service module provides an API interface and batch processing capabilities, supporting system integration with external platforms (such as enterprise financial systems and financial institution audit platforms). It also allows users to batch import multiple financial statements for centralized processing, improving the system's practicality and scalability. The API interface enables cross-platform data exchange, adapting to the integration needs of different scenarios. The batch processing function significantly improves processing efficiency. For example, on an 8-core server, the average processing time for 1,000 statements is only 1.5 minutes per statement, which is more than 20 times the efficiency of manual processing.

[0171] It can be seen that in the above embodiments of the present application, the modules collaborate to jointly achieve high accuracy (matching rate of more than 95%), strong error tolerance (85% error processing rate) and high efficiency (20 times improvement in manual efficiency) in the standardized processing of financial statements, providing reliable technical support for the automated analysis and application of financial data.

[0172] Figure 4 A comparison diagram of matching effects according to an exemplary embodiment is shown. Figure 4 It can be seen that the effect of the matching rules adopted in this application can reach 96% compared with traditional exact matching (accuracy of 65%) and simple module matching (accuracy of 78%). The accuracy of traditional exact matching is improved by 47.7%, and the processing effect is improved by 20 times, which has huge advantages in matching effect and processing time.

[0173] The preferred features of the above embodiments can be used alone in any embodiment, or in any combination without conflict. In addition, parts not described in detail in the embodiments can be implemented using existing technologies.

[0174] Based on the above description of the technical solution, the present application is further explained below in combination with specific application examples. It should be understood that the following are only some examples and are not intended to limit the present application.

[0175] Application Example 1: Balance Sheet Normalization

[0176] Input data: A company's 2023 balance sheet (partial)

[0177] 1. Monetary capital: 15,234,567.89

[0178] 2. Accounts receivable (net): 8,456,234.12

[0179] 3. Inventory: 12,345,678.90

[0180] Transactional financial consulting: 3,456,789.01

[0181] Processing input data using the intelligent normalization processing method for financial statements in the above embodiment of the present application may include the following steps:

[0182] S11, building a hierarchical standard knowledge base:

[0183] The standard item names built into the system include: cash and cash equivalents, accounts receivable, inventory, and trading financial assets;

[0184] The common variant library includes item names with serial numbers, such as "1. Monetary Funds" and "I. Monetary Funds."

[0185] The potential error database includes common incorrect expressions such as "accounts receivable" and "transactional financial consulting";

[0186] Define a feature vector for each item. For example, the keyword features of "Money and Funds" are "Currency" and "Funds," and the position feature is the first item in the balance sheet, usually adjacent to the "Short-term Investments" item.

[0187] S12, intelligent matching based on improved sequence alignment:

[0188] (1) Preprocessing operations:

[0189] Remove the serial number from "1. Monetary Funds" and process it as "Monetary Funds";

[0190] Remove the brackets from "Accounts Receivable (Net Amount)" and convert it to "Accounts Receivable";

[0191] The original content of "Transactional Financial Consulting" is retained for subsequent processing (because it may contain OCR errors).

[0192] (2) Sequence alignment and matching:

[0193] "Cash and Cash" is compared with the standard item "Cash and Cash", and receives a basic score of 1.0 due to its exact match;

[0194] "Accounts receivable" is compared with the standard item "Accounts receivable". Since "账" and "账" are similar in shape, it receives a basic score of 0.875;

[0195] "Trading financial consulting and production" is compared with the standard item "Trading financial assets". Since "顾问" and "资" are similar in shape, it receives a basic score of 0.75.

[0196] (3) Calculation of comprehensive matching score:

[0197] The combined score of "monetary funds" based on sequence alignment score (1.0), keyword exact match (1.0), and position match (1.0) is 0.5×1.0+0.3×1.0+0.2×1.0=0.98.

[0198] The comprehensive score of "Accounts Receivable" is 0.5×0.875+0.3×1.0+0.2×0.9=0.85;

[0199] The comprehensive score of "transactional financial consulting" is 0.5×0.75+0.3×0.9+0.2×0.8=0.78;

[0200] S13, context-aware rule matching:

[0201] Single item matching enhancement: The system detects that "Money" (position 1) and "Inventory" (position 3) have been matched, and the intermediate item "Accounts Receivable" (position 2) passes context verification and is confirmed to match "Accounts Receivable";

[0202] Multiple intelligent matching: Use the Hungarian algorithm to optimize the overall matching of multiple projects to ensure the best overall matching effect;

[0203] Machine learning verification: Verify the matching results of "Transactional Financial Consulting and Production." The system identifies an OCR error where "顾问" is replaced by "资" and confirms that the match is valid.

[0204] S14, generate normalized output:

[0205] (1) Construct a mapping relationship table: Create the corresponding relationships between "1. Monetary Funds" → "Monetary Funds", "Accounts Receivable (Net Amount)" → "Accounts Receivable", and "Trading Financial Assets" → "Trading Financial Assets";

[0206] (2) Convert the original financial statements:

[0207] Standardized balance sheet (part):

[0208] Cash and cash equivalents: 15,234,567.89

[0209] Accounts receivable: 8,456,234.12

[0210] Inventory: 12,345,678.90

[0211] Trading financial assets: 3,456,789.01

[0212] Furthermore, a quality assessment report was generated: the matching success rate was 98%, with the matching confidence level of "trading financial assets" being 0.78, and manual review was recommended.

[0213] In this application example, by combining a sequence alignment algorithm with contextual rules, the matching accuracy rate was increased from 65% with traditional methods to over 95%. It can effectively handle problems such as typos and OCR recognition errors, with an error tolerance rate of 85%. Compared with manual processing, efficiency is increased by more than 20 times, and the processing time for a single report is reduced from 30 minutes to 1.5 minutes. Through continuous optimization of matching strategies through machine learning models, system performance continues to improve with usage time. It reduces the need for manual intervention by 90%, saving millions of yuan in labor costs annually.

[0214] Application Example 2: Income Statement Batch Processing and Performance Testing

[0215] In this application example, a batch processing test is performed on 1,000 corporate income statements in different formats. The test environment is a server with 8-core CPUs and 16GB of memory.

[0216] In this application example, the application system for intelligent standardization of financial statements of this application is adopted, and the processing flow is as follows:

[0217] S21, batch import 1,000 income statements in different formats, including reports from different companies and different years;

[0218] S22, the system automatically performs preprocessing:

[0219] All types of serial number formats are uniformly removed, including digital serial numbers (1., 2.), Chinese character serial numbers (one, two), symbol serial numbers (●, ■), etc.;

[0220] Clean up bracketed content and special marks, such as "(10,000 yuan)", "*" and other supplementary information;

[0221] Character normalization, including full-width and half-width conversion, uppercase and lowercase unification, and standardization of special symbols;

[0222] Batch matching processing: The system executes a sequence alignment algorithm on 1,000 reports in parallel, matching them against a standard knowledge base. Contextual rules are applied to optimize matching results, such as leveraging the fact that "operating income" typically comes before "operating costs." Machine learning models automatically verify low-confidence matches and flag items requiring manual review.

[0223] S23, generates standardized income statement and quality assessment report.

[0224] The test results of this embodiment are as follows:

[0225] Total processing time: 25 minutes (average processing time per report: 1.5 minutes);

[0226] Matching accuracy: 96.3%;

[0227] Items requiring manual review: 3.7%.

[0228] Application Example 3: Special Item Processing in Cash Flow Statement

[0229] Input data: Cash flow statement of a listed company (partial)

[0230] Subtotal of cash inflow from operating activities: 56,234,123.45

[0231] Cash outflow from investing activities: 23,567,890.12

[0232] Net cash flow from financing activities: -8,765,432.10

[0233] Impact of exchange rate changes on cash and cash equivalents: 123,456.78

[0234] The processing process of this application example is as follows:

[0235] S31, pre-processing:

[0236] Remove the "-" symbol from the item name and process it as "Subtotal of cash inflows from operating activities";

[0237] Other items remain in basic format, with only character normalization performed.

[0238] S32, matching process:

[0239] "Subtotal of cash inflow from operating activities" matches the standard items, with an overall score of 0.96;

[0240] "Cash outflow from investing activities" lacks the word "generated", and the score is 0.82 after sequence alignment, and 0.85 after combining the position characteristics;

[0241] "Net cash flows from financing activities" is a perfect match and has a score of 1.0;

[0242] When comparing "the impact of exchange rate changes on cash and cash equivalents" with the standard item "the impact of exchange rate changes on cash and cash equivalents", due to the additional word "amount", the sequence comparison score is 0.92 and the overall score is 0.90.

[0243] S33, context verification:

[0244] The system identifies the correlation between three types of projects: "operating activities," "investing activities," and "financing activities." It confirms the overall structural matching through multiple matching rules, thereby improving the matching confidence of individual projects.

[0245] S34, output result:

[0246] All projects were successfully matched, with a matching success rate of 100%;

[0247] Generate a standardized cash flow statement, with project names fully complying with accounting standards;

[0248] The quality assessment report shows that no manual review is required.

[0249] It can be seen from the above embodiments that the intelligent normalization processing method and system for financial statements provided in this application can effectively handle the normalization problems of different types of financial statements. Whether it is the detailed processing of a single report or the efficient processing of large batches of reports, it can maintain high accuracy and high efficiency, which is significantly better than traditional methods and has broad application prospects and practical value.

[0250] Application Example 4: System Application Deployment

[0251] When deploying specific applications, refer to Figure 5 As shown in the figure, in an application example, the financial statement intelligent standardization processing application system is divided into the interface service layer, core processing layer, support service layer, and infrastructure layer from top to bottom. Specifically:

[0252] 1. The interface service layer, also known as the front-end application layer, serves as the entry point for the system to interact with the outside world and provides three types of interfaces:

[0253] WebAPI: supports calling system functions through web interfaces to meet lightweight interactive needs such as online query and real-time processing.

[0254] Batch processing interface: For large-scale financial report data, it supports one-time import and processing of multiple reports, improving batch operation efficiency.

[0255] Result output interface: used to output normalized results (such as standardized reports, matching reports, etc.) and connect to downstream systems or storage tools.

[0256] 2. The core processing layer, also known as the business logic layer, includes data preprocessing, intelligent matching engine, quality control, and result generation, enabling full-process processing from raw data to standardized results. Specifically:

[0257] (1) Data preprocessing: performing standardized cleaning on the input original financial statements, including format normalization, feature extraction, data cleaning, etc.

[0258] (2) Intelligent matching engine is the core matching logic of the system, integrating multiple technologies to achieve accurate matching, including sequence alignment, rule matching, ML verification (machine learning verification), etc.

[0259] (3) Quality control to ensure the accuracy and reliability of output results.

[0260] Quality control includes confidence assessment, anomaly detection, and manual review. Confidence assessment quantifies the reliability of matching results and flags low-confidence items. Anomaly detection identifies anomalies such as data format errors and logical inconsistencies. Manual review triggers the manual review process for high-risk, low-confidence matches, enabling collaborative human-machine error correction.

[0261] (4) Generate results and complete the final normalized output.

[0262] Mapping conversion: Build a mapping relationship between the original project and the standard project, and convert the report format.

[0263] Report generation: Output quality assessment report (matching success rate, error type distribution, etc.).

[0264] Data export: Supports exporting standardized reports to Excel, JSON and other formats for connection to downstream systems.

[0265] 3. The supporting service layer, namely the data layer, includes the knowledge management library, cache system, log and monitoring, providing basic services and resource support for the core processing layer.

[0266] (1) Knowledge base management, storing standard reference data required for matching, including:

[0267] Standard project library: maintains standard financial project names such as "cash and cash equivalents" and "accounts receivable".

[0268] Alias ​​library: records common variants of items (such as "1. Monetary funds") and incorrect expressions (such as "Monetary consulting funds").

[0269] Feature vector library: stores the keywords, positions, values ​​and other features of the project to provide a benchmark for matching.

[0270] (2) Cache system to improve system performance and response speed:

[0271] Matching cache: Temporarily stores intermediate matching results to avoid repeated calculations and speed up batch processing.

[0272] Result cache: stores historical normalized results and supports fast query and reuse.

[0273] (3) Logging and monitoring to ensure stable system operation and problem tracing:

[0274] Operation log: records detailed logs such as user operations and matching processes for troubleshooting.

[0275] Performance monitoring: Real-time monitoring of system resources (CPU, memory), processing efficiency, and early warning of performance bottlenecks.

[0276] 4. Infrastructure layer, which provides the hardware and network foundation for system operation, including data servers, application servers, and message queues.

[0277] Database server: stores persistent data such as knowledge base data, original reports, and normalized results.

[0278] Application server: deploys the core system programs and runs matching algorithms, rule engines and other logic.

[0279] Message queue: implements asynchronous communication between modules (such as queue scheduling for batch processing tasks), decouples system dependencies, and improves concurrent processing capabilities.

[0280] The above describes some specific embodiments of the present application. It should be understood that the present application is not limited to the specific embodiments described above, and those skilled in the art may make various variations or modifications within the scope of the claims, which do not affect the substantive content of the present application. The above preferred features may be used in any combination as long as they do not conflict with each other.

Claims

1. A method for intelligent standardization of financial statements, characterized in that: include: Constructing a hierarchical standard knowledge base, wherein the standard knowledge base includes a reference system of standard project names and feature vectors; Using the standard knowledge base as a reference, a weighted sequence alignment algorithm is used to match the original financial statement item names with the standard item names to achieve a preliminary match; Based on the preliminary matching, context-aware rule matching is used to optimize the matching results, taking into account the structural characteristics of the financial statements and the relationships between items; Based on the optimization matching results, a mapping relationship table between original items and standard items is constructed, and the original financial statements are converted accordingly to complete the normalization processing of the financial statements.

2. The method for intelligent standardization of financial statements according to claim 1, characterized in that: The standard knowledge base includes establishing a three-layer standard project name system and defining project feature vectors, where: The three-tier standard project name system is specifically: First level: standard project name; Second layer: common variant collection; The third layer: potential error set; The project feature vector is specifically: Keyword features: extract the core words of the project name; Position characteristics: the typical position of the item in the report; Numerical characteristics: the numerical range and characteristics of the item; Association characteristics: association relationship with other items.

3. The method for intelligent standardization of financial statements according to claim 1, characterized in that: The weighted sequence alignment algorithm is used to match the original financial statement item names with the standard item names to achieve a preliminary match, including: Pre-process the original financial statement item names to unify the format; The pre-processed financial statements are matched using a weighted sequence alignment algorithm to obtain the matching score of each item's feature vector; A comprehensive matching score is calculated based on the matching scores of the feature vectors of each item to obtain a preliminary matching result.

4. The method for intelligent standardization of financial statements according to claim 3, characterized in that: The pre-processing of the original financial statement item names includes: Remove all serial numbers from project names; Remove the brackets and the contents within the brackets from the project name; Normalize characters.

5. The method for intelligent standardization of financial statements according to claim 1, characterized in that: The pre-processed financial statements are matched using a weighted sequence alignment algorithm to obtain a matching score for each item's feature vector, including: Create a score matrix with the initial value set to 0; The matrix is ​​filled using a dynamic programming method, wherein for each character position in the two strings to be compared, the matching score of the current position is calculated and the maximum value is taken according to the following three conditions: The score of the current character match between the two strings plus the score of the previous position; The score of the character at the current position in the first string and the gap in the second string plus the score of the previous position; The score of the missing character in the first string and the character at the current position in the second string plus the score of the character at the left position; Finally, the final score of the matrix is ​​divided by the length of the longer of the two strings to obtain a normalized similarity score.

6. The method for intelligent standardization of financial statements according to claim 3, characterized in that: The comprehensive matching score is calculated as follows: Score=α×SeqScore+β×KeywordScore+γ×PositionScore Among them, αβ and γ are weight coefficients, and α+β+γ=1; SeqScore is the sequence alignment score, KeywordScore is the keyword matching score, and PositionScore is the position matching score.

7. The method for intelligent standardization of financial statements according to claim 1, characterized in that: The context-aware rule matching includes sequential execution of single matching enhancement rules, multiple intelligent matching rules, and matching verification based on machine learning, wherein: The multiple intelligent matching rules use the Hungarian algorithm to solve the optimal matching solution and obtain matching pairs; The machine learning validation uses a pre-trained classification model to assess the reliability of the matching pairs.

8. The method for intelligent standardization of financial statements according to claim 7, characterized in that: The completion of the standardized processing of financial statements includes: Construct a mapping table between original project names and standard project names; Convert the original financial statements according to the mapping relationship table and unify the project name expressions.

9. An intelligent standardization processing system for financial statements, characterized by: include: A standard knowledge base construction module, wherein the constructed standard knowledge base includes a standard project name reference system and feature vectors; A preliminary matching module, using the standard knowledge base as a reference, uses a weighted sequence alignment algorithm to match the original financial statement item names with the standard item names to achieve preliminary matching; A matching optimization module, based on the preliminary matching, combines the structural characteristics of the financial statements and the relationship between items, and uses context-aware rule matching to optimize the matching results; The normalized output module constructs a mapping relationship table between original items and standard items based on the optimized matching results, converts the original financial statements accordingly, and completes the normalized processing of the financial statements.

10. An application system for intelligent standardization processing of financial statements, characterized by: include: Knowledge base management module, used to maintain the hierarchical standard project name system and feature vector library; An intelligent matching engine for performing sequence alignment, rule matching, and machine learning verification; the intelligent matching engine employing the intelligent normalization processing method for financial statements described in any one of claims 1 to 8 or the intelligent normalization processing system for financial statements described in claim 9; Data preprocessing module, used to format and extract features of input data; A quality control module to assess matching quality and generate audit reports; The interface service module is used to provide API interface and batch processing functions.

Citation Information

Cited By

  • Multilink financial subject mapping method and system fusing semantic model and rule matching

    CN121579453A