A Machine Learning-Based Method and System for Automatically Generating Financial Statements

By using machine learning technology to perform format and semantic matching of financial data from multiple systems, a data logic association model is constructed, and business processes are dynamically reorganized. This solves the problem of low efficiency in the generation of financial statements in existing technologies and achieves efficient and accurate automatic generation of financial statements.

CN120996007BActive Publication Date: 2026-04-03GEOLOGICAL & NATURAL DISASTER PREVENTION & CONTROL INST GANSU ACADEMY OF SCI
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-06
Publication Date
2026-04-03

AI Technical Summary

Technical Problem

Existing technologies struggle to standardize data processing and establish logical connections across business processes when dealing with complex, multi-system, and multi-source corporate financial data, resulting in inefficient and inaccurate financial statement generation.

Method used

By employing a machine learning-based approach, data parsing and semantic matching are performed through a pre-established format mapping rule base. This constructs a data logical association model, dynamically reorganizes business processes, verifies and integrates data, and ultimately generates compliant financial statements.

Benefits of technology

It enables intelligent integration of data from multiple sources and automatic generation of financial statements, improving the efficiency and accuracy of cross-system data processing and ensuring the integrity and consistency of financial data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120996007B_ABST
    Figure CN120996007B_ABST
Patent Text Reader

Abstract

This application provides a machine learning-based method and system for automatically generating financial statements, comprising: based on preliminary data format difference classification results, using intelligent recognition technology to match and transform the semantics of different fields to generate a standardized data structure set and obtain a unified field mapping relationship; through a complete business data link, using a process tracing algorithm to dynamically reorganize cross-system business processes, obtaining the temporal relationship and data flow path of each business link, and obtaining a complete process tracing map; through the repaired process data set, obtaining key financial indicators and related data of each link, using a data integration module to perform multi-dimensional summarization, and obtaining a preliminary financial data report framework; and through the final financial data content, obtaining a preset report template and output rules to generate a compliant financial report output result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of information technology, and in particular relates to a method and system for automatically generating financial statements based on machine learning. Background Technology

[0002] Financial statement generation is an indispensable part of enterprise management, directly affecting the accuracy of business decisions and the rationality of resource allocation; its importance is self-evident. In modern enterprises, financial data comes from a wide range of complex sources, involving multiple business systems and data formats. Therefore, efficiently integrating and generating accurate reports has become a key area in enterprise digital transformation.

[0003] However, many current solutions often struggle to adapt to the dynamic changes in data from multiple systems and sources within an enterprise when faced with complex data environments. Existing methods rely heavily on manual intervention or fixed rules, failing to flexibly address differences in data structures between different systems and struggling to uncover deep-seated correlations during data integration. This limitation often leads to inefficiency and inaccuracy when generating financial statements for enterprises.

[0004] Against this backdrop, the core challenges facing this field are becoming increasingly apparent. The most pressing issue is the difficulty of data standardization. Due to significant differences in data formats and recording methods across various business systems within an enterprise, systems struggle to automatically identify and unify this heterogeneous data, leading to a complex and error-prone integration process. This problem further exacerbates the difficulty of tracing cross-system business processes. Lacking a unified data foundation, systems cannot accurately establish logical connections between data from different sources, making it exceptionally difficult to trace the full picture of business processes. These two intertwined technical factors jointly hinder the automated transformation from disparate data to unified financial statements.

[0005] Therefore, how to achieve standardized processing of multi-source data through intelligent means, and on this basis, build logical connections across system business processes to complete the automated integration and generation of financial statements, has become a key issue that this research urgently needs to address. Summary of the Invention

[0006] The technical problem to be solved by the present invention is to provide a method and system for automatically generating financial statements based on machine learning.

[0007] To achieve the above objectives, the present invention adopts the following technical solution:

[0008] This invention provides a machine learning-based method for automatically generating financial statements, comprising:

[0009] By using a pre-established format mapping rule library, heterogeneous data from multiple systems is initially analyzed to obtain the field structure and encoding rules of each system's data, and to determine the preliminary data format difference classification results.

[0010] Based on the preliminary data format difference classification results, the semantics of different fields are matched and transformed to generate a standardized set of data structures and obtain a unified field mapping relationship;

[0011] For a standardized set of data structures, the dependencies and business rules between the data of each system are obtained. By constructing a data logic association model, it is determined whether there are any missing or conflicting data across systems. If there are missing or conflicting data, a supplementary rule base is triggered to fill the gaps and determine the complete business data link.

[0012] By using complete business data links, cross-system business processes are dynamically reorganized to obtain the temporal relationships and data flow paths of each business link, resulting in a complete process tracing map.

[0013] Based on the complete process tracing map, the data flow path of each business link is verified. If the path is interrupted or abnormal, the preset threshold rules are called to repair it and determine the repaired process data set.

[0014] By using the repaired process data set, key financial indicators and related data of each stage are obtained, and the data integration module is used to summarize them from multiple dimensions to obtain a preliminary financial data report framework.

[0015] Based on the preliminary financial data reporting framework, a second verification is performed on the completeness and consistency of the summarized data. If a data deviation is detected, a backtracking mechanism is triggered to re-match the business data links and determine the final financial data content.

[0016] Based on the final financial data, pre-defined report templates and output rules are obtained to generate compliant financial report outputs.

[0017] Preferably, the preliminary analysis of heterogeneous data from multiple systems using a pre-established format mapping rule library is performed to obtain the field structure and encoding rules of each system's data, and to determine the preliminary data format difference classification results, including:

[0018] By using a pre-defined format mapping rule library, an initial parsing operation is performed on heterogeneous data from multiple system sources to obtain the field structure and encoding rules of each system's data and determine the preliminary format difference classification results.

[0019] Extract key metadata from the field structure obtained from the initial parsing, determine the completeness and consistency of the field structure, and if there are missing or inconsistent fields, perform field completion and standardization processing through a preset rule base to obtain standardized field structure data.

[0020] Based on the standardized field structure data, a layer-by-layer comparison and analysis is performed on the encoding rules. If there is a deviation between the encoding rules and the preset rule library, the encoding is converted through the mapping rules to obtain a unified encoding rule dataset.

[0021] By using a unified encoding rule dataset, the data formats are compared in depth. If there are significant differences between the data formats, a preset format conversion template is used for adjustment to determine a standardized set of data formats.

[0022] After obtaining a standardized set of data formats, the differences are further refined, and the format differences are automatically classified to obtain detailed difference classification results.

[0023] Based on the detailed difference classification results, a corresponding data integration mapping table is generated, and the format of heterogeneous data from multiple system sources is standardized to obtain the final integrated data result.

[0024] Preferably, based on the preliminary data format difference classification results, the semantics of different fields are matched and transformed to generate a standardized data structure set, resulting in a unified field mapping relationship, including:

[0025] Initial data format information is obtained from the classification results. Preliminary analysis is performed on data from different sources. A pre-established rule base is used to identify the data format, resulting in a preliminary set of format classifications.

[0026] Based on the initial set of format classifications, in-depth analysis of field semantics is performed. The meaning of each field is labeled using a semantic analysis model to determine the semantic category of the field.

[0027] Based on the semantic category of the field, the semantic association between different fields is compared. If the semantic similarity between fields exceeds a preset threshold, they are classified into the same semantic group to obtain the semantic grouping result.

[0028] Based on the semantic grouping results, a transformation process is implemented to map fields within different semantic groups to a standardized data structure. Field formats are adjusted according to preset transformation rules to construct a standardized data set.

[0029] For standardized datasets, analyze field relationships and use association analysis to identify logical dependencies between fields. If a field has a strong correlation with other fields, incorporate it into a unified mapping framework to determine the field mapping relationship.

[0030] Based on the field mapping relationship, the data structure and semantic analysis results are integrated to generate the final unified mapping scheme and obtain a complete set of field mappings;

[0031] For the complete set of field mappings, data validation and optimization are performed. A consistency check tool is used to detect the accuracy of the mapping results. If mapping deviations are found, they are corrected by adjusting the rules to obtain the final standard data structure.

[0032] Preferably, for a standardized data structure set, the dependencies and business rules between systems are obtained. By constructing a data logical association model, it is determined whether there are any missing or conflicting data across systems. If missing or conflicting data exists, a supplementary rule base is triggered to fill in the gaps, thus determining the complete business data link, including:

[0033] By parsing the data structure, we can obtain the data dependencies and business rules between various systems, build an initial logical relationship framework, and obtain a preliminary data dependency graph.

[0034] Based on the preliminary data dependency map, the data interaction points across systems are scanned to determine whether there is any missing data or data conflict. If missing data or conflict is detected, the corresponding inter-system interaction nodes are recorded to determine the list of problem nodes.

[0035] By using a list of problem nodes, we can analyze the specific types of missing or conflicting data, combine them with business rules, call the pre-established rule library, obtain the corresponding supplementary rules, and get an appropriate filling solution.

[0036] By adapting the filling solution, data is supplemented or conflicts are adjusted for problem nodes, a corrected data dependency graph is generated, and the adjusted business data link is established.

[0037] Based on the revised data dependency graph, the data consistency between systems is verified. If data conflicts or missing data are still found, the rule base is iteratively called to fill the gaps a second time, so as to obtain the final consistent data link.

[0038] Obtain a consistent data link, combine it with the logical model to generate a complete business data flow path, determine whether the data interaction between each system conforms to the preset business rules, and determine the final business data closed loop.

[0039] Preferably, the process involves dynamically reorganizing cross-system business processes through a complete business data link to obtain the temporal relationships and data flow paths of each business link, resulting in a complete process tracing map, including:

[0040] Raw business data is obtained from a cross-system environment, organized into a structured dataset, and a preliminary dataset is obtained.

[0041] Based on the initial dataset, a process tracing algorithm is used to analyze the data links, identify the data flow path, and determine the flow order between each stage.

[0042] Based on the data flow path and temporal relationships, a dynamic reorganization model of the business process is constructed to obtain the reorganized process structure;

[0043] If there are inconsistencies in the timing relationships in the restructured process, a preset threshold will be used to check whether there are any abnormal steps.

[0044] Based on the verification results, data backtracking is performed on the abnormal steps to extract relevant business data from the data flow path and obtain the corrected process segment.

[0045] By integrating comprehensive process information through the revised process fragments, a complete tracking map is constructed to determine the final business process view.

[0046] Preferably, based on the complete process tracing graph, the data flow path of each business link is verified. If a path interruption or anomaly is found, a preset threshold rule is invoked for repair, and the repaired process data set is determined, including:

[0047] By analyzing the flowchart, we can obtain the data flow path information of each business link, organize the complete path mapping relationship, and determine the initial data flow structure.

[0048] Based on the initial data flow structure, path verification is performed for each business link to determine whether there is a path interruption or abnormal situation. If an interruption or abnormality is detected, the corresponding link identifier is recorded to obtain a list of abnormal links.

[0049] By using a list of abnormal steps and calling preset threshold rules, the data flow path of each abnormal step is compared and analyzed. If the data deviation exceeds the threshold range, a repair mechanism is triggered to determine the repair priority sequence.

[0050] By repairing the priority sequence, the corresponding repair mechanism logic is obtained, and data is reconstructed for path interruption or abnormal conditions to obtain the repaired path data group.

[0051] Based on the repaired path data group, verify the continuity of data flow between business processes. If the continuity verification fails, the threshold rules are called again for a second repair to determine the final path consistency result.

[0052] Obtain path consistency results, organize the repaired data of all business links, generate a complete set of process data, and determine the final closed-loop structure of data flow.

[0053] Preferably, the process involves obtaining key financial indicators and related data for each stage through the repaired process data set, and then using a data integration module to perform multi-dimensional summarization to obtain a preliminary financial data report framework, including:

[0054] By extracting the contents of the repair set from the process data and using data cleaning tools to preprocess the data, a standardized basic dataset is obtained.

[0055] Based on the standardized basic dataset, key indicators and financial processes are classified and mapped, and the set of indicators corresponding to each process is determined by using preset classification rules.

[0056] If there are missing values ​​in the indicator set, the missing information is obtained from the relevant data source through the supplementation mechanism of associated data to determine whether the missing part has been completely filled.

[0057] By using the completed set of indicators, the data integration module is used to process the multidimensional summary and obtain the summarized structured data results.

[0058] Based on the structured data results, the data is formatted and transformed to meet the requirements of the preliminary report. A standardized report framework is obtained by using a preset report template.

[0059] If outliers exist in the standardized report framework, a preset threshold detection mechanism will be used to filter them and determine the distribution range of the outlier data.

[0060] By analyzing the distribution range of abnormal data, regression analysis algorithms are used to correct outliers, resulting in the final financial data framework.

[0061] Preferably, based on the preliminary financial data reporting framework, a secondary verification is performed on the completeness and consistency of the summarized data. If data deviation is detected, a backtracking mechanism is triggered to re-match the business data links and determine the final financial data content, including:

[0062] The following technical processes are constructed for attributes such as financial data, aggregated data, completeness verification, consistency verification, data deviation, backtracking process, business chain, data matching, final confirmation, secondary verification, and data content;

[0063] Discard report frame attributes because they have weak correlation with other attributes in the context of logic and cannot form a tight chain of thought.

[0064] Obtain initial financial and summary data, and conduct a preliminary analysis of the data integrity using a pre-established data verification model to determine whether there are any missing or abnormal items.

[0065] If missing or abnormal data is detected, it is marked as data to be processed, and preliminary verification results are obtained.

[0066] For the data to be processed in the preliminary verification results, consistency verification rules are used to compare the field correspondence between the summary data and the source data. If inconsistencies in fields or deviations in values ​​are found, the deviation details are recorded and the range of deviation data is determined.

[0067] Based on the range of deviation data, the backtracking process is initiated to extract the relevant original records from the business chain. Data comparison tools are used to trace the data source layer by layer to obtain the backtracked data set.

[0068] Using the backtracked dataset, a data matching operation is performed, comparing the extracted original records with the deviation data one by one. If a match is found, the deviation data is updated, and the corrected data content is determined.

[0069] For the corrected data content, a second verification process is implemented. The random forest algorithm is used to perform in-depth detection of the data integrity and consistency. If the detection result meets the preset threshold, the data accuracy is confirmed and the final verification result is obtained.

[0070] Based on the final verification results, all data content is integrated to generate a unified financial data view. Through automated processing tools, the corrected data and the original data are classified and stored to determine the final financial data content.

[0071] The final financial data is used to perform data consistency verification. If the verification passes, the data is archived to the system database, the archived data identifier is obtained, and the entire processing flow is completed.

[0072] Preferably, the step of obtaining a preset report template and output rules from the final financial data content, and generating a compliant financial report output result includes:

[0073] By obtaining raw financial data from the source system and using data cleaning tools to remove outliers and duplicates, a pre-processed financial dataset is obtained.

[0074] Based on the pre-processed financial dataset, the preset business logic is applied to classify and summarize the data to determine the distribution characteristics of each financial indicator.

[0075] If the distribution characteristics exceed the preset threshold range, the data verification mechanism is triggered to perform a second check on the abnormal indicators to determine whether there are potential errors.

[0076] After verifying the financial dataset, a preset report template is obtained, and template matching technology is used to match the data with the template fields to obtain preliminary mapping results;

[0077] Based on the preliminary mapping results, the data format and content are adjusted according to preset rules to determine the final report structure;

[0078] By adjusting the report structure, the standardized output module generates financial statement content that conforms to the standards, completing the full conversion from data to reports;

[0079] If format deviations are found in the generated financial statements during verification, the mapping results are traced back to correct them, resulting in the final error-free output.

[0080] This invention also provides a machine learning-based automatic financial statement generation system, comprising:

[0081] The format mapping parsing module is used to perform preliminary parsing of heterogeneous data from multiple systems using a pre-established format mapping rule library, to obtain the field structure and encoding rules of the data from each system, and to determine the preliminary data format difference classification results.

[0082] The semantic matching and conversion module is used to match and convert the semantics of different fields based on the preliminary data format difference classification results, using intelligent recognition technology to generate a standardized set of data structures and obtain a unified field mapping relationship;

[0083] The data link construction module is used to obtain the data dependencies and business rules between systems for a standardized set of data structures. By building a data logic association model, it determines whether there are any missing or conflicting data across systems. If there are missing or conflicting data, it triggers the supplementary rule base to fill in the gaps and determine the complete business data link.

[0084] The process reengineering and tracing module is used to dynamically reengineer cross-system business processes through a complete business data link and a process tracing algorithm, to obtain the temporal relationship and data flow path of each business link, and to obtain a complete process tracing map.

[0085] The path verification and repair module is used to verify the data flow path of each business link based on the complete process tracing map. If a path interruption or abnormality is found, the preset threshold rules are called to repair it and determine the repaired process data set.

[0086] The financial data integration module is used to obtain key financial indicators and related data of each link through the repaired process data set, and to summarize them from multiple dimensions to obtain a preliminary financial data report framework.

[0087] The data consistency verification module is used to perform a secondary verification of the completeness and consistency of the summarized data based on the initial financial data report framework. If a data deviation is detected, a backtracking mechanism is triggered to rematch the business data links and determine the final financial data content.

[0088] The report generation and output module is used to obtain preset report templates and output rules from the final financial data content, and generate financial report output results that meet the standards.

[0089] This invention performs preliminary parsing using a pre-established format mapping rule base, employs intelligent recognition technology to match and transform the semantics of different fields, and generates a standardized set of data structures. It constructs a data logical association model to identify missing or conflicting cross-system data and triggers supplementary rule bases to fill in the gaps. A process tracing algorithm dynamically reorganizes cross-system business processes to obtain a complete process tracing graph. For abnormal data flow paths, preset threshold rules are invoked for repair. This invention uses a data integration module to perform multi-dimensional aggregation, generate a preliminary financial data report framework, and performs secondary verification, triggering a backtracking mechanism to re-match business data links, ultimately generating compliant financial report outputs. This invention achieves intelligent integration of heterogeneous data and automatic generation of financial reports, improving the efficiency and accuracy of cross-system data processing. Attached Figure Description

[0090] Figure 1 This is a flowchart of the machine learning-based automatic financial statement generation method of the present invention. Detailed Implementation

[0091] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be described in detail below with reference to the accompanying drawings and specific embodiments.

[0092] Example 1:

[0093] like Figure 1 As shown, this embodiment of the invention provides a method for automatically generating financial statements based on machine learning, including:

[0094] Step S101: Using a pre-established format mapping rule library, perform preliminary analysis on heterogeneous data from multiple systems to obtain the field structure and encoding rules of each system's data, and determine the preliminary data format difference classification results.

[0095] Using a pre-defined format mapping rule base, initial parsing operations are performed on heterogeneous data from multiple system sources to obtain the field structure and encoding rules of each system's data, determining preliminary format difference classification results. Hierarchical parsing technology is employed to extract key metadata from the field structures obtained from the initial parsing, assessing the completeness and consistency of the field structures. If missing or inconsistent field structures exist, field completion and standardization are performed using the pre-defined rule base to obtain standardized field structure data. Based on the standardized field structure data, a layer-by-layer comparative analysis of encoding rules is conducted. If the encoding rules deviate from the pre-defined rule base, encoding conversion is performed using mapping rules to obtain a unified encoding rule dataset. Using the unified encoding rule dataset, a deep comparison of data formats is performed. If significant differences exist between data formats, adjustments are made using a pre-defined format conversion template to determine a standardized data format set. After obtaining the standardized data format set, the difference classification is refined using a decision tree algorithm for automated classification of format differences, yielding detailed difference classification results. Based on the detailed difference classification results, a corresponding data integration mapping table is generated, and format unification processing is performed on the heterogeneous data from multiple system sources to obtain the final integrated data result.

[0096] For example, when processing heterogeneous data from multiple system sources, the initial parsing operation is a crucial first step. Suppose we are dealing with procurement data from three different enterprise resource planning (ERP) systems, each with different field names and formats. System A uses "PurchaseID" as the purchase number field, System B uses "OrderNo," and System C uses "BuyCode." Using a pre-defined format mapping rule base, we can initially parse the field structures of each system and map these fields to a unified "Purchase Number" field, determining the classification results of format differences. This process helps to quickly identify differences in field naming, laying the foundation for subsequent processing.

[0097] Specifically, layered parsing technology can be used to extract key metadata and determine the completeness of field structures. Suppose system A lacks a "purchase date" field, while systems B and C include this field. Using a pre-defined rule base, the system will automatically complete the missing field in system A and standardize the date format to "YYYY-MM-DD," similarly adjusting system B's "MM / DD / YYYY" format to the standard format. This standardization ensures the consistency of field structures, providing a reliable data foundation for subsequent coding rule comparisons. In coding rule comparison analysis, if system C's purchase number coding rule is found to be a combination of letters and numbers, such as "P12345," while the standard rule uses purely numeric coding, it will be converted to "12345" using a mapping rule, achieving coding rule unification. This conversion ensures the comparability and consistency of data during integration, avoiding data matching errors caused by coding differences.

[0098] For example, for deep comparison of data formats, if system A stores data in CSV format and system B in JSON format, a preset format conversion template can be used to adjust both to a unified JSON format, ensuring the standardization of the data format set. This adjustment significantly improves the compatibility of data processing and reduces obstacles to data interaction between systems.

[0099] Specifically, the detailed processing of discrepancy classification can be automated using decision tree algorithms. Assuming that based on features such as field length, data type, and encoding rules, the system automatically categorizes discrepancies into three types: "missing fields," "inconsistent format," and "encoding deviation." This classification provides clear guidance for subsequent integration, reducing the cost of manual intervention. Finally, a data integration mapping table is generated based on the discrepancy classification results, unifying the data from multiple systems into a standard format.

[0100] For example, the purchase number fields of systems A, B, and C are ultimately mapped to a unified "ProcurementID" and integrated into a single dataset. This process not only improves data consistency but also provides high-quality data support for enterprise data analysis and decision-making, significantly improving business efficiency.

[0101] Step S102: Based on the preliminary data format difference classification results, intelligent recognition technology is used to match and transform the semantics of different fields to generate a standardized data structure set and obtain a unified field mapping relationship.

[0102] Step 1: Obtain initial data format information from the classification results. Perform preliminary analysis on data from different sources, and use a pre-established rule base to identify data formats, obtaining a preliminary set of format classifications. Step 2: Based on the preliminary set of format classifications, use intelligent recognition technology to perform in-depth semantic analysis of the fields. Label the meaning of each field using a semantic analysis model to determine its semantic category. Step 3: For the semantic categories of the fields, use matching technology to compare the semantic relationships between different fields. If the semantic similarity between fields exceeds a preset threshold, they are classified into the same semantic group, obtaining semantic grouping results. Step 4: Based on the semantic grouping results, implement a transformation process, mapping fields within different semantic groups to a standardized data structure. Adjust the field format using preset transformation rules to construct a standardized dataset. Step 5: For the standardized dataset, analyze field relationships. Use association analysis methods to identify logical dependencies between fields. If a field has a strong correlation with other fields, include it in the unified mapping framework to determine the field mapping relationship. Step 6: Based on the field mapping relationships, integrate the data structure and semantic analysis results. Use structure construction technology to generate the final unified mapping scheme, obtaining a complete set of field mappings. Step 7: For the complete set of field mappings, perform data validation and optimization. Use a consistency check tool to check the accuracy of the mapping results. If mapping deviations are found, correct them by adjusting the rules to obtain the final standard data structure.

[0103] For example, when processing heterogeneous data from multiple systems, obtaining initial data format information is a crucial step. Preliminary analysis can be performed using a pre-established rule base for data from different sources. For instance, suppose there is data from three different business systems, involving customer information, order records, and inventory management, all falling under the category of internal enterprise data integration. The rule base might contain basic rules such as field length and data type. Through comparison, it can be found that the data fields in the customer information system are primarily text-based, while the data in the order record system is mostly numerical. Therefore, the initial classification set is divided into two groups: text-based and numerical-based.

[0104] For example, in the deep semantic analysis phase of fields, intelligent recognition technology can help understand the meaning of fields. Suppose a field in a customer information system is named "Cust_ID". Through a semantic analysis model, combined with context and historical data, its meaning is labeled as "customer number," and it is categorized as an identity identifier. This analysis can be further extended to "Order_No" in order records, which is also categorized as an identity identifier, laying the foundation for subsequent associations.

[0105] For example, in semantic category comparison, matching techniques are used to determine the semantic relationships between fields. Suppose the fields "Cust_ID" and "Order_Cust" achieve a similarity of 85% through semantic similarity calculation, exceeding a preset threshold of 80%, then they are grouped into the same semantic group, indicating that both are related to customer identity. This grouping helps reduce redundant fields and improves data integration efficiency. Mapping the semantic grouping results to a standardized data structure is a crucial step in the transformation process. Suppose "Cust_ID" and "Order_Cust" are uniformly mapped to the standard field "CustomerID," and the field format is adjusted according to preset rules, such as standardizing the length to 10 characters, to ensure compatibility in subsequent processing. For field relationship analysis, association analysis methods can identify logical dependencies. Suppose a strong correlation is found between "Order_Amount" in order records and "Stock_Value" in inventory management; changes in order amount affect inventory value. Then, both are included in a unified mapping framework to ensure data logic consistency. When integrating data structures and semantic results, structure building techniques are used to generate a unified mapping scheme. For example, all customer-related fields can be integrated into a complete mapping set, including categories such as identity identifiers and contact information, forming a clear field mapping relationship for easy subsequent retrieval. Finally, during the data validation and optimization phase, consistency checking tools can detect the accuracy of the mapping results. If an error is found in the mapping of "CustomerID" in some records, it can be corrected by adjusting the rules, ensuring the reliability of the final standard data structure. This validation mechanism effectively improves the quality of data integration.

[0106] Step S103: For the standardized data structure set, obtain the data dependencies and business rules between the systems, and determine whether there are any missing or conflicting data across systems by constructing a data logic association model. If there are missing or conflicting data, trigger the supplementary rule base to fill in the gaps and determine the complete business data link.

[0107] By parsing the data structure, the data dependencies and business rules between systems are obtained, and an initial logical framework is constructed to obtain a preliminary data dependency graph. Based on this preliminary graph, cross-system data interaction points are scanned to determine if any data is missing or conflicting. If missing or conflicting data is detected, the corresponding system interaction nodes are recorded, identifying a list of problem nodes. Using this list, the specific types of missing or conflicting data are analyzed. Combined with business rules, a pre-established rule base is invoked to obtain corresponding supplementary rules, resulting in an appropriate filling solution. Using this solution, data is supplemented or conflicts are adjusted at the problem nodes, generating a revised data dependency graph and establishing the adjusted business data links. Based on the revised graph, data consistency between systems is verified. If data conflicts or missing data are still found, the rule base is iteratively invoked for secondary filling to obtain a final consistent data link. This final consistent data link, combined with the logical model, generates a complete business data flow path, determines whether the data interactions between systems conform to preset business rules, and establishes the final business data closed loop.

[0108] For example, when parsing data structures to obtain data dependencies between systems, one can first extract field information from the core database of the business system and analyze the calling relationships between fields. Suppose that in a supply chain management system, the order number field in the order table has a direct dependency on the inventory number field in the inventory table. By tracing the data flow path between the two tables, a preliminary logical framework for the relationship between order generation and inventory deduction can be constructed. This approach helps to clearly present the logic of data flow between systems.

[0109] For example, when scanning cross-system data interaction points, one can focus on the data transmission between the order system and the payment system. Suppose the order system generates order data, but the payment system does not receive the corresponding payment status update; this data loss will be recorded as a problem node. During the scan, log comparison can identify the specific time point and interaction interface where the data is missing, thus creating a list of problem nodes and providing a basis for subsequent processing.

[0110] For example, when analyzing data loss or conflict types, if the payment system is found to be missing order amount data, business rules can be used to determine if this is due to an interface transmission interruption. After calling a pre-established rule base, the system will match a supplementary rule, which retrieves the amount data from the order system and fills it into the payment system. This adaptation approach can quickly locate the root cause of the problem and provide a targeted solution.

[0111] For example, to implement a solution for filling in missing data points, a scheduled task mechanism can be used to automatically trigger the data replenishment process after a missing data point is detected. For instance, the payment system checks data integrity hourly; if missing amount data is found, it calls the order system interface to fill it in. This approach effectively reduces manual intervention and improves data processing efficiency.

[0112] For example, when verifying the consistency of the corrected data dependency graph, the key field values ​​of the order system and the payment system can be compared to determine if there are any discrepancies in amounts. If conflicts are still found, the rule base is iteratively invoked to find more granular adjustment rules, such as prioritizing the amount data from the order system as the benchmark. This iterative approach can gradually approach the data consistency goal.

[0113] For example, when generating the final business data flow path, a logical model can be used to outline the complete chain from order creation and payment confirmation to inventory deduction. If verification reveals a data update delay in the payment confirmation stage, data real-time performance can be ensured by optimizing the frequency of API calls. This closed-loop design guarantees the smooth execution of the business process while providing data support for subsequent optimizations.

[0114] Step S104: Through the complete business data link, the process tracing algorithm is used to dynamically reorganize the cross-system business process, obtain the temporal relationship and data flow path of each business link, and obtain a complete process tracing map.

[0115] The business data acquisition module obtains raw business data from the cross-system environment and organizes it into a structured dataset, resulting in a preliminary dataset. Based on this preliminary dataset, a process tracing algorithm is used to analyze the data flow, identify the data transfer path, and determine the flow order between each stage. For the data transfer path, combined with temporal relationships, a dynamic reconfiguration model of the business process is constructed to obtain the reconfigured process structure. If inconsistencies in temporal relationships exist in the reconfigured process structure, a preset threshold is used for verification to determine if any abnormal stages exist. Based on the verification results, data backtracking is performed on the abnormal stages, extracting relevant business data from the data transfer path to obtain corrected process segments. Using these corrected process segments, comprehensive process stage information is integrated to construct a complete tracing graph, determining the final business process view.

[0116] For example, in the application of the business data acquisition module, raw business data can be obtained across system environments. Taking a financial trading platform as an example, suppose the platform involves multiple subsystems, such as a user account system, a transaction record system, and a payment settlement system. The raw data may include basic user information, transaction details, and payment status, which are scattered across different systems. Through the acquisition module, this data is integrated into a structured dataset, for example, by unifying fields such as user ID, transaction amount, and timestamp into a table format, facilitating subsequent analysis. The key to this step is ensuring the integrity and consistency of the data, laying the foundation for subsequent process tracking.

[0117] For example, when using process tracing algorithms to analyze data flows, the entire data path from a user placing an order to payment completion can be traced. Taking a financial transaction platform as an example, data flow may involve stages such as user order placement, order confirmation, payment request, and payment result feedback. Process tracing algorithms determine the data flow order based on timestamps and event correlations; for example, if the order is generated at 10:00, the payment request at 10:02, and the payment completion at 10:05, a clear timeline is formed. This helps identify potential delays or breakpoints, improving the visualization and management of processes.

[0118] For example, when building a dynamic business process reengineering model, the process structure can be adjusted based on temporal relationships. In the aforementioned financial transaction platform, if the payment request process is found to be taking too long, the dynamic reengineering model might process the payment request and order confirmation processes in parallel, generating a new process structure. This adjustment, based on the analysis of time-series data, aims to optimize business efficiency, reduce user waiting time, and thus improve the overall experience.

[0119] For example, when verifying inconsistencies in the timing relationships, a preset threshold can be set for judgment. Assuming the normal time threshold for payment request to payment completion is 5 minutes, if a transaction takes more than 10 minutes, it is marked as an abnormal step. This verification method can quickly locate problematic nodes, providing a basis for subsequent backtracking and helping to promptly identify and resolve issues.

[0120] For example, data backtracking for abnormal processes can extract relevant business data from the data flow path. In a financial transaction platform, if a transaction is marked as abnormal, it can be traced back to the payment request stage to extract relevant log data, such as the request initiation time and interface response status, to analyze whether it was caused by network latency or system failure, and then generate a corrected process segment. This approach can accurately pinpoint the root cause of the problem and improve problem-solving efficiency.

[0121] For example, when integrating comprehensive process information and constructing a tracking map, all revised segments can be aggregated to form a complete business process view. Taking a financial transaction platform as an example, the final map might display the entire chain from a user placing an order to payment completion, including the time and status of each step. This view provides an intuitive basis for business optimization, helping managers quickly grasp the overall process and formulate improvement strategies.

[0122] Step S105: Based on the complete process tracing map, verify the data flow path of each business link. If a path interruption or abnormality is found, call the preset threshold rules to repair it and determine the repaired process data set.

[0123] By analyzing the flowchart, data flow path information for each business step is obtained, a complete path mapping relationship is established, and the initial data flow structure is determined. Based on this initial structure, path verification is performed for each business step to determine if any path interruptions or anomalies exist. If an interruption or anomaly is detected, the corresponding step identifier is recorded, resulting in a list of abnormal steps. Using this list, preset threshold rules are invoked to compare and analyze the data flow path for each abnormal step. If the data deviation exceeds the threshold range, a repair mechanism is triggered, and a repair priority sequence is determined. The corresponding repair mechanism logic is obtained through the repair priority sequence, and data is reconstructed for path interruptions or anomalies, resulting in a repaired path data set. Based on the repaired path data set, the continuity of data flow between business steps is verified. If the continuity verification fails, the threshold rules are invoked again for a second repair, and the final path consistency result is determined. The path consistency result is obtained, and the repaired data for all business steps is compiled to generate a complete process data set, determining the final closed-loop data flow structure.

[0124] For example, when analyzing a business process diagram, one can first analyze the data flow path information from an overall perspective. Taking the order processing flow of an e-commerce platform as an example, assume that the complete link from a user placing an order to order delivery involves four stages: order generation, inventory verification, payment processing, and logistics delivery. By mapping the data flow path of each stage, the flow structure of order data from generation to delivery can be preliminarily determined. For example, after an order is generated, it needs to flow to the inventory system to verify the product inventory, then enter the payment system to complete the transaction, and finally push it to the logistics system to arrange delivery.

[0125] For example, regarding the topic of path verification.

[0126] In one possible implementation, an interruption can be determined by comparing the input and output data of each step. For example, in the e-commerce order process described above, if the output data of the inventory verification step is not successfully transmitted to the payment processing step, the payment system may be unable to obtain the order status. In this case, the inventory verification step will be recorded as an abnormal step and added to the list of abnormal steps. This verification method helps to quickly locate the problematic step.

[0127] For example, a common method for handling a list of abnormal processes is to compare and analyze data using preset threshold rules. If the data flow delay in the payment processing stage exceeds a preset 5-minute threshold, it is considered an anomaly, triggering a repair mechanism with high priority. This time-threshold-based rule effectively filters out critical issues affecting user experience, ensuring that repair resources are prioritized for the most urgent anomalies.

[0128] For example, in implementing the repair mechanism logic, data reconstruction can be performed to address path interruptions. Taking the e-commerce order process as an example, if the payment processing is interrupted due to data loss, the payment request data packet can be reconstructed by tracing back the data from the order generation stage. The repaired path data group will then restore the continuity of the flow. This approach can effectively reduce business interruptions caused by data loss.

[0129] For example, regarding data flow continuity verification, if the repaired path data group still shows data synchronization issues in the logistics and delivery process, the threshold rules can be invoked again for a second repair. If the second verification finds that the logistics system's response time exceeds the preset 10-minute limit, the data push frequency is adjusted to ensure data flow consistency. This second repair mechanism further guarantees the integrity of the process.

[0130] For example, when generating the final process dataset, all the corrected data can be integrated into a closed-loop structure. Taking the e-commerce order process as an example, the final dataset will clearly show the status of each step from order placement to delivery, ensuring no business link is missed. This closed-loop structure facilitates subsequent business optimization and monitoring.

[0131] For example, the results of path consistency can be organized using visualization tools to present the closed-loop structure of the repaired data flow.

[0132] For example, displaying the time and status of each step in the e-commerce order process in chart form allows business personnel to quickly understand the health of the chain. This approach can improve the efficiency and transparency of business management.

[0133] Step S106: Through the repaired process data set, obtain the key financial indicators and related data of each link, and use the data integration module to summarize them in multiple dimensions to obtain a preliminary financial data report framework.

[0134] By extracting the contents of the repair set from the process data and preprocessing the data using data cleaning tools, a standardized basic dataset is obtained. Based on this standardized dataset, key indicators and financial processes are categorized and mapped using preset classification rules to determine the corresponding indicator sets for each process. If missing values ​​exist in the indicator sets, a data supplementation mechanism is used to obtain filling information from relevant data sources to determine if the missing parts have been completely filled. Using the supplemented indicator sets, a data integration module processes the multidimensional summarization to obtain the summarized structured data results. Based on the structured data results, formatting and transformation are performed according to the requirements of the preliminary reports, using a preset report template to obtain a standardized report framework. If outliers exist in the standardized report framework, a preset threshold detection mechanism is used to filter and determine the distribution range of the outlier data. Based on the distribution range of the outlier data, a regression analysis algorithm is used to correct the outliers, resulting in the final financial data framework.

[0135] For example, when processing repair sets in workflow data, it's helpful to first understand the principles behind data cleaning tools. Data cleaning tools primarily remove redundancy, errors, or inconsistencies from data to ensure the accuracy of subsequent analysis.

[0136] In one possible implementation, assuming that a financial process dataset contains duplicate transaction records, a data cleaning tool would compare transaction numbers and timestamps to remove duplicates, ultimately resulting in a standardized base dataset, such as reducing the original 10,000 records to 9,800 valid records. This processing method effectively improves data quality and lays the foundation for subsequent classification and mapping.

[0137] For example, the classification mapping of key performance indicators (KPIs) and financial processes can be achieved by using pre-defined rules to associate indicators such as revenue and costs with specific business processes. Suppose a company's revenue indicator needs to be linked to the sales process; the classification rules will categorize all sales-related transaction data under the revenue indicator, forming a set of indicators that includes information such as monthly sales volume and customer distribution. This classification method helps to clearly show the financial performance of each process, making it easier for managers to quickly identify problem areas.

[0138] For example, when dealing with missing values ​​in a set of indicators, a supplementation mechanism is particularly important. Suppose a set of indicators is missing some cost data; this can be filled by linking to data sources, such as records from a procurement system, and extracting the corresponding cost information. If 30% of the cost data for a certain month is missing, supplementation can improve the completeness to over 95%. This approach ensures the comprehensiveness of data analysis and avoids decision-making biases caused by missing data.

[0139] For example, the data integration module's multi-dimensional aggregation processing can unify and organize data from different sources. Suppose it needs to aggregate sales, inventory, and cost data; the integration module will categorize this data by dimensions such as time and region, forming structured data results, such as generating a table comparing sales and costs across different regions. This integration helps to gain insights into business conditions from multiple perspectives.

[0140] For example, for the formatting and conversion of preliminary reports, preset report templates can quickly transform data into a standard format. Suppose a financial report needs to be displayed quarterly; the template allows data to be directly populated into the corresponding quarterly columns, forming a standardized report framework. This approach improves the standardization and readability of the reports.

[0141] For example, threshold detection mechanisms can be useful when detecting outliers in a standardized reporting framework. Suppose a revenue statistic is 50% higher than the historical average, exceeding a preset threshold. The system will mark it as an anomaly and determine its distribution range, such as concentration in a specific region. This filtering method helps to quickly identify potential problems.

[0142] For example, regression analysis algorithms can be used to smooth data to correct outliers. Suppose an outlier in revenue is 5 million yuan, while the historical trend is around 3 million yuan. The system will adjust the figure based on historical trends, ultimately correcting it to a value close to the trend. This correction method ensures the rationality of the financial data framework and provides a reliable basis for subsequent decision-making.

[0143] Step S107: Based on the preliminary financial data report framework, a second verification is performed on the completeness and consistency of the summarized data. If a data deviation is detected, a backtracking mechanism is triggered to re-match the business data links and determine the final financial data content.

[0144] The following technical process is constructed to address attributes such as financial data, summary data, integrity verification, consistency verification, data deviation, backtracking process, business chain, data matching, final confirmation, secondary verification, and data content. Report framework attributes are discarded because they have weak logical connections with other attributes and cannot form a tight thought chain. Initial financial and summary data are obtained. A pre-established data verification model is used to perform a preliminary analysis of data integrity, determining if there are any missing or abnormal items. If missing or abnormal items are detected, they are marked as data to be processed, yielding preliminary verification results. For the data to be processed in the preliminary verification results, consistency verification rules are applied to compare the field correspondence between the summary data and the source data. If inconsistencies or numerical deviations are found, deviation details are recorded, and the range of deviation data is determined. Based on the deviation data range, a backtracking process is initiated, extracting relevant original records from the business chain. Data comparison tools are used to trace the data source layer by layer, obtaining a backtracked data set. Using the backtracked data set, a data matching operation is performed, comparing the extracted original records with the deviation data one by one. If a match is successful, the deviation data is updated, and the corrected data content is determined. For the corrected data, a secondary verification process is implemented. A random forest algorithm is used to perform deep testing on the data's integrity and consistency. If the test results meet a preset threshold, the data accuracy is confirmed, and the final verification result is obtained. Based on the final verification result, all data content is integrated to generate a unified financial data view. Using automated processing tools, the corrected data and the original data are categorized and stored to determine the final financial data content. Data consistency verification is then performed on the final financial data content. If the verification passes, the data content is archived to the system database, and the archived data identifier is obtained, completing the entire processing flow.

[0145] For example, in processing financial and aggregated data, obtaining initial data is fundamental to the entire process. Suppose a company needs to compile its sales revenue data for the previous quarter. Financial data might include detailed revenue statements for each department, while aggregated data represents the sum of all departmental revenues. During initial analysis, a pre-established data validation model can quickly identify missing values ​​in a department's revenue statements, such as missing sales figures for a particular product line. The system will mark this as data awaiting processing. This approach allows for timely problem detection, ensuring that subsequent processing is targeted and effective.

[0146] For example, when applying consistency verification rules to check the aggregated data against the source data, it might be discovered that a department's aggregated revenue is 5 million yuan, while the total of the source data is only 4.8 million yuan, resulting in a discrepancy of 200,000 yuan. The system will record the details of this discrepancy and determine the specific product line or time period involved in the discrepancy. This verification method helps to accurately pinpoint the root cause of the problem and provides a clear direction for subsequent corrections.

[0147] For example, in the backtracking process, the system extracts relevant original records from the business chain. Suppose the deviation data involves sales records for a certain product line in a certain month, the system will trace back to original documents such as sales orders and invoices, comparing the data sources layer by layer, ultimately forming a backtracked data set. This method can trace the data flow path from the source, ensuring the reliability of the basis for correction.

[0148] For example, data matching is crucial for correcting discrepancies. Suppose a missed order is found in the backtracking data; by comparing the original record with the discrepancy data, the system updates the order amount in the details, thus correcting the data. This step-by-step comparison method effectively reduces errors.

[0149] For example, in the secondary verification process, a random forest algorithm is used to perform deep testing on the corrected data. Assuming the corrected departmental revenue is 5 million yuan, the system will combine historical data trends and business rules to determine whether it meets the expected threshold. This deep testing further ensures data accuracy.

[0150] For example, after the final verification results are obtained, all data is integrated to generate a financial data view. Assuming the view clearly displays the revenue share and trends of each department, the automated processing tool will categorize and store the corrected data along with the original data for easy retrieval later. This integration method improves data readability.

[0151] For example, during the data consistency verification and archiving process, the system will reconfirm that the data content is correct before archiving it to the database and generating a unique data identifier, such as "2023Q3-FIN-001". This approach facilitates data tracking and management, ensuring the security and traceability of long-term storage.

[0152] Step S108: Based on the final financial data, obtain the preset report template and output rules, and generate a financial report output that meets the specifications.

[0153] Raw financial data is obtained from the source system, and data cleaning tools are used to remove outliers and duplicates, resulting in a pre-processed financial dataset. Based on this dataset, pre-defined business logic is applied for classification and summarization to determine the distribution characteristics of each financial indicator. If the distribution characteristics exceed a pre-defined threshold, a data verification mechanism is triggered to perform a secondary check on the outlier indicators to determine if any potential errors exist. Using the verified financial dataset, a pre-defined report template is obtained, and template matching technology is used to map the data to the template fields, resulting in a preliminary mapping result. Based on the preliminary mapping result, pre-defined rules are applied to adjust the data format and content, determining the final report structure. Using the adjusted report structure, a standardized output module generates standard-compliant financial report content, completing the full data-to-report conversion. If format deviations are found in the generated financial report content during verification, the mapping result is revisited for correction, resulting in the final error-free output content.

[0154] For example, when processing raw financial data, core data such as revenue, costs, and profits can be extracted from the source system. Suppose a company's monthly revenue records contain duplicate transactions and outliers, such as a negative revenue amount, which clearly does not conform to business logic. Data cleaning tools will identify and remove these outliers and duplicates, ensuring that the initially processed financial dataset accurately reflects the actual business situation. Such a cleaning process contributes to the reliability of subsequent analysis.

[0155] For example, when classifying and summarizing the initially processed financial dataset, revenue can be divided into two categories based on business type: product sales and service revenue. Assuming product sales revenue accounts for 60% and service revenue for 40%, the system will calculate the sum of each category and analyze its distribution characteristics through pre-defined business logic. If monthly fluctuations in product sales revenue exceed a preset threshold, such as ±20%, a data verification mechanism is triggered to perform a secondary check on the abnormal indicators. This approach can promptly identify potential data entry errors or business anomalies.

[0156] For example, during the secondary verification process, the identification of abnormal indicators can be combined with the analysis of historical data trends. Suppose that product sales revenue suddenly increases by 50% in a certain month, while the historical average growth rate is only 5%, the system will mark this data as a potential error and confirm whether there was a data entry mistake by verifying the original vouchers and transaction records. This method can effectively pinpoint the root cause of the problem.

[0157] For example, when retrieving a preset report template and mapping data, the revenue field in the financial dataset can be mapped to the revenue column of the template. Suppose the template requires revenue data to be in units of ten thousand yuan, while the original data is in units of yuan; the system will automatically adjust the data format to ensure the mapping result conforms to the template requirements. This matching technology can improve the efficiency of report generation.

[0158] For example, when adjusting data format and content to determine the final report structure, the system may find inconsistencies between subtotals and totals in some fields. In such cases, subtotals can be recalculated using preset rules to ensure data consistency and ultimately create a standardized report structure. This adjustment ensures the accuracy of the report.

[0159] For example, when generating standard-compliant financial statements, the standardized output module arranges the data according to financial principles, producing clear income statements and balance sheets. If a cost data point is formatted incorrectly during output, the system will backtrack to the mapping results to correct it, ensuring the final output is error-free. This approach enhances the professionalism of the reports.

[0160] For example, when validating financial statements, if formatting discrepancies are found, such as incorrect decimal places, the system will automatically backtrack and correct them. This backtracking mechanism ensures the compliance of the final statements and provides a reliable basis for corporate decision-making.

[0161] Example 2:

[0162] This invention provides a machine learning-based automatic financial statement generation system, comprising:

[0163] The format mapping parsing module is used to perform preliminary parsing of heterogeneous data from multiple systems using a pre-established format mapping rule library, to obtain the field structure and encoding rules of the data from each system, and to determine the preliminary data format difference classification results.

[0164] The semantic matching and conversion module is used to match and convert the semantics of different fields based on the preliminary data format difference classification results, using intelligent recognition technology to generate a standardized set of data structures and obtain a unified field mapping relationship;

[0165] The data link construction module is used to obtain the data dependencies and business rules between systems for a standardized set of data structures. By building a data logic association model, it determines whether there are any missing or conflicting data across systems. If there are missing or conflicting data, it triggers the supplementary rule base to fill in the gaps and determine the complete business data link.

[0166] The process reengineering and tracing module is used to dynamically reengineer cross-system business processes through a complete business data link and a process tracing algorithm, to obtain the temporal relationship and data flow path of each business link, and to obtain a complete process tracing map.

[0167] The path verification and repair module is used to verify the data flow path of each business link based on the complete process tracing map. If a path interruption or abnormality is found, the preset threshold rules are called to repair it and determine the repaired process data set.

[0168] The financial data integration module is used to obtain key financial indicators and related data of each link through the repaired process data set, and to summarize them from multiple dimensions to obtain a preliminary financial data report framework.

[0169] The data consistency verification module is used to perform a secondary verification of the completeness and consistency of the summarized data based on the initial financial data report framework. If a data deviation is detected, a backtracking mechanism is triggered to rematch the business data links and determine the final financial data content.

[0170] The report generation and output module is used to obtain preset report templates and output rules from the final financial data content, and generate financial report output results that meet the standards.

[0171] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A method for automatically generating financial statements based on machine learning, characterized in that, include: By using a pre-established format mapping rule library, heterogeneous data from multiple systems is initially analyzed to obtain the field structure and encoding rules of each system's data, and to determine the preliminary data format difference classification results. Based on the preliminary data format difference classification results, the semantics of different fields are matched and transformed to generate a standardized set of data structures and obtain a unified field mapping relationship; For a standardized set of data structures, the dependencies and business rules between the data of each system are obtained. By constructing a data logic association model, it is determined whether there are any missing or conflicting data across systems. If there are missing or conflicting data, a supplementary rule base is triggered to fill the gaps and determine the complete business data link. By using complete business data links, cross-system business processes are dynamically reorganized to obtain the temporal relationships and data flow paths of each business link, resulting in a complete process tracing map. Based on the complete process tracing map, the data flow path of each business link is verified. If the path is interrupted or abnormal, the preset threshold rules are called to repair it and determine the repaired process data set. By using the repaired process data set, key financial indicators and related data of each stage are obtained, and the data integration module is used to summarize them from multiple dimensions to obtain a preliminary financial data report framework. Based on the preliminary financial data reporting framework, a second verification is performed on the completeness and consistency of the summarized data. If a data deviation is detected, a backtracking mechanism is triggered to re-match the business data links and determine the final financial data content. Based on the final financial data, obtain the preset report templates and output rules to generate compliant financial report output results; The standardized data structure set is used to obtain the data dependencies and business rules between systems. By constructing a data logical association model, it is determined whether there are any missing or conflicting data between systems. If missing or conflicting data exists, a supplementary rule base is triggered to fill in the gaps, thus determining the complete business data chain, including: By parsing the data structure, we can obtain the data dependencies and business rules between various systems, build an initial logical relationship framework, and obtain a preliminary data dependency graph. Based on the preliminary data dependency map, the data interaction points across systems are scanned to determine whether there is any missing data or data conflict. If missing data or conflict is detected, the corresponding inter-system interaction nodes are recorded to determine the list of problem nodes. By using a list of problem nodes, we can analyze the specific types of missing or conflicting data, combine them with business rules, call the pre-established rule library, obtain the corresponding supplementary rules, and get an appropriate filling solution. By adapting the filling solution, data is supplemented or conflicts are adjusted for problem nodes, a corrected data dependency graph is generated, and the adjusted business data link is established. Based on the revised data dependency graph, the data consistency between systems is verified. If data conflicts or missing data are still found, the rule base is iteratively called to fill the gaps a second time, so as to obtain the final consistent data link. Obtain a consistent data link, combine it with the logical model to generate a complete business data flow path, determine whether the data interaction between each system conforms to the preset business rules, and determine the final business data closed loop.

2. The method for automatically generating financial statements based on machine learning according to claim 1, characterized in that, The process involves using a pre-established format mapping rule base to perform preliminary analysis on heterogeneous data from multiple systems, obtaining the field structure and encoding rules of the data from each system, and determining preliminary data format difference classification results, including: By using a pre-defined format mapping rule library, an initial parsing operation is performed on heterogeneous data from multiple system sources to obtain the field structure and encoding rules of each system's data and determine the preliminary format difference classification results. Extract key metadata from the field structure obtained from the initial parsing, determine the completeness and consistency of the field structure, and if there are missing or inconsistent fields, perform field completion and standardization processing through a preset rule base to obtain standardized field structure data. Based on the standardized field structure data, a layer-by-layer comparison and analysis is performed on the encoding rules. If there is a deviation between the encoding rules and the preset rule library, the encoding is converted through the mapping rules to obtain a unified encoding rule dataset. By using a unified encoding rule dataset, the data formats are compared in depth. If there are differences between the data formats, a preset format conversion template is used for adjustment to determine a standardized set of data formats. After obtaining a standardized set of data formats, the differences are further refined, and the format differences are automatically classified to obtain detailed difference classification results. Based on the detailed difference classification results, a corresponding data integration mapping table is generated, and the format of heterogeneous data from multiple system sources is standardized to obtain the final integrated data result.

3. The method for automatically generating financial statements based on machine learning according to claim 2, characterized in that, Based on the preliminary data format difference classification results, the semantics of different fields are matched and transformed to generate a standardized data structure set, resulting in a unified field mapping relationship, including: Initial data format information is obtained from the classification results. Preliminary analysis is performed on data from different sources. A pre-established rule base is used to identify the data format, resulting in a preliminary set of format classifications. Based on the initial set of format classifications, in-depth analysis of field semantics is performed. The meaning of each field is labeled using a semantic analysis model to determine the semantic category of the field. Based on the semantic category of the field, the semantic association between different fields is compared. If the semantic similarity between fields exceeds a preset threshold, they are classified into the same semantic group to obtain the semantic grouping result. Based on the semantic grouping results, a transformation process is implemented to map fields within different semantic groups to a standardized data structure. Field formats are adjusted according to preset transformation rules to construct a standardized data set. For standardized datasets, analyze field relationships and use association analysis to identify logical dependencies between fields. If a field has a strong correlation with other fields, incorporate it into a unified mapping framework to determine the field mapping relationship. Based on the field mapping relationship, the data structure and semantic analysis results are integrated to generate the final unified mapping scheme and obtain a complete set of field mappings; For the complete set of field mappings, data validation and optimization are performed. A consistency check tool is used to detect the accuracy of the mapping results. If mapping deviations are found, they are corrected by adjusting the rules to obtain the final standard data structure.

4. The method for automatically generating financial statements based on machine learning according to claim 1, characterized in that, The process involves dynamically reorganizing cross-system business processes through a complete business data link, obtaining the temporal relationships and data flow paths of each business link, and generating a complete process tracing graph, including: Raw business data is obtained from a cross-system environment, organized into a structured dataset, and a preliminary dataset is obtained. Based on the initial dataset, a process tracing algorithm is used to analyze the data links, identify the data flow path, and determine the flow order between each stage. Based on the data flow path and temporal relationships, a dynamic reorganization model of the business process is constructed to obtain the reorganized process structure; If there are inconsistencies in the timing relationships in the restructured process, a preset threshold will be used to check whether there are any abnormal steps. Based on the verification results, data backtracking is performed on the abnormal steps to extract relevant business data from the data flow path and obtain the corrected process segment. By revising the process fragments and integrating comprehensive process information, a complete tracking map is constructed to determine the final business process view.

5. The method for automatically generating financial statements based on machine learning according to claim 1, characterized in that, The process involves verifying the data flow path of each business step based on a complete process tracing graph. If a path interruption or anomaly is detected, a preset threshold rule is invoked for repair, and the repaired process data set is determined, including: By analyzing the flowchart, we can obtain the data flow path information of each business link, organize the complete path mapping relationship, and determine the initial data flow structure. Based on the initial data flow structure, path verification is performed for each business link to determine whether there is a path interruption or abnormal situation. If an interruption or abnormality is detected, the corresponding link identifier is recorded to obtain a list of abnormal links. By using a list of abnormal steps and calling preset threshold rules, the data flow path of each abnormal step is compared and analyzed. If the data deviation exceeds the threshold range, a repair mechanism is triggered to determine the repair priority sequence. By repairing the priority sequence, the corresponding repair mechanism logic is obtained, and data is reconstructed for path interruption or abnormal conditions to obtain the repaired path data group. Based on the repaired path data group, verify the continuity of data flow between business processes. If the continuity verification fails, the threshold rules are called again for a second repair to determine the final path consistency result. Obtain path consistency results, organize the repaired data of all business links, generate a complete set of process data, and determine the final closed-loop structure of data flow.

6. The method for automatically generating financial statements based on machine learning according to claim 1, characterized in that, The process involves obtaining key financial indicators and related data for each stage through the repaired process data set, and then using a data integration module to perform multi-dimensional aggregation to obtain a preliminary financial data reporting framework, including: By extracting the contents of the repair set from the process data and using data cleaning tools to preprocess the data, a standardized basic dataset is obtained. Based on the standardized basic dataset, key indicators and financial processes are classified and mapped, and the set of indicators corresponding to each process is determined by using preset classification rules. If there are missing values ​​in the indicator set, the missing information is obtained from the relevant data source through the supplementation mechanism of associated data to determine whether the missing part has been completely filled. By using the completed set of indicators, the data integration module is used to process the multi-dimensional summary and obtain the summarized structured data results. Based on the structured data results, the data is formatted and transformed to meet the requirements of the initial report, and a standardized report framework is obtained by using a preset report template. If outliers exist in the standardized report framework, a preset threshold detection mechanism will be used to filter them and determine the distribution range of the outlier data. By analyzing the distribution range of abnormal data, regression analysis algorithms are used to correct outliers, resulting in the final financial data framework.

7. The method for automatically generating financial statements based on machine learning according to claim 1, characterized in that, Based on the preliminary financial data reporting framework, a secondary verification is performed on the completeness and consistency of the summarized data. If data deviation is detected, a backtracking mechanism is triggered to re-match the business data links and determine the final financial data content, including: The following technical processes are constructed for financial data, aggregated data, completeness verification, consistency verification, data deviation, backtracking process, business chain, data matching, final confirmation, secondary verification, and data content attributes; Discard report frame attributes because they have weak connections with other attributes in the context of logic and cannot form a tight chain of thought. Obtain initial financial and summary data, and conduct a preliminary analysis of the data integrity using a pre-established data verification model to determine whether there are any missing or abnormal items. If missing or abnormal data is detected, it is marked as data to be processed, and preliminary verification results are obtained. For the data to be processed in the preliminary verification results, consistency verification rules are used to compare the field correspondence between the summary data and the source data. If inconsistencies in fields or deviations in values ​​are found, the deviation details are recorded and the range of deviation data is determined. Based on the range of deviation data, the backtracking process is initiated to extract the relevant original records from the business chain. Data comparison tools are used to trace the data source layer by layer to obtain the backtracked data set. Using the backtracked dataset, a data matching operation is performed, comparing the extracted original records with the deviation data one by one. If a match is found, the deviation data is updated, and the corrected data content is determined. For the corrected data content, a second verification process is implemented. The random forest algorithm is used to perform in-depth detection of the data integrity and consistency. If the detection result meets the preset threshold, the data accuracy is confirmed and the final verification result is obtained. Based on the final verification results, all data content is integrated to generate a unified financial data view. Through automated processing tools, the corrected data and the original data are classified and stored to determine the final financial data content. The final financial data is used to perform data consistency verification. If the verification passes, the data is archived to the system database, the archived data identifier is obtained, and the entire processing flow is completed.

8. The method for automatically generating financial statements based on machine learning according to claim 1, characterized in that, The process of obtaining preset report templates and output rules from the final financial data content, and generating compliant financial report output results includes: By obtaining raw financial data from the source system and using data cleaning tools to remove outliers and duplicates, a pre-processed financial dataset is obtained. Based on the pre-processed financial dataset, the preset business logic is applied to classify and summarize the data to determine the distribution characteristics of each financial indicator. If the distribution characteristics exceed the preset threshold range, the data verification mechanism is triggered to perform a second check on the abnormal indicators to determine whether there are potential errors. After verifying the financial dataset, a preset report template is obtained, and template matching technology is used to match the data with the template fields to obtain preliminary mapping results; Based on the preliminary mapping results, the data format and content are adjusted according to preset rules to determine the final report structure; By adjusting the report structure, the standardized output module generates financial statement content that conforms to the standards, completing the full conversion from data to reports; If format deviations are found in the generated financial statements during verification, the mapping results are traced back to correct them, resulting in the final error-free output.

9. A machine learning-based automatic financial statement generation system that implements the machine learning-based automatic financial statement generation method of claim 1, characterized in that, include: The format mapping parsing module is used to perform preliminary parsing of heterogeneous data from multiple systems using a pre-established format mapping rule library, to obtain the field structure and encoding rules of the data from each system, and to determine the preliminary data format difference classification results. The semantic matching and conversion module is used to match and convert the semantics of different fields based on the preliminary data format difference classification results, using intelligent recognition technology to generate a standardized set of data structures and obtain a unified field mapping relationship; The data link construction module is used to obtain the data dependencies and business rules between systems for a standardized set of data structures. By building a data logic association model, it determines whether there are any missing or conflicting data across systems. If there are missing or conflicting data, it triggers the supplementary rule base to fill in the gaps and determine the complete business data link. The process reengineering and tracing module is used to dynamically reengineer cross-system business processes through a complete business data link and a process tracing algorithm, to obtain the temporal relationship and data flow path of each business link, and to obtain a complete process tracing map. The path verification and repair module is used to verify the data flow path of each business link based on the complete process tracing map. If a path interruption or abnormality is found, the preset threshold rules are called to repair it and determine the repaired process data set. The financial data integration module is used to obtain key financial indicators and related data of each link through the repaired process data set, and to summarize them from multiple dimensions to obtain a preliminary financial data report framework. The data consistency verification module is used to perform a secondary verification of the completeness and consistency of the summarized data based on the initial financial data report framework. If a data deviation is detected, a backtracking mechanism is triggered to rematch the business data links and determine the final financial data content. The report generation and output module is used to obtain preset report templates and output rules from the final financial data content, and generate financial report output results that meet the standards.

Citation Information

Patent Citations

  • Intelligent financial statement generation method and device, equipment and medium

    CN118886405A

  • Audit report automatic generation method based on natural language processing

    CN120124612A