Finance and tax agency accounting intelligent system and method based on machine learning

The intelligent accounting and tax agency system based on machine learning solves the problems of data dispersion and low efficiency of manual processing in traditional accounting and tax agency systems. It realizes automated data processing and efficient and accurate generation of accounting and tax vouchers, ensuring compliance.

CN121504641APending Publication Date: 2026-02-10HANGZHOU HUICAI NETWORK TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511669531.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-14
Publication Date
2026-02-10

AI Technical Summary

Technical Problem

Traditional accounting and tax outsourcing suffers from problems such as data fragmentation, low efficiency of manual processing, difficulty in detecting anomalies, poor data accuracy, and difficulty in ensuring compliance.

Method used

The system employs a machine learning-based intelligent accounting and bookkeeping system. It generates a structured set of financial and tax features through multimodal data fusion processing, uses machine learning models for anomaly detection, combines dynamic verification processes and a historical compliance case library for data correction, generates accounting vouchers, and performs logical conflict detection.

Benefits of technology

It has achieved automated and efficient data processing, improved data accuracy and compliance, reduced manual intervention, and enhanced the timeliness and accuracy of financial and tax processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121504641A_ABST
    Figure CN121504641A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of finance and taxation intelligent accounting, and discloses a finance and taxation agency accounting intelligent system and method based on machine learning. The method comprises the following steps: acquiring enterprise invoice information, bank statement, contract text, tax declaration and other finance and tax original data; performing multi-modal data fusion processing on the original data to generate a structured finance and tax feature set; classifying the scenes into an income confirmation scene, a cost collection scene and a tax accounting scene according to a preset rule; performing anomaly detection on each scene feature set by using a machine learning model, and outputting an anomaly feature mark; triggering a dynamic verification process based on the mark and generating a result, and correcting abnormal data in combination with a historical compliance case library; inputting the corrected feature set into a finance and tax rule engine to generate an accounting voucher draft; and performing logic conflict detection on the draft, outputting a report, adjusting the subject mapping relation according to the report, and finally generating an accounting voucher. According to the method, finance and tax agency accounting automation and precision are realized, and the process normalization and efficiency are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of intelligent accounting technology for finance and taxation, specifically to an intelligent system and method for accounting outsourcing based on machine learning. Background Technology

[0002] In the current field of accounting and tax outsourcing, the types of financial and tax data generated by enterprises in their daily operations are diverse and varied in format, covering paper and electronic invoices, bank statements issued by different banks, various business contract texts, and various forms of tax returns. Under the traditional processing model, this data mainly relies on manual collection, sorting, and entry, which not only consumes a lot of manpower but is also prone to data omissions or errors due to human error. For example, when manually entering invoice information, there may be misalignment of amounts, discrepancies between bank statements and business records, and other problems that affect the accuracy of subsequent accounting data.

[0003] In traditional financial and tax processing, data integration lacks an effective multimodal fusion mechanism. It's difficult to establish efficient correlations between structured invoice data (such as amounts and tax rates), unstructured information from contract texts (such as payment terms and service periods), and time-series bank transaction data. This hinders the full extraction of data value and makes it difficult to form a unified and complete financial and tax data view. This data fragmentation forces finance personnel to repeatedly switch between multiple data sources when conducting revenue recognition and cost aggregation, significantly increasing workload complexity and reducing overall processing efficiency.

[0004] In terms of business scenario segmentation, traditional methods rely heavily on the experience and judgment of finance personnel, lacking standardized segmentation rules. The boundary between revenue recognition and cost allocation is sometimes blurred, and data screening in tax accounting scenarios also lacks unified standards. This can easily lead to the same transaction being incorrectly classified into different scenarios, resulting in deviations in subsequent accounting treatment and affecting the authenticity and compliance of financial statements. Furthermore, the anomaly detection process also relies on manual screening. Finance personnel can usually only judge based on common anomaly characteristics (such as excessively large or small amounts). It is difficult to effectively identify anomalies hidden in data relationships (such as invoice amounts not matching contract agreements, or bank statements not matching tax declaration data), allowing abnormal data to flow into subsequent accounting stages and increasing financial and tax risks.

[0005] When potential anomalies are detected, traditional verification processes often involve fixed steps and lack dynamic adjustment capabilities. They cannot flexibly adapt verification methods based on factors such as anomaly type and data source, resulting in some anomalies failing to pinpoint their root causes and limiting verification effectiveness. Furthermore, during the anomaly data correction phase, the lack of effective correlation with historical compliance cases means the correction process often relies on personal experience, making it difficult to guarantee that the correction results comply with industry standards and tax policies. This can lead to situations where the corrected data still harbors compliance risks.

[0006] In the traditional process of generating accounting vouchers, vouchers need to be manually prepared based on the compiled data. This is not only time-consuming but also prone to logical conflicts due to incorrect account mapping (such as mistakenly recording expense accounts as asset accounts). Furthermore, there is a lack of a systematic logical conflict detection mechanism after the vouchers are generated, and problems are often only discovered in the subsequent review stage. Adjustments made at this point increase rework costs, further slowing down the accounting process and failing to meet enterprises' needs for timeliness and accuracy in financial and tax processing. Summary of the Invention

[0007] The purpose of this invention is to provide an intelligent system and method for financial and tax agency bookkeeping based on machine learning, so as to solve the problems mentioned in the background art.

[0008] To achieve the above objectives, this invention provides an intelligent method for financial and tax agency bookkeeping based on machine learning, the method comprising: Collect raw financial and tax data from enterprises, including invoice information, bank statements, contract texts, and tax returns. The original data is subjected to multimodal data fusion processing to generate a structured set of fiscal and tax features; According to the preset financial and tax business scenario classification rules, the structured financial and tax feature set is classified into revenue recognition scenario, cost collection scenario and tax accounting scenario. A machine learning model is used to perform anomaly detection on the structured financial and tax feature set for each scenario, and anomaly feature labels are output. Based on the aforementioned abnormal feature markers, a dynamic verification process is triggered to generate verification results; Based on the matching degree between the verification results and the historical compliance case library, the abnormal data in the structured financial and tax feature set is corrected; The revised set of structured financial and tax features is input into the financial and tax rules engine to generate a draft accounting voucher; Perform logical conflict detection on the draft accounting voucher and output a conflict detection report; Adjust the account mapping relationship of the draft accounting voucher based on the conflict detection report, and generate the final accounting voucher.

[0009] Preferably, the multimodal data fusion processing of the original data includes: Extract the price-tax separation feature from the invoice information to generate a correlation matrix between the tax-exclusive amount and the tax amount; Analyze the time-series characteristics of transaction counterparties and amounts in the bank statements to construct a fund flow topology graph; Entity recognition is performed on the contract text to extract payment terms and payment period constraints. The structured financial and tax feature set is generated by aligning the correlation matrix, the fund flow topology, and the payment terms and payment period constraints along the time axis.

[0010] Preferably, the anomaly detection using a machine learning model on the structured financial and tax feature set for each scenario includes: For the structured financial and tax feature set in the aforementioned revenue recognition scenario, the isolated forest algorithm is used to detect anomalies in revenue recognition across periods. For the structured financial and tax feature set in the cost aggregation scenario, a clustering algorithm is used to identify deviations in the cost allocation ratio; For the structured set of financial and tax features in the aforementioned tax accounting scenario, a time-series prediction model is used to compare the difference between the actual tax burden and the theoretical tax burden.

[0011] Preferably, the dynamic verification process triggered based on the abnormal feature marker includes: When an anomaly in the revenue recognition across periods is detected, the associated contract text and the bank statement are retrieved to verify the consistency of receipt and payment. When a deviation in the cost allocation ratio is detected, the matching relationship between the material code and the project number in the invoice information is traced. When the difference between the actual tax burden and the theoretical tax burden exceeds a threshold, the tax reduction and exemption clause reference records in the tax return form are checked.

[0012] Preferably, the correction of abnormal data in the structured financial and tax feature set includes: Replace the erroneous fields in the abnormal data with the account processing records of similar scenarios in the historical compliance case library; The replaced fields are validated for compliance with tax rules. If the validation fails, the value range of the fields is iteratively adjusted until compliance is achieved.

[0013] Preferably, the generation logic of the tax and finance rules engine includes: Encode accounting standard entries and tax regulation provisions as rule nodes; Activate the corresponding rule node based on the business type of the structured financial and tax feature set; Debit and credit account matching is performed according to the priority order of the rule nodes.

[0014] Preferably, the logical conflict detection of the draft accounting voucher includes: Iterate through the account balance direction in the draft accounting voucher and the constraints of the rule nodes; When a debit and credit entry for the same account simultaneously meets the reversal condition, it is marked as a redundant entry conflict. When a tax item is detected that does not match the aforementioned tax regulation clause, it is marked as a clause reference conflict.

[0015] Preferably, adjusting the account mapping relationship of the draft accounting voucher based on the conflict detection report includes: For the aforementioned redundant journal entries, the debit and credit amounts of related accounts are merged and the net value is recalculated; In case of any conflict of reference to the aforementioned clauses, supplement the draft accounting voucher with the special handling rules in the aforementioned tax regulations.

[0016] Preferably, the method further includes: Real-time monitoring of updated tax and financial regulations documents, and extraction of keywords related to clause changes; When the keyword for the clause change is related to a rule node in the tax and finance rule engine, the version iteration of the rule node is triggered.

[0017] Preferably, the present invention also includes a machine learning-based intelligent accounting system for financial and tax agency services. The system includes a memory, a processor, and a computer program stored in the memory and running on the processor. When the processor executes the computer program, it implements the steps of the machine learning-based intelligent accounting method for financial and tax agency services described above.

[0018] Compared with the prior art, the beneficial effects of the present invention are: In the data collection and processing stage, we break through the limitations of traditional manual collection and organization. By automatically collecting various types of raw data such as enterprise invoices, bank statements, contract texts, and tax returns, we reduce operational errors caused by manual intervention. At the same time, by using multimodal data fusion processing, we transform different formats and types of financial and tax data into a unified set of structured financial and tax features, breaking down the barriers between data and allowing scattered data to form an organic whole. This facilitates efficient use of data in subsequent stages and eliminates the need for financial personnel to repeatedly switch between multiple data sources, simplifying the data processing workflow.

[0019] In terms of business scenario classification, this method, based on the preset financial and tax business scenario classification rules, clearly classifies the structured financial and tax feature set into revenue recognition scenario, cost collection scenario, and tax accounting scenario, replacing the traditional fuzzy classification method that relies on manual experience. This makes the data boundaries of each business scenario clear, and the subsequent processing work carried out for different scenarios is more targeted, avoiding deviations in accounting processing direction caused by scenario confusion, and ensuring that the processing of various financial and tax businesses conforms to the corresponding business logic and standards.

[0020] In the anomaly detection stage, machine learning models are used to detect the structured financial and tax feature sets of various scenarios. Compared with the limitations of traditional manual investigation, which can only identify common anomalies, machine learning models can mine hidden correlation anomalies from massive amounts of data, such as subtle deviations between invoice amounts and contractual agreements, and implicit mismatches between bank statements and tax declaration data. This allows for a more comprehensive discovery of anomalies in the data, reducing the possibility of abnormal data flowing into subsequent stages and lowering financial and tax risks.

[0021] Once an anomaly is detected, the method triggers a dynamic verification process based on the anomaly markers, rather than using traditional, fixed verification steps. This dynamic verification process can flexibly adjust the verification methods and steps according to the specific type of anomaly, data source, and other factors, more accurately locating the root cause of the anomaly, improving the effectiveness of the verification work, and avoiding the shortcomings of traditional verification processes that cannot thoroughly investigate some anomalies.

[0022] During the abnormal data correction phase, this method combines the matching degree between the verification results and the historical compliance case library for correction. It no longer relies on personal experience, but uses a large number of historical compliance cases as a reference to ensure that the corrected structured financial and tax feature set meets industry norms and tax policy requirements, reducing the possibility of compliance risks in the correction results and making the data correction process more objective and standardized.

[0023] In the accounting voucher generation stage, the revised structured financial and tax feature set is input into the financial and tax rule engine to automatically generate draft accounting vouchers. This replaces the tedious process of manually preparing vouchers, significantly shortening voucher generation time and reducing potential account mapping errors that may occur during manual preparation. Subsequently, the draft voucher undergoes logical conflict detection and a conflict detection report is generated. This allows for the timely identification of logical conflicts such as account mapping issues in the draft, rather than discovering problems only during the subsequent review stage. This avoids increased rework costs and further improves the efficiency of the accounting process.

[0024] Based on the conflict detection report, the subject mapping relationship is adjusted to generate the final accounting voucher, ensuring that the final voucher is logically rigorous and the data is accurate. The entire process, from data collection to voucher generation, forms a complete intelligent processing chain, reducing manual intervention while improving the overall standardization and processing efficiency of financial and tax agency bookkeeping, and better meeting the needs of enterprises for timeliness and accuracy of financial and tax processing. Attached Figure Description

[0025] Figure 1 This is a schematic diagram illustrating the working principle of the intelligent financial and tax agency bookkeeping method based on machine learning described in this invention. Figure 2 A flowchart for multimodal data fusion processing; Figure 3 This is a flowchart for anomaly detection in a machine learning model. Figure 4 A bar chart for monitoring cost deviation rates in each department. Detailed Implementation

[0026] The following description, in conjunction with the accompanying drawings of the embodiments of the present invention, will clearly and completely describe the technical solutions of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0027] Please see Figure 1 This invention provides an intelligent method for financial and tax agency accounting based on machine learning. The method includes: Original financial and tax data for enterprises originating from invoice information, bank statements, contract texts, and tax returns are collected in real-time through standardized interfaces and stored in a distributed database. A multimodal data fusion processing module cleans and aligns the original data, eliminating format differences and missing values, and generating a structured financial and tax feature set with a unified timestamp. Financial and tax business scenario segmentation rules, based on predefined thresholds in accounting standards and tax regulations, map the structured financial and tax feature set to revenue recognition scenarios, cost aggregation scenarios, and tax accounting scenarios. A machine learning model loads pre-trained parameters for different scenarios, performs pattern recognition on the structured financial and tax feature set, and outputs abnormal feature markers. A dynamic verification process calls related data sources for cross-validation based on the anomaly type, generating a verification result log. A historical compliance case library stores approved accounting cases, retrieves the best correction scheme through a similarity matching algorithm, and replaces or reconstructs abnormal fields in the structured financial and tax feature set. A financial and tax rule engine parses the business attributes of the structured financial and tax feature set, activates the corresponding accounting standards and tax regulations nodes, and generates a draft accounting voucher with debit and credit balance. The logical conflict detection module iterates through the account relationships in the draft accounting voucher, identifies redundant entries or missing clauses, and outputs a conflict detection report. The account mapping relationship adjustment module recalculates the account net value or supplements tax rules according to the conflict type, generates the final accounting voucher, and synchronizes it to the financial system.

[0028] Example 1: See Figure 2The company's raw financial and tax data includes invoice information, bank statements, contract texts, and tax returns. This data is collected from enterprise resource planning systems, online banking systems, and e-invoice platforms through pre-configured application programming interfaces (APIs). The data collection process faces challenges due to the diversity of formats. Invoice information may include scanned copies of paper invoices in image format and structured electronic data files; bank statements are text files with specific delimiters; contract texts are mostly unstructured natural language documents; and tax returns are semi-structured tabular data. The multimodal data fusion processing module needs to deploy a series of data parsers to handle this complexity. The invoice information parser integrates an optical character recognition (OCR) engine for processing image invoices. The OCR engine is optimized for Chinese financial documents and can accurately identify key fields such as invoice code, number, invoice date, buyer and seller information, product and service names, quantity, unit price, amount, tax rate, and tax amount. Price-tax separation feature calculation is the core step in invoice information processing. The system reads the identified amount and tax rate fields and performs arithmetic operations to generate the amount excluding tax and the tax amount. The relationship matrix between the tax-exclusive amount and the tax amount is constructed based on the invoice line items. Each line of goods or services details generates a matrix row. The columns of the matrix contain three values: tax-exclusive amount, tax amount, and tax rate. Multiple lines of details from the same invoice together form a two-dimensional numerical matrix, which preserves the complete logic of the price-tax relationship in the original invoice.

[0029] The analysis of bank statements focuses on understanding the temporal patterns and relationships of fund movements. Counterparty feature extraction uses pattern matching algorithms to analyze statement summary information, identifying the counterparty's account name, account number, and transaction type. The temporal features of the amounts are sorted by the timestamps of the transactions, forming a time-series dataset. The generation of the fund flow topology uses graph theory concepts, abstracting each participating bank account as a graph node and each fund transfer as a directed edge, with the edge's direction from the payer's account to the payee's account. The edge weight is the transfer amount, and edge attributes include the transaction date and summary. The fund flow topology dynamically displays the enterprise's financial trajectory over a period of time, revealing the main fund transfer partners and the circulation path of funds. Contract text processing relies on natural language processing technology. The entity recognition model is built based on a pre-trained financial domain language model, enabling accurate location of key entities from lengthy contract clauses, such as the total contract amount, payer, payee, payment node, payment ratio, payment period start and end dates, and penalty clauses. Extracting payment terms and credit period constraints requires a deeper level of semantic understanding. The system analyzes the dependency syntax tree of contract statements to identify phrase structures that represent time conditions, payment terms, and obligations. These unstructured language descriptions are then transformed into structured data fields. For example, "95% of the contract amount will be paid within 30 days after the goods are received and accepted" can be parsed as payment terms of "goods received and accepted", credit period of 30 days, and payment ratio of 95%.

[0030] Aligning the association matrix, fund flow topology, and payment terms with payment period constraints along a timeline is a crucial data integration step. The timeline alignment algorithm establishes a unified timeline based on natural days. Each data record is assigned a precise timestamp: for invoice information, the timestamp is the invoice date; for bank statements, the timestamp is the transaction date; and for contract terms, the timestamp is the agreed payment due date or service delivery date. Dynamic time warping is used to handle minor timestamp discrepancies that may exist between different data sources. For example, the arrival date of a bank statement may be delayed by one or two days from the agreed payment date in a contract. The dynamic time warping algorithm can tolerate this discrepancy within this time window and treat it as a valid alignment. The aligned data units form a multidimensional data point, each containing invoice and tax information, fund flow information, and contractual agreement information within the same time window. The generation process of the structured financial and tax feature set involves feature encoding of the aligned multidimensional data points. The feature encoder normalizes numerical data to a specific interval, performs one-hot encoding on categorical data, and extracts graph-structured data into node degree centrality, edge weights, and isograph features. The resulting structured financial and tax feature set is a highly structured dataset, typically existing in tabular or tensor form. Each row represents a unique point in time or business event, and each column represents a fused financial and tax feature, such as "current period sales revenue excluding tax," "current period certified input tax," "amount received from core customers," "amount paid to major suppliers," "contractual receivables," and "contractual payables." This structured financial and tax feature set provides standardized, high-quality data input for subsequent financial and tax business scenario segmentation and machine learning model analysis, enabling raw data from different modalities to be consistently understood and processed within a unified feature space.

[0031] The commodity code in the invoice information is associated with the transaction summary in the bank statement through semantic similarity calculation. The semantic similarity model maps the text description to a high-dimensional vector, and determines whether two records describe the same business substance by calculating the cosine similarity between the vectors. The payment period in the contract text is fuzzily matched with the arrival date in the bank statement. The matching algorithm allows for a configurable time tolerance, thus accurately linking the contractually agreed creditor-debtor relationship with the actual fund receipt and payment behavior. The entire multimodal data fusion processing flow is executed on a distributed computing framework to meet the needs of large-scale data processing. Data flows through multiple stages in the pipeline, including parsing, cleaning, transformation, integration, and encoding. Each stage has a data verification mechanism to ensure data integrity and accuracy. The generated structured financial and tax feature set is ultimately output as a standardized data exchange format, such as a columnar storage format with time-series indexing, facilitating efficient storage and fast retrieval, laying a solid data foundation for subsequent intelligent methods of financial and tax agency accounting based on machine learning. The structured financial and tax feature set not only contains the numerical information of the original data, but more importantly, it contains the deep business logic relationships between the data through fusion processing. This enables machine learning models to learn patterns and make decisions based on the essence of the business rather than just the numerical surface.

[0032] Example 2: See Figure 3The structured financial and tax feature set is categorized into revenue recognition scenarios, cost aggregation scenarios, and tax accounting scenarios according to preset financial and tax business scenario classification rules. The feature data for each scenario is then routed to the corresponding dedicated machine learning analysis module. The structured financial and tax feature set for revenue recognition scenarios mainly includes feature dimensions related to sales revenue, such as contract signing date, goods or service delivery date, invoice issuance date, amount received, and payment time—both temporal and numerical features. For the anomaly detection task in revenue recognition scenarios, the system loads a pre-trained Isolation Forest algorithm model. The Isolation Forest algorithm is an unsupervised learning method suitable for anomaly detection in high-dimensional data. Its principle is to construct multiple isolation trees by randomly selecting features and split points. Each isolation tree recursively and randomly splits the data space, so that anomalous data points, due to differences in features from the normal group, usually require fewer splits to be isolated. The model calculates the average path length of each data instance in the forest; the shorter the path length, the easier the instance is to be isolated, thus being judged as an anomaly. In revenue recognition scenarios, anomalies primarily manifest as revenue recognition across accounting periods, meaning the timing of revenue recognition deviates from the accrual basis principle stipulated by accounting standards. Examples include recognizing revenue belonging to the next accounting period prematurely or delaying the recognition of revenue that should be recognized in the current period. The Isolation Forest algorithm model analyzes the characteristic relationship patterns between the revenue recognition date and the contractually agreed delivery date, service completion date, or invoice issuance date to identify anomalous data points that significantly deviate from the mainstream pattern in the time series, and marks these points as revenue recognition anomalies across accounting periods.

[0033] The structured financial and tax feature set for cost allocation scenarios focuses on the occurrence and attribution information of costs and expenses. Features include cost occurrence date, cost amount, cost type, benefiting department or project code, and cost allocation basis. Anomaly detection in cost allocation scenarios employs clustering algorithms, specifically K-means clustering. K-means clustering aggregates cost data points with similar characteristics into several clusters, with the center of each cluster representing a typical cost allocation pattern. The system trains a clustering model based on historical compliant data to determine the optimal number of clusters and the coordinates of the center points of each cluster. When processing new cost allocation data, the model calculates the distance from each cost data point to all cluster centers and assigns it to the nearest cluster. Anomalies in cost allocation ratios manifest as a cost data point's Euclidean distance to its nearest cluster center being significantly greater than the average distance from other data points within that cluster to the cluster center, or the data point being assigned to a very small cluster with abnormal characteristics. This deviation may indicate that costs have been incorrectly allocated to unrelated departments or projects, or that the allocation ratio does not comply with established internal management regulations or accounting standards. After the clustering algorithm identifies these outliers, the system will mark them with abnormal features indicating deviations in cost sharing ratios.

[0034] The structured tax feature set in the tax accounting scenario involves characteristics related to the calculation of a company's tax payable, such as historical time-series data on taxable income, applicable tax rates, tax reductions and exemptions, prepaid taxes, and declared tax payable for various tax types. Anomaly detection in the tax accounting scenario relies on time-series prediction models, such as the Autoregressive Integral Moving Average (ARIMA) model. ARIMA models can capture trends, seasonality, and periodicity in time-series data. The system trains the ARIMA model using historical theoretical and actual tax burden data for the company. The theoretical tax burden is the tax payable calculated according to tax laws, while the actual tax burden is the tax actually declared and paid by the company. The trained model can predict a reasonable range for the theoretical and actual tax burden of the current period based on historical data. The anomaly detection module compares the actual tax burden value declared for the current period with the theoretical tax burden value predicted by the model and calculates the difference between the two. When the absolute value or relative percentage of the difference exceeds the preset dynamic threshold, the system determines that there is an abnormality in tax accounting. The difference exceeding the threshold indicates that there is a significant inconsistency between the actual tax burden and the theoretical value calculated based on historical patterns and current tax laws. This may be due to calculation errors, improper application of policies, or deliberate tax avoidance. The system will mark this as an abnormal feature.

[0035] The output of the anomaly feature markers triggers a dynamic verification process, a targeted data checking and verification phase. When the Isolation Forest algorithm detects anomalies in revenue recognition across periods, the dynamic verification process initiates a payment consistency verification sub-process. This sub-process automatically retrieves the original contract text and bank statement records associated with the abnormal revenue record. The contract text provides the basis for the rights and obligations of revenue recognition, while the bank statement provides objective evidence of fund receipts and payments. The system cross-verifies the payment conditions and times stipulated in the contract terms with the actual amount and date received in the bank statement to check for inconsistencies such as whether the revenue recognition conditions stipulated in the contract have been met but the funds have not arrived, or whether the funds have arrived but the contractual obligations have not been fulfilled. The verification results are recorded in the verification result log.

[0036] When the clustering algorithm identifies an abnormal deviation in the cost allocation ratio, the dynamic verification process triggers a cost tracing sub-process. The system traces the original invoice information based on key information from the abnormal cost records, such as invoice numbers and supplier information. By analyzing the detailed content of the invoice information, especially material codes, service descriptions, and associated project numbers, the system verifies the rationality of cost allocation. For example, it checks whether the material code matches the material requirements of the benefiting project, and whether the service description matches the actual work content of the project, thereby determining whether the cost allocation object and ratio are correct. This tracing and matching result constitutes an important part of the verification result. When the time-series prediction model finds that the difference between the actual tax burden and the theoretical tax burden exceeds a threshold, the dynamic verification process initiates a tax clause verification sub-process. The system will specifically verify the reference records in the current tax return regarding special tax treatment clauses such as tax reductions and exemptions, preferential tax rates, and additional deductions. By comparing the tax reduction and exemption bases filled in the declaration form with the currently effective tax preferential policies issued by the State Taxation Administration item by item, the system checks for issues such as incorrect citation of clauses, unmet applicable conditions, or inaccurate calculation bases. Any discrepancies or omissions discovered are recorded in detail in the verification results. The entire dynamic verification process relies on a business rule engine; different anomaly markers are mapped to different sets of verification rules, ensuring the accuracy and efficiency of the verification process. The verification results ultimately guide subsequent correction operations for abnormal data in the structured financial and tax feature set.

[0037] See Figure 4 In a machine learning-based intelligent accounting and tax agency system, the cost deviation monitoring module visualizes the cost execution deviations of each department through bar charts. Specifically, the cost deviation rate is calculated as the relative percentage difference between actual costs and the budget baseline, with negative deviations indicating cost savings and positive deviations indicating cost overruns. The vertical axis of the chart is sequentially arranged by department name, and the horizontal axis quantifies the percentage deviation rate. Warning and abnormal thresholds are integrated as dynamic monitoring trigger points. In the parameter configuration, the warning threshold is typically set within a ±5% range based on historical compliance data, and the abnormal threshold is set at a ±10% critical value to identify significant deviations requiring priority intervention. The deviation rate data for each department originates from the machine learning clustering analysis output of the cost aggregation scenario. Departments with deviation rates exceeding the threshold (such as the sales department at 11.9% and the R&D department at 10.7%) automatically trigger abnormal feature marking and are linked to the cost traceability sub-process in the dynamic verification process, verifying the rationality of cost allocation through invoice information parsing and project code matching. The visualization elements use dark gray bar charts to encode the deviation range, helping to quickly locate the negative deviation of -7.6% in the finance department and the slight positive deviation of 3.0% in the administration department. The overall structure supports horizontal comparison and trend warning across multiple departments.

[0038] Example 3: The verification results generated by the dynamic verification process accurately pinpoint the specific abnormal fields and their potential error nature within the structured financial and tax feature set. The historical compliance case library, as a knowledge base, stores a large number of audited and verified correct accounting cases and their corresponding business background data. The construction of the historical compliance case library relies on the structured processing of historically archived compliant accounting vouchers and their attachments. Each case is represented as a high-dimensional feature vector, with dimensions including business type code, transaction amount range, core accounting subjects involved, tax processing method, associated contract type, and index of applicable tax policy clauses. The correction process begins with similarity matching calculation. The system also represents the abnormal structured financial and tax feature set to be corrected as a feature vector, and finds the most similar compliant case by calculating its distance to the feature vectors of all cases in the historical compliance case library. The distance metric uses a weighted Euclidean distance formula, which fully considers the differences in the importance of different features in the business substance.

[0039] in: This represents the distance between example i to be amended and example j in the historical compliance case library, where n is the total dimension of the feature vector. It is the value of the example to be amended on the k-th feature. It is the value of historical case j on the k-th feature. These are the weight coefficients assigned to the k-th feature. The values ​​are pre-set by domain experts based on the importance of the features in reflecting the business essence. Key features such as business type and core subjects have higher weights, while some auxiliary descriptive features have lower weights. The system calculates the examples to be amended and all cases in the library. Choose the value. The first historical compliance case with the smallest value is selected as the candidate set of similar scenarios.

[0040] After the replacement operation, the structured financial and tax feature set enters the tax rule compliance verification stage, which is performed by a rule checker component. The rule checker embeds a set of logical rules abstracted from tax laws and regulations. These rules are encoded as executable judgment conditions, such as "the input tax deduction amount shall not exceed the current period's output tax amount" and "the deduction amount for business entertainment expenses shall not exceed five per thousand of the annual sales revenue." The rule checker traverses the corrected structured financial and tax feature set, checking whether the value of each field related to tax processing meets all relevant rule conditions. The verification result has two states: verification pass means the corrected data is compliant in tax processing and can proceed to the next stage. Verification failure triggers an iterative adjustment process, which analyzes the specific rule constraints violated by the fields that failed verification. The system accesses the historical compliance case library to statistically analyze the distribution range of compliant values ​​for this field in similar business scenarios, such as statistically analyzing the proportion of actual business entertainment expenses to revenue in all similar cases. Based on this statistical distribution, the system dynamically calculates a new, narrower value range. This range must simultaneously satisfy both historical distribution patterns and the upper limit requirements of tax law rules. Then, the system replaces the original field value with the median of this new range or a randomly generated, reasonable value falling within the range, followed by a new round of tax rule compliance verification. This iterative process continues until the field value passes all rule checks or reaches a preset maximum iteration threshold. Through this iterative adjustment based on historical case statistics and rule constraints, the system can find a reasonable correction value for abnormal data while complying with tax laws.

[0041] The revised and verified structured financial and tax feature set is fed into the financial and tax rule engine to generate draft accounting vouchers. The engine's generation logic is built on a rule node network. Rule nodes are the engine's basic processing units, each corresponding to a specific accounting standard or tax regulation clause. The encoding process of rule nodes transforms the natural language descriptions of standards and regulations into machine-readable logical expressions and triggering conditions. For example, a rule node might be encoded as "IF business type is 'sales' AND invoiced THEN: Debit: Accounts Receivable, Credit: Main Business Revenue, Taxes Payable - VAT Payable (Output Tax)". Rule nodes are connected by directed edges, forming a complex dependency network; the execution of some rules requires the results of other rules as a prerequisite. The business type parsing module analyzes the input structured financial and tax feature set, extracting core business attributes, such as whether the transaction is a purchase, sales, expense reimbursement, or asset acquisition. Business attributes are used to activate the corresponding rule node subgraph in the financial and tax rule engine; only rule nodes related to the current business are loaded into the working memory for execution. Activated rule nodes are not executed randomly. A rule node priority sorting mechanism ensures the accuracy of processing logic. Priority sorting is based on the topological order of rule node dependencies. Nodes without preconditions are executed first, while nodes that depend on the output of other nodes are executed later. The execution of high-priority rule nodes may provide crucial contextual information to low-priority nodes.

[0042] The core step in generating journal entries is matching debit and credit accounts according to the priority order of rule nodes. The engine traverses the active and ready rule nodes in its working memory. Each rule node checks whether the feature values ​​in the structured financial and tax feature set meet its condition. If the condition is met, the rule node is triggered, and its action part is executed. The action part of the rule node usually contains instructions to generate accounting entries, specifying the debit account, credit account, and the calculation method for the amount. The amount calculation method may be directly taken from the feature values, or it may be based on the arithmetic result of several feature values. The engine maintains a temporary voucher draft, continuously adding the journal entries generated by the triggered rule nodes. Throughout the matching process, the engine continuously performs debit and credit balance checks, checking whether the sum of the debit amounts of all generated entries is equal to the sum of the credit amounts. If an imbalance occurs, the engine will backtrack the execution process, attempting to adjust the triggering order of the rule nodes, or activate some auxiliary rule nodes with balancing functions to supplement the difference journal entries. After a series of rule triggers, account matching, and balance checks, the tax and finance rule engine outputs a preliminary draft accounting voucher. This draft contains complete accounting entries, account codes, amounts, and business summary information, preparing for subsequent logical conflict detection.

[0043] Example 4: Referring to Table 1, the draft accounting voucher is a preliminary set of accounting entries automatically generated by the tax and finance rules engine based on the modified structured tax and finance feature set. The task of the logic conflict detection module is to scan the internal consistency and compliance of these entries. The detection process first traverses each accounting subject in the draft accounting voucher and checks its account balance direction attribute. The account balance direction is a basic attribute that specifies whether an account is added to the debit or credit side. For example, an increase in an asset account is recorded as a debit, and an increase in a liability account is recorded as a credit. The logic conflict detection module has a built-in constraint condition knowledge base for rule nodes. The knowledge base stores rules extracted from accounting standards and internal accounting policies. These rules define the reconciliation, correspondence, and prohibition relationships between accounts. The detection of redundant entry conflicts focuses on the identity of economic substance. The system analyzes whether there are multiple entries in the draft accounting voucher pointing to the same economic transaction and whether their debit and credit accounts have overlapping or reversible relationships. The detection algorithm groups the entries by business order number or timestamp and performs a debit and credit account correlation analysis on all entries within the same group. When the system detects that a debit entry in the same group has the exact same debit account as another entry in the same group, and that the amounts are equal or similar, the system determines that these two entries meet the reversal criteria. Entries meeting the reversal criteria mean that they may be duplicate records of the same transaction, or that an erroneous record and its correction exist simultaneously. Marking these as redundant entries improves the simplicity and accuracy of accounting information. For tax treatments and other transactions strictly governed by regulations, the system maintains a mapping table between tax accounts and tax regulations (see Table 1). This mapping table clearly specifies the specific tax laws and administrative regulations that must be referenced in accounting vouchers when dealing with specific tax treatments. The logical conflict detection module checks the usage records of each tax account in the draft accounting voucher to verify whether it carries a regulatory clause reference identifier that meets the requirements of the mapping table. Tax account usage records where no matching reference identifier is found are marked as clause reference conflicts, indicating that the tax treatment may lack a clear legal basis.

[0044] Table 1: Mapping Relationship Between Tax Items and Regulatory Clauses Tax Account Code Tax Item Name Business Scenarios Required legal clause number Effective date of terms 22210105 Taxes Payable - Value Added Tax - Output Tax Sales of goods under the general tax calculation method Article 1 of the Provisional Regulations on Value-Added Tax 2017-11-19 22210108 Taxes Payable - Value Added Tax - Input Tax Calculating tax deductions for purchased agricultural products Article 2 of the "Pilot Implementation Measures for the Determination and Deduction of Input Tax on Agricultural Products" 2012-04-06 222105 Taxes Payable - Income Tax Payable Tax incentives for high-tech enterprises Article 28 of the Enterprise Income Tax Law 2008-01-01 222112 Taxes Payable - Individual Income Tax Payable Annual bonus taxable Article 1 of the "Notice on Issues Concerning the Transition of Preferential Policies After the Amendment of the Individual Income Tax Law" 2018-12-27 The conflict detection report is the output of the logical conflict detection module. The report lists all detected conflicts in a structured format. Each conflict record includes a unique identifier for the conflict, the conflict type, the draft accounting voucher number involved, the specific account code related to the conflict, a conflict description, and a summary of the system's suggested resolution strategy. The conflict detection report provides clear targets for subsequent adjustment operations, making the correction process more targeted. Adjusting the account mapping relationship of the draft accounting voucher based on the conflict detection report is a fine-grained correction stage. The adjustment strategy for redundant entry conflicts is to merge the debit and credit amounts of related accounts and recalculate the net value. The system identifies all groups of entries marked as redundant and summarizes the debit and credit amounts of these entries separately. For a set of redundant entries, the system calculates the net value of its total debit and credit amounts. The algebraic sign of the net value determines the posting direction of the new entry. After the net value calculation is completed, the system deletes the original set of redundant entries and replaces it with a new entry. The new entry's account is taken from the core business account in the original entry, and the amount is the calculated net value. This consolidation process eliminates duplicate entries in accounting, ensuring that accounting vouchers accurately reflect the net economic impact. For example, the original draft might contain an entry debiting accounts receivable and crediting main business revenue, followed by another entry debiting bank deposits and crediting accounts receivable with the same amount. The system detects that these two entries meet the reversal criteria in the accounts receivable account and consolidates them into a single entry debiting bank deposits and crediting main business revenue.

[0045] The adjustment strategy for conflicting tax regulations involves supplementing missing tax regulation reference information. Based on the specific tax accounts and business scenarios indicated in the conflict detection report, the system queries the mapping table between tax accounts and tax regulations to find the required tax regulation number and its specific processing rule description. The system appends the corresponding tax regulation number as metadata to the relevant entry line in the draft accounting voucher. For some complex clauses, special processing rules (such as limited deductions or installment deductions) may require generating additional memorandum records or modifying the calculation logic of the amount. Supplementing the reference information ensures that every tax-related entry is legally sound and verifiable, enhancing the compliance and auditability of accounting records. After the account mapping relationship is adjusted, before generating the final accounting voucher, the system will again activate the logic conflict detection module to review the adjusted draft voucher. The review process ensures that all conflict markers have been cleared, the draft accounting voucher meets the basic principles of double-entry bookkeeping—that every debit must have a corresponding credit and that debits must equal credits—and that all necessary clause references are complete. The final generated accounting voucher is a complete, logically consistent set of accounting entries that comply with accounting standards and tax regulations. Its format is standardized and includes elements such as voucher date, voucher number, summary, account code, account name, debit amount, credit amount, and number of attachments. It can be seamlessly integrated with financial software systems for posting processing, completing the intelligent conversion from raw data to compliant accounting vouchers.

[0046] Example 5: The fiscal and tax policy environment is constantly changing. Authoritative institutions such as the State Taxation Administration and the Ministry of Finance will periodically issue announcements, notices, or documents to revise, supplement, or interpret existing accounting standards and tax laws and regulations. The real-time monitoring module connects to officially designated policy release platforms via application programming interfaces (APIs). These platforms include the "Tax Policy" section of the State Taxation Administration's official website, the "Government Information Disclosure" section of the Ministry of Finance's official website, and authoritative information sources such as the China Government Procurement Network. The monitoring module accesses these interfaces at a preset frequency to scan for new policy documents. The scanning cycle can be set to once per hour or several times per day to balance system load and the timeliness of policy response. After a new policy document is captured, it enters the parsing and keyword extraction process. Policy documents are usually published in PDF, WORD, or HTML formats. The file parser first converts them into plain text format for machine processing. The text preprocessing step removes formatting marks, headers, footers, and irrelevant announcement header information, retaining the core content of the policy document. The keyword extraction model is built upon a professional dictionary and terminology database in the fields of finance and taxation. This database covers core terms such as accounting subject names, tax types, tax processing actions, tax rates, limits, and conditional constraints. The extraction process combines precise matching based on the terminology database with a statistical TF-IDF algorithm. The TF-IDF algorithm measures the importance of a word in a single policy document. A word that appears frequently in a single document but infrequently in historical policy documents has a high TF-IDF weight and is more likely to represent a key point in the current policy update. For example, suppose the Ministry of Finance issues a document titled "Announcement on Improving the Policy of Additional Deduction for R&D Expenses." The document parser extracts the full text. With the assistance of the terminology database, the keyword extraction model identifies high-weight keywords such as "R&D expenses," "additional deduction," "intangible assets," "manufacturing," and "amortization." These keywords are extracted and formed into a structured keyword list, with each keyword associated with its context and position in the original text.

[0047] The correlation detection between the changed policy document keywords and the rule nodes of the tax and finance rule engine is the logical judgment step that triggers the update. Each rule node in the tax and finance rule engine is associated with a set of tags, which describe the business scenarios covered by the rule node, the accounting subjects involved, and the tax policy clause numbers referenced. The system calculates the similarity between the extracted policy document keyword list and the tag set of all rule nodes in the tax and finance rule engine. The similarity calculation can use a vector space model, representing the keyword list and the rule node tag set as vectors in a high-dimensional space, and evaluating their semantic relevance by calculating the cosine value of the two vectors. When the similarity between the tag of a rule node and the policy document keywords exceeds a set threshold, the system determines that this new policy document is related to the rule node of the tax and finance rule engine. The correlation detection results generate a list of rule nodes to be evaluated and trigger the version iteration process of the rule nodes.

[0048] The version iteration of rule nodes is a rigorous update process. The system does not immediately overwrite existing rules with new policies, but follows version management principles. The system creates a copy of the current version for the affected rule nodes, marking it as a version to be updated. The full text of the new policy document and the key information parsed are associated with this version. Domain experts or system administrators will be notified and need to review and confirm this rule iteration. The review interface will display the current version logic of the rule node, the relevant clauses of the new policy document, and possible modification suggestions initially analyzed by the system based on the new policy, side-by-side. After expert confirmation, the formal update of the rule node begins. Rule node version iteration includes updating node attributes, such as node name, trigger conditions, logical expressions, and output actions. For the example of R&D expense deduction, attributes in the rule node regarding the deduction ratio, applicable industry scope, and expense collection scope need to be updated according to the new announcement. The connection edges between rule nodes may also need to be adjusted to reflect the changes in business processes brought about by policy changes. After the version iteration is completed, the new version of the rule node is deployed to the financial and tax rule engine.

[0049] The tax and accounting rules engine automatically invokes the new version of the rule nodes when processing related business and generating draft accounting vouchers. For example, when the system processes an expense reimbursement form with the "R&D expense" attribute, the activated rule node is no longer the one that calculates the additional deduction at 75% in the old version, but rather the one that calculates it at 100% in the new version. This means that in the generated draft accounting voucher, the amounts of income tax expense and deferred income tax asset will be automatically and accurately calculated according to the latest policies. The entire real-time monitoring and iteration process ensures that the machine learning-based intelligent accounting and accounting method can continuously adapt to changes in the external regulatory environment without human intervention, and the generated accounting records always comply with the latest compliance requirements. The system records complete monitoring logs, including the source of the policy document, capture time, extracted keywords, affected rule nodes, iteration version number, iteration time, and operator identifier, forming a complete audit trail to facilitate tracing the impact of policy changes on specific accounting treatments.

[0050] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus.

[0051] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.

Claims

1. A machine learning-based intelligent method for financial and tax agency bookkeeping, characterized in that, include: Collect raw financial and tax data from enterprises, including invoice information, bank statements, contract texts, and tax returns. The original data is subjected to multimodal data fusion processing to generate a structured set of fiscal and tax features; According to the preset financial and tax business scenario classification rules, the structured financial and tax feature set is classified into revenue recognition scenario, cost collection scenario and tax accounting scenario. A machine learning model is used to perform anomaly detection on the structured financial and tax feature set for each scenario, and anomaly feature labels are output. Based on the aforementioned abnormal feature markers, a dynamic verification process is triggered to generate verification results; Based on the matching degree between the verification results and the historical compliance case library, the abnormal data in the structured financial and tax feature set is corrected; The revised set of structured financial and tax features is input into the financial and tax rules engine to generate a draft accounting voucher; Perform logical conflict detection on the draft accounting voucher and output a conflict detection report; Adjust the account mapping relationship of the draft accounting voucher based on the conflict detection report, and generate the final accounting voucher.

2. The intelligent method for financial and tax agency bookkeeping based on machine learning according to claim 1, characterized in that, The multimodal data fusion processing of the original data includes: Extract the price-tax separation feature from the invoice information to generate a correlation matrix between the tax-exclusive amount and the tax amount; Analyze the time-series characteristics of transaction counterparties and amounts in the bank statements to construct a fund flow topology graph; Entity recognition is performed on the contract text to extract payment terms and payment period constraints. The structured financial and tax feature set is generated by aligning the correlation matrix, the fund flow topology, and the payment terms and payment period constraints along the time axis.

3. The intelligent method for financial and tax agency bookkeeping based on machine learning according to claim 1, characterized in that, The anomaly detection using a machine learning model on the structured financial and tax feature set for each scenario includes: For the structured financial and tax feature set in the aforementioned revenue recognition scenario, the isolated forest algorithm is used to detect anomalies in revenue recognition across periods. For the structured financial and tax feature set in the cost aggregation scenario, a clustering algorithm is used to identify deviations in the cost allocation ratio; For the structured set of financial and tax features in the aforementioned tax accounting scenario, a time-series prediction model is used to compare the difference between the actual tax burden and the theoretical tax burden.

4. The intelligent method for financial and tax agency bookkeeping based on machine learning according to claim 3, characterized in that, The dynamic verification process triggered based on the abnormal feature markers includes: When an anomaly in the revenue recognition across periods is detected, the associated contract text and the bank statement are retrieved to verify the consistency of receipt and payment. When a deviation in the cost allocation ratio is detected, the matching relationship between the material code and the project number in the invoice information is traced. When the difference between the actual tax burden and the theoretical tax burden exceeds a threshold, the tax reduction and exemption clause reference records in the tax return form are checked.

5. The intelligent method for financial and tax agency bookkeeping based on machine learning according to claim 1, characterized in that, The correction of abnormal data in the structured financial and tax feature set includes: Replace the erroneous fields in the abnormal data with the account processing records of similar scenarios in the historical compliance case library; The replaced fields are validated for compliance with tax rules. If the validation fails, the value range of the fields is iteratively adjusted until compliance is achieved.

6. The intelligent method for financial and tax agency bookkeeping based on machine learning according to claim 1, characterized in that, The generation logic of the tax and finance rules engine includes: Encode accounting standard entries and tax regulation provisions as rule nodes; Activate the corresponding rule node based on the business type of the structured financial and tax feature set; Debit and credit account matching is performed according to the priority order of the rule nodes.

7. The intelligent method for financial and tax agency bookkeeping based on machine learning according to claim 6, characterized in that, The logical conflict detection of the draft accounting voucher includes: Iterate through the account balance direction in the draft accounting voucher and the constraints of the rule nodes; When a debit and credit entry for the same account simultaneously meets the reversal condition, it is marked as a redundant entry conflict. When a tax item is detected that does not match the aforementioned tax regulation clause, it is marked as a clause reference conflict.

8. The intelligent method for financial and tax agency bookkeeping based on machine learning according to claim 7, characterized in that, The adjustment of the account mapping relationship of the draft accounting voucher based on the conflict detection report includes: For the aforementioned redundant journal entries, the debit and credit amounts of related accounts are merged and the net value is recalculated; In case of any conflict of reference to the aforementioned clauses, supplement the draft accounting voucher with the special handling rules in the aforementioned tax regulations.

9. The intelligent method for financial and tax agency bookkeeping based on machine learning according to claim 1, characterized in that, Also includes: Real-time monitoring of updated tax and financial regulations documents, and extraction of keywords related to clause changes; When the keyword for the clause change is related to a rule node in the tax and finance rule engine, the version iteration of the rule node is triggered.

10. A machine learning-based intelligent accounting and tax agency system, comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the intelligent financial and tax agency bookkeeping method based on any one of claims 1 to 9.