Automatic accounting and tax declaration processing system based on AI large model
Through the automatic accounting and tax filing processing system based on the AI big model, efficient and accurate data integration and tax compliance analysis are achieved, solving the problems of inefficiency and high error rate in traditional financial and tax management, reducing corporate tax risks and improving operational efficiency.
Patent Information
- Application Number
- CN202510801433.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-16
- Publication Date
- 2025-09-19
AI Technical Summary
Traditional financial and tax management methods are inefficient and error-prone in the process of data integration, bookkeeping and tax filing. They are difficult to adapt to complex business models and constantly updated tax policies, resulting in insufficient data accuracy and timeliness, and increasing corporate tax risks.
An automatic accounting and tax filing processing system based on a large AI model is used to obtain data through multi-source interfaces. Heterogeneous data alignment technology and multimodal graph neural networks are used for data cleaning and feature association analysis. The hierarchical reinforcement learning architecture is combined to generate optimal accounting rules and tax filing plans, and compliance is ensured through a dynamic compliance verification model.
It improves data processing efficiency and accuracy, reduces reliance on manual operations, reduces error rates, enhances corporate operational efficiency and tax compliance, reduces tax risks, and protects the company's tax credit.
Smart Images

Figure CN120672492A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence and financial and tax processing technology, and specifically to an automatic accounting and tax filing processing system based on an AI large model. Background Art
[0002] In today's digital age, financial and tax management has become increasingly complex and critical. As businesses continue to expand, the volume of financial data is exploding, and traditional manual bookkeeping and tax filing methods face numerous challenges.
[0003] From a data processing perspective, companies' financial data comes from a wide range of sources, including electronic invoices, bank statements, payroll information, and various tax policy documents. This data not only comes in a variety of formats—including structured, semi-structured, and unstructured—but is also stored across disparate systems and platforms. For example, electronic invoices exist as electronic documents, bank statements are provided in specific tabular formats, and payroll data and tax policy documents each have their own unique formats. This makes data integration and processing extremely difficult. Manual data collection and organization is not only inefficient but also prone to omissions and errors. Faced with this massive amount of data, manually verifying its accuracy and completeness becomes a nearly impossible task, directly impacting the quality of subsequent bookkeeping and tax filing.
[0004] Traditional bookkeeping relies on accountants to classify, record, and calculate financial data based on accounting standards and established experience. This requires deep expertise and extensive practical experience, but even with this expertise, errors can still occur due to negligence or a misunderstanding of complex business processes. With the emergence of new business models, such as online transactions and cross-border e-commerce, transaction structures and financial relationships have become increasingly complex, making it difficult for traditional bookkeeping methods to accurately reflect the economic substance of these businesses. For some emerging financial instruments and business innovations, accounting treatment methods are still underdeveloped, and manual bookkeeping often falls short, failing to meet businesses' requirements for accurate and timely financial information.
[0005] When it comes to tax filing, tax policies are frequently updated, and tax regulations vary from region to region. Companies need to stay up-to-date on the latest policies and accurately apply them to the tax filing process. However, manually interpreting and applying tax policies is not only time-consuming and labor-intensive, but can also lead to tax filing errors due to a lack of understanding of the policies, which in turn can lead to tax risks. For example, when applying preferential tax policies, if a company fails to accurately grasp the policy conditions, it may miss out on the opportunity to enjoy preferential treatment and increase its tax burden. Conversely, if preferential policies are incorrectly applied, it may face penalties from the tax authorities. Furthermore, the tax filing process involves multiple links and departments, and manual operations are prone to problems such as non-standard processes and untimely filing, which can affect a company's tax credit rating.
[0006] To address these issues, some companies have attempted to introduce simple financial software. However, existing financial software is largely single-function, capable of only basic bookkeeping functions and lacking intelligent tax policy analysis and automated tax filing capabilities. Furthermore, these software programs are limited in their ability to handle complex business processes and multi-source, heterogeneous data, failing to meet the growing demand for intelligent financial and tax management. With the rapid development of artificial intelligence technology, its application in the field of corporate bookkeeping and tax filing has become a key area for addressing these challenges, providing an opportunity for the research of this invention. Summary of the Invention
[0007] The purpose of the present invention is to provide an automatic accounting and tax filing processing system based on an AI large model to solve the problems raised in the above background technology.
[0008] To achieve the above objectives, the present invention provides the following technical solution: an automatic accounting and tax filing processing system based on an AI large model, the system comprising:
[0009] Data collection module: used to obtain enterprise financial data, tax policy documents and transaction flow information through multi-source interfaces;
[0010] Data preprocessing module: cleans and standardizes the data acquired by the data collection module based on heterogeneous data alignment technology to generate structured financial and tax data sets;
[0011] Intelligent analysis module: uses a multimodal graph neural network to perform feature correlation analysis on the structured financial and taxation dataset to extract tax compliance features;
[0012] Decision generation module: inputs the tax compliance features into the pre-trained tax strategy generation model and outputs the optimal accounting rules and tax filing plan;
[0013] Execution feedback module: Build a dynamic compliance verification model based on the tax reporting solution, and realize the atomic execution and status feedback of financial and tax operations through a distributed transaction management framework.
[0014] Preferably, the acquiring of enterprise financial data through a multi-source interface includes:
[0015] Multi-source interfaces include electronic invoice interface, bank reconciliation interface, payroll management interface and tax declaration interface;
[0016] Align the timestamps of electronic invoice data and bank transaction data to build a transaction graph; perform semantic analysis on salary data and tax policy data to generate a policy correlation matrix;
[0017] A two-branch feature fusion model is constructed. The first branch uses a graph attention network to extract the topological relationship of the transaction graph, and the second branch uses a temporal convolutional network to extract the version evolution features of the policy association matrix.
[0018] The topological relationship and version evolution characteristics are integrated through a cross-modal gating mechanism to generate a multi-dimensional fiscal and taxation feature vector; the multi-dimensional fiscal and taxation feature vector is dynamically updated based on a gated recurrent unit, and a comprehensive analysis result including transaction link integrity, policy matching and risk warning level is output.
[0019] Preferably, the tax strategy generation model adopts a hierarchical reinforcement learning architecture to generate a tax filing plan based on a multi-objective optimization mechanism; the hierarchical reinforcement learning architecture includes:
[0020] Construct a fiscal and taxation-policy interaction graph, where nodes include tax type nodes, enterprise nodes, regulation nodes, and audit nodes. Node attributes include tax rates, deduction items, and declaration deadlines.
[0021] A two-stage attention mechanism is used. In the first stage, the spatial graph convolution layer is used to calculate the dependency weights between tax type nodes and associated nodes. In the second stage, the temporal graph convolution layer is used to screen the compliance of historical declaration records.
[0022] Node features are iteratively optimized based on a multi-head strategy network. Each strategy head integrates node attributes with real-time audit rules; ultimately, a dynamic tax filing strategy that meets the tax policies of multiple regions is output.
[0023] Preferably, the dynamic compliance verification model integrates rule reasoning and anomaly detection strategies, including:
[0024] The compliance verification problem is modeled as a multi-constraint satisfaction problem, where the decision variables include the declared amount, deduction ratio, and declaration time window;
[0025] Initialize the rule engine and load the tax regulations knowledge base, using a dynamic priority mechanism to adjust rule weights based on the frequency of policy updates;
[0026] In the inference phase, the constraint propagation algorithm is used to verify the logical consistency of the declared data; in the detection phase, the isolation forest algorithm is used to identify abnormal transaction patterns;
[0027] The atomicity of the multi-threaded verification process is guaranteed through a distributed lock mechanism, and an audit log with a timestamp is generated.
[0028] Preferably, the heterogeneous data alignment technology adopts a semantic embedding and entity linking method, including:
[0029] Perform paragraph-level word segmentation and named entity recognition on unstructured financial documents to extract accounting subjects and amount entities;
[0030] Build a domain knowledge graph and match the extracted entities with the standard subject nodes in the graph based on semantic similarity;
[0031] A generative adversarial network is used to correct the matching results and generate a high-confidence alignment data table.
[0032] Preferably, the construction of the transaction graph is achieved through temporal relationship coding, including: defining the temporal dependency between transaction nodes, including payment cycle, invoice issuance time and arrival delay; modeling the dynamic correlation strength between nodes through time-aware graph convolutional network, and generating a weighted transaction path sequence.
[0033] Preferably, the rule engine is implemented in a logic programming language, including: converting tax regulations into predicate logic rules and compiling them into executable inference trees; embedding fuzzy logic operators in the inference trees, processing grayscale intervals in policy terms, capturing rule conflict events through a real-time monitoring module, and triggering a manual review process.
[0034] Preferably, the adversarial generative network adopts a conditional adversarial training framework, including: the generator reconstructs the alignment data table through a variational autoencoder, and the discriminator judges the authenticity of the data through a convolutional neural network; a domain classifier is introduced during the training process to force the generated data to be aligned with the distribution of the target tax account.
[0035] Preferably, the time-aware graph convolutional network is enhanced by relative time coding, including: mapping the time difference of transaction nodes into a sinusoidal position coding vector and splicing it with the node attributes; introducing a time decay factor in each layer of graph convolution operation to dynamically adjust the contribution weight of historical transactions.
[0036] Preferably, an automatic accounting and tax filing processing method based on an AI large model is applied to the automatic accounting and tax filing processing system as described above, and the method comprises the following steps:
[0037] Step 1: Obtain enterprise financial data, tax policy documents, and transaction flow information through the multi-source interface of the data acquisition module;
[0038] Step 2: Using the data preprocessing module and heterogeneous data alignment technology, the acquired corporate financial data, tax policy documents, and transaction flow information are cleaned and standardized to generate a structured financial and tax data set.
[0039] Step 3: Use the multimodal graph neural network in the intelligent analysis module to perform feature correlation analysis on the structured fiscal and taxation dataset to extract tax compliance features;
[0040] Step 4: Input the extracted tax compliance features into the pre-trained tax strategy generation model in the decision generation module to output the optimal accounting rules and tax filing plan;
[0041] Step 5: Build a dynamic compliance verification model based on the tax filing plan through the execution feedback module, and use the distributed transaction management framework to achieve atomic execution and status feedback of financial and tax operations.
[0042] Compared with the prior art, the present invention has the following beneficial effects:
[0043] In terms of data processing efficiency and accuracy, the data acquisition module can quickly obtain corporate financial data, tax policy documents and transaction flow information through multi-source interfaces, covering multiple channels such as electronic invoice interfaces and bank reconciliation interfaces. The integration of multi-source data greatly improves the comprehensiveness of data acquisition and avoids data omissions that may occur in manual collection. The data preprocessing module uses heterogeneous data alignment technology to perform paragraph-level word segmentation and named entity recognition on unstructured financial documents, extract accounting subjects and amount entities, and then perform semantic similarity matching by constructing a domain knowledge graph. With the help of adversarial generative network error correction, a high-confidence aligned data table is generated. This series of operations realizes the efficient cleaning and standardization of multi-source heterogeneous data. Compared with traditional manual processing methods, it greatly reduces data processing time and keeps the data error rate at an extremely low level, providing an accurate and reliable data foundation for subsequent accounting and tax reporting work.
[0044] The intelligence level of bookkeeping and tax filing has been significantly improved. The intelligent analysis module uses a multimodal graph neural network to perform feature correlation analysis on structured fiscal and tax datasets, which can deeply explore the potential connections between data and accurately extract tax compliance features. The decision generation module inputs these features into a pre-trained tax strategy generation model. This model uses a hierarchical reinforcement learning architecture and a multi-objective optimization mechanism, considering multiple factors such as tax type, enterprise, regulations, and audits. It constructs a fiscal and tax-policy interaction graph, and through iterative optimization using a two-stage attention mechanism and a multi-head strategy network, ultimately outputs the optimal bookkeeping rules and tax filing solutions. This process is entirely based on intelligent algorithms, breaking away from excessive reliance on human experience. It can adapt to complex and changing business scenarios and constantly updated tax policies, providing enterprises with accurate and efficient bookkeeping and tax filing guidance.
[0045] Effectively reduce tax risks. The dynamic compliance verification model integrates rule reasoning and anomaly detection strategies, modeling the compliance verification problem as a multi-constraint satisfaction problem. By initializing the rule engine and loading the tax regulations knowledge base, a dynamic priority mechanism is used to adjust rule weights. During the reasoning phase, a constraint propagation algorithm is used to verify the logical consistency of the declared data. During the detection phase, an isolation forest algorithm is used to identify abnormal transaction patterns. At the same time, a distributed locking mechanism ensures the atomicity of the multi-threaded verification process and generates a timestamped audit log. This enables companies to promptly identify and correct potential tax risk points during the tax filing process, avoiding financial losses such as fines and late payment fees due to tax violations, maintaining their tax credit rating, and ensuring their legal and compliant operations.
[0046] Improve business operational efficiency and save costs. Traditional accounting and tax filing methods require a large amount of manpower. From data collection and organization to accounting and tax filing, each link requires accountants to spend a lot of time and energy. However, this invention realizes the automation of accounting and tax filing, greatly reducing the manual operation links and reducing the company's dependence on the number of professional accountants, thereby saving labor costs. At the same time, automated processing reduces the duplication of work and tax risk management costs caused by human error, improves the overall operational efficiency of the enterprise, enables enterprises to invest more resources in the development of core businesses, and enhances the market competitiveness of enterprises. BRIEF DESCRIPTION OF THE DRAWINGS
[0047] Figure 1 This is a diagram showing the overall working principle of the automatic accounting and tax filing processing system based on the AI big model described in the present invention;
[0048] Figure 2 A working diagram for obtaining enterprise financial data through multi-source interfaces;
[0049] Figure 3 A diagram showing how the model works for dynamic compliance verification;
[0050] Figure 4 The figure shows the working principle of heterogeneous data alignment technology. DETAILED DESCRIPTION
[0051] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0052] See also Figures 1-4 The present invention provides an automatic accounting and tax filing processing system based on an AI big model, and its specific implementation method is described in detail below.
[0053] Data Collection Module: This module acquires enterprise financial data, tax policy documents, and transaction flow information through multi-source interfaces. These interfaces encompass various types of interfaces, connecting to various data sources related to an enterprise's finances and taxes to ensure comprehensive and accurate data collection. For example, the e-invoice interface connects to the enterprise's system for issuing and receiving e-invoices, acquiring e-invoice data in real time; the bank reconciliation interface connects to the enterprise's bank account system to acquire bank transaction data; the payroll management interface connects to the enterprise's internal payroll management system to acquire payroll data; and the tax filing interface connects to the tax department's filing system to acquire tax filing-related data. These multi-source interfaces enable data collection from diverse channels and formats, providing a rich data foundation for subsequent processing.
[0054] Data preprocessing module: Based on heterogeneous data alignment technology, the data obtained by the data collection module is cleaned and standardized to generate a structured financial and taxation data set. Since the data collected from multiple source interfaces has various formats and standards, and there is a large amount of heterogeneous data, specific technologies are required for processing. Heterogeneous data alignment technology is mainly implemented through semantic embedding and entity linking methods. First, paragraph-level word segmentation and named entity recognition are performed on unstructured financial documents to extract accounting subjects and amount entities; then, a domain knowledge graph is constructed to match the extracted entities with the standard subject nodes in the graph for semantic similarity; finally, a generative adversarial network is used to correct the matching results and generate a high-confidence alignment data table, thereby completing the cleaning and standardization of the data and generating a structured financial and taxation data set that meets the requirements of subsequent processing.
[0055] Intelligent Analysis Module: This module uses a multimodal graph neural network to perform feature correlation analysis on the structured fiscal and taxation dataset and extract tax compliance features. Multimodal graph neural networks can integrate multiple data features and fully explore the relationships between data within the structured fiscal and taxation dataset. Through in-depth data analysis, it extracts tax compliance-related features from various perspectives. These features provide important insights for subsequent decision-making, helping the system determine whether a company's tax operations comply with relevant regulations and policies.
[0056] The decision generation module inputs the tax compliance characteristics into a pre-trained tax strategy generation model and outputs the optimal accounting rules and tax filing plan. The tax strategy generation model, trained on extensive data, uses the input tax compliance characteristics and the tax policies of multiple regions. Based on a hierarchical reinforcement learning architecture, it generates the optimal accounting rules and tax filing plan through a multi-objective optimization mechanism. This plan comprehensively considers factors such as the company's actual financial situation, tax policy requirements, and potential risks, ensuring that the company's accounting and tax filing operations comply with regulations while maximizing its interests.
[0057] Execution Feedback Module: A dynamic compliance verification model is constructed based on the tax filing solution, enabling atomic execution and status feedback of financial and tax operations through a distributed transaction management framework. The dynamic compliance verification model integrates rule-based reasoning and anomaly detection strategies to monitor and verify the execution of the tax filing solution in real time. During execution, the distributed transaction management framework ensures the atomicity of financial and tax operations, ensuring that all operations are either successfully executed or rolled back, ensuring data consistency and integrity. The system also provides real-time feedback on the status of operations, allowing businesses to keep abreast of tax filing progress and take timely action to address any anomalies.
[0058] Example 1:
[0059] This example details the specific implementation of a multi-source interface in the data acquisition module to acquire enterprise financial data. In real-world applications, a company's financial data comes from a wide range of sources and is complex, making multi-source interfaces crucial. For example, a medium-sized manufacturing company's operations span multiple processes, including raw material procurement, product production and sales, and employee payroll, generating a large amount of diverse financial data.
[0060] The multi-source interface includes an electronic invoice interface, a bank reconciliation interface, a payroll management interface, and a tax filing interface. The electronic invoice interface connects the company's electronic invoice issuance and receipt system with this automated accounting and tax filing system. During the raw material procurement process, electronic invoices issued by suppliers are transmitted to the system in real time through this interface. Electronic invoice data includes details such as the invoice number, invoice date, buyer information, seller information, name of the goods or taxable services, amount, tax rate, and tax amount. The bank reconciliation interface connects to the company's bank account system to obtain bank statement data. Bank statement data records the inflow and outflow of every enterprise's funds, including transaction date, transaction amount, counterparty account information, and transaction summary. For example, when a company receives sales payment from a customer, the bank statement will record the receipt of this payment in detail.
[0061] The payroll management interface connects to the company's internal payroll management system to obtain payroll data. This data covers employees' base salary, performance-based pay, bonuses, social security and provident fund contributions, and individual income tax deductions. For sales staff, performance-based pay is calculated based on sales performance and transmitted to the system via the payroll management interface. The tax filing interface connects to the tax department's filing system to obtain tax filing data, including the types of taxes declared, the amounts declared, and the deadlines for filing.
[0062] After acquiring the data, the electronic invoice data and the bank transaction data are time-stamped and aligned to construct a transaction graph. The timing dependency between transaction nodes is defined, including the payment cycle, invoice issuance time, and account arrival delay. Assume that the payment cycle is T and the invoice issuance time is t 发票 , the arrival delay is Δt, then the time relationship between transaction nodes can be expressed as follows: If the payment cycle of a transaction is 30 days, the invoice is issued on the 10th day, and the arrival delay is 5 days, then the time from invoice issuance to fund arrival is t 到账 =t 发票 +Δt, and within this payment cycle, this transaction forms a specific time series relationship with other related transactions. The dynamic correlation strength between nodes is modeled through a time-aware graph convolutional network, and a weighted transaction path sequence is generated. The time-aware graph convolutional network is enhanced by relative time coding, mapping the time difference of transaction nodes into a sinusoidal position coding vector and concatenating it with node attributes; a time decay factor α (0<α<1) is introduced in each layer of graph convolution operation to dynamically adjust the contribution weight of historical transactions. Let the time difference between transaction nodes i and j be τ ij , the sinusoidal position encoding vector is PE(τ ij ), the node attribute is X i , then the node feature after relative time coding enhancement is In the graph convolution operation, the updated feature H of node i is i It can be expressed as: Where σ is the activation function, N(i) is the set of neighbor nodes of node i, c i is the normalization constant, W is the weight matrix, is the feature of the previous iteration of node i. In this way, the time-aware graph convolutional network can more accurately reflect the dynamic relationship between transaction nodes and construct a more effective transaction graph.
[0063] Perform semantic parsing on salary data and tax policy data to generate a policy correlation matrix. Semantically match the terms in the tax policy document with the content in the salary data. For example, associate the relevant policy terms for individual income tax with the taxable income portion of an employee's salary. For example, if the tax policy includes provisions for special additional deductions, the employee's special additional deduction information (such as children's education deductions and housing loan interest deductions) in the salary data corresponds to this policy provision. Through this semantic parsing, the degree of correlation between different salary items and tax policy provisions is determined and expressed in matrix form, forming a policy correlation matrix.
[0064] A dual-branch feature fusion model is constructed. The first branch uses a graph attention network to extract the topological relationship of the transaction graph. The graph attention network determines the importance of each node to other nodes by calculating the attention weights between nodes. Let the node in the transaction graph be v i , which is characterized by h i , the attention coefficient e between nodes i and j ij It can be expressed as: Where W is the learnable weight matrix, [h i ||h j ] means concatenating the features of nodes i and j, and LeakyReLU is the activation function. In this way, the topological relationship between nodes in the transaction graph can be extracted, highlighting key nodes and important connections. The second branch uses a temporal convolutional network to extract the version evolution characteristics of the policy association matrix. The temporal convolutional network can capture the dynamic changes in time series data. For the policy association matrix, its correlation relationship will change as the tax policy is updated. The temporal convolutional network extracts features from different time versions of the policy association matrix through convolution operations. Suppose the input policy association matrix sequence is X = [x1, x2,…, x T ], the convolution kernel is k, then the output feature after the convolution operation is Where K is the size of the convolution kernel, and t is the time step. The topological relationship and version evolution features are fused through a cross-modal gating mechanism to generate a multi-dimensional fiscal and taxation feature vector. The cross-modal gating mechanism adaptively adjusts the fusion ratio according to the importance of the two features. Let the topological relationship feature be F1, the version evolution feature be F2, and the gating weight be g, then the fused multi-dimensional fiscal and taxation feature vector F can be expressed as: F=g·F1+(1-g)·F2. The multi-dimensional fiscal and taxation feature vector is dynamically updated based on the gated recurrent unit, and a comprehensive analysis result including transaction link integrity, policy matching and risk warning level is output. The gated recurrent unit can dynamically update the feature vector based on the current input and the state of the previous moment. Let the current input be x t, the hidden state at the previous moment is h t-1 , update gate to z t , reset the gate to r t , then the update formula of the gated recurrent unit is: t =σ(W z ·[x t ||h t-1 ]), r t =σ(W r ·[x t ||h t-1 ]), Where W z 、W r , W is the weight matrix, σ is the sigmoid activation function, and tanh is the hyperbolic tangent activation function. Through continuous updates of the gated recurrent unit, the final output includes a comprehensive analysis of transaction link integrity, policy compatibility, and risk warning level, providing strong support for subsequent accounting and tax filing decisions.
[0065] Example 2:
[0066] This example focuses on a tax strategy generation model. In real-world tax scenarios, tax policies vary across regions, and businesses also vary widely. Therefore, a model is needed that comprehensively considers multiple factors to generate tax filing plans. For example, consider a multi-regional chain enterprise with stores in multiple cities. Each region has different tax policies regarding tax types, tax rates, deductions, and filing deadlines.
[0067] The tax strategy generation model uses a hierarchical reinforcement learning architecture to generate tax filing plans based on a multi-objective optimization mechanism. First, a fiscal and taxation-policy interaction graph is constructed. The nodes in the graph include tax type nodes, enterprise nodes, regulation nodes, and audit nodes. Node attributes contain important information such as tax rates, deductions, and filing deadlines. For tax type nodes, such as the VAT node, attributes include the tax rates corresponding to different business types (e.g., 13% for sales of goods and 6% for provision of services), the range of deductible input tax, etc.; enterprise nodes contain information such as the company's business scope, sales, and costs; regulation nodes record relevant tax regulations, such as tax incentives and filing requirements; and audit nodes are used to monitor and assess the compliance of a company's tax filings.
[0068] A two-stage attention mechanism is adopted. In the first stage, the dependency weights of the tax type node and the associated nodes are calculated through the spatial graph convolution layer. The spatial graph convolution layer can perform convolution operations on graph structure data to extract the spatial relationship between nodes. Let the tax type node be v i , which is characterized by h i, the set of neighbor nodes is N(i), the output features of the spatial graph convolution layer It can be expressed as: Where σ is the activation function, c i is a normalization constant, and W is a weight matrix. In this way, the dependency weights between the tax type node and other related nodes (such as enterprise nodes, regulation nodes, etc.) are calculated to determine which nodes have a greater impact on the tax type node. In the second stage, the historical declaration records are screened for compliance through the time series graph convolution layer. The time series graph convolution layer can process time series graph data. For the historical declaration records of an enterprise, tax policies and business conditions may change over time. Suppose the graph sequence composed of historical declaration records is G = [G1, G2, ..., G T ], each graph G t The node feature in is h t,i , the output features of the temporal graph convolution layer It can be expressed as: Where K is the time window size, c t,i The time series graph convolution layer can filter out historical tax declaration records that are consistent with current tax policies and the actual situation of the enterprise, providing a reference for generating a reasonable tax declaration plan.
[0069] Based on the multi-head strategy network, node features are iteratively optimized. Each strategy head integrates node attributes and real-time audit rules. Assume that there are M strategy heads. For the mth strategy head, its input is node feature h and real-time audit rule r, and the output optimization feature is It can be expressed as: where f m is the optimization function of the mth strategy head. Different strategy heads can employ different optimization methods, such as those based on different neural network structures or algorithms. Ultimately, the outputs of multiple strategy heads are integrated to produce a dynamic tax filing strategy that meets the tax policies of multiple regions. Through this hierarchical reinforcement learning architecture and multi-objective optimization mechanism, the tax strategy generation model can generate optimal tax filing plans based on the company's actual situation and local tax policies, helping companies to rationally plan their taxes and mitigate tax risks.
[0070] Example 3:
[0071] This example focuses on a dynamic compliance verification model. Ensuring the compliance of declared data is crucial during corporate tax filing. This dynamic compliance verification model enables comprehensive and rigorous monitoring and verification of the implementation of tax filing plans. For example, a large corporate group, with diverse business types and complex tax processes, requires a highly accurate compliance verification mechanism.
[0072] The dynamic compliance verification model integrates rule reasoning and anomaly detection strategies. First, the compliance verification problem is modeled as a multi-constraint satisfaction problem, with decision variables including the declared amount, deduction ratio, and declaration time window. The declared amount A must meet the tax calculation requirements of tax regulations for different business types, the deduction ratio R must comply with the deduction scope and ratio restrictions stipulated by relevant policies, and the declaration time window [t start ,t end ] must be within the declaration period specified by the tax department.
[0073] The rule engine is initialized and loaded with the tax regulations knowledge base. A dynamic priority mechanism adjusts rule weights based on the frequency of policy updates. Implemented in a logic programming language, the rule engine converts tax regulations into predicate logic rules and compiles them into executable inference trees. For example, for corporate income tax filings, tax regulations specify the calculation method for taxable income. This can be expressed as a predicate logic rule: "If a company's revenue is I, its costs are C, and its allowable deductions are D, then its taxable income E = ICD." Fuzzy logic operators are embedded in the inference tree to handle grayscale ranges within policy clauses. Because some tax policies contain ambiguous provisions, such as the flexibility of deduction standards for certain expenses, fuzzy logic operators provide more flexible handling of these situations. Furthermore, a real-time monitoring module detects rule conflicts and triggers manual review. When different tax rules are applied to the same filing data and conflict or inconsistency arises, the real-time monitoring module promptly detects and notifies relevant personnel for manual review.
[0074] In the inference phase, the logical consistency of the declared data is verified based on the constraint propagation algorithm. The constraint propagation algorithm passes constraint information between variables, gradually narrowing the value range of the variables, and thus judging whether the declared data meets all the constraints. Suppose the variable set in the declared data is V = {v1, v2, ..., v n}, the constraint set is C={c1,c2,…,c m}, for each constraint c i , which defines a relationship between variables. For example, in a VAT declaration, the output tax amount T 销项 , input tax T 进项 and the tax payable T 应纳 There is a constraint relationship T 应纳 =T 销项 -T 进项. By continuously propagating these constraint information between variables, if all variables can find values that satisfy all constraints, the declared data is logically consistent; otherwise, there is a logical error. In the detection stage, the isolation forest algorithm is used to identify abnormal transaction patterns. The isolation forest algorithm is a tree-based anomaly detection algorithm that evaluates the degree of isolation of samples by constructing an isolation tree. For each transaction in the declared data, its features are input as samples into the isolation forest model. If the path length of a transaction in the isolation tree is significantly shorter than the path length of a normal transaction, the transaction is judged to be an abnormal transaction. For example, in a company's sales transactions, if a certain sales amount is much higher or much lower than the normal level, and other features of the transaction are also significantly different from normal transactions, the isolation forest algorithm can identify that the transaction may be abnormal.
[0075] A distributed locking mechanism ensures the atomicity of the multi-threaded verification process and generates a timestamped audit log. In large enterprise groups, due to the massive volume of data, multi-threaded verification is often used to improve verification efficiency. This distributed locking mechanism ensures that only one thread can verify specific data at a time, avoiding data conflicts and inconsistencies. Simultaneously, a timestamped audit log is generated, recording the time, content, and results of each verification operation to facilitate subsequent queries and audits. The audit log format can be designed as: [timestamp, operation content, verification result], for example: [2024-10-01 10:00:00, Verification of corporate income tax declaration data, passed]. This audit log provides enterprises with detailed verification records, facilitating the tracing and management of the tax declaration process.
[0076] Example 4:
[0077] This example focuses on heterogeneous data alignment technology. In actual business operations, financial data comes from diverse sources, with vastly different formats and standards. Heterogeneous data alignment technology is crucial for efficient data processing. For example, consider a diversified group company. Its subsidiaries span manufacturing, services, and other sectors. Each subsidiary utilizes a different financial system, generating data in varying formats and standards. This places high demands on heterogeneous data alignment.
[0078] Heterogeneous data alignment technology uses semantic embedding and entity linking. First, paragraph-level segmentation and named entity recognition are performed on unstructured financial documents to extract accounting items and amount entities. In unstructured financial documents such as procurement contracts for manufacturing subsidiaries, paragraph-level segmentation is performed using natural language processing technology to break the text into meaningful words or phrases. For example, in the sentence "On October 5, 2024, our company purchased 1,000 pieces of raw materials from XX supplier at a unit price of 50 yuan for a total amount of 50,000 yuan for the production of XX product," after segmentation, the resulting words include "our company," "October 5, 2024," "XX supplier," "1,000 pieces," "raw materials," "unit price," "50 yuan," "total amount," "50,000 yuan," "production," and "XX product." Then, named entity recognition is used to extract accounting items and amount entities. For example, "raw materials" is an inventory accounting item, and "50,000 yuan" is an amount entity.
[0079] Next, a domain knowledge graph is constructed to match the extracted entities with the standard account nodes in the graph based on semantic similarity. The domain knowledge graph contains various finance-related concepts, entities, and the relationships between them. In the knowledge graph of this group company, "inventory" is a standard accounting subject node with clear definitions and attributes. For the extracted "raw materials" entity, its semantic similarity with the "inventory" node is calculated to determine whether it matches. There are many ways to calculate semantic similarity, such as cosine similarity calculation based on word vectors. Assuming that the word vector of "raw materials" is V1 and the word vector of "inventory" is V2, their cosine similarity Sim can be expressed as: If the similarity exceeds a certain threshold, the entity is considered to match the "inventory" node.
[0080] Finally, a generative adversarial network is used to correct the matching results and generate a high-confidence alignment data table. The generative adversarial network adopts a conditional adversarial training framework, and the generator reconstructs the alignment data table through a variational autoencoder. The variational autoencoder can encode the input data into a latent vector, and then decode the latent vector into reconstructed data through the decoder. Let the input alignment data be x and the output latent vector of the encoder be z. The mapping relationship of the encoder can be expressed as z = Encoder (x), and the mapping relationship of the decoder is x ′ =Decoder(z), the goal of the generator is to minimize the reconstruction error, that is, Loss gen =|xx ′The discriminator uses a convolutional neural network to determine data authenticity, aiming to distinguish generated data from real data. A domain classifier is introduced during training to align the generated data with the target tax category distribution. The domain classifier determines the tax category to which the data belongs based on its characteristics. By adjusting the generator parameters, the generated data aligns with the real data distribution in terms of tax category. After training and error correction using the generative adversarial network, a highly confident aligned data table is generated, providing an accurate and unified data foundation for subsequent financial data processing.
[0081] Example 5:
[0082] This example details the specific implementation of temporal relationship encoding in transaction graph construction. In a company's daily operations, various transactions occur frequently, and there are complex temporal dependencies between them. Accurately constructing a transaction graph is crucial for analyzing a company's financial status and tax compliance. For example, an e-commerce company involves a large number of transactions, including sales, purchases, and payments. These transactions are temporally interconnected, forming a complex transaction chain.
[0083] The transaction graph is constructed through temporal relationship encoding, which defines the temporal dependencies between transaction nodes, including payment cycles, invoice issuance times, and payment delays. In the sales operations of an e-commerce company, a transaction node is generated when a consumer places an order for goods. Assuming the company's payment cycle is 7 days, the consumer must complete payment within 7 days of placing the order. The invoice issuance time depends on the company's invoicing policy and may be issued within 1-3 days of payment completion. Payment delay refers to the time between the consumer's successful payment and the company's actual receipt of the funds. For example, due to the payment channel's clearing process, the payment delay may be 1-2 days.
[0084] The dynamic correlation strength between nodes is modeled through a time-aware graph convolutional network, and a weighted transaction path sequence is generated. The time-aware graph convolutional network is enhanced by relative time encoding, mapping the time difference of transaction nodes into a sinusoidal position encoding vector and concatenating it with the node attributes. Suppose there are two transaction nodes i and j, and the time difference between them is τ ij , the calculation formula of the sinusoidal position encoding vector is: Where d is the dimension of the encoding vector, k = 0, 1, ..., d / 2-1. This sinusoidal position encoding vector is combined with the attribute X of node i. i Splicing to obtain enhanced node attributes Introduce a time decay factor α (0<α<1) in each layer of graph convolution operation to dynamically adjust the contribution weight of historical transactions. Let the neighbor node set of node i be N(i), and the updated feature H of node i in the graph convolution operation be i It can be expressed as: Where σ is the activation function, c i is the normalization constant, W is the weight matrix, is the feature of the previous iteration of node i.
[0085] This time-aware graph convolutional network processing can more accurately reflect the dynamic connections between transaction nodes. For example, when analyzing sales trends for an e-commerce company, if the delay in payment arrival at a large number of transaction nodes suddenly increases within a certain period of time, the time-aware graph convolutional network can capture this change. Due to the time decay factor, this recent abnormal change will be more prominent, helping companies to promptly identify potential financial risks. Ultimately, based on these calculations and processing, a weighted transaction path sequence is generated. This transaction path sequence clearly demonstrates the temporal sequence and correlation strength between different transactions, providing strong support for companies' financial analysis and tax reporting. For example, it can help companies accurately determine the timing of revenue recognition and ensure the accuracy of tax filings.
[0086] Example 6:
[0087] This example describes in detail the rules engine and GAN. In a company's tax processing, these two technologies play a key role in ensuring tax compliance and data accuracy. For example, a complex financial enterprise, involved in trading multiple financial products and complex tax processing, requires a precise rules engine and an efficient GAN to ensure smooth tax processing.
[0088] The rule engine is implemented using a logic programming language, converting tax regulations into predicate logic rules and compiling them into executable inference trees. Regarding the tax treatment of interest income for financial enterprises, tax regulations specify the tax calculation methods for different types of interest income. For example, for corporate bond interest income, it is stipulated that it must be taxed at a specific rate after deducting certain expenses. Converting this into a predicate logic rule can be expressed as follows: "If the income earned by a financial enterprise is corporate bond interest income I bond , the deductible expense is D, the applicable tax rate is r, then the tax payable T=(I bond-D)×r". These rules are compiled into executable inference trees, and fuzzy logic operators are embedded in the inference trees to handle the grayscale intervals in policy clauses. There are some ambiguous provisions in the tax policies of the financial industry. For example, there is a certain degree of flexibility in the tax treatment of certain financial innovation products. By embedding fuzzy logic operators, such as fuzzy "and" and "or" operators, in the inference tree, these situations can be handled more flexibly. At the same time, the real-time monitoring module captures rule conflict events and triggers the manual review process. When different tax rules conflict when applied to the tax treatment of the same financial business, for example, there are multiple and conflicting interpretations of the income tax rules for a certain financial derivative, the real-time monitoring module will promptly detect and notify the tax specialist for manual review to ensure the accuracy of the tax treatment.
[0089] The Generative Adversarial Network uses a conditional adversarial training framework and plays an important role in processing the complex financial data alignment of financial enterprises. The generator reconstructs the alignment data table through a variational autoencoder. For various types of transaction data of financial enterprises, such as stock trading data and bond trading data, the variational autoencoder encodes the input raw data into a latent vector and then reconstructs the data through the decoder. Let the input data be x and the latent vector output by the encoder be z, then the encoder is z = Encoder(x) and the decoder is x ′ =Decoder(z), the goal of the generator is to minimize the reconstruction error Loss gen =|xx ′ The discriminator determines the authenticity of the data using a convolutional neural network. Convolutional neural networks can extract data features, thereby distinguishing generated data from real data. A domain classifier is introduced during the training process to align the generated data with the distribution of the target tax category. Financial enterprises are subject to a variety of tax categories, such as value-added tax and corporate income tax, and different transaction data corresponds to different tax categories. The domain classifier determines the tax category to which the data belongs based on its characteristics and adjusts the generator parameters through a backpropagation algorithm to ensure that the generated data aligns with the distribution of real data in terms of tax category. Through the collaborative work of this adversarial generative network training and the rule engine, financial enterprises can more accurately process financial data, ensure the compliance and accuracy of tax declarations, and effectively reduce tax risks.
[0090] It should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that includes a list of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus.
[0091] While embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions, and variations may be made to these embodiments without departing from the principles and spirit of the invention, and that the scope of the invention is defined by the appended claims and their equivalents.
Claims
1. An automatic accounting and tax filing processing system based on AI big model, characterized by: include: Data collection module: used to obtain enterprise financial data, tax policy documents and transaction flow information through multi-source interfaces; Data preprocessing module: cleans and standardizes the data acquired by the data collection module based on heterogeneous data alignment technology to generate structured financial and tax data sets; Intelligent analysis module: uses a multimodal graph neural network to perform feature correlation analysis on the structured financial and taxation dataset to extract tax compliance features; Decision generation module: inputs the tax compliance features into the pre-trained tax strategy generation model and outputs the optimal accounting rules and tax filing plan; Execution feedback module: Build a dynamic compliance verification model based on the tax reporting solution, and realize the atomic execution and status feedback of financial and tax operations through a distributed transaction management framework.
2. The automatic accounting and tax filing processing system according to claim 1, characterized in that: The acquisition of enterprise financial data through a multi-source interface includes: Multi-source interfaces include electronic invoice interface, bank reconciliation interface, payroll management interface and tax declaration interface; Align the timestamps of electronic invoice data and bank transaction data to build a transaction graph; perform semantic analysis on salary data and tax policy data to generate a policy correlation matrix; A two-branch feature fusion model is constructed. The first branch uses a graph attention network to extract the topological relationship of the transaction graph, and the second branch uses a temporal convolutional network to extract the version evolution features of the policy association matrix. The topological relationship and version evolution characteristics are integrated through a cross-modal gating mechanism to generate a multi-dimensional fiscal and taxation feature vector; the multi-dimensional fiscal and taxation feature vector is dynamically updated based on a gated recurrent unit, and a comprehensive analysis result including transaction link integrity, policy matching and risk warning level is output.
3. The automatic accounting and tax filing processing system according to claim 1, characterized in that: The tax strategy generation model adopts a hierarchical reinforcement learning architecture to generate tax filing plans based on a multi-objective optimization mechanism; The hierarchical reinforcement learning architecture includes: Construct a fiscal and taxation-policy interaction graph, where nodes include tax type nodes, enterprise nodes, regulation nodes, and audit nodes. Node attributes include tax rates, deduction items, and declaration deadlines. A two-stage attention mechanism is used. In the first stage, the spatial graph convolution layer is used to calculate the dependency weights between tax type nodes and associated nodes. In the second stage, the temporal graph convolution layer is used to screen the compliance of historical declaration records. Node features are iteratively optimized based on a multi-head strategy network. Each strategy head integrates node attributes with real-time audit rules; ultimately, a dynamic tax filing strategy that meets the tax policies of multiple regions is output.
4. The automatic accounting and tax filing processing system according to claim 1, characterized in that: The dynamic compliance verification model integrates rule reasoning and anomaly detection strategies, including: The compliance verification problem is modeled as a multi-constraint satisfaction problem, where the decision variables include the declared amount, deduction ratio, and declaration time window; Initialize the rule engine and load the tax regulations knowledge base, using a dynamic priority mechanism to adjust rule weights based on the frequency of policy updates; In the inference phase, the constraint propagation algorithm is used to verify the logical consistency of the declared data; in the detection phase, the isolation forest algorithm is used to identify abnormal transaction patterns; The atomicity of the multi-threaded verification process is guaranteed through a distributed lock mechanism, and an audit log with a timestamp is generated.
5. The automatic accounting and tax filing processing system according to claim 1 is characterized in that: The heterogeneous data alignment technology adopts semantic embedding and entity linking methods, including: Perform paragraph-level word segmentation and named entity recognition on unstructured financial documents to extract accounting subjects and amount entities; Build a domain knowledge graph and match the extracted entities with the standard subject nodes in the graph based on semantic similarity; A generative adversarial network is used to correct the matching results and generate a high-confidence alignment data table.
6. The automatic accounting and tax filing processing system according to claim 2, characterized in that: The construction of the transaction graph is achieved through temporal relationship encoding, including: defining the temporal dependencies between transaction nodes, including payment cycle, invoice issuance time and arrival delay; modeling the dynamic correlation strength between nodes through time-aware graph convolutional network, and generating a weighted transaction path sequence.
7. The automatic accounting and tax filing processing system according to claim 4, characterized in that: The rule engine is implemented using a logic programming language, including: converting tax regulations into predicate logic rules and compiling them into executable inference trees; embedding fuzzy logic operators in the inference trees, processing grayscale intervals in policy terms, capturing rule conflict events through a real-time monitoring module, and triggering a manual review process.
8. The automatic accounting and tax filing processing system according to claim 5, characterized in that: The adversarial generative network adopts a conditional adversarial training framework, including: the generator reconstructs the aligned data table through a variational autoencoder, and the discriminator judges the authenticity of the data through a convolutional neural network; a domain classifier is introduced during the training process to force the generated data to be aligned with the distribution of the target tax account.
9. The automatic accounting and tax filing processing system according to claim 6, characterized in that: The time-aware graph convolutional network is enhanced through relative time encoding, including: mapping the time difference of transaction nodes into a sinusoidal position encoding vector and concatenating it with node attributes; introducing a time decay factor in each layer of graph convolution operation to dynamically adjust the contribution weight of historical transactions.
10. An automatic accounting and tax filing processing method based on an AI large model, applied to the automatic accounting and tax filing processing system according to any one of claims 1 to 9, characterized in that: The following steps are involved: Step 1: Obtain enterprise financial data, tax policy documents, and transaction flow information through the multi-source interface of the data acquisition module; Step 2: Using the data preprocessing module and heterogeneous data alignment technology, the acquired corporate financial data, tax policy documents, and transaction flow information are cleaned and standardized to generate a structured financial and tax data set. Step 3: Use the multimodal graph neural network in the intelligent analysis module to perform feature correlation analysis on the structured fiscal and taxation dataset to extract tax compliance features; Step 4: Input the extracted tax compliance features into the pre-trained tax strategy generation model in the decision generation module to output the optimal accounting rules and tax filing plan; Step 5: Build a dynamic compliance verification model based on the tax filing plan through the execution feedback module, and use the distributed transaction management framework to achieve atomic execution and status feedback of financial and tax operations.
Citation Information
Cited By
Tax processing method based on AI intelligent question and answer of large language model
CN121681726A
Taxation system automatic interaction method and system based on interface integration
CN122023046A