Enterprise tax return data intelligent comparison and automatic verification optimization method and system

By extracting multi-level features and constructing cross-dimensional related feature sets from corporate tax declaration data, and combining a hybrid architecture of unsupervised clustering and supervised classification algorithms with a causal inference mechanism, the shortcomings of existing technologies in identifying anomalies in tax declaration data are solved. This enables efficient and accurate anomaly tracing and automatic correction, thereby improving the level of automation in tax management.

CN121481753BActive Publication Date: 2026-05-19YICHENG (BEIJING) INFORMATION SERVICE CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
YICHENG (BEIJING) INFORMATION SERVICE CO LTD
Filing Date
2025-11-14
Publication Date
2026-05-19

AI Technical Summary

Technical Problem

Existing methods for identifying anomalies in tax return data fail to fully explore the cross-dimensional relationships between different tax types and different return items, making it difficult to accurately pinpoint the true source of anomalies. Furthermore, they lack the ability to conduct in-depth analysis of the mechanisms by which anomalies are formed, leading to missed detections or misjudgments and increasing the workload of tax officials.

Method used

By extracting multi-level features from corporate tax declaration data, a cross-dimensional associated feature set is constructed. A hybrid architecture combining unsupervised clustering and supervised classification algorithms is used to identify anomaly type identifiers. Furthermore, a directed causal dependency graph is constructed through a causal inference mechanism to trace the source path of anomalies and automatically adjust the tax declaration data to achieve correction.

Benefits of technology

It significantly improves the accuracy and comprehensiveness of anomaly identification, reduces the false negative and false positive rates, provides interpretable analytical basis, enhances the pertinence and efficiency of anomaly handling, realizes a closed-loop processing flow from anomaly identification to correction, reduces labor costs, and shortens the tax data quality management cycle.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121481753B_ABST
    Figure CN121481753B_ABST
Patent Text Reader

Abstract

The application provides an enterprise tax declaration data intelligent comparison and automatic checking optimization method and system, relates to the tax data processing technical field, and comprises the following steps: multi-level feature extraction is carried out on enterprise tax declaration data; a cross-dimension correlation feature set is constructed based on industry business rules; constraint conditions and threshold boundaries are extracted from historical compliance and non-compliance samples through an association rule mining algorithm to construct a compliance constraint rule set; a hybrid architecture of unsupervised clustering and supervised classification is adopted to perform abnormal mode matching and deviation calculation and identify abnormal types; a directed causal dependence graph is constructed and an abnormal traceability path is traced back through a causal inference mechanism to analyze feature contribution weights; and the data to be corrected is located according to the abnormal traceability path and the compliance constraint rule and is corrected.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of tax data processing technology, and in particular to a method and system for intelligent comparison and automatic verification optimization of corporate tax declaration data. Background Technology

[0002] With the continuous improvement of my country's tax collection and administration system and the deepening of digital transformation, corporate tax declaration data is characterized by its large scale, multiple dimensions, and complex relationships. Tax authorities need to process massive amounts of corporate tax declaration data annually, involving multiple taxes such as value-added tax, corporate income tax, and stamp duty, with complex interrelationships between these taxes. In actual tax declaration processes, abnormal tax declaration data frequently occurs due to reasons such as inaccurate understanding of tax policies by corporate financial personnel, errors in the declaration system, data entry errors, or improper tax processing caused by changes in corporate business. These abnormal data not only affect the accuracy and efficiency of tax collection and administration but also bring tax risks and economic losses to enterprises.

[0003] Existing methods for identifying tax anomalies primarily rely on pre-defined fixed rules or simple statistical analysis models. These methods often focus only on single-dimensional or localized features when extracting anomalies, failing to fully explore the cross-dimensional relationships between different tax types and declaration items within the tax return data. This results in insufficient ability to identify complex anomaly patterns, easily leading to missed detections or misjudgments, especially when dealing with complex anomalies involving multiple related fields, making it difficult to accurately pinpoint the true source of the anomaly. After identifying anomalies, current technologies typically only provide anomaly alerts or warnings, lacking the ability to deeply analyze the anomaly formation mechanism. They cannot accurately determine the causal relationships and degree of influence between multiple anomaly features, nor can they effectively distinguish which features are the root cause of the anomaly and which are secondary anomalies derived from the root cause. This deficiency forces tax officials to invest significant effort in manual source tracing analysis when processing anomaly data, making it difficult to quickly locate the core of the problem. Summary of the Invention

[0004] This invention provides a method and system for intelligent comparison and automatic verification optimization of enterprise tax declaration data, which can solve the problems in the prior art.

[0005] A first aspect of this invention provides a method for intelligent comparison and automatic verification optimization of enterprise tax declaration data, comprising:

[0006] Multi-level feature extraction is performed on the tax declaration data of enterprises, and a cross-dimensional related feature set is constructed based on the business rule constraints of the industry to which the enterprise belongs;

[0007] Based on compliant and non-compliant samples in historical tax declaration data, the constraints and threshold boundaries between tax declaration data are extracted through association rule mining algorithms to construct a set of compliance constraint rules; based on a hybrid architecture of unsupervised clustering algorithm and supervised classification algorithm, abnormal pattern matching and deviation calculation are performed on the cross-dimensional association feature set to identify abnormal type identifiers.

[0008] Locate the feature subset corresponding to the anomaly type identifier in the cross-dimensional associated feature set, identify the contribution weight of each feature in the feature subset to the formation of the anomaly through the causal inference mechanism, construct a directed causal dependency graph between features based on the contribution weight, trace the path along the direction of decreasing contribution weight in the directed causal dependency graph, and extract the maximum weight path from the normal state node to the abnormal state node as the anomaly source tracing path.

[0009] Based on the anomaly tracing path and the set of compliance constraint rules, the tax declaration data that needs to be corrected and the corresponding correction direction are located. According to the correction direction, the tax declaration data that needs to be corrected is adjusted numerically or its status is changed to obtain the corrected tax declaration data.

[0010] Multi-level feature extraction is performed on the enterprise's tax declaration data, and a cross-dimensional correlation feature set is constructed based on the business rule constraints of the enterprise's industry, including:

[0011] The tax declaration data of enterprises is analyzed in a multi-dimensional and layered manner. The periodic change characteristics of financial indicators are extracted from the time dimension, the revenue and cost structure characteristics corresponding to different business types are extracted from the business dimension, and the correlation relationship characteristics between tax declaration items are extracted from the account dimension, forming an initial feature set.

[0012] The standard financial indicator distribution range and revenue-cost structure ratio range of the industry to which the enterprise belongs are used as the industry standard reference framework. The periodic change characteristics and the revenue-cost structure characteristics are mapped to the numerical space corresponding to the industry standard reference framework. By calculating the distance between the feature value of the tax declaration data and the boundary of the industry standard reference framework, the industry adaptability deviation characteristics are generated.

[0013] Based on the correlation relationship features, the association and dependency between each tax data item are determined. A row index is created based on the association and dependency. The numerical ratio between any two tax data items is used as a matrix element to construct a mutual dependency constraint matrix. The mutual dependency constraint matrix is ​​used to perform correlation fusion on the features of each dimension in the initial feature set. The contribution of each dimension feature in the fusion is adjusted by the industry adaptability deviation feature to generate the cross-dimensional correlation feature set.

[0014] Based on compliant and non-compliant samples from historical tax return data, an association rule mining algorithm is used to extract constraints and threshold boundaries between tax return data, constructing a set of compliance constraint rules including:

[0015] Based on the historical tax audit results, the historical tax declaration data is divided into a compliant sample set and a non-compliant sample set. The numerical distribution characteristics of the compliant sample set and the non-compliant sample set are extracted, the distribution boundary differences of the numerical distribution characteristics on each tax data item are calculated, and the numerical boundary threshold for distinguishing between compliant and non-compliant samples is determined.

[0016] In the set of compliant samples, identify combinations of tax data items whose co-occurrence frequency exceeds a preset frequency threshold, perform statistical analysis on the numerical relationships between each tax data item, calculate the numerical ratio relationship and state logic dependency relationship between each data item in the combination of tax data items, and generate numerical ratio constraints and logical consistency conditions.

[0017] The numerical boundary threshold is used as a single constraint condition, and the numerical ratio constraint and the logical consistency condition are used as multiple associated constraints conditions. The single constraint condition and the multiple associated constraints are classified and merged according to the type of tax data item. A mapping relationship between trigger conditions and constraint boundaries is established for the constraints under each category to generate the compliance constraint rule set.

[0018] Based on a hybrid architecture combining unsupervised clustering and supervised classification algorithms, anomaly pattern matching and deviation calculation are performed on the cross-dimensional associated feature set to identify anomaly type identifiers, including:

[0019] A feature representation space is constructed using each feature dimension in the cross-dimensional associated feature set as the coordinate axis. Cluster analysis is performed using the cross-dimensional associated feature set as input data. Based on the distance measurement and density distribution characteristics between sample points in the feature representation space, the cross-dimensional associated feature set is divided into multiple clusters, and the boundary range and outlier sample judgment threshold of each cluster are determined.

[0020] Using historical tax declaration data with labeled anomaly types as training samples, an anomaly type recognizer is constructed by learning the feature patterns corresponding to different anomaly types through a gradient boosting tree-based classification algorithm.

[0021] Calculate the Euclidean distance from each sample to be identified to the central feature vector of each cluster. If the Euclidean distance exceeds the outlier determination threshold, the sample to be identified is determined to be an anomalous sample. Calculate the deviation of the anomalous sample from the boundary range of the nearest cluster. Input the feature vector corresponding to the anomalous sample into the anomaly type identifier to obtain the anomaly type probability distribution. Select the anomaly type with the highest probability value as the anomaly type identifier. Use the deviation as the confidence adjustment parameter for the anomaly type identifier.

[0022] Locating the feature subset corresponding to the anomaly type identifier in the cross-dimensional correlation feature set, identifying the contribution weight of each feature in the feature subset to the anomaly formation through a causal inference mechanism, and constructing a directed causal dependency graph between features based on the contribution weights includes:

[0023] In the cross-dimensional correlation feature set, identify samples labeled with the anomaly type identifier, extract the corresponding feature vectors, calculate the difference ratio between the variance of each feature dimension in the anomaly sample and the variance in the normal sample, and select the feature dimension set that has the ability to distinguish the anomaly type identifier as the feature subset based on the difference ratio.

[0024] For any two feature dimensions in the feature subset, the temporal order of the two feature dimensions in the data generation process is determined by temporal sequence analysis. The conditional independence test is used to determine whether a change in the value of one feature dimension causes a change in the value of the other feature dimension. Based on the temporal order and the conditional independence test results, it is determined whether there is a causal relationship between the two feature dimensions. If a causal relationship exists, the feature dimension that appears earlier in time is marked as the causal feature, and the feature dimension that appears later in time is marked as the effect feature. An integral operation is performed within the value range of the causal feature to obtain the contribution weight of the causal feature to the anomaly type identification.

[0025] Using each feature dimension in the feature subset as a node, and taking the nodes corresponding to the causal feature and the effect feature as the starting and ending points of the directed edges respectively, and using the contribution weight as the edge weight, a directed causal dependency graph between features is constructed.

[0026] In the directed causal dependency graph, path tracing is performed along the direction of decreasing contribution weight, and the path with the maximum weight from the normal state node to the abnormal state node is extracted as the anomaly tracing path, including:

[0027] In the directed causal dependency graph, nodes whose corresponding feature dimension values ​​are within the compliant range are marked as normal state nodes, and the remaining nodes are marked as abnormal state nodes; the directed causal dependency graph is traversed, and node pairs whose starting node is a normal state node and whose ending node is an abnormal state node are identified as target node pairs.

[0028] Based on the target node pair, a reverse traversal is performed from the abnormal state node, and the previous nodes in the path are visited sequentially in the opposite direction of the directed edges. During the reverse traversal, the edge weight corresponding to each directed edge is extracted, and it is determined whether the sequence of edge weights extracted sequentially along the reverse traversal direction shows a decreasing trend.

[0029] In a connected path where the edge weight sequence shows a decreasing trend, the edge weights of each directed edge are summed to obtain the path weight of the connected path.

[0030] Among the connected paths that start from the same normal state node and reach the same abnormal state node and whose edge weight sequence shows a decreasing trend, the connected path with the largest path weight is selected as the maximum weight path. All nodes in the maximum weight path are arranged in chronological order to generate the abnormal source tracing path.

[0031] Based on the anomaly tracing path and the set of compliance constraint rules, the tax declaration data requiring correction and the corresponding correction direction are located. According to the correction direction, the tax declaration data requiring correction is adjusted numerically or its status is changed, including:

[0032] Obtain all abnormal state nodes and corresponding feature dimension identifiers in the abnormal source tracing path, locate the related tax declaration data items in the tax declaration data according to the feature dimension identifiers, and mark them as the tax declaration data that needs to be corrected;

[0033] Extract the identifiers of tax declaration data items involved in the set of compliance constraint rules, determine whether the identifiers of the tax declaration data items match the data item names of the tax declaration data to be corrected, extract the compliant value ranges defined in the successfully matched constraint rules, calculate the deviation between the current value of the tax declaration data to be corrected and the boundary value of the compliant value range, and determine the correction direction based on the sign of the deviation.

[0034] Calculate the numerical adjustment range based on the absolute value of the deviation, add or subtract the numerical adjustment range to the current value of the tax declaration data to be corrected according to the correction direction, and generate the corrected tax declaration data value; record the status compliance transition path of the tax declaration data under different correction directions, and perform a status change operation on the tax declaration data to be corrected according to the status compliance transition path.

[0035] A second aspect of this invention provides an intelligent comparison and automatic verification optimization system for enterprise tax declaration data, comprising:

[0036] The first unit is used to extract multi-level features from the tax declaration data of enterprises and to construct a cross-dimensional related feature set based on the business rule constraints of the industry to which the enterprise belongs.

[0037] The second unit is used to extract constraints and threshold boundaries between tax declaration data based on compliant and non-compliant samples in historical tax declaration data through association rule mining algorithms, and construct a set of compliance constraint rules; based on a hybrid architecture of unsupervised clustering algorithm and supervised classification algorithm, it performs abnormal pattern matching and deviation calculation on the cross-dimensional association feature set to identify abnormal type identifiers;

[0038] The third unit is used to locate the feature subset corresponding to the anomaly type identifier in the cross-dimensional associated feature set, identify the contribution weight of each feature in the feature subset to the formation of the anomaly through a causal inference mechanism, construct a directed causal dependency graph between features based on the contribution weight, trace the path along the direction of decreasing contribution weight in the directed causal dependency graph, and extract the maximum weight path from the normal state node to the abnormal state node as the anomaly tracing path.

[0039] The fourth unit is used to locate the tax declaration data that needs to be corrected and the corresponding correction direction based on the abnormality tracing path and the set of compliance constraint rules, and to adjust the value or change the status of the tax declaration data that needs to be corrected according to the correction direction to obtain the corrected tax declaration data.

[0040] A third aspect of the present invention,

[0041] An electronic device is provided, comprising:

[0042] processor;

[0043] Memory used to store processor-executable instructions;

[0044] The processor is configured to invoke instructions stored in the memory to execute the aforementioned method.

[0045] Fourth aspect of the embodiments of the present invention,

[0046] A computer-readable storage medium is provided, having stored thereon computer program instructions that, when executed by a processor, implement the aforementioned method.

[0047] The beneficial effects of this application are as follows:

[0048] This invention extracts multi-level features from corporate tax declaration data and constructs a cross-dimensional correlation feature set by combining industry business rules. It can comprehensively capture the deep-level correlations in tax data. Compared with traditional identification methods that are based on only a single dimension or surface features, it significantly improves the ability to perceive complex abnormal patterns, making anomaly identification more accurate and comprehensive, and effectively reducing the false negative and false positive rates.

[0049] This invention employs a hybrid architecture combining unsupervised clustering and supervised classification algorithms for anomaly pattern matching. By constructing a directed causal dependency graph between features through a causal inference mechanism, it can not only identify anomalous data but also trace the root cause of the anomalies and reveal the internal logical chain of their formation. Compared with traditional methods that can only identify anomalies but cannot explain their causes, this invention provides tax management personnel with interpretable analytical basis, improving the pertinence and efficiency of anomaly handling.

[0050] This invention automatically determines the correction direction and executes data correction based on the anomaly tracing path and compliance constraint rule set, realizing a closed-loop processing flow from anomaly identification to anomaly correction. Compared with the traditional method that requires manual item-by-item verification and manual correction, it significantly reduces labor costs, shortens the cycle of tax data quality management, and ensures the compliance of correction results through rule constraints, thereby improving the automation level and data quality of enterprise tax management. Attached Figure Description

[0051] Figure 1 This is a flowchart illustrating the intelligent comparison and automatic verification optimization method for enterprise tax declaration data according to an embodiment of the present invention.

[0052] Figure 2 A schematic diagram illustrating the process of constructing a cross-dimensional associated feature set. Detailed Implementation

[0053] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0054] The technical solution of the present invention will be described in detail below with reference to specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments.

[0055] Figure 1This is a flowchart illustrating the intelligent comparison and automatic verification optimization method for enterprise tax declaration data according to an embodiment of the present invention. Figure 1 As shown, the method includes:

[0056] Multi-level feature extraction is performed on the tax declaration data of enterprises, and a cross-dimensional related feature set is constructed based on the business rule constraints of the industry to which the enterprise belongs;

[0057] Based on compliant and non-compliant samples in historical tax declaration data, the constraints and threshold boundaries between tax declaration data are extracted through association rule mining algorithms to construct a set of compliance constraint rules; based on a hybrid architecture of unsupervised clustering algorithm and supervised classification algorithm, abnormal pattern matching and deviation calculation are performed on the cross-dimensional association feature set to identify abnormal type identifiers.

[0058] Locate the feature subset corresponding to the anomaly type identifier in the cross-dimensional associated feature set, identify the contribution weight of each feature in the feature subset to the formation of the anomaly through the causal inference mechanism, construct a directed causal dependency graph between features based on the contribution weight, trace the path along the direction of decreasing contribution weight in the directed causal dependency graph, and extract the maximum weight path from the normal state node to the abnormal state node as the anomaly source tracing path.

[0059] Based on the anomaly tracing path and the set of compliance constraint rules, the tax declaration data that needs to be corrected and the corresponding correction direction are located. According to the correction direction, the tax declaration data that needs to be corrected is adjusted numerically or its status is changed to obtain the corrected tax declaration data.

[0060] In one optional implementation, multi-level feature extraction is performed on the enterprise's tax declaration data, and a cross-dimensional correlation feature set is constructed based on the business rule constraints of the enterprise's industry, including:

[0061] The tax declaration data of enterprises is analyzed in a multi-dimensional and layered manner. The periodic change characteristics of financial indicators are extracted from the time dimension, the revenue and cost structure characteristics corresponding to different business types are extracted from the business dimension, and the correlation relationship characteristics between tax declaration items are extracted from the account dimension, forming an initial feature set.

[0062] The standard financial indicator distribution range and revenue-cost structure ratio range of the industry to which the enterprise belongs are used as the industry standard reference framework. The periodic change characteristics and the revenue-cost structure characteristics are mapped to the numerical space corresponding to the industry standard reference framework. By calculating the distance between the feature value of the tax declaration data and the boundary of the industry standard reference framework, the industry adaptability deviation characteristics are generated.

[0063] Based on the correlation relationship features, the association and dependency between each tax data item are determined. A row index is created based on the association and dependency. The numerical ratio between any two tax data items is used as a matrix element to construct a mutual dependency constraint matrix. The mutual dependency constraint matrix is ​​used to perform correlation fusion on the features of each dimension in the initial feature set. The contribution of each dimension feature in the fusion is adjusted by the industry adaptability deviation feature to generate the cross-dimensional correlation feature set.

[0064] like Figure 2 As shown, the method includes:

[0065] Retrieve the target company's tax filing records for the past three years from the database. These records include VAT returns, corporate income tax returns, surcharge tax returns, and financial statement data. Taking a manufacturing company as an example, its monthly filing data from 2021 to 2023 includes 56 basic data items such as sales revenue, output VAT, input VAT, tax payable, operating costs, and period expenses. Store this data in three independent data tables according to time, business, and account dimensions.

[0066] The time-dimensional feature extraction uses months as the basic unit, calculating the fluctuation of each financial indicator over a continuous period. For the manufacturing company's sales revenue data, the monthly revenues from January to December 2023 were recorded as 8.2 million yuan, 7.5 million yuan, 9.1 million yuan, 8.8 million yuan, 9.5 million yuan, 9.2 million yuan, 10.3 million yuan, 9.9 million yuan, 10.5 million yuan, 11 million yuan, 10.8 million yuan, and 11.5 million yuan, respectively. The magnitude of revenue changes in adjacent months was calculated, yielding 11 rate-of-change values. Significant growth inflection points were identified in the 7th and 10th months. Furthermore, comparing the data with the same period last year, a year-on-year revenue growth trend was identified. These cyclical fluctuations, growth inflection points, and year-on-year change rates were used as the time-dimensional feature vector.

[0067] Business dimension feature extraction categorizes revenue sources based on the different business types declared by the enterprise. This manufacturing enterprise's business types include product sales, processing services, and technical consulting. Extracting full-year data for 2023, product sales revenue accounted for 73% of total revenue, processing services revenue accounted for 22%, and technical consulting revenue accounted for 5%. Further analysis of the cost structure corresponding to each business type revealed a cost ratio of 65% for product sales, 58% for processing services, and 42% for technical consulting. The revenue share, cost ratio, and gross profit margin for each business type are combined to form a business structure feature vector, resulting in nine elements of business dimension feature data.

[0068] Subject-based feature extraction identifies the reconciliation relationships between items in the tax return. Taking VAT declaration as an example, the ratio of output VAT to sales revenue is extracted. For this manufacturing company, a 13% tax rate applies. In December 2023, sales revenue of 11.5 million yuan corresponded to output VAT of 1.495 million yuan, a ratio of 0.13. Simultaneously, the relationship between input VAT and procurement costs is extracted. Procurement costs of 7.2 million yuan in the same month corresponded to input VAT of 936,000 yuan, also a ratio of 0.13. The relationship between tax payable and output VAT / input VAT is identified. The tax payable of 559,000 yuan in the current month equals output VAT of 1.495 million yuan minus input VAT of 936,000 yuan. These numerical ratios, differences, and cumulative relationships are organized into a reconciliation relationship feature set.

[0069] Standard financial indicator ranges for the manufacturing industry were retrieved from the industry database. The standard gross profit margin range for the manufacturing industry is 18% to 35%, the debt-to-equity ratio range is 40% to 70%, the current ratio range is 1.2 to 2.5, and the inventory turnover days range is 45 to 90 days. The target company's actual indicators were compared with these standard ranges. In 2023, the company's overall gross profit margin was 28%, its debt-to-equity ratio was 55%, its current ratio was 1.8, and its inventory turnover days were 62 days. The distance of each indicator value from the range boundaries was calculated. A gross profit margin of 28% is 10 percentage points away from the lower boundary (18%) and 7 percentage points away from the upper boundary (35%). This indicates that the indicator is biased towards the upper limit of the range, thus obtaining the industry standard reference framework.

[0070] The company's revenue and cost structure was matched with industry standards. The standard range for direct material costs in the manufacturing industry is 45% to 60%, direct labor costs 15% to 25%, and manufacturing overhead 10% to 20%. This company's actual direct material cost ratio is 52%, direct labor cost is 19%, and manufacturing overhead is 14%, all three percentages falling within the standard range. The deviation of each cost item's percentage from the central value of the range was calculated, and the deviation was used as an industry fit characteristic. For indicators exceeding the standard range, the extent of the deviation was recorded and marked as an abnormal deviation.

[0071] Based on the interrelationships, a relational structure is created between data items. Six core data items—sales revenue, output tax, input tax, tax payable, operating costs, and gross profit—are selected as row and column indices of the matrix. The numerical ratios between any two data items are calculated: the ratio of output tax to sales revenue is 0.13, the ratio of input tax to operating costs is 0.13, and the ratio of gross profit to sales revenue is 0.28. These ratios are then filled into the corresponding matrix positions, forming a six-row, six-column interdependent constraint matrix. A diagonal element of 1 indicates a self-relationship.

[0072] This study integrates time-dimension, business-dimension, and account-dimension features using a dependency constraint matrix. The revenue growth rate feature extracted from the time dimension is correlated with the revenue structure feature from the business dimension. By analyzing the ratio of sales revenue to revenue for each business type in the matrix, the single growth rate feature is expanded into a growth rate feature vector segmented by business type. Similarly, the tax rate feature from the account dimension is correlated with the cost rate feature from the business dimension. The comprehensive tax burden feature is generated by analyzing the ratio of tax amount to cost in the matrix. The contribution of each dimension feature is adjusted based on industry adaptability deviation features. Features with larger deviations are assigned higher weight coefficients. For example, the inventory turnover feature of this manufacturing company has a relatively small deviation, so its weight coefficient is set to 0.8, while the accounts receivable turnover feature has a larger deviation, so its weight coefficient is set to 1.3. After fusion processing, a cross-dimensional correlation feature set containing 124 elements is generated. This feature set simultaneously reflects the temporal patterns, business structure, account correlations, and industry adaptability of the company's tax data.

[0073] In one optional implementation, based on compliant and non-compliant samples in historical tax return data, a set of compliance constraint rules is constructed by extracting constraints and threshold boundaries between tax return data using an association rule mining algorithm, including:

[0074] Based on the historical tax audit results, the historical tax declaration data is divided into a compliant sample set and a non-compliant sample set. The numerical distribution characteristics of the compliant sample set and the non-compliant sample set are extracted, the distribution boundary differences of the numerical distribution characteristics on each tax data item are calculated, and the numerical boundary threshold for distinguishing between compliant and non-compliant samples is determined.

[0075] In the set of compliant samples, identify combinations of tax data items whose co-occurrence frequency exceeds a preset frequency threshold, perform statistical analysis on the numerical relationships between each tax data item, calculate the numerical ratio relationship and state logic dependency relationship between each data item in the combination of tax data items, and generate numerical ratio constraints and logical consistency conditions.

[0076] The numerical boundary threshold is used as a single constraint condition, and the numerical ratio constraint and the logical consistency condition are used as multiple associated constraints conditions. The single constraint condition and the multiple associated constraints are classified and merged according to the type of tax data item. A mapping relationship between trigger conditions and constraint boundaries is established for the constraints under each category to generate the compliance constraint rule set.

[0077] Retrieve all tax return records from the past three years from the database. Each record includes data items such as company identifier, filing period, operating revenue, operating costs, pre-tax profit, tax payable, input tax, output tax, total assets, and total liabilities. Based on the compliance flag field in the historical audit results, data is filtered, with records marked "Audit Passed" placed in the compliant sample set, and records marked "Audit Failed" or "Tax Payment Required" placed in the non-compliant sample set. Assuming 8,000 compliant samples and 2,000 non-compliant samples are obtained, an initial dataset is formed.

[0078] The data processing module performs statistical analysis on the compliant and non-compliant sample sets separately. For each numerical tax data item, it calculates its distribution characteristics in both sets. Taking the tax burden rate as an example, the minimum tax burden rate in the compliant sample set is calculated to be 0.8%, the maximum to be 12.5%, the average to be 4.3%, and the median to be 3.9%. In the non-compliant sample set, the minimum tax burden rate is calculated to be 0.1%, the maximum to be 23.7%, the average to be 8.6%, and the median to be 7.2%. A large number of outliers (below 1% or above 15%) are identified in the non-compliant samples. By comparing the distribution boundaries of the tax burden rate data item in the two sets, the quantile difference method is used to determine the cutoff thresholds. The 5th percentile of the compliant samples (0.9%) is used as the lower threshold, and the 95th percentile (11.2%) is used as the upper threshold, forming the compliant range constraint for the tax burden rate.

[0079] The same distribution feature extraction and threshold calculation process is applied to other numerical data items such as operating cost ratio, profit margin, and debt-to-equity ratio. The compliant sample distribution range for operating cost ratio is 35% to 85%, while non-compliant samples show numerous extreme values ​​exceeding 95% or falling below 20%, thus determining the compliant threshold range for operating cost ratio to be 40% to 90%. Profit margin in compliant samples is concentrated in the 5% to 30% range, while non-compliant samples exhibit negative values ​​or exceed 50%, setting the normal range for profit margin to be 3% to 35%. Individual constraints are established for each numerical data item using this method, including upper and lower limits of the value range, forming a first-type constraint rule set.

[0080] Frequent itemset mining was performed on the compliant sample set to identify frequently co-occurring combinations of tax data items. A support threshold of 60% was set, meaning that a data item combination needed to co-occur in at least 4800 compliant samples to be considered a frequent combination. The mining results showed that operating revenue and operating costs co-occurred 7600 times, operating revenue and tax payable co-occurred 7900 times, and the combination of operating costs, operating revenue, and profit co-occurred 7200 times. These high-frequency co-occurring data item combinations were extracted as objects for association analysis.

[0081] For the data item combination of operating revenue and operating cost, the numerical ratio between the two is calculated. By iterating through each record in the compliant sample set, the operating cost is divided by the operating revenue to obtain the cost-to-revenue ratio, and the cost-to-revenue ratio of all samples is statistically analyzed to form a distribution sequence. The analysis results show that 92% of the compliant samples have a cost-to-revenue ratio between 0.4 and 0.9, with a median of 0.65. Using 0.35 as the lower threshold and 0.95 as the upper threshold, the constraint rule "the ratio of operating cost to operating revenue should be between 0.35 and 0.95" is generated. This rule reflects the normal numerical ratio constraint relationship between operating cost and operating revenue.

[0082] There is also a numerical proportional relationship between operating revenue and tax payable. The actual tax burden rate is calculated by dividing the tax payable by the operating revenue. In the compliant samples, this ratio mainly falls within the range of 0.01 to 0.12, with over 85% of the samples falling within this range. Combining tax type information and industry classification attributes, the constraints are further refined: the actual tax burden rate range is set at 0.02 to 0.08 for manufacturing enterprises and 0.03 to 0.1 for service enterprises. This hierarchical constraint based on business characteristics improves the accuracy of the rules.

[0083] The study identifies a logical calculation relationship between output tax, input tax, and VAT payable. In compliant samples, it verifies whether each record satisfies the equation that output tax minus input tax equals VAT payable. Statistical results show that 98.5% of compliant samples satisfy this equation, with an error range within ±2%. A logical consistency condition is generated: "VAT payable should equal the difference between output tax and input tax, with an allowable error of no more than 2%." For non-compliant samples, the satisfaction rate of this logical relationship is only 43%, indicating a significant data inconsistency issue.

[0084] There are also logical constraints among the data items in the balance sheet: total assets should equal the sum of total liabilities and owners' equity. In the compliance samples tested, this balance was satisfied in 99.2% of the tests. This constraint was extracted, and a tolerance range was set, generating the rule: "The deviation between total assets and the sum of total liabilities and owners' equity should be less than 1% of total assets." In the income statement, there is a recursive calculation relationship between operating profit, total profit, and net profit. By verifying whether operating profit plus non-operating income minus non-operating expenses equals total profit, and whether total profit minus income tax expense equals net profit, multi-level logical consistency constraints are formed.

[0085] The generated individual constraints and multiple related constraints are categorized and organized according to the type of tax data item. Constraints for revenue data items include the range of operating revenue values, the ratio of operating revenue to operating costs, and the ratio of operating revenue to tax payable. Constraints for cost and expense data items include the range of operating costs values, the threshold boundary of the operating cost ratio, and the matching relationship between costs and revenue. Constraints for tax amount data items include the range of various tax amounts, the logical calculation relationship between tax amounts, and the reasonable range of the tax burden rate. Constraints for asset and liability data items include the asset-liability ratio threshold, the balance sheet balance, and the range constraints of the current ratio and quick ratio.

[0086] When the declared revenue exceeds 10 million, the tax burden rate constraint for high-income enterprises is triggered, requiring a tax burden rate of no less than 2.5%. When the enterprise type is identified as a small-scale taxpayer, the simplified tax collection rate constraint is triggered, limiting the VAT collection rate to 3%. When the declaration period is a quarterly declaration, the quarterly data item constraint rule is triggered, requiring quarterly revenue to not exceed 30% of the annual budget. Different triggering conditions activate different constraint boundary parameters, forming a dynamically adaptable rule execution mechanism, ultimately constructing a complete set of compliance constraint rules that includes single-item threshold constraints, numerical ratio constraints, and logical consistency constraints.

[0087] In one optional implementation, based on a hybrid architecture of unsupervised clustering and supervised classification algorithms, anomaly pattern matching and deviation calculation are performed on the cross-dimensional associated feature set, and anomaly type identification includes:

[0088] A feature representation space is constructed using each feature dimension in the cross-dimensional associated feature set as the coordinate axis. Cluster analysis is performed using the cross-dimensional associated feature set as input data. Based on the distance measurement and density distribution characteristics between sample points in the feature representation space, the cross-dimensional associated feature set is divided into multiple clusters, and the boundary range and outlier sample judgment threshold of each cluster are determined.

[0089] Using historical tax declaration data with labeled anomaly types as training samples, an anomaly type recognizer is constructed by learning the feature patterns corresponding to different anomaly types through a gradient boosting tree-based classification algorithm.

[0090] Calculate the Euclidean distance from each sample to be identified to the central feature vector of each cluster. If the Euclidean distance exceeds the outlier determination threshold, the sample to be identified is determined to be an anomalous sample. Calculate the deviation of the anomalous sample from the boundary range of the nearest cluster. Input the feature vector corresponding to the anomalous sample into the anomaly type identifier to obtain the anomaly type probability distribution. Select the anomaly type with the highest probability value as the anomaly type identifier. Use the deviation as the confidence adjustment parameter for the anomaly type identifier.

[0091] Each feature dimension in the cross-dimensional correlation feature set is used as an independent coordinate axis for spatial mapping. Assuming the cross-dimensional correlation feature set includes six dimensions: annual turnover, input tax, output tax, filing frequency, number of invoices issued, and cash flow cycle, a six-dimensional feature representation space is constructed. Each tax declaration record to be analyzed is represented as a sample point in this space, with its coordinates composed of the feature values ​​of each dimension. For example, the coordinates of a sample point for a certain enterprise are: turnover of 5 million yuan, input tax of 850,000 yuan, output tax of 650,000 yuan, filing frequency of 12 times per year, number of invoices issued of 240, and cash flow cycle of 45 days.

[0092] A density-based clustering method is used to group sample points in the feature representation space. The distance metric between any two sample points is calculated by taking the square root of the sum of the squares of the differences between the two sample points across each feature dimension. The neighborhood radius parameter is set to 5% of the overall feature space scale, and the minimum number of points to include is set to 2% of the total number of samples. Each sample point in the feature space is scanned, and the number of sample points within its neighborhood radius is counted. When the number of sample points in a sample point's neighborhood exceeds the minimum number of points to include, that sample point is marked as a core point. Starting from each core point, all reachable sample points within its neighborhood are grouped into the same cluster. After this traversal process, all sample points are divided into several clusters. For example, 10,000 tax declaration records are divided into eight different clusters, such as a normal business cluster, a high-frequency transaction cluster, and a periodic fluctuation cluster.

[0093] Calculate the central feature vector of all sample points within a cluster. Each dimension of this central feature vector is the arithmetic mean of the corresponding eigenvalues ​​of all sample points within the cluster. Calculate the distance metric from each sample point within the cluster to the central feature vector and statistically analyze the distribution characteristics of these distance metrics. Sort the distance metrics from smallest to largest, and select the distance metric located at the 95th percentile as the boundary radius of the cluster. This boundary radius defines the boundary range of the cluster. Multiply the boundary radius by an expansion factor of 1.5 to obtain the outlier threshold. For example, a cluster of 4500 sample points with a normal operating business might have the following central feature vector: turnover of 3.8 million yuan, input tax of 640,000 yuan, output tax of 490,000 yuan, tax declaration frequency of 12 times per year, number of invoices issued of 180, and cash flow cycle of 38 days. Its boundary radius is 85 units, and the outlier threshold is 127.5 units.

[0094] We collected sample records from historical tax returns that had been manually reviewed and labeled with anomalies. These anomaly types included specific categories such as false invoices, tax evasion, inaccurate declarations, related-party transaction anomalies, and abnormal fund flows. We extracted the cross-dimensional correlation feature set for each historical anomaly sample to form a training sample set. The training sample set contains 5000 labeled samples, including 800 false invoices, 1200 tax evasion samples, 1500 inaccurate declarations, 900 related-party transaction anomalies, and 600 abnormal fund flow samples.

[0095] An anomaly type identifier was trained using the gradient boosting tree algorithm. The number of decision trees was set to 200, with a maximum depth of 8 layers per tree, a learning rate of 0.1, and a subsampling ratio of 0.8. During training, decision trees were constructed sequentially. Each new tree was trained by fitting the residuals of its predecessor trees. The splitting nodes of each decision tree were selected by traversing the feature dimensions and candidate splitting points, choosing the splitting scheme that maximized the purity of the anomaly types in the resulting sample set. For example, the splitting condition of the root node of a certain decision tree might be whether the ratio of input tax to output tax is greater than 1.8. The left subtree might then determine whether the number of invoices issued exceeds 300, and the right subtree might determine whether the cash flow cycle is less than 20 days. After ensemble training of 200 decision trees, a classifier model capable of identifying five anomaly types was formed.

[0096] When determining anomalies in a sample, the cross-dimensional correlation feature vector of the sample is extracted. For example, the feature vector of a sample to be identified might be: turnover of 6.2 million yuan, input tax of 1.8 million yuan, output tax of 750,000 yuan, declaration frequency of 18 times per year, number of invoices issued of 520, and cash flow cycle of 15 days. The distance metric from the sample point to the feature vectors of each cluster center is calculated. The distance from the sample to the center of the normal operating cluster is 356 units, to the center of the high-frequency transaction cluster is 198 units, and to the center of the periodic fluctuation cluster is 287 units. Comparing these distance values ​​with the outlier determination threshold of each cluster, it is found that the distance of 198 units from the nearest high-frequency transaction cluster center exceeds the outlier determination threshold of 165 units for that cluster. Therefore, the sample is determined to be an anomaly.

[0097] When calculating the deviation of anomaly samples, the actual distance (198 units) from the sample to the nearest cluster (i.e., the high-frequency trading cluster) is extracted, and the boundary radius (110 units) of that cluster is extracted. The value exceeding the boundary radius is calculated to be 88 units. Dividing this excess value by the boundary radius yields a deviation ratio of 0.8. This deviation ratio reflects the severity of the anomaly sample's deviation from the normal pattern.

[0098] The feature vectors of anomalous samples are input into a trained anomaly type identifier. These feature vectors are then processed sequentially through 200 decision trees, with each tree outputting an anomaly type prediction. The outputs of all decision trees are then aggregated, and the number of votes for each anomaly type is counted. The following anomaly type received 35 votes: fictitious invoices, tax evasion, false declarations, related-party transactions, and fund flows. Dividing the number of votes for each anomaly type by the total number of decision trees (200) yields the anomaly type probability distribution: fictitious invoices 0.175, tax evasion 0.06, false declarations 0.04, related-party transactions 0.6, and fund flows 0.125. The related-party transaction anomaly type with the highest probability is selected as the anomaly type identifier for that sample.

[0099] The confidence level of the anomaly type identifier is adjusted using a deviation ratio. The baseline confidence level is set at the anomaly type probability value of 0.6, and adjusted based on the deviation ratio of 0.8. A higher deviation ratio indicates a more severe deviation from the normal pattern, and a higher degree of reliability in the anomaly determination. Multiplying the deviation ratio by a weighting factor of 0.3 yields a confidence increment of 0.24. Adding the baseline confidence level to the confidence increment gives an adjusted confidence level of 0.84. This adjusted confidence level serves as the final confidence score for identifying the anomaly type of related-party transactions, and is used for subsequent risk assessment and early warning decisions.

[0100] In one optional implementation, locating a subset of features corresponding to the anomaly type identifier within the cross-dimensional correlation feature set, identifying the contribution weight of each feature in the feature subset to the anomaly formation through a causal inference mechanism, and constructing a directed causal dependency graph between features based on the contribution weights includes:

[0101] In the cross-dimensional correlation feature set, identify samples labeled with the anomaly type identifier, extract the corresponding feature vectors, calculate the difference ratio between the variance of each feature dimension in the anomaly sample and the variance in the normal sample, and select the feature dimension set that has the ability to distinguish the anomaly type identifier as the feature subset based on the difference ratio.

[0102] For any two feature dimensions in the feature subset, the temporal order of the two feature dimensions in the data generation process is determined by temporal sequence analysis. The conditional independence test is used to determine whether a change in the value of one feature dimension causes a change in the value of the other feature dimension. Based on the temporal order and the conditional independence test results, it is determined whether there is a causal relationship between the two feature dimensions. If a causal relationship exists, the feature dimension that appears earlier in time is marked as the causal feature, and the feature dimension that appears later in time is marked as the effect feature. An integral operation is performed within the value range of the causal feature to obtain the contribution weight of the causal feature to the anomaly type identification.

[0103] Using each feature dimension in the feature subset as a node, and taking the nodes corresponding to the causal feature and the effect feature as the starting and ending points of the directed edges respectively, and using the contribution weight as the edge weight, a directed causal dependency graph between features is constructed.

[0104] For each sample in the abnormal sample group, its feature vector data is extracted. Assuming each sample's feature vector contains 12 feature dimensions: temperature sensor reading, current intensity, rotational speed, vibration frequency, pressure value, humidity, running time, load rate, voltage fluctuation amplitude, heat dissipation efficiency, ambient temperature, and coolant flow rate, the feature vectors of 200 samples in the abnormal sample group are read sequentially to form a 200-row, 12-column abnormal feature matrix. Similarly, the feature vectors of 4800 samples in the normal sample group are read to form a 4800-row, 12-column normal feature matrix.

[0105] Calculate the variance of each feature dimension in the outlier samples. Taking temperature sensor readings as an example, extract 200 values ​​for that dimension from the outlier feature matrix, calculate the sum of squared deviations between these 200 values ​​and their mean, and divide the sum of squared deviations by the number of samples to obtain a variance of 825 for temperature sensor readings in the outlier samples. For normal samples, extract 4800 values ​​for the temperature sensor reading dimension from the normal feature matrix, and calculate the variance of that dimension in the normal samples using the same method, obtaining a variance of 120. Divide the outlier sample variance of 825 by the normal sample variance of 120 to obtain a variance ratio of 6.875.

[0106] The variance and difference ratio calculation processes described above were performed on each of the 12 feature dimensions. The difference ratio for the current intensity dimension was 5.32, the difference ratio for the rotational speed dimension was 1.08, the difference ratio for the vibration frequency dimension was 7.91, the difference ratio for the pressure value dimension was 2.45, the difference ratio for the humidity dimension was 1.15, the difference ratio for the running time dimension was 0.93, the difference ratio for the load rate dimension was 6.02, the difference ratio for the voltage fluctuation amplitude dimension was 4.78, the difference ratio for the heat dissipation efficiency dimension was 8.36, the difference ratio for the ambient temperature dimension was 7.24, and the difference ratio for the coolant flow rate dimension was 5.67.

[0107] A difference ratio threshold of 3.0 was set, and feature dimensions with difference ratios greater than the threshold were selected. These included temperature sensor readings (6.875), current intensity (5.32), vibration frequency (7.91), load rate (6.02), voltage fluctuation amplitude (4.78), heat dissipation efficiency (8.36), ambient temperature (7.24), and coolant flow rate (5.67). These eight feature dimensions formed a feature subset for subsequent causal inference analysis.

[0108] Two feature dimensions, ambient temperature and temperature sensor readings, were selected from the feature subset to determine causal relationships. Through temporal sequence analysis, time-series data containing these two feature dimensions were extracted, and the trend of feature value changes in 200 outlier samples along the time axis was observed. In the data segment from time 0 to time 10, the ambient temperature began to rise at time 0, increasing from 25 degrees to 32 degrees, while the temperature sensor reading began to rise at time 3, increasing from 40 degrees to 55 degrees, showing a lag of 3 time units. Cross-correlation analysis was used to calculate the correlation coefficient of the two feature dimensions at different time offsets. It was found that the correlation coefficient reached its maximum value of 0.89 when the ambient temperature led the temperature sensor reading by 3 time units, confirming that the ambient temperature changed before the temperature sensor reading in time sequence.

[0109] To perform the conditional independence test, a set of control variables was introduced, comprising six characteristic dimensions other than ambient temperature and temperature sensor readings. With fixed values ​​for these six dimensions—coolant flow rate, heat dissipation efficiency, current intensity, vibration frequency, load rate, and voltage fluctuation amplitude—the effect of ambient temperature changes on temperature sensor readings was observed. A subset of samples with ambient temperatures between 28 and 35 degrees Celsius was selected. The results showed that, while keeping other characteristic dimensions constant, for every 1-degree increase in ambient temperature, the temperature sensor reading increased by an average of 2.3 degrees Celsius, indicating a significant conditional dependency. The chi-square test statistic yielded a value of 45.6, significantly higher than the critical value of 3.84, thus rejecting the conditional independence hypothesis and concluding that changes in ambient temperature cause changes in temperature sensor readings.

[0110] Based on the time sequence and conditional independence tests, a causal relationship was established between ambient temperature and temperature sensor readings. Ambient temperature was considered the causal feature, and temperature sensor readings the effect feature. An integral was performed within the range of ambient temperature (20°C to 40°C). This range was divided into 200 equally spaced intervals, each 0.1°C wide. Within each interval, the number of samples whose ambient temperature fell within that interval was counted. The first interval (20°C to 20.1°C) contained 0 samples, the 50th interval (24.9°C to 25°C) contained 3 samples, and the 120th interval (31.9°C to 32°C) contained 18 samples. Multiplying the number of samples in each interval by the interval width of 0.1°C and summing the products of all intervals yielded an integral result of 540.2. Dividing this integral result by the total number of abnormal samples (200) yielded a contribution weight of 2.701 for ambient temperature in identifying the high-temperature anomaly type.

[0111] Following the same method, pairwise combinations of the remaining feature dimensions in the feature subset were processed. Regarding the relationship between coolant flow rate and heat dissipation efficiency, coolant flow rate precedes heat dissipation efficiency in time sequence. The conditional independence test rejected the independence hypothesis, classifying coolant flow rate as the causal feature and heat dissipation efficiency as the effect feature. The calculated contribution weight of coolant flow rate was 1.893. Regarding the relationship between current intensity and temperature sensor readings, current intensity precedes temperature sensor readings, indicating a causal relationship. The contribution weight of current intensity was 2.345. Regarding the relationship between load rate and current intensity, load rate precedes current intensity, indicating a causal relationship. The contribution weight of load rate was 1.678. Regarding the relationship between vibration frequency and heat dissipation efficiency, vibration frequency precedes heat dissipation efficiency, indicating a causal relationship. The contribution weight of vibration frequency was 1.456.

[0112] Eight feature dimensions from the feature subset are each represented as graph nodes, labeled as ambient temperature, temperature sensor reading, coolant flow rate, heat dissipation efficiency, current intensity, load rate, vibration frequency, and voltage fluctuation amplitude. A directed edge is created between the ambient temperature node and the temperature sensor reading node, starting at the ambient temperature node and ending at the temperature sensor reading node, with a contribution weight of 2.701 assigned to the edge weight attribute. A directed edge is also created between the coolant flow rate node and the heat dissipation efficiency node, starting at the coolant flow rate node and ending at the heat dissipation efficiency node, with an edge weight of 1.893. A directed edge is created between the load rate node and the current intensity node, starting at the load rate node and ending at the current intensity node, with an edge weight of 1.678. A directed edge is also created between the current intensity node and the temperature sensor reading node, starting at the current intensity node and ending at the temperature sensor reading node, with an edge weight of 2.345. A directed edge is created between the vibration frequency node and the heat dissipation efficiency node, starting at the vibration frequency node and ending at the heat dissipation efficiency node, with an edge weight of 1.456. After creating directed edges corresponding to all causal relationships, a directed causal dependency graph between features containing 8 nodes and 13 directed edges is formed. This graph clearly shows the causal transmission path between each feature dimension and its contribution strength to the formation of anomalies.

[0113] In one optional implementation, path tracing is performed in the directed causal dependency graph along the direction of decreasing contribution weight, and the path with the maximum weight from the normal state node to the abnormal state node is extracted as the anomaly tracing path, including:

[0114] In the directed causal dependency graph, nodes whose corresponding feature dimension values ​​are within the compliant range are marked as normal state nodes, and the remaining nodes are marked as abnormal state nodes; the directed causal dependency graph is traversed, and node pairs whose starting node is a normal state node and whose ending node is an abnormal state node are identified as target node pairs.

[0115] Based on the target node pair, a reverse traversal is performed from the abnormal state node, and the previous nodes in the path are visited sequentially in the opposite direction of the directed edges. During the reverse traversal, the edge weight corresponding to each directed edge is extracted, and it is determined whether the sequence of edge weights extracted sequentially along the reverse traversal direction shows a decreasing trend.

[0116] In a connected path where the edge weight sequence shows a decreasing trend, the edge weights of each directed edge are summed to obtain the path weight of the connected path.

[0117] Among the connected paths that start from the same normal state node and reach the same abnormal state node and whose edge weight sequence shows a decreasing trend, the connected path with the largest path weight is selected as the maximum weight path. All nodes in the maximum weight path are arranged in chronological order to generate the abnormal source tracing path.

[0118] The process of extracting anomaly tracing paths in a directed causal dependency graph begins with classifying the states of all nodes in the graph. This requires obtaining the feature dimensions and actual values ​​of each node and comparing them with pre-defined compliance ranges. The compliance range is typically a threshold interval derived from statistical analysis of historical normal operation data. For example, the compliance range for CPU utilization of a data center server might be set at 20% to 70%, memory usage at 30% to 80%, and network bandwidth usage at 10% to 60%. The feature values ​​of each node in the graph are checked one by one. When node A's CPU utilization is 45%, this value is within the compliance range, so node A is marked as a normal state node, and a status identifier of "normal" is added to its node attributes. When node B's memory usage reaches 95%, this value exceeds the upper limit of the compliance range, so node B is marked as an abnormal state node, and a status identifier of "abnormal" is added to its node attributes. After completing the state marking by traversing all nodes in the graph, two node sets are established: one set stores the identifiers of all normal state nodes, and the other set stores the identifiers of all abnormal state nodes.

[0119] After classifying node states, the process begins identifying target node pairs using a double-loop traversal mechanism. The outer loop traverses the set of normal-state nodes, while the inner loop traverses the set of abnormal-state nodes. For each normal-state node, it checks if a connected path exists from that node through several directed edges to a specific abnormal-state node. This is done by performing a depth-first search or breadth-first search starting from the normal-state node, traversing the graph along the directed edges, and recording all nodes visited during the traversal. When the search reaches an abnormal-state node, a connection is confirmed between the normal and abnormal nodes, and this node pair is recorded as a target node pair. For example, if node C is a normal-state node, and nodes D and E are both abnormal-state nodes, graph traversal reveals paths from node C to node D and from node C to node E. Therefore, node pairs CD and CE are both recorded as target node pairs and added to the target node pair list.

[0120] After identifying the target node pair, reverse path tracing is performed for each target node pair. Taking the target node pair CD as an example, the reverse traversal starts from the abnormal state node D. Reverse traversal refers to visiting nodes along the reverse direction of the directed edges. Specifically, it involves finding all directed edges pointing to the current node and visiting the starting nodes of these edges. Node D receives directed edges from nodes F and G. The edge weight from F to D is extracted as 0.8, and the edge weight from G to D is extracted as 0.6. The edge weight represents the causal contribution of the preceding node to the current node; a larger value indicates a more significant contribution. These edge weights are stored in the edge weight sequence according to the reverse traversal order. Continuing the reverse traversal from node F, a directed edge from node C to node F is found, and the edge weight of this edge is extracted as 0.9. At this point, the edge weight sequence is 0.9 and 0.8. It is necessary to determine whether this sequence shows a decreasing trend.

[0121] The weights of adjacent edges in the sequence are compared to determine a decreasing trend. Starting from the first element, the numerical relationship between the current element and the next element is compared sequentially. For an edge weight sequence of 0.9 and 0.8, 0.9 is greater than 0.8, satisfying the decreasing condition. Continuing the reverse traversal to node G, if the directed edge weight from node C to node G is 0.7, then the edge weight sequence of another path is 0.7 and 0.6, also satisfying the decreasing condition. This decreasing trend ensures that the causal strength along the anomaly propagation path gradually decays, conforming to the physical law of anomalies spreading downstream from their source. If a path's edge weight sequence contains values ​​like 0.5, 0.7, and 0.4, since 0.5 is less than 0.7, the decreasing condition is not met, and this path is excluded as a valid candidate for anomaly tracing.

[0122] For connected paths that satisfy the decreasing edge weight condition, the path weight is calculated to quantify the overall causal contribution of the path. The path weight is calculated by summing the edge weights of all directed edges in the path. For the path C to F to D, the edge weights are 0.9 and 0.8 respectively, and the summation operation yields a path weight of 1.7. For the path C to G to D, the edge weights are 0.7 and 0.6 respectively, and the summation yields a path weight of 1.3. The calculated path weights are associated with and stored with the corresponding path information, which includes the identifiers of all nodes traversed in the path and their order. When multiple paths exist from the same normal state node to the same abnormal state node that satisfy the decreasing edge weight condition, the optimal path needs to be selected.

[0123] The path with the highest weight is selected by comparing the path weights of all candidate paths. This involves extracting the path weights of all valid paths from the normal state node C to the abnormal state node D, and then comparing their values ​​to find the maximum value. In the example above, the path weight 1.7 is greater than 1.3, so the path C to F to D is selected as the path with the highest weight. This path represents the propagation chain with the strongest causal contribution during the evolution from the normal operating state to the abnormal state, and has the highest reference value for locating the root cause of the anomaly. The complete node sequence corresponding to this path with the highest weight is recorded as C, F, D.

[0124] The nodes in the maximum weighted path are arranged and output according to the causal propagation time sequence. Since the reverse traversal traces from the abnormal node to the normal node, the node access order is D, F, C. This order needs to be reversed to obtain the forward causal propagation order C, F, D. The reversed node sequence is output as the anomaly tracing path. This path clearly shows how the anomaly started from node C in an initially normal state, passed through the intermediate node F, and finally led to the abnormal state of node D. The output anomaly tracing path also includes detailed characteristic information of each node. For example, node C's CPU utilization is 45%, which is normal; node F's network latency starts to show slight abnormal signs at 50 milliseconds; and node D's response time exceeds 3 seconds, reaching a severely abnormal state. This information provides maintenance personnel with a complete anomaly evolution chain and detailed state change data support.

[0125] In one optional implementation, based on the anomaly tracing path and the set of compliance constraint rules, the tax declaration data requiring correction and the corresponding correction direction are located. According to the correction direction, the tax declaration data requiring correction is adjusted numerically or its status is changed, including:

[0126] Obtain all abnormal state nodes and corresponding feature dimension identifiers in the abnormal source tracing path, locate the related tax declaration data items in the tax declaration data according to the feature dimension identifiers, and mark them as the tax declaration data that needs to be corrected;

[0127] Extract the identifiers of tax declaration data items involved in the set of compliance constraint rules, determine whether the identifiers of the tax declaration data items match the data item names of the tax declaration data to be corrected, extract the compliant value ranges defined in the successfully matched constraint rules, calculate the deviation between the current value of the tax declaration data to be corrected and the boundary value of the compliant value range, and determine the correction direction based on the sign of the deviation.

[0128] Calculate the numerical adjustment range based on the absolute value of the deviation, add or subtract the numerical adjustment range to the current value of the tax declaration data to be corrected according to the correction direction, and generate the corrected tax declaration data value; record the status compliance transition path of the tax declaration data under different correction directions, and perform a status change operation on the tax declaration data to be corrected according to the status compliance transition path.

[0129] The anomaly tracing path stores the complete anomaly propagation chain from the root node to the leaf node. Each node carries an anomaly status identifier and a corresponding feature dimension identifier. Traversing all nodes in this path, the anomaly status identifier is extracted. When a node's status attribute value is "abnormal" or "non-compliant," that node is included in the anomaly status node set. Each anomaly status node is associated with one or more feature dimension identifiers. For example, a feature dimension identifier of "revenue_amount" represents the income amount dimension, and a feature dimension identifier of "tax_rate_type" represents the tax rate type dimension. All data items in the tax return data table are read, and the field names of each data item are compared one by one with the feature dimension identifiers. When a data item's field name completely matches the feature dimension identifier, that data item is marked as tax return data requiring correction.

[0130] In a real-world case, the anomaly tracing path contained three anomalous state nodes, identified by the feature dimensions "taxable_income", "deduction_amount", and "final_tax_payable". The data item with the field name "taxable_income" was located in the tax return data table, with a current value of 1,500,000 yuan, and was added to the list of data requiring correction. Simultaneously, data items with the field name "deduction_amount" (current value 200,000 yuan) and the field name "final_tax_payable" (current value 325,000 yuan) were also located and added to the list of data requiring correction.

[0131] Load the compliance constraint rule set, which contains multiple rules, each defining a tax return data item identifier, a compliant value range, and state transition conditions. Parse the first constraint rule, extracting the tax return data item identifier "taxable_income". Compare this identifier with the data item name in the data list to be corrected; a data item named "taxable_income" is found, indicating a successful match. Further extract the compliant value range defined in this constraint rule, which specifies that the compliant value range for this data item is between 800,000 yuan and 1,200,000 yuan. Obtain the current value of the data item to be corrected, "taxable_income," which is 1,500,000 yuan. Calculate the deviation between the current value and the upper boundary of the compliant value range, 1,200,000 yuan. The deviation equals the current value minus the upper boundary value, resulting in 300,000 yuan. Since the deviation is positive, the correction direction is determined to be downward, meaning the value of this data item needs to be reduced.

[0132] Continuing with the second constraint rule, the tax return data item identified as "deduction_amount" matches the data item name in the list requiring correction. This constraint rule defines a compliant value range of 50,000 to 150,000 yuan. The current value is 200,000 yuan. The deviation between the current value and the upper limit of the compliant range of 150,000 yuan is calculated to be 50,000 yuan, which is positive. Therefore, the correction direction is also downward. For the third data item "final_tax_payable", the constraint rule stipulates a compliant value range of 180,000 to 260,000 yuan. The current value of 325,000 yuan exceeds the upper limit, with a deviation of 65,000 yuan. The correction direction is also downward.

[0133] The adjustment range is calculated based on the absolute value of the deviation. For the data item "taxable_income", the absolute value of the deviation is 300,000 yuan. With an adjustment factor of 80%, the adjustment range equals the absolute value of the deviation multiplied by the adjustment factor, resulting in 240,000 yuan. The current value of 1,500,000 yuan is then adjusted downwards by subtracting the adjustment range of 240,000 yuan, resulting in a corrected value of 1,260,000 yuan. This value is still slightly above the upper limit of the compliant range, so a second round of adjustment is performed. A new deviation of 60,000 yuan is calculated, multiplied by the adjustment factor to obtain an adjustment range of 48,000 yuan, and then subtracted again to obtain a corrected value of 1,212,000 yuan. After iterative adjustments, the final value is adjusted to 1,200,000 yuan, bringing it within the compliant range.

[0134] For the data item "deduction_amount", the absolute value of the deviation of 50,000 yuan is multiplied by the adjustment factor of 80% to obtain an adjustment range of 40,000 yuan. The current value of 200,000 yuan is reduced by the adjustment range of 40,000 yuan to obtain a corrected value of 160,000 yuan. This value still exceeds the upper boundary, so the adjustment continues until the value drops to 150,000 yuan. For the data item "final_tax_payable", the absolute value of the deviation of 65,000 yuan is calculated to obtain an adjustment range of 52,000 yuan. The current value of 325,000 yuan is reduced by 52,000 yuan to obtain 273,000 yuan. After multiple rounds of iteration, it is finally adjusted to 260,000 yuan.

[0135] While adjusting the values, the compliance transition path is recorded. For the data item "taxable_income", the initial state is recorded as "Exceeding the limit, abnormal". After the first downward adjustment, the state changes to "Approaching the limit". After the second adjustment, the state changes to "Fully compliant". The compliance transition path is stored as a state sequence of "Exceeding the limit - Approaching the limit - Fully compliant". Based on this transition path, a state change operation is performed on the data item, updating the data item's state attribute field to "Fully compliant", and the historical state sequence is written to the state change log table.

[0136] When anomalies in certain tax return data items do not involve numerical ranges but rather status compliance, a status change operation is performed. For example, if the data item "tax_declaration_status" is currently in the "Unreviewed" state, but the constraint rule requires that the status must be "Approved," a compliance transition path is parsed, defined as "Unreviewed - Under Review - Approved." The status change is executed step-by-step according to the transition path, updating the data item's status from "Unreviewed" to "Under Review," triggering the review process, and then updating the status to "Approved," completing the compliance transition. The corrected tax return data is written to the corrected data table for use in subsequent filing processes.

[0137] This invention provides an intelligent comparison and automatic verification optimization system for enterprise tax declaration data, the system comprising:

[0138] The first unit is used to extract multi-level features from the tax declaration data of enterprises and to construct a cross-dimensional related feature set based on the business rule constraints of the industry to which the enterprise belongs.

[0139] The second unit is used to extract constraints and threshold boundaries between tax declaration data based on compliant and non-compliant samples in historical tax declaration data through association rule mining algorithms, and construct a set of compliance constraint rules; based on a hybrid architecture of unsupervised clustering algorithm and supervised classification algorithm, it performs abnormal pattern matching and deviation calculation on the cross-dimensional association feature set to identify abnormal type identifiers;

[0140] The third unit is used to locate the feature subset corresponding to the anomaly type identifier in the cross-dimensional associated feature set, identify the contribution weight of each feature in the feature subset to the formation of the anomaly through a causal inference mechanism, construct a directed causal dependency graph between features based on the contribution weight, trace the path along the direction of decreasing contribution weight in the directed causal dependency graph, and extract the maximum weight path from the normal state node to the abnormal state node as the anomaly tracing path.

[0141] The fourth unit is used to locate the tax declaration data that needs to be corrected and the corresponding correction direction based on the abnormality tracing path and the set of compliance constraint rules, and to adjust the value or change the status of the tax declaration data that needs to be corrected according to the correction direction to obtain the corrected tax declaration data.

[0142] A third aspect of the present invention provides an electronic device, comprising:

[0143] processor;

[0144] Memory used to store processor-executable instructions;

[0145] The processor is configured to invoke instructions stored in the memory to execute the aforementioned method.

[0146] A fourth aspect of the present invention provides a computer-readable storage medium having stored thereon computer program instructions that, when executed by a processor, implement the aforementioned method.

[0147] This invention can be a method, apparatus, system, and / or computer program product. The computer program product may include a computer-readable storage medium having computer-readable program instructions loaded thereon for performing various aspects of the invention.

[0148] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. An optimized method for intelligent comparison and automatic verification of enterprise tax declaration data, characterized in that: include: Multi-level feature extraction is performed on the tax declaration data of enterprises, and a cross-dimensional related feature set is constructed based on the business rule constraints of the industry to which the enterprise belongs; Based on compliant and non-compliant samples in historical tax declaration data, the constraints and threshold boundaries between tax declaration data are extracted through association rule mining algorithms to construct a set of compliance constraint rules; based on a hybrid architecture of unsupervised clustering algorithm and supervised classification algorithm, abnormal pattern matching and deviation calculation are performed on the cross-dimensional association feature set to identify abnormal type identifiers. Locate the feature subset corresponding to the anomaly type identifier in the cross-dimensional associated feature set, identify the contribution weight of each feature in the feature subset to the formation of the anomaly through the causal inference mechanism, construct a directed causal dependency graph between features based on the contribution weight, trace the path along the direction of decreasing contribution weight in the directed causal dependency graph, and extract the maximum weight path from the normal state node to the abnormal state node as the anomaly source tracing path. Based on the anomaly tracing path and the set of compliance constraint rules, the tax declaration data that needs to be corrected and the corresponding correction direction are located. According to the correction direction, the tax declaration data that needs to be corrected is adjusted numerically or its status is changed to obtain the corrected tax declaration data.

2. The method according to claim 1, characterized in that, Multi-level feature extraction is performed on the enterprise's tax declaration data, and a cross-dimensional correlation feature set is constructed based on the business rule constraints of the enterprise's industry, including: The tax declaration data of enterprises is analyzed in a multi-dimensional and layered manner. The periodic change characteristics of financial indicators are extracted from the time dimension, the revenue and cost structure characteristics corresponding to different business types are extracted from the business dimension, and the correlation relationship characteristics between tax declaration items are extracted from the account dimension, forming an initial feature set. The standard financial indicator distribution range and revenue-cost structure ratio range of the industry to which the enterprise belongs are used as the industry standard reference framework. The periodic change characteristics and the revenue-cost structure characteristics are mapped to the numerical space corresponding to the industry standard reference framework. By calculating the distance between the feature value of the tax declaration data and the boundary of the industry standard reference framework, the industry adaptability deviation characteristics are generated. Based on the correlation relationship features, the association and dependency between each tax data item are determined. A row index is created based on the association and dependency. The numerical ratio between any two tax data items is used as a matrix element to construct a mutual dependency constraint matrix. The mutual dependency constraint matrix is ​​used to perform correlation fusion on the features of each dimension in the initial feature set. The contribution of each dimension feature in the fusion is adjusted by the industry adaptability deviation feature to generate the cross-dimensional correlation feature set.

3. The method according to claim 1, characterized in that, Based on compliant and non-compliant samples from historical tax return data, an association rule mining algorithm is used to extract constraints and threshold boundaries between tax return data, constructing a set of compliance constraint rules including: Based on the historical tax audit results, the historical tax declaration data is divided into a compliant sample set and a non-compliant sample set. The numerical distribution characteristics of the compliant sample set and the non-compliant sample set are extracted, the distribution boundary differences of the numerical distribution characteristics on each tax data item are calculated, and the numerical boundary threshold for distinguishing between compliant and non-compliant samples is determined. In the set of compliant samples, identify combinations of tax data items whose co-occurrence frequency exceeds a preset frequency threshold, perform statistical analysis on the numerical relationships between each tax data item, calculate the numerical ratio relationship and state logic dependency relationship between each data item in the combination of tax data items, and generate numerical ratio constraints and logical consistency conditions. The numerical boundary threshold is used as a single constraint condition, and the numerical ratio constraint and the logical consistency condition are used as multiple associated constraints conditions. The single constraint condition and the multiple associated constraints are classified and merged according to the type of tax data item. A mapping relationship between trigger conditions and constraint boundaries is established for the constraints under each category to generate the compliance constraint rule set.

4. The method according to claim 1, characterized in that, Based on a hybrid architecture combining unsupervised clustering and supervised classification algorithms, anomaly pattern matching and deviation calculation are performed on the cross-dimensional associated feature set to identify anomaly type identifiers, including: A feature representation space is constructed using each feature dimension in the cross-dimensional associated feature set as the coordinate axis. Cluster analysis is performed using the cross-dimensional associated feature set as input data. Based on the distance measurement and density distribution characteristics between sample points in the feature representation space, the cross-dimensional associated feature set is divided into multiple clusters, and the boundary range and outlier sample judgment threshold of each cluster are determined. Using historical tax declaration data with labeled anomaly types as training samples, an anomaly type recognizer is constructed by learning the feature patterns corresponding to different anomaly types through a gradient boosting tree-based classification algorithm. Calculate the Euclidean distance from each sample to be identified to the central feature vector of each cluster. If the Euclidean distance exceeds the outlier determination threshold, the sample to be identified is determined to be an anomalous sample. Calculate the deviation of the anomalous sample from the boundary range of the nearest cluster. Input the feature vector corresponding to the anomalous sample into the anomaly type identifier to obtain the anomaly type probability distribution. Select the anomaly type with the highest probability value as the anomaly type identifier. Use the deviation as the confidence adjustment parameter for the anomaly type identifier.

5. The method according to claim 1, characterized in that, Locating the feature subset corresponding to the anomaly type identifier in the cross-dimensional correlation feature set, identifying the contribution weight of each feature in the feature subset to the anomaly formation through a causal inference mechanism, and constructing a directed causal dependency graph between features based on the contribution weights includes: In the cross-dimensional correlation feature set, identify samples labeled with the anomaly type identifier, extract the corresponding feature vectors, calculate the difference ratio between the variance of each feature dimension in the anomaly sample and the variance in the normal sample, and select the feature dimension set that has the ability to distinguish the anomaly type identifier as the feature subset based on the difference ratio. For any two feature dimensions in the feature subset, the temporal order of the two feature dimensions in the data generation process is determined by temporal sequence analysis. The conditional independence test is used to determine whether a change in the value of one feature dimension causes a change in the value of the other feature dimension. Based on the temporal order and the conditional independence test results, it is determined whether there is a causal relationship between the two feature dimensions. If a causal relationship exists, the feature dimension that appears earlier in time is marked as the causal feature, and the feature dimension that appears later in time is marked as the effect feature. An integral operation is performed within the value range of the causal feature to obtain the contribution weight of the causal feature to the anomaly type identification. Using each feature dimension in the feature subset as a node, and taking the nodes corresponding to the causal feature and the effect feature as the starting and ending points of the directed edges respectively, and using the contribution weight as the edge weight, a directed causal dependency graph between features is constructed.

6. The method according to claim 1, characterized in that, In the directed causal dependency graph, path tracing is performed along the direction of decreasing contribution weight, and the path with the maximum weight from the normal state node to the abnormal state node is extracted as the anomaly tracing path, including: In the directed causal dependency graph, nodes whose corresponding feature dimension values ​​are within the compliant range are marked as normal state nodes, and the remaining nodes are marked as abnormal state nodes; the directed causal dependency graph is traversed, and node pairs whose starting node is a normal state node and whose ending node is an abnormal state node are identified as target node pairs. Based on the target node pair, a reverse traversal is performed from the abnormal state node, and the previous nodes in the path are visited sequentially in the opposite direction of the directed edges. During the reverse traversal, the edge weight corresponding to each directed edge is extracted, and it is determined whether the sequence of edge weights extracted sequentially along the reverse traversal direction shows a decreasing trend. In a connected path where the edge weight sequence shows a decreasing trend, the edge weights of each directed edge are summed to obtain the path weight of the connected path. Among the connected paths that start from the same normal state node and reach the same abnormal state node and whose edge weight sequence shows a decreasing trend, the connected path with the largest path weight is selected as the maximum weight path. All nodes in the maximum weight path are arranged in chronological order to generate the abnormal source tracing path.

7. The method according to claim 1, characterized in that, Based on the anomaly tracing path and the set of compliance constraint rules, the tax declaration data requiring correction and the corresponding correction direction are located. According to the correction direction, the tax declaration data requiring correction is adjusted numerically or its status is changed, including: Obtain all abnormal state nodes and corresponding feature dimension identifiers in the abnormal source tracing path, locate the related tax declaration data items in the tax declaration data according to the feature dimension identifiers, and mark them as the tax declaration data that needs to be corrected; Extract the identifiers of tax declaration data items involved in the set of compliance constraint rules, determine whether the identifiers of the tax declaration data items match the data item names of the tax declaration data to be corrected, extract the compliant value ranges defined in the successfully matched constraint rules, calculate the deviation between the current value of the tax declaration data to be corrected and the boundary value of the compliant value range, and determine the correction direction based on the sign of the deviation. Calculate the numerical adjustment range based on the absolute value of the deviation, add or subtract the numerical adjustment range to the current value of the tax declaration data to be corrected according to the correction direction, and generate the corrected tax declaration data value; record the status compliance transition path of the tax declaration data under different correction directions, and perform a status change operation on the tax declaration data to be corrected according to the status compliance transition path.

8. An intelligent comparison and automatic verification optimization system for enterprise tax declaration data, used to implement the method as described in any one of claims 1-7, characterized in that, include: The first unit is used to extract multi-level features from the tax declaration data of enterprises and to construct a cross-dimensional related feature set based on the business rule constraints of the industry to which the enterprise belongs. The second unit is used to extract constraints and threshold boundaries between tax declaration data based on compliant and non-compliant samples in historical tax declaration data through association rule mining algorithms, and construct a set of compliance constraint rules; based on a hybrid architecture of unsupervised clustering algorithm and supervised classification algorithm, it performs abnormal pattern matching and deviation calculation on the cross-dimensional association feature set to identify abnormal type identifiers; The third unit is used to locate the feature subset corresponding to the anomaly type identifier in the cross-dimensional associated feature set, identify the contribution weight of each feature in the feature subset to the formation of the anomaly through a causal inference mechanism, construct a directed causal dependency graph between features based on the contribution weight, trace the path along the direction of decreasing contribution weight in the directed causal dependency graph, and extract the maximum weight path from the normal state node to the abnormal state node as the anomaly tracing path. The fourth unit is used to locate the tax declaration data that needs to be corrected and the corresponding correction direction based on the abnormality tracing path and the set of compliance constraint rules, and to adjust the value or change the status of the tax declaration data that needs to be corrected according to the correction direction to obtain the corrected tax declaration data.

9. An electronic device, characterized in that, include: processor; Memory used to store processor-executable instructions; The processor is configured to invoke instructions stored in the memory to execute the method according to any one of claims 1 to 7.

10. A computer-readable storage medium having computer program instructions stored thereon, characterized in that, When the computer program instructions are executed by the processor, they implement the method described in any one of claims 1 to 7.