A tax declaration manuscript management system and method based on rule engine
Through the tax declaration draft management system based on the rules engine, combined with the hybrid model architecture of long-term and short-term memory network and autoencoder, the problems of inadequate rule requirements for complex business scenarios and insufficient data tampering risk warnings in the existing technology are solved, and efficient, safe processing and risk warnings are achieved for tax declaration data.
Patent Information
- Application Number
- CN202510213179.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-26
- Publication Date
- 2025-05-20
- Estimated Expiration
- 2045-02-26
AI Technical Summary
The existing technology lacks the need for dynamic rules for complex business scenarios in the processing of tax declaration data, fails to effectively warn of potential data tampering risks and tax logic vulnerabilities, and there are problems such as single-dimensional analysis and weight determination in feature extraction and model construction.
The tax declaration paper management system based on the rules engine is adopted, and multi-dimensional tax declaration data is extracted through the development of data interfaces and ETL tools, a Drools workspace is built and the rule engine is configured, and tax compliance, data logic consistency and risk warning rules are entered. Combining the hybrid model architecture of long and short-term memory networks and autoencoders, feature extraction and model training are carried out to judge the tampering risk of tax declaration papers and provide early warning.
It has achieved the satisfaction of dynamic rules requirements for complex business scenarios, early warning of tax risks, improved the accuracy and security of tax declaration data processing, and ensured the scientific and rationality of the weight system.
Smart Images

Figure CN119719205B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and specifically to a tax return draft management system and method based on a rule engine. Background Art
[0002] In today's digital age, the accuracy and security of tax return data processing are of crucial importance. With the continuous expansion of enterprise business scale and the increasingly strict tax supervision requirements, traditional tax return data processing methods face many difficulties.
[0003] In terms of rule application, most of the existing technologies roughly check based on basic tax regulations statically, without fully considering the rule requirements in complex and changeable actual business scenarios. The lack of in-depth exploration of potential risks, such as the lack of early warning mechanisms for data tampering risks and hidden tax logic loopholes, makes enterprises passive when facing tax audits. In feature extraction and model construction, existing technologies usually only simply analyze a single data dimension in isolation, without deeply exploring the combined value between different data features. Moreover, in the method of determining weights, traditional means either simply rely on the simple average of historical data, ignoring the complex factors behind data fluctuations; or rely solely on subjective experience judgment, lacking scientific and rigorous quantitative analysis, resulting in unreasonable weight settings for each indicator. Summary of the Invention
[0004] The purpose of the present invention is to provide a tax return draft management system and method based on a rule engine to solve the problems raised in the existing technology.
[0005] To achieve the above purpose, the present invention provides the following technical solutions:
[0006] A tax return draft management method based on a rule engine, the method comprising the following steps:
[0007] Develop a data interface to connect to the data sources within the enterprise, and extract tax return data, including financial data, invoice data, inventory data, and human resource data, at a predetermined frequency through an ETL tool for data preprocessing and integration;
[0008] Build a Drools workspace, create a rule library directory for storing rule files, and configure the running parameters of the rule engine; enter rules based on financial data and invoice data, the rules including tax compliance rules, data logical consistency rules, and risk warning rules;
[0009] Extract inventory features based on inventory data and human resource features based on human resource data; calculate the first operational efficiency human comprehensive feature and the second operational efficiency human comprehensive feature according to the inventory features and the human resource features;
[0010] For the first comprehensive human resource characteristics of operational efficiency and the second comprehensive human resource characteristics of operational efficiency, perform time series expansion and data annotation based on the rules of the Drools rule engine; select a hybrid model architecture that combines a long short-term memory network and an autoencoder, divide the annotated time series data set into a training set, a validation set, and a test set, and perform model training, evaluation, and optimization;
[0011] When there is new data update or submission in the tax return working papers, input the newly generated feature sequence into the trained hybrid model to determine whether there is a risk of tampering in the current tax return working papers. If there is a risk of tampering, give a warning to the staff.
[0012] Further, after the enterprise's financial department performs the daily closing operation every day, use the scheduling function of the ETL tool to obtain the latest financial data; write an associated query statement to integrate the data in the invoice management system, associate the detail table from the invoice main table to obtain invoice data; for the inventory system database, use the query statement to associate and query the real-time status of the inventory from the inbound table, outbound table, and inventory balance table to obtain inventory data; query information from the human resource management system database to obtain human resource data;
[0013] Perform data cleaning and data transformation on the extracted financial data, invoice data, inventory data, and human resource data, create a database based on tax return requirements, construct dimension tables and fact tables therein, and realize the integration of data from different data sources through the primary key-foreign key association mechanism; associate the financial data with the invoice data, the inventory data with the financial data, and the human resource data with the financial data.
[0014] Further, set aside an area for Drools-related development and management, and divide it into multiple subdirectories according to functions, including: a rule library directory, a model directory, and a test directory; among them, the rule library directory is used to store rule files related to tax returns, which define constraint conditions and behavior logics; the model directory stores domain model objects related to tax returns, which are abstractions and simplifications of the underlying data structures and are presented in the form of Java classes; the test directory writes unit test and integration test cases, inputs different test data, verifies whether the rules can be triggered correctly, and outputs the expected results;
[0015] In the "kmodule.xml" in the Drools workspace, define various configurations of the knowledge session, which is used to ensure that different types of sessions meet different business requirements; perform rule loading and update parameters, and set execution priorities, conflict resolution parameters, and performance optimization parameters.
[0016] According to the conventional technology in this field, the revenue recognition rules in financial data are set; for enterprises adopting different sales models, the rules for when and at what amount the revenue under each model should be recognized in the financial statements are clarified; based on invoice data, the rules for issuing, authenticating and deducting special VAT invoices and ordinary invoices are standardized;
[0017] Establish a logical connection between financial income data and invoice data. The rule is set that the main business income in the financial statements should fluctuate within a predefined difference range with the tax-inclusive sales amount issued by invoices in the same period. The difference range is determined through statistical analysis of historical data; set rules for the logical coherence of invoice data;
[0018] For risk warning rules, we will conduct warnings on abnormal data fluctuations, tax rate deviations and related-party transaction risks.
[0019] Based on inventory data, inventory features are extracted, including dynamic features, trend features and correlation features. The dynamic features include inventory turnover rate and inventory change rate. The trend features include moving average and seasonal index. The correlation feature is the correlation between inventory and sales. Based on human resource data, human resource features are extracted, including salary structure features and personnel change features. The salary structure features include salary composition ratio and salary level distribution. The personnel change features include employee turnover rate and new employee ratio.
[0020] Furthermore, for the inventory dynamic characteristics, the inventory turnover rate and the inventory change rate are standardized and weighted and summed according to the first weight, which is determined according to the importance of the two parameters in the inventory dynamic characteristics on the operating efficiency, to obtain the inventory operating efficiency value; for the inventory trend characteristics, the moving average and the seasonal index are standardized and weighted and summed according to the second weight, which is weighted according to the relative importance of the two parameters in the inventory trend characteristics in reflecting the inventory trend, to obtain the inventory trend comprehensive value; for the salary structure characteristics, the basic salary ratio, performance salary ratio, bonus ratio and welfare ratio are weighted and summed according to their respective proportions in the total salary to obtain the salary structure comprehensive value; the inventory operating efficiency value, the inventory trend comprehensive value and the salary structure comprehensive value are weighted and summed according to the third weight, which is determined through actual business verification, to obtain the first operating efficiency human comprehensive characteristics;
[0021] The correlation between inventory and sales is standardized to obtain a standardized correlation value; for the personnel change characteristics, the employee turnover rate and the proportion of new employees are standardized, and weighted summed according to the fourth weight, which is determined according to the importance of the two parameters in the personnel change characteristics on the operational efficiency, to obtain a comprehensive value of personnel change; the standardized correlation value and the comprehensive value of personnel change are added according to half the weight to obtain the second comprehensive human resource characteristics of operational efficiency.
[0022] Further, the historical data analysis regression method is used to determine the first weight, and the inventory turnover rate of the enterprise in the past m quarters is collected , the inventory change rate and the corresponding operation efficiency indicators ; A multiple linear regression model is established: , where is the intercept, is the first regression coefficient, is the second regression coefficient, is the random error term;
[0023] By performing regression analysis on historical data, the estimated values of the regression coefficients and are obtained. The calculation formula for the weight of the inventory turnover rate in the first weight is: , and the calculation formula for the weight of the inventory change rate is: ;
[0024] For the second weight and the fourth weight, the expert scoring method is used for determination;
[0025] For the third weight, the analytic hierarchy process is used for determination; a hierarchical structure model is constructed. The target layer is to determine the comprehensive characteristics of the first operation efficiency human resources. The criterion layer includes inventory operation efficiency, inventory trend, and salary structure. The scheme layer is the weight allocation for each criterion; through pairwise comparison of the relative importance between the factors in the criterion layer by experts, a judgment matrix is constructed; the maximum eigenvalue of the judgment matrix and its corresponding eigenvector are calculated, and the eigenvector is normalized to obtain the normalized vector elements as the third weight.
[0026] Further, multi-period inventory data and human resource data are retrieved from the enterprise's database, and the collected data is arranged in chronological order to form a structured time series dataset; for the first operation efficiency human resources comprehensive characteristics and the second operation efficiency human resources comprehensive characteristics, sliding window statistics are calculated; seasonal decomposition is used to split the time series of each comprehensive characteristic into a trend component, a seasonal component, and a residual component;
[0027] When the operation efficiency human resources comprehensive characteristic data after the time series expansion flows into the Drools rule engine, the engine sequentially starts rule matching; each data point is checked for compliance rules one by one. When a certain compliance rule condition is met, the corresponding compliance risk label is marked on the data; risk warning rule matching is performed, and when triggered, the risk warning label is superimposed and marked; the data logic consistency rule is executed, and if the condition is met, the data logic consistency label is added.
[0028] Further, based on the labeled time series dataset, the dataset is divided into a training set, a validation set, and a test set according to a ratio; in the hybrid model architecture, the long short-term memory network processes the input time series data sequentially, and each long short-term memory network receives the enterprise operation feature vector at the current time step, including the first operation efficiency human comprehensive feature, the second operation efficiency human comprehensive feature, and each sub-feature included therein; combined with the hidden state of the previous time step, through the input gate, the forget gate, and the output gate, it is determined which information needs to be retained, updated, or forgotten; the autoencoder is used to learn the feature representation of the data and consists of an encoder and a decoder; the encoder compresses the high-dimensional input data into a low-dimensional coding space and extracts the key features of the data; during the training process, the model, based on the labeled information, aims to minimize the loss function between the prediction result and the true label, and continuously adjusts the weight parameters of the model;
[0029] At each stage of model training and after the training is finally completed, the model is evaluated; according to the results of model evaluation, the hyperparameters of the long short-term memory network and the autoencoder are adjusted.
[0030] The newly generated feature sequence is input into a hybrid model combined with a pre-trained long short-term memory network and an autoencoder. The model, based on the normal and abnormal sample feature patterns accumulated during the training stage, compares the currently input feature sequence, comprehensively judges whether there is a risk of tampering in the current tax return draft, and gives the corresponding risk probability value;
[0031] For the high-risk level, a pop-up window appears on the tax return operation interface, with a risk prompt in the pop-up window content, attaching a key abnormal data comparison chart to show the difference between the current data and the normal range, and giving a recommended verification direction; for the medium-risk level, the warning information is sent to the relevant staff in the form of an email; for the low-risk level, the warning information is published in the system internal notification bar.
[0032] A tax return draft management system based on a rule engine, including:
[0033] Data collection and integration module: including: a data collection unit and a data preprocessing unit; among them, the data collection unit develops data interfaces to connect to the data sources within the enterprise, extracts tax return data, including financial data, invoice data, inventory data, and human resource data, at a predetermined frequency through an ETL tool, and the data integration unit performs data preprocessing and integration;
[0034] Rule Engine Building Module: It includes: Drools Workspace Building Unit and Rule Engine Configuration Unit; among them, the Drools Workspace Building Unit builds a Drools workspace, creates a rule library directory for storing rule files, and configures the running parameters of the rule engine; the Rule Engine Configuration Unit inputs rules based on financial data and invoice data, and the rules include tax compliance rules, data logic consistency rules, and risk warning rules;
[0035] Feature Engineering Module: It includes: Inventory Feature Extraction Unit, Human Resource Feature Extraction Unit, and Comprehensive Feature Calculation Unit; among them, the Inventory Feature Extraction Unit extracts inventory features based on inventory data, and the Human Resource Feature Extraction Unit extracts human resource features based on human resource data; the Comprehensive Feature Calculation Unit calculates the first operation efficiency human comprehensive feature and the second operation efficiency human comprehensive feature according to the inventory features and the human resource features;
[0036] Time Series Processing and Model Training Module: It includes: Time Series Expansion Unit and Model Training Unit; among them, the Time Series Expansion Unit performs time series expansion on the first operation efficiency human comprehensive feature and the second operation efficiency human comprehensive feature, and performs data annotation based on the rules of the Drools rule engine; the Model Training Unit selects a hybrid model architecture combining long short-term memory network and autoencoder, divides the labeled time series data set into training set, validation set, and test set, and performs model training, evaluation, and optimization;
[0037] Risk Warning Module: It includes: Risk Judgment Unit and Warning Push Unit; among them, when there is new data update or submission in the tax return draft, the Risk Judgment Unit inputs the newly generated feature sequence into the trained hybrid model to judge whether there is a risk of tampering in the current tax return draft. If there is a risk of tampering, the Warning Push Unit warns the staff.
[0038] Compared with the prior art, the beneficial effects of the present invention are:
[0039] 1. The present invention builds a Drools workspace and scientifically configures the running parameters of the rule engine, laying a solid foundation for the efficient execution of rules. On this basis, comprehensive and detailed rules are input based on multi-dimensional data, deeply insight into the internal logic errors of the data, and prevent tax risks in advance.
[0040] 2. In terms of model construction, the present invention selects a hybrid model architecture combining long short-term memory network and autoencoder, which is specifically optimized for the time series characteristics of tax return data. This model can effectively capture the long-term dependence relationship of the data and has strong feature self-learning and noise reduction capabilities.
[0041] 3. In determining the weights, the present invention combines multiple scientific methods. For some weights, the historical data analysis and regression method is used, and for other weights, the expert scoring method and the analytic hierarchy process are respectively adopted, fully absorbing professional knowledge and experience wisdom, comprehensively weighing the influences of different factors, ensuring the scientific and reasonable weight system, and comprehensively improving the accuracy of tax return risk judgment. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] Figure 1 is a schematic diagram of the steps of a method for managing tax return work sheets based on a rule engine according to the present invention;
[0043] Figure 2 is a system structure diagram of a system for managing tax return work sheets based on a rule engine according to the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0044] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0045] Embodiment: As Figure 1 - Figure 2 shown, the present invention provides a technical solution.
[0046] According to an embodiment of the present invention, as Figure 1 shown in the schematic diagram of the steps of a method for managing tax return work sheets based on a rule engine, a method for managing tax return work sheets based on a rule engine, the method includes the following steps:
[0047] Develop a data interface to connect to the data sources within the enterprise, extract tax return data, including financial data, invoice data, inventory data, and human resource data, at a predetermined frequency through an ETL tool, and perform data preprocessing and integration;
[0048] Build a Drools workspace, create a rule library directory for storing rule files, and configure the running parameters of the rule engine; enter rules based on financial data and invoice data, and the rules include tax compliance rules, data logic consistency rules, and risk warning rules;
[0049] Extract inventory features based on inventory data and human resource features based on human resource data; calculate a first comprehensive human feature of operational efficiency and a second comprehensive human feature of operational efficiency according to the inventory features and the human resource features;
[0050] For the first comprehensive human characteristics of operational efficiency and the second comprehensive human characteristics of operational efficiency, perform time series expansion and data annotation based on the rules of the Drools rule engine; select a hybrid model architecture that combines a long short-term memory network and an autoencoder, divide the annotated time series dataset into a training set, a validation set, and a test set, and perform model training, evaluation, and optimization;
[0051] When there is new data update or submission in the tax return working papers, input the newly generated feature sequence into the trained hybrid model to determine whether there is a risk of tampering in the current tax return working papers. If there is a risk of tampering, give a warning to the staff.
[0052] Furthermore, after the enterprise's financial department performs the daily closing operation every day, use the scheduling function of the ETL tool to obtain the latest financial data; write an associated query statement to integrate the data in the invoice management system, associate the detail table from the invoice main table to obtain invoice data; for the inventory system database, use a query statement to associate and query the real-time status of the inventory from the inbound table, outbound table, and inventory balance table to obtain inventory data; query information from the human resource management system database to obtain human resource data;
[0053] Perform data cleaning and data transformation on the extracted financial data, invoice data, inventory data, and human resource data, create a database based on tax return requirements, construct dimension tables and fact tables in it, and realize the integration of data from different data sources through the primary key-foreign key association mechanism; associate the financial data with the invoice data, the inventory data with the financial data, and the human resource data with the financial data.
[0054] Furthermore, set aside an area for Drools-related development and management, and divide it into multiple subdirectories according to functions, including: a rule library directory, a model directory, and a test directory; among them, the rule library directory is used to store rule files related to tax returns, which define constraint conditions and behavior logics; the model directory stores domain model objects related to tax returns, which are abstractions and simplifications of the underlying data structure and are presented in the form of Java classes; the test directory writes unit test and integration test cases, inputs different test data, verifies whether the rules can be triggered correctly, and outputs the expected results;
[0055] In the "kmodule.xml" in the Drools workspace, define various configurations of the knowledge session, which is used to ensure that different types of sessions meet different business requirements; perform rule loading and update parameters, and set execution priorities, conflict resolution parameters, and performance optimization parameters.
[0056] According to the conventional technology in this field, the revenue recognition rules in financial data are set; for enterprises adopting different sales models, the rules for when and at what amount the revenue under each model should be recognized in the financial statements are clarified; based on invoice data, the rules for issuing, authenticating and deducting special VAT invoices and ordinary invoices are standardized;
[0057] Establish a logical connection between financial income data and invoice data. The rule is set that the main business income in the financial statements should fluctuate within a predefined difference range with the tax-inclusive sales amount issued by invoices in the same period. The difference range is determined through statistical analysis of historical data; set rules for the logical coherence of invoice data;
[0058] For risk warning rules, we will conduct warnings on abnormal data fluctuations, tax rate deviations and related-party transaction risks.
[0059] Based on inventory data, inventory features are extracted, including dynamic features, trend features and correlation features. The dynamic features include inventory turnover rate and inventory change rate. The trend features include moving average and seasonal index. The correlation feature is the correlation between inventory and sales. Based on human resource data, human resource features are extracted, including salary structure features and personnel change features. The salary structure features include salary composition ratio and salary level distribution. The personnel change features include employee turnover rate and new employee ratio.
[0060] In this example, we study a medium-sized manufacturing company that mainly produces seasonal consumer goods, and the peak sales season for products is concentrated in the third and fourth quarters of each year. The company's relevant data for the past 20 quarters (m=20) has been fully collected, covering inventory turnover rate, inventory change rate, operating efficiency index (here, return on assets ROA is used as an indicator to measure operating efficiency), moving average, seasonal index, proportion of each component of the salary structure, employee turnover rate, proportion of new employees, and inventory-sales correlation.
[0061] For the first weight, the inventory turnover rate (ITR) data for the past 20 quarters are as follows:
[0062] [1.2,1.3,1.1,0.9,1.4,1.5,1.3,1.2,0.8,0.7,1.6,1.8,2.0,1.9,1.7,1.5,1.3,1.1,1.0,1.2];
[0063] Inventory Change Rate (IVR) data are as follows:
[0064] [0.1, -0.05, 0.08, -0.12, 0.15, 0.2, -0.08, 0.05, -0.15, -0.2, 0.25, 0.3, -0.1, 0.05, -0.05, 0.1, -0.08, 0.03, -0.05, 0.08];
[0065] The corresponding return on assets (ROA) data is as follows:
[0066] [0.1, 0.12, 0.08, 0.06, 0.15, 0.18, 0.13, 0.1, 0.05, 0.04, 0.2, 0.22, 0.18, 0.16, 0.14, 0.12, 0.1, 0.08, 0.06, 0.08];
[0067] The calculated weight of inventory turnover is , and the weight of inventory change rate is 1 - 0.625 = 0.375.
[0068] For the second weight, 5 senior inventory management experts and operation experts within the enterprise were invited to score (on a scale of 1 - 10) the relative importance of the moving average value (MAV) and the seasonal index (SI) in reflecting inventory trends. The scoring results are as follows: Expert 1: MAV - 8 points, SI - 6 points; Expert 2: MAV - 7 points, SI - 5 points; Expert 3: MAV - 9 points, SI - 7 points; Expert 4: MAV - 8 points, SI - 6 points; Expert 5: MAV - 7 points, SI - 5 points. The calculated weight of the moving average value is 0.557, and the weight of the seasonal index is 0.443.
[0069] Similarly, for the fourth weight, the description of the experts' scoring is omitted, and the calculated weight of the employee turnover rate is 0.603, and the weight of the proportion of new employees is 0.397.
[0070] For the third weight, the target layer: determine the first comprehensive human characteristics of operational efficiency; the criterion layer: inventory operation efficiency (O), inventory trend (T), salary structure (S); 7 experts (including those in the fields of finance, operation, human resources, etc.) were used to make pairwise comparisons of the relative importance among the factors in the criterion layer to construct the judgment matrix A. The mathematical software MATLAB was used to calculate the maximum eigenvalue of the judgment matrix and its corresponding eigenvector. After calculation, the normalized eigenvector corresponding to the maximum eigenvalue was obtained. The weight of the inventory operation efficiency value is 0.54, the weight of the comprehensive inventory trend value is 0.16, and the weight of the comprehensive salary structure value is 0.30.
[0071] Further, for the dynamic characteristics of inventory, the inventory turnover rate and the inventory change rate are standardized and summed up by weighted according to the first weight, where the first weight is determined according to the importance of the two parameters in the dynamic characteristics of inventory to the operation efficiency, and the inventory operation efficiency value is obtained; for the trend characteristics of inventory, the moving average value and the seasonal index are standardized and summed up by weighted according to the second weight, where the second weight is assigned according to the relative importance of the two parameters in the trend characteristics of inventory in reflecting the inventory trend, and the comprehensive inventory trend value is obtained; for the salary structure characteristics, the proportion of basic salary, performance salary, bonus and welfare in the total salary is summed up by weighted according to their respective proportions in the total salary, and the comprehensive salary structure value is obtained; the inventory operation efficiency value, the comprehensive inventory trend value and the comprehensive salary structure value are summed up by weighted according to the third weight, where the third weight is determined through actual business verification, and the first operation efficiency human comprehensive characteristic is obtained as 0.07;
[0072] The correlation degree between inventory and sales is standardized to obtain the standardized correlation degree value; for the personnel change characteristics, the employee turnover rate and the proportion of new employees are standardized and summed up by weighted according to the fourth weight, where the fourth weight is determined according to the importance of the two parameters in the personnel change characteristics to the operation efficiency, and the comprehensive personnel change value is obtained; the standardized correlation degree value and the comprehensive personnel change value are added together with a weight of 0.5 each, and the second operation efficiency human comprehensive characteristic is obtained as 0.33.
[0073] Further, multi-period inventory data and human resource data are retrieved from the enterprise's database, and the collected data are arranged in chronological order to form a structured time series data set; for the first operation efficiency human comprehensive characteristic and the second operation efficiency human comprehensive characteristic, the sliding window statistic is calculated; using seasonal decomposition, the time series of each comprehensive characteristic is decomposed into a trend component, a seasonal component and a residual component;
[0074] When the data of the operation efficiency human comprehensive characteristic after the time series is extended flows into the Drools rule engine, the engine starts rule matching in sequence; each data point is checked for compliance rules one by one, and when a certain compliance rule condition is met, the corresponding compliance risk label is marked on the data; the risk warning rule matching is carried out, and when triggered, the risk warning label is superimposed and marked; the data logic consistency rule is executed, and if the condition is met, the data logic consistency label is added.
[0075] Further, based on the labeled time series dataset, the dataset is split into a training set, a validation set, and a test set according to a ratio; in the hybrid model architecture, the long short-term memory network processes the input time series data sequentially. Each long short-term memory network receives the enterprise operation feature vector at the current time step, including the first operation efficiency human comprehensive feature, the second operation efficiency human comprehensive feature, and each sub-feature included therein; combined with the hidden state of the previous time step, through the input gate, the forget gate, and the output gate, it determines which information needs to be retained, updated, or forgotten; the autoencoder is used to learn the feature representation of the data and consists of an encoder and a decoder; the encoder compresses the high-dimensional input data into a low-dimensional coding space and extracts the key features of the data; during the training process, the model, based on the labeled information, aims to minimize the loss function between the prediction result and the true label, and continuously adjusts the weight parameters of the model.
[0076] At each stage of model training and after the training is finally completed, the model is evaluated; according to the results of model evaluation, the hyperparameters of the long short-term memory network and the autoencoder are adjusted.
[0077] The newly generated feature sequence is input into the hybrid model combined with the pre-trained long short-term memory network and autoencoder. The model, based on the normal and abnormal sample feature patterns accumulated during the training stage, compares the currently input feature sequence, comprehensively judges whether there is a risk of tampering in the current tax return draft, and gives the corresponding risk probability value.
[0078] For a high risk level, a pop-up window appears on the tax return operation interface. The content of the pop-up window is a risk reminder, accompanied by a key abnormal data comparison chart, showing the difference between the current data and the normal range, and giving the recommended verification direction; for a medium risk level, the warning information is sent to the relevant staff in the form of an email; for a low risk level, the warning information is published in the system internal notification bar.
[0079] According to another embodiment of the present invention, as Figure 2 shown in the system structure diagram of a tax return draft management system based on a rule engine, a tax return draft management system based on a rule engine includes:
[0080] Data collection and integration module: including: a data collection unit and a data preprocessing unit; wherein, the data collection unit develops data interfaces to connect to the data sources within the enterprise, extracts tax return data, including financial data, invoice data, inventory data, and human resource data, at a predetermined frequency through an ETL tool, and the data integration unit performs data preprocessing and integration.
[0081] Rule engine building module: including: Drools workspace building unit and rule engine configuration unit; among them, the Drools workspace building unit builds a Drools workspace, creates a rule library directory for storing rule files, and configures the running parameters of the rule engine; the rule engine configuration unit inputs rules based on financial data and invoice data, and the rules include tax compliance rules, data logic consistency rules, and risk warning rules;
[0082] Feature engineering module: including: inventory feature extraction unit, human resource feature extraction unit, and comprehensive feature calculation unit; among them, the inventory feature extraction unit extracts inventory features based on inventory data, and the human resource feature extraction unit extracts human resource features based on human resource data; the comprehensive feature calculation unit calculates the first operation efficiency human comprehensive feature and the second operation efficiency human comprehensive feature according to the inventory features and the human resource features;
[0083] Time series processing and model training module: including: time series extension unit and model training unit; among them, the time series extension unit performs time series extension on the first operation efficiency human comprehensive feature and the second operation efficiency human comprehensive feature, and performs data annotation based on the rules of the Drools rule engine; the model training unit selects a hybrid model architecture combining long short-term memory network and autoencoder, divides the labeled time series data set into training set, validation set, and test set, and performs model training, evaluation, and optimization;
[0084] In this embodiment, the sliding window size is set to 4 quarters, and the sliding window statistics are calculated for the first operation efficiency human comprehensive feature FEATURE1 and the second operation efficiency human comprehensive feature FEATURE2. Taking FEATURE1 as an example, the mean value of the first window (quarters 1-4) is calculated as: 0.1875, the standard deviation is: 0.06, the maximum value is 0.3, and the minimum value is 0.1. Similarly, these statistics are calculated for each window and recorded as extended features.
[0085] The time series decomposition method is used to decompose FEATURE1 and FEATURE2. Taking FEATURE1 as an example, through the seasonal decomposition algorithm, the trend component TREND1 shows an increasing trend year by year, the seasonal component SEASON1 shows certain seasonal fluctuations, the feature values in the peak season are higher, and those in the off-season are lower, and the residual component RESIDUAL1 contains some random fluctuations that cannot be explained by the trend and season. Similar decomposition operations are also performed on FEATURE2. Data annotation is performed based on the rules of the Drools rule engine.
[0086] The labeled time series dataset is divided into a training set, a validation set, and a test set in the ratio of 70%, 15%, and 15%. The training set contains approximately 8 quarters of data, and the validation set and the test set each contain approximately 2 quarters of data. A hybrid model architecture combining a long short-term memory network (LSTM) and an autoencoder is selected. For the LSTM part, 2 hidden layers are set, with 64 neurons in each layer, to handle the long-term and short-term dependencies of time series data. For the autoencoder part, the encoding dimension is set to 8. The high-dimensional input data is compressed into a low-dimensional encoding space by the encoder and then reconstructed by the decoder to learn the key features of the data. The training set data is input into the model for training, aiming to minimize the cross-entropy loss (for classification annotation tasks) and the mean squared error loss (for prediction tasks of some numerical features) between the prediction results and the true annotations. The stochastic gradient descent optimization algorithm is used, and the learning rate is set to 0.001, and 100 epochs are trained.
[0087] After multiple trainings and optimizations, the final performance of the model on the test set is as follows: for the classification tasks of various risk labels, the average accuracy reaches 0.85, the recall rate reaches 0.8, and the F1 value reaches 0.82, which can accurately identify potential operation risks and compliance issues. For the prediction tasks of numerical features, such as predicting the values of FEATURE1 and FEATURE2 in the future quarter, the RMSE is reduced to 0.03, and the MAE is reduced to 0.02, indicating that the model can better fit the data trend and has a certain prediction ability for future operation conditions.
[0088] Risk warning module: including: a risk judgment unit and a warning push unit; among them, when there is new data update or submission in the tax return draft, the risk judgment unit inputs the newly generated feature sequence into the trained hybrid model to judge whether there is a risk of tampering in the current tax return draft. If there is a risk of tampering, the warning push unit warns the staff.
[0089] For those skilled in the art, it is obvious that the present invention is not limited to the details of the above exemplary embodiments, and can be implemented in other specific forms without departing from the spirit or basic characteristics of the present invention. Therefore, from any point of view, the embodiments should be regarded as exemplary and non-limiting. The scope of the present invention is defined by the appended claims rather than the above description. Therefore, all changes falling within the meaning and scope of the equivalent elements of the claims are intended to be included in the present invention. Any reference signs in the claims should not be regarded as limiting the claims involved.
Claims
1. A tax declaration manuscript management method based on rule engine, characterized in that: The method comprises the following steps: Develop data interfaces to connect to the enterprise's internal data sources, extract tax return data at a predetermined frequency through ETL tools, including financial data, invoice data, inventory data and human resources data, and perform data preprocessing and integration; Build a Drools workspace, create a rule base directory for storing rule files, and configure the operating parameters of the rule engine; enter rules based on financial data and invoice data, including tax compliance rules, data logic consistency rules, and risk warning rules; Extracting inventory features based on inventory data, and extracting human resource features based on human resource data; calculating a first operational efficiency human resource comprehensive feature and a second operational efficiency human resource comprehensive feature based on the inventory features and the human resource features; Extract inventory features based on inventory data, including dynamic features, trend features and correlation features. The dynamic features include inventory turnover rate and inventory change rate. The trend features include moving average and seasonal index. The correlation feature is the correlation between inventory and sales. Extract human resource features based on human resource data, including salary structure features and personnel change features. The salary structure features include salary composition ratio and salary level distribution. The personnel change features include employee turnover rate and new employee ratio. For inventory dynamic characteristics, the inventory turnover rate and inventory change rate are standardized and weighted and summed according to the first weight, which is determined according to the importance of the two parameters in the inventory dynamic characteristics on the operating efficiency, to obtain the inventory operating efficiency value; for inventory trend characteristics, the moving average and the seasonal index are standardized and weighted and summed according to the second weight, which is weighted according to the relative importance of the two parameters in the inventory trend characteristics in reflecting the inventory trend, to obtain the inventory trend comprehensive value; for salary structure characteristics, the basic salary ratio, performance salary ratio, bonus ratio and welfare ratio are weighted and summed according to their respective proportions in the total salary to obtain the salary structure comprehensive value; the inventory operating efficiency value, the inventory trend comprehensive value and the salary structure comprehensive value are weighted and summed according to the third weight, which is determined through actual business verification, to obtain the first operating efficiency human comprehensive characteristics; The correlation between inventory and sales is standardized to obtain a standardized correlation value; for the personnel change feature, the employee turnover rate and the new employee ratio are standardized, and weighted summed according to the fourth weight, wherein the fourth weight is determined according to the importance of the two parameters in the personnel change feature on the operational efficiency, to obtain a comprehensive value of personnel change; the standardized correlation value and the comprehensive value of personnel change are added according to the half weight to obtain the second comprehensive human resource feature of operational efficiency; For the first and second operational efficiency human comprehensive characteristics, time series expansion is performed, and data is labeled based on the rules of the Drools rule engine; a hybrid model architecture combining long short-term memory network and autoencoder is selected, and the labeled time series data set is divided into training set, validation set and test set for model training, evaluation and optimization; When new data is updated or submitted in the tax return draft, the newly generated feature sequence is input into the trained hybrid model to determine whether there is a risk of tampering in the current tax return draft. If there is a risk of tampering, the staff will be warned.
2. According to claim 1, a tax declaration manuscript management method based on rule engine is characterized by: When the company's financial department completes the daily closing operation, it uses the scheduling function of the ETL tool to obtain the latest financial data; Write associated query statements to integrate the data in the invoice management system, associate the invoice master table with the detail table, and obtain the invoice data; For the inventory system database, use query statements to query the real-time status of inventory from the inventory entry table, inventory exit table and inventory balance table to obtain inventory data; Query information from the human resource management system database to obtain human resource data; The extracted financial data, invoice data, inventory data and human resources data are cleaned and converted, and a database based on tax declaration requirements is created. Dimension tables and fact tables are constructed in it, and data from different data sources are integrated through the primary key-foreign key association mechanism; financial data is associated with invoice data, inventory data is associated with financial data, and human resources data is associated with financial data.
3. The tax declaration manuscript management method based on rule engine according to claim 1 is characterized by: Open up an area for Drools-related development and management, and divide it into multiple subdirectories according to functions, including: rule base directory, model directory and test directory; the rule base directory is used to store rule files related to tax declaration, defining constraints and behavior logic; the model directory stores domain model objects related to tax declaration, which is an abstract simplification of the underlying data structure; the test directory is used to write unit test and integration test cases, input different test data, verify whether the rules can be triggered correctly, and output the expected results; In the "kmodule.xml" in the Drools workspace, define various configurations of knowledge sessions, which are used to ensure that different types of sessions meet different business needs; perform rule loading and update parameters, set execution priority, conflict resolution parameters, and performance optimization parameters.
4. The tax declaration manuscript management method based on rule engine according to claim 3 is characterized by: Set up rules for revenue recognition in financial data; for enterprises adopting different sales models, clarify the rules for when and at what amount revenue should be recognized in the financial statements under each model; standardize the rules for issuing, certifying and deducting special VAT invoices and ordinary invoices based on invoice data; Establish a logical connection between financial income data and invoice data. The rule is set that the main business income in the financial statements should fluctuate within a predefined difference range with the tax-inclusive sales amount issued by invoices during the same period. The difference range is determined through statistical analysis of historical data. Set rules for the logical coherence of invoice data. For risk warning rules, warnings are issued for abnormal data fluctuations, tax rate deviations and related-party transaction risks.
5. The tax declaration manuscript management method based on rule engine according to claim 1 is characterized by: Use historical data analysis and regression method to determine the first weight and collect the inventory turnover rate (ITR) of the company in the past m quarters. j , Inventory Change Rate IVR j And the corresponding operational efficiency indicator OE j ; Establish multiple linear regression model: OE j =α+β1*ITR j +β2*IVR j +∈ j , where α is the intercept, β1 is the first regression coefficient, β2 is the second regression coefficient, ∈ j is the random error term; By performing regression analysis on historical data, we can obtain the estimated values of regression coefficients β1 and β2. The calculation formula for the inventory turnover rate weight in the first weight is: The weight calculation formula for inventory change rate is: IVR =1-w ITR ; The second weight and the fourth weight are determined by using an expert scoring method; The third weight is determined by using the hierarchical analysis method; a hierarchical structure model is constructed, wherein the target layer is to determine the comprehensive human resource characteristics of the first operating efficiency, the criterion layer includes inventory operating efficiency, inventory trend and salary structure, and the scheme layer is to allocate weights to each criterion; experts compare the relative importance of each factor in the criterion layer in pairs to construct a judgment matrix; The maximum eigenvalue of the judgment matrix and its corresponding eigenvector are calculated, and the eigenvector is normalized. The normalized vector element obtained is the third weight.
6. The tax declaration manuscript management method based on rule engine according to claim 1 is characterized by: Obtain multiple periods of inventory data and human resource data from the company's database, and arrange the collected data in chronological order to form a structured time series data set; Calculate sliding window statistics for the first operational efficiency manpower comprehensive feature and the second operational efficiency manpower comprehensive feature; Using seasonal decomposition, the time series of each composite feature is split into trend component, seasonal component and residual component; When the time series expanded operational efficiency manpower comprehensive feature data flows into the Drools rule engine, the engine starts rule matching in sequence; the compliance rule verification is performed on each data point one by one, and when a compliance rule condition is met, the corresponding compliance risk label is marked on the data; Match risk warning rules and add risk warning labels when triggered; Execute data logic consistency rules and add data logic consistency tags if they meet the conditions; Based on the annotated time series data set, the data set is divided into training set, validation set and test set according to the proportion; in the hybrid model architecture, the long short-term memory network processes the input time series data in sequence, and each long short-term memory network receives the enterprise operation feature vector of the current time step, including the first operation efficiency human comprehensive feature, the second operation efficiency human comprehensive feature and its contained sub-features; during the training process, the model continuously adjusts the weight parameters of the model based on the annotation information, with the goal of minimizing the loss function between the predicted result and the actual annotation; The model is evaluated at each stage of model training and after the training is finally completed; based on the results of the model evaluation, the hyperparameters of the long short-term memory network and the autoencoder are adjusted.
7. A tax declaration manuscript management method based on rule engine according to claim 6, characterized in that: The newly generated feature sequence is input into a hybrid model that combines a pre-trained long short-term memory network with an autoencoder. The model compares the currently input feature sequence with the normal and abnormal sample feature patterns accumulated during the training phase, comprehensively determines whether the current tax return draft has a risk of tampering, and gives a corresponding risk probability value. For high-risk levels, a pop-up window will pop up on the tax declaration operation interface with risk warnings, key abnormal data comparison charts, showing the difference between the current data and the normal range, and providing suggested verification directions; for medium-risk levels, warning information will be sent to relevant staff via email; For low risk levels, warning information is published in the system’s internal notification bar.
8. A tax declaration draft management system based on a rule engine, using a tax declaration draft management method based on a rule engine according to any one of claims 1 to 7, comprising: Data collection and integration module: including: data collection unit and data preprocessing unit; the data collection unit develops data interface, connects to the data source within the enterprise, extracts tax declaration data at a predetermined frequency through ETL tools, including financial data, invoice data, inventory data and human resources data, and the data integration unit performs data preprocessing and integration; Rule engine building module: including: Drools workspace building unit and rule engine configuration unit; wherein, the Drools workspace building unit builds the Drools workspace, creates a rule base directory for storing rule files, and configures the running parameters of the rule engine; the rule engine configuration unit enters rules based on financial data and invoice data, and the rules include tax compliance rules, data logic consistency rules, and risk warning rules; Feature engineering module: including: an inventory feature extraction unit, a human resource feature extraction unit and a comprehensive feature calculation unit; wherein the inventory feature extraction unit extracts inventory features based on inventory data, and the human resource feature extraction unit extracts human resource features based on human resource data; the comprehensive feature calculation unit calculates a first operational efficiency human resource comprehensive feature and a second operational efficiency human resource comprehensive feature according to the inventory feature and the human resource feature; Time series processing and model training module: including: time series expansion unit and model training unit; wherein the time series expansion unit performs time series expansion for the first operational efficiency human comprehensive characteristics and the second operational efficiency human comprehensive characteristics, and performs data annotation based on the rules of the Drools rule engine; the model training unit uses a hybrid model architecture combining a long short-term memory network and an autoencoder, divides the annotated time series data set into a training set, a validation set and a test set, and performs model training, evaluation and optimization; Risk warning module: including: risk judgment unit and warning push unit; among them, when there is new data update or submission of tax declaration draft, the risk judgment unit inputs the newly generated feature sequence into the trained hybrid model to judge whether there is tampering risk in the current tax declaration draft. If there is tampering risk, the warning push unit will warn the staff.
Citation Information
Patent Citations
Tax calculation method and system based on enterprise custom rule
CN116051299A
System and method for tax declaration and risk early warning
CN117934183A