Electronic Taxation Bureau Docking Method Based on General Report Technology and Big Data Knowledge Base
The integration of AI-generated universal tax forms and Flink stream processing addresses inefficiencies in tax reporting software, enhancing compliance and security while standardizing data across regions.
Patent Information
- Application Number
- CN202111354853.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-11-16
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2041-11-16
AI Technical Summary
The existing financial software has simple functions and is prone to data loss. The high-end software is expensive and the declaration forms of electronic tax bureaus in various provinces and cities are not unified, resulting in the large workload and low efficiency of the electronic tax bureaus to change and upgrade.
Using a method based on general reporting technology and big data knowledge base, general reports are generated through artificial intelligence machine learning, combined with Flink streaming calculation, data preprocessing and calculation are realized, and the tax declaration form formats in different regions are unified.
It has improved the work efficiency of the Electronic Taxation Bureau's changes and upgrades, solved the problems of diversity of bureau-end interface access and regional differentiation and unified standardization, and improved the efficiency and security of agency accounting and tax filing.
Smart Images

Figure CN113988038B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data processing, and specifically relates to a method for docking an electronic tax bureau based on general report technology and a big data knowledge base. Background Art
[0002] Currently, various traditional bookkeeping and tax filing software, such as financial software like Yunzhangfang and Inspur Cloud, have simple functions and are generally stand-alone versions, and data is prone to loss. Although high-end financial software provides cloud computing technology and mobile application technology, the price is relatively high. In addition, the tax return forms of electronic tax bureaus in each province and city are mostly the same but not completely unified. When facing the change and upgrade of the electronic tax bureau, the change workload is large and it is difficult to complete efficiently.
[0003] Therefore, there is a need for a general report service that is efficient, secure, and powerful to improve the efficiency of agency bookkeeping and tax filing, solve the problem of improving the efficiency of the change and upgrade work of the electronic tax bureau, and the unified standardization problems of the diversity of bureau-side interface access and regional differences. Summary of the Invention
[0004] In view of the defects in the prior art, the present invention provides a method for docking an electronic tax bureau based on general report technology and a big data knowledge base.
[0005] The method for docking an electronic tax bureau based on general report technology and a big data knowledge base includes the following steps:
[0006] S1: Obtain the tax return form header data, and generate a general report through artificial intelligence machine learning. Among them, the tax return form header data includes the headers of different tax return forms in each province and city, and then execute step S2;
[0007] S2: Obtain the tax declaration data of the target enterprise, preprocess the tax declaration data, and then execute step S3;
[0008] S3: Calculate the preprocessed tax declaration data by using the Flink streaming calculation method, and apply the calculation result to the general report to obtain the target tax return form.
[0009] Preferably, in step S1, the method for generating a general report through artificial intelligence machine learning includes the following steps:
[0010] S11: Generate a training data set based on the tax return form header data;
[0011] S12: Separate an evaluation data set from the training data set. Among them, the headers of different tax return forms in each province and city include common headers and unique headers, and the unique headers of each province and city are used as the evaluation data set;
[0012] S13: Construct an algorithm model that meets the evaluation criteria, and use the training set to optimize and train the algorithm model to obtain an optimized algorithm model;
[0013] S14: Use the evaluation data set to verify the optimized algorithm model; serialize the optimized algorithm model as a general report.
[0014] Preferably, in step S2, the method for obtaining the tax declaration data of the target enterprise includes the following steps:
[0015] S31: Connect the target enterprise to the general report; after the connection is completed, obtain the first tax declaration data of the target enterprise through the API interface;
[0016] S32: Connect the general report to the electronic tax bureau; after the connection is completed, obtain the second tax declaration data of the target enterprise through the API interface;
[0017] S33: Summarize the first tax declaration data and the second tax declaration data to obtain the tax declaration data of the target enterprise.
[0018] Preferably, in step S31, the method for connecting the target enterprise to the general report is: the target enterprise signs an online or offline agreement with a third party, and verifies the identity of the target enterprise through API interface calls. Only when the verification is successful is it allowed to call the first tax data of the target enterprise, where the third party is configured with a general template.
[0019] Preferably, in step S32, the method for connecting the general report to the electronic tax bureau is: the third party signs an access agreement with the electronic tax bureau, and verifies the identity of the target enterprise through API interface calls. Only when the verification is successful is it allowed to call the second tax data of the target enterprise.
[0020] Preferably, in step S2, the method for preprocessing the tax declaration data includes the following steps:
[0021] S61: Perform data cleaning on the error items, missing items, duplicate items and redundant items in the tax declaration data;
[0022] S62: Perform feature selection on the cleaned tax declaration data using the coefficients of the regression model;
[0023] S63: Convert the tax declaration data after feature selection into a unified format.
[0024] Preferably, in step S3, the method for calculating the preprocessed tax declaration data using the Flink streaming calculation method and applying the calculation result to the general report to obtain the target tax return form includes the following steps:
[0025] Convert the preprocessed tax declaration data into a tax data stream;
[0026] Fill the tax data stream into the corresponding positions of the general report through the configuration document to obtain the target tax return form, and complete the docking between the target enterprise and the e-tax bureau.
[0027] The beneficial effects of the present invention are reflected in that: the e-tax bureau docking method based on the general report technology and the big data knowledge base proposed by the present invention collects enterprise invoice data, financial data, initial declaration data at the tax bureau end, etc.; performs data preprocessing on the collected data; for the preprocessed data, generates a general report through artificial intelligence machine learning; uses the Flink streaming computing method for data calculation, and applies the calculation results to the general report. It effectively solves the problem of improving the efficiency of the e-tax bureau change and upgrade work and the unified standardization problems of the diversity of bureau-side interface access, the diversity of requirements, and regional differences. Description of the Drawings
[0028] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for use in the description of the specific embodiments or the prior art. In all the drawings, similar elements or parts are generally identified by similar reference numerals. In the drawings, the elements or parts are not necessarily drawn to scale.
[0029] Figure 1 It is a flowchart of the e-tax bureau docking method based on the general report technology and the big data knowledge base of the present invention. Detailed Embodiments
[0030] The following will describe in detail the embodiments of the technical solutions of the present invention with reference to the drawings. The following embodiments are only used to more clearly illustrate the technical solutions of the present invention, so they are only examples and cannot be used to limit the protection scope of the present invention.
[0031] It should be noted that unless otherwise specified, the technical terms or scientific terms used in this application should have the ordinary meaning understood by those skilled in the art to which the present invention belongs.
[0032] As Figure 1 shown, a method for docking an e-tax bureau based on the general report technology and the big data knowledge base provided by an embodiment of the present invention includes the following steps:
[0033] S1: Obtain the tax return form header data and generate a general report through artificial intelligence machine learning. Among them, the tax return form header data includes the headers of different tax return forms for each province and city, and perform step S2;
[0034] Specifically, in this embodiment, the process of generating a general report through artificial intelligence machine learning includes:
[0035] Data preparation mainly refers to the header data of various declaration forms, and pre-collects the headers of different tax declaration forms in each province and city.
[0036] The evaluation algorithm is mainly to find the best algorithm subset. Specifically, it includes:
[0037] ①Separate the evaluation data set to facilitate model verification. Use the collected header data (including the same and different header data. There are common parts and unique parts in the declaration form headers of each province and city. The unique parts are used as specific headers) as a training data set, and use K-fold cross-validation to separate the data set. Use the different header data as an evaluation data set to achieve separation.
[0038] ②Define the model evaluation criteria for evaluating the algorithm model. Specifically, the evaluation criteria include whether the selected data headers from the evaluation data set are accurate.
[0039] ③According to the evaluation criteria, compare the accuracy of the evaluation algorithm to obtain an algorithm with sufficient accuracy.
[0040] Optimize the model, adjust the parameters for each algorithm to obtain the best results, and use the ensemble algorithm to improve the accuracy of the algorithm model.
[0041] Result deployment: Verify the optimized model through the verification data set, generate the model through the entire data set, and serialize the model as a general report.
[0042] S2: Obtain the tax declaration data of the target enterprise, preprocess the tax declaration data, and execute step S3;
[0043] Among them, the tax declaration data includes the first tax declaration data and the second tax declaration data. The first tax declaration data includes data such as enterprise invoice data and financial data, and the second tax declaration data includes the initial declaration data at the tax bureau end.
[0044] It should be noted that the method for obtaining the tax declaration data of the target enterprise includes the following steps:
[0045] S31: Connect the target enterprise with the general report; after the connection is completed, obtain the first tax declaration data of the target enterprise through the API interface.
[0046] S32: Connect the general report with the electronic tax bureau. After the connection is completed, obtain the second tax declaration data of the target enterprise through the API interface.
[0047] S32: Aggregate the first tax declaration data and the second tax declaration data to obtain the tax declaration data of the target enterprise.
[0048] It should be noted that the method for connecting the target enterprise with the general report is as follows: The target enterprise signs an online or offline agreement with a third party, and verifies the identity of the target enterprise through API interface calls. Only when the verification is successful is it allowed to call the first tax data of the target enterprise. Among them, the third party is configured with a general template.
[0049] It should be noted that in step S32, the method for connecting the general report with the electronic tax bureau is as follows: The third party signs an access agreement with the electronic tax bureau, and verifies the identity of the target enterprise through API interface calls. Only when the verification is successful is it allowed to call the second tax data of the target enterprise.
[0050] It should be noted that the preprocessing of the tax declaration data includes:
[0051] For the collected data, data cleaning is first performed, that is, data processing is carried out on data redundancy, data duplication, and data inconsistency. Processing method: If highly suspected samples are adjacent, a sliding window can be used for comparison. In order to make similar records adjacent, a hash key can be generated for each record and sorted according to the key.
[0052] After data cleaning, feature selection is carried out using the coefficients of the regression model. The more important the feature, the larger the corresponding coefficient in the model, and the closer the coefficient corresponding to the feature that is less relevant to the output variable is to 0. Exclude the feature attributes with small coefficients according to the coefficient size, and select the feature attributes with large coefficients.
[0053] After feature selection, data conversion is carried out. For non-numerical types, category conversion is performed, that is, non-numerical types are converted into numerical types. For nominal types, one-hot encoding is used, and for ordinal types, ordinal encoding is used. This is convenient for subsequent processing.
[0054] In this embodiment, enterprise invoice data and financial data are obtained through enterprise financial software, accounting software, etc., and the initial declaration data at the tax bureau end is obtained through access to the electronic tax bureau. The collected data is preprocessed to ensure the authenticity of the business and the accuracy of the data.
[0055] S3: The preprocessed tax declaration data is calculated using the Flink streaming calculation method, and the calculation results are applied to the general report to obtain the target tax return form.
[0056] The tax declaration data source continuously generates data to form a stream, generates a new stream through calculation, and continuously updates the target data source to achieve real-time update of the result data. According to the preset trigger conditions, data transmission and corresponding calculations are carried out according to the trigger conditions. Among them, the preset trigger conditions specifically include:
[0057] When a new data stream is generated by the data source, the data stream is divided into windows and the windows are aggregated. Windows are generated according to time, and whenever the sliding step size is met, a calculation is performed on the windows to generate a new stream.
[0058] Through the configuration document, the data stream is corresponded with the general report to form a mapping relationship comparison table. In the configuration document, the mapping relationship comparison table formed in advance according to the data characteristics and types and the general report is: [characteristic (type), location]. When the new data stream arrives, it is automatically filled into the corresponding location of the general report through the settings of the configuration document.
[0059] In summary, the present invention generates a general report through artificial intelligence machine learning; adopts the Flink streaming computing method for data calculation, and applies the calculation result to the general report. It effectively solves the problem of improving the efficiency of the change and upgrade work of the electronic tax bureau, as well as the unified standardization problems of the diversity of bureau-side interface access, the diversity of requirements, and regional differences.
[0060] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: it is still possible to modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present invention, and they should all be covered by the scope of the claims and the description of the present invention.
Claims
1. An electronic tax bureau docking method based on general report technology and big data knowledge base, characterized in that It includes the following steps: S1: Obtain the tax return header data, and generate a general report through artificial intelligence machine learning. Among them, the tax return header data includes the headers of different tax return forms in each province and city, and execute step S2; S2: Obtain the tax declaration data of the target enterprise, preprocess the tax declaration data, and execute step S3; S3: Calculate the preprocessed tax declaration data using the Flink streaming computing method, and apply the calculation result to the general report to obtain the target tax return form; In step S1, the method of generating a general report through artificial intelligence machine learning includes the following steps: S11: Generate a training data set based on the tax return header data; S12: Separate the evaluation data set from the training data set. Among them, the headers of different tax return forms in each province and city include common headers and unique headers, and the unique headers of each province and city are used as the evaluation data set; S13: Construct an algorithm model that meets the evaluation criteria, and optimize and train the algorithm model using the training data set to obtain an optimized algorithm model; S14: Verify the optimized algorithm model using the evaluation data set; serialize the optimized algorithm model as a general report; In step S2, the method of obtaining the tax declaration data of the target enterprise includes the following steps: S31: Connect the target enterprise with the general report; after the connection is completed, obtain the first tax declaration data of the target enterprise through the API interface; S32: Connect the general report with the electronic tax bureau. After the connection is completed, obtain the second tax declaration data of the target enterprise through the API interface; S33: Summarize the first tax declaration data and the second tax declaration data to obtain the tax declaration data of the target enterprise; In step S3, the method of calculating the preprocessed tax declaration data using the Flink streaming computing method and applying the calculation result to the general report to obtain the target tax return form includes the following steps: Convert the preprocessed tax declaration data into a tax data stream; Fill the tax data stream into the corresponding position of the general report through the configuration document to obtain the target tax return form, and complete the connection between the target enterprise and the electronic tax bureau; Through the configuration document, the data stream and the general report are corresponded to form a mapping relationship comparison table. The mapping relationship comparison table between the data characteristics and types and the general report is pre-configured in the configuration document as: [characteristic (type), position]. When a new data stream arrives, it is automatically filled into the corresponding position of the general report through the settings of the configuration document.
2. The method for the connection of the electronic tax bureau based on the general report technology and the big data knowledge base according to claim 1, wherein, In step S31, the method of connecting the target enterprise with the general report is: the target enterprise signs an online or offline agreement with a third party, and verifies the identity of the target enterprise through API interface calls. Only when the verification is successful is it allowed to call the first tax data of the target enterprise. Among them, the third party is configured with a general template.
3. The method for the e-tax bureau docking based on the general report technology and the big data knowledge base according to claim 2, wherein, In step S32, the method of connecting the general report with the electronic tax bureau is: the third party signs an access agreement for the electronic tax bureau with the electronic tax bureau, and verifies the identity of the target enterprise through API interface calls. Only when the verification is successful is it allowed to call the second tax data of the target enterprise.
4. The method for connecting the electronic tax bureau based on the general report technology and the big data knowledge base according to claim 2, wherein In step S2, the method for preprocessing the tax declaration data includes the following steps: S61: Perform data cleaning on the error items, missing items, duplicate items, and redundant items in the tax declaration data; S62: Perform feature selection on the cleaned tax declaration data using the coefficients of the regression model; S63: Convert the tax declaration data after feature selection into a unified format.
Citation Information
Patent Citations
System and method for tax declaration
CN111192125A
An intelligent engineering cost data analysis method and system based on deep learning
CN113010503A