Financial and tax data processing method, system, device and storage medium
By performing thematic domain summary, dimensional modeling and personalized statistics of fiscal and tax data of small and medium-sized enterprises, the problem of incomplete data analysis and lack of a unified perspective in fiscal and tax data processing is solved, and the rapid cleaning and efficient use of data is achieved.
Patent Information
- Application Number
- CN202310631256.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-05-30
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2043-05-30
AI Technical Summary
Small and medium-sized enterprises face the incomplete comprehensiveness of semi-structured/unstructured data analysis in fiscal and taxation data processing, and lack of a unified global perspective on data definition, data classification, data subject domain division and data model maintenance, resulting in customers being unable to quickly retrieve their own required data.
By obtaining the original fiscal and taxation data, inputting the preset topic domain induction model for topic domain division, dimensional modeling, business modeling and topic modeling form an aggregated data table, and personalized data statistics are performed based on the aggregated data table.
It realizes rapid cleaning and unified processing of fiscal and tax data, and forms an aggregated data table to improve the reusability of public indicators, reduces duplicate processing work in personalized data statistics, shortens query response time, and improves data usage efficiency.
Smart Images

Figure CN116843486B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of software technology, and in particular to a financial and tax data processing method, system, device and storage medium. Background Art
[0002] Under the background of the new normal of the economy, innovation and intelligence of fiscal and taxation technology have gradually become the consensus of the industry and a new challenge for enterprises. Behind this, a large amount of reasonably structured data is needed as the support for data analysis. The further close integration of finance and taxation with big data will promote the pace of smart tax reform.
[0003] Currently, small and medium-sized enterprises use a variety of invoicing and accounting systems, and the names of financial and tax data of the same dimension are inconsistent. The analysis of semi-structured / unstructured data for small and medium-sized enterprises is not comprehensive enough, and there is a lack of a unified global perspective on data definition, data classification, data subject domain division, and data model maintenance, so that customers cannot quickly retrieve the data they need. Summary of the invention
[0004] In view of this, the embodiments of the present invention provide a method, system, device and storage medium for processing financial and tax data, so as to provide a convenient and unified means for processing financial and tax data, and facilitate small and medium-sized enterprises to conduct financial and tax data analysis at low cost.
[0005] In a first aspect, an embodiment of the present invention provides a method for processing financial and tax data, including:
[0006] Obtain original financial and tax data;
[0007] Inputting the original financial and tax data into a preset subject domain induction model to determine at least one subject domain according to the original financial and tax data, and dividing the original financial and tax data into the subject domains;
[0008] Performing dimensional modeling, business modeling and topic modeling on the original financial and taxation data in each of the subject domains to obtain an aggregated data table;
[0009] Perform personalized data statistics according to the aggregated data table to generate personalized statistical results.
[0010] Optionally, in some embodiments, obtaining original financial and tax data includes:
[0011] The original financial and tax data of the source system can be obtained through the preset data interface, or the original financial and tax data can be obtained by identifying paper financial and tax reports through the OCR recognition model.
[0012] Optionally, in some embodiments, inputting the original financial and tax data into a preset subject domain induction model to determine at least one subject domain according to the original financial and tax data includes:
[0013] Pairing the original financial and tax data with the preset topics to calculate the topic relevance;
[0014] If the topic relevance meets the preset requirements, the preset topic is used as the topic domain;
[0015] If the subject relevance does not meet the preset requirements, the original financial and tax data is integrated, classified and analyzed to generate at least one subject domain.
[0016] Optionally, in some embodiments, before performing dimensional modeling, business modeling and topic modeling on the original financial and tax data in each subject domain to obtain an aggregated data table, the method further includes:
[0017] The original financial and tax data is cleaned of abnormal data and stored in a standardized manner.
[0018] Optionally, in some embodiments, the dimensional modeling includes:
[0019] Based on the original financial and tax data, dimension attributes are obtained through logical processing, dimension attributes are obtained through multi-table association, dimension attributes are obtained through mixed processing of different fields of a single table, and dimension attributes are obtained by parsing designated fields of a single table;
[0020] A dimension table is constructed according to the dimension attributes, and the original financial and tax data is organized into the dimension table.
[0021] Optionally, in some embodiments, after performing personalized statistics of data according to the aggregated data table to generate personalized statistical results, the method further includes:
[0022] Matching the personalized statistical results with the subject domain to calculate the relevance;
[0023] The subject domain induction model is reversely adjusted according to the relevance.
[0024] In a second aspect, an embodiment of the present invention further provides a financial and tax data processing system, the system comprising:
[0025] Data acquisition module, used to obtain original financial and tax data;
[0026] A subject domain division module, used for inputting the original financial and tax data into a preset subject domain induction model to generate at least one subject domain according to the original financial and tax data, and dividing the original financial and tax data into the subject domains;
[0027] A data aggregation module, used to perform dimensional modeling, business modeling and topic modeling on the original financial and tax data in each subject domain to obtain an aggregated data table;
[0028] The personalized statistics module is used to generate personalized statistical results by performing personalized data statistics according to the aggregated data table.
[0029] In a third aspect, an embodiment of the present invention further provides an electronic device, including a memory and a processor, wherein the memory stores a computer program that can be executed on the processor, and the processor implements the aforementioned financial and tax data processing method when executing the computer program.
[0030] In a fourth aspect, an embodiment of the present invention provides a computer-readable storage medium, wherein the storage medium stores a computer program, wherein the computer program includes program instructions, and when the program instructions are executed, the aforementioned financial and tax data processing method is implemented.
[0031] The technical solution provided by the embodiment of the present invention performs data warehouse stratification on the acquired original financial and tax data. First, the subject domain is divided with the help of the subject domain induction model, and the subject domain is based on business planning. In this way, the data required for financial and tax related businesses can be quickly cleaned out from the huge financial and tax data, and then the cleaned financial and tax data are subjected to dimensional modeling, business modeling and subject modeling. In the process of modeling, a series of aggregated data tables are formed. The aggregated data tables improve the reusability of common indicators, and can reduce repeated processing work when performing personalized data statistics at the end. The query is directly connected to the user end, which greatly shortens the query response time and makes data use more efficient. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] Figure 1 It is a flow chart of a method for processing financial and tax data in Embodiment 1 of the present invention;
[0033] Figure 2 It is a sub-flow chart of a method for processing financial and tax data in Embodiment 1 of the present invention;
[0034] Figure 3 It is a sub-flow chart of a method for processing financial and tax data in Embodiment 1 of the present invention;
[0035] Figure 4 is a flow chart of another method for processing financial and tax data in Embodiment 1 of the present invention;
[0036] Figure 5 is a schematic diagram of a financial and tax data processing system in Embodiment 2 of the present invention;
[0037] Figure 6 It is a schematic diagram of an electronic device in Embodiment 3 of the present invention. DETAILED DESCRIPTION
[0038] The present invention will be further described in detail below in conjunction with the accompanying drawings and embodiments. It is to be understood that the specific embodiments described herein are only used to explain the present invention, rather than to limit the present invention. It should also be noted that, for ease of description, only parts related to the present invention, rather than all structures, are shown in the accompanying drawings.
[0039] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as those generally understood by those skilled in the art of the technical field of the present invention. The terms used herein in the specification of the present invention are only for the purpose of describing specific embodiments and are not intended to limit the present invention. The term "and / or" used herein includes any and all combinations of one or more related listed items. In the description of the present invention, the meaning of "plurality" is at least two, such as two, three, etc., unless otherwise clearly and specifically limited.
[0040] In addition, the terms "first", "second", etc. can be used in this article to describe various directions, actions, steps or elements, etc., but these directions, actions, steps or elements are not limited by these terms. These terms are only used to distinguish a first direction, action, step or element from another direction, action, step or element. For example, without departing from the scope of the present invention, the first speed difference can be referred to as the second speed difference, and similarly, the second speed difference can be referred to as the first speed difference. Both the first application and the second application are applications, but they are not the same application. The terms "first", "second", etc. cannot be understood as indicating or implying relative importance or implicitly indicating the number of indicated technical features. Thus, the features defined as "first" and "second" can explicitly or implicitly include one or more of the features. In the description of the present invention, the meaning of "multiple" is at least two, such as two, three, etc., unless otherwise clearly and specifically defined. It should be noted that when a part is referred to as "fixed to" another part, it can be directly on the other part or there can be a central part. When a part is considered to be "connected" to another part, it can be directly connected to the other part or there may be a central part at the same time. The terms "vertical", "horizontal", "left", "right" and similar expressions used herein are for illustrative purposes only and do not represent the only implementations.
[0041] It should be mentioned before discussing the exemplary embodiments in more detail that some exemplary embodiments are described as processes or methods depicted as flow charts. Although the flow charts describe the steps as sequential processes, many of the steps therein can be implemented in parallel, concurrently or simultaneously. In addition, the order of the steps can be rearranged. The process can be terminated when its operation is completed, but can also have additional steps not included in the accompanying drawings. The process can correspond to a method, function, procedure, subroutine, subprogram, etc.
[0042] Embodiment 1
[0043] Figure 1 This is a flowchart of a method for processing financial and tax data provided in Embodiment 1 of the present invention. The method can be executed by a terminal or a server, or can be completed through interaction between a terminal and a server. This embodiment is described by taking the application of the method for processing financial and tax data to a server as an example. Specifically, Figure 1 The method shown comprises the following steps:
[0044] S110. Obtain original financial and tax data.
[0045] The original fiscal and taxation data is obtained in different ways. Currently, there are a large number of semi-structured / unstructured fiscal and taxation data in small and medium-sized enterprises. In addition, different enterprises use different invoicing and accounting systems (called source systems, such as Manyoucai, Huijietong, Manhaorong, etc.), and the taxation data interfaces provided by relevant institutions and departments are not unified, such as government big data, customs, and taxation interfaces. Therefore, in this embodiment, obtaining the original fiscal and taxation data requires the application of different types of data interfaces, and even requires data collection for some paper fiscal and taxation reports. Specifically, in this embodiment, step 110 specifically includes: obtaining the original fiscal and taxation data of the source system through a preset data interface (including interfaces for different source systems), or identifying paper fiscal and taxation reports through an OCR recognition model to obtain the original fiscal and taxation data.
[0046] OCR (Optical Character Recognition) refers to the process of analyzing and identifying the input image to obtain the text information in the image. It has a wide range of application scenarios, such as scene image text recognition, document image recognition, card recognition (such as identity card, bank card, social security card), bill recognition, etc. Its main recognition methods include three methods: 1) method based on connected domain; 2) method based on sliding window; 3) method based on deep learning. Considering that there is often some handwritten information in paper financial and tax reports, and this part is quite different, the methods based on connected domain and sliding window have a high accuracy rate in recognizing standard fonts, but the accuracy rate for handwritten text recognition is low. Therefore, in this embodiment, the OCR recognition method based on deep learning is preferably adopted. Specifically, a neural network-based OCR recognition model is pre-built, part of the paper financial and tax report image data is obtained, and the image data is marked. The OCR recognition model is trained using the marked image data. When the recognition accuracy of the OCR recognition model meets the end condition, the training is terminated to obtain the trained OCR recognition model, and the trained OCR recognition model is used to collect the original financial and tax data of the paper financial and tax reports. It should be understood that the OCR recognition model is not always the same, but is continuously trained and optimized during use. The recognition method based on deep learning uses more robust high-level semantic features, uses more data to fit more complex and more generalized models, and can continuously evolve according to actual needs to adapt to more complex scenarios with higher accuracy.
[0047] S120: Input the original financial and tax data into a preset subject domain induction model to determine at least one subject domain according to the original financial and tax data, and divide the original financial and tax data into the subject domains.
[0048] The division of subject domains is an abstract concept for integrating, classifying, analyzing and utilizing data in enterprise information systems at a higher level. Subject domains are usually a collection of closely related data topics, which are divided and abstractly classified from the perspective of business needs analysis. Each topic basically corresponds to a macro analysis field.
[0049] Optionally, in some specific embodiments, step 120 is as follows Figure 2 The specific steps shown include steps 121-123:
[0050] S121. Pair the original financial and taxation data with preset topics to calculate topic relevance.
[0051] S122: If the topic relevance meets the preset requirement, the preset topic is used as the topic domain.
[0052] S123. If the subject relevance does not meet the preset requirements, the original financial and tax data is integrated, classified and analyzed to generate at least one subject domain.
[0053] The preset subject is a pre-classified subject domain, including: basic subject, used to store enterprise basic business information, legal person information and other data; tax subject, to store company tax interface data; financial subject, to store enterprise accounting service data, such as balance sheet, income statement, account balance sheet, etc.; business subject, to store enterprise sales loan orders, customer information and other data. Of course, the above subject domains may not meet all needs. For example, in some enterprises, they need to store too many financial and tax data in a single subject domain for many business sectors. Therefore, they can be further divided into subject domains. For example, in a taxi-hailing software, it can be divided into personal moving subject domain and enterprise moving subject domain. When there is no preset subject matching a large amount of financial and tax data, this embodiment can also dynamically parse and summarize the original financial and tax data and automatically generate subject domains. The process can be to analyze the business nodes of the enterprise, cluster the financial and tax data corresponding to the business with multiple common nodes, and then abstractly describe the clustered financial and tax data to obtain a corresponding subject domain. It will not be repeated here.
[0054] S130, performing dimensional modeling, business modeling and topic modeling on the original financial and taxation data in each of the subject domains to obtain an aggregated data table.
[0055] The aggregate data table includes dimension tables, business tables and subject tables. The DIM dimension table is based on the concept of dimensional modeling to establish consistent dimensions for the entire enterprise and reduce the risk of inconsistent data calculation caliber and algorithm. The DWD business table uses the business process as the modeling driver, and builds the most fine-grained detail layer fact table based on the characteristics of each specific business process. It can be combined with the data usage characteristics of the enterprise to make some important dimensional attribute fields of the detail fact table appropriately redundant, that is, wide table processing. The DWS subject table uses the subject object of analysis as the modeling driver, and builds a summary indicator fact table of a common granularity based on the indicator requirements of upper-level applications and products, and physicalizes the model by means of wide table. The purpose of step 130 is to construct statistical indicators with standardized naming and consistent caliber, provide public indicators for the upper layer, and establish summary wide tables and detailed fact tables.
[0056] For ease of understanding, the implementation principle of the aggregate data table in this embodiment is further explained by taking the dimension table as an example. Figure 3 As shown, step 130 specifically includes steps 131-132:
[0057] S131. Based on the original financial and tax data, dimension attributes are obtained through logical processing, dimension attributes are obtained through multi-table association, dimension attributes are obtained through mixed processing of different fields of a single table, and dimension attributes are obtained by parsing designated fields of a single table.
[0058] S132. Construct a dimension table according to the dimension attributes, and organize the original financial and tax data into the dimension table.
[0059] Dimension is a measurement environment used to reflect a type of business attribute. The collection of such attributes constitutes a dimension, which can also be called an entity object. The dimension belongs to a data domain, such as the geographic dimension (including content at the country, region, province, and city levels) and the time dimension (including content at the year, quarter, month, week, and day levels).
[0060] S140: Perform personalized data statistics according to the aggregated data table to generate personalized statistical results.
[0061] Personalized data statistics actually involve data project development based on aggregated data, such as KPI reports, corporate portraits, risk prevention and control, and financial and tax warnings.
[0062] The embodiment of the present invention provides a method for processing financial and tax data, which performs data warehouse stratification on the acquired original financial and tax data. First, the subject domain is divided with the help of a subject domain induction model, and the subject domain is based on business planning. In this way, the data required for financial and tax related businesses can be quickly cleaned out from the huge financial and tax data, and then the cleaned financial and tax data are subjected to dimensional modeling, business modeling and subject modeling. In the process of modeling, a series of aggregated data tables are formed. The aggregated data tables improve the reusability of common indicators, and can reduce repeated processing work when performing personalized data statistics at the end. The query is directly connected to the user end, which greatly shortens the query response time and makes data use more efficient.
[0063] More specifically, in some embodiments, before step 130, step 100 (not shown) is further included:
[0064] S100: Clean the original financial and tax data to remove abnormal data, and perform data normalization and storage.
[0065] Before data aggregation analysis, exception handling is required. For data with special meanings, abnormal characters are removed or replaced. After extraction, the data format is standardized to design data accuracy, processing performance, and business expansion. In terms of data accuracy, the original data often contains some abnormal characters such as "." and spaces in Chinese fields such as names and abbreviations due to input errors. In terms of performance optimization, in order to improve access efficiency, invoice data and financial data are partitioned and optimized for storage with time as the partition key, and indexes are established on the relevant corresponding fields.
[0066] Optionally, in some embodiments, Figure 4 As shown, after step 140, steps 150-160 are also included:
[0067] S150: Match the personalized statistical result with the subject domain to calculate the relevance.
[0068] S160: Reversely adjust the subject domain induction model according to the relevance.
[0069] The relevance is actually calculating whether the division of the subject domain is reasonable. If it is reasonable, it means that the current division of the subject domain is conducive to the user's personalized statistical results, which has a positive impact on the efficiency and accuracy of data analysis. On the contrary, it means that the unreasonable division of the subject domain hinders the user's personalized data statistics. At this time, it is necessary to use the personalized statistical results to reversely influence the subject domain induction model, so that the divided subject domain is more reasonable and fits the user's actual needs. For example, the user needs to perform personalized statistics around the KPI report, and the subject domains divided in step 120 are all around the basic theme, which is definitely unreasonable. Therefore, it needs to be adjusted to business themes and cores. Examples are not given here one by one.
[0070] Embodiment 2
[0071] Figure 5 FIG. 3 is a schematic diagram of a financial and tax data processing system 300 provided in Embodiment 2 of the present invention. The specific structure of the system includes:
[0072] Data acquisition module 310, used to acquire original financial and tax data;
[0073] The subject domain division module 320 is used to input the original financial and tax data into a preset subject domain induction model to generate at least one subject domain according to the original financial and tax data, and divide the original financial and tax data into the subject domains;
[0074] A data aggregation module 330 is used to perform dimensional modeling, business modeling and topic modeling on the original financial and tax data in each subject domain to obtain an aggregated data table;
[0075] The personalized statistics module 340 is used to generate personalized statistics results by performing personalized statistics on the data according to the aggregated data table.
[0076] Optionally, in some embodiments, the system further comprises:
[0077] The model adjustment module is used to match the personalized statistical results with the subject domain to calculate the relevance, and reversely adjust the subject domain induction model according to the relevance.
[0078] Optionally, in some embodiments:
[0079] The data acquisition module 310 is specifically used to acquire original financial and tax data from a source system through a preset data interface, or to obtain original financial and tax data by recognizing paper financial and tax reports through an OCR recognition model.
[0080] Optionally, in some embodiments, the subject domain division module 320 is specifically used to:
[0081] Pairing the original financial and tax data with the preset topics to calculate the topic relevance;
[0082] If the topic relevance meets the preset requirements, the preset topic is used as the topic domain;
[0083] If the subject relevance does not meet the preset requirements, the original financial and tax data is integrated, classified and analyzed to generate at least one subject domain.
[0084] Optionally, in some embodiments, the data aggregation module 330 is specifically used to:
[0085] Based on the original financial and tax data, dimension attributes are obtained through logical processing, dimension attributes are obtained through multi-table association, dimension attributes are obtained through mixed processing of different fields of a single table, and dimension attributes are obtained by parsing designated fields of a single table;
[0086] A dimension table is constructed according to the dimension attributes, and the original financial and tax data is organized into the dimension table.
[0087] This embodiment further provides a financial and taxation data processing system, which performs data warehouse stratification on the acquired original financial and taxation data. First, the subject domain is divided with the help of the subject domain induction model, and the subject domain is based on business planning. In this way, the data required for financial and taxation related businesses can be quickly cleaned out from the huge amount of financial and taxation data, and then the cleaned financial and taxation data is subjected to dimensional modeling, business modeling and subject modeling. In the process of modeling, a series of aggregated data tables are formed. The aggregated data tables improve the reusability of public indicators, and can reduce repetitive processing work when performing personalized data statistics at the end. The query can be directly connected to the user end, which greatly shortens the query response time and makes data use more efficient.
[0088] The embodiment of the present invention provides a financial and taxation data processing system which can execute any financial and taxation data processing method provided by the aforementioned embodiment of the present invention, and has the corresponding functional modules and beneficial effects of the execution method.
[0089] Embodiment 3
[0090] Figure 6 A schematic diagram of the structure of an electronic device 400 provided in Embodiment 3 of the present invention is shown in FIG. Figure 6 As shown, the terminal includes a memory 410 and a processor 420. The number of processors 420 in the terminal may be one or more. Figure 6 A processor 420 is taken as an example; the memory 410 and the processor 420 in the terminal can be connected via a bus or other means. Figure 6 The example of connecting through bus is taken in the following.
[0091] The memory 40, as a computer-readable storage medium, can be used to store software programs, computer executable programs and modules, such as program instructions / modules corresponding to the fiscal and taxation data processing method in the embodiment of the present invention (for example, the data acquisition module 310, the subject domain division module 320, the data aggregation module 330 and the personalized statistics module 340 in the fiscal and taxation data processing system). The processor 420 executes various functional applications and data processing of the terminal by running the software programs, instructions and modules stored in the memory 410, that is, realizes the above-mentioned fiscal and taxation data processing method.
[0092] Among them, the processor 420 is used to run the computer executable program stored in the memory 410 to implement the following steps: Step 110, obtain original financial and tax data; Step 120, input the original financial and tax data into a preset subject domain induction model to determine at least one subject domain based on the original financial and tax data, and divide the original financial and tax data into the subject domains; Step 130, perform dimensional modeling, business modeling and subject modeling on the original financial and tax data in each of the subject domains to obtain an aggregated data table; Step 140, perform personalized data statistics based on the aggregated data table to generate personalized statistical results.
[0093] Of course, an electronic device provided by an embodiment of the present invention is not limited to the method operation described above, and can also perform related operations in the financial and tax data processing method provided by any embodiment of the present invention.
[0094] The memory 410 may mainly include a program storage area and a data storage area, wherein the program storage area may store an operating system, application programs required for functions; the data storage area may store data created according to the use of the terminal, etc. In addition, the memory 410 may include a high-speed random access memory, and may also include a non-volatile memory, such as a disk storage device, a flash memory device, or other non-volatile solid-state storage devices. In some instances, the memory 410 may further include a memory remotely arranged relative to the processor 420, and these remote memories may be connected to the terminal via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0095] The above-mentioned device can execute the method provided by any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of the execution method.
[0096] Embodiment 4
[0097] Embodiment 4 of the present invention further provides a storage medium containing computer executable instructions, wherein the computer executable instructions are used to execute a financial and tax data processing method when executed by a computer processor, and the financial and tax data processing method includes:
[0098] Obtain original financial and tax data;
[0099] Inputting the original financial and tax data into a preset subject domain induction model to determine at least one subject domain according to the original financial and tax data, and dividing the original financial and tax data into the subject domains;
[0100] Performing dimensional modeling, business modeling and topic modeling on the original financial and taxation data in each of the subject domains to obtain an aggregated data table;
[0101] Perform personalized data statistics according to the aggregated data table to generate personalized statistical results.
[0102] Of course, the storage medium containing computer executable instructions provided in an embodiment of the present invention, whose computer executable instructions are not limited to the method operations described above, can also execute related operations in the financial and tax data processing method provided in any embodiment of the present invention.
[0103] Through the above description of the implementation methods, the technicians in the relevant field can clearly understand that the present invention can be implemented by means of software and necessary general hardware, and of course it can also be implemented by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product, and the computer software product can be stored in a computer-readable storage medium, such as a computer floppy disk, a read-only memory (ROM), a random access memory (RAM), a flash memory (FLASH), a hard disk or an optical disk, etc., including a number of instructions for a computer device (which can be a personal computer, a terminal, or a network device, etc.) to execute the methods described in each embodiment of the present invention.
[0104] It is worth noting that in the above-mentioned embodiment of the financial and taxation data processing system, the various units and modules included are only divided according to functional logic, but are not limited to the above-mentioned division, as long as the corresponding functions can be achieved; in addition, the specific names of the functional units are only for the convenience of distinguishing each other, and are not used to limit the scope of protection of the present invention.
[0105] Note that the above are only preferred embodiments of the present invention and the technical principles used. Those skilled in the art will understand that the present invention is not limited to the specific embodiments described herein, and that various obvious changes, readjustments and substitutions can be made by those skilled in the art without departing from the scope of protection of the present invention. Therefore, although the present invention has been described in more detail through the above embodiments, the present invention is not limited to the above embodiments, and may include more other equivalent embodiments without departing from the concept of the present invention, and the scope of the present invention is determined by the scope of the appended claims.
Claims
1. A method for processing financial and tax data, characterized in that: include: Obtain original financial and tax data; Inputting the original financial and tax data into a preset subject domain induction model to determine at least one subject domain according to the original financial and tax data, and dividing the original financial and tax data into the subject domains; Performing dimensional modeling, business modeling and topic modeling on the original financial and taxation data in each of the subject domains to obtain an aggregated data table; Performing personalized statistics on data according to the aggregated data table to generate personalized statistical results; The obtaining of original financial and tax data includes: obtaining original financial and tax data of a source system through a preset data interface, or obtaining original financial and tax data by recognizing paper financial and tax reports through an OCR recognition model; The inputting the original financial and tax data into a preset subject domain induction model to determine at least one subject domain according to the original financial and tax data includes: pairing the original financial and tax data with a preset subject to calculate subject relevance; if the subject relevance meets the preset requirements, the preset subject is used as the subject domain; if the subject relevance does not meet the preset requirements, synthesizing, classifying and analyzing the original financial and tax data to generate at least one subject domain; Before the original financial and tax data in each subject domain is subjected to dimensional modeling, business modeling and subject modeling to obtain an aggregated data table, the method further includes: cleaning abnormal data from the original financial and tax data and performing data normalization and storage; The dimensional modeling includes: based on the original financial and tax data, obtaining dimensional attributes through logical processing, obtaining dimensional attributes through multi-table association, obtaining dimensional attributes through mixed processing of different fields of a single table, and obtaining dimensional attributes through parsing designated fields of a single table; constructing a dimension table according to the dimensional attributes, and arranging the original financial and tax data into the dimension table; After performing personalized statistics on data according to the aggregated data table to generate personalized statistical results, the method further includes: matching the personalized statistical results with the subject domain to calculate the relevance; and reversely adjusting the subject domain induction model according to the relevance.
2. A financial and tax data processing system, applying the financial and tax data processing method according to claim 1, characterized in that: include: Data acquisition module, used to obtain original financial and tax data; A subject domain division module, used for inputting the original financial and tax data into a preset subject domain induction model to generate at least one subject domain according to the original financial and tax data, and dividing the original financial and tax data into the subject domains; A data aggregation module, used to perform dimensional modeling, business modeling and topic modeling on the original financial and tax data in each subject domain to obtain an aggregated data table; The personalized statistics module is used to generate personalized statistical results by performing personalized data statistics according to the aggregated data table.
3. The financial and tax data processing system according to claim 2, characterized in that: Also includes: The model adjustment module is used to match the personalized statistical results with the subject domain to calculate the relevance, and reversely adjust the subject domain induction model according to the relevance.
4. An electronic device, characterized in that: It includes a memory and a processor, the memory stores a computer program that can be run on the processor, and the processor implements the financial and tax data processing method as claimed in claim 1 when executing the computer program.
5. A computer-readable storage medium, characterized in that: The storage medium stores a computer program, which includes program instructions. When the program instructions are executed, the financial and tax data processing method according to claim 1 is implemented.
Citation Information
Patent Citations
Integrated tax administration platform
CN105184642A