Tax data processing analysis method and device based on large model
Through the big model-based tax data processing and analysis method, the problems of cross-department sharing obstacles and inefficiency in tax data processing are solved, and efficient tax data processing and analysis are achieved to support the accuracy of tax decisions.
Patent Information
- Application Number
- CN202411915901.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-24
- Publication Date
- 2025-05-23
AI Technical Summary
In the prior art, tax data processing methods have problems such as cross-department sharing barriers and inefficient data operation when facing massive tax data.
The tax data processing and analysis method based on the big model is adopted, and the data is marked and statistically analyzed by obtaining the financial and tax business data, using the preset financial and tax vertical field big model, generating the tax business statistical analysis result data, and interactive visual analysis and report generation with users.
This method effectively reduces the obstacles to cross-departmental data sharing, improves the efficiency and accuracy of tax data processing, supports tax staff to timely understand tax collection and management loopholes and policy implementation deviations, and make effective decisions.
Smart Images

Figure CN120030112A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of intelligent AI technology, and more specifically, to a tax data processing and analysis method and device based on a large model. Background Art
[0002] With the rapid advancement of digital tax management, the amount of data such as taxpayer information, transaction data, declaration data, and consultation feedback data in tax business scenarios has increased dramatically. These data are diverse in source, highly professional, and complex in format, which brings huge challenges to data processing. At the same time, when tax policies are adjusted, the tax bureau must update the data processing and analysis logic in a timely manner to adapt to the changes. When facing the processing of massive tax data, the tax data processing methods in the existing technology have problems such as obstacles to cross-departmental sharing and low data operation efficiency. Summary of the invention
[0003] In order to solve the technical problems of cross-departmental sharing barriers and low data operation efficiency in the existing tax data processing methods when facing the processing of massive tax data, the present invention provides a tax data processing and analysis method and device based on a large model.
[0004] According to one aspect of the present invention, the present invention provides a tax data processing and analysis method based on a large model, comprising:
[0005] Obtain financial and tax business data;
[0006] Use the preset financial and taxation vertical field model to analyze the financial and taxation business data. Annotation , generate annotation business data and Store in index , wherein the labels for annotating business data include business labels and category labels;
[0007] According to the natural language instructions input by the user, the financial and taxation vertical field big model is used to perform statistical analysis on the marked business data to generate taxation business statistical analysis result data, and interactive visual analysis is carried out with the user;
[0008] According to the question and demand words input by the user, the financial and taxation vertical field big model is used to filter the taxation business statistical analysis result data to generate a taxation business analysis report.
[0009] According to another aspect of the present invention, the present invention provides a tax data processing and analysis device based on a large model, the device comprising:
[0010] The data acquisition module is used to obtain financial and tax business data, as well as natural language instructions and question demand words input by users;
[0011] A data labeling module, used to label the financial and tax business data using a preset financial and tax vertical field model, generate labeled business data and store it in an index library, wherein the labels of the labeled business data include business labels and category labels;
[0012] An intelligent analysis module, which is used to perform statistical analysis on the annotated business data using the financial and taxation vertical field big model according to the natural language instructions input by the user to generate taxation business statistical analysis result data, and to conduct interactive visual analysis with the user;
[0013] The report generation module is used to generate a tax business analysis report by filtering the tax business statistical analysis result data based on the question requirements input by the user using the finance and taxation vertical field big model.
[0014] According to another aspect of the present invention, the present invention provides a computer-readable storage medium, wherein the storage medium stores a computer program, and the computer program is used to execute the method described in any one of the above aspects of the present invention.
[0015] According to another aspect of the present invention, an electronic device is provided, comprising: a processor; a memory for storing instructions executable by the processor; the processor is configured to read the executable instructions from the memory and execute the instructions to implement the method described in any one of the above aspects of the present invention.
[0016] The tax data processing and analysis method and device based on the big model of the present invention, wherein the method includes obtaining financial and tax business data; using the preset financial and tax vertical field big model to mark the financial and tax business data, generating marked business data and storing it in the index library, wherein the label of the marked business data includes business labels and category labels; according to the natural language instructions input by the user, using the financial and tax vertical field big model to perform statistical analysis on the marked business data to generate tax business statistical analysis result data, and conducting interactive visual analysis with the user; according to the question demand words input by the user, using the financial and tax vertical field big model to filter from the tax business statistical analysis result data to generate a tax business analysis report. The method uses the trained financial and tax vertical field big model to assist the daily work of financial and tax staff, reduce the work pressure of front-line staff, and improve work efficiency; solve the problem that there are obstacles to cross-departmental data sharing, thereby affecting the efficiency of tax work and the accuracy of decision-making, improve the comprehensive operation capability of tax big data, and can timely and comprehensively understand tax collection and management loopholes, inadequate tax services, policy implementation deviations and other problems through natural language commands, and make effective decisions. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] A more complete understanding of exemplary embodiments of the present invention may be obtained by referring to the following drawings:
[0018] Figure 1 A flowchart of a tax data processing and analysis method based on a large model according to a preferred embodiment of the present invention;
[0019] Figure 2 It is a structural schematic diagram of a tax data processing and analysis device based on a large model according to a preferred embodiment of the present invention;
[0020] Figure 3 Schematic diagram of the structure of an electronic device according to a preferred embodiment of the present invention. DETAILED DESCRIPTION
[0021] Now, exemplary embodiments of the present invention are described with reference to the accompanying drawings. However, the present invention can be implemented in many different forms and is not limited to the embodiments described herein. These embodiments are provided to disclose the present invention in detail and completely and to fully convey the scope of the present invention to those skilled in the art. The terms used in the exemplary embodiments shown in the accompanying drawings are not intended to limit the present invention. In the accompanying drawings, the same units / elements are marked with the same reference numerals.
[0022] Unless otherwise specified, the terms (including technical terms) used herein have the commonly understood meanings to those skilled in the art. In addition, it is understood that the terms defined in commonly used dictionaries should be understood to have the same meanings as those in the context of the relevant fields, and should not be understood as idealized or overly formal meanings.
[0023] Exemplary Methods
[0024] Figure 1 FIG. 1 is a flow chart of a tax data processing and analysis method based on a large model according to a preferred embodiment of the present invention. Figure 1 As shown, the tax data processing and analysis method based on a large model described in this preferred embodiment starts from step 101.
[0025] In step 101, financial and taxation business data is obtained.
[0026] Preferably, before acquiring the financial and tax business data, a large model of the financial and tax vertical field is established, including:
[0027] Establish an initial financial and taxation model based on the historical financial and taxation knowledge data obtained;
[0028] The initial finance and taxation big model is tuned and learned according to the finance and taxation knowledge data updated in real time to generate an optimized finance and taxation big model;
[0029] Obtain historical financial and tax business data;
[0030] Clustering the historical finance and taxation business data using an open source data clustering method;
[0031] According to the clustering results, the optimized finance and taxation big model is used to generate a category title and description for each category of historical finance and taxation business data, and a classification label is annotated for each category of historical finance and taxation business data based on the category title;
[0032] Based on a multi-level business label system customized according to business scenarios, the historical finance and taxation business data that has been labeled with classification labels are labeled with business labels to generate labeled historical business data;
[0033] The optimized finance and taxation big model is trained for business data labeling based on the annotated historical business data to generate a finance and taxation vertical field big model that can label business labels and classification labels at the same time, and the annotated historical business data is stored in the index library.
[0034] Preferably, the obtaining of financial and tax business data includes:
[0035] Obtain financial and tax business data entered by users; and / or
[0036] Get unlabeled finance and tax business data added in batches from the data source.
[0037] In this preferred embodiment, first, in order to discover the inherent patterns and structures of the data, the grouping of historical financial and tax business data is automatically found through open source data clustering methods such as K-means. In order to describe the generated clusters, the present invention innovatively proposes to use big model technology to analyze and summarize the data of each cluster through a pre-established big model of the financial and tax vertical field, and generate concise and clear category titles and descriptions, which help data analysts to more easily understand the clusters and their main features, and support subsequent rapid retrieval for each category, as well as in-depth analysis of the common features of each category.
[0038] Secondly, in order to accurately locate data of a specific type or category through rapid retrieval through labels, help better understand the data to support the training of models for data annotation in the large model of the finance and taxation vertical field, and more easily perform visualization to more intuitively display data features, it is necessary to perform multi-level label design and analysis indicator setting in the process of processing historical finance and taxation business data, and perform data annotation on historical finance and taxation business data according to the designed multi-level business labels. Traditional data labeling relies on manual operation, consumes a lot of manpower costs and cannot guarantee the accuracy of labeling, thus affecting the quality of subsequent data analysis. The present invention proposes to optimize the data labeling capabilities of the large model of the finance and taxation vertical field that has learned finance and taxation knowledge and finance and taxation policies through the study of a large amount of labeled data, so as to assist data labeling work, reduce the burden of manual labeling, and improve labeling efficiency and quality. For example, for work order data, it is necessary to mark the types of tax issues that users consult (system operation type, tax collection and management process type, policy and business type, etc.), and manually label some data for optimization of the big model in the finance and taxation vertical field. The big model can then perform semantic understanding and analysis on subsequent user demands such as "How does a taxpayer verify the information of real-name tax payers online?", classify it as a system operation type and complete data labeling.
[0039] In step 102, the preset finance and taxation vertical field big model is used to label the finance and taxation business data, generate labeled business data and store it in the index library, wherein the labels of the labeled business data include business labels and category labels.
[0040] In step 103, according to the natural language instructions input by the user, the finance and taxation vertical domain big model is used to perform statistical analysis on the annotated business data to generate taxation business statistical analysis result data, and interactive visual analysis is carried out with the user.
[0041] Preferably, according to the natural language instructions input by the user, the big model of the finance and taxation vertical field is used to perform statistical analysis on the annotated business data to generate taxation business statistical analysis result data, and interactive visual analysis is carried out with the user, including:
[0042] According to the natural language data statistical analysis query instructions input by the user, the finance and taxation vertical field big model converts it into SQL query statements using NL2SQL technology, performs statistical analysis on the annotated business data in the index library, and generates taxation business statistical analysis result data after screening out the annotated business data that meets the query instructions, and displays it on the BI visualization screen;
[0043] Based on the statistical results of the annotated business data displayed on the BI visualization screen, the user uses the finance and taxation vertical field big model to interact with the BI visualization screen through natural language instructions to adjust the analysis dimensions and indicators of the annotated business data in real time.
[0044] In terms of specific applications, new fiscal and taxation policies emerge in an endless stream, and taxation staff need to timely and comprehensively grasp the regulations and policies related to tax incentives, tax collection and management, and tax audits to improve the compliance, efficiency, and scientific nature of taxation work. The fiscal and taxation vertical field big model learns a wide range of fiscal and taxation policies and regulations and fiscal and taxation knowledge, and updates them in real time. It can be used as an intelligent analysis assistant and embedded in various taxation work websites in the form of a floating window, so that taxation staff can obtain the latest and most comprehensive fiscal and taxation policy information from the big model anytime and anywhere through natural language dialogue. The big model can also be used to provide in-depth interpretation and personalized consultation on new fiscal and taxation policies and knowledge.
[0045] In addition, tax staff need to obtain the latest multi-dimensional and multi-level tax statistics anytime and anywhere through convenient query methods, so as to grasp the actual situation of tax work in real time and make timely adjustments and optimizations to tax work. The large model NL2SQL technology can convert the natural language questions entered by users into query statements (SQL), so that tax staff with non-technical backgrounds can also easily and quickly obtain tax statistics filtered and analyzed by business tags of different dimensions such as time, region, and industry through natural language interaction, without having to wait for technical personnel to convert the analysis results into visual charts, which significantly improves work efficiency and decision-making timeliness.
[0046] Finally, in order to more intuitively display the trend, distribution and manageability of tax data and facilitate in-depth analysis and mining by tax personnel, BI charts and large-scale visualization screens are generally used to display statistical data. The large model of the finance and taxation vertical field makes the BI large-scale visualization screen more interactive. Tax personnel can interact with the large screen through natural language and adjust the analysis dimensions and indicators in real time, thereby obtaining deeper data insights or data results that are more in line with their own analysis needs.
[0047] In step 104, based on the question and demand words input by the user, the taxation vertical domain big model is used to filter the taxation business statistical analysis result data to generate a taxation business analysis report.
[0048] Preferably, the method uses the big model of the finance and taxation vertical field to filter the taxation business statistical analysis result data based on the question demand words input by the user to generate a taxation business analysis report, including:
[0049] The finance and taxation vertical field big model is used to annotate the question demand words input by the user, and the business tags and classification tags related to the question demand are determined;
[0050] Filter the tax business statistical analysis result data according to the business tags and classification tags related to the problem requirements;
[0051] Draw BI visualization charts based on the filtered tax business statistical analysis results;
[0052] The tax business statistical analysis result data and BI visualization charts obtained by comprehensive screening of the finance and taxation vertical field big model are used to assist in generating an analysis report based on a preset report template, wherein the analysis report includes problem analysis conclusions and suggestions for the problems.
[0053] In terms of practical application, due to the high complexity and importance of tax work, regular output of analysis reports can comprehensively and systematically reflect the actual situation of tax work, provide accurate and timely information for tax workers and superior departments, and analysis reports are also an important means of supervising the implementation of tax work, which helps to discover deficiencies and problems in the work, such as loopholes in tax collection and management, inadequate tax services, policy implementation deviations, etc., and provide support for the decision-making of the tax department through analysis reports. The large model in the vertical field of finance and taxation has powerful data processing capabilities, and can assist in inputting, integrating and analyzing tax data by learning a large amount of tax data and professional knowledge, discover problems, trends and potential risks in tax work, and use language generation capabilities to assist in generating analysis reports based on preset templates and data analysis results, thereby greatly simplifying the report preparation process for tax workers and improving preparation efficiency.
[0054] The tax data processing and analysis method based on the big model described in this preferred embodiment generates the title and description of each cluster of tax business data by using the big model of the finance and taxation vertical field established by the big model technology, and assists in data annotation, builds a finance and taxation database, effectively reduces the workload of manual data processing and annotation, and improves the efficiency and accuracy of tax data processing, thereby supporting the mining of the deep value of tax business data, discovering potential tax risks, and identifying abnormal patterns. Furthermore, by using the big model of the finance and taxation vertical field, statistical data can be queried through natural language anytime and anywhere, and visual analysis results can be interactively viewed, so that financial workers who do not understand technology can also quickly obtain tax data analysis results and guide decision-making as needed; finally, by combining the big model of the finance and taxation vertical field with BI visual charts, various tax analysis reports can be quickly generated according to the analysis topics and problem requirements specified by the user, and risk warnings and prevention suggestions can be provided, providing intuitive and comprehensive data support for tax decision-making.
[0055] Exemplary Devices
[0056] Figure 2 FIG. 1 is a schematic diagram of the structure of a tax data processing and analysis device based on a large model according to a preferred embodiment of the present invention. Figure 2 As shown, the tax data processing and analysis device 200 based on the big model described in this preferred embodiment includes:
[0057] The data acquisition module 201 is used to acquire financial and tax business data, as well as natural language instructions and question demand words input by users;
[0058] The data labeling module 202 is used to label the financial and tax business data using a preset financial and tax vertical field model, generate labeled business data and store it in an index library, wherein the labels of the labeled business data include business labels and category labels;
[0059] Intelligent analysis module 203, used to perform statistical analysis on the marked business data to generate tax business statistical analysis result data according to the natural language instructions input by the user, and to conduct interactive visual analysis with the user;
[0060] The report generation module 204 is used to generate a tax business analysis report by filtering the tax business statistical analysis result data based on the question requirements input by the user using the finance and taxation vertical field big model.
[0061] Preferably, the device further comprises a model building module for building a large model in the vertical field of finance and taxation, including:
[0062] Establish an initial financial and taxation model based on the historical financial and taxation knowledge data obtained;
[0063] The initial finance and taxation big model is tuned and learned according to the finance and taxation knowledge data updated in real time to generate an optimized finance and taxation big model;
[0064] Obtain historical financial and tax business data;
[0065] Clustering the historical finance and taxation business data using an open source data clustering method;
[0066] According to the clustering results, the optimized finance and taxation big model is used to generate a category title and description for each category of historical finance and taxation business data, and a classification label is annotated for each category of historical finance and taxation business data based on the category title;
[0067] Based on a multi-level business label system customized according to business scenarios, the historical finance and taxation business data that has been labeled with classification labels are labeled with business labels to generate labeled historical business data;
[0068] The optimized finance and taxation big model is trained for business data labeling based on the annotated historical business data to generate a finance and taxation vertical field big model that can label business labels and classification labels at the same time, and the annotated historical business data is stored in the index library.
[0069] Preferably, the data acquisition module 201 acquires the financial and taxation business data including:
[0070] Obtain financial and tax business data entered by users; and / or
[0071] Get unlabeled finance and tax business data added in batches from the data source.
[0072] Preferably, the intelligent analysis module 203 uses the finance and taxation vertical field big model to perform statistical analysis on the annotated business data according to the natural language instructions input by the user to generate taxation business statistical analysis result data, and conducts interactive visual analysis with the user, including:
[0073] According to the natural language data statistical analysis query instructions input by the user, the finance and taxation vertical field big model converts it into SQL query statements using NL2SQL technology, performs statistical analysis on the annotated business data in the index library, and generates taxation business statistical analysis result data after screening out the annotated business data that meets the query instructions, and displays it on the BI visualization screen;
[0074] Based on the annotated business data displayed on the BI visualization screen, the user uses the finance and taxation vertical field big model to interact with the BI visualization screen through natural language instructions to adjust the analysis dimensions and indicators of the annotated business data in real time.
[0075] Preferably, the report generation module 204 uses the taxation vertical field big model to filter the taxation business statistical analysis result data according to the question demand words input by the user, and generates a taxation business analysis report, including:
[0076] The finance and taxation vertical field big model is used to annotate the question demand words input by the user, and the business tags and classification tags related to the question demand are determined;
[0077] Filter the tax business statistical analysis result data according to the business tags and classification tags related to the problem requirements;
[0078] Draw BI visualization charts based on the filtered tax business statistical analysis results;
[0079] The tax business statistical analysis result data and BI visualization charts obtained by comprehensive screening of the finance and taxation vertical field big model are used to assist in generating an analysis report based on a preset report template, wherein the analysis report includes problem analysis conclusions and suggestions for the problems.
[0080] The big model-based tax data processing and analysis device and the big model-based tax data processing and analysis method described in this preferred embodiment use a trained big model of the finance and taxation vertical field to label business data, perform statistical analysis of business data and generate analysis reports. The steps are the same, and the technical effects achieved are also the same, which will not be repeated here.
[0081] Exemplary Electronic Devices
[0082] Figure 3 The electronic device may be any one or both of the first device and the second device, or a stand-alone device independent of them, and the stand-alone device may communicate with the first device and the second device to receive the collected input signals from them. Figure 3 FIG. 1 is a block diagram of an electronic device according to an embodiment of the present disclosure. Figure 3 As shown, the electronic device includes one or more processors 301 and a memory 302 .
[0083] The processor 301 may be a central processing unit (CPU) or other forms of processing units having data processing capabilities and / or instruction execution capabilities, and may control other components in the electronic device to perform desired functions.
[0084] The memory 302 may include one or more computer program products, which may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. The volatile memory may include, for example, random access memory (RAM) and / or cache memory (cache), etc. The non-volatile memory may include, for example, read-only memory (ROM), hard disk, flash memory, etc. One or more computer program instructions may be stored on the computer-readable storage medium, and the processor 301 may run the program instructions to implement the energy consumption anomaly diagnosis method based on the enterprise energy consumption space of each disclosed embodiment described above and / or other desired functions. In one example, the electronic device may also include: an input device 303 and an output device 304, which are interconnected via a bus system and / or other forms of connection mechanisms (not shown).
[0085] In addition, the input device 303 may also include, for example, a keyboard, a mouse, and the like.
[0086] The output device 304 can output various information to the outside, and can include, for example, a display, a speaker, a printer, a communication network and a remote output device connected thereto.
[0087] Of course, to simplify, Figure 3 Only some of the components related to the present disclosure in the electronic device are shown, and components such as a bus, an input / output interface, etc. are omitted. In addition, according to specific application situations, the electronic device may further include any other appropriate components.
[0088] Exemplary computer program products and computer-readable storage media
[0089] In addition to the above-mentioned methods and devices, an embodiment of the present disclosure may also be a computer program product, which includes computer program instructions, which, when executed by a processor, enable the processor to execute the steps of the big model-based tax data processing and analysis method according to various embodiments of the present disclosure described in the above-mentioned "Exemplary Method" section of this specification.
[0090] The computer program product may be written in any combination of one or more programming languages to write program code for performing the operations of the disclosed embodiments, including object-oriented programming languages such as Java, C++, etc., and conventional procedural programming languages such as "C" or similar programming languages. The program code may be executed entirely on the user computing device, partially on the user device, as a separate software package, partially on the user computing device and partially on a remote computing device, or entirely on a remote computing device or server.
[0091] In addition, an embodiment of the present disclosure may also be a computer-readable storage medium on which computer program instructions are stored. When the computer program instructions are executed by a processor, the processor executes the steps of the tax data processing and analysis method based on a large model according to various embodiments of the present disclosure described in the above "Exemplary Method" section of this specification.
[0092] The computer readable storage medium can adopt any combination of one or more readable media. The readable medium can be a readable signal medium or a readable storage medium. The readable storage medium can include, for example, but is not limited to, a system, device or device of electricity, magnetism, light, electromagnetic, infrared, or semiconductor, or any combination of the above. More specific examples (non-exhaustive list) of readable storage media include: an electrical connection with one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.
[0093] The basic principles of the present disclosure are described above in conjunction with specific embodiments. However, it should be noted that the advantages, strengths, effects, etc. mentioned in the present disclosure are only examples and not limitations, and it cannot be considered that these advantages, strengths, effects, etc. are required by each embodiment of the present disclosure. In addition, the specific details disclosed above are only for the purpose of illustration and ease of understanding, and are not limitations. The above details do not limit the present disclosure to the necessity of adopting the above specific details to be implemented.
[0094] Each embodiment in this specification is described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the embodiments can be referred to each other. For the system embodiment, since it basically corresponds to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.
[0095] The block diagrams of the devices, apparatuses, equipment, and systems involved in this disclosure are only illustrative examples and are not intended to require or imply that they must be connected, arranged, and configured in the manner shown in the block diagrams. As will be appreciated by those skilled in the art, these devices, apparatuses, equipment, and systems can be connected, arranged, and configured in any manner. Words such as "including," "comprising," "having," and the like are open words, referring to "including but not limited to," and can be used interchangeably therewith. The words "or" and "and" used herein refer to the words "and / or," and can be used interchangeably therewith, unless the context clearly indicates otherwise. The word "such as" used herein refers to the phrase "such as but not limited to," and can be used interchangeably therewith.
[0096] The apparatus and method of the present disclosure may be implemented in many ways. For example, the apparatus and method of the present disclosure may be implemented by software, hardware, firmware, or any combination of software, hardware, and firmware. The above order of steps for the method is for illustration only, and the steps of the method of the present disclosure are not limited to the order specifically described above, unless otherwise specifically stated. In addition, in some embodiments, the present disclosure may also be implemented as a program recorded in a recording medium, which includes machine-readable instructions for implementing the method according to the present disclosure. Therefore, the present disclosure also covers a recording medium storing a program for executing the method according to the present disclosure.
[0097] It should also be noted that in the apparatus, equipment and method of the present disclosure, each component or each step can be decomposed and / or recombined. These decompositions and / or recombinations should be regarded as equivalent schemes of the present disclosure. The above description of the disclosed aspects is provided to enable any technician in the field to make or use the present disclosure. Various modifications to these aspects are very obvious to those skilled in the art, and the general principles defined herein can be applied to other aspects without departing from the scope of the present disclosure. Therefore, the present disclosure is not intended to be limited to the aspects shown here, but to the widest scope consistent with the principles and novel features disclosed herein.
[0098] The above description has been given for the purpose of illustration and description. In addition, this description is not intended to limit the embodiments of the present disclosure to the forms disclosed herein. Although multiple example aspects and embodiments have been discussed above, those skilled in the art will recognize certain variations, modifications, changes, additions and sub-combinations thereof.
Claims
1. A tax data processing and analysis method based on a large model, characterized in that: The method comprises: Obtain financial and tax business data; The financial and tax business data are labeled using a preset financial and tax vertical field big model, and labeled business data is generated and stored in an index library, wherein the labels of the labeled business data include business labels and category labels; According to the natural language instructions input by the user, the financial and taxation vertical field big model is used to perform statistical analysis on the marked business data to generate taxation business statistical analysis result data, and interactive visual analysis is carried out with the user; According to the question and demand words input by the user, the financial and taxation vertical field big model is used to filter the taxation business statistical analysis result data to generate a taxation business analysis report.
2. The method according to claim 1, characterized in that Before acquiring financial and tax business data, a large model for the financial and tax vertical field is established, including: Establish an initial financial and taxation model based on the historical financial and taxation knowledge data obtained; The initial finance and taxation big model is tuned and learned according to the finance and taxation knowledge data updated in real time to generate an optimized finance and taxation big model; Obtain historical financial and tax business data; Clustering the historical finance and taxation business data using an open source data clustering method; According to the clustering results, the optimized finance and taxation big model is used to generate a category title and description for each category of historical finance and taxation business data, and a classification label is annotated for each category of historical finance and taxation business data based on the category title; Based on a multi-level business label system customized according to business scenarios, the historical finance and taxation business data that has been labeled with classification labels are labeled with business labels to generate labeled historical business data; The optimized finance and taxation big model is trained for business data labeling based on the annotated historical business data to generate a finance and taxation vertical field big model that can label business labels and classification labels at the same time, and the annotated historical business data is stored in the index library.
3. The method according to claim 1, characterized in that: The acquisition of financial and tax business data includes: Obtain financial and tax business data entered by users; and / or Get unlabeled finance and tax business data added in batches from the data source.
4. The method according to claim 1, characterized in that: According to the natural language instructions input by the user, the big model of the finance and taxation vertical field is used to perform statistical analysis on the marked business data to generate taxation business statistical analysis result data, and interactive visual analysis is carried out with the user, including: According to the natural language data statistical analysis query instructions input by the user, the finance and taxation vertical field big model converts it into SQL query statements using NL2SQL technology, performs statistical analysis on the annotated business data in the index library, and generates taxation business statistical analysis result data after screening out the annotated business data that meets the query instructions, and displays it on the BI visualization screen; Based on the annotated business data displayed on the BI visualization screen, the user uses the finance and taxation vertical field big model to interact with the BI visualization screen through natural language instructions to adjust the analysis dimensions and indicators of the annotated business data in real time.
5. The method according to claim 1, characterized in that The method uses the big model of the finance and taxation vertical field to filter the taxation business statistical analysis result data based on the question and demand words input by the user to generate a taxation business analysis report, including: The finance and taxation vertical field big model is used to annotate the question demand words input by the user, and the business tags and classification tags related to the question demand are determined; Filter the tax business statistical analysis result data according to the business tags and classification tags related to the problem requirements; Draw BI visualization charts based on the filtered tax business statistical analysis results; The tax business statistical analysis result data and BI visualization charts obtained by comprehensive screening of the finance and taxation vertical field big model are used to assist in generating an analysis report based on a preset report template, wherein the analysis report includes problem analysis conclusions and suggestions for the problems.
6. A tax data processing and analysis device based on a large model, characterized in that: The device comprises: The data acquisition module is used to obtain financial and tax business data, as well as natural language instructions and question demand words input by users; A data labeling module, used to label the financial and tax business data using a preset financial and tax vertical field model, generate labeled business data and store it in an index library, wherein the labels of the labeled business data include business labels and category labels; An intelligent analysis module, which is used to perform statistical analysis on the annotated business data using the finance and taxation vertical field big model according to the natural language instructions input by the user to generate taxation business statistical analysis result data, and to conduct interactive visual analysis with the user; The report generation module is used to generate a tax business analysis report by filtering the tax business statistical analysis result data based on the question requirements input by the user using the finance and taxation vertical field big model.
7. The device according to claim 6, characterized in that The device also includes a model building module for building a large model in the vertical field of finance and taxation, including: Establish an initial financial and taxation model based on the historical financial and taxation knowledge data obtained; The initial finance and taxation big model is tuned and learned according to the finance and taxation knowledge data updated in real time to generate an optimized finance and taxation big model; Obtain historical financial and tax business data; Clustering the historical finance and taxation business data using an open source data clustering method; According to the clustering results, the optimized finance and taxation big model is used to generate a category title and description for each category of historical finance and taxation business data, and a classification label is annotated for each category of historical finance and taxation business data based on the category title; Based on a multi-level business label system customized according to business scenarios, the historical finance and taxation business data that has been labeled with classification labels are labeled with business labels to generate labeled historical business data; The optimized finance and taxation big model is trained for business data labeling based on the annotated historical business data to generate a finance and taxation vertical field big model that can label business labels and classification labels at the same time, and the annotated historical business data is stored in the index library.
8. The device according to claim 6, characterized in that The data acquisition module acquires the financial and taxation business data including: Obtain financial and tax business data entered by users; and / or Get unlabeled finance and tax business data added in batches from the data source.
9. A computer-readable storage medium, characterized in that: The storage medium stores a computer program, and the computer program is used to execute the method according to any one of claims 1 to 5.
10. An electronic device, characterized in that: The electronic device comprises: processor; a memory for storing instructions executable by the processor; The processor is used to read the executable instructions from the memory and execute the instructions to implement the method described in any one of claims 1 to 5.
Citation Information
Patent Citations
Method and device for processing finance and tax data based on double-label model, medium and equipment
CN116245670A
Financial bill and non-tax collection intelligent data analysis platform
CN118964399A
Operation management data intelligent analysis system and method based on AIGC
CN118982175A
Tax field-oriented knowledge map construction method and system
WO2021196520A1