Finance and tax training data processing method and system for large language model

Through the collection, preprocessing, marking classification and quality inspection methods of fiscal and taxation data, combined with the collaboration between manual and machine, the problem of difficult to obtain high-quality training data in the financial and taxation field is solved, the timeliness and accuracy of large-scale training data is achieved, and labor costs are reduced.

CN120030342APending Publication Date: 2025-05-23AISINO CORPORATION
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202411915892.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-24
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

It is difficult to obtain high-quality fiscal and tax training data in the existing technology, especially when the data in the fiscal and taxation field is scarce and professional, and fiscal and tax policies change quickly, making it difficult to update in a timely manner.

Method used

A method for financial and tax training data processing for large language models is proposed, including data collection, data preprocessing, multi-dimensional marking classification, initial training data construction and quality detection. Through manual and machine collaboration, closed-loop iteration is formed to gradually improve data quality.

Benefits of technology

The full-process data construction strategy is implemented, which solves the problem of difficult to construct high-quality training data in the fiscal and taxation industry, ensures the timeliness and accuracy of large-scale training data, and reduces labor costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120030342A_ABST
    Figure CN120030342A_ABST
Patent Text Reader

Abstract

The invention discloses a finance and taxation training data processing method and system for a large language model, and the method comprises the steps: collecting finance and taxation data for finance and taxation training, so as to obtain original finance and taxation data; performing data preprocessing on the original finance and taxation data, exploring finance and taxation processing data obtained after data preprocessing, and performing marking classification on the finance and taxation processing data in a multi-dimensional manner to obtain finance and taxation label data; according to different types of training tasks, initial training data are constructed based on the finance and tax label data; and performing quality detection on the initial training data to obtain finance and tax training data which meets training requirements and is used for large language model training. According to the method, from data acquisition to final data quality inspection, a whole-process data construction strategy is realized, and the problem that high-quality training data in the finance and tax industry is difficult to construct is solved; and through continuously updated assembly line work, the training data of the large model closely follows the change condition of finance and taxation laws and policies, and outdated judgment and reply are avoided.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data processing, and more specifically, to a method and system for processing fiscal and tax training data for large language models. Background Art

[0002] To build powerful large language models, it is necessary to collect massive amounts of data from diverse data sources for training. Existing large language models mainly mix various publicly available text data, such as web pages, books, news, scientific texts, code data, and other data as pre-training corpora.

[0003] According to different sources, pre-training data is mainly divided into two types: general text data and specialized text data. General text data covers web pages, books, and conversation texts, etc. Large language models will collect a large amount of general text data to enhance their language modeling capabilities. In addition, in order to further improve the performance of large language models in specific professional tasks, the scope of pre-training corpora has been extended to more specialized data sets, such as policy interpretations, exams, reading comprehensions, etc. in the fiscal and tax fields.

[0004] However, data in professional fields outside the Internet industry is relatively scarce, especially in the fiscal and tax fields, where there is little ready-made or open-source data. A small part of the data is scattered on the Internet in a structured form, and quite a lot of data is stored in a semi-structured or even unstructured manner, making it difficult to obtain high-quality training data.

[0005] Moreover, fiscal and tax-related policies change relatively quickly, and there are significant differences between new and old policies. If it is necessary to speed up the frequency of regularly updating policy announcements of countries, regions, and relevant departments, knowledge optimization should be carried out in a timely manner to avoid interference from old information to users.

[0006] Finally, the data in the fiscal and tax fields is relatively professional, and general text quality detection cannot identify errors or even harmful information in it, and a large amount of professional manpower is required for processing.

[0007] Combining the above problems, there is an urgent need for a method for processing fiscal and tax training data for large language models. Summary of the Invention

[0008] The present invention proposes a method and system for processing fiscal and tax training data for large language models to solve the problem of how to obtain high-quality fiscal and tax training data.

[0009] To solve the above problems, according to one aspect of the present invention, there is provided a method for processing fiscal and tax training data for large language models, the method comprising:

[0010] Collecting fiscal and tax data for fiscal and tax training to obtain original fiscal and tax data;

[0011] Preprocessing the original financial and tax data, exploring the financial and tax processing data obtained after the data preprocessing, and labeling and classifying the financial and tax processing data in multiple dimensions to obtain financial and tax label data;

[0012] According to different types of training tasks, constructing initial training data based on the financial and tax label data;

[0013] The initial training data is subjected to quality inspection to obtain financial and taxation training data for large language model training that meets training requirements.

[0014] Preferably, the collecting of the financial and taxation data for financial and taxation training to obtain the original financial and taxation data includes:

[0015] Capturing industry data at preset time intervals, analyzing the industry data to obtain text data and attachment data, and performing document analysis on the attachment data;

[0016] Obtaining business documents uploaded by the business department and performing document analysis on the business documents;

[0017] The text data, the attachment data and the business documents that have undergone document analysis are used as the original financial and taxation data.

[0018] Preferably, the data preprocessing of the original financial and tax data includes:

[0019] Clean the original financial and tax data and unify the format to obtain clean data;

[0020] Performing deduplication processing on the clean data;

[0021] The clean data after deduplication processing is filtered to obtain the financial and tax processing data.

[0022] Preferably, the construction of initial training data based on the fiscal and taxation label data according to different types of training tasks includes:

[0023] Obtain complete QA data from the tax label data and use it directly as initial training data;

[0024] When it is determined based on the tax label data that the training data does not meet the quantity required by the training task, the LLM large model and prompt data enhancement technology are used to expand the data to obtain the initial training data;

[0025] When training data cannot be directly obtained based on financial and tax label data, task data is generated based on the self-question-answering method with the help of the LLM model to obtain initial training data.

[0026] Preferably, the quality detection of the initial training data to obtain the fiscal and taxation training data for large language model training that meets the training requirements includes:

[0027] The initial training data is manually labeled and reviewed, and the screened bad cases are sent to the quality inspection algorithm to train the quality inspection algorithm based on the bad cases, and the model is used to evaluate the data effect. After the target expectations are achieved, the data construction is started, and the pipeline is iterated. The data that has undergone multiple quality inspections is output as SFT training data, which is the financial and tax training data for large language model training that meets the training needs.

[0028] Preferably, the method further comprises:

[0029] Use quality control algorithms for machine testing to perform secondary screening of manually reviewed data;

[0030] The model is evaluated and the evaluation score is fed back to the manual labeling, quality inspection algorithm, and task decomposition parts for joint optimization of manual labeling and the model.

[0031] According to another aspect of the present invention, a financial and tax training data processing system for a large language model is provided, the system comprising:

[0032] A data collection unit, used for collecting financial and tax data for financial and tax training to obtain original financial and tax data;

[0033] A data processing unit, used to perform data preprocessing on the original financial and tax data, and to detect the financial and tax processed data obtained after the data preprocessing, and to label and classify the financial and tax processed data in multiple dimensions to obtain financial and tax label data;

[0034] A data construction unit, used for constructing initial training data based on the financial and tax label data according to different types of training tasks;

[0035] The quality inspection unit is used to perform quality inspection on the initial training data to obtain financial and tax training data for large language model training that meets training requirements.

[0036] Preferably, the data collection unit collects the financial and tax data used for financial and tax training to obtain the original financial and tax data, including:

[0037] Capturing industry data at preset time intervals, analyzing the industry data to obtain text data and attachment data, and performing document analysis on the attachment data;

[0038] Obtaining business documents uploaded by the business department and performing document analysis on the business documents;

[0039] The text data and the attachment data and business documents that have undergone document analysis are used as the original financial and taxation data.

[0040] Preferably, the data processing unit performs data preprocessing on the original financial and tax data, including:

[0041] Clean the original financial and tax data and unify the format to obtain clean data;

[0042] Performing deduplication processing on the clean data;

[0043] The clean data after deduplication processing is filtered to obtain the financial and tax processing data.

[0044] Preferably, the data construction unit constructs initial training data based on the fiscal and taxation label data according to different types of training tasks, including:

[0045] Obtain complete QA data from the tax label data and use it directly as initial training data;

[0046] When it is determined based on the tax label data that the training data does not meet the quantity required by the training task, the LLM large model and prompt data enhancement technology are used to expand the data to obtain the initial training data;

[0047] When training data cannot be directly obtained based on financial and tax label data, the self-question-answering system based on self-fQA uses the LLM model to generate task data to obtain initial training data.

[0048] Preferably, the quality inspection unit performs quality inspection on the initial training data to obtain financial and tax training data for large language model training that meets the training requirements, including:

[0049] The initial training data is manually labeled and reviewed, and the screened bad cases are sent to the quality inspection algorithm to train the quality inspection algorithm based on the bad cases, and the model is used to evaluate the data effect. After the target expectations are achieved, the data construction is started, and the pipeline is iterated. The data that has undergone multiple quality inspections is output as SFT training data, which is the financial and tax training data for large language model training that meets the training needs.

[0050] Preferably, the quality inspection unit is further used for:

[0051] Use quality control algorithms for machine testing to perform secondary screening of manually reviewed data;

[0052] The model is evaluated and the evaluation score is fed back to the manual labeling, quality inspection algorithm, and task decomposition parts for joint optimization of manual labeling and the model.

[0053] The present invention provides a method and system for processing tax training data for a large language model, including: collecting tax data for tax training to obtain original tax data; preprocessing the original tax data, and exploring the tax processing data obtained after data preprocessing, labeling and classifying the tax processing data in multiple dimensions to obtain tax label data; constructing initial training data based on the tax label data according to the type of training task; performing quality inspection on the initial training data to obtain tax training data for large language model training that meets the training requirements. The method of the present invention implements a full-process data construction strategy from data collection to the final data quality inspection, solves the problem that high-quality training data in the tax industry is difficult to construct; and through continuously updated pipeline operations, ensures that the training data of the large model keeps up with the changes in tax laws and regulations and policies, and avoids outdated judgments and replies; forms a closed loop of labeling, quality inspection, classification, and evaluation, and forms a flywheel iteration through manual + machine collaboration, gradually improves the quality of data labeling, reduces the labor cost of quality inspection, and finally realizes automated data construction and quality inspection processes, reducing labor costs by more than 80%. BRIEF DESCRIPTION OF THE DRAWINGS

[0054] A more complete understanding of exemplary embodiments of the present invention may be obtained by referring to the following drawings:

[0055] Figure 1 is a flowchart of a method 100 for processing finance and taxation training data for a large language model according to an embodiment of the present invention;

[0056] Figure 2 It is an overall flow chart of the financial and tax training data processing process according to an embodiment of the present invention;

[0057] Figure 3 A schematic diagram of the evolution of manual + machine annotation quality inspection according to an embodiment of the present invention;

[0058] Figure 4 Schematic diagram of the structure of a finance and taxation training data processing system 400 for a large language model according to an embodiment of the present invention. DETAILED DESCRIPTION

[0059] Now, exemplary embodiments of the present invention are described with reference to the accompanying drawings. However, the present invention can be implemented in many different forms and is not limited to the embodiments described herein. These embodiments are provided to disclose the present invention in detail and completely and to fully convey the scope of the present invention to those skilled in the art. The terms used in the exemplary embodiments shown in the accompanying drawings are not intended to limit the present invention. In the accompanying drawings, the same units / elements are marked with the same reference numerals.

[0060] Unless otherwise specified, the terms (including technical terms) used herein have the commonly understood meanings to those skilled in the art. In addition, it is understood that the terms defined in commonly used dictionaries should be understood to have the same meanings as those in the context of the relevant fields, and should not be understood as idealized or overly formal meanings.

[0061] Figure 1 FIG. 1 is a flow chart of a method 100 for processing tax training data for a large language model according to an embodiment of the present invention. Figure 1 As shown, the method for processing financial and tax training data for a large language model provided by the embodiment of the present invention realizes a full-process data construction strategy from data collection to the final data quality inspection, and solves the problem that high-quality training data in the financial and tax industry is difficult to construct; and through the continuous update of the pipeline operation, it ensures that the training data of the large model keeps up with the changes in financial and tax laws and policies, and avoids outdated judgments and replies; it forms a closed loop of labeling, quality inspection, classification, and evaluation, and the human + machine collaboration forms a flywheel iteration, gradually improves the quality of data labeling, reduces the labor cost of quality inspection, and finally realizes the automated data construction and quality inspection process, reducing the labor cost by more than 80%. The method 100 for processing financial and tax training data for a large language model provided by the embodiment of the present invention starts from step 101. In step 101, the financial and tax data used for financial and tax training are collected to obtain the original financial and tax data.

[0062] Preferably, the collecting of the financial and taxation data for financial and taxation training to obtain the original financial and taxation data includes:

[0063] Capturing industry data at preset time intervals, analyzing the industry data to obtain text data and attachment data, and performing document analysis on the attachment data;

[0064] Obtaining business documents uploaded by the business department and performing document analysis on the business documents;

[0065] The text data and the attachment data and business documents that have undergone document analysis are used as the original financial and taxation data.

[0066] Combination Figure 2As shown, in the present invention, in the data collection stage, relevant business data is mainly obtained through two channels: the Internet and the business department.

[0067] In order to ensure that the training data is updated in a timely manner, part of the data will be regularly captured from the industry data on the Internet and web data analysis (CC). Crawler, web crawling and other technologies can be used to obtain the original data from the official website of the competent department and the website of the relevant business system in a timely manner, and then handed over to the web data analysis (C-WA) module for analysis, and divided into two categories: text (CT) and attachment (C-ATT). The attachment data (C-ATT) will enter the document analysis module (C-ATT-A) and be handed over to the subsequent data processing module together with CT.

[0068] Because it involves professional knowledge and related systems, some professional documents are not available on the public network. Therefore, in order to complement the data on the public network, another part of the data is the business documents of the company's business departments. This type of data will be timely sorted and uploaded through the business system (CU) and handed over to C-ATT-A for processing and analysis. At the same time, it will also be retained in the company's knowledge base (CK) to provide a basis for future knowledge base retrieval.

[0069] In step 102, the original financial and tax data is preprocessed, and the financial and tax processing data obtained after the data preprocessing is explored, and the financial and tax processing data is labeled and classified in multiple dimensions to obtain financial and tax label data.

[0070] Preferably, the data preprocessing of the original financial and tax data includes:

[0071] Clean the original financial and tax data and unify the format to obtain clean data;

[0072] Performing deduplication processing on the clean data;

[0073] The clean data after deduplication processing is filtered to obtain the financial and tax processing data.

[0074] Combination Figure 2As shown, in the present invention, in the data processing stage, the collected data first enters the data cleaning module (EW). Data cleaning is the processing of all text data, removing useless and erroneous data, processing it into a unified format, and outputting it as clean data. Then, it enters the deduplication module (E-DEP). Because there are many sources of data, there are a large number of similar or even identical data, and it is necessary to deduplicate such data to ensure the uniqueness and accuracy of the data. Because there are still a large number of useless fields in the deduplicated data, such as collection metadata, the deduplicated data is filtered (E-FT) to retain fields that are useful for task analysis and model training. Finally, the data can be explored (EL), and the data can be labeled and classified according to multiple dimensions such as business direction, core capabilities, and business capabilities, so as to facilitate subsequent data construction according to tasks.

[0075] After data processing, there are basically no problems with the data format, semantics, and repeatability, and data can be constructed according to tasks.

[0076] In step 103, initial training data is constructed based on the financial and tax label data according to different types of training tasks.

[0077] Preferably, the construction of initial training data based on the fiscal and taxation label data according to different types of training tasks includes:

[0078] Obtain complete QA data from the tax label data and use it directly as initial training data;

[0079] When it is determined based on the tax label data that the training data does not meet the quantity required by the training task, the LLM large model and prompt data enhancement technology are used to expand the data to obtain the initial training data;

[0080] When training data cannot be directly obtained based on financial and tax label data, task data is generated based on the self-question-answering method with the help of the LLM model to obtain initial training data.

[0081] Combination Figure 2 As shown, in the present invention, in the data construction stage, the training tasks are decomposed (DP) according to the data exploration and business capability tasks.

[0082] Among them, according to different training tasks, the methods of constructing data are mainly divided into three categories:

[0083] 1. Question-Answering Data (D-QA): After being acquired and cleaned, some data becomes complete QA data and can be used directly as training data, but it requires manual labeling and inspection.

[0084] Prompt data augmentation (DA): Some tasks have small data volumes and require the use of external large models 2. (D-LLM) and prompt data augmentation techniques to further improve the quality and richness of the data. A method to expand training data by generating or modifying input prompts. The basic idea of ​​this method is to generate more training instances by designing different prompts, thereby enhancing the model's capabilities and generalization.

[0085] 3. Self-QA (D-SQA): Some data is not convenient to obtain or construct directly. Self-QA is a self-questioning method that also generates training data with the help of an external large model (D-LLM). This method generates new training samples by letting the model ask and answer questions by itself. Subsequently, manual evaluation or automatic evaluation methods can be used to select high-quality question-answer pairs.

[0086] After data construction processing, initial training data that meets business and format requirements is formed.

[0087] In step 104, the initial training data is quality checked to obtain financial and tax training data for large language model training that meets training requirements.

[0088] Preferably, the quality detection of the initial training data to obtain the fiscal and taxation training data for large language model training that meets the training requirements includes:

[0089] The initial training data is manually labeled and reviewed, and the screened bad cases are sent to the quality inspection algorithm to train the quality inspection algorithm based on the bad cases, and the model is used to evaluate the data effect. After the target expectations are achieved, the data construction is started, and the pipeline is iterated. The data that has undergone multiple quality inspections is output as SFT training data, which is the financial and tax training data for large language model training that meets the training needs.

[0090] Preferably, the method further comprises:

[0091] Use quality control algorithms for machine testing to perform secondary screening of manually reviewed data;

[0092] The model is evaluated and the evaluation score is fed back to the manual labeling, quality inspection algorithm, and task decomposition parts for joint optimization of manual labeling and the model.

[0093] Combination Figure 2 As shown, in the present invention, the data after data collection, data processing and data construction may have poor data sources, cleaning errors, and policy changes. The data that is finally delivered for large model training still needs to be labeled by professionals and reviewed by experts to ensure the accuracy of the data.

[0094] However, in actual operation, the cost of using a large number of professionals for labeling is very high and impractical when resources are insufficient. It will seriously slow down the speed and frequency of model training and become a speed bottleneck for the entire model training.

[0095] In this link, the invention of this article proposes to use a combination of manual and machine collaboration to improve efficiency, and gradually iterate multiple times to form a flywheel to improve accuracy.

[0096] First, the data will go through manual labeling (QA) and review stages (Q-AUD). At this time, there is only an initial quality inspection algorithm (Q-ALG) that is not very accurate, so the bad cases screened out manually need to be given to the quality inspection algorithm (Q-ALG) to help Q-ALG train with the manually provided data so that it can classify the data.

[0097] In the initial stage, all labeling and quality inspection work needs to be completed by business experts. After the classification information is given to Q-ALG to complete the training, the model evaluation (QE) is used to evaluate the data effect. After the target expectation is achieved, the data construction pipeline begins to iterate. In the initial stage, 10%-20% of the data is tested by Q-ALG, and the rest is mainly tested manually. At the same time, manual feedback of relevant bad cases will be continuously given to Q-ALG. After multiple rounds of optimization and iteration, while ensuring the accuracy of Q-ALG, Q-ALG`, Q-ALG``, etc. are iterated, and the manual workload can be gradually reduced in stages. In the end, the general manual labeling and quality inspection are changed to random inspection, which can greatly reduce labor costs while ensuring accuracy, and at the same time improve the speed of quality inspection. For details, see Figure 4 .

[0098] At the same time, the quality inspection algorithm (Q-ALG) can be used for machine inspection (Q-MD) to perform secondary screening on the manually reviewed data to avoid mislabeling or missing labels due to manual fatigue or inconsistent labeling standards among multiple people.

[0099] The data that has undergone multiple quality checks is output as SFT training data (QD), which can be used as training data for the model. Because the quality of the training data will directly affect the training effect of the model, in order to further optimize the data quality, the model needs to be evaluated (QE). The evaluation score will be fed back to manual labeling, quality inspection algorithms, and task decomposition to help people and models optimize together.

[0100] The quality of data and models is continuously improved through iterative feedback from quality control algorithms (Q-ALG) and manual audits (Q-AUD). The high-quality data outputted will be used for model training and other applications to ensure the performance and accuracy of the model.

[0101] The key points of the present invention are:

[0102] 1. Based on the actual situation and cost assessment, we designed a full-process high-quality training data method for large models in the finance and taxation industry, forming a high-quality training data construction pipeline, which can update the latest industry data in a timely manner and ensure that the model effect is aligned with the latest laws and regulations;

[0103] 2. Optimize the links of quality inspection that can improve efficiency and quality, such as using external large models to build data and using classification models to reduce labor costs while ensuring quality; and propose a strategy that combines manual and machine processing to form a closed loop with manual processing, gradually improve quality and efficiency, and further unify the standards for labeling and quality inspection to avoid increasing data quality variance due to different standards among multiple people. Finally, a flywheel iteration is formed to gradually and efficiently improve the efficiency of data labeling and quality inspection.

[0104] Figure 4 FIG. 4 is a schematic diagram of a structure of a tax training data processing system 400 for a large language model according to an embodiment of the present invention. Figure 4 As shown, the tax training data processing system 400 for a large language model provided by an embodiment of the present invention includes: a data collection unit 401, a data processing unit 402, a data construction unit 403 and a quality inspection unit 404.

[0105] Preferably, the data collection unit 401 is used to collect financial and tax data for financial and tax training to obtain original financial and tax data.

[0106] Preferably, the data collection unit 401 collects the financial and tax data used for financial and tax training to obtain the original financial and tax data, including:

[0107] Capturing industry data at preset time intervals, analyzing the industry data to obtain text data and attachment data, and performing document analysis on the attachment data;

[0108] Obtaining business documents uploaded by the business department and performing document analysis on the business documents;

[0109] The text data, the attachment data and the business documents that have undergone document analysis are used as the original financial and taxation data.

[0110] Preferably, the data processing unit 402 is used to perform data preprocessing on the original financial and tax data, and to detect the financial and tax processing data obtained after the data preprocessing, and to label and classify the financial and tax processing data in multiple dimensions to obtain financial and tax label data.

[0111] Preferably, the data processing unit 402 performs data preprocessing on the original financial and tax data, including:

[0112] Clean the original financial and tax data and unify the format to obtain clean data;

[0113] Performing deduplication processing on the clean data;

[0114] The clean data after deduplication processing is filtered to obtain the financial and tax processing data.

[0115] Preferably, the data construction unit 403 is used to construct initial training data based on the financial and tax label data according to different types of training tasks.

[0116] Preferably, the data construction unit 403 constructs initial training data based on the fiscal and taxation label data according to different types of training tasks, including:

[0117] Obtain complete QA data from the tax label data and use it directly as initial training data;

[0118] When it is determined based on the tax label data that the training data does not meet the quantity required by the training task, the LLM large model and prompt data enhancement technology are used to expand the data to obtain the initial training data;

[0119] When training data cannot be directly obtained based on financial and tax label data, the self-question-answering system based on self-fQA uses the LLM model to generate task data to obtain initial training data.

[0120] Preferably, the quality inspection unit 404 is used to perform quality inspection on the initial training data to obtain financial and tax training data for large language model training that meets training requirements.

[0121] Preferably, the quality inspection unit 404 performs quality inspection on the initial training data to obtain financial and tax training data for large language model training that meets the training requirements, including:

[0122] The initial training data is manually labeled and reviewed, and the screened bad cases are sent to the quality inspection algorithm to train the quality inspection algorithm based on the bad cases, and the model is used to evaluate the data effect. After the target expectations are achieved, the data construction is started, and the pipeline is iterated. The data that has undergone multiple quality inspections is output as SFT training data, which is the financial and tax training data for large language model training that meets the training needs.

[0123] Preferably, the quality inspection unit 404 is further used for:

[0124] Use quality control algorithms for machine testing to perform secondary screening of manually reviewed data;

[0125] The model is evaluated and the evaluation score is fed back to the manual labeling, quality inspection algorithm, and task decomposition parts for joint optimization of manual labeling and the model.

[0126] The finance and taxation training data processing system 400 for a large language model according to an embodiment of the present invention corresponds to the finance and taxation training data processing method 100 for a large language model according to another embodiment of the present invention, and will not be described in detail here.

[0127] The invention has been described above with reference to a few embodiments. However, it is readily apparent to a person skilled in the art that other embodiments than the ones disclosed above are equally within the scope of the invention, as defined by the appended patent claims.

[0128] Generally, all terms used in the claims are to be interpreted according to their ordinary meaning in the technical field, unless explicitly defined otherwise therein. All references to "a / said / the [means, components, etc.]" are to be openly interpreted as at least one instance of said means, components, etc., unless explicitly stated otherwise. The steps of any method disclosed herein do not necessarily have to be performed in the exact order disclosed, unless explicitly stated otherwise.

[0129] Those skilled in the art will appreciate that embodiments of the present invention may be provided as methods, systems, or computer program products. Therefore, the present invention may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0130] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0131] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.

[0132] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.

[0133] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the above embodiments, ordinary technicians in the relevant field should understand that the specific implementation methods of the present invention can still be modified or replaced by equivalents. Any modification or equivalent replacement that does not depart from the spirit and scope of the present invention should be covered within the scope of protection of the claims of the present invention.

Claims

1. A method for processing financial and tax training data for a large language model, characterized in that: The method comprises: Collect the financial and taxation data used for financial and taxation training to obtain the original financial and taxation data; Preprocessing the original financial and tax data, exploring the financial and tax processing data obtained after the data preprocessing, and labeling and classifying the financial and tax processing data in multiple dimensions to obtain financial and tax label data; According to different types of training tasks, constructing initial training data based on the financial and tax label data; The initial training data is subjected to quality inspection to obtain financial and taxation training data for large language model training that meets the training requirements.

2. The method according to claim 1, characterized in that The collecting of the financial and taxation data for financial and taxation training to obtain the original financial and taxation data includes: Capturing industry data at preset time intervals, analyzing the industry data to obtain text data and attachment data, and performing document analysis on the attachment data; Obtaining business documents uploaded by the business department and performing document analysis on the business documents; The text data and the attachment data and business documents that have undergone document analysis are used as the original financial and taxation data.

3. The method according to claim 1, characterized in that: The data preprocessing of the original financial and tax data includes: Clean the original financial and tax data and unify the format to obtain clean data; Performing deduplication processing on the clean data; The clean data after deduplication processing is filtered to obtain the financial and tax processing data.

4. The method according to claim 1, characterized in that According to different types of training tasks, the initial training data is constructed based on the financial and tax label data, including: Obtain complete QA data from the tax label data and use it directly as initial training data; When it is determined based on the tax label data that the training data does not meet the quantity required by the training task, the LLM large model and prompt data enhancement technology are used to expand the data to obtain the initial training data; When training data cannot be directly obtained based on financial and tax label data, the selfQA self-question-answering method is used to generate task data with the help of the LLM model to obtain initial training data.

5. The method according to claim 1, characterized in that: The performing of quality inspection on the initial training data to obtain financial and tax training data for large language model training that meets the training requirements includes: The initial training data is manually labeled and reviewed, and the screened bad cases are sent to the quality inspection algorithm to train the quality inspection algorithm based on the bad cases, and the model is used to evaluate the data effect. After the target expectations are achieved, the data construction is started, and the pipeline is iterated. The data that has undergone multiple quality inspections is output as SFT training data, which is the financial and tax training data for large language model training that meets the training needs.

6. The method according to claim 5, characterized in that The method further comprises: Use quality control algorithms for machine testing to perform secondary screening of manually reviewed data; The model is evaluated and the evaluation score is fed back to the manual labeling, quality inspection algorithm, and task decomposition parts for joint optimization of manual labeling and the model.

7. A financial and tax training data processing system for a large language model, characterized in that: The system comprises: A data collection unit, used for collecting financial and tax data for financial and tax training to obtain original financial and tax data; A data processing unit, used to perform data preprocessing on the original financial and tax data, and to detect the financial and tax processed data obtained after the data preprocessing, and to label and classify the financial and tax processed data in multiple dimensions to obtain financial and tax label data; A data construction unit, used for constructing initial training data based on the financial and tax label data according to different types of training tasks; The quality inspection unit is used to perform quality inspection on the initial training data to obtain financial and tax training data for large language model training that meets training requirements.

8. The system according to claim 7, characterized in that The data collection unit collects the financial and tax data used for financial and tax training to obtain the original financial and tax data, including: Capturing industry data at preset time intervals, analyzing the industry data to obtain text data and attachment data, and performing document analysis on the attachment data; Obtaining business documents uploaded by the business department and performing document analysis on the business documents; The text data and the attachment data and business documents that have undergone document analysis are used as the original financial and taxation data.

9. The system according to claim 7, characterized in that The data processing unit performs data preprocessing on the original financial and tax data, including: Clean the original financial and tax data and unify the format to obtain clean data; Performing deduplication processing on the clean data; The clean data after deduplication processing is filtered to obtain the financial and tax processing data.

10. The system according to claim 7, characterized in that The data construction unit constructs initial training data based on the financial and tax label data according to different types of training tasks, including: Obtain complete QA data from the tax label data and use it directly as initial training data; When it is determined based on the tax label data that the training data does not meet the quantity required by the training task, the LLM large model and prompt data enhancement technology are used to expand the data to obtain the initial training data; When training data cannot be directly obtained based on financial and tax label data, the selfQA self-question-answering system uses the LLM model to generate task data to obtain initial training data.

11. The system according to claim 7, characterized in that The quality inspection unit performs quality inspection on the initial training data to obtain financial and tax training data for large language model training that meets training requirements, including: The initial training data is manually labeled and reviewed, and the screened bad cases are sent to the quality inspection algorithm to train the quality inspection algorithm based on the bad cases, and the model is used to evaluate the data effect. After the target expectations are achieved, the data construction is started, and the pipeline is iterated. The data that has undergone multiple quality inspections is output as SFT training data, which is the financial and tax training data for large language model training that meets the training needs.

12. The system according to claim 11, characterized in that The quality inspection unit is also used for: Use quality control algorithms for machine testing to perform secondary screening of manually reviewed data; The model is evaluated and the evaluation score is fed back to the manual labeling, quality inspection algorithm, and task decomposition parts for joint optimization of manual labeling and the model.

Citation Information

Patent Citations

  • Enterprise intelligent question-answering system data set acquisition method and device

    CN117743540A

  • Training data processing method and device and related equipment

    CN117973566A

  • E-commerce vertical domain large language model training method and device

    CN118394905A

  • Intelligent tax handling method based on tax industry large model

    CN118941400A