A method and system for establishing an information extraction model for fiscal and tax preferential policies

Through the two-level Bert+CRF model, the human resource consumption problem in the extraction of information of fiscal and tax preferential policies has been alleviated, and the efficient identification of custom entity types has been achieved, and the efficiency of information extraction has been improved.

CN114444483BActive Publication Date: 2025-07-25AISINO CORPORATION
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202111639139.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-29
Publication Date
2025-07-25
Estimated Expiration
2041-12-29

AI Technical Summary

Technical Problem

In the prior art, the extraction of information on fiscal and tax preferential policies requires a large amount of labeled data to train models, resulting in huge human resources consumption and it is difficult to efficiently identify custom entity types.

Method used

The two-level Bert+CRF model is adopted to obtain and train the optimal first-level and second-level information extraction model, and use Bert base to connect to the CRF layer to alleviate the data sparse problem and improve the information extraction efficiency.

Benefits of technology

It effectively solves the problem of data sparseness caused by the large number of custom entity types and few labeled data, reduces the human resources required for training the model, and improves the efficiency of information extraction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114444483B_ABST
    Figure CN114444483B_ABST
Patent Text Reader

Abstract

An embodiment of the present invention discloses a method and system for establishing an information extraction model for fiscal and tax preferential policies. The method includes: obtaining a first labeled data set, generating an optimal first-level information extraction model according to the first labeled data set; obtaining a second labeled data set, generating an optimal second-level information extraction model according to the second labeled data set, and using the combination of the optimal first-level information extraction model and the optimal second-level information extraction model as the information extraction model for fiscal and tax preferential policies. Among them, both the optimal first-level information extraction model and the optimal second-level information extraction model are based on a fine-tuned Bert base followed by a CRF layer. The method and system design a two-level Bert+CRF model for the extraction of fiscal and tax preferential policy information, effectively solving the data sparsity problem caused by a large number of custom types and few labeled data when identifying information, and effectively improving the efficiency of information extraction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of information extraction, and in particular, to a method and system for establishing an information extraction model for fiscal and tax preferential policies. Background Art

[0002] At present, taxpayers have great difficulties in judging the applicable conditions and applying for fiscal and tax preferential policies. That is to say, how to match the key taxpayer conditions in the fiscal and tax preferential policies with the taxpayer portraits, so as to accurately recommend the fiscal and tax preferential policies to the taxpayers who are applicable to the policies has become an urgent problem to be solved. In this problem, how to extract the information in the fiscal and tax preferential policies to obtain the above key taxpayer conditions is the most crucial step. Its main function is to extract the more important information in the policy text, such as tax types, industries, regions, taxpayer credit ratings, etc. Currently, common methods include entity recognition models such as LSTM+CRF and lattice LSTM, which can extract three entity types: personal names, place names, and organization names. For the information extraction of policy content, since it is not limited to the above three entity types, but also includes custom entities related to business, such as tax types and tax rates. Therefore, using network models such as LSTM+CRF and lattice LSTM to identify custom entities requires a large amount of labeled data for training, and the network can achieve better results, which consumes a huge amount of human resources. Summary of the Invention

[0003] In order to solve the technical problem in the prior art that a large amount of labeled data is required to train a model for extracting information in fiscal and tax preferential policies, resulting in a huge consumption of human resources, embodiments of the present invention provide a method and system for establishing an information extraction model for fiscal and tax preferential policies.

[0004] According to one aspect of the embodiments of the present invention, a method for establishing an information extraction model for fiscal and tax preferential policies is provided. The method includes:

[0005] Step 101, obtain a first labeled data set, where the first labeled data set is a data set generated after labeling fiscal and tax preferential policy information according to a first information extraction content set in advance;

[0006] Step 102: Input the first labeled dataset into the initial first-level information extraction model for model training to generate the optimal first-level information extraction model. Among them, the initial first-level information extraction model is the publicly available pre-trained model Bert base followed by the initial first CRF layer. The optimal first-level information extraction model is the optimal first pre-trained model Bert base followed by the optimal first CRF layer. The optimal first pre-trained model Bert base is the pre-trained model Bert base obtained by fine-tuning the publicly available pre-trained model Bert base. The optimal first CRF layer is the CRF layer obtained by adjusting the parameters of the initial first CRF layer;

[0007] Step 103: Obtain the second labeled dataset. Among them, the second labeled dataset is the dataset generated after labeling the fiscal and tax preferential policy information according to the preset second information extraction content;

[0008] Step 104: Input the second labeled dataset into the initial second-level information extraction model for model training to generate the optimal second-level information extraction model. Among them, the initial second-level information extraction model is the initial second pre-trained model Bert base followed by the initial second CRF layer. The initial second pre-trained model Bert base is the pre-trained model Bert base obtained by migrating the training parameters of the first N layers of the optimal first pre-trained model Bert base to the first N layers of the publicly available pre-trained model Bert base. The optimal second-level information extraction model is the optimal second pre-trained model Bert base followed by the optimal second CRF layer. The optimal second pre-trained model Bert base is the pre-trained model Bert base obtained by fine-tuning the initial second pre-trained model Bert base. The optimal second CRF layer is the CRF layer obtained by adjusting the parameters of the initial second CRF layer, and N is a natural number;

[0009] Step 105: Use the combination of the optimal first-level information extraction model and the optimal second-level information extraction model as the information extraction model for fiscal and tax preferential policies.

[0010] Optionally, in each of the above method embodiments of the present invention, before obtaining the first labeled dataset, it further includes setting the first information extraction content and the second information extraction content of the fiscal and tax preferential policy information. Among them, the first information extraction content and the second information extraction content are in the key-value structure, and the key value of the second information extraction content belongs to the value value of the first extraction information content.

[0011] Optionally, in each of the above method embodiments of the present invention, the number of network layers L of the publicly disclosed pre-trained model Bertbase adopted by the method is 12, the number of hidden layer nodes H is 768, and the number of self-attention heads A is 12.

[0012] Optionally, in each of the above method embodiments of the present invention, the initial second pre-trained model Bertbase is a pre-trained model Bertbase obtained by migrating the training parameters of the first N layers of the optimal first pre-trained model Bertbase to the first N layers of the publicly disclosed pre-trained model Bertbase, where the value of N is 6.

[0013] According to another aspect of the embodiments of the present invention, a system for establishing an information extraction model for fiscal and tax preferential policies is provided. The system includes:

[0014] A first data module, configured to obtain a first labeled data set, where the first labeled data set is a data set generated by labeling fiscal and tax preferential policy information according to pre-set first information extraction content;

[0015] A first model module, configured to input the first labeled data set into an initial first-level information extraction model for model training to generate an optimal first-level information extraction model, where the initial first-level information extraction model is the publicly disclosed pre-trained model Bertbase followed by an initial first CRF layer, the optimal first-level information extraction model is the optimal first pre-trained model Bertbase followed by an optimal first CRF layer, the optimal first pre-trained model Bertbase is a pre-trained model Bertbase obtained by fine-tuning the publicly disclosed pre-trained model Bertbase, and the optimal first CRF layer is a CRF layer obtained by adjusting the parameters of the initial first CRF layer;

[0016] A second data module, configured to obtain a second labeled data set, where the second labeled data set is a data set generated by labeling fiscal and tax preferential policy information according to pre-set second information extraction content;

[0017] The second model module is used to input the second labeled dataset into the initial second-level information extraction model for model training to generate an optimal second-level information extraction model. Among them, the initial second-level information extraction model is an initial second pre-trained model Bert base followed by an initial second CRF layer. The initial second pre-trained model Bert base is a pre-trained model Bert base obtained by migrating the training parameters of the first N layers of the optimal first pre-trained model Bert base to the first N layers of the publicly available pre-trained model Bert base. The optimal second-level information extraction model is an optimal second pre-trained model Bert base followed by an optimal second CRF layer. The optimal second pre-trained model Bert base is a pre-trained model Bert base obtained by fine-tuning the initial second pre-trained model Bert base. The optimal second CRF layer is a CRF layer obtained by adjusting the parameters of the initial second CRF layer. N is a natural number;

[0018] The model generation module is used to combine the optimal first-level information extraction model and the optimal second-level information extraction model as the information extraction model for fiscal and tax preferential policies.

[0019] Optionally, in each of the above system embodiments of the present invention, the system further includes a parameter setting module for setting the first information extraction content and the second information extraction content of the fiscal and tax preferential policy information. Among them, the first information extraction content and the second information extraction content are in a key-value structure, and the key value of the second information extraction content belongs to the value value of the first extraction information content.

[0020] Optionally, in each of the above system embodiments of the present invention, the publicly available pre-trained model Bert base adopted by the first model module has a network layer number L = 12, a hidden layer node number H = 768, and a self-attention head number A = 12.

[0021] Optionally, in each of the above system embodiments of the present invention, the initial second pre-trained model Bert base in the second model module is a pre-trained model Bert base obtained by migrating the training parameters of the first N layers of the optimal first pre-trained model Bert base to the first N layers of the publicly available pre-trained model Bert base, where the value of N is 6.

[0022] Based on the method and system for establishing an information extraction model of fiscal and tax preferential policies provided in the above embodiments of the present invention, the method includes: obtaining a first labeled data set, inputting the first labeled data set into an initial first-level information extraction model for model training to generate an optimal first-level information extraction model; obtaining a second labeled data set, inputting the second labeled data set into an initial second-level information extraction model for model training to generate an optimal second-level information extraction model, and using the combination of the optimal first-level information extraction model and the optimal second-level information extraction model as the information extraction model of fiscal and tax preferential policies. Among them, both the optimal first-level information extraction model and the optimal second-level information extraction model are based on a fine-tuned Bert base followed by a CRF layer. The method and system design a two-level Bert+CRF model for the extraction of fiscal and tax preferential policy information, effectively solving the problem of data sparsity caused by a large number of custom types and few labeled data when identifying information, avoiding the huge consumption of human resources caused by the need for a large amount of labeled data when training the model, and effectively improving the efficiency of information extraction.

[0023] The technical solution of the present invention will be further described in detail below with reference to the drawings and embodiments. Description of the Drawings

[0024] By describing the embodiments of the present invention in more detail in conjunction with the drawings, the above and other objects, features, and advantages of the present invention will become more obvious. The drawings are used to provide a further understanding of the embodiments of the present invention and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention and do not constitute a limitation to the present invention. In the drawings, the same reference numerals generally represent the same components or steps.

[0025] Figure 1 It is a flowchart of a method for establishing an information extraction model of fiscal and tax preferential policies provided by an exemplary embodiment of the present invention;

[0026] Figure 2 It is a schematic diagram of information extraction from a policy text using an information extraction model provided by an exemplary embodiment of the present invention;

[0027] Figure 3 It is a schematic diagram of the structure of a system for establishing an information extraction model of fiscal and tax preferential policies provided by an exemplary embodiment of the present invention. Detailed Embodiments

[0028] Next, exemplary embodiments according to the present invention will be described in detail with reference to the drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments of the present invention. It should be understood that the present invention is not limited by the exemplary embodiments described herein.

[0029] It should be noted that: Unless otherwise specifically stated, the relative arrangements, numerical expressions, and numerical values of the components and steps set forth in these embodiments do not limit the scope of the present invention.

[0030] Those skilled in the art can understand that terms such as "first", "second", etc. in the embodiments of the present invention are only used to distinguish different steps, devices, or modules, etc., and neither represent any specific technical meaning nor indicate an inevitable logical order between them.

[0031] It should also be understood that in the embodiments of the present invention, "a plurality of" may refer to two or more, and "at least one" may refer to one, two, or more.

[0032] It should also be understood that for any component, data, or structure mentioned in the embodiments of the present invention, in the absence of a clear limitation or a contrary indication in the context, it can generally be understood as one or more.

[0033] In addition, the term "and / or" in the present invention is merely a description of the association relationship between associated objects, indicating that three relationships can exist. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " in the present invention generally represents an "or" relationship between the associated objects before and after.

[0034] It should also be understood that the present invention emphasizes the differences between the various embodiments, and their similarities or similarities can be referred to each other. For the sake of brevity, they will not be elaborated one by one.

[0035] At the same time, it should be understood that for the sake of description, the dimensions of the various parts shown in the drawings are not drawn in accordance with the actual proportional relationship.

[0036] The following description of at least one exemplary embodiment is actually merely illustrative and in no way restricts the present invention and its application or use.

[0037] Techniques, methods, and devices known to those of ordinary skill in the relevant art may not be discussed in detail, but where appropriate, the techniques, methods, and devices should be regarded as part of the specification.

[0038] It should be noted that: Similar reference numerals and letters denote similar items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further discussed in subsequent drawings.

[0039] Embodiments of the present invention can be applied to electronic devices such as terminal devices, computer systems, servers, etc., which can operate with many other general or special computing system environments or configurations. Examples of well-known terminal devices, computer systems, environments, and / or configurations suitable for use with electronic devices such as terminal devices, computer systems, servers, etc. include, but are not limited to: personal computer systems, server computer systems, thin clients, thick clients, handheld or laptop devices, microprocessor-based systems, set-top boxes, programmable consumer electronics, network personal computers, minicomputer systems, mainframe computer systems, and distributed cloud computing technology environments including any of the above systems, and so on.

[0040] Electronic devices such as terminal devices, computer systems, servers, etc. can be described in the general context of computer system-executable instructions (such as program modules) executed by a computer system. Generally, program modules can include routines, programs, object programs, components, logics, data structures, etc., which perform specific tasks or implement specific abstract data types. The computer system / server can be implemented in a distributed cloud computing environment where tasks are executed by remote processing devices linked through a communication network. In a distributed cloud computing environment, program modules can be located on local or remote computing system storage media including storage devices.

[0041] Exemplary method

[0042] Figure 1 It is a schematic flowchart of a method for establishing an information extraction model for fiscal and tax preferential policies provided by an exemplary embodiment of the present invention. This embodiment can be applied to an electronic device, such as Figure 1 As shown, the method for establishing an information extraction model for fiscal and tax preferential policies in this embodiment includes the following steps:

[0043] Step 101, obtain a first labeled data set, where the first labeled data set is a data set generated by labeling fiscal and tax preferential policy information according to a preset first information extraction content;

[0044] Optionally, before obtaining the first labeled data set, it further includes setting a first information extraction content and a second information extraction content for fiscal and tax preferential policy information, where the first information extraction content and the second information extraction content are in a key-value structure, and the key value of the second information extraction content belongs to the value of the first extraction information content.

[0045] In one embodiment, the content to be extracted from the fiscal and tax preferential policy information is set from a business perspective. Since tax preferential policies involve different tax types and different industries and regions, in order to narrow the extraction scope, this patent only focuses on the preferential policies related to the value-added tax type. Combining with the knowledge of the business field, the policy information related to value-added tax is divided into two levels in the form of key-value, as shown in Table 1 for example.

[0046] Table 1 Example of two-level extraction content (key-value) of policy information

[0047]

[0048]

[0049] The purpose of dividing the extracted information into two levels is to solve the problem of data sparsity. Because if only level-2 is considered, there are hundreds of entity types that the network needs to extract, and the value types corresponding to each entity type are relatively single, which will lead to insufficient data required for model training. After generating the policy text for the collected fiscal and tax preferential policy information, the policy text is labeled according to the above two levels respectively, and the first labeled dataset and the second labeled dataset can be obtained.

[0050] Step 102: Input the first labeled dataset into the initial first-level information extraction model for model training to generate the optimal first-level information extraction model. Among them, the initial first-level information extraction model is the publicly available pre-trained model Bert base followed by the initial first CRF layer. The optimal first-level information extraction model is the optimal first pre-trained model Bert base followed by the optimal first CRF layer. The optimal first pre-trained model Bert base is the pre-trained model Bert base obtained by fine-tuning the publicly available pre-trained model Bert base. The optimal first CRF layer is the CRF layer obtained by adjusting the parameters of the initial first CRF layer.

[0051] Optionally, the publicly available pre-trained model Bert base adopted by the method has a network layer number L = 12, a hidden layer node number H = 768, and a self-attention head number A = 12.

[0052] In one embodiment, for multiple publicly available pre-trained models Bert base, in this embodiment, the publicly available pre-trained model Bert base with a network layer number L = 12, a hidden layer node number H = 768, and a self-attention head number A = 12 is selected.

[0053] It is known to those skilled in the art to train an initial first information extraction model with a publicly available pre-trained model Bert base followed by a CRF layer on the premise of having an annotated data set, and to adjust the training parameters of the publicly available pre-trained model Bert base and the CRF layer by analyzing the output results to obtain an optimal first-level information extraction model, which will not be elaborated here.

[0054] Step 103: Obtain a second annotated data set, where the second annotated data set is a data set generated by annotating the fiscal and tax preferential policy information according to the preset second information extraction content.

[0055] Step 104: Input the second annotated data set into the initial second-level information extraction model for model training to generate an optimal second-level information extraction model, where the initial second-level information extraction model is an initial second pre-trained model Bert base followed by an initial second CRF layer, the initial second pre-trained model Bert base is a pre-trained model Bert base obtained by migrating the training parameters of the first N layers of the optimal first pre-trained model Bert base to the first N layers of the publicly available pre-trained model Bert base, the optimal second-level information extraction model is an optimal second pre-trained model Bert base followed by an optimal second CRF layer, the optimal second pre-trained model Bert base is a pre-trained model Bert base obtained by fine-tuning the initial second pre-trained model Bert base, and the optimal second CRF layer is a CRF layer obtained by adjusting the parameters of the initial second CRF layer, and N is a natural number;

[0056] Optionally, the initial second pre-trained model Bert base is a pre-trained model Bert base obtained by migrating the training parameters of the first N layers of the optimal first pre-trained model Bert base to the first N layers of the publicly available pre-trained model Bert base, where the value of N is 6.

[0057] In one embodiment, when the value of N is selected as 6, the information key values in level-2 of Table 1 are summarized into 6 categories (level-1 key values). The first-level information extraction model is used to learn the 6 categories of entity key-values in level-1 (key:value), and the general features learned by the first 6 layers of the pre-trained model Bert base in the first-level information extraction model are migrated to the first 6 layers of the pre-trained model Bert base in the first-level information extraction model for Fined Bert+CRF to learn the hundreds of entity types in level-2. By means of migrating network parameters, the problem of data sparsity during the training of the second-level information extraction model is alleviated.

[0058] Step 105: Combine the optimal first-level information extraction model and the optimal second-level information extraction model as the information extraction model for the fiscal and tax preferential policies.

[0059] Figure 2 It is a schematic diagram of information extraction from policy texts using an information extraction model provided by an exemplary embodiment of the present invention. As Figure 2 shown, in one embodiment, through the generated two-level information extraction model, for the same policy text generated according to the collected fiscal and tax preferential policy information, by respectively inputting the optimal first-level information extraction model and the optimal second-level information extraction model in the information extraction model, then for the same policy text, extraction information at two levels is obtained. Among them, the first level includes the tax credit rating under the conditional element and the tax rate under the business element, and the second level includes two levels of A / B in the tax credit rating and two cases of 3% and 5% in the tax rate.

[0060] As can be seen from the above embodiments, if only one level is used to extract information, since the number of custom entity types to be extracted is large, such as hundreds of types, and the value corresponding to each entity type is relatively single, it will inevitably cause insufficient data required for model training. When using a two-level model to extract information, since the number of entity types set in the first level is small (only 6 categories in this embodiment), the data required for model training is relatively small. However, since the general features learned by the first 6 layers of Bert base in the first-level information extraction model are migrated to the first 6 layers of Bert base in the second-level information extraction model in the form of training parameters and used for learning hundreds of custom entity types involved in the second level, it can better alleviate the data sparsity problem during model training of the second-level information extraction model.

[0061] Exemplary system

[0062] Figure 3 It is a schematic structural diagram of a system for establishing an information extraction model for fiscal and tax preferential policies provided by an exemplary embodiment of the present invention. As Figure 3 shown, the system for establishing an information extraction model for fiscal and tax preferential policies described in this embodiment includes:

[0063] The first data module 301 is used to obtain a first labeled data set, where the first labeled data set is a data set generated by labeling fiscal and tax preferential policy information according to the first information extraction content set in advance;

[0064] The first model module 302 is used to input the first labeled dataset into the initial first-level information extraction model for model training to generate an optimal first-level information extraction model. Among them, the initial first-level information extraction model is the publicly available pre-trained model Bert base followed by an initial first CRF layer. The optimal first-level information extraction model is the optimal first pre-trained model Bert base followed by an optimal first CRF layer. The optimal first pre-trained model Bert base is a pre-trained model Bert base obtained by fine-tuning the publicly available pre-trained model Bert base. The optimal first CRF layer is a CRF layer obtained by adjusting the parameters of the initial first CRF layer;

[0065] The second data module 303 is used to obtain a second labeled dataset. Among them, the second labeled dataset is a dataset generated by labeling the fiscal and tax preferential policy information according to the pre-set second information extraction content;

[0066] The second model module 304 is used to input the second labeled dataset into the initial second-level information extraction model for model training to generate an optimal second-level information extraction model. Among them, the initial second-level information extraction model is the initial second pre-trained model Bert base followed by an initial second CRF layer. The initial second pre-trained model Bert base is a pre-trained model Bert base obtained by migrating the training parameters of the first N layers of the optimal first pre-trained model Bert base to the first N layers of the publicly available pre-trained model Bert base. The optimal second-level information extraction model is the optimal second pre-trained model Bert base followed by an optimal second CRF layer. The optimal second pre-trained model Bert base is a pre-trained model Bert base obtained by fine-tuning the initial second pre-trained model Bert base. The optimal second CRF layer is a CRF layer obtained by adjusting the parameters of the initial second CRF layer, and N is a natural number;

[0067] The model generation module 305 is used to use the combination of the optimal first-level information extraction model and the optimal second-level information extraction model as the information extraction model for fiscal and tax preferential policies.

[0068] Optionally, the system further includes a parameter setting module 306 for setting the first information extraction content and the second information extraction content of the fiscal and tax preferential policy information. Among them, the first information extraction content and the second information extraction content are in the key-value structure, and the key value of the second information extraction content belongs to the value of the first extraction information content.

[0069] Optionally, the publicly available pre-trained model Bert base adopted by the first model module 302 has a network layer number L = 12, a hidden layer node number H = 768, and a self-attention head number A = 12.

[0070] Optionally, the initial second pre-trained model Bert base in the second model module 304 is a pre-trained model Bert base obtained by migrating the training parameters of the first N layers of the optimal first pre-trained model Bert base to the first N layers of the publicly available pre-trained model Bert base, where the value of N is 6.

[0071] Exemplary computer program product and computer-readable storage medium

[0072] In addition to the above methods and devices, an embodiment of the present disclosure may also be a computer program product, which includes computer program instructions. When the computer program instructions are run by a processor, the processor is caused to execute the steps in the method for establishing an information extraction model for fiscal and tax preferential policies according to various embodiments of the present disclosure described in the "Exemplary Method" section of this specification.

[0073] The computer program product can be written in any combination of one or more programming languages for programming code to perform the operations of the embodiments of the present disclosure. The programming languages include object-oriented programming languages such as Java, C++, etc., and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computing device, partially on the user's device, executed as an independent software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server.

[0074] In addition, an embodiment of the present disclosure may also be a computer-readable storage medium, on which computer program instructions are stored. When the computer program instructions are run by a processor, the processor is caused to execute the steps in the method for establishing an information extraction model for fiscal and tax preferential policies according to various embodiments of the present disclosure described in the "Exemplary Method" section of this specification.

[0075] The computer-readable storage medium may adopt any combination of one or more readable media. The readable media may be a readable signal medium or a readable storage medium. The readable storage medium may include, for example, but is not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (a non-exhaustive list) of the readable storage medium include: an electrical connection having one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.

[0076] The basic principles of the present disclosure have been described above in conjunction with specific embodiments. However, it should be noted that the advantages, benefits, effects, etc. mentioned in the present disclosure are only examples and not limitations. It cannot be considered that these advantages, benefits, effects, etc. are essential for each embodiment of the present disclosure. In addition, the above-mentioned specific details are only for illustrative and facilitating understanding purposes and are not limitations. The above details do not limit the present disclosure to necessarily adopt the above specific details for implementation.

[0077] Each embodiment in this specification is described in a progressive manner. The key point of each embodiment is to illustrate the differences from other embodiments. The same or similar parts among the embodiments may be referred to each other. For system embodiments, since they basically correspond to method embodiments, the description is relatively simple. The relevant parts may refer to the partial description of the method embodiments.

[0078] The block diagrams of the devices, apparatuses, equipment, and systems involved in the present disclosure are only illustrative examples and do not intend to require or imply that they must be connected, arranged, and configured in the manner shown in the block diagrams. As those skilled in the art will recognize, these devices, apparatuses, equipment, and systems can be connected, arranged, and configured in any way. Words such as "including", "comprising", "having", etc. are open-ended terms, meaning "including but not limited to", and can be used interchangeably with each other. The words "or" and "and" used herein refer to the word "and / or" and can be used interchangeably with each other, unless the context clearly indicates otherwise. The word "such as" used herein refers to the phrase "such as but not limited to" and can be used interchangeably with each other.

[0079] The methods and apparatuses of the present disclosure may be implemented in many ways. For example, the methods and apparatuses of the present disclosure may be implemented by software, hardware, firmware, or any combination of software, hardware, and firmware. The above order of steps for the methods is for illustration only, and the steps of the methods of the present disclosure are not limited to the specific order described above, unless otherwise specifically stated. In addition, in some embodiments, the present disclosure may also be implemented as a program recorded in a recording medium, and these programs include machine-readable instructions for implementing the methods according to the present disclosure. Thus, the present disclosure also covers a recording medium storing a program for executing the methods according to the present disclosure.

[0080] It should also be noted that in the apparatuses, devices, and methods of the present disclosure, each component or each step can be decomposed and / or recombined. These decompositions and / or recombinations should be regarded as equivalent solutions of the present disclosure. The above description of the disclosed aspects is provided to enable any person skilled in the art to make or use the present disclosure. Various modifications to these aspects are very obvious to those skilled in the art, and the general principles defined herein can be applied to other aspects without departing from the scope of the present disclosure. Therefore, the present disclosure is not intended to be limited to the aspects shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

[0081] The above description has been given for purposes of illustration and description. In addition, this description is not intended to limit the embodiments of the present disclosure to the forms disclosed herein. Although multiple example aspects and embodiments have been discussed above, those skilled in the art will recognize certain variations, modifications, alterations, additions, and sub-combinations thereof.

Claims

1. A method for establishing an information extraction model for fiscal and tax preferential policies, characterized in that, The method includes: Step 101, set the first information extraction content and the second information extraction content of the fiscal and tax preferential policy information. Among them, the first information extraction content and the second information extraction content are in the key-value structure, and the key value of the second information extraction content belongs to the value of the first extraction information content. Step 102, obtain the first labeled dataset. Among them, the first labeled dataset is a dataset generated after labeling the fiscal and tax preferential policy information according to the first information extraction content. Step 103, input the first labeled dataset into the initial first-level information extraction model for model training to generate the optimal first-level information extraction model. Among them, the initial first-level information extraction model is the publicly available pre-trained model Bertbase followed by the initial first CRF layer. The optimal first-level information extraction model is the optimal first pre-trained model Bertbase followed by the optimal first CRF layer. The optimal first pre-trained model Bert base is a pre-trained model Bert base obtained by fine-tuning the publicly available pre-trained model Bertbase. The optimal first CRF layer is a CRF layer obtained by adjusting the parameters of the initial first CRF layer. Step 104, obtain the second labeled dataset. Among them, the second labeled dataset is a dataset generated after labeling the fiscal and tax preferential policy information according to the second information extraction content. Step 105, input the second labeled dataset into the initial second-level information extraction model for model training to generate the optimal second-level information extraction model. Among them, the initial second-level information extraction model is the initial second pre-trained model Bert base followed by the initial second CRF layer. The initial second pre-trained model Bert base is a pre-trained model Bert base obtained by migrating the training parameters of the first N layers of the optimal first pre-trained model Bert base to the first N layers of the publicly available pre-trained model Bert base. The optimal second-level information extraction model is the optimal second pre-trained model Bert base followed by the optimal second CRF layer. The optimal second pre-trained model Bert base is a pre-trained model Bert base obtained by fine-tuning the initial second pre-trained model Bert base. The optimal second CRF layer is a CRF layer obtained by adjusting the parameters of the initial second CRF layer, and N is a natural number. Step 106, use the combination of the optimal first-level information extraction model and the optimal second-level information extraction model as the information extraction model for fiscal and tax preferential policies.

2. The method according to claim 1, characterized in that, The publicly available pre-trained model Bertbase adopted by the method has a network layer number L = 12, a hidden layer node number H = 768, and a self-attention head number A = 12.

3. The method according to claim 2, wherein The initial second pre-trained model Bert base is a pre-trained model Bert base obtained by migrating the training parameters of the first N layers of the optimal first pre-trained model Bert base to the first N layers of the publicly available pre-trained model Bert base, where the value of N is 6.

4. A system for establishing an information extraction model for fiscal and tax preferential policies, characterized in that, The system includes: A parameter setting module for setting the first information extraction content and the second information extraction content of the fiscal and tax preferential policy information, where the first information extraction content and the second information extraction content are in the key-value structure, and the key value of the second information extraction content belongs to the value of the first extraction information content; A first data module for obtaining a first labeled data set, where the first labeled data set is a data set generated by labeling the fiscal and tax preferential policy information according to the first information extraction content; A first model module for inputting the first labeled data set into the initial first-level information extraction model for model training to generate an optimal first-level information extraction model, where the initial first-level information extraction model is the publicly available pre-trained model Bert base followed by an initial first CRF layer, the optimal first-level information extraction model is the optimal first pre-trained model Bert base followed by an optimal first CRF layer, the optimal first pre-trained model Bert base is a pre-trained model Bert base obtained by fine-tuning the publicly available pre-trained model Bert base, and the optimal first CRF layer is a CRF layer obtained by adjusting the parameters of the initial first CRF layer; A second data module for obtaining a second labeled data set, where the second labeled data set is a data set generated by labeling the fiscal and tax preferential policy information according to the second information extraction content; A second model module for inputting the second labeled data set into the initial second-level information extraction model for model training to generate an optimal second-level information extraction model, where the initial second-level information extraction model is the initial second pre-trained model Bert base followed by an initial second CRF layer, the initial second pre-trained model Bert base is a pre-trained model Bert base obtained by migrating the training parameters of the first N layers of the optimal first pre-trained model Bert base to the first N layers of the publicly available pre-trained model Bert base, the optimal second-level information extraction model is the optimal second pre-trained model Bert base followed by an optimal second CRF layer, the optimal second pre-trained model Bert base is a pre-trained model Bert base obtained by fine-tuning the initial second pre-trained model Bert base, and the optimal second CRF layer is a CRF layer obtained by adjusting the parameters of the initial second CRF layer, and N is a natural number; A model generation module for using the combination of the optimal first-level information extraction model and the optimal second-level information extraction model as the information extraction model for fiscal and tax preferential policies.

5. The system according to claim 4, wherein The number of network layers L of the publicly available pre-trained model Bert base adopted by the first model module is 12, the number of hidden layer nodes H is 768, and the number of self-attention heads A is 12.

6. The system according to claim 5, wherein The initial second pre-trained model Bert base in the second model module is a pre-trained model Bert base obtained by migrating the training parameters of the first N layers of the optimal first pre-trained model Bert base to the first N layers of the publicly available pre-trained model Bert base, where the value of N is 6.

Citation Information

Patent Citations

  • Policy file information extraction method based on extended corpus neural network

    CN112257442A