A method for identifying policy information

By defining sentence templates and using language models, the problem of low accuracy in the recognition of similarity of policy documents in the prior art is solved, and the accurate identification of key information in policy documents is achieved and the identification process is simplified.

CN115906842BActive Publication Date: 2025-06-10天道金科股份有限公司
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211232088.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-10-08
Publication Date
2025-06-10
Estimated Expiration
2042-10-08

AI Technical Summary

Technical Problem

The prior art is not very accurate when identifying documents with similar content in policy documents, and is prone to misjudgment, resulting in little reference value for judgment results.

Method used

Provide a policy information identification method, which can accurately identify entity-level content in policy document paragraphs by defining a collection of sentence templates, a policy document element system and a language model. Specific steps include defining sentence templates, filling in text fragments and labels, calculating probability scores using language models, and using the text fragment with the highest score as key information entities.

Benefits of technology

It realizes accurate identification of key information in policy documents, simplifies the difficulty of identifying text entities, and can show excellent recognition effects on small-scale annotation data sets.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115906842B_ABST
    Figure CN115906842B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for identifying policy information, belonging to the technical field of natural language processing. The present invention constructs a complete system of policy document elements, clearly dividing each different element in the policy document. Subsequently, based on this system, it is possible to more accurately classify each paragraph in the policy document and extract key information from the text paragraphs at the entity level. In addition, the provided policy information identifier simplifies the difficulty of text entity identification by predicting missing content labels under the constructed policy text element system, can more accurately extract useful key information from the text based on the constructed policy document element system, and has excellent performance in the case of a small scale of labeled training data set.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of natural language processing, and in particular to a method for identifying policy information. Background Art

[0002] Generally, the text structure of policy documents has a standard to follow, and even the wording has a unified standard. Automatically identifying and analyzing the content and structure of policy documents is particularly important for improving the analysis efficiency of policy documents. In recent years, natural language processing technology has developed rapidly and is mainly applied to machine translation, public opinion monitoring, automatic summarization, opinion extraction, text classification, question answering, text semantic comparison, speech recognition, Chinese OCR, etc. Therefore, for policy documents with structured text content, natural language processing technology is an effective means to analyze the text content of policy documents.

[0003] Policy name, formulating department, implementing department, release time, etc. are important contents in policy documents. Identifying these important contents is of great significance for policy document duplicate checking, analyzing the applicable scope and expiration date of policy documents, and avoiding the introduction of policy documents with overlapping contents. However, how to accurately identify policy documents with similar contents from a large number of various policy documents has become a technical problem to be solved urgently. In the prior art, policy document duplicate checking is usually carried out by content matching. For example, by matching the similarity of policy names of two documents, when the similarity is higher than a preset similarity threshold, it is determined that the two documents are similar, or by matching the content similarity of the documents, when the content similarity of the two documents is greater than the similarity threshold, it is determined that the contents of the two documents are similar. However, in practical applications, it is found that the above two existing matching methods have low accuracy, are prone to misjudgment, and the reference value of the judgment results is not great. Summary of the Invention

[0004] The present invention aims to achieve accurate identification of entity-level content in policy document paragraphs, and provides a method for identifying policy information.

[0005] To achieve this purpose, the present invention adopts the following technical solutions:

[0006] Provide a method for identifying policy information, and the steps include:

[0007] S1, define a set of sentence templates , a set of tag words for entity recognition in the policy document element system , and a language model a set of tags for entity recognition , the set of sentence templates includes sentence templates of entity types and non-entity types , the sentence template It contains two words to be filled in the blanks. The first blank is a text fragment intercepted from the input paragraph The second blank is a category label for classifying the intercepted text fragment. Each label in the label set has a mapping relationship with a label word in the label word set ;

[0008] S2, for each text fragment intercepted from the paragraph and each label corresponding to the label word set in the label words are respectively filled into the first blank and the second blank in each sentence template in the sentence template set . Then, the language model is used to calculate the probability scores of these filled sentences ;

[0009] S3, the text fragment filled in with the highest score is used as the key information entity, and the corresponding type label is mapped to the label word and then used as the corresponding entity type, jointly constituting the key information of the paragraph .

[0010] Preferably, the calculation method is expressed by the following formula (1):

[0011] Formula (1)

[0012] In formula (1), represents the sentence obtained by using the candidate text fragment and the label word having a mapping relationship with the label to fill in the sentence template ;

[0013] represents the sequence length of the sentence ;

[0014] represents the th item in the word sequence of the sentence ;

[0015] represents the sentence from the first item to the​ item;

[0016] representing the paragraph input into the language model thereof;

[0017] representing the first item to the paragraph and the sentence item in the case of the word sequence, the probability that the model predicts the c-th item as is calculated by the pre-trained language model thereof;

[0018] Preferably, the language model is a BART model.

[0019] Preferably, the pre-constructed system of policy document elements includes sentence-level elements and entity-level elements. The sentence-level elements include any one or more of 8 categories and 27 sub-categories, namely policy objectives, application review, supply-type policy tools, environment-type policy tools, demand-type policy tools, fund management, supervision and evaluation, and access conditions.

[0020] Among them, under the category of supply-type policy tools, there are any one or more of 4 sub-categories, namely talent cultivation, financial support, technical support, and public services;

[0021] under the category of environment-type policy tools, there are any one or more of 6 sub-categories, namely regulatory control, target planning, tax incentives, financial support, organizational construction, and policy publicity;

[0022] under the category of demand-type policy tools, there are any one or more of 3 sub-categories, namely government procurement, corporate cooperation, and overseas cooperation;

[0023] under the category of supervision and evaluation, there are 2 sub-categories, namely supervision and management and / or assessment;

[0024] under the category of fund management, there are 2 sub-categories, namely source of funds and / or management principles;

[0025] The entity-level elements include any one or more of 7 categories, namely policy name, policy document number, release area, formulating department, implementing department, release time, and implementation period.

[0026] Preferably, in step S1, the classified paragraph Further extract the key information at the entity level, specifically by using a pre-trained policy text classifier to classify the paragraph The method steps include:

[0027] L1. For the paragraph in the given policy document , use the template function to convert it into the input of the language model . , Add the prompt language for the classification task to the original paragraph . The prompt language contains the masked positions where the labels need to be predicted and filled in;

[0028] L2. The language model predicts the labels to be filled in the masked positions .

[0029] L3. The label converter maps the labels to the corresponding label words in the label word set of the pre-constructed policy document element system as the type of the predicted paragraph .

[0030] Preferably, the method steps for training the language model include:

[0031] A1. For each used as a training sample, calculate the probability score of each label word in the label word set being filled in the masked position . The calculation method of

[0032] is expressed by the following formula (2):

[0033] A2. Calculate the probability distribution through the softmax function . The softmax function (3) is calculated as follows:

[0034] Formula (3)

[0035] In formulas (2) and (3), represents the label in the label set that has a mapping relationship with the label word ; ​​

[0036] Represents the label set for the text classification task;

[0037] A3. According to and , and using the constructed loss function, calculate the model prediction loss. The constructed loss function is expressed by the following formula (4):

[0038] Formula (4)

[0039] In formula (4), represents the fine-tuning coefficient;

[0040] represents the gap between the distribution predicted by the model and the true distribution;

[0041] represents the score predicted by the model and the gap with the true score;

[0042] A4. Determine whether the termination condition for the iterative training of the model is reached.

[0043] If so, terminate the iteration and output the language model ;

[0044] If not, adjust the model parameters and return to step A1 to continue the iterative training.

[0045] Preferably, the language model is a fusion language model formed by fusing several language sub-models . The method for training the fusion language model includes the steps:

[0046] B1. Define a set of template functions . The set of template functions contains several different template functions ;

[0047] B2. For each used as a training sample, through the corresponding language sub-model , calculate the probability score of each label word in the label word set filled into the mask position. The calculation method of

[0048] Formula (5)

[0049] B3. For each associated template function of are fused to obtain , which is obtained by fusing through the following formula (6):

[0050] Formula (6)

[0051] In formula (6), represents the set of said template functions and the said template function in it;

[0052] represents the said template function in the calculation when the weight occupied;

[0053] B4, calculate the probability distribution through the softmax function , The calculation method of which is expressed by the following formula (7):

[0054] Formula (7)

[0055] In formulas (5), (6), and (7), represents the label set and the label that has a mapping relationship with the said label word;

[0056] represents the label set of the text classification task;

[0057] B5, according to and , and using the constructed loss function, calculate the model prediction loss, and the constructed said loss function is expressed by the following formula (8):

[0058] Formula (8)

[0059] In formula (8), represents the fine-tuning coefficient;

[0060] represents the distribution of the model prediction and the gap between the true distribution;

[0061] represents the score of the model prediction and the gap between the true score;

[0062] B6, determine whether the termination condition of the model iterative training is reached,

[0063] If so, terminate the iteration and output the fused language model;

[0064] If not, adjust the model parameters and return to step B2 to continue the iterative training.

[0065] Preferably, the fine-tuning coefficient .

[0066] The present invention has the following beneficial effects:

[0067] 1. A complete system of policy document elements is constructed, clearly separating each different element in the policy document. Based on this system, it is possible to more accurately classify each paragraph in the policy document and extract key information from text paragraphs at the entity level.

[0068] 2. The provided policy information recognizer simplifies the difficulty of text entity recognition by predicting two missing content labels under the constructed system of policy document elements. It can more accurately extract useful key information from the text based on the constructed system of policy document elements and has excellent performance when the scale of the labeled training dataset is small. BRIEF DESCRIPTION OF THE DRAWINGS

[0069] To more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required to be used in the embodiments of the present invention. Obviously, the following described drawings are only some embodiments of the present invention, and those of ordinary skill in the art can also obtain other drawings based on these drawings without creative efforts.

[0070] Figure 1 is a schematic diagram of the system of policy document elements constructed in an embodiment of the present invention;

[0071] Figure 2 is a logic block diagram for predicting the paragraph category of a policy document provided in an embodiment of the present invention;

[0072] Figure 3 is a logic block diagram of a policy information recognizer based on prompt learning provided in an embodiment of the present invention;

[0073] Figure 4 is a logic block diagram of a policy information recognizer based on pre-training and fine-tuning for comparison provided in an embodiment of the present invention;

[0074] Figure 5 is an implementation step diagram of the policy information recognition method provided in an embodiment of the present invention;

[0075] Figure 6 is the method implementation step diagram for a policy text classifier to predict the type of a paragraph . Detailed implementation manners

[0076] The technical solution of the present invention will be further described below with reference to the accompanying drawings and through specific implementation manners.

[0077] Among them, the accompanying drawings are only for illustrative purposes, showing only schematic diagrams rather than physical diagrams, and should not be construed as a limitation of this patent; in order to better illustrate the embodiments of the present invention, some components in the accompanying drawings will be omitted, enlarged or reduced, and do not represent the dimensions of actual products; for those skilled in the art, it is understandable that some well-known structures and their descriptions in the accompanying drawings may be omitted.

[0078] In the accompanying drawings of the embodiments of the present invention, the same or similar reference numerals correspond to the same or similar components; in the description of the present invention, it should be understood that if terms such as "upper", "lower", "left", "right", "inner", "outer", etc. are used to indicate the orientation or positional relationship, it is based on the orientation or positional relationship shown in the accompanying drawings, and is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation. Therefore, the terms describing the positional relationship in the accompanying drawings are only for illustrative purposes and should not be construed as a limitation of this patent. For those of ordinary skill in the art, the specific meanings of the above terms can be understood according to specific circumstances.

[0079] In the description of the present invention, unless otherwise clearly specified and defined, if terms such as "connection" are used to indicate the connection relationship between components, this term should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or integrated; it can be a mechanical connection or an electrical connection; it can be directly connected or indirectly connected through an intermediate medium, and can be the communication inside two components or the interaction relationship between two components. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.

[0080] In the embodiments of the present invention, the applicant has collected a certain number of policy documents as a reference for the construction of the policy document element system and the model training data for the subsequent policy text classifier and policy information recognizer. These policy documents cover various fields such as agriculture, industry, commerce, and service industries, and the applicable objects of the policy documents include individuals, enterprises, institutions, etc. The policy document element system constructed in this embodiment is as Figure 1As shown, according to the character length in the text paragraph, the elements in the system are divided into sentence level and entity level. The elements at the sentence level generally cover the entire sentence in the paragraph. For example, "For enterprises that successfully go public, a reward of 2 million yuan with urban-rural linkage will be given to the management team." This sentence is a complete sentence and is thus recognized as being at the sentence level. While the elements at the entity level generally are included in words with specific meanings in the paragraph, such as policy names, policy document numbers, release regions, formulating departments, etc.

[0081] Furthermore, the elements at the sentence level are further divided into general form and "subject-relation-domain" form. The elements at the sentence level in the general form are used to distinguish the content composition of the paragraph in the policy text, such as Figure 1 the policy objectives, application review, policy tools, supervision and evaluation, fund management, etc. in Figure 1 The elements at the sentence level in the "subject-relation-domain" form are used to structurally represent the access conditions of the policy, such as the access condition related to the enterprise registration location "Enterprise registration location - belongs to - Shanghai". Specifically, as

[0082] 1. The elements at the entity level include: 7 categories of policy name, policy document number, release region, formulating department, implementing department, release time, and implementation period;

[0083] 2. The elements at the sentence level in the general form include: 5 major categories of policy objectives, application review, policy tools, supervision and evaluation, and fund management. Among them, supervision and evaluation are further divided into 2 sub-categories of supervision management and assessment evaluation. Fund management is further divided into 2 sub-categories of fund source and management rules. The policy tools are further divided into the following 3 types with a total of 13 sub-categories:

[0084] Supply-type policy tools (i.e., policy tools - supply-type), including talent cultivation (establishing talent development plans, actively improving various education systems and training systems, etc.), financial support (providing financial support, such as R & D funds and infrastructure construction funds, etc.), technical support (technical guidance and consultation, strengthening technical infrastructure construction, etc.), and public services (improving relevant supporting facilities, policy environment, etc.).

[0085] Environmental policy tools (i.e., policy tools - environmental type), including regulatory control (formulating regulations and standards, regulating market order, and strengthening supervision), target planning (top - level design, providing corresponding policy support services), tax incentives (policy incentives such as tax reduction, exemption, and rebate, including investment tax credit, accelerated depreciation, tax exemption, and tax deduction), financial support (providing loans, subsidies, venture capital, credit guarantees, funds, risk control, and other financial support for enterprises through financial institutions), organizational construction (establishing leadership, supervision, service, etc. organizations and team building to promote the healthy development of industries), and policy publicity (publicizing relevant policies to promote industrial development).

[0086] Demand - type policy tools (i.e., policy tools - demand type), including government procurement (the government purchases products from relevant enterprises), public - private partnership (the government and multiple social entities jointly participate in relevant activities of industrial development, such as joint investment, joint technology research, and development planning research, etc.), and overseas cooperation (introducing foreign investment and conducting cooperation and exchanges with overseas governments, enterprises, or research institutions in aspects such as technology generation and standard setting).

[0087] The sentence - level elements in the form of "subject - relationship - domain" include access conditions, which can be further divided into 8 sub - categories: place of registration, property rights requirements, business scope, employee composition, legal person qualification, enterprise type, operation requirements, and R & D requirements.

[0088] Before classifying paragraphs and identifying key information in policy texts, first split the text content of the policy document into paragraphs. There are many existing methods for splitting the text content of policy documents, and the way of splitting paragraphs is not within the scope of the claims of this invention, so the specific method of paragraph splitting will not be described here.

[0089] After completing the paragraph splitting, enter the paragraph classification and key information identification process. In this embodiment, a pre - trained policy text classifier is used to classify paragraphs, and then analyze the content composition and document structure of the policy document. In this embodiment, select Figure 1 the general - form sentence - level elements in the policy document element system shown in Figure 1There are seven categories in total, including policy objectives, application review, policy tools - supply-oriented, policy tools - environmental, policy tools - demand-oriented, fund management and regulatory assessment; another classification granularity is 17 subcategories after the expansion of the three major categories of policy tools, regulatory assessment, and fund management, and 19 categories in total for the two major categories of policy objectives and application review. When classifying a paragraph, the policy text classifier will also determine whether the paragraph does not belong to any of these categories, that is, whether it is a meaningless paragraph.

[0090] The following is a detailed description of the method for classifying an input paragraph using a pre-trained policy text classifier in this embodiment:

[0091] In this embodiment, the technical core of classifying input paragraphs is to adopt the idea of ​​prompt learning, which can simplify the classification process, improve classification efficiency, and has higher classification superiority for small-scale data sets. Specifically, in order to give full play to the powerful question-answering and reading comprehension capabilities of the policy text classifier, and to mine the deeper information contained in the annotated small-scale policy document text data set, the input paragraph text is processed according to a specific pattern and task prompt language is added to it, making it more suitable for the question-answering form of the language model. The principle of paragraph recognition by the policy text classifier based on prompt learning is as follows:

[0092] set up is a pre-trained language model (preferably a BERT language model), is a set of label words in the policy document element system, mask words Used to fill in the language model The masked positions in the input content, and make It is the label set of the text classification task (paragraph classification task). After tokenizing each policy text paragraph, we get the input language model The word sequence , and then use the custom template function Will Convert to language model Input , exist The prompt language for classification tasks has been added, which contains the mask position that needs to be predicted and filled in with the label. After conversion, the paragraph type prediction problem can be converted into a cloze problem, that is, the language model In the form of a cloze question As input, the most suitable word to fill in the mask position is predicted as Classification prediction results of the expressed paragraph.

[0093] It should be emphasized that based on the idea of prompt learning, this application makes better use of the question-answering and reading comprehension capabilities of the language model , and at the same time, since the classification problem is converted into a cloze problem, the prediction process is simpler, improving the classification efficiency of the policy text classifier. Further, in this embodiment, a mapping from the label set of the text classification task to the label word set in the policy document element system is defined as a label converter . For example, for the label in , this label converter maps it to the label word such as = Figure 1 the policy objective shown in, and "policy objective" is the predicted paragraph category

[0094] Figure 2 is the logical block diagram for predicting the paragraph category of the policy document provided by the embodiment of the present invention. It should be emphasized that for each template function and label converter , this embodiment classifies the paragraph through the following steps:

[0095] Given an input paragraph (preferably the word sequence of the original paragraph), use the template function to convert into the input of the language model , and the language model will predict the most suitable label at the masked position in , then use the label converter to map this label to the label word in the policy document element system , and use it as the classification of the paragraph . Preferably, this embodiment uses a pre-trained Chinese BERT model as the language model , and its prediction method for the masked position follows the pre-training task of the BERT model, that is, uses its output corresponding to the masked position in

[0096] to predict the label at the masked position (the prediction method is the same as the Masked Language Model pre-training task of the BERT model and will not be elaborated). , assuming it is defined as " Overall, this is a policy text paragraph regarding _____. Among them, "_____" represents the masked position, thus adding a hint language for the classification task to the original text paragraph For example, for the sentence "For enterprises that have successfully gone public, a reward of 2 million yuan will be given to the management team through the cooperation between the urban and rural areas", after adding the above hint language to this paragraph the classification task of the language model is to predict the label of the masked position "_____" in "For enterprises that have successfully gone public, a reward of 2 million yuan will be given to the management team through the cooperation between the urban and rural areas. Overall, this is a policy text paragraph regarding _____. After predicting the label of the masked position, map the predicted label to the corresponding label word in the set of label words in the policy document element system as the type of the predicted paragraph .

[0097] The following describes the method for training the language model in this embodiment :

[0098] The language model preferably uses the BERT model. There are many existing training methods for the BERT model, and these existing training methods can be applied to this application to train the language model . The difference is that the samples used to train the language model in this embodiment are those obtained through the template function conversion and the set of label words obtained through the label converter conversion, as well as the loss function for evaluating the model performance improved in this application to improve the classification accuracy.

[0099] When training the language model , this application randomly divides the sample dataset into a training set and a validation set according to a ratio of 7:3. The training process is as follows:

[0100] For each sequence containing only one masked position generated from a policy text paragraph , calculate a score for the probability of filling each label word in the set of label words in the policy document element system into this masked position (since the label has a corresponding mapped label word in the set of label words in the set of label words , so predicting the label ​​​​​​ The probability score filled in this masked position is equivalent to predicting the corresponding labeled word (the probability score filled in this masked position), which is predicted by the language model and represents the possibility that the predicted labeled word can be filled in this masked position. More specifically, for a sequence , the present application calculates the probability score of the labeled word in the label set for the text classification task filled in this masked position, and the method is expressed by the following formula (1): Formula (1)

[0101] In formula (1)

[0102] represents the probability score of the label filled in the masked position. Since the label has a mapping relationship with the labeled word set in the policy document element system , so is equivalent to representing the probability score of the labeled word filled in the masked position; represents the label in the labeled word set that has a mapping relationship with the labeled word

[0103] . For example, the label of the labeled word "policy objective" in can be mapped to , and the label of the labeled word "application review" can be mapped to Figure 1 . By establishing such a mapping relationship, the task of assigning a meaningless label to the input sentence is changed to selecting the word most likely to be filled in the masked position. After calculating the scores of all labeled words in filled in the same masked position, a probability distribution is obtained through the softmax function, and the specific calculation method is expressed by the following formula (2):

[0104] In formula (2) represents the label set of the text classification task;

[0105] Formula (2)

[0106] In formula (2) represents the label set of the text classification task;

[0107] Then, according to and , and using the constructed loss function, the model prediction loss is calculated, and the constructed loss function is expressed by the following formula (3):

[0108] Formula (3)

[0109] In Formula (3), represents the fine-tuning coefficient (preferably 0.0001);

[0110] represents the distribution predicted by the model and the gap with the true one-hot vector distribution;

[0111] represents the score predicted by the model and the gap with the true score;

[0112] Finally, determine whether the termination condition of the model iterative training is reached.

[0113] If so, terminate the iteration and output the language model ;

[0114] If not, adjust the model parameters and continue the iterative training.

[0115] To further improve the model training effect and thus improve the classification performance of the language model Preferably, the language model is a fused language model formed by fusing a number of language sub-models The method for training the fused language model is as follows:

[0116] First, define a set of template functions , and the set of template functions contains a number of different template functions , for example, " . What is this policy text paragraph related to? _____", and another example, "What is this policy text paragraph related to? It is related to _____", and so on. For different template functions , in this embodiment, the following method is used to train the fused language model:

[0117] For each used as a training sample, calculate the probability score of each label word in the label word set filled into the masked position through the corresponding language sub-model . The calculation method is expressed by the following formula (4):

[0118] Formula (4)

[0119] Fuse the associated with each template function to obtain , which is specifically expressed by the following formula (5):

[0120] Formula (5)

[0121] In formula (5), represents the number of template functions in the set of template functions ; represents the weight of the template function

[0122] in the calculation of , , . In this embodiment, according to the accuracies obtained by each language sub-model on the training set and the validation set, the weights of each are determined.

[0123] Then, the probability distribution is calculated through the softmax function, and the calculation method is expressed by the following formula (6):

[0124] Formula (6)

[0125] In formulas (4), (5), and (6), represents the label in the set of label words that has a mapping relationship with the label word ; represents the set of labels for the text classification task;

[0126] Finally, based on and , and using the constructed loss function, the model prediction loss is calculated. The constructed loss function is expressed by the following formula (7):

[0127] Formula (7)

[0128] In formula (7), represents the fine-tuning coefficient (preferably 0.0001);

[0129] represents the gap between the distribution predicted by the model and the true distribution;

[0130] represents the gap between the score predicted by the model and the true score.

[0131] This application provides a with a prompt language as the language model The input mask position label prediction method has excellent prediction performance when the size of the labeled training dataset is small. To verify its excellent performance when the training data is scarce, this application also designs multiple policy text classifiers based on fully supervised learning for performance comparison. The specific methods include:

[0132] (1) For a policy document paragraph , use a word segmentation tool to obtain a word sequence, denoted as , represents the word sequence in the -th word. Then, use the word vector representation model pre-trained on a large-scale comprehensive domain corpus to perform distributed representation on each word after word segmentation. In this embodiment, static word vectors are used, and each word is represented as a 300-dimensional pre-trained vector , represents the word sequence in the -th word. After obtaining the feature representation of the paragraph through the word vector, input the feature representation of the paragraph into a multi-classifier to predict the probability that each paragraph belongs to each category. The prediction process is expressed as: , , is the feature representation function, represents the probability that the paragraph belongs to the -th category. Select the category with the highest probability as the category described for the paragraph.

[0133] (2) In the multi-classifier part, this application selects methods based on statistical machine learning and deep learning to perform fully supervised learning on the multi-classifier. Among them, the multi-classifier based on statistical machine learning is designed based on the support vector machine model and the XGBoost model; the multi-classifier based on deep learning is designed based on the TextCNN model and the Bi-LSTM+Attention model.

[0134] 1) In the multi-classifier based on statistical machine learning, for a policy text paragraph , take the average value of each dimension of the 300-dimensional distributed representation of all words in the word-segmented paragraph, and concatenate the two features of the length of the paragraph and its relative position in the entire policy document (the index value of the paragraph in the document / the total number of segments in the document) to obtain a 302-dimensional feature vector , and input it into the multi-classifier to output the classification label of the paragraph.

[0135] 2) In the multi-classifier based on deep learning, for a policy text paragraph , concatenate the distributed representations of all words in the segmented paragraph into a matrix, and use three different-sized convolutional kernels to extract features. The sizes of the three convolutional kernels can be 3×3, 4×4, and 5×5 respectively. After convolution, perform max pooling. Finally, concatenate the features extracted by the convolutional kernels of different sizes into a feature vector, input it into the softmax activation function, and output the label for the classification of this paragraph.

[0136] 3) In another multi-classifier based on deep learning, for a policy text paragraph , the 300-dimensional distributed representations of all words in the segmented paragraph are input forward into the LSTM long short-term memory network to obtain , and input backward into the LSTM to obtain , and add the elements corresponding to the same time sequence of the two to obtain the output vector for each time sequence . Then, through the Attention mechanism, calculate the weights for each time sequence and perform weighted summation of the vectors for all time sequences as the feature vector. Finally, use the softmax function for classification.

[0137] The following shows the multi-classifiers trained by the four methods (1) and (1), (2), (3) in method (2) on a small-scale training dataset and the language model trained by the policy text classification method based on prompt language and masked position label prediction provided by the embodiments of the present invention for Figure 1 the comparison table of the classification effects of the paragraphs in the two different granularity policy documents of the 7 major categories of "policy objectives, application review, policy tools - supply type, policy tools - environmental type, policy tools - demand type, regulatory evaluation, and fund management" and the 19 categories of "policy objectives, application review, talent cultivation, financial support, technical support, public services, regulatory control, goal planning, tax incentives, financial support, organizational construction, policy publicity, government procurement, public-private partnership, overseas cooperation, supervision and management, assessment and evaluation, source of funds, and management principles". The evaluation index is the accuracy rate on the test set. It can be seen from Table a below: The language model trained in this embodiment in the paragraph the paragraph text classification method that adds classification task prompt language for masked position label prediction shows better paragraph classification performance than the multi-classifiers trained by the other four methods on a small-scale dataset, proving the superiority of the language model trained in this embodiment in predicting paragraph categories on a small-scale dataset.

[0138]

[0139] Table a

[0140] After completing the paragraph classification of the policy text, it is sometimes necessary to automatically identify the key information in each paragraph. In this application, a policy information recognizer based on prompt learning is used to identify the key information in the policy document. In this application, the entity-level elements in the policy document element system shown in Figure 1 are defined as the set of key information categories of the policy, that is, Figure 1 the seven categories of "policy name, policy document number, release area, formulating department, implementing department, release time, and implementation period" shown in

[0141] The following specifically elaborates on the method of extracting the key information in each paragraph by the policy information recognizer based on prompt learning: in

[0142] Generally speaking, in this application, each paragraph is regarded as a character sequence, and the policy information recognizer is used to identify whether each bit in the character sequence is an entity boundary and the type of the entity. Specifically, as Figure 3 shown, set as the pre-trained language model. In the model , is the set of label words for entity recognition in the policy document element system, and let be the label set for the entity recognition task. Each label in the label set has a label word with a mapping relationship in the label word set , and define the sentence template . The template contains two vacancies for filling in words. The content filled in the first vacancy is the text fragment intercepted from the input paragraph, and these fragments are regarded as candidate entities. The second vacancy is the entity category label of the text fragment to be filled in and predicted. For each label word in the set of label words for entity recognition in the policy document element system , for the entity type represented by this label word , fill in this entity type into to define a new template. For example, define the sentence template as "[text fragment] is a [entity type] policy entity". Then, for the "formulating department" entity type in the set of label words for entity recognition , fill it into the template After that, a new template can be defined, for example, "[Candidate entity] is an entity that formulates department policies". In addition, in order to handle the case where the text fragment is not an entity, a sentence template of the "non-entity" type is defined, that is, "[Text fragment] is not a policy entity". In this way, multiple sentence templates of different entity types and the sentence template of the non-entity type form a set of sentence templates 。

[0143] Each text fragment intercepted from the paragraph is filled into each sentence template in the set of sentence templates Then, the probability scores of these filled sentences are calculated using the language model (also preferably the BART model), and the calculation method is expressed by the following formula (8): Formula (8)

[0144] In formula (8),

[0145] represents the sentence obtained after filling the candidate text fragment and the label word with a mapping relationship with the label into the sentence template ; ;

[0146] represents the sequence length of the sentence ;

[0147] represents the -th item in the word sequence of the sentence ;

[0148] represents the sentence from the 1st item to the -th item in the word sequence;

[0149] represents the text sequence input into the language model ;

[0150] represents the probability that the model predicts the -th item to be given the input text and the word sequence of the sentence template from the 1st item to the -th item, and this probability is calculated by the pre-trained generative language model.

[0151] Through the above process, the language model A probability score for filling the tag word in the second blank is calculated for each sentence template of entity types and non-entity types, and then each candidate text segment is classified into the type corresponding to the sentence template with the highest score, which may also be "non-entity". The text segments assigned with entity types are the entities identified in this text, and their entity types are the assigned entity types.

[0152] The following briefly describes the method for training the policy information recognizer:

[0153] With and the corresponding true tag word as the model training samples, the sample dataset is randomly divided into a training set and a validation set according to a ratio of 7:3. For the data in the training set, if the entity type of the text segment is , then and are respectively filled into the first blank and the second blank of the sentence template of the entity type . If the text segment is not an entity, then is filled into the sentence template of the non-entity type , and a filled sentence is also obtained. In addition, all entity samples in the training set are used in this application to fill the sentence templates containing entities, and the non-entity sentence templates are filled by randomly sampling from the remaining non-entity type words. The ratio of the two is preferably 1:1.5 to increase the interference of non-entity sentence templates on the recognition of entity sentence templates, thereby improving the key information extraction accuracy of the policy information recognizer.

[0154] It should be emphasized that in this application, the language model is preferably the BART model. The principle of the BART model for calculating the score of the sentence template is as follows:

[0155] Given a policy text paragraph and a set of sentence templates , is input into the encoder of the BART model to obtain the feature representation of the paragraph . In each step of the decoder of the BART model, and the output before the decoder are used together as the input of the current step, and the attention method is used to obtain the feature representation of the current step. After linearly transforming this feature representation, the softmax function is used to obtain the output word The conditional probability of (the probability distribution of the cth item given the previous c-1 items and the input paragraph) is calculated as ,in is the model parameter.

[0156] In training the BART model, the cross entropy loss function is used to calculate the gap between the decoder output and the true template, which is used as the basis for adjusting the model parameters. After adjusting the model parameters, the BART model is iteratively trained until the model convergence conditions are reached.

[0157] The policy information extraction method based on prompt learning provided in this application has excellent recognition effect on small-scale data sets. In order to verify its performance when the training data set is small, this application also designs a variety of policy information identifiers based on pre-training and fine-tuning to compare their performance on the same data set. The specific method is as follows: Figure 4 As shown, including:

[0158] In the distributed feature representation part of the input data of the policy information identifier, both the vocabulary-level and character-level distributed feature representations are used. The distributed feature representation of each word at the vocabulary level is realized by a word vector representation model pre-trained on a large-scale comprehensive domain corpus, while the distributed features of each character at the character level are realized by a pre-trained Chinese RoBERTa model. Since the process of distributed feature representation of input data by the word vector representation model and the Chinese RoBERTa model is not within the scope of the rights claimed in this application, the specific process is not described.

[0159] The context encoding layer of the policy information identifier receives the output of the distributed representation layer and further models the text semantics and the dependencies between words. In this embodiment, multi-layer perceptron, Transformer and Flat-Lattice Transforme are used. The structures and construction methods of the three models are briefly described as follows:

[0160] In the context encoding layer based on the multi-layer perceptron, a structure of linear layer-ReLU function layer-linear layer is adopted.

[0161] In the Transformer-based context encoding layer, the Transformer Encoder is used to encode the text features.

[0162] In the context encoding layer based on Flat-Lattice Transformer (FLAT), a variant of Transformer, FLAT, is used. The distributed representation of text characters and words is used at the same time. The position encoding in Transformer is further expanded, and the relative position of text characters and words is introduced to try to better overcome the problem of imbalanced length of policy document entities. The calculation method of relative position encoding of FLAT text fragments is expressed by the following formula (9):

[0163] Formula (9)

[0164] In formula (9), and Respectively represent The position index of the first and last characters of a text segment in the original sequence. For a character, the position index of the first and last characters is the same (head and tail are used to indicate where the text segment starts and ends. For example, in the text "The validity period of the policy is 3 years", the head and tail of "policy" are 1 and 2 respectively; and for the character "政", the head and tail are both 1). is a learnable parameter, include , The calculation method of is expressed by the following formulas (10) and (11):

[0165] Formula (10)

[0166] Formula (11)

[0167] In formulas (10) and (11), include , , , Any of the following; Represents the length of the vector of input models.

[0168] The decoding layer of the policy information identifier uses a conditional random field model. The decoding process uses the Viterbi algorithm based on dynamic programming to obtain higher decoding efficiency, and uses the conditional random field loss function for optimization.

[0169] The following shows the policy information identifier based on pre-training and fine-tuning and the policy information identifier based on prompt learning provided by the embodiment of the present invention when the scale of the labeled training data set is small. Figure 1Comparison table of the extraction effects of the seven entity-level policy information of "policy name, policy document number, release area, formulating department, implementing department, release time, and implementation period" shown in the figure. The evaluation index is the F1 score on the test set. It can be seen from Table b below that the language model N trained in this embodiment shows better performance than the policy information recognizers trained by other methods on a small-scale training dataset, proving the superiority of the language model N trained in this embodiment in identifying policy key information when there are few labeled training datasets.

[0170]

[0171] Table b

[0172] In summary, the policy information recognition method provided by the embodiment of the present invention, as Figure 5 shown, the steps include:

[0173] S1, define a set of sentence templates , a set of labeled words for entity recognition in the policy document element system , and a language model a set of labels for entity recognition , the set of sentence templates includes sentence templates of entity types and non-entity types , the sentence template contains two blanks to be filled with words, where the first blank is a text fragment intercepted from the input paragraph , and the second blank is a category label for classifying the intercepted text fragment. Each label in the label set has a corresponding labeled word in the set of labeled words ;

[0174] S2, fill each text fragment intercepted from the paragraph and each label with the corresponding labeled word in the set of labeled words into the first blank and the second blank of each sentence template in the set of sentence templates , and then use the language model to calculate the probability scores of these filled sentences;

[0175] S3, take the text fragment filled with the highest score as the key information entity, and map the corresponding type label to the labeled word and then use it as the corresponding entity type, and jointly form the paragraph key information.

[0176] More specifically, step S1 is to sort the paragraphs that have been classified. Further extract the key information at the entity level, such as Figure 6 As shown, the paragraph is classified by the pre-trained policy text classifier Classification, the method steps include:

[0177] L1, for a given paragraph in a policy document , using template functions Will Convert to language model Input , In the original paragraph Added prompt language for classification tasks, which contains the mask position that needs to be predicted and filled in with labels;

[0178] L2, language model Predict the label of the mask position ;

[0179] L3, label converter Label Mapped to a set of label words of the pre-built policy document element system The corresponding label words in As predicted paragraph Type.

[0180] In summary, the present invention has the following beneficial effects:

[0181] 1. A complete policy document element system was built to clearly divide the different elements in the policy document. Based on this system, the classification of each paragraph type in the policy document and the extraction of key information of the text paragraph at the entity level can be achieved more accurately.

[0182] 2. By in the original paragraph A prompt language for classification tasks is added, which contains the mask positions that need to be predicted and filled in with labels, converting the paragraph classification problem into a classification prediction problem similar to cloze test, simplifying the paragraph classification prediction process, and being able to more accurately parse the policy document text from the perspective of content composition and file structure based on the constructed complete policy document element system, and dig out deeper information, and has excellent performance when the scale of the annotated training data set is small.

[0183] 3. The provided policy information recognizer simplifies the difficulty of text entity recognition by predicting two missing content tags under the constructed policy document element system, can more accurately extract useful key information from the text based on the constructed policy document element system, and has excellent performance when the scale of the labeled training data set is small.

[0184] It should be noted that the above specific implementation manners are only the preferred embodiments of the present invention and the applied technical principles. Those skilled in the art should understand that various modifications, equivalent replacements, changes, etc. can be made to the present invention. However, as long as these transformations do not deviate from the spirit of the present invention, they should be within the protection scope of the present invention. In addition, some terms used in the specification and claims of this application are not restrictive, but are only for the convenience of description.

Claims

1. A method for identifying policy information, characterized in that the steps include: S1, define a set of sentence templates , a set of labeled words for entity recognition in the policy document element system , and a language model , a set of labels for entity task recognition in , the set of sentence templates includes sentence templates of entity types and non-entity types , the sentence template contains two blanks to be filled with words. The first blank is a text fragment intercepted from the input paragraph , and the second blank is a category label for classifying the intercepted text fragment. Each label in the set of labels has a corresponding labeled word with a mapping relationship in the set of labeled words in ; ; S2, for each text segment intercepted from the paragraph and each label corresponding to the label word set fill the label words into the first vacancy and the second vacancy in each sentence template in the sentence template set respectively, and then use the language model to calculate the probability scores of these filled sentences ; ; S3. Take the filled text fragment with the highest score as the key information entity, and map the corresponding type label to the label word and then use it as the corresponding entity type to jointly form the key information of the paragraph; The text segment using the candidate is indicated and the label and the label word with a mapping relationship are filled into the sentence template to obtain the sentence 2. The method for identifying policy information according to claim 1, characterized in that The calculation method is expressed by the following formula (1): Formula (1) In formula (1), represents using the candidate text segment and the label word with a mapping relationship with the label filled into the sentence template to obtain the sentence; Indicates the sentence sequence length; denote the n-th item in the word sequence of the represent the first item to the nth item in the word sequence of the said sentence;​ represent the paragraphs input into the language model thereof; represents the first item to the given paragraph of the input and the sentence in the word sequence, in the case of the c-th item, the probability predicted by the model is calculated by the pre-trained language model.

3. The method for identifying policy information according to claim 1, characterized in that The language model is the BART model.

4. The method for identifying policy information according to claim 1, characterized in that The pre-constructed policy document element system includes sentence-level elements and entity-level elements. The sentence-level elements include any one or more of 8 categories and 27 sub-categories such as policy objectives, application review, supply-type policy tools, environment-type policy tools, demand-type policy tools, fund management, supervision and evaluation, and access conditions. Among them, under the category of supply-type policy tools, it includes any one or more of 4 sub-categories such as talent cultivation, financial support, technical support, and public services; Under the category of environment-type policy tools, it includes any one or more of 6 sub-categories such as regulatory control, target planning, tax incentives, financial support, organizational construction, and policy publicity; Under the category of demand-type policy tools, it includes any one or more of 3 sub-categories such as government procurement, corporate cooperation, and overseas cooperation; Under the category of supervision and evaluation, it includes 2 sub-categories of supervision and management and / or assessment and evaluation; Under the category of fund management, it includes 2 sub-categories of fund source and / or management principle; The entity-level elements include any one or more of 7 categories such as policy name, policy document number, release area, formulating department, implementing department, release time, and implementation period.

5. The method for identifying policy information according to any one of claims 1-4, characterized in that In step S1, the classified passage is Further extract the key information at the entity level. Specifically, use a pre-trained policy text classifier to classify the passage The method steps include: L1, for the paragraph in the given policy document , use the template function to convert it into the input of the language model , , add the prompt language for the classification task to the original paragraph , and the prompt language contains the masked positions that need to be predicted and filled with labels; L2, the language model predicts the label to be filled in the masked position ; L3, Label Converter Map the said label to the set of label words in the pre-constructed policy document element system corresponding label words as the type of the predicted paragraph of.

6. The method for identifying policy information according to claim 5, characterized in that Training the language model The method steps include: A1, for each of the training samples , calculate each label word in the set of label words and fill in the probability score of the label word at the masked position . The calculation method is expressed by the following formula (2): Formula (2) A2, calculate the probability distribution through the softmax function , calculate through the softmax function (3): Formula (3) In Formulas (2) and (3), represents the label set and the label word has a mapping relationship with the label; Represents the label set for the text classification task; A3, according to and , and using the constructed loss function, calculate the model prediction loss. The constructed loss function is expressed by the following formula (4): Formula (4) In formula (4), represents the fine-tuning coefficient; Represents the distribution predicted by the model The gap between the true distribution; The score predicted by the model The gap with the true score; A4. Judge whether the termination condition of model iterative training is reached. If so, terminate the iteration and output the language model ; If not, adjust the model parameters and return to step A1 to continue iterative training.

7. The method for identifying policy information according to claim 5, characterized in that The language model is a fused language model formed by fusing a number of language sub-models The method for training the fused language model includes the steps of: B1, define a set of template functions , the set of template functions includes several different ones of the said template functions ; B2, for each of the training samples , through the corresponding language sub-model , calculate each label word in the set of label words The probability score filled in the masked position , The calculation method is expressed by the following formula (5):​ Formula (5) B3, for associating each of the said template functions of are fused to obtain , obtained by fusing with the following formula (6): Formula (6) In formula (6), represents the number of template functions in the set of the template functions; representing the said template function during the calculation of the weight occupied B4. Calculate the probability distribution through the softmax function , The calculation method is expressed by the following formula (7): Formula (7) In Formulas (5), (6), and (7), represents the label set and the label word has a mapping relationship; Represents the label set for the text classification task; B5, according to and , and using the constructed loss function, calculate the model prediction loss, and the constructed loss function is expressed by the following formula (8): Formula (8) In formula (8), represents the fine-tuning coefficient; Indicates the distribution predicted by the model The gap with the true distribution; The score predicted by the model The gap with the true score; B6. Judge whether the termination condition of model iterative training is reached. If so, terminate the iteration and output the fusion language model; If not, adjust the model parameters and return to step B2 to continue iterative training.

8. The method for identifying policy information according to claim 6 or 7, characterized in that Fine-tuning coefficient .

Citation Information

Patent Citations

  • Method and device for intelligent triage, electronic device and storage medium

    CN110197730A

  • Method and terminal for automatically generating emotion text

    CN115017876A