Policy text annotation method and system based on joint pre-training and graph neural network
By jointly pre-training the language model and graph neural network, the structural and semantic information of the policy text is obtained, and the problem of insufficient subjectivity and accuracy in policy text annotation is solved, and efficient and accurate policy text annotation is achieved.
Patent Information
- Application Number
- CN202211116359.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-09-14
- Publication Date
- 2025-08-22
- Estimated Expiration
- 2042-09-14
AI Technical Summary
The existing technology has strong subjectivity and unstable labeling accuracy in policy text annotation, and cannot effectively utilize the structural and semantic information of policy text, resulting in insufficient classification accuracy.
The combined pre-training and graph neural network method is adopted to obtain structural information of policy text through graph neural network, pre-training language models to obtain semantic information, and fusion operations are combined to construct text-level graph structures, extract semantic and structural features, and finally determine the annotation results.
It improves the accuracy and efficiency of policy text annotation, reduces the cost of manual annotation, reduces the waste of computing resources, can process large amounts of text without reducing the accuracy rate, and avoids errors caused by personal subjectivity.
Smart Images

Figure CN115374792B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data processing technology, and in particular to a policy text annotation method and system using joint pre-training and graph neural network. Background Art
[0002] The statements in this section merely mention background art related to the present invention and do not necessarily constitute prior art.
[0003] Current policy text information is complex and diverse in structure, with varying lengths, high information density, and inconsistent classification systems. Currently, policy text annotation primarily relies on subjective manual annotation. This drawback is that, for certain issues without clear annotation standards, different judgments can be made based on subjective factors. Another approach is to use a single model from deep learning, but this method underrepresents policy text information and fails to integrate its structural and semantic context, limiting its accuracy. Therefore, analyzing and annotating policy text information from the perspective of semantic integrity and accuracy is crucial for accurately annotating policy text. Summary of the Invention
[0004] To address the shortcomings of existing technologies, this invention provides a policy text annotation method and system that combines pre-training and graph neural networks. This method comprehensively utilizes the structural and semantic information of policy texts. The graph neural network is used to capture the text's structural information, while the pre-trained language model focuses on capturing the text's semantic information. This allows for better extraction, classification, and annotation of the basic information in policy documents, providing business owners with more accurate information about their respective policies. This addresses the subjectivity inherent in manual annotation and the instability of annotation accuracy inherent in existing annotation systems, which suffer from the single, single-algorithm model.
[0005] In the first aspect, the present invention provides a policy text annotation method that combines pre-training and graph neural networks;
[0006] The policy text annotation method of joint pre-training and graph neural network includes:
[0007] Obtain the policy text to be annotated and preprocess the policy text to be annotated;
[0008] Input the preprocessed policy text into the trained policy text annotation model and output the annotation results of the policy text;
[0009] Among them, the working principle of the trained policy text annotation model includes: extracting word vectors and sentence vectors for the processed policy text; constructing a text-level graph structure based on the preprocessed policy text, and obtaining the adjacency matrix corresponding to the text-level graph structure; extracting the semantic features of the policy text based on word vectors and sentence vectors; extracting the structural features of the policy text based on word vectors and adjacency matrix; and determining the policy text annotation results based on semantic features and structural features.
[0010] In a second aspect, the present invention provides a policy text annotation system that combines pre-training and graph neural networks;
[0011] A policy text annotation system that combines pre-training and graph neural networks, including:
[0012] An acquisition module is configured to: acquire the policy text to be annotated and pre-process the policy text to be annotated;
[0013] An annotation module is configured to: input the preprocessed policy text into the trained policy text annotation model and output the annotation results of the policy text;
[0014] Among them, the working principle of the trained policy text annotation model includes: extracting word vectors and sentence vectors for the processed policy text; constructing a text-level graph structure based on the preprocessed policy text, and obtaining the adjacency matrix corresponding to the text-level graph structure; extracting the semantic features of the policy text based on word vectors and sentence vectors; extracting the structural features of the policy text based on word vectors and adjacency matrix; and determining the policy text annotation results based on semantic features and structural features.
[0015] In a third aspect, the present invention further provides an electronic device, comprising:
[0016] a memory for non-transitory storage of computer-readable instructions; and
[0017] a processor for executing said computer-readable instructions,
[0018] When the computer-readable instructions are executed by the processor, the method described in the first aspect is executed.
[0019] In a fourth aspect, the present invention further provides a storage medium that non-temporarily stores computer-readable instructions, wherein when the non-temporary computer-readable instructions are executed by a computer, the instructions of the method described in the first aspect are executed.
[0020] In a fifth aspect, the present invention further provides a computer program product, comprising a computer program, wherein the computer program is used to implement the method described in the first aspect when running on one or more processors.
[0021] Compared with the prior art, the present invention has the following beneficial effects:
[0022] (1) The present invention jointly trains a pre-trained language model and a graph convolutional network in a learnable manner, and simultaneously learns the structural information and semantic information of the text; wherein, the graph neural network is used to obtain the structural information of the text, and the pre-trained language model focuses on obtaining the semantic information of the text, and then the results of the graph neural network and pre-training are output through a fusion operation.
[0023] (2) The present invention only uses words to construct an isomorphic graph for each policy text, and can simultaneously perform small-batch inductive training on the pre-trained language model and the graph convolutional network, reducing the occupation of system resources and enabling the induction of new words and new texts;
[0024] (3) The pre-trained language model and graph convolutional network are jointly trained in a learnable way, and the structural and semantic information of the text are learned at the same time. At the same time, a homogeneous graph is constructed for each text. Small batch induction training can be performed on the pre-trained language model and the graph convolutional network at the same time, which reduces the occupation of system resources and can achieve the induction of new words and new texts.
[0025] (4) The cost of manual labeling of policy classification is greatly reduced, while also avoiding excessive waste of computing resources. The labeling method of the present invention is more accurate and efficient than manual standards. It will not reduce the labeling accuracy due to excessive text information. It also avoids errors in the labeling of policy documents caused by different personal subjectivities when using manual labeling. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] The accompanying drawings, which constitute a part of the present invention, are used to provide a further understanding of the present invention. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute improper limitations on the present invention.
[0027] Figure 1 is a flow chart of a model method according to a first embodiment of the present invention;
[0028] Figure 2 This is a model structure diagram of the first embodiment of the present invention;
[0029] Figure 3 This is a structural diagram of the feature joint output layer in Example 1 of the present invention;
[0030] Figure 4 This is a diagram of the neural network layer structure in Example 1 of the present invention. DETAILED DESCRIPTION
[0031] It should be noted that the following detailed descriptions are exemplary and intended to provide further explanation of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which the present invention belongs.
[0032] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit exemplary embodiments according to the present invention. As used herein, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. In addition, it should be understood that the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0033] In the absence of conflict, the embodiments of the present invention and the features thereof may be combined with each other.
[0034] All data in this embodiment is obtained in compliance with laws and regulations and based on the consent of the user, and is used legally.
[0035] Example 1
[0036] This embodiment provides a policy text annotation method that combines pre-training and graph neural networks;
[0037] like Figure 1 As shown in Figure 2, the policy text annotation method of joint pre-training and graph neural network includes:
[0038] S101: Obtain the policy text to be annotated and pre-process the policy text to be annotated;
[0039] S102: Input the pre-processed policy text into the trained policy text annotation model, and output the annotation results of the policy text;
[0040] The working principle of the trained policy text annotation model includes the following:
[0041] Extract word vectors and sentence vectors from the processed policy text;
[0042] Construct a text-level graph structure based on the preprocessed policy text and obtain the adjacency matrix corresponding to the text-level graph structure;
[0043] Extract semantic features of policy text based on word vectors and sentence vectors;
[0044] Extract the structural features of policy text based on word vectors and adjacency matrix;
[0045] Determine the policy text annotation results based on semantic features and structural features.
[0046] Furthermore, if Figure 2 As shown in FIG, the trained policy text annotation model has a model structure including:
[0047] First pre-trained language model and text-level graph construction layer;
[0048] The input end of the first pre-trained language model and the input end of the text-level graph construction layer are both used to input the pre-processed policy text; the output of the first pre-trained language model is the word vector of the 1st, 5th, 9th, and 12th hidden layers; the output end of the text-level graph construction is the adjacency matrix represented by the graph structure of a single policy text;
[0049] The output of the first pre-trained language model is connected to the input of the second pre-trained language model, the output of the second pre-trained language model is connected to the input of the fully connected layer, the output of the fully connected layer is connected to the input of the first sigmoid activation function layer, and the output of the first sigmoid activation function layer is connected to the input of the joint output layer;
[0050] The output of the text-level graph construction layer and the output of the first pre-trained language model are both connected to the input of the graph neural network layer; the output of the graph neural network layer is connected to the input of the maximum pooling layer, the output of the maximum pooling layer is connected to the input of the sigmoid activation function layer, and the output of the second sigmoid activation function layer is connected to the input of the joint output layer;
[0051] The output end of the joint output layer is used to output the policy text annotation results.
[0052] Exemplarily, the first pre-trained language model and the second pre-trained language model are both implemented by a Bert model.
[0053] It should be understood that in terms of pre-trained language models, the first pre-trained language model is mainly used to extract word vectors and sentence vectors from policy texts, and input the word vectors into the graph neural network layer as the initial feature vectors of word nodes in the policy text graph structure, which is an indispensable prerequisite for improving the classification and labeling accuracy of the graph neural network layer. The second pre-trained language model is used for fine-tuning to obtain the final contextual semantic feature representation of the policy text. Using the first pre-trained language model to extract vectors and the second pre-trained language model for fine-tuning can not only improve the semantic understanding of the policy text, but also can perfectly parallelize operations with the graph neural network, ultimately achieving the most advanced performance in policy text classification and labeling.
[0054] Furthermore, the text-level graph construction layer uses a sliding window to slide the word segmentation results of the policy text. The length of the sliding window is N words, and the sliding step of the sliding window is M words. N and M are both positive integers. Each word is regarded as a node, and the weight between any two nodes in the window is calculated according to the content of the sliding window. When the weight is positive, a connecting edge is set between the two nodes. When the weight is zero or negative, no connecting edge is set. After the sliding is completed, the constructed text-level graph structure is obtained, and the corresponding adjacency matrix is obtained according to the text-level graph structure.
[0055] It should be understood that the use of a text-level graph construction layer can fully understand the structural information of each policy text, while also avoiding the negative impact of corpus-level graph construction, retaining the individual features of each policy text, and is more conducive to inductive training of graph neural networks.
[0056] Furthermore, if Figure 4 As shown, the graph neural network layer includes a first graph neural network GCN sublayer and a second graph neural network GCN sublayer connected in series, wherein the input end of the first graph neural network GCN sublayer is respectively connected to the output end of the text-level graph construction layer and the output end of the first pre-trained language model; wherein the output end of the second graph neural network GCN sublayer is connected to the input end of the maximum pooling layer.
[0057] The word vectors encoded by the first pre-trained language model layer are mapped to word nodes on the text-level policy text graph and input into the graph neural network layer together with the policy text adjacency matrix for inductive training.
[0058] It should be understood that graph neural networks perform inductive training, that is, only train the document information of the policy text in the training set, and have strong generalization capabilities for new policy texts. At the same time, they can also perform small-batch learning, which greatly reduces the complexity of time and space. Non-corpus-level graph construction for direct training requires re-learning and can only learn the entire corpus-level graph.
[0059] At the same time, the policy text content is input into the second pre-trained language model layer (i.e., Bert layer) for fine-tuning training in the same small batch as the graph neural network;
[0060] It should be understood that BERT applies the self-attention mechanism to construct a multi-layer self-attention network, which can achieve parallel operation and can be easily transferred to the policy text annotation task through fine-tuning, and can more effectively extract the semantic features of each policy text;
[0061] Furthermore, the joint output layer is used to perform weighted summation on the semantic features and the structural features to obtain fused features and output the policy text annotation results.
[0062] Furthermore, the training process of the trained policy text annotation model includes:
[0063] Constructing a training set; the training set is a policy text with known annotation results;
[0064] Preprocess the training set;
[0065] The preprocessed training set is input into the policy text annotation model, and the model is trained. When the loss function value no longer decreases, the training is stopped to obtain the trained policy text annotation model.
[0066] Among them, the loss function is the cross entropy function.
[0067] Furthermore, the step S101: obtaining the policy text to be annotated and preprocessing the policy text to be annotated specifically includes:
[0068] Use regular expressions to remove HTML tags and non-text content from the policy text to be annotated;
[0069] Perform word segmentation on the policy text to be annotated;
[0070] Remove stop words from the segmented vocabulary.
[0071] For example, the step S101 of obtaining a policy text to be annotated and preprocessing the policy text to be annotated specifically includes:
[0072] Clean the original policy text. Since the file contains HTML tags and non-text content, use Python regular expressions (re) to remove them;
[0073] For example, take a policy document published on a website as an example: <p style="text-indent:2em;font-family:宋体,simsun;font-size:16px;line-height:2em;text-indent:2em;"> In accordance with the requirements of the "Measures for the Disclosure of Information on Construction Projects of the Ministry of Agriculture (Trial)", the approval of the feasibility study reports (implementation plans) for 15 animal and plant protection capacity enhancement projects, including the construction project of the Key Laboratory of Rice Biology and Genetic Breeding of Southwest China of the Ministry of Agriculture and Rural Affairs, the construction of the medium-term bank of germplasm resources of Southwest China, and the construction of special facilities for animal epidemic prevention in pastoral areas of Ganzi Prefecture - vaccination pens, are now made public. ", use Python regular expression (re) to remove the HTML tags and non-text content of the example, and the result is "In accordance with the Ministry of Agriculture's construction project information disclosure measures, the approval of the feasibility study reports of the Ministry of Agriculture and Rural Affairs' Southwest Rice Biology and Genetic Breeding Key Laboratory Construction Project, Southwest Characteristic Crop Germplasm Resources Mid-term Bank Construction Project, and other modern seed industry improvement projects, such as the Ganzi Prefecture Pastoral Area Animal Epidemic Prevention Special Facilities and Epidemic Prevention Injection Fence Construction, and other animal and plant protection capacity improvement projects are now disclosed."
[0074] Use the jieba.cut() function in the jieba word segmentation library to segment the cleaned policy text in precise mode;
[0075] For example, after using jieba to segment the cleaned text, it becomes: 'In accordance with the requirements of the Ministry of Agriculture's construction project information disclosure measures, the approval of the feasibility study reports of several modern seed industry improvement projects, such as the construction project of the Key Laboratory of Rice Biology and Genetic Breeding in Southwest China, the construction of the mid-term bank of germplasm resources for characteristic crops in Southwest China, and the construction of special facilities for animal epidemic prevention and vaccination pens in pastoral areas of Ganzi Prefecture, and other animal and plant protection capacity improvement projects, are now made public.'
[0076] We use the Chinese policy text stop word list to delete stop words from the text. We also count the words that appear less frequently in the text and remove them. This avoids unnecessary information redundancy and reduces the computational overhead when building the text-level graph.
[0077] For example, after removing low-frequency words and stop words from the segmented text, it becomes: "The Ministry of Agriculture's Measures for the Disclosure of Construction Project Information now disclose the feasibility study report approved by the Ministry of Agriculture and Rural Affairs for the construction of the Southwest Key Laboratory of Rice Biology and Genetics Breeding, the construction of the Southwest Characteristic Crop Germplasm Resources Mid-term Bank, the Seed Industry Improvement Project, the Ganzi Prefecture Pastoral Area Animal Epidemic Prevention Special Facilities and Injection Pens, and the Animal and Plant Protection Capacity Improvement Project."
[0078] Furthermore, the word vectors and sentence vectors are extracted from the processed policy text. Specifically, the first pre-trained language model is used to extract word vectors for the processed policy text, and the output values of the 1st, 5th, 9th and 12th hidden layers of the pre-trained language model are used as word vectors; or the output values of the 1st, 5th, 9th and 12th hidden layers are combined as sentence vectors, so that the policy text annotation effect obtained is more accurate.
[0079] Exemplarily, the hidden layer features of the 1st, 5th, 9th, and 12th layers of the first pre-trained language model are extracted as high-dimensional semantic features of the policy text;
[0080] Among them, the representation of the policy text vector is:
[0081]
[0082] Where K is the number of policy documents, n is the number of words in a single document, and d is the dimension of the word vector.
[0083] Furthermore, the method of constructing a text-level graph structure based on the preprocessed policy text and obtaining an adjacency matrix corresponding to the text-level graph structure means: sliding the word segmentation results of the policy text using a sliding window, the length of the sliding window is N words, and the sliding step of the sliding window is M words; N and M are both positive integers; treating each word as a node, and calculating the weight between any two nodes in the window according to the content of the sliding window. When the weight is positive, a connecting edge is set between the two nodes, and when the weight is zero or negative, no connecting edge is set. Finally, the constructed text-level graph structure is obtained, and the corresponding adjacency matrix is obtained according to the text-level graph structure.
[0084] For example, the text-level graph structure is to build a separate graph structure for each text, and each node in the graph represents a word in the text;
[0085] Use PMI (Pointwise Mutual Information) to calculate the weight between two word nodes and add edges to the graph;
[0086] The calculation formula of PMI is expressed as:
[0087]
[0088]
[0089]
[0090] Because the policy texts contain a large number of words after data preprocessing, a vocabulary sliding window size of 20 is selected, where #W(i) is the number of sliding windows in a single policy text containing word i, #W(i,j) is the number of sliding windows containing both words i and j, and #W is the total number of sliding windows in a single policy text. A positive PMI value indicates a high degree of semantic relevance between words in the policy text, while a negative PMI value indicates almost no semantic relevance in the policy text. Therefore, edges are added only between word pairs with positive PMI values, and an adjacency matrix is constructed from this:
[0091]
[0092] Among them, K is the number of texts, n is the number of word nodes; A K is the adjacency matrix of the entire policy text graph structure.
[0093] Furthermore, the semantic features of the policy text are extracted based on the word vectors and sentence vectors, specifically by using a second pre-trained language model to extract the semantic features of the policy text from the word vectors and sentence vectors.
[0094] Furthermore, the second pre-trained language model is used to extract semantic features of the policy text from word vectors and sentence vectors, specifically including:
[0095] The preprocessed policy text is input into the second pre-trained language model layer (i.e., Bert layer) to obtain the semantic information of a single text and more semantic context features.
[0096] The sentence vector output by the second pre-trained language model layer is used as the feature vector of the policy text and is input into the fully connected layer. The output of the fully connected layer is the number of categories for each policy text. The sigmoid function is used to squeeze each dimension value of the policy text vector to within (0,1). The entire Bert layer formula is as follows:
[0097] Z Bert =sigmoid(f Bert (W2S)
[0098] Among them, W2 is the parameter of Bert, S is the content of each policy text, and f Bert Represents the second pre-trained language model.
[0099] Furthermore, the structural features of the policy text are extracted based on the word vectors and the adjacency matrix. Specifically, the word vectors and the adjacency matrix are processed using a graph neural network layer to extract the structural features of the policy text.
[0100] Furthermore, a graph neural network layer is used to process word vectors and adjacency matrices to extract the structural features of the policy text. Specifically,
[0101] The word vector and the adjacency matrix The data is fed into two GCN sub-layers in batches, which then acquire the one-hop and two-hop neighbor information of each word node in the adjacency matrix and learn fine-grained vocabulary representations of local structures in a single policy text graph.
[0102] Among them, after the first GCN sublayer, the rectified linear unit ReLU activation function is used to correct the linearity of the hidden unit features output by the first GCN sublayer to avoid gradient disappearance, and at the same time serve as the input of the second GCN sublayer;
[0103] Among them, the output of the second graph neural network GCN sublayer is the number of categories of policy text, expressed as:
[0104]
[0105] Among them, K represents the number of policy texts, n represents the number of word nodes, and m represents the number of categories of policy texts. The final output vector of the second graph neural network.
[0106] The maximum pooling layer performs a maximum pooling operation on the word vector of each policy text to obtain the maximum value of the feature in each category dimension, which is used as the text feature vector of each policy text. The feature vector of the word node is compressed into a text feature vector, which is expressed as
[0107] The sigmoid function is used to squeeze each dimension value of the policy text vector into (0, 1). The formula of the entire graph neural network layer is as follows:
[0108] Z GCN =sigmoid(Maxpooling(A*ReLU(AH (0) W0)W1))
[0109] Among them, W0 is the first graph neural network GCN sublayer parameter, W1 is the second graph neural network GCN sublayer parameter, H (0) is the initial word node embedding vector of the text-level graph structure, A is the adjacency matrix of the text-level graph structure, and ReLU is the activation function.
[0110] Furthermore, if Figure 3 As shown, the policy text annotation result is determined based on semantic features and structural features, which is to input both semantic features and structural features into the joint output layer, and the joint output layer fuses the two features to obtain the policy text annotation result after fusion.
[0111] It should be understood that if Figure 3 As shown, the policy text annotation result is determined based on semantic features and structural features, which is to convert the output result Z in the graph neural network layer into GCN and the output Z of the second pre-trained language model Bert The input is fed into the joint layer to dynamically combine the text structural representation learned by the graph neural network layer and the contextual semantic representation learned by the second pre-trained language model.
[0112] Use two learnable parameters to learn the ratio of the output value of the second pre-trained language model and the output value of the graph neural network during the model training phase. For example, the learned value is 0.7*the output value of the second pre-trained language model + 0.3*the output value of the graph neural network model.
[0113] It should be understood that is the final representation of the semantic information of the i-th policy text, is the final representation of the structural information of the i-th policy text.
[0114] A learnable weighting approach is used to balance the predictions between the graph neural network layer and the second pre-trained language model. Unlike fixed interpolation, the weights used in the joint output layer are multi-dimensional 1×1 convolutional kernels pre-processed with softmax. The weights of the aggregation layer are learnable during training, effectively transforming fixed interpolation into dynamic interpolation. Before each aggregation, the weights of the aggregation layer undergo softmax pre-processing to control the weighting of the prediction results within a ratio of 1. This can control the weighting of the two predictions, allowing the model to achieve a matching combination of the two-hop neighborhood structure information output by the second graph neural network sublayer and the textual semantic information learned by the second pre-trained language model. By combining different representations of policy text, the model achieves excellent results, improving its retrieval and classification capabilities. Furthermore, during training, the best representation of the two predictions can be learned, allowing for better model optimization.
[0115] The joint output layer has the following formula:
[0116] W Conv1 , W Conv2 =softmax(0.1);
[0117] Z=W Conv1 Z GCN +W Conv2 Z Bert ;
[0118] Among them, W Conv1 , W Conv2 are all learnable parameters, Z is the final output of the joint pre-training and graph neural network model, denoted as Z∈R k×m , k is the number of policy texts in the input batch, m is the number of categories of policy texts; softmax represents the activation function, Z GCN Represents structural features, Z Bert Represents semantic features.
[0119] It should be understood that the multi-dimensional convolution kernel parameters Conv are divided into the first W Conv1 、Second W Conv2 The two aggregation weights are initialized to 0.1 respectively. After the softmax operation, the value becomes 0.5. That is, at the beginning of training, the first and second aggregation weights have the same ratio, both W Conv1 Z GcN +W Conv2 Z Bert After the i-th round of training, the first and second aggregation weights are dynamically changed according to whether a policy text is biased towards structural information or semantic information, and thus adjusted accordingly. If the policy text is biased towards structural information, the first aggregation weight W Conv1 The proportion increases, the second aggregation weight WConv2 The ratio is reduced, and vice versa. Since the aggregation weights are pre-processed by softmax each time, the dynamic adjustment ratio range is limited to W Conv1 +W Conv2 =1, which ensures both the dynamic adjustment range of the aggregation weight and its balance invariance.
[0120] The policy text is classified according to a pre-set threshold. If the value in the vector m is higher than the threshold, it is set to 1; if it is lower than the threshold, it is set to 0.
[0121] The parameters of the joint output layer are divided into the first and second aggregation weights. After training, the first and second aggregation weights are dynamically adjusted according to whether a policy text is biased towards structural information or semantic information.
[0122] If the policy text is biased towards structural information, the first aggregation weight ratio increases and the second aggregation weight ratio decreases, and vice versa.
[0123] Since the aggregation weights are preprocessed by softmax each time, the dynamic adjustment range is limited to the sum of the first and second aggregation weights being 1. This ensures both the dynamic adjustment range of the aggregation weights and their balance invariance.
[0124] The present invention discloses a policy text annotation method and system that combines pre-training and graph neural networks, including: pre-processing the information contained in the policy text; converting the obtained data set into corresponding word vectors after pre-processing; using word co-occurrence information to construct and initialize a text-level graph structure; extracting structural features through a graph neural network, and extracting semantic features through a pre-trained language model, and inputting the above two features into a feature joint output layer to classify and annotate the graph structure information accordingly; compared with the existing technology, the present invention uses a learnable method to jointly train a pre-trained language model and a graph convolutional network, while learning the structural information and semantic information of the text, and constructing an isomorphic graph for each text, and can also achieve the induction of new words and new texts. Compared with manual standards, it is more accurate and efficient, and avoids the waste of excessive computing resources.
[0125] Example 2
[0126] This embodiment provides a policy text annotation system that combines pre-training and graph neural networks;
[0127] A policy text annotation system that combines pre-training and graph neural networks, including:
[0128] An acquisition module is configured to: acquire the policy text to be annotated and pre-process the policy text to be annotated;
[0129] An annotation module is configured to: input the preprocessed policy text into the trained policy text annotation model and output the annotation results of the policy text;
[0130] Among them, the working principle of the trained policy text annotation model includes: extracting word vectors and sentence vectors for the processed policy text; constructing a text-level graph structure based on the preprocessed policy text, and obtaining the adjacency matrix corresponding to the text-level graph structure; extracting the semantic features of the policy text based on word vectors and sentence vectors; extracting the structural features of the policy text based on word vectors and adjacency matrix; and determining the policy text annotation results based on semantic features and structural features.
[0131] It should be noted that the acquisition module and the annotation module described above correspond to steps S101 to S102 in Example 1. The examples and application scenarios implemented by the modules and the corresponding steps are the same, but are not limited to the contents disclosed in Example 1. It should be noted that the modules described above, as part of a system, can be executed in a computer system, such as a set of computer-executable instructions.
[0132] The descriptions of the various embodiments in the above embodiments have different focuses. For parts not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0133] The proposed system can be implemented in other ways. For example, the system embodiment described above is merely illustrative. For example, the above module division is only a logical function division. In actual implementation, other division methods may be used. For example, multiple modules can be combined or integrated into another system, or some features can be ignored or not implemented.
[0134] Example 3
[0135] This embodiment also provides an electronic device, comprising: one or more processors, one or more memories, and one or more computer programs; wherein the processor is connected to the memory, and the one or more computer programs are stored in the memory. When the electronic device is running, the processor executes the one or more computer programs stored in the memory, so that the electronic device executes the method described in the above embodiment one.
[0136] It should be understood that in this embodiment, the processor may be a central processing unit (CPU), or may be other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), off-the-shelf field-programmable gate arrays (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor, etc.
[0137] The memory may include a read-only memory and a random access memory, and provides instructions and data to the processor. A portion of the memory may also include a non-volatile random access memory. For example, the memory may also store information about the device type.
[0138] During implementation, each step of the above method may be completed by an integrated logic circuit of hardware in a processor or by instructions in the form of software.
[0139] The method in Example 1 can be directly implemented as being executed by a hardware processor, or by a combination of hardware and software modules within the processor. The software module can be located in a storage medium well-established in the art, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, or registers. The storage medium is located in the memory, and the processor reads the information in the memory and, in conjunction with its hardware, completes the steps of the above method. To avoid repetition, a detailed description is not given here.
[0140] Those skilled in the art will appreciate that the units and algorithm steps of the various examples described in conjunction with this embodiment can be implemented using electronic hardware or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the present invention.
[0141] Example 4
[0142] This embodiment further provides a computer-readable storage medium for storing computer instructions. When the computer instructions are executed by a processor, the method described in the first embodiment is performed.
[0143] The foregoing description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Those skilled in the art will readily appreciate that various modifications and variations of the present invention are possible. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present invention are intended to be within the scope of protection of the present invention.
Claims
1. A policy text annotation method based on joint pre-training and graph neural network, characterized by: include: Obtain the policy text to be annotated and preprocess the policy text to be annotated; Input the preprocessed policy text into the trained policy text annotation model and output the annotation results of the policy text; The working principle of the trained policy text annotation model includes the following steps: extracting word vectors and sentence vectors from the processed policy text; constructing a text-level graph structure based on the preprocessed policy text and obtaining the adjacency matrix corresponding to the text-level graph structure; extracting the semantic features of the policy text based on the word vectors and sentence vectors; extracting the structural features of the policy text based on the word vectors and adjacency matrix; and determining the policy text annotation results based on the semantic features and structural features. The method of extracting semantic features of the policy text based on word vectors and sentence vectors specifically uses a second pre-trained language model to extract semantic features of the policy text from word vectors and sentence vectors, specifically including: Input the pre-processed policy text into the second pre-trained language model layer, the Bert layer, to obtain the semantic information of the single text and more semantic context features; The sentence vector output by the second pre-trained language model layer is used as the feature vector of the policy text and is input into the fully connected layer. The output of the fully connected layer is the number of categories for each policy text. The sigmoid function is used to squeeze each dimension value of the policy text vector to within (0,1). The entire Bert layer formula is as follows: in, is the parameter of Bert, S is the content of each policy text, represents the second pre-trained language model; Based on word vectors and adjacency matrices, the structural features of the policy text are extracted, including: The word vector and the adjacency matrix The data is fed into two GCN sub-layers in batches, which then acquire the one-hop and two-hop neighbor information of each word node in the adjacency matrix and learn fine-grained vocabulary representations of local structures in a single policy text graph. Among them, after the first GCN sublayer, the rectified linear unit ReLU activation function is used to correct the linearity of the hidden unit features output by the first GCN sublayer to avoid gradient disappearance, and at the same time serve as the input of the second GCN sublayer; Among them, the output of the second graph neural network GCN sublayer is the number of categories of policy text, expressed as: , ; Among them, K represents the number of policy texts, n represents the number of word nodes, and m represents the number of categories of policy texts. The final output vector of the second graph neural network; The joint output layer has the following formula: ; ; in, are all learnable parameters, and Z is the final output result of the joint pre-training and graph neural network model, expressed as , k is the number of policy texts in the input batch, and m is the number of categories of policy texts; represents the activation function, Represents structural features, Represents semantic features; The method of determining the policy text annotation result based on semantic features and structural features is to input both the semantic features and the structural features into a joint output layer, and the joint output layer fuses the two features to obtain the policy text annotation result after fusion.
2. The policy text annotation method of the joint pre-training and graph neural network according to claim 1 is characterized in that: The trained policy text annotation model has a model structure including: First pre-trained language model and text-level graph construction layer; The input end of the first pre-trained language model and the input end of the text-level graph construction layer are both used to input the pre-processed policy text; the output of the first pre-trained language model is the word vector of the 1st, 5th, 9th, and 12th hidden layers; the output end of the text-level graph construction is the adjacency matrix represented by the graph structure of a single policy text; The output of the first pre-trained language model is connected to the input of the second pre-trained language model, the output of the second pre-trained language model is connected to the input of the fully connected layer, the output of the fully connected layer is connected to the input of the first sigmoid activation function layer, and the output of the first sigmoid activation function layer is connected to the input of the joint output layer; The output of the text-level graph construction layer and the output of the first pre-trained language model are both connected to the input of the graph neural network layer; the output of the graph neural network layer is connected to the input of the maximum pooling layer, the output of the maximum pooling layer is connected to the input of the sigmoid activation function layer, and the output of the second sigmoid activation function layer is connected to the input of the joint output layer; The output end of the joint output layer is used to output the policy text annotation results.
3. The policy text annotation method of joint pre-training and graph neural network according to claim 2 is characterized in that: The text-level graph construction layer uses a sliding window to slide the word segmentation results of the policy text, the length of the sliding window is N words, and the sliding step of the sliding window is M words; N and M are both positive integers; Consider each word as a node, and calculate the weight between any two nodes in the sliding window based on the content of the window. When the weight is positive, a connecting edge is set between the two nodes. When the weight is zero or negative, no connecting edge is set. After the sliding is completed, the constructed text-level graph structure is obtained, and the corresponding adjacency matrix is obtained based on the text-level graph structure.
4. The policy text annotation method of the joint pre-training and graph neural network according to claim 2 is characterized in that: The graph neural network layer includes a first graph neural network GCN sublayer and a second graph neural network GCN sublayer connected in series, wherein the input end of the first graph neural network GCN sublayer is respectively connected to the output end of the text-level graph construction layer and the output end of the first pre-trained language model; wherein the output end of the second graph neural network GCN sublayer is connected to the input end of the maximum pooling layer.
5. The policy text annotation method of joint pre-training and graph neural network according to claim 1 is characterized in that: Obtain the policy text to be annotated and preprocess the policy text to be annotated, specifically including: using regular expressions to remove HTML tags and non-text content in the policy text to be annotated; performing word segmentation on the policy text to be annotated; and removing stop words from the words after word segmentation.
6. A policy text annotation system that combines pre-training and graph neural networks, featuring: An acquisition module is configured to: acquire the policy text to be annotated and pre-process the policy text to be annotated; An annotation module is configured to: input the preprocessed policy text into the trained policy text annotation model and output the annotation results of the policy text; The working principle of the trained policy text annotation model includes the following steps: extracting word vectors and sentence vectors from the processed policy text; constructing a text-level graph structure based on the preprocessed policy text and obtaining the adjacency matrix corresponding to the text-level graph structure; extracting the semantic features of the policy text based on the word vectors and sentence vectors; extracting the structural features of the policy text based on the word vectors and adjacency matrix; and determining the policy text annotation results based on the semantic features and structural features. The method of extracting semantic features of the policy text based on word vectors and sentence vectors specifically uses a second pre-trained language model to extract semantic features of the policy text from word vectors and sentence vectors, specifically including: Input the pre-processed policy text into the second pre-trained language model layer, the Bert layer, to obtain the semantic information of the single text and more semantic context features; The sentence vector output by the second pre-trained language model layer is used as the feature vector of the policy text and is input into the fully connected layer. The output of the fully connected layer is the number of categories for each policy text. The sigmoid function is used to squeeze each dimension value of the policy text vector to within (0,1). The entire Bert layer formula is as follows: in, is the parameter of Bert, S is the content of each policy text, represents the second pre-trained language model; Based on word vectors and adjacency matrices, the structural features of the policy text are extracted, including: The word vector and the adjacency matrix The data is fed into two GCN sub-layers in batches, which then acquire the one-hop and two-hop neighbor information of each word node in the adjacency matrix and learn fine-grained vocabulary representations of local structures in a single policy text graph. Among them, after the first GCN sublayer, the rectified linear unit ReLU activation function is used to correct the linearity of the hidden unit features output by the first GCN sublayer to avoid gradient disappearance, and at the same time serve as the input of the second GCN sublayer; Among them, the output of the second graph neural network GCN sublayer is the number of categories of policy text, expressed as: , ; Among them, K represents the number of policy texts, n represents the number of word nodes, and m represents the number of categories of policy texts. The final output vector of the second graph neural network; The joint output layer has the following formula: ; ; in, are all learnable parameters, and Z is the final output result of the joint pre-training and graph neural network model, expressed as , k is the number of policy texts in the input batch, and m is the number of categories of policy texts; represents the activation function, Represents structural features, Represents semantic features; The method of determining the policy text annotation result based on semantic features and structural features is to input both the semantic features and the structural features into a joint output layer, and the joint output layer fuses the two features to obtain the policy text annotation result after fusion.
7. An electronic device, comprising: a memory for non-transitory storage of computer-readable instructions; as well as a processor for executing said computer-readable instructions, When the computer-readable instructions are executed by the processor, the method according to any one of claims 1 to 5 is executed.
8. A storage medium, characterized in that: Computer-readable instructions are non-transitory stored, wherein when the non-transitory computer-readable instructions are executed by a computer, the instructions of the method according to any one of claims 1 to 5 are executed.
Citation Information
Patent Citations
Systems and methods for polygon object annotation and a method of training an object annotation system
CA3091035A1
English-Mian bilingual parallel sentence pair extraction method and device fusing pre-trained language model and structural features
CN112287688A