A standard content keyword recognition method based on fused column-atrous convolution

By adopting a BERT-based fusion column hole convolution method in standard content keyword recognition, combined with BiLSTM-Fusion and CRF models, the problem of local and global feature extraction balance in the prior art is solved, and the extraction ability of long-range dependency information and the accuracy of keyword recognition are improved.

CN114757175BActive Publication Date: 2025-05-06BEIJING UNIV OF POSTS & TELECOMM
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210492445.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-04-29
Publication Date
2025-05-06
Estimated Expiration
2042-04-29

AI Technical Summary

Technical Problem

The prior art is difficult to balance the extraction of local and global features in standard content keyword recognition, resulting in insufficient extraction capability of long-range dependent features, and excessive pooling layers lead to loss of data space information.

Method used

The fusion column cavity convolution method based on BERT pre-trained model is adopted, combined with BiLSTM-Fusion and CRF models, local features are extracted through column cavity convolution, and spliced ​​with the feature information of BiLSTM-Fusion, and sent to the CRF layer for optimal labeling sequence generation.

Benefits of technology

It improves the model's ability to extract long-range dependency information, retains the spatial information of text, and enhances the accuracy of keyword recognition and the integrity of information extraction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114757175B_ABST
    Figure CN114757175B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for identifying keywords in the content of a standard based on fused column-drilled convolution, the steps of which are: (1) determining a serialized vector of the standard content text; (2) extracting the local features of each word; (3) determining the word context weight information; (4) obtaining the final conditional distribution of the annotation sequence; and (5) optimizing the parameters to obtain the optimal annotation sequence. The present invention combines column-drilled convolution with BiLSTM-Fusion, and uses column-drilled convolution to extract local feature information, effectively improving the model's ability to extract long-range dependent information, while retaining the spatial information of the text, providing a high-accuracy extraction method for keyword extraction of standard content.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of keyword recognition and standard digitization, and specifically, mainly relates to a method for recognizing keywords in standard content based on fused column dilated convolution. Background Art

[0002] Keyword recognition is a necessary step in extracting key information in the digitalization of standards. Accurate recognition of keywords in standard content is helpful for the construction of standard knowledge base and standard classification. In recent years, keyword recognition models have gradually shifted from traditional machine learning-based models to deep learning-based models. The emergence of pre-trained models such as BERT and GPT-3 has greatly improved the accuracy of entity naming recognition, but their generality in the field is poor. At present, deep learning-based models have achieved the ability to extract high-order features of text, especially for the extraction of long-distance dependent features. In order to obtain high-order features of text, traditional convolutional networks usually stack multiple convolutional layers and pooling layers to expand the network's receptive field, but the addition of too many pooling layers will lead to the loss of data spatial information. There is an irreconcilable contradiction between the field of view of text information and the accurate extraction of spatial information.

[0003] Standard content text has the characteristics of a mixture of long- and short-distance dependency features. Therefore, a standard content keyword recognition method is needed that can accurately extract standard keywords and has good sensitivity to local keywords and global keywords. This provides strong support for the construction of standard knowledge bases and subsequent digitalization. Summary of the invention

[0004] In order to solve the problems existing in the above-mentioned prior art, the present invention provides a method for identifying keywords in the content of the Standard based on fused column dilated convolution. The specific process is as follows: Figure 1 shown.

[0005] The steps for implementing the technical solution are as follows:

[0006] (1) Determine the serialization vector E of the standard content text:

[0007] Get the representation of a sentence in the text

[0008] X=[x1,x2,…,x n ] T

[0009] Where X is the vector representation of the sentence, x i Represents the i-th word in the sentence text. By inputting the text X into the layer for BERT serialization operation, the serialized text vector is obtained

[0010] F=[F1,F2,…,F n ] T

[0011] Where F represents the character array after the sentence text is serialized, F i The serialized word representing the i-th word in the text.

[0012] (2) Extract the local features m of each word:

[0013] A hole is formed in the word direction to form a column-hole convolution. The column-hole convolution kernel with a convolution kernel size of 3 and a hole rate of 2 is as follows

[0014]

[0015] Convolution kernel k∈R 3×l , where the first and third rows have convolution parameters, and the second row does not participate in the convolution operation and does not update the parameters.

[0016]

[0017] F∈R n×l represents the text input matrix, which is a two-dimensional matrix of n words embedded in the text, l is the word embedding dimension, and the convolution kernel k is selected with a size of W×l. W is the convolution kernel width, l is the embedding length of the word vector, and the convolution operation is performed in the direction of word splicing. F o represents the sum of the products of the convolution kernel parameters and the elements of the text embedding matrix, where the hollow part of the convolution kernel does not participate in the calculation, W (S,r) It represents the width of the convolution kernel with size S and dilation rate r. The output f[x][y] after processing by the column-dilated convolution kernel represents the y-th feature of the x-th row element.

[0018] Introduce parallel convolution layer and stacked pooling layer. The parallel convolution layer and the column-hole convolution are run in parallel, which is used to extract the features of the original data; the stacked pooling layer stacks the features of the parallel convolution layer and the column-hole convolution vertically, fuses the features of the two aspects, and finally obtains the feature information through maximum pooling, that is,

[0019] M′=max(M[0][y],M[1][y],…,M[G+1-W (S,r) ][y]),1≤y≤l.

[0020] (3) Determine word context weight information

[0021] The dimensional feature information after column-hole convolution is fused with word embedding to obtain

[0022] E i =[F i T ,M′] T

[0023] Ε=[E1,E2,…,E n ]

[0024] E i represents the word vector representation after the i-th column-dilated convolution and BERT word embedding, and Ε represents the word vector representation of the sentence.

[0025] BiLSTM-fusion fuses the BiLSTM input and BiLSTM output layers again in the final output layer. The formula is as follows.

[0026]

[0027]

[0028] Above are the three gates of the forward LSTM, namely the input gate, the forget gate, and the output gate. are the three gates of the backward LSTM. These six gates can control the flow of information. In the forward LSTM, the hidden layer state r t-1 R t The update has an impact on the backward LSTM, the hidden layer state l t+1 Yes t W is the weight matrix; b is the bias term; σ is the sigmoid activation function; c is the state variable, which together with the output gate controls the final hidden layer state; * is the Hadamard product; tanh is the hyperbolic tangent function; is the concatenation operation of the vector. The character array with context information after BiLSTM-fusion processing is

[0029]

[0030] (4) Obtain the final label sequence conditional distribution P(y|x):

[0031] Obtained by BiLSTM-Fusion

[0032]

[0033] In the CRF model, the conditional distribution probability P(y|x) of the labeled sequence is

[0034]

[0035] and Represents the label pair (y i-1 ,y i )’s state transfer matrix and bias parameters, which can be continuously optimized.

[0036] (5) Optimize the parameters to obtain the optimal labeling sequence Y = (y1, y2, ..., y n ):

[0037] Finally, the maximum likelihood function L is selected as the objective function

[0038]

[0039] X is all training samples, Y is the set of all labels, and the training uses the Adam optimizer.

[0040] After parameter adjustment, the optimal labeling sequence is generated

[0041] Y=(y1,y2,…,y n ).

[0042] The present invention has the following advantages over the prior art:

[0043] (1) The method of the present invention is based on the BERT pre-training set and the standard content keyword recognition task, using a combination of column-drilled convolution and BiLSTM-Fusion.

[0044] (2) The present invention uses column-dilated convolution to extract local feature information, effectively improving the model's ability to extract long-range dependent information while retaining the spatial information of the text.

[0045] (3) The present invention concatenates the information extracted by the column-drilled convolution with the feature information extracted by the BiLSTM-Fusion and sends them to the CRF layer. This method provides more feature information and further ensures the accuracy of information extraction. BRIEF DESCRIPTION OF THE DRAWINGS

[0046] In order to better understand the present invention, further description is given below in conjunction with the accompanying drawings.

[0047] Figure 1 It is a flowchart of the steps of the method for identifying keywords in the content of the Standard based on fused column-atrous convolution.

[0048] Figure 2 This is an algorithm flow chart of the keyword recognition method for the Standard content based on fused column-cavitated convolution.

[0049] Figure 3 This is a schematic diagram of the network model of the "Standard" content keyword recognition method based on fused column void convolution.

[0050] Figure 4 This is a comparison chart of the effects of the "Standard" content keyword recognition method based on fused column void convolution.

[0051] Figure 5This is a comparison chart of the effects of the "Standard" content keyword recognition method based on fused column void convolution. DETAILED DESCRIPTION

[0052] The present invention is further described in detail below through implementation cases.

[0053] In this implementation case, two standard data sets, gas accident standards and hazardous chemicals accident standards, were selected for testing. Manual labeling was used to verify the model performance. The two standard sets each contained 300 standards in total.

[0054] The algorithm flow of the method for identifying keywords in the content of the Standard based on fused column dilated convolution provided by the present invention is as follows: Figure 2 As shown, the specific steps are as follows:

[0055] (1) Determine the serialization vector E of the standard content text:

[0056] Taking the gas accident handling standard dataset as an example, the average number of words in the gas accident handling standard dataset is 20, and the text representation of the corresponding sentence is

[0057] X=[x1,x2,…,x n ] T

[0058] Where X is the vector representation of the sentence, x i Represents the i-th word in the sentence text. By inputting the text X into the layer for BERT serialization operation, the serialized text vector is obtained

[0059] F=[F1,F2,…,F 20 ] T

[0060] Where F represents the character array after the sentence text is serialized, F i Represents the serialized word of the i-th word in the text, and the word embedding dimension is 768. The overall algorithm flow is as follows Figure 2 shown.

[0061] (2) Extract the local features m of each word:

[0062] A hole is formed in the word direction to form a column-wise atrous convolution. Two atrous convolution kernels are selected, one of which satisfies the convolution kernel size of 3 and the atrous rate of 2, and the other satisfies the convolution kernel size of 4 and the atrous rate of 3.

[0063] F∈R 20×768represents the text input matrix, which is a two-dimensional matrix of 20 words embedded in the text. The word embedding dimension is 768. In the first column convolution, two convolution kernels are selected, 3 of each convolution kernel are selected, and the convolution kernel parameter initialization distribution satisfies N(0,1). Each convolution kernel of the same size is convolved in the word splicing direction. Each output element corresponding to the convolution kernel size of 3 and 4 is

[0064]

[0065] F o represents the sum of the products of the convolution kernel parameters and the elements of the text embedding matrix, where the hollow part of the convolution kernel does not participate in the calculation, W (S,r) The sum width of the convolution kernel with size S and dilation rate r is represented. The output f[x][y] after column dilation convolution kernel processing represents the yth feature of the xth row element. After the first convolution, the representation information dimension of the text satisfies F′ 1,i ∈R 18×768 , 1≤i≤3, F′ 2,i ∈R 17×768 , 1≤i≤3, the feature information matrix obtained after the first convolution is sent to the convolution kernel to satisfy the convolution kernel size of 3, the void rate of 2 and the convolution kernel k′∈R 2×768 The convolution kernel output is obtained from the ordinary convolution kernel, where the convolution kernel parameter initialization distribution satisfies N(0,1)

[0066]

[0067] After the second convolution, the representation information of the text satisfies F′ 1,i ∈R 16×768 , 1≤i≤3, F′ 2,i ∈R 16×768 , 1≤i≤3. Then the representation information obtained by different convolution kernels is mapped to R 16×768 Space, get F l Feature information.

[0068] Simply using column-level atrous convolutions may make the model unstable, so parallel convolution layers and stacked pooling layers are introduced.

[0069] The convolutional layer uses the kernel k′∈R 5×768 The convolution kernel of the convolution kernel is obtained, and the convolution kernel output space satisfies R 16×768 , the convolution kernel parameters satisfy the N(0,1) distribution.

[0070] The parallel convolution layer and the column-atrous convolution run in parallel and are used to extract the features of the original data. The stacked pooling layer stacks the features of the parallel convolution layer and the column-atrous convolution vertically, fuses the two features, and finally obtains the feature information through maximum pooling, that is, M′=max(M[1][y],M[2][y],…,M

[16] [y]),1≤y≤768.

[0071] (3) Determine word context weight information

[0072] The dimensional feature information after column-hole convolution is fused with word embedding to obtain

[0073] E i =[F i ,M′],1≤i≤20

[0074] Ε=[E1,E2,…,E 20 ]

[0075] E i represents the word vector representation after the i-th column-dilated convolution and BERT word embedding, and Ε represents the word vector representation of the sentence.

[0076] BiLSTM-fusion fuses the BiLSTM input and BiLSTM output layer again in the final output layer.

[0077] The formula is as follows:

[0078]

[0079]

[0080] Above are the three gates of the forward LSTM, namely the input gate, the forget gate, and the output gate. are the three gates of the backward LSTM. These six gates can control the flow of information. In the forward LSTM, the hidden layer state r t-1 R t The update has an impact on the backward LSTM, the hidden layer state l t+1 Yes t The update has an impact. W is the weight matrix, and the parameter initialization satisfies the distribution b is the bias term, and the parameter is initialized to 0; σ is the sigmoid activation function; c is the state variable, which together with the output gate controls the final hidden layer state; * is the Hadamard product; tanh is the hyperbolic tangent function; is the concatenation operation of the vector. The character array with context information after BiLSTM-fusion processing is, The network model diagram is as follows Figure 3 shown.

[0081] (4) Obtain the final label sequence conditional distribution P(y|x):

[0082] Obtained by BiLSTM-Fusion

[0083]

[0084] The conditional distribution probability P(y|x) of the labeled sequence is

[0085]

[0086] and Represents the label pair (y i-1 ,y i )’s state transfer matrix and bias parameters, where The initialization parameters satisfy the N(0,1) distribution. The parameters are initialized to 0 and can be continuously optimized.

[0087] (5) Optimize the parameters to obtain the optimal labeling sequence Y = (y1, y2, ..., y n ):

[0088] Finally, the maximum likelihood function L is selected as the objective function

[0089]

[0090] X is all training samples, Y is the set of all labels, and the training uses the Adam optimizer.

[0091] After the CRF model, the optimal labeling sequence is generated

[0092] Y=(y1,y2,…,y 20 ).

[0093] In order to verify the effectiveness of the present invention, a keyword extraction experiment was conducted on the present invention. The results are as follows: Figure 4 , 5 As shown. Figure 4 , 5 It can be seen that the performance of the model using this method in the dataset is improved compared with other methods.

Claims

1. A standard content keyword recognition method based on fused column dilated convolution, characterized in that: The following steps are involved: Step 1: Determine the serialization vector E of the standard content text: Get a representation of a sentence in text: X=[x1,x2,…,x n ] T ; Where X is the vector representation of the sentence, x i Represents the i-th word in the sentence text. By inputting the text X into the layer for BERT serialization operation, the serialized text vector is obtained; F=[F1,F2,…,F n ] T ; Where F represents the character array after the sentence text is serialized, F i The serialized word representing the i-th word in the text; Step 2: Extract the local features m of each word: A hole is formed in the word direction to form a column-hole convolution. The column-hole convolution kernel with a convolution kernel size of 3 and a hole rate of 2 is as follows: Convolution kernel k∈R 3×l , where the first and third rows have convolution parameters, and the second row does not participate in the convolution operation and does not update the parameters; F∈R n×l represents the text input matrix, which is a two-dimensional matrix of n words embedded in the text, l is the word embedding dimension, and the convolution kernel k is selected with a size of W×l. W is the convolution kernel width, l is the embedding length of the word vector, and the convolution operation is performed in the direction of word splicing; F o represents the sum of the products of the convolution kernel parameters and the elements of the text embedding matrix, where the hollow part of the convolution kernel does not participate in the calculation, W (S,r) represents the width of the convolution kernel with size S and dilation rate r. The output f[x][y] after processing by the column-dilated convolution kernel represents the yth feature of the xth row element. Parallel convolutional layers and stacked pooling layers are introduced. The parallel convolutional layers and the column-hole convolutions are run in parallel and are used to extract the features of the original data. The stacked pooling layer stacks the features of the parallel convolutional layers and the column-hole convolutions vertically, fuses the features of the two aspects, and finally obtains the feature information through maximum pooling, that is, M′=max(M[0][y],M[1][y],…,M[G+1-W (S,r) ][y]),1≤y≤l; Step 3: Determine word context weight information The dimensional feature information after column-drilled convolution is fused with word embedding to obtain: E i =[F i T ,M′] T ; Ε=[E1,E2,…,E n ]: E i represents the word vector representation of the i-th word after column-dilution convolution and BERT word embedding, and Ε represents the word vector representation of the sentence; BiLSTM-fusion fuses the BiLSTM input and BiLSTM output layer again in the final output layer; the formula is as follows: Above are the three gates of the forward LSTM, namely the input gate, the forget gate, and the output gate. are the three gates of the backward LSTM. These six gates can control the flow of information. In the forward LSTM, the hidden layer state r t-1 R t The update has an impact on the backward LSTM, the hidden layer state l t+1 Yes t The update of has an impact; W is the weight matrix; b is the bias term; σ is the sigmoid activation function; c is the state variable, which controls the final hidden layer state together with the output gate; * is the Hadamard product; tanh is the hyperbolic tangent function; For the concatenation operation of the vector, the character array with context information after BiLSTM-fusion processing is: Step 4: Get the final label sequence conditional distribution P(y|x): After BiLSTM-Fusion, we get: In the CRF model, the conditional distribution probability P(y|x) of the labeled sequence is: and Represents the label pair (y i-1 ,y i )’s state transfer matrix and bias parameters, which can be continuously optimized; Step 5: Optimize parameters to obtain the optimal labeling sequence Y = (y1, y2, ..., y n ): Finally, the maximum likelihood function L is selected as the objective function: X is all training samples, Y is the set of all labels, and the training uses the Adam optimizer; After parameter adjustment, the optimal labeling sequence is generated, Y = (y1, y2, ..., y n ).

Citation Information

Patent Citations

  • Text multi-label classification method based on semantic unit information

    CN109582789A

  • Chinese emotion analysis method based on bidirectional time convolution network

    CN110059188A