A standard content text summary generation method that integrates lexical encoding and structural encoding

By combining vocabulary coding and structural coding methods, combined with BERT, TextCNN, TreeLSTM and Att-LSTM models, the existing digest generation methods are solved, and more efficient and accurate digest generation is achieved.

CN114925195BActive Publication Date: 2025-05-27BEIJING UNIV OF POSTS & TELECOMM
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210475184.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-04-29
Publication Date
2025-05-27
Estimated Expiration
2042-04-29

AI Technical Summary

Technical Problem

The existing digest generation methods are less efficient in the training and generation stages, and it is difficult to effectively capture text structure features, resulting in low accuracy and high repetition rate of generated digests.

Method used

The method of fusion vocabulary coding and structural coding is adopted, word embedding is performed through the BERT model, the TextCNN model is performed for vocabulary coding, the TreeLSTM model is performed for structural coding, and decoding is combined with the Att-LSTM model, and finally the cross entropy loss function is used for model training.

Benefits of technology

It improves the accuracy and efficiency of the summary, can better capture the local information and structural characteristics of the text, reduces the repetition rate, and improves the overall performance of summary generation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114925195B_ABST
    Figure CN114925195B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for generating a standard content text abstract by integrating lexical encoding and structural encoding. The steps are as follows: (1) determining a serialized vector of the standard content; (2) performing lexical encoding output through processing by a TextCNN model; (3) performing structural encoding output through processing by a TreeLSTM model; (4) performing decoding through processing by an Att-LSTM model; (5) determining a loss function. Compared with traditional encoding, the present invention can extract more accurate local information and syntactic structure information, further strengthen the core vocabulary and key grammar in the text in the abstract expression, and effectively improve the accuracy of generating the standard content text abstract.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of text summarization generation and standard digitization. Specifically, it is mainly a method for generating a text summary of "Standard" content that integrates lexical encoding and structural encoding. Background Art

[0002] The extraction of the summary of standard content is an essential link in the process of standard digitization and also a necessary link in standard digital management. Accurately extracting the summary of standard content can greatly improve the efficiency of users retrieving corresponding standards. Currently, text summarization tasks can be divided into two categories according to the implementation method: extractive summarization and generative summarization. Extractive summarization uses specific scoring rules and sorting methods to extract several important sentences from the original text to form a summary. Generative summarization, on the other hand, automatically generates a semantically coherent short text by understanding the context semantic information. Compared with extractive summarization, generative summarization is more in line with the human's language cognitive habits. In text generation tasks, recurrent neural networks are often used as encoders and decoders. By taking advantage of their ability to process sequences word by word, they can effectively and accurately understand the information expressed in the source text and convert it into another form, achieving good results in the generation field.

[0003] In the process of implementing the embodiments of the present invention, the inventors found that there are at least the following defects in the background art. Since RNN and its variants must wait for the output of the previous neuron as the input of the current neuron, it is difficult to achieve parallel computing, resulting in low efficiency in the training and generation stages, and problems such as low accuracy and high repetition rate of the generated summary; although traditional summary generation methods can extract local information well, they completely ignore the structural features of the text and cannot generate satisfactory summaries.

[0004] Standard content has the characteristics of obvious cross-references, with concepts overlapping with each other, and obvious key differences between different words. Therefore, a summary extraction method that can better consider the structural features of the text while completing the local information extraction task is needed for standard content summary extraction, so as to improve the accuracy of summary extraction and provide support for standard content text summary generation. Summary of the Invention

[0005] In view of the problems existing in the above-mentioned existing technologies, the technical problem to be solved by the present invention is to provide a method for generating a text summary of "Standard" content that integrates lexical encoding and structural encoding, and its specific process is as Figure 1 shown.

[0006] The implementation steps of the technical solution are as follows:

[0007] (1) Determine the serialized vector E of the standard content:

[0008] Use the word embedding vectors pre-trained by the BERT model to represent the words in the input text. The BERT model can obtain low-dimensional and dense word vectors through training on large-scale text data. These word vectors can represent the semantic information of words and have good performance in terms of the accuracy of semantic expression information.

[0009] Obtain the sentence representation in the text

[0010] W = [w 1 , w 2 ,..., w N

[0011] After passing through the word embedding layer, the text representation is converted to

[0012] E = [e 1 , e 2 ,..., e N , e i ∈R d

[0013] where E represents the character array of the sentence text after preprocessing, and e i represents the serialized characters of the i-th word in the text, and d is the dimension of the word vector.

[0014] (2) Process through the TextCNN model for lexical encoding to output r:

[0015] The present invention uses a TextCNN neural network to implement the lexical encoder. The TextCNN uses multiple convolutional kernels of different scales to extract key information in the sentence, so as to better capture local correlations. Its input vector sequence is E = [e 1 , e 2 ,..., e N , with a size of (N×d), where N represents the number of text words and d represents the dimension of the word vector.

[0016] Set equal-width convolutional kernels of lengths (2, 3, 4), where each length of convolutional kernel includes M. The process of the convolutional operation is:

[0017] conv i = f(w·e i:i+h-1 + b)

[0018] where h represents the length of the convolutional kernel; w is the weight of the convolutional kernel; b is the bias term; the function f represents the non-linear activation function ReLU; K is the width of the convolutional kernel. For a convolutional kernel of size (h×K), the feature map size obtained through the convolutional operation is (N - h + 1, 1).

[0019] The feature map vector representation is:

[0020] ​c = [c 1 , c 2 ,..., c N-h+1

[0021] Through the max-pooling layer, the maximum value of a single feature map is taken to obtain

[0022]

[0023] The maximum response values of all feature maps are concatenated to obtain the text representation vector r. The optimization parameters include the word vector matrix W, the convolutional kernel weight w, and the bias term b:

[0024]

[0025] (3) Processed by the TreeLSTM model for structure encoding to output h:

[0026] Early abstract generation algorithms completely ignored the structural features of the text and could not generate satisfactory text abstracts. The present invention uses a variant TreeLSTM model optimized based on LSTM to extract text structure information for encoding. The method is as follows:

[0027] The input vector sequence E = [e 1 , e 2 ,..., e N , where N represents the number of text words;

[0028] The hidden layer variable h(t) is updated through the following formula in the loop structure, and h(t) is calculated as follows:

[0029] h t = f(h t-1 , e t )

[0030] f t = σ(W f · [h t-1 , e t + b f )

[0031] i t = σ(W t · [h t-1 , e t + b i )

[0032] C t = f t × C t-1 + i t × tanh(W f · [h t-1 , e​t +b c )

[0033] o t =σ(W 0 ·[h t-1 ,e t +b 0 )

[0034] h t =o t ·tanh(C t )

[0035] Among them, f is the mapping of the TreeLSTM unit in the structure encoder; h t-1 is the hidden state variable of the previous time node; W f and b f are the weight matrix and bias vector of the input gate; W f and b f are the weight matrix and bias vector of the forget gate; W o and b o are the weight matrix and bias vector of the input gate; σ and tanh are the activation functions of the model, and each parameter is solved through supervised training.

[0036] The input vector is encoded and mapped by the structure encoder, and finally the hidden layer state is formed:

[0037] h=[h 1 ,h 2 ,...,h n

[0038] (4) Decoding is performed through the Att-LSTM model:

[0039] The decoder is built by an LSTM neural network based on a hybrid attention mechanism, and the probability distribution of each predicted summary word is output in turn according to the text representation vector r and the hidden layer state vector h of the encoder:

[0040] p(y i |y 1 ,y 2 ,...,y i-1 )=g(y i-1 ,s i ,c i )

[0041]

[0042] g is a non-linear transformation used to predict the probability distribution of the generated word y i ; s i represents the hidden layer output vector of the current stage of the decoder; c i ​It is the environmental vector defined in the attention mechanism, and the calculation method is as follows.

[0043]

[0044] c i is the weighted sum of the text representation vector r of the vocabulary encoder, the hidden layer output vector h of the structure encoder, and the attention weight coefficients; r j , h j are respectively the text representation vector of the vocabulary encoder and the hidden layer output vector of the structure encoder; α ij , β ij are respectively the attention coefficients of the vocabulary encoder and the structure encoder.

[0045]

[0046] q ij = a(s i-1 , r j )

[0047] The above formula is the calculation method of the attention coefficient α ij corresponding to the vocabulary encoder, where q ij is the alignment degree between the text representation vector r of the vocabulary encoder and the decoder state vector s i-1 , and a is the similarity calculation function in the attention mechanism.

[0048]

[0049] q ij = a(s i-1 , h j )

[0050] The above formula is the calculation method of the attention coefficient β ij corresponding to the structure encoder, and q' ij is the alignment degree between the hidden layer state vector h of the structure encoder and the decoder state vector s i-1 , and a is the similarity calculation function in the attention mechanism.

[0051] (5) Determine the loss function H:

[0052] Define the loss function for model training using minimized cross-entropy:

[0053]

[0054] N is the number of samples, l is the length of the target summary, represents the j-th vocabulary in the i-th generated summary. The minimum value of H(y) is optimized using the gradient descent method.

[0055] Advantages of the present invention over the prior art:

[0056] (1) Two independent sub-encoders are used in the present invention to extract key information on vocabulary and sentence structure. Compared with traditional encoding, more accurate local information and syntactic structure information can be extracted, greatly improving the accuracy of the abstract.

[0057] (2) The method of the present invention provides a method of hybrid attention mechanism to further strengthen the core vocabulary and key grammar in the text in the abstract expression. BRIEF DESCRIPTION OF THE DRAWINGS

[0058] To better understand the present invention, the following further description is made in conjunction with the accompanying drawings.

[0059] Figure 1 is a flowchart of the steps of the method for generating a text abstract of the "Standard" content that combines vocabulary encoding and structure encoding;

[0060] Figure 2 is a flowchart of the algorithm of the method for generating a text abstract of the "Standard" content that combines vocabulary encoding and structure encoding;

[0061] Figure 3 is a schematic diagram of the network model of the method for generating a text abstract of the "Standard" content that combines vocabulary encoding and structure encoding;

[0062] Figure 4 is a comparison chart of the accuracy of the method for generating a text abstract of the "Standard" content that combines vocabulary encoding and structure encoding; DETAILED DESCRIPTION OF THE INVENTION

[0063] The following further detailed description of the present invention is made through implementation cases.

[0064] In this implementation case, two standard data sets of gas accident standards and hazardous chemical accident standards are selected for testing. Each standard set contains 150 standards. Assuming the maximum abstract length is 40 and the input original text length is 100 characters.

[0065] The method for generating a text abstract of the "Standard" content that combines vocabulary encoding and structure encoding provided by the present invention has an algorithm flow as Figure 2 shown, and the specific steps are as follows:

[0066] (1) Determine the serialized vector E of the standard content:

[0067] Use the word embedding vectors pre-trained by the BERT model to represent the words of the input text. The BERT model can obtain low-dimensional and dense word vectors through training on large-scale text data. These word vectors can represent the semantic information of words and have good effects in terms of the accuracy of semantic expression information.

[0068] Obtain sentence representations in the text

[0069] W = [w 1 , w 2 ,..., w 100

[0070] After passing through the word embedding layer, the text representation is converted to

[0071] E = [e 1 , e 2 ,..., e 100 , e i ∈ R d

[0072] where E represents the character array of the sentence text after preprocessing, e i represents the serialized characters of the i-th word in the text, and d is the word vector dimension; the word vector dimension d = 512.

[0073] (2) Perform lexical encoding output r through the processing of the TextCNN model:

[0074] The present invention uses a TextCNN neural network to implement a lexical encoder. The TextCNN uses multiple convolutional kernels of different scales to extract key information in the sentence, so as to better capture local correlations. Its input vector sequence is E = [e 1 , e 2 ,..., e 100 , with a size of (N × d), where N represents the number of text words and d represents the dimension of the word vector.

[0075] Set equal-width convolutional kernels with lengths of (2, 3, 4). The process of the convolutional operation is as follows:

[0076] conv i = f(w · e i:i+h-1 + b), 1 ≤ i ≤ 100, 2 ≤ h ≤ 4

[0077] where h represents the length of the convolutional kernel; w is the weight of the convolutional kernel, which is a learnable parameter matrix, b is the bias term, with an initial value of 0.01, and the linear activation function ReLU; d is the width of the convolutional kernel. For a convolutional kernel of size (h × d), the size of the feature map obtained through the convolutional operation is (N - h + 1, 1).

[0078] The vector representation of the feature map is:

[0079] c = [c 1 , c 2 ,..., c N-h+1

[0080] ​​Through the max pooling layer, the maximum value of a single feature map is taken to obtain

[0081]

[0082] the maximum response values of all feature maps are concatenated to obtain the text representation vector r. The optimization parameters include the word vector matrix W, the convolutional kernel weight w, and the bias term b:

[0083]

[0084] (3) After being processed by the TreeLSTM model, structural encoding is performed to output h:

[0085] Early abstract generation algorithms completely ignored the structural features of the text and could not generate satisfactory text abstracts. The present invention uses a variant TreeLSTM model optimized based on LSTM to extract text structure information for encoding. The method is as follows:

[0086] The input vector sequence E = [e 1 , e 2 ,..., e 100 , where N represents the number of text words;

[0087] The hidden layer variable h(t) is updated through the following formula in the loop structure, and h(t) is calculated as follows:

[0088] h t = f(h t-1 , e t ), 1 ≤ t ≤ 100

[0089] f t = σ(W f · [h t-1 , e t + b f ), 1 ≤ t ≤ 100

[0090] i t = σ(W t · [h t-1 , e t + b i ), 1 ≤ t ≤ 100

[0091] C t = f t × C t-1 + i t × tanh(W f · [h t-1 , e t + b c ), 1 ≤ t ≤ 100

[0092] o t = σ(W 0 · [h t-1 , e t + b 0 ), 1 ≤ t ≤ 100

[0093] h t = o t · tanh(C t ), 1 ≤ t ≤ 100

[0094] where f is the mapping of the TreeLSTM cell in the structure encoder; h t-1 is the hidden state variable of the previous time node; W f and b f are the weight matrix and bias vector of the input gate, and the matrix initialization distribution satisfies 's random distribution, and the initial value of b is 0; W f and b f are the weight matrix and bias vector of the forget gate, and the weight matrix W initialization distribution satisfies 's random distribution, and b is initialized to 0; W o and b o are the weight matrix and bias vector of the input gate, and the weight matrix W initialization distribution satisfies 's random distribution, and b is initialized to 0; σ and tanh are the activation functions of the model, and each parameter is solved through supervised training.

[0095] The input vector is encoded and mapped by the structure encoder, and finally the hidden layer state is formed:

[0096] h = [h 1 , h 2 ,..., h n , n = 100.

[0097] (4) Decoding is performed through the Att-LSTM model:

[0098] The decoder is built by an LSTM neural network based on a hybrid attention mechanism, and the probability distribution of each predicted summary word is output in turn according to the text representation vector r and the hidden layer state vector h of the encoder:

[0099] p(y i |y 1 , y 2 ,..., y i-1 ) = g(y i-1 , s i , c i ), 1 ≤ i ≤ 100

[0100]

[0101] g is a non - linear transformation used to predict the probability distribution of the generated word y i ; s i represents the hidden - layer output vector of the decoder at the current stage; c i is the context vector defined in the attention mechanism, and the calculation method is as follows

[0102]

[0103] c i is the weighted sum of the text representation vector r of the word encoder, the hidden - layer output vector h of the structure encoder, and the attention weight coefficients; r j , h j are respectively the text representation vector of the word encoder and the hidden - layer output vector of the structure encoder; α ij , β ij are respectively the attention coefficients of the word encoder and the structure encoder

[0104]

[0105] q ij = a(s i-1 , r j )

[0106] The above formula is the calculation method of the attention coefficient α ij corresponding to the word encoder, where q ij is the alignment degree between the text representation vector r of the word encoder and the decoder state vector s i-1 , and a is the similarity calculation function in the attention mechanism

[0107]

[0108] q ij = a(s i-1 , h j )

[0109] The above formula is the calculation method of the attention coefficient β ij corresponding to the structure encoder, q' ij is the alignment degree between the hidden - layer state vector h of the structure encoder and the decoder state vector s i-1 , and a is the similarity calculation function in the attention mechanism. The schematic diagram of the network model is as Figure 3 shown

[0110] (5) Determine the loss function H:

[0111] Define the loss function for model training using minimized cross - entropy

[0112]

[0113] N is the number of samples, and l is the length of the target abstract. represents the j-th word in the i-th generated abstract. The minimum value of H(y) is optimized using the gradient descent method.

[0114] To verify the effectiveness of the present invention, an abstract generation experiment was conducted on the present invention, and the experimental results are as Figure 4 shown. As Figure 4 can be seen, the performance of the model using this method on the dataset is significantly improved compared to other models.

Claims

1. A method for generating a standard content text summary by integrating lexical encoding and structural encoding, characterized in that, it includes the following steps: Step 1: Determine the serialized vector E of the standard content: Use the word embedding vectors pre-trained by the BERT model to represent the words of the input text in vectors; Obtain the sentence representation in the text: W = [w 1 , w 2 ,..., w N ; After passing through the word embedding layer, the text representation is converted to: E = [e 1 , e 2 ,..., e N , e i ∈ R d ; Among them, E represents the character array after preprocessing the sentence text, and e i represents the serialized characters of the i-th word in the text, and d is the dimension of the word vector; Step 2: Process through the TextCNN model for lexical encoding and output r: Use convolutional kernels of multiple different scales through TextCNN to extract key information in sentences, and its input vector sequence is E = [e 1 , e 2 ,..., e N ; Step 3: Process through the TreeLSTM model for structural encoding and output h: Use a variant TreeLSTM model optimized based on LSTM to extract text structure information for encoding; Step 4: Process through the Att-LSTM model for decoding: The decoder is built by an LSTM neural network based on a hybrid attention mechanism, and successively outputs the probability distribution of each predicted summary word according to the text representation vector r and the hidden layer state vector h of the encoder: p(y i |y 1 ,y 2 ,...,y i-1 ) = g(y i-1 , s i , c i ); g is a non - linear transformation used to predict the probability distribution of the generated word y i ; s i represents the hidden layer output vector at the current stage of the decoder; c i is the context vector defined in the attention mechanism, and the calculation method is as follows; c i is the weighted sum of the text representation vector r of the lexical encoder, the hidden layer output vector h of the structure encoder, and the attention weight coefficients; r j , h j are respectively the text representation vector of the lexical encoder and the hidden layer output vector of the structure encoder; α ij , β ij are respectively the attention coefficients of the lexical encoder and the structure encoder; q ij = a(s i-1 , r j ); The above formula is the calculation method of the attention coefficient α corresponding to the lexical encoder, where q ij is the alignment degree between the text representation vector r of the lexical encoder and the decoder state vector s ij in the attention mechanism, and a is the similarity calculation function in the attention mechanism; i-1 ​ q ij = a(s i-1 , h j ); The above formula is the calculation method of the attention coefficient β corresponding to the structure encoder. ij For q i,j it is the alignment degree between the hidden layer state vector h of the structure encoder and the decoder state vector s i-1 and a is the similarity calculation function in the attention mechanism. Step 5: Determine the loss function H: Use the minimized cross-entropy to define the loss function for model training: N is the number of samples, and l is the length of the target abstract. represents the j-th word in the i-th generated abstract; the minimum value of H(y) is optimized using the gradient descent method.

2. The method according to claim 1, characterized in that, Step 2 further includes: Set equal-width convolutional kernels with lengths of (2, 3, 4), where each length of convolutional kernel includes M. The process of the convolutional operation is as follows: conv i = f(w·e i:i+h-1 + b); Among them, h represents the length of the convolutional kernel; w is the weight of the convolutional kernel; b is the bias term; the function f represents the non-linear activation function ReLU; K is the width of the convolutional kernel; for a convolutional kernel of size (h×K), the size of the feature map obtained through the convolutional operation is (N - h + 1, 1); The vector representation of the feature map is: c = [c 1 , c 2 ,..., c N-h+1 ; Through the max pooling layer, the largest value of a single feature map is taken to obtain Concatenate the maximum response values of all feature maps to obtain the text representation vector r. The optimization parameters include the word vector matrix W, the convolutional kernel weight w, and the bias term b:

3. The method according to claim 1, characterized in that, Step 3 further includes: The method is as follows: The input vector sequence E = [e 1 , e 2 ,..., e N , where N represents the number of text words; The loop structure updates the hidden layer variable h(t) through the following formula, and h(t) is calculated as follows: h t = f(h t-1 , e t ); f t = σ(W f · [h t-1 , e t + b f ); i t = σ(W t · [h t-1 , e t + b i ); C t = f t × C t-1 + i t × tanh(W f · [h t-1 , e t + b c ) o t = σ(W 0 ·[h t-1 ,e t +b 0 ); h t = o t ·tanh(C t ) where f is the mapping of the TreeLSTM unit in the structural encoder; h t-1 is the hidden state variable of the previous time node; W f and b f are the weight matrix and bias vector of the input gate; W f and b f are the weight matrix and bias vector of the forget gate; W o and b o are the weight matrix and bias vector of the input gate; σ and tanh are the activation functions of the model, and each parameter is solved through supervised training; The input vector is encoded and mapped through the structure encoder, and finally forms the hidden layer state: h = [h 1 , h 2 ,..., h n 。

Citation Information

Patent Citations

  • A generation type text abstract method fusing a sequence grammar annotation framework

    CN109948162A

  • Text abstract and sentiment classification combined training method

    CN110929030A