Cross-language automatic summary generation method, device, computer equipment and storage medium

By combining convolutional neural networks and recurrent neural networks for global encoding, combined with Transformer networks and self-attention mechanisms for decoding, and combining bundled search algorithms for constraint scoring, the error propagation and data dependence problems in traditional cross-language digest generation methods are solved, and the accuracy and efficiency of digest generation are improved.

CN112711661BActive Publication Date: 2025-06-06华润数字科技(西安)有限公司
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202011642808.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-12-30
Publication Date
2025-06-06
Estimated Expiration
2040-12-30

AI Technical Summary

Technical Problem

The traditional cross-language digest generation method has serious problems of error propagation and data dependence, and the multi-task learning method takes a long time.

Method used

Convolutional neural network and recurrent neural network are used for global encoding, combined with multi-layer Transformer network and self-attention mechanism for decoding, and constrained scoring of candidate digests is used through a cluster search algorithm to generate the final digest text.

Benefits of technology

It improves the accuracy and efficiency of cross-language text summary generation, reduces dependence on data, and is suitable for small-scale cross-language summary generation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN112711661B_ABST
    Figure CN112711661B_ABST
Patent Text Reader

Abstract

The present invention discloses a cross-language automatic summary generation method, device, computer equipment and storage medium, the method comprising: obtaining a bilingual text for which a summary is to be generated, and preprocessing the bilingual text to obtain a text data set; globally encoding the context information in the text data set based on a convolutional neural network and a recurrent neural network to obtain a summary state sequence of the text data set; decoding the summary state sequence using a multi-layer Transformer network, and calculating the decoded result using a self-attention mechanism, and then using the obtained calculation result as a candidate text summary; constraining the candidate text summary through a beam search, thereby scoring the sentences in the candidate text summary, and selecting the sentence with the highest score from the scored candidate text summary as the final summary text. The present invention can effectively improve the accuracy and efficiency of summary generation for cross-language texts.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of natural language processing technology, and in particular to a method, device, computer equipment and storage medium for cross-language automatic summary generation. Background Art

[0002] With the continuous development of Internet technology, international exchanges are becoming more and more frequent, and the information people are exposed to is growing exponentially. Faced with a huge amount of network information flow, how to effectively select the information you need becomes increasingly important, and text automatic summarization technology is one of the means to solve this problem, especially cross-language automatic summarization technology helps people quickly browse massive amounts of international news and literature, and can help people effectively understand the main points of articles written in unfamiliar foreign languages.

[0003] Traditional cross-language summary generation methods generally adopt the method of generating a summary first and then translating it, or the method of translating first and then generating a summary. This pipeline method is intuitive and simple, but due to the lack of parallel data, the above method will face serious error propagation problems, which greatly restricts the quality of the summary. At the same time, due to the difficulty in obtaining cross-language summary datasets, some previous studies have focused on zero-shot learning, that is, using machine translation or single-language summaries or both methods to train cross-language summary generation systems.

[0004] In addition, in current related technologies, a back-and-forth translation strategy is used to obtain large-scale cross-language summary datasets, that is, machine translation and single-language summary generation are merged into the training of cross-language summaries through multi-task learning to improve the quality of summaries. However, this method has two problems: (1) Multi-task methods use ultra-large-scale parallel data from other tasks, resulting in serious dependence on data, which makes it more difficult to migrate to languages ​​with fewer resources. (2) Multi-task methods either need to train cross-language summaries and single-language summaries at the same time, or alternately train cross-language summaries and machine translation, which is extremely time-consuming. Summary of the invention

[0005] The embodiments of the present invention provide a cross-language automatic summary generation method, apparatus, computer equipment and storage medium, aiming to improve the summary generation accuracy and summary generation efficiency for cross-language texts.

[0006] In a first aspect, an embodiment of the present invention provides a cross-language automatic summary generation method, comprising:

[0007] Obtaining a bilingual text for which a summary is to be generated, and preprocessing the bilingual text to obtain a text dataset;

[0008] Globally encoding the context information in the text dataset based on a convolutional neural network and a recurrent neural network to obtain a summary state sequence of the text dataset;

[0009] Decoding the summary state sequence using a multi-layer Transformer network, calculating the decoded result using a self-attention mechanism, and then using the calculated result as a candidate text summary;

[0010] The candidate text summaries are constrained by beam search, so as to score the sentences in the candidate text summaries, and the sentence with the highest score is selected from the scored candidate text summaries as the final summary text.

[0011] In a second aspect, an embodiment of the present invention provides a cross-language automatic summary generation device, including:

[0012] An acquisition unit, used for acquiring a bilingual text for which a summary is to be generated, and preprocessing the bilingual text to obtain a text data set;

[0013] A global encoding unit, used for globally encoding the context information in the text data set based on a convolutional neural network and a recurrent neural network to obtain a summary state sequence of the text data set;

[0014] A first decoding unit is used to decode the summary state sequence using a multi-layer Transformer network, calculate the decoded result using a self-attention mechanism, and then use the calculated result as a candidate text summary;

[0015] The constraint scoring unit is used to constrain the candidate text summary through a beam search, thereby scoring the sentences in the candidate text summary, and select the sentence with the highest score from the scored candidate text summary as the final summary text.

[0016] In a third aspect, an embodiment of the present invention provides a computer device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the cross-language automatic summary generation method as described in the first aspect is implemented.

[0017] In a fourth aspect, an embodiment of the present invention provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the automatic testing method for the WiFi module as described in the first aspect is implemented.

[0018] The embodiment of the present invention provides a method, device, computer equipment and storage medium for cross-language automatic summary generation, the method comprising: obtaining a bilingual text for which a summary is to be generated, and preprocessing the bilingual text to obtain a text data set; globally encoding the context information in the text data set based on a convolutional neural network and a recurrent neural network to obtain a summary state sequence of the text data set; decoding the summary state sequence using a multi-layer Transformer network, and calculating the decoded result using a self-attention mechanism, and then using the calculated result as a candidate text summary; constraining the candidate text summary through a beam search, thereby scoring the sentences in the candidate text summary, and selecting the sentence with the highest score from the scored candidate text summary as the final summary text. The embodiment of the present invention globally encodes the bilingual text by fusing a convolutional neural network and a recurrent neural network, and combines the attention mechanism and the Transformer network to effectively capture the information of the text context, and at the same time introduces a beam search algorithm for summary decoding, thereby improving the accuracy and efficiency of summary generation for cross-language text. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other accompanying drawings can be obtained based on these accompanying drawings without paying any creative work.

[0020] Figure 1 A schematic diagram of a flow chart of a cross-language automatic summary generation method provided by an embodiment of the present invention;

[0021] Figure 2 A schematic diagram of a sub-flow of step S102 in a cross-language automatic summary generation method provided by an embodiment of the present invention;

[0022] Figure 3 A schematic diagram of a sub-flow of step S204 in a cross-language automatic summary generation method provided by an embodiment of the present invention;

[0023] Figure 4 A schematic diagram of a sub-flow of step S103 in a cross-language automatic summary generation method provided by an embodiment of the present invention;

[0024] Figure 5 A schematic diagram of a sub-flow of step S104 in a cross-language automatic summary generation method provided by an embodiment of the present invention;

[0025] Figure 6 Another schematic diagram of a cross-language automatic summary generation method provided by an embodiment of the present invention;

[0026] Figure 7 A schematic block diagram of a cross-language automatic summary generation device provided by an embodiment of the present invention;

[0027] Figure 8 A schematic sub-block diagram of a global encoding unit 702 in a cross-language automatic summary generation device provided by an embodiment of the present invention;

[0028] Fig. 9 A schematic sub-block diagram of a screening unit 804 in a cross-language automatic summary generation device provided by an embodiment of the present invention;

[0029] Fig.10 A sub-schematic block diagram of a first decoding unit 703 in a cross-language automatic summary generation device provided by an embodiment of the present invention;

[0030] Fig.11 A schematic sub-block diagram of a constraint scoring unit 704 in a cross-language automatic summary generation device provided by an embodiment of the present invention;

[0031] Fig.12 Another schematic block diagram of a cross-language automatic summary generation device provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0032] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0033] It should be understood that when used in this specification and the appended claims, the terms "include" and "comprises" indicate the presence of described features, integers, steps, operations, elements and / or components, but do not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or combinations thereof.

[0034] It should also be understood that the terms used in this specification of the present invention are only for the purpose of describing specific embodiments and are not intended to limit the present invention. As used in the specification of the present invention and the appended claims, unless the context clearly indicates otherwise, the singular forms "a", "an" and "the" are intended to include plural forms.

[0035] It should be further understood that the term "and / or" used in the present description and the appended claims refers to any and all possible combinations of one or more of the associated listed items, and includes these combinations.

[0036] See below Figure 1 , Figure 1 A schematic flow chart of a cross-language automatic summary generation method provided by an embodiment of the present invention specifically includes: steps S101 to S104.

[0037] S101, obtaining a bilingual text for which a summary is to be generated, and preprocessing the bilingual text to obtain a text dataset;

[0038] S102, globally encoding the context information in the text data set based on a convolutional neural network and a recurrent neural network to obtain a summary state sequence of the text data set;

[0039] S103, using a multi-layer Transformer network to decode the summary state sequence, and using a self-attention mechanism to calculate the decoded result, and then using the calculated result as a candidate text summary;

[0040] S104: Constraining the candidate text summaries through a beam search, thereby scoring the sentences in the candidate text summaries, and selecting the sentence with the highest score from the scored candidate text summaries as the final summary text.

[0041] In this embodiment, the convolutional neural network and the recurrent neural network are integrated to globally encode the context information in the text data set formed by the bilingual texts for which summaries are to be generated, so as to obtain a corresponding summary state sequence. The summary state sequence is decoded and calculated through the multi-layer Transformer network and the self-attention mechanism, so as to obtain a candidate text summary. Then, the candidate text summary is constrained and scored using a beam search algorithm, so as to obtain the final summary text.

[0042] This embodiment fuses the convolutional neural network and the recurrent neural network to globally encode the text data set, and combines the Transformer network and the attention mechanism to effectively capture contextual information. In addition, the summary decoding constraints are performed through the beam search algorithm, thereby improving the accuracy and efficiency of the final generated summary. It should be noted that the cross-language automatic summary generation method provided in this embodiment is particularly suitable for cross-language summary generation on a smaller scale, thereby reducing dependence on data and improving summary generation efficiency. In addition, the bilingual text described in this embodiment can be a bilingual text containing Chinese and English, and of course it can also be a bilingual text containing other languages. The content of the bilingual text can be, for example, a paper, a report, etc.

[0043] In a specific embodiment, the step of obtaining a bilingual text for which a summary is to be generated and preprocessing the bilingual text to obtain a text dataset includes:

[0044] Obtain a bilingual text of the abstract to be generated according to actual application requirements, such as obtaining a bilingual paper for which an abstract is to be generated;

[0045] Deleting useless symbols from the bilingual text;

[0046] Using a word segmentation tool to perform word segmentation processing on the Chinese text in the bilingual text, and performing part-of-speech restoration and part-of-speech tagging processing on the English text in the bilingual text;

[0047] Stop words in the bilingual text are removed, and then the documents in the bilingual text are converted into vectors, and the text dataset is constructed according to the obtained vectors.

[0048] In another specific embodiment, the En2ZhSum dataset is used as the bilingual text for generating summaries. The En2ZhSum dataset is an English-to-Chinese summary dataset, which contains 370,687 English documents (755 entries per document on average) and Chinese summaries (96 Chinese characters per document on average). Further, the En2ZhSum dataset is divided into 364,687 training pairs, 3,000 validation pairs, and 3,000 test pairs.

[0049] In one embodiment, if Figure 2 As shown, step S102 includes: steps S201 to S204.

[0050] S201, fusing the convolutional neural network and the recurrent neural network into an encoder;

[0051] S202, inputting the context information in the text data set into the encoder as a text input sequence of the encoder;

[0052] S203, mapping the text input sequence into a hidden sequence through the encoder;

[0053] S204: Screen the hidden sequence to obtain the summary state sequence.

[0054] In this embodiment, the convolutional neural network and the recurrent neural network are fused to obtain an encoder for globally encoding the context information in the text data set. The encoder is used to finally encode and map the context information in the text data set into a summary state sequence.

[0055] In one embodiment, if Figure 3 As shown, the step S204 includes: steps S301 to S304.

[0056] S301, calculating the state summary sequence s of the text input sequence according to the following formula:

[0057]

[0058] In the formula, is the forward hidden state vector obtained by mapping, is the backward hidden state vector obtained by mapping the encoder;

[0059] S302, calculate the element according to the following formula Information gain IG i :

[0060]

[0061] Where tanh(·) is the activation function, W g and U g is the weight matrix, v g is the weight vector, b g is the bias vector;

[0062] S303, according to the following formula To filter:

[0063]

[0064] S304, discard The summary state of

[0065] In this embodiment, when the hidden sequence is screened to obtain the summary state sequence, the corresponding state summary sequence is calculated according to the text input sequence, and then the information gain is calculated for the elements in the state summary sequence. At the same time, the elements in the state summary sequence are screened, and the summary states of the elements whose gain is less than or equal to 0 are discarded, and then the remaining summary states in the summary state sequence are retained, that is, the final summary state sequence.

[0066] In one embodiment, if Figure 4 As shown, the step S103 includes: steps S401 to S404.

[0067] S401, using a multi-layer Transformer network to decode the summary state sequence to obtain a decoded sequence;

[0068] S402, performing a linear transformation on the decoding sequence according to the following formula:

[0069]

[0070] Where Q represents the query vector of the decoding sequence, K represents the key vector of the decoding sequence, V represents the value vector of the decoding sequence, and W 1 represents the query vector matrix of the first layer Transformer network, represents the key vector matrix of the first layer Transformer network, Represents the value vector matrix of the first layer Transformer network, H l-1 Indicates the result of the previous layer of Transformer network operation;

[0071] S403, perform an attention mechanism operation on the result of the linear transformation according to the following formula to obtain the output result of the self-attention model:

[0072]

[0073] Where, d k Denotes the dimensions of Q and K, H l-1 represents the result after the previous layer of Transformer network operation, softmax represents the softmax function, and A represents the final result after the self-attention model;

[0074] S404: According to the following formula, the output result of the self-attention model is connected and projected using the feedforward layer to obtain the candidate text summary:

[0075] MultiHead(Q,K,V)=Concat(head 1 , …, head b )W O

[0076]

[0077] Where W O , and are all learnable matrices.

[0078] In this embodiment, a Transformer network is used as a decoder, and a multi-layer bidirectional Transformer encoder module is stacked, each layer contains a multi-head self-attention mechanism block, and information is obtained from different positions representing different subspaces. After the decoding result is obtained through the multi-layer Transformer network, Q (query vector), K (key vector), and V (value vector) are first linearly transformed, and then the linear transformation result is scaled dot product Attention operation h times (h represents the number of multi-head self-attention mechanism blocks), and then the result of the scaled dot product Attention operation is spliced, and then the spliced ​​result is linearly transformed for the second time to obtain the output result of the self-attention model, and then the output result of the self-attention model is connected and projected by the feedforward layer to obtain the final value, that is, the candidate text summary.

[0079] In one embodiment, the step S401 includes:

[0080] The average value of the attention coefficient of the multi-head self-attention mechanism block in each layer of the Transformer network is calculated according to the following formula:

[0081]

[0082] In the formula, α t is the average of the attention coefficients of the multi-head self-attention mechanism blocks in each layer of the Transformer network, α t h is a multi-head attention mechanism block, and h is the h-th multi-head attention mechanism block.

[0083] In this embodiment, the encoder-decoder attention distribution α is used t h To focus on some prominent words in the summary state sequence. And because α t h is a multi-head attention mechanism block, so this embodiment uses the average value as the attention coefficient on the multi-head attention mechanism block.

[0084] In one embodiment, if Figure 5 As shown, the step S104 includes: steps S501 to S506.

[0085] S501. Score each sentence in the candidate text summary according to the following scoring formula, and select the first B sentences with the highest scores as the sentences to be expanded:

[0086]

[0087] Where x represents the characters in the candidate text summary sentence, y tIndicates the word generated at the current moment, Y t-1 represents the sequence Y of candidate text summary sentences expanded up to time t-1 t-1 ={y 1 y 2 ...y t-1};

[0088] S502, performing cyclic expansion on each of the sentences to be expanded according to the preset parameters θ of the convolutional neural network and the recurrent neural network, the beam search width B, and the maximum step length T of sentence expansion, to obtain B×B candidate sentences;

[0089] S503, for each round of expansion, determining whether the number of the sentence-end expansion generation symbols of each candidate sentence reaches B;

[0090] S504, if the number of the end-of-sentence extension generation symbols of the candidate sentence reaches B, then jump out of the current loop, and score the candidate sentence with the number of the end-of-sentence extension generation symbols reaching B using the scoring formula;

[0091] S505, if the number of symbols generated by the end-of-sentence extension of the candidate sentence does not reach B, continue the cyclic extension until the cyclic extension reaches the maximum step length T of the sentence extension, and score the candidate sentence that reaches the maximum step length T of the sentence extension by using the scoring formula;

[0092] S506: Select the candidate sentence with the highest score as the final summary text.

[0093] This embodiment uses a beam search algorithm to perform a heuristic search on the candidate summary texts to reduce the search scope and the complexity of the problem, thereby reducing space consumption and improving time efficiency.

[0094] Specifically, first, score each sentence in the candidate summary text according to the scoring formula, and select multiple sentences with the highest scores as sentences to be expanded, for example, select the top 5 sentences with the highest scores as sentences to be expanded. Then expand the selected sentences to be expanded to obtain corresponding candidate sentences, for example, expand the selected 5 sentences to be expanded by 5×5, so as to obtain 25 candidate sentences. When the number of symbols generated by the end-of-sentence expansion of the candidate sentence meets the requirements, the expansion can be stopped, that is, the loop expansion is jumped out. When the number of symbols generated by the end-of-sentence expansion of the candidate sentence does not meet the requirements, it is necessary to continue to expand until the loop expansion reaches the maximum step length of the sentence expansion. Of course, it can be understood that in the process of continuing to expand, if the number of symbols generated by the end-of-sentence expansion of the candidate sentence meets the requirements, the expansion can also be stopped. Then score the candidate sentences whose number of symbols generated by the end-of-sentence expansion meets the requirements and the candidate sentences whose loop expansion reaches the maximum step length of the sentence expansion, and select the candidate sentence with the highest score as the final summary text.

[0095] In one embodiment, if Figure 6 As shown, the cross-language automatic summary generation method also includes: steps S601 to S607.

[0096] S601, obtaining a bilingual dictionary, and aligning the words in the bilingual dictionary using a fast alignment tool to obtain a bilingual parallel corpus;

[0097] S602, performing machine translation from source sequence to target sequence and from target sequence to source sequence on the words in the bilingual parallel corpus;

[0098] S603: Obtain the probability of the direction from the source sequence to the target sequence and the average value of the direction from the target sequence to the source sequence by maximum likelihood estimation Among them, w 1 is the source sequence, w 2 is the target sequence;

[0099] S604, normalizing the obtained average value to obtain a probabilistic bilingual dictionary;

[0100] S605: In the probabilistic bilingual dictionary, the probability of the direction from the source sequence to the target sequence and the average value of the direction from the target sequence to the source sequence are obtained according to the following formula: The translation probability P T :

[0101]

[0102] In the formula, w j is the jth target sequence;

[0103] S606: Parameter distribution and translation probability P in the convolutional neural network and the recurrent neural network T Performing weighted summation to obtain a target distribution, then reordering the phrases in the bilingual parallel corpus according to the target distribution, and using the ordered results as candidate sequences;

[0104] S607, using the candidate sequence to train the convolutional neural network and the recurrent neural network, and maximizing the target sequence of the trained parameters according to the following formula:

[0105]

[0106] In the formula, y t is a random variable representing N words, and P is the translation probability P T , X is the candidate sequence, θ is the parameter of the convolutional neural network and the recurrent neural network, and t is the constraint on the output result.

[0107] In this embodiment, a reordering mechanism of key phrases is introduced to train the convolutional neural network and the recurrent neural network to improve the performance of the convolutional neural network and the recurrent neural network. Specifically, a preset bilingual dictionary is first obtained, and the bilingual dictionary is aligned using a fast alignment tool to obtain a bilingual parallel corpus containing phrases, and then machine translation of the source sequence to the target sequence and the target sequence to the source sequence is performed on the bilingual parallel corpus. Preferably, in order to improve the quality of word alignment, alignment in only two directions (e.g., Chinese to English and English to Chinese) can be maintained.

[0108] Next, the average dictionary translation probability of the source sequence to the target sequence and the target sequence to the source sequence is obtained by maximum likelihood estimation. Furthermore, the obtained average value of the dictionary translation probability is normalized to obtain a probabilistic bilingual dictionary.

[0109] Then, the translation probability of the probabilistic bilingual dictionary is calculated based on the average value of the dictionary translation probability, and the final target distribution is calculated according to the translation probability and the parameter distribution of the convolutional neural network and the parameter distribution of the recurrent neural network. The reordering of key phrases (i.e., phrases in the bilingual parallel corpus) can be achieved according to the target distribution.

[0110] Figure 7 A schematic block diagram of a cross-language automatic summary generation device 700 provided in an embodiment of the present invention, the device 700 includes:

[0111] An acquisition unit 701 is used to acquire a bilingual text for which a summary is to be generated, and preprocess the bilingual text to obtain a text dataset;

[0112] A global encoding unit 702, configured to globally encode the context information in the text data set based on a convolutional neural network and a recurrent neural network to obtain a summary state sequence of the text data set;

[0113] A first decoding unit 703 is used to decode the summary state sequence using a multi-layer Transformer network, calculate the decoded result using a self-attention mechanism, and then use the calculated result as a candidate text summary;

[0114] The constraint scoring unit 704 is used to constrain the candidate text summary through a beam search, thereby scoring the sentences in the candidate text summary, and select the sentence with the highest score from the scored candidate text summary as the final summary text.

[0115] In one embodiment, if Figure 8 As shown, the global encoding unit 702 includes:

[0116] A fusion unit 801, used to fuse the convolutional neural network and the recurrent neural network into an encoder;

[0117] An input unit 802, configured to input the context information in the text data set into the encoder as a text input sequence of the encoder;

[0118] A mapping unit 803, configured to map the text input sequence into a hidden sequence through the encoder;

[0119] The screening unit 804 is used to screen the hidden sequence to obtain the summary state sequence.

[0120] In one embodiment, if Fig. 9 As shown, the screening unit 804 includes:

[0121] The first calculation unit 901 is used to calculate the state summary sequence s of the text input sequence according to the following formula:

[0122]

[0123] In the formula, is the forward hidden state vector obtained by mapping, is the backward hidden state vector obtained by mapping the encoder;

[0124] The second calculation unit 902 is used to calculate the element according to the following formula Information gain IG i :

[0125]

[0126] Where tanh(·) is the activation function, W g and U g is the weight matrix, v g is the weight vector, b g is the bias vector;

[0127] The third calculation unit 903 is used to calculate the element according to the following formula To filter:

[0128]

[0129] A discarding unit 904 is used to discard The summary state of

[0130] In one embodiment, if Fig.10 As shown, the first decoding unit 703 includes:

[0131] A second decoding unit 1001 is used to decode the summary state sequence using a multi-layer Transformer network to obtain a decoded sequence;

[0132] The linear transformation unit 1002 is used to perform a linear transformation on the decoding sequence according to the following formula:

[0133]

[0134] Where Q represents the query vector of the decoding sequence, K represents the key vector of the decoding sequence, V represents the value vector of the decoding sequence, and W 1 represents the query vector matrix of the first layer Transformer network, represents the key vector matrix of the first layer Transformer network, Represents the value vector matrix of the first layer Transformer network, H l-1 Indicates the result of the previous layer of Transformer network operation;

[0135] The attention operation unit 1003 is used to perform an attention mechanism operation on the result of the linear transformation according to the following formula to obtain the output result of the self-attention model:

[0136]

[0137] Where, d k Denotes the dimensions of Q and K, H l-1 represents the result after the previous layer of Transformer network operation, softmax represents the softmax function, and A represents the final result after the self-attention model;

[0138] The connection and projection unit 1004 is used to connect and project the output results of the self-attention model using the feedforward layer according to the following formula, so as to obtain the candidate text summary:

[0139] MultiHead(Q,K,V)=Concat(head 1 , …, head b )W O

[0140]

[0141] Where W O , and are all learnable matrices.

[0142] In one embodiment, the second decoding unit 1001 includes:

[0143] The average calculation unit is used to calculate the average of the attention coefficients of the multi-head self-attention mechanism blocks in each layer of the Transformer network according to the following formula:

[0144]

[0145] In the formula, α t is the average of the attention coefficients of the multi-head self-attention mechanism blocks in each layer of the Transformer network, α t h is a multi-head attention mechanism block, and h is the h-th multi-head attention mechanism block.

[0146] In one embodiment, if Fig.11 As shown, the constraint scoring unit 704 includes:

[0147] The selection unit 1101 is used to score each sentence in the candidate text summary according to the following scoring formula, and select the first B sentences with the highest scores as the sentences to be expanded:

[0148]

[0149] Where x represents the characters in the candidate text summary sentence, y t Indicates the word generated at the current moment, Y t-1 represents the sequence Y of candidate text summary sentences expanded up to time t-1 t-1 ={y 1 ,y 2 ...y t-1};

[0150] The cyclic extension unit 1102 is used to perform cyclic extension on each of the sentences to be extended according to the preset convolutional neural network and recurrent neural network parameters θ, the beam search width B and the maximum step length T of sentence extension to obtain B×B candidate sentences;

[0151] A judging unit 1103 is used to judge, for each round of expansion, whether the number of the sentence-end expansion generation symbols of each candidate sentence reaches B;

[0152] The jump unit 1104 is used to jump out of the current loop if the number of the sentence-end extension generation symbols of the candidate sentence reaches B, and score the candidate sentence whose number of the sentence-end extension generation symbols reaches B by using the scoring formula;

[0153] A cyclic scoring unit 1105, configured to continue cyclic extension until the cyclic extension reaches a maximum step length T of sentence extension if the number of symbols generated by the end extension of the candidate sentence does not reach B, and score the candidate sentence that reaches the maximum step length T of sentence extension by using the scoring formula;

[0154] The selection unit 1106 is used to select the candidate sentence with the highest score as the final summary text.

[0155] In one embodiment, if Fig.12 As shown, the cross-language automatic summary generation device 700 also includes:

[0156] An alignment unit 1201 is used to obtain a bilingual dictionary and align words in the bilingual dictionary using a fast alignment tool to obtain a bilingual parallel corpus;

[0157] A machine translation unit 1202, configured to perform machine translation from a source sequence to a target sequence and from a target sequence to a source sequence on the words in the bilingual parallel corpus;

[0158] The maximum likelihood estimation unit 1203 is used to obtain the probability of the direction from the source sequence to the target sequence and the average value of the direction from the target sequence to the source sequence through maximum likelihood estimation. Among them, w 1 is the source sequence, w 2 is the target sequence;

[0159] A normalization unit 1204 is used to normalize the obtained average value to obtain a probabilistic bilingual dictionary;

[0160] The probability acquisition unit 1205 is used to obtain the probability of the direction from the source sequence to the target sequence and the average value of the direction from the target sequence to the source sequence in the probabilistic bilingual dictionary according to the following formula: The translation probability P T :

[0161]

[0162] In the formula, w j is the jth target sequence;

[0163] The reordering unit 1206 is used to reorder the parameter distribution and translation probability P in the convolutional neural network and the recurrent neural network. T Performing weighted summation to obtain a target distribution, then reordering the phrases in the bilingual parallel corpus according to the target distribution, and using the ordered results as candidate sequences;

[0164] The training unit 1207 is used to train the convolutional neural network and the recurrent neural network using the candidate sequence, and maximize the target sequence of the trained parameters according to the following formula:

[0165]

[0166] In the formula, y t is a random variable representing N words, and P is the translation probability P T , X is the candidate sequence, θ is the parameter of the convolutional neural network and the recurrent neural network, and t is the constraint on the output result.

[0167] Since the embodiments of the apparatus part correspond to the embodiments of the method part, please refer to the description of the embodiments of the method part for the embodiments of the apparatus part, which will not be repeated here.

[0168] The embodiment of the present invention further provides a computer-readable storage medium on which a computer program is stored, and when the computer program is executed, the steps provided in the above embodiment can be implemented. The storage medium may include: a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, and other media that can store program codes.

[0169] The embodiment of the present invention also provides a computer device, which may include a memory and a processor, wherein a computer program is stored in the memory, and when the processor calls the computer program in the memory, the steps provided in the above embodiment may be implemented. Of course, the computer device may also include various network interfaces, power supplies and other components.

[0170] The various embodiments in the specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same and similar parts between the various embodiments can be referred to each other. For the system disclosed in the embodiment, since it corresponds to the method disclosed in the embodiment, the description is relatively simple, and the relevant parts can be referred to the method part description. It should be pointed out that for ordinary technicians in this technical field, without departing from the principles of this application, several improvements and modifications can be made to the present application, and these improvements and modifications also fall within the scope of protection of the claims of this application.

[0171] It should also be noted that, in this specification, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "comprises", "comprising" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, an element defined by the statement "comprising a ..." does not exclude the presence of other identical elements in the process, method, article or device including the element.

Claims

1. A cross-language automatic summary generation method, It is characterized in that include: Obtaining a bilingual text for which a summary is to be generated, and preprocessing the bilingual text to obtain a text dataset; Globally encoding the context information in the text dataset based on a convolutional neural network and a recurrent neural network to obtain a summary state sequence of the text dataset; Decoding the summary state sequence using a multi-layer Transformer network, calculating the decoded result using a self-attention mechanism, and then using the calculated result as a candidate text summary; Constraining the candidate text summaries through a beam search, thereby scoring the sentences in the candidate text summaries, and selecting the sentence with the highest score from the scored candidate text summaries as the final summary text; Also includes: Obtain a bilingual dictionary, and use a fast alignment tool to align words in the bilingual dictionary to obtain a bilingual parallel corpus; Performing machine translation from source sequence to target sequence and from target sequence to source sequence on the words in the bilingual parallel corpus; The probability of the direction from the source sequence to the target sequence and the average value of the direction from the target sequence to the source sequence are obtained through maximum likelihood estimation. Among them, w 1 is the source sequence, w 2 is the target sequence; The obtained average values ​​are normalized to obtain a probabilistic bilingual dictionary; In the probabilistic bilingual dictionary, the probability of the direction from the source sequence to the target sequence and the average value of the direction from the target sequence to the source sequence are obtained according to the following formula: The translation probability P T : In the formula, w j is the jth target sequence; The parameter distribution and translation probability P in the convolutional neural network and the recurrent neural network T Performing weighted summation to obtain a target distribution, then reordering the phrases in the bilingual parallel corpus according to the target distribution, and using the ordered results as candidate sequences; The candidate sequence is used to train the convolutional neural network and the recurrent neural network, and the trained parameters are maximized according to the following formula: In the formula, y t is a random variable representing N words, and P is the translation probability P T , X is the candidate sequence, θ is the parameter of the convolutional neural network and the recurrent neural network, and t is the constraint on the output result.

2. The cross-language automatic summary generation method according to claim 1, It is characterized in that The globally encoding the context information in the text data set based on the convolutional neural network and the recurrent neural network to obtain a summary state sequence of the text data set includes: fusing the convolutional neural network and the recurrent neural network into an encoder; Inputting the context information in the text data set into the encoder as a text input sequence of the encoder; Mapping the text input sequence into a hidden sequence by the encoder; The hidden sequence is screened to obtain the summary state sequence.

3. The cross-language automatic summary generation method according to claim 2, It is characterized in that The step of screening the hidden sequence to obtain the summary state sequence includes: The state summary sequence s of the text input sequence is calculated according to the following formula: In the formula, is the forward hidden state vector obtained by mapping, is the backward hidden state vector obtained by mapping the encoder; Calculate the elements according to the following formula Information gain IG i : Where tanh(·) is the activation function, W g and U g is the weight matrix, v g is the weight vector, b g is the bias vector; According to the following formula, the elements To filter: throw away The summary state of 4. The cross-language automatic summary generation method according to claim 1, It is characterized in that The method of decoding the summary state sequence using a multi-layer Transformer network, calculating the decoded result using a self-attention mechanism, and then using the calculated result as a candidate text summary includes: Decoding the summary state sequence using a multi-layer Transformer network to obtain a decoded sequence; The decoded sequence is linearly transformed according to the following formula: Where Q represents the query vector of the decoding sequence, K represents the key vector of the decoding sequence, V represents the value vector of the decoding sequence, and W 1 Represents the query vector matrix of the first layer Transformer network, W l K represents the key vector matrix of the first layer Transformer network, Represents the value vector matrix of the first layer Transformer network, H l-1 Indicates the result of the previous layer of Transformer network operation; Perform the attention mechanism operation on the result of the linear transformation according to the following formula to obtain the output result of the self-attention model: Where, d k Denotes the dimensions of Q and K, H l-1 represents the result after the previous layer of Transformer network operation, softmax represents the softmax function, and A represents the final result after the self-attention model; According to the following formula, the output results of the self-attention model are connected and projected using the feedforward layer to obtain the candidate text summary: MultiHead(Q,K,V)=Concat(head 1 ,…,head b )W O Where W O , W i Q , and are all learnable matrices.

5. The cross-language automatic summary generation method according to claim 4, It is characterized in that The method of decoding the summary state sequence using a multi-layer Transformer network to obtain a decoded sequence includes: The average value of the attention coefficient of the multi-head self-attention mechanism block in each layer of the Transformer network is calculated according to the following formula: In the formula, α t is the average of the attention coefficients of the multi-head self-attention mechanism blocks in each layer of the Transformer network, α t h is a multi-head attention mechanism block, and h is the h-th multi-head attention mechanism block.

6. The cross-language automatic summary generation method according to claim 1, It is characterized in that The constraining the candidate text summary by the beam search, thereby scoring the sentences in the candidate text summary, and selecting the sentence with the highest score from the scored candidate text summary as the final summary text, includes: Each sentence in the candidate text summary is scored according to the following scoring formula, and the top B sentences with the highest scores are selected as the sentences to be expanded: Where x represents the characters in the candidate text summary sentence, y t Indicates the word generated at the current moment, Y t-1 represents the sequence Y of candidate text summary sentences expanded up to time t-1 t-1 ={y 1 y 2 ...y t-1 }; According to the preset parameters θ of the convolutional neural network and the recurrent neural network, the beam search width B and the maximum step length T of the sentence expansion, each of the sentences to be expanded is cyclically expanded to obtain B×B candidate sentences; For each round of expansion, determining whether the number of the sentence-end expansion generated symbols of each candidate sentence reaches B; If the number of the end-of-sentence extension generation symbols of the candidate sentence reaches B, then the current loop is jumped out, and the candidate sentence whose number of the end-of-sentence extension generation symbols reaches B is scored by the scoring formula; If the number of symbols generated by the end-of-sentence extension of the candidate sentence does not reach B, continue the cyclic extension until the cyclic extension reaches the maximum step length T of the sentence extension, and score the candidate sentence that reaches the maximum step length T of the sentence extension by using the scoring formula; The candidate sentence with the highest score is selected as the final summary text.

7. A cross-language automatic summary generation device, It is characterized in that include: An acquisition unit, used for acquiring a bilingual text for which a summary is to be generated, and preprocessing the bilingual text to obtain a text data set; A global encoding unit, used for globally encoding the context information in the text data set based on a convolutional neural network and a recurrent neural network to obtain a summary state sequence of the text data set; A first decoding unit is used to decode the summary state sequence using a multi-layer Transformer network, calculate the decoded result using a self-attention mechanism, and then use the calculated result as a candidate text summary; A constraint scoring unit, used to constrain the candidate text summary through a beam search, thereby scoring the sentences in the candidate text summary, and selecting the sentence with the highest score from the scored candidate text summary as the final summary text; Also includes: An alignment unit, used for acquiring a bilingual dictionary, and aligning the words in the bilingual dictionary using a fast alignment tool to obtain a bilingual parallel corpus; A machine translation unit, used for performing machine translation from source sequence to target sequence and from target sequence to source sequence on words in the bilingual parallel corpus; The maximum likelihood estimation unit is used to obtain the probability of the direction from the source sequence to the target sequence and the average value of the direction from the target sequence to the source sequence through maximum likelihood estimation. Among them, w 1 is the source sequence, w 2 is the target sequence; A normalization unit, used for normalizing the obtained average value to obtain a probabilistic bilingual dictionary; A probability acquisition unit is used to obtain the probability of the direction from the source sequence to the target sequence and the average value of the direction from the target sequence to the source sequence in the probabilistic bilingual dictionary according to the following formula: The translation probability P T : In the formula, w j is the jth target sequence; A reordering unit is used to reorder the parameter distribution and translation probability P in the convolutional neural network and the recurrent neural network. T Performing weighted summation to obtain a target distribution, then reordering the phrases in the bilingual parallel corpus according to the target distribution, and using the ordered results as candidate sequences; A training unit is used to train the convolutional neural network and the recurrent neural network using the candidate sequence, and maximize the target sequence of the trained parameters according to the following formula: In the formula, y t is a random variable representing N words, and P is the translation probability P T , X is the candidate sequence, θ is the parameter of the convolutional neural network and the recurrent neural network, and t is the constraint on the output result.

8. A computer device, It is characterized in that The invention comprises a memory, a processor and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the cross-language automatic summary generation method according to any one of claims 1 to 6 is implemented.

9. A computer-readable storage medium, It is characterized in that The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the cross-language automatic summary generation method according to any one of claims 1 to 6 is implemented.

Citation Information

Patent Citations

  • Automatic text abstraction method

    CN110390010A

  • Two-stage text abstract generation method for long document

    CN111651589A

  • Mongolian online handwritten form recognition method

    CN111695527A