Model training method, text abstract generation method, device, equipment and medium

CN115934930BActive Publication Date: 2026-09-18BEIJING JINGDONG SHANGKE INFORMATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111210416.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-10-18
Publication Date
2026-09-18
Estimated Expiration
2041-10-18

AI Technical Summary

Benefits of technology

[0027] By encoding the training text using an encoder in the summarization model to obtain the encoded hidden state, and then using a decoder in the same model to determine the first decoded hidden state, the summarization model determines the decoding probability of the first predicted word. In response to the first predicted word being included in the summary information annotated in the training text, the summarization model is trained to maximize the decoding probability. Therefore, training the summarization model based on the decoding probability of each predicted word and whether each predicted word is included in the summary information annotated in the training text can improve the prediction performance of the summarization model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115934930B_ABST
    Figure CN115934930B_ABST
Patent Text Reader

Abstract

The application provides a model training method, a text summary generation method, a device, equipment and a medium. The method comprises the following steps: encoding a training text by using an encoder in a summary generation model to obtain an encoded hidden layer state; determining an output first decoding hidden layer state by using a decoder in the summary generation model according to the encoded hidden layer state; determining a decoding probability of a first predicted word according to the first decoding hidden layer state; and training the summary generation model in response to the first predicted word being contained in annotated summary information of the training text, so as to maximize the decoding probability. Thus, the summary generation model is trained based on the decoding probability of each predicted word and whether each predicted word is contained in the annotated summary information of the training text, so that the prediction effect of the summary generation model can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, and in particular to a model training method, a text summarization generation method, an apparatus, a device, and a medium. Background Technology

[0002] Text summarization refers to the technology of using computer equipment to analyze text, summarize its content, and automatically generate summaries to express the main content of the text in a concise form, thereby improving readers' reading speed and quality. Currently, deep learning-based summarization models can be used to extract summaries from text. Training these models is crucial for improving their predictive performance. Summary of the Invention

[0003] This application aims to at least partially address one of the technical problems in the related art.

[0004] This application proposes a model training method, a text summarization generation method, an apparatus, a device, and a medium to train a summarization generation model and improve its prediction performance.

[0005] The first aspect of this application proposes a model training method, including:

[0006] Obtain training text, wherein the training text is annotated with summary information;

[0007] The training text is encoded using an encoder in a summary generation model to obtain the encoded hidden state;

[0008] Based on the encoded hidden layer state, the decoder in the digest generation model is used to determine the first decoded hidden layer state of the output.

[0009] The decoding probability of the first predicted word is determined based on the first decoding hidden layer state;

[0010] In response to the first predicted word being included in the summary information, the summary generation model is trained to maximize the decoding probability.

[0011] A second aspect of this application provides a text summarization method, comprising:

[0012] Get the input text;

[0013] A trained summary generation model is used to predict the input text to obtain summary information corresponding to the input text; wherein the summary generation model is trained using the model training method described in the first aspect embodiment of this application.

[0014] A third aspect of this application provides a model training apparatus, comprising:

[0015] The acquisition module is used to acquire training text, wherein the training text is annotated with summary information;

[0016] The encoding module is used to encode the training text using the encoder in the summary generation model to obtain the encoded hidden layer state;

[0017] The first determining module is used to determine the first decoding hidden layer state of the output using the decoder in the digest generation model based on the encoded hidden layer state;

[0018] The second determining module is used to determine the decoding probability of the first predicted word based on the first decoding hidden layer state;

[0019] A training module is used to train the summary generation model in response to the first predicted word being contained in the summary information, so as to maximize the decoding probability.

[0020] A fourth aspect of this application provides a text summarization apparatus, comprising:

[0021] The acquisition module is used to acquire the input text;

[0022] The prediction module is used to predict the input text using a trained summary generation model to obtain summary information corresponding to the input text; wherein the summary generation model is trained using the model training device as described in the third aspect embodiment of this application.

[0023] A fifth aspect of this application provides a computer device, including: a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, it implements a model training method as proposed in the first aspect of this application, or a text summarization generation method as proposed in the second aspect of this application.

[0024] A sixth aspect of this application provides a non-transitory computer-readable storage medium storing a computer program that, when executed by a processor, implements the model training method proposed in the first aspect of this application, or the text summarization generation method proposed in the second aspect of this application.

[0025] A seventh aspect of this application provides a computer program product that, when executed by a processor, performs a model training method as proposed in a first aspect of this application, or a text summarization method as proposed in a second aspect of this application.

[0026] One embodiment of this application described above has at least the following advantages or beneficial effects:

[0027] By encoding the training text using an encoder in the summarization model to obtain the encoded hidden state, and then using a decoder in the same model to determine the first decoded hidden state, the summarization model determines the decoding probability of the first predicted word. In response to the first predicted word being included in the summary information annotated in the training text, the summarization model is trained to maximize the decoding probability. Therefore, training the summarization model based on the decoding probability of each predicted word and whether each predicted word is included in the summary information annotated in the training text can improve the prediction performance of the summarization model.

[0028] Additional aspects and advantages of this application will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of this application. Attached Figure Description

[0029] The above and / or additional aspects and advantages of this application will become apparent and readily understood from the following description of the embodiments taken in conjunction with the accompanying drawings, wherein:

[0030] Figure 1 This is a flowchart illustrating the model training method provided in Embodiment 1 of this application;

[0031] Figure 2 This is a schematic flowchart of the model training method provided in Embodiment 2 of this application;

[0032] Figure 3 This is a schematic flowchart of the model training method provided in Embodiment 3 of this application;

[0033] Figure 4 This is a schematic flowchart of the model training method provided in Embodiment 4 of this application;

[0034] Figure 5 This is a flowchart illustrating the text digest generation method provided in Embodiment 5 of this application;

[0035] Figure 6 This is a schematic diagram of the model training device provided in Embodiment Six of this application;

[0036] Figure 7 This is a schematic diagram of the text summarization device provided in Embodiment 7 of this application;

[0037] Figure 8 A block diagram of an exemplary computer device suitable for implementing embodiments of the present application is shown. Detailed Implementation

[0038] The embodiments of this application are described in detail below. Examples of these embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and intended to explain this application, and should not be construed as limiting this application.

[0039] Text generation technology refers to the techniques used by computer devices to automatically generate a valid and fluent text. For example, a bidirectional RNN (Recurrent Neural Network) model can be used as a decoder to encode the input text, while a unidirectional RNN model, based on a copying mechanism, can generate a summary of the input text. The copying mechanism refers to directly selecting certain words from the input text and copying them into the output text when generating the output text (i.e., the summary of the input text).

[0040] Traditional copying mechanisms rely solely on the current state of the decoder when deciding which words to copy. However, in practice, it's common to copy a segment of text consecutively, meaning the current copying decision should be related to the previous decision. Therefore, traditional copying mechanisms may lead to low accuracy and reliability in model predictions.

[0041] Therefore, in order to address the above problems, this application proposes a model training method, apparatus, device, and storage medium.

[0042] The model training method, text summarization method, apparatus, device, and medium of embodiments of this application are described below with reference to the accompanying drawings.

[0043] Figure 1 This is a flowchart illustrating the model training method provided in Embodiment 1 of this application.

[0044] This application example illustrates the model training method configured in a model training device, which can be applied to any electronic device to enable the electronic device to perform model training functions.

[0045] Among them, electronic devices can be any device with computing capabilities, such as personal computers, mobile terminals, servers, etc. Mobile terminals can be hardware devices with various operating systems, touch screens and / or displays, such as mobile phones, tablets, personal digital assistants, wearable devices, etc.

[0046] like Figure 1 As shown, the model training method may include the following steps:

[0047] Step 101: Obtain the training text, which is annotated with summary information.

[0048] In one possible implementation of this application embodiment, the training text can be obtained from an existing dataset, test set, or training set.

[0049] In another possible implementation of this application embodiment, training text can be collected online. For example, some news articles and other texts can be collected online using web crawler technology as training text.

[0050] In another possible implementation of this application, training text can be collected offline, for example, by taking an image containing text information by a user holding an image acquisition device, and then recognizing the text information in the image based on OCR (Optical Character Recognition) technology to obtain training text.

[0051] In another possible implementation of this application's embodiments, the training text can also be obtained through manual input by the user. Input methods include, but are not limited to, touch input (such as swiping, clicking, etc.), keyboard input, and voice input.

[0052] Optionally, to improve the model's prediction performance, the summary information in the training text can be manually annotated.

[0053] Step 102: The training text is encoded using the encoder in the summary generation model to obtain the encoded hidden state.

[0054] In the embodiments of this application, the specific type of the summary generation model is not limited. For example, the summary generation model can be an RNN model, or it can be other models, and there is no limitation on this.

[0055] In this embodiment of the application, the summary generation model may include an encoder and a decoder. The encoder in the summary generation model can be used to encode the training text to obtain the encoded hidden layer state.

[0056] One possible approach is to segment the training text to obtain individual words, and then use an encoder to encode each word to obtain the corresponding hidden state of the encoding layer.

[0057] As an example, let the training text be labeled X, and the i-th word segmented in the training text be x. i Then encoder f enc For the word segmentation vocabulary x i Encoding is performed, and the generated hidden state can be:

[0058] h i =f enc (x i,h i-1 (1)

[0059] Among them, h i The word x represents the participle. i The corresponding encoded hidden state, h i-1 This indicates that the word in the training text is located in the segmented vocabulary x. i Previous adjacent word x i-1 The corresponding coded hidden state.

[0060] In other words, in this application, for any segmented word x in the training text i It can be based on the word x in the training text that is located in the word segmentation vocabulary. i Previous adjacent word x i-1 The corresponding coded hidden state h i-1 Encoder f enc For the word segmentation vocabulary x i Encode the word segmentation vocabulary x i The corresponding coded hidden state h i .

[0061] Taking the summary generation model as an example, using an LSTM (Long Short Term Memory) model, given the input training text X (e.g., a complete news article), and the text generation task y, the encoder f based on a bidirectional LSTM is used to generate the summary. enc Encode each word segment in the training text to obtain the corresponding hidden state of each word segment as shown in formula (1).

[0062] Step 103: Based on the encoded hidden layer state, the decoder in the digest generation model is used to determine the first decoded hidden layer state of the output.

[0063] In this embodiment of the application, the encoding hidden state h corresponding to each word segmentation vocabulary can be used as a basis. i The decoder in the summary generation model is used to determine the first decoded hidden state of the output.

[0064] As an example, a decoder f based on a unidirectional LSTM can be used. dec The first decoding hidden state is generated using a copying mechanism.

[0065] Step 104: Determine the decoding probability of the first predicted word based on the first decoding hidden layer state.

[0066] In this embodiment of the application, the first predicted word refers to the predicted word output in this application.

[0067] In this embodiment of the application, the decoding probability of the first predicted word can be determined according to the first decoding hidden layer state. The smaller the decoding probability, the lower the importance of the first predicted word. Conversely, the larger the decoding probability, the more important the first predicted word is and the more likely it is to be a summary word.

[0068] Step 105: In response to the first predicted vocabulary being contained in the summary information, the summary generation model is trained to maximize the decoding probability.

[0069] In the embodiments of this application, when the first predicted word is included in the summary information, that is, the first predicted word is a summary word, the summary generation model can be trained to maximize the decoding probability.

[0070] The model training method of this application embodiment encodes the training text using an encoder in a summarization generation model to obtain an encoded hidden state; based on the encoded hidden state, a decoder in the summarization generation model determines the output first decoded hidden state; based on the first decoded hidden state, the decoding probability of a first predicted word is determined; in response to the first predicted word being included in the summary information annotated in the training text, the summarization generation model is trained to maximize the decoding probability. Therefore, by training the summarization generation model based on the decoding probability of each predicted word and whether each predicted word is included in the summary information annotated in the training text, the prediction performance of the summarization generation model can be improved.

[0071] To clearly illustrate how the first decoding hidden layer state is determined based on the encoding hidden layer state in this application, this embodiment provides another model training method.

[0072] Figure 2 This is a flowchart illustrating the model training method provided in Embodiment 2 of this application.

[0073] like Figure 2 As shown, the model training method may include the following steps:

[0074] Step 201: Obtain the training text, which is annotated with summary information.

[0075] Step 202: The training text is encoded using the encoder in the summary generation model to obtain the encoded hidden state.

[0076] The execution process of steps 201 to 202 can be referred to the execution process of steps 101 to 102 in the above embodiment, and will not be repeated here.

[0077] Step 203: Generate the decoding vector corresponding to the first predicted word based on the encoding hidden state of each word segment in the training text and the attention weight of the first predicted word to each word segment.

[0078] In this embodiment, the encoding hidden state h of each word segmentation word in the training text can be used as a basis. i The attention weights of the first predicted word to each segmented word are used to generate the decoding vector (or decoding context vector) corresponding to the first predicted word.

[0079] As an example, the attention weight of the first predicted word to each segmented word is α. t,i The decoding vector corresponding to the first predicted word is c. t Then we can get:

[0080] c t =∑ i α t,i h i (1)

[0081] Among them, the attention weight α of the first predicted word to each segmented word t,i The attention weight matrix α t The elements in the attention weight matrix α t It can be calculated using the following formula:

[0082] α t =softmax(e t (2)

[0083] Among them, e t The elements e in t,i It can be calculated using the following formula:

[0084]

[0085] Among them, u a W a V a All are model parameter vectors or matrices, s t-1 This refers to the second decoded hidden state of the most recent (previous) output, or, in the case that this output is the first output, s t-1 This refers to setting the initial second decoding hidden layer state.

[0086] Step 204: Obtain the second decoded hidden state and the corresponding second predicted word from the most recent output.

[0087] In this embodiment, when the current output is not the first output, the second predicted word refers to the predicted word of the most recent (or previous) output. However, when the current output is the first output, the second predicted word can be a pre-defined word. For example, the second predicted word could be "Tsinghua University," and the first predicted word could be "university."

[0088] In the embodiments of this application, when the current output is not the first output, the second decoding hidden state and the corresponding second prediction word of the most recent output can be obtained, while when the current output is the first output, the initial second decoding hidden state and the corresponding second prediction word can be obtained.

[0089] Step 205: The decoder determines the first decoding hidden state of the current output based on the decoding vector corresponding to the first predicted word, the second decoding hidden state, and the second predicted word.

[0090] In this embodiment of the application, a decoder can be used based on the decoding vector c corresponding to the first predicted word. t Second decoding hidden state s t-1 The first decoded hidden state of the current output is determined by the second predicted vocabulary.

[0091] As an example, the second predicted word is labeled y. t-1 The first decoded hidden state output this time is s t Then, the first decoded hidden state s of this output can be determined according to the following formula. t :

[0092] s t =f dec (s t-1 ,y t-1 ,c t (4)

[0093] Step 206: Determine the decoding probability of the first predicted word based on the first decoding hidden layer state.

[0094] Step 207: In response to the first predicted vocabulary being contained in the summary information, the summary generation model is trained to maximize the decoding probability.

[0095] The execution process of steps 206 to 207 can be found in the execution process of any embodiment of this application, and will not be described in detail here.

[0096] The model training method in this application determines the current output hidden layer state by combining the second decoding hidden layer state of the most recent output. This enables the determination of the current state based on the historical state of the decoder, thereby improving the model training effect.

[0097] To clearly illustrate how the decoding probability of the first predicted word is determined based on the first decoding hidden layer state in any of the above embodiments of this application, this embodiment provides another model training method.

[0098] Figure 3 This is a flowchart illustrating the model training method provided in Embodiment 3 of this application.

[0099] like Figure 3 As shown, the model training method may include the following steps:

[0100] Step 301: Obtain the training text, which is annotated with summary information.

[0101] Step 302: The encoder in the summary generation model is used to encode the training text to obtain the encoded hidden state.

[0102] Step 303: Based on the encoded hidden layer state, the decoder in the digest generation model is used to determine the first decoded hidden layer state of the output.

[0103] The execution process of steps 301 to 303 can be found in the execution process of any embodiment of this application, and will not be described in detail here.

[0104] Step 304: Determine the generation probability of the first predicted word based on the first decoded hidden layer state.

[0105] In this embodiment of the application, the generation probability of the first predicted word can be determined based on the first decoded hidden layer state of the current output.

[0106] As one possible implementation, the decoding vector c corresponding to the first predicted word can be obtained. t (For the specific calculation process, please refer to step 203 in the above embodiment), based on the decoding vector c corresponding to the first predicted word... t And the first decoded hidden state s of this output t Determine the generation probability of the first predicted word.

[0107] As an example, let the first predicted word be labeled w, and let the generation probability of the first predicted word w be P. vocab (w), then we can obtain:

[0108] P vocab (w) = softmax(W) b s t +V b c t (5)

[0109] Among them, W b and V b This refers to a vector or matrix of model parameters.

[0110] Step 305: Generate the initial copy probability of the first predicted word based on the attention weight of the first predicted word to each word segment in the training text.

[0111] It should be noted that the attention weight α of the first predicted word to each word segment in the training text t,iIt can be calculated using formulas (2) and (3) in step 203, which will not be elaborated here.

[0112] Optionally, the initial replication probability of the first predicted word w is P. copy_original Then we can get:

[0113]

[0114] Step 306: Determine the target replication probability corresponding to the first predicted word based on the initial replication probability of the first predicted word and the target replication probability corresponding to the most recently output second predicted word.

[0115] In this embodiment of the application, when the current output is not the first output, the target replication probability corresponding to the second predicted word in the most recent output can be obtained, based on the initial replication probability P of the first predicted word. copy_original The target replication probability corresponding to the first predicted word is determined by the target replication probability corresponding to the second predicted word in the most recent output.

[0116] In this embodiment of the application, when this output is the first output, the target replication probability corresponding to the initial second predicted word can be obtained, based on the initial replication probability P of the first predicted word. copy_original The target replication probability corresponding to the first predicted word is determined by setting the initial target replication probability for the second predicted word.

[0117] Optionally, the target replication probability corresponding to the first predicted word is P. copy_new .

[0118] Step 307: Determine the decoding probability corresponding to the first predicted word based on the generation probability and the target replication probability corresponding to the first predicted word.

[0119] In this embodiment of the application, the generation probability P corresponding to the first predicted word can be used as a basis. vocab And the target replication probability P copy_new Determine the decoding probability corresponding to the first predicted word.

[0120] As an example, if we label the decoding probability corresponding to the first predicted word w as P(w), then we can obtain:

[0121] P(w)=p gen *P vocab (w)+(1-p gen )*P copy_new (w); (7)

[0122] Where, p gen It can be obtained through the following formula:

[0123]

[0124] Among them, w c w s w x and scalar b c Here are the model parameters, σ is the activation function (i.e., the sigmoid function), and y is the model parameter. t The predicted character for this output is w, which is the first predicted word.

[0125] Step 308: In response to the first predicted vocabulary being contained in the summary information, the summary generation model is trained to maximize the decoding probability.

[0126] The execution process of step 308 can be found in any embodiment of this application, and will not be described in detail here.

[0127] The model training method in this application determines the target replication probability of the first predicted word in the current output based on the target replication probability corresponding to the second predicted word in the most recent (i.e., previous) output. This can associate replication decisions between decoding times, thereby optimizing the replication mechanism and improving the training effect of the model.

[0128] To clearly illustrate how the target replication probability corresponding to the first predicted word is determined in the above embodiments of this application, this embodiment provides another model training method.

[0129] Figure 4 This is a flowchart illustrating the model training method provided in Embodiment 4 of this application.

[0130] like Figure 4 As shown, the model training method may include the following steps:

[0131] Step 401: Obtain the training text, which is annotated with summary information.

[0132] Step 402: The encoder in the summary generation model is used to encode the training text to obtain the encoded hidden state.

[0133] Step 403: Based on the encoded hidden layer state, the decoder in the digest generation model is used to determine the first decoded hidden layer state of the output.

[0134] Step 404: Determine the generation probability of the first predicted word based on the first decoded hidden layer state.

[0135] Step 405: Generate the initial copy probability of the first predicted word based on the attention weight of the first predicted word to each word segment in the training text.

[0136] The execution process of steps 401 to 405 can be found in the execution process of any embodiment of this application, and will not be described in detail here.

[0137] Step 406: Obtain the target replication probability corresponding to the second predicted word in the most recent output.

[0138] In this embodiment of the application, when the current output is not the first output, the target replication probability corresponding to the second predicted word of the most recent output can be obtained, while when the current output is the first output, the target replication probability corresponding to the second predicted word that is set initially can be obtained.

[0139] Optionally, the target replication probability corresponding to the second predicted word is P. copyprevious .

[0140] Step 407: Determine the correlation between the second predicted word and the first predicted word.

[0141] In the embodiments of this application, the relevance represents the degree of association between two words. For example, the relevance between "spring breeze" and "gentle breeze" is higher than that between "winter wind" and "gentle breeze".

[0142] In this embodiment of the application, the correlation between the second predicted word and the first predicted word can be determined based on the correlation calculation formula.

[0143] It is understandable that for any two words, the closer the distance between them (such as text distance or dependency tree distance), the greater the correlation between the two words.

[0144] Therefore, in one possible implementation of this application embodiment, the text length of the training text can be determined, and the relative position difference between the second predicted word and the first predicted word in the training text can be determined, so that the correlation between the second predicted word and the first predicted word can be determined based on the above-mentioned text length, relative position difference and first decoding hidden layer state.

[0145] As an example, the correlation between any two words in the training text, such as word v and word w, can be calculated using the following formula: rel(v,w)

[0146]

[0147] Where pos(v,w) refers to the relative position difference between words v and w in the training text, and δ can be calculated using the following formula:

[0148]

[0149] Among them, w s"X" refers to the model parameters, and "X" refers to the length of the training text.

[0150] Step 408: Based on the aforementioned correlation, the initial replication probability of the first predicted word, and the target replication probability of the second predicted word, determine the target replication probability corresponding to the first predicted word.

[0151] In this embodiment of the application, in order to encourage the model to coherently reproduce the input text, when calculating the target reproduction probability corresponding to the first predicted word in the current output, the target reproduction probability corresponding to the second predicted word in the most recent output, as well as the correlation between the second predicted word and the first predicted word, can be combined to determine the target reproduction probability corresponding to the first predicted word.

[0152] Optionally, the target replication probability corresponding to the first predicted word is P. copy_new The target replication probability corresponding to the first predicted word is P. copy_new The target replication probability P can be determined using the following formula. copy_new :

[0153] P copy_new (w)=p t *P copy_original (w)+(1-p t )*∑ v∈X (P copyprevious (w)*rel(v,w)); (11)

[0154] Where, p t It can be obtained through the following formula:

[0155]

[0156] Among them, w p and scalar b p For model parameters, This refers to the target replication probability corresponding to the second predicted word in the most recent output, while rel(v,w) refers to the replication probability of the target replication probability of the second predicted word in the most recent output when it is transferred to the first predicted word w. rel(v,w) refers to the correlation between word v and word w.

[0157] Step 409: Determine the decoding probability corresponding to the first predicted word based on the generation probability and target replication probability corresponding to the first predicted word.

[0158] Step 410: In response to the first predicted vocabulary being contained in the summary information, the summary generation model is trained to maximize the decoding probability.

[0159] The execution process of steps 409 to 410 can be found in the execution process of any embodiment of this application, and will not be described in detail here.

[0160] The model training method of this application embodiment, since the correlation degree represents the degree of correlation between different words, determines the target replication probability of the second predicted word output this time based on the correlation degree between the second predicted word and the first predicted word, and the target replication probability corresponding to the most recently output second predicted word. This not only realizes the correlation between replication decisions at each decoding time, but also realizes the determination of the target replication probability corresponding to the second predicted word output this time based on the degree of correlation or importance between each word. This can further optimize the replication mechanism and thus improve the training effect of the model.

[0161] The above describes the training method for the summary generation model. This application also proposes an application method for the summary generation model.

[0162] Figure 5 This is a flowchart illustrating the text digest generation method provided in Embodiment 5 of this application.

[0163] like Figure 5 As shown, the text summarization method may include the following steps:

[0164] Step 501: Obtain the input text.

[0165] In one possible implementation of this application embodiment, the input text can be obtained from an existing dataset or test set.

[0166] In another possible implementation of this application embodiment, the input text can be collected online. For example, web crawler technology can be used to collect some news, articles and other texts online as input text.

[0167] In another possible implementation of this application, training text can be collected offline, for example, by taking an image containing text information by a user holding an image acquisition device, and then recognizing the text information in the image based on OCR technology to obtain the input text.

[0168] In another possible implementation of this application embodiment, the input text can also be obtained by the user manually inputting it.

[0169] Step 502: The trained summary generation model is used to predict the input text to obtain the summary information corresponding to the input text; wherein, the summary generation model is trained using the model training method proposed in any of the above embodiments.

[0170] In this embodiment of the application, a summary generation model trained using the model training method proposed in any of the above embodiments can be obtained. The trained summary generation model is then used to predict the input text in order to obtain the summary information corresponding to the input text.

[0171] The text summarization method of this application employs a trained summarization model to predict the input text and obtain the corresponding summary information. The summarization model is trained using the model training method proposed in any of the above embodiments. Therefore, the deep learning-based summarization model improves the accuracy and reliability of summarization extraction.

[0172] With the above Figures 1 to 4 Corresponding to the model training method provided in the embodiments, this application also provides a model training apparatus. Since the model training apparatus provided in the embodiments of this application is different from the one described above... Figures 1 to 4 The model training method provided in the embodiments corresponds to the model training device provided in the embodiments of this application, and will not be described in detail in the embodiments of this application.

[0173] Figure 6 This is a schematic diagram of the model training device provided in Embodiment Six of this application.

[0174] like Figure 6 As shown, the model training device 600 may include: an acquisition module 610, an encoding module 620, a first determination module 630, a second determination module 640, and a training module 650.

[0175] The acquisition module 610 is used to acquire training text, which is annotated with summary information.

[0176] The encoding module 620 is used to encode the training text using the encoder in the summary generation model to obtain the encoded hidden state.

[0177] The first determining module 630 is used to determine the first decoded hidden state of the output using the decoder in the digest generation model based on the encoded hidden state.

[0178] The second determining module 640 is used to determine the decoding probability of the first predicted word based on the first decoding hidden layer state.

[0179] Training module 650 is used to train the summary generation model in response to the first predicted words contained in the summary information, so as to maximize the decoding probability.

[0180] In one possible implementation of this application embodiment, the first determining module 630 is specifically used to: generate a decoding vector corresponding to the first predicted word based on the encoding hidden state of each word segment in the training text and the attention weight of the first predicted word to each word segment; obtain the second decoding hidden state and the corresponding second predicted word of the most recent output; and use a decoder to determine the first decoding hidden state of the current output based on the decoding vector corresponding to the first predicted word, the second decoding hidden state, and the second predicted word.

[0181] In one possible implementation of this application embodiment, the second determining module 640 includes:

[0182] The first determining unit is used to determine the generation probability of the first predicted word based on the first decoded hidden layer state.

[0183] The generation module is used to generate the initial copy probability of the first predicted word based on the attention weight of the first predicted word to each segmented word.

[0184] The second determining unit is used to determine the target replication probability corresponding to the first predicted word based on the initial replication probability of the first predicted word and the target replication probability corresponding to the most recently output second predicted word.

[0185] The third determining unit is used to determine the decoding probability corresponding to the first predicted word based on the generation probability and the target replication probability corresponding to the first predicted word.

[0186] In one possible implementation of this application embodiment, the first determining unit is specifically used to: determine the generation probability of the first predicted word based on the decoding vector corresponding to the first predicted word and the first decoding hidden layer state.

[0187] In one possible implementation of this application, the second determining unit is specifically used to: determine the correlation between the second predicted word and the first predicted word; and determine the target replication probability corresponding to the first predicted word based on the correlation, the initial replication probability of the first predicted word, and the target replication probability corresponding to the second predicted word.

[0188] In one possible implementation of this application, the second determining unit is specifically used for: determining the text length of the training text; determining the relative position difference between the second predicted word and the first predicted word in the training text; and determining the correlation between the second predicted word and the first predicted word based on the text length, the relative position difference, and the first decoding hidden layer state.

[0189] In one possible implementation of this application embodiment, the encoding module 620 is specifically used to: perform word segmentation on the training text to obtain each word segmentation word; and use an encoder to encode each word segmentation word to obtain the encoding hidden state corresponding to each word segmentation word.

[0190] In one possible implementation of this application embodiment, the encoding module 620 is specifically used to: for any one-segment word, according to the encoding hidden state corresponding to the adjacent word in the training text that precedes the one-segment word, use an encoder to encode the one-segment word to obtain the encoding hidden state corresponding to the one-segment word.

[0191] The model training apparatus of this application embodiment encodes the training text using an encoder in a summarization generation model to obtain an encoded hidden state; based on the encoded hidden state, a decoder in the summarization generation model determines the output first decoded hidden state; based on the first decoded hidden state, the decoding probability of a first predicted word is determined; and in response to the first predicted word being included in the summary information annotated in the training text, the summarization generation model is trained to maximize the decoding probability. Therefore, by training the summarization generation model based on the decoding probability of each predicted word and whether each predicted word is included in the summary information annotated in the training text, the prediction performance of the summarization generation model can be improved.

[0192] With the above Figure 5 Corresponding to the text summarization method provided in the embodiments, this application also provides a text summarization apparatus. Since the text summarization apparatus provided in the embodiments of this application is similar to the one described above... Figure 5 The text summarization method provided in the embodiments corresponds to the text summarization method provided in the embodiments of this application. Therefore, the implementation of the text summarization method is also applicable to the text summarization apparatus provided in the embodiments of this application, and will not be described in detail in the embodiments of this application.

[0193] Figure 7 This is a schematic diagram of the text summarization device provided in Embodiment 7 of this application.

[0194] like Figure 7 As shown, the text summarization generation device 700 may include an acquisition module 710 and a prediction module 720.

[0195] The acquisition module 710 is used to acquire the input text.

[0196] The prediction module 720 is used to predict the input text using a trained summary generation model to obtain the summary information corresponding to the input text; wherein the summary generation model is trained using the model training device proposed in the aforementioned embodiments.

[0197] The text summarization generation apparatus of this application employs a trained summarization generation model to predict the input text and obtain the corresponding summary information. The summarization generation model is trained using the model training method proposed in any of the above embodiments. Therefore, the deep learning-based summarization generation model can improve the accuracy and reliability of summarization extraction.

[0198] To implement the above embodiments, this application also proposes a computer device, including: a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, it implements the model training method proposed in any of the foregoing embodiments of this application, or the text summarization generation method proposed in the foregoing embodiments of this application.

[0199] To implement the above embodiments, this application also proposes a non-transitory computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the model training method proposed in any of the foregoing embodiments of this application, or implements the text summarization generation method proposed in the foregoing embodiments of this application.

[0200] To implement the above embodiments, this application also proposes a computer program product, which, when the instructions in the computer program product are executed by a processor, executes the model training method proposed in any of the foregoing embodiments of this application, or executes the text summarization generation method proposed in the foregoing embodiments of this application.

[0201] Figure 8 A block diagram of an exemplary computer device suitable for implementing embodiments of the present application is shown. Figure 8 The computer device 12 shown is merely an example and should not impose any limitation on the functionality and scope of use of the embodiments of this application.

[0202] like Figure 8 As shown, the computer device 12 is represented in the form of a general-purpose computing device. The components of the computer device 12 may include, but are not limited to: one or more processors or processing units 16, system memory 28, and a bus 18 connecting different system components (including system memory 28 and processing unit 16).

[0203] Bus 18 represents one or more of several bus architectures, including a memory bus or memory controller, a peripheral bus, a graphics acceleration port, a processor, or a local bus using any of the various bus architectures. Examples of these architectures include, but are not limited to, the Industry Standard Architecture (ISA) bus, the Micro Channel Architecture (MAC) bus, the Enhanced ISA bus, the Video Electronics Standards Association (VESA) local bus, and the Peripheral Component Interconnect (PCI) bus.

[0204] Computer device 12 typically includes a variety of computer system readable media. These media can be any available media that can be accessed by computer device 12, including volatile and non-volatile media, removable and non-removable media.

[0205] Memory 28 may include computer system readable media in the form of volatile memory, such as Random Access Memory (RAM) 30 and / or cache memory 32. Computer device 12 may further include other removable / non-removable, volatile / non-volatile computer system storage media. By way of example only, storage system 34 may be used to read and write non-removable, non-volatile magnetic media (…). Figure 8 Not shown; usually referred to as a "hard drive"). Although Figure 8 Not shown, a disk drive for reading and writing to a removable non-volatile disk (e.g., a "floppy disk") and an optical disc drive for reading and writing to a removable non-volatile optical disc (e.g., a compact disc read-only memory (CD-ROM), a digital video disc read-only memory (DVD-ROM), or other optical media) may be provided. In these cases, each drive may be connected to bus 18 via one or more data media interfaces. Memory 28 may include at least one program product having a set (e.g., at least one) of program modules configured to perform the functions of the embodiments of this application.

[0206] A program / utility 40 having a set (at least one) of program modules 42 may be stored, for example, in memory 28. Such program modules 42 include, but are not limited to, an operating system, one or more application programs, other program modules, and program data. Each or some combination of these examples may include an implementation of a network environment. Program modules 42 typically perform the functions and / or methods described in the embodiments of this application.

[0207] Computer device 12 can also communicate with one or more external devices 14 (e.g., keyboard, pointing device, display 24, etc.), and with one or more devices that enable a user to interact with computer device 12, and / or with any device that enables computer device 12 to communicate with one or more other computing devices (e.g., network card, modem, etc.). This communication can be performed via input / output (I / O) interface 22. Furthermore, computer device 12 can also communicate with one or more networks (e.g., local area network (LAN), wide area network (WAN), and / or public networks, such as the Internet) via network adapter 20. As shown, network adapter 20 communicates with other modules of computer device 12 via bus 18. It should be understood that, although not shown in the figures, other hardware and / or software modules can be used in conjunction with computer device 12, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems.

[0208] The processing unit 16 executes various functional applications and data processing by running programs stored in the system memory 28, such as implementing the methods mentioned in the foregoing embodiments.

[0209] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., refer to specific features, structures, materials, or characteristics described in connection with that embodiment or example, which are included in at least one embodiment or example of this application. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of different embodiments or examples.

[0210] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of that feature. In the description of this application, "multiple" means at least two, such as two, three, etc., unless otherwise explicitly specified.

[0211] Any process or method description in the flowchart or otherwise herein can be understood as representing a module, segment, or portion of code comprising one or more executable instructions for implementing custom logic functions or processes, and the scope of the preferred embodiments of this application includes additional implementations in which functions may be performed not in the order shown or discussed, including substantially simultaneously or in reverse order depending on the functions involved, as should be understood by those skilled in the art to which embodiments of this application pertain.

[0212] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (such as a computer-based system, a processor-included system, or other system that can fetch and execute instructions from, an instruction execution system, apparatus, or device). For the purposes of this specification, "computer-readable medium" can be any means that can contain, store, communicate, propagate, or transmit programs for use by, or in conjunction with, an instruction execution system, apparatus, or device. More specific examples (a non-exhaustive list) of computer-readable media include: an electrical connection having one or more wires (electronic device), a portable computer disk drive (magnetic device), random access memory (RAM), read-only memory (ROM), erasable and editable read-only memory (EPROM or flash memory), fiber optic devices, and portable optical disc read-only memory (CDROM). Alternatively, the computer-readable medium may be paper or other suitable media on which the program can be printed, since the program can be obtained electronically, for example, by optically scanning the paper or other medium, followed by editing, interpreting, or otherwise processing as necessary, and then stored in a computer memory.

[0213] It should be understood that various parts of this application can be implemented using hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented using software or firmware stored in memory and executed by a suitable instruction execution system. For example, if implemented in hardware as in another embodiment, it can be implemented using any one or a combination of the following techniques known in the art: discrete logic circuits having logic gates for implementing logical functions on data signals, application-specific integrated circuits (ASICs) having suitable combinational logic gates, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.

[0214] Those skilled in the art will understand that all or part of the steps of the methods in the above embodiments can be implemented by a program instructing related hardware. The program can be stored in a computer-readable storage medium, and when executed, the program includes one or a combination of the steps of the method embodiments.

[0215] Furthermore, the functional units in the various embodiments of this application can be integrated into a processing module, or each unit can exist physically separately, or two or more units can be integrated into a module. The integrated module can be implemented in hardware or as a software functional module. If the integrated module is implemented as a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium.

[0216] The storage medium mentioned above can be a read-only memory, a disk, or an optical disk, etc. Although embodiments of this application have been shown and described above, it is understood that the above embodiments are exemplary and should not be construed as limiting this application. Those skilled in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of this application.

Claims

1. A model training method, characterized in that, The method includes the following steps: Obtain training text, wherein the training text is annotated with summary information; The training text is encoded using an encoder in a summary generation model to obtain the encoded hidden state; Based on the encoded hidden layer state, the decoder in the summary generation model determines the first decoded hidden layer state of the output, including: generating a decoding vector corresponding to the first predicted word based on the encoded hidden layer state of each word segmentation in the training text and the attention weight of the first predicted word to each word segmentation word; obtaining the second decoded hidden layer state and the corresponding second predicted word of the most recent output; and using the decoder to determine the first decoded hidden layer state of the current output based on the decoding vector corresponding to the first predicted word, the second decoded hidden layer state, and the second predicted word. Determining the decoding probability of the first predicted word based on the first decoding hidden layer state includes: determining the generation probability of the first predicted word based on the first decoding hidden layer state; generating the initial copy probability of the first predicted word based on the attention weight of the first predicted word to each of the segmented words; determining the target copy probability corresponding to the first predicted word based on the initial copy probability of the first predicted word and the target copy probability corresponding to the most recently output second predicted word; and determining the decoding probability corresponding to the first predicted word based on the generation probability and the target copy probability corresponding to the first predicted word. In response to the first predicted word being included in the summary information, the summary generation model is trained to maximize the decoding probability.

2. The method according to claim 1, characterized in that, Determining the generation probability of the first predicted word based on the first decoded hidden layer state includes: The generation probability of the first predicted word is determined based on the decoding vector corresponding to the first predicted word and the first decoding hidden layer state.

3. The method according to claim 1, characterized in that, The step of determining the target replication probability corresponding to the first predicted word based on the initial replication probability of the first predicted word and the target replication probability corresponding to the most recently output second predicted word includes: Determine the correlation between the second predicted term and the first predicted term; The target replication probability corresponding to the first predicted word is determined based on the correlation degree, the initial replication probability of the first predicted word, and the target replication probability corresponding to the second predicted word.

4. The method according to claim 3, characterized in that, Determining the correlation between the second predicted word and the first predicted word includes: Determine the text length of the training text; Determine the relative position difference between the second predicted word and the first predicted word in the training text; The correlation between the second predicted word and the first predicted word is determined based on the text length, the relative position difference, and the first decoding hidden layer state.

5. The method according to any one of claims 1-4, characterized in that, The step of encoding the training text using the encoder in the summary generation model to obtain the encoded hidden state includes: The training text is segmented into words to obtain the segmented vocabulary; The encoder is used to encode each of the segmented words to obtain the encoding hidden state corresponding to each of the segmented words.

6. The method according to claim 5, characterized in that, The step of encoding each segmented word using the encoder to obtain the encoding hidden state corresponding to each segmented word includes: For any segmented word, the encoder is used to encode the segmented word based on the encoding hidden state of the adjacent words in the training text that precede the segmented word, thereby obtaining the encoding hidden state corresponding to the segmented word.

7. A text summarization method, characterized in that, The method includes: Get the input text; A trained summary generation model is used to predict the input text to obtain summary information corresponding to the input text; wherein the summary generation model is trained using the model training method described in any one of claims 1-6.

8. A model training device, characterized in that, include: The acquisition module is used to acquire training text, wherein the training text is annotated with summary information; The encoding module is used to encode the training text in the acquisition module using the encoder in the summary generation model to obtain the encoding hidden state. Specifically, it is used to: generate the decoding vector corresponding to the first predicted word based on the encoding hidden state of each word segment in the training text and the attention weight of the first predicted word to each word segment; obtain the second decoding hidden state and the corresponding second predicted word of the most recent output; and use the decoder to determine the first decoding hidden state of the current output based on the decoding vector corresponding to the first predicted word, the second decoding hidden state, and the second predicted word. The first determining module is used to determine the first decoding hidden state of the output using the decoder in the digest generation model based on the encoding hidden state in the encoding module. The second determining module is used to determine the decoding probability of the first predicted word based on the first decoding hidden layer state in the first determining module. The second determining module includes: a first determining unit, used to determine the generation probability of the first predicted word based on the first decoding hidden layer state; a generation module, used to generate an initial copy probability of the first predicted word based on the attention weight of the first predicted word to each segmented word; a second determining unit, used to determine the target copy probability corresponding to the first predicted word based on the initial copy probability of the first predicted word and the target copy probability corresponding to the most recently output second predicted word; and a third determining unit, used to determine the decoding probability corresponding to the first predicted word based on the generation probability and the target copy probability corresponding to the first predicted word. A training module is configured to train the summary generation model in response to the first predicted word being contained in the summary information in the second determining module, so as to maximize the decoding probability.

9. The apparatus according to claim 8, characterized in that, The first determining unit is specifically used for: The generation probability of the first predicted word is determined based on the decoding vector corresponding to the first predicted word and the first decoding hidden layer state.

10. The apparatus according to claim 8, characterized in that, The second determining unit is specifically used for: Determine the correlation between the second predicted term and the first predicted term; The target replication probability corresponding to the first predicted word is determined based on the correlation degree, the initial replication probability of the first predicted word, and the target replication probability corresponding to the second predicted word.

11. The apparatus according to claim 10, characterized in that, The second determining unit is specifically used for: Determine the text length of the training text; Determine the relative position difference between the second predicted word and the first predicted word in the training text; The correlation between the second predicted word and the first predicted word is determined based on the text length, the relative position difference, and the first decoding hidden layer state.

12. The apparatus according to any one of claims 8-11, characterized in that, The encoding module is specifically used for: The training text is segmented into words to obtain the segmented vocabulary; The encoder is used to encode each of the segmented words to obtain the encoding hidden state corresponding to each of the segmented words.

13. The apparatus according to claim 12, characterized in that, The encoding module is specifically used for: For any segmented word, the encoder is used to encode the segmented word based on the encoding hidden state of the adjacent words in the training text that precede the segmented word, thereby obtaining the encoding hidden state corresponding to the segmented word.

14. A text summarization generation apparatus, characterized in that, include: The acquisition module is used to acquire the input text; The prediction module is used to predict the input text using a trained summary generation model to obtain summary information corresponding to the input text; wherein the summary generation model is trained using the model training device as described in any one of claims 8-13.

15. A computer device, characterized in that, It includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, it implements the model training method as described in any one of claims 1-6, or the text summarization generation method as described in claim 7.

16. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When executed by a processor, the program implements the model training method as described in any one of claims 1-6, or the text summarization generation method as described in claim 7.

17. A computer program product, characterized in that, When the instructions in the computer program product are executed by the processor, the model training method as described in any one of claims 1-6 is executed, or the text summarization generation method as described in claim 7 is executed.

Citation Information

Patent Citations

  • A neural network question generation method based on answers and answer position information

    CN109684452A

  • Neural network automatic abstract model based on semantic alignment

    CN111552801A