Text Abstract Generation Method, Apparatus, Electronic Device, and Storage Medium

By performing sentence breaking and triplet extraction processing on historical reference articles, and using classification model to generate text summary, the problem of insufficient accuracy in text summary generation in the existing technology is solved, and more efficient text summary generation is achieved.

CN115357681BActive Publication Date: 2025-06-20PING AN TECH (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210824644.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-07-13
Publication Date
2025-06-20
Estimated Expiration
2042-07-13

AI Technical Summary

Technical Problem

The existing text summary generation methods have insufficient accuracy, resulting in insufficient focus on the summary content and lack of bias towards important content of the article.

Method used

By obtaining historical reference articles, perform sentence breaking processing, extracting and deduplication processing triples, labeling and training classification models, generating standard classification results, filtering and splicing triples, and entering a bidirectional long and short-term memory network to generate an article summary.

Benefits of technology

Improve the accuracy of text summary generation, so that the generated summary contains more accurately important information in the article.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115357681B_ABST
    Figure CN115357681B_ABST
Patent Text Reader

Abstract

The present invention relates to artificial intelligence and discloses a method for generating a text summary, including: performing sentence segmentation on historical reference articles to obtain a plurality of reference sentences; performing triple extraction and duplicate removal processing on the plurality of reference sentences to obtain a plurality of standard triples and performing label marking, using the data with label marking as a training data set, training a classification model using the training data set to obtain a standard classification model; inputting the article to be processed into the standard classification model to obtain a standard classification result; using the triples that meet the screening conditions in the standard classification result as target triples and performing triple splicing processing to obtain an input sequence, inputting the input sequence into a bidirectional long short-term memory network to obtain an article summary. In addition, the present invention also relates to blockchain technology, and the standard triples can be stored in the nodes of the blockchain. The present invention also proposes a text summary generating device, an electronic device, and a storage medium. The present invention can improve the accuracy of text summary generation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence, and in particular, to a method, apparatus, electronic device, and storage medium for generating text summaries. Background Art

[0002] Internet technology has made the collection and dissemination of information faster, enabling people to enter an era of information explosion. On the one hand, the rich and diverse information resources bring great convenience to people's lives, but the massive amount of information also brings great troubles to people. How to quickly obtain the information one wants from the trillions of information on the Internet has become a challenging task. Therefore, researching a text summarization method that can be used to extract key information from text can improve the information query efficiency and reading efficiency of users, facilitating people's work and life.

[0003] Existing methods for generating text summaries usually use a trained language model to generate summaries of text. This method usually encodes and then decodes the summary with the entire text as the input. Although summary information can be obtained, the content of the summary is not focused enough, and there is a lack of bias towards the important content of the article. As a result, the accuracy of text summary generation is not high enough. Summary of the Invention

[0004] The present invention provides a method, apparatus, electronic device, and storage medium for generating text summaries, and its main purpose is to improve the accuracy of text summary generation.

[0005] To achieve the above object, a method for generating a text summary provided by the present invention includes:

[0006] Obtain historical reference articles, perform sentence segmentation processing on the historical reference articles according to preset sentence segmentation rules to obtain a plurality of reference sentences;

[0007] Perform triple extraction and deduplication processing on the plurality of reference sentences to obtain a plurality of standard triples;

[0008] Perform label marking on the plurality of standard triples, and use the data with label marking as a training data set to train a preset classification model to obtain a standard classification model;

[0009] Obtain an article to be processed, input the article to be processed into the standard classification model to obtain a standard classification result;

[0010] Use the triples that meet the preset screening conditions in the standard classification result as target triples and perform triple splicing processing to obtain an input sequence, and input the input sequence into a preset bidirectional long short-term memory network to obtain an article summary.

[0011] Optionally, the tagging of the multiple standard triples includes:

[0012] Obtain the labeled triples in the labeled data summary, and compare the multiple standard triples with the labeled triples in the labeled data summary;

[0013] Assign positive labels to the triples in the multiple standard triples that are consistent with the labeled triples, and assign negative labels to the triples in the multiple standard triples that are inconsistent with the labeled triples;

[0014] Summarize the positive labels, negative labels, and the corresponding labeled triples to obtain the tagged data.

[0015] Optionally, the training of the preset classification model using the training data set to obtain the standard classification model includes:

[0016] Obtain a preset number of random vectors, and combine the random vectors with the training data in the training data set to obtain a training sequence;

[0017] Input the training sequence into the classification model to obtain an output vector;

[0018] Perform vector mapping processing on the output vector using a preset multi-layer perceptron to obtain a final mapped vector, and input the final mapped vector into an activation function to obtain an activation value;

[0019] Obtain a predicted classification result according to the activation value and a preset reference table, and compare the predicted classification result with the tagged labels in the training data set based on a preset cross-entropy loss function to obtain an error value;

[0020] Backpropagate the error value to train the classification model until the convergence condition is met to obtain the standard classification model.

[0021] Optionally, taking the triples that meet the preset screening conditions in the standard classification result as target triples and performing triple splicing processing to obtain an input sequence includes:

[0022] Identify the occurrence order of the multiple target triples in the article to be processed respectively, and sort the multiple target triples according to the occurrence order to obtain sorted triples;

[0023] Perform front-back splicing processing on the multiple entities and the relationships between the entities in the spliced triples to obtain an input sequence.

[0024] Optionally, inputting the input sequence into a preset bidirectional long short-term memory network to obtain an article summary includes:

[0025] Calculate the state value of the input sequence through the input gate in the bidirectional long short-term memory network;

[0026] Calculate the activation value of the input sequence through the forget gate in the bidirectional long short-term memory network;

[0027] Calculate the state update value of the input sequence according to the state value and the activation value;

[0028] Use the output gate in the bidirectional long short-term memory network to calculate the article abstract corresponding to the state update value.

[0029] Optionally, the triple extraction and deduplication processing of the multiple reference sentences to obtain multiple standard triples includes:

[0030] Input the multiple reference sentences into a preset triple template extraction tool to obtain multiple initial triples;

[0031] Determine whether the entities and the relationships between entities among the multiple initial triples are consistent, and merge the initial triples with consistent entities and relationships between entities to obtain multiple standard triples.

[0032] Optionally, the sentence segmentation processing of the historical reference article according to a preset sentence segmentation rule to obtain multiple reference sentences includes:

[0033] Extract multiple sentence segmentation punctuation marks from a preset punctuation mark library;

[0034] Respectively use the multiple sentence segmentation punctuation marks as division nodes to segment the historical reference article to obtain multiple reference sentences.

[0035] To solve the above problems, the present invention also provides a text summary generation device, and the device includes:

[0036] A triple generation module, configured to obtain a historical reference article, perform sentence segmentation processing on the historical reference article according to a preset sentence segmentation rule to obtain multiple reference sentences, perform triple extraction and deduplication processing on the multiple reference sentences to obtain multiple standard triples;

[0037] A model training module, configured to perform label marking on the multiple standard triples, use the data with label marking as a training data set, and train a preset classification model with the training data set to obtain a standard classification model;

[0038] A standard classification module, configured to obtain an article to be processed, input the article to be processed into the standard classification model to obtain a standard classification result;

[0039] The abstract extraction module is used to take the triples that meet the preset screening conditions in the standard classification result as target triples and perform triple splicing processing to obtain an input sequence, and input the input sequence into a preset bidirectional long short-term memory network to obtain an article abstract.

[0040] To solve the above problems, the present invention also provides an electronic device, which includes:

[0041] At least one processor; and,

[0042] A memory communicatively connected to the at least one processor; wherein,

[0043] The memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the above-mentioned text abstract generation method.

[0044] To solve the above problems, the present invention also provides a storage medium, in which at least one computer program is stored, and the at least one computer program is executed by a processor in an electronic device to implement the above-mentioned text abstract generation method.

[0045] In the embodiments of the present invention, by performing sentence segmentation on historical reference articles, a plurality of reference sentences are obtained, and triple extraction and deduplication processing are performed on the plurality of reference sentences, so that the obtained standard triples are more accurate. By performing label marking on the plurality of standard triples and using the data with label marking as a training data set, the standard classification model trained using the training data set is more accurate. Input the article to be processed into the standard classification model to obtain a standard classification result; take the triples that meet the preset screening conditions in the standard classification result as target triples and perform triple splicing processing to obtain an input sequence, and input the input sequence into a preset bidirectional long short-term memory network to obtain an article abstract. Since the input sequence contains important information in the article, the obtained article abstract has higher accuracy. Therefore, the text abstract generation method, device, electronic device and storage medium proposed by the present invention can solve the problem of low accuracy in text abstract generation. Description of the Drawings

[0046] Figure 1 It is a schematic flowchart of a text abstract generation method provided by an embodiment of the present invention;

[0047] Figure 2 For Figure 1 A detailed implementation flowchart of one of the steps;

[0048] Figure 3 For Figure 1Schematic diagram of the detailed implementation process of one of the steps;

[0049] Figure 4 is Figure 1 Schematic diagram of the detailed implementation process of one of the steps;

[0050] Figure 5 is Figure 1 Schematic diagram of the detailed implementation process of one of the steps;

[0051] Figure 6 is Figure 1 Schematic diagram of the detailed implementation process of one of the steps;

[0052] Figure 7 is Figure 1 Schematic diagram of the detailed implementation process of one of the steps;

[0053] Figure 8 Functional module diagram of the text summary generation device provided by an embodiment of the present invention;

[0054] Figure 9 Schematic diagram of the structure of an electronic device for implementing the text summary generation method provided by an embodiment of the present invention.

[0055] The realization, functional characteristics and advantages of the object of the present invention will be further described with reference to the embodiments and the accompanying drawings. Specific embodiments

[0056] It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0057] An embodiment of the present application provides a text summary generation method. The execution subject of the text summary generation method includes, but is not limited to, at least one of electronic devices such as a server, a terminal, etc. that can be configured to execute the method provided by the embodiment of the present application. In other words, the text summary generation method can be executed by software or hardware installed on a terminal device or a server device, and the software can be a blockchain platform. The server includes, but is not limited to: a single server, a server cluster, a cloud server or a cloud server cluster, etc. The server can be an independent server or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, Content Delivery Network (CDN), and big data and artificial intelligence platforms.

[0058] Refer to Figure 1As shown in the figure, it is a schematic flowchart of a text abstract generation method provided by an embodiment of the present invention. In this embodiment, the text abstract generation method includes the following steps S1 - S5:

[0059] S1. Obtain historical reference articles, and perform sentence segmentation processing on the historical reference articles according to preset sentence segmentation rules to obtain multiple reference sentences.

[0060] In the embodiment of the present invention, the historical reference article refers to an article for reference on any website or in a database. For example, it can be article A in the financial field, article B in the sports field, or article C in current affairs news.

[0061] Specifically, referring to Figure 2 As shown in the figure, the step of performing sentence segmentation processing on the historical reference article according to preset sentence segmentation rules to obtain multiple reference sentences includes the following steps S11 - S12:

[0062] S11. Extract multiple sentence segmentation punctuation marks from a preset punctuation mark library;

[0063] S12. Use the multiple sentence segmentation punctuation marks as division nodes respectively to perform sentence division on the historical reference article to obtain multiple reference sentences.

[0064] In detail, the punctuation mark library contains multiple different punctuation marks. For example, full stops, commas, colons, exclamation marks, semicolons, double quotation marks, single quotation marks, etc. Extract multiple sentence segmentation punctuation marks from the preset punctuation mark library. The multiple sentence segmentation punctuation marks in this solution can be full stops, exclamation marks, question marks, and semicolons. Search for multiple sentence segmentation punctuation marks in the historical reference article, and use the multiple sentence segmentation punctuation marks as division nodes to perform sentence division to obtain multiple reference sentences.

[0065] S2. Perform triple extraction and deduplication processing on the multiple reference sentences to obtain multiple standard triples.

[0066] In the embodiment of the present invention, referring to Figure 3 As shown in the figure, the step of performing triple extraction and deduplication processing on the multiple reference sentences to obtain multiple standard triples includes the following steps S21 - S22:

[0067] S21. Input the multiple reference sentences into a preset triple template extraction tool to obtain multiple initial triples;

[0068] S22. Determine whether the entities and the relationships between the entities among the multiple initial triples are consistent, and perform merging processing on the initial triples with consistent entities and relationships between the entities to obtain multiple standard triples.

[0069] In detail, the triple template extraction tools are Stanford CoreNLP OpenIE and UW OpenIE, where the defined triple template is <entity, relationship, entity>. Multiple reference sentences are input into these two triple template extraction tools, i.e., the tool code library. If the reference sentences cover entities and the relationships between entities, they will be extracted as initial triples.

[0070] Preferably, the main purpose of using two triple template extraction tools for extraction is to identify the triples in the multiple reference sentences as much as possible. At the same time, there will be many triples that are extracted repeatedly. Therefore, after extracting multiple initial triples, it is necessary to perform deduplication processing and merge the same initial triples. Among them, the initial triples with consistent entities and entity relationships are determined to be the same triples and merged.

[0071] S3. Label the plurality of standard triples, and use the labeled data as a training data set, and use the training data set to train a preset classification model to obtain a standard classification model.

[0072] In the embodiment of the present invention, refer to Figure 4 As shown, the labeling of the plurality of standard triples includes the following steps S31-S33:

[0073] S31, obtaining annotated triples in the annotated data summary, and performing triple comparison between the plurality of standard triples and the annotated triples in the annotated data summary;

[0074] S32, assigning positive labels to the triples in the plurality of standard triples that are consistent with the annotated triples, and assigning negative labels to the triples in the plurality of standard triples that are inconsistent with the annotated triples;

[0075] S33: Summarize the positive labels, the negative labels and the corresponding annotation triples to obtain label-marked data.

[0076] Specifically, the labeled triples in the labeled data summary are triple2, and the multiple standard triples extracted from Article A are triple1, triple2, and triple3. Then, the triple that is consistent with the labeled triple among the multiple standard triples is triple2. Therefore, a positive label is assigned to triple2, and the triples that are inconsistent with the labeled triple among the multiple standard triples are triple1 and triple3. Thus, negative labels are assigned to triple1 and triple3. The positive label, the negative label, and the corresponding labeled triples are aggregated to obtain the labeled data, which is [Article A + triple1: negative label], [Article A + triple2: positive label], and [Article A + triple3: negative label]. Then, the labeled data contains data with both positive and negative labels.

[0077] Specifically, referring to Figure 5 as shown, training the preset classification model using the training data set to obtain the standard classification model includes the following steps S301 - S305:

[0078] S301. Obtain a preset number of random vectors, and combine the random vectors with the training data in the training data set to obtain a training sequence;

[0079] S302. Input the training sequence into the classification model to obtain an output vector;

[0080] S303. Use a preset multi - layer perceptron to perform vector mapping processing on the output vector to obtain a final mapped vector, and input the final mapped vector into an activation function to obtain an activation value;

[0081] S304. Obtain a predicted classification result according to the activation value and a preset reference table, and compare the predicted classification result with the labels in the training data set based on a preset cross - entropy loss function to obtain an error value;

[0082] S305. Back - propagate the error value to train the classification model until the convergence condition is met to obtain the standard classification model.

[0083] Specifically, the preset number of random vectors are [CLS] and [SEP]. The random vectors are combined with the training data in the training dataset to obtain a training sequence as [CLS] Article 1 [SEP] triple1. The classification model is a BERT model. The preset Multilayer Perceptron (MLP) is a feedforward artificial neural network model that maps multiple input datasets to a single output dataset. The activation function is the sigmoid function, and the sigmoid function outputs values from 0 to 1.

[0084] S4. Obtain the article to be processed, and input the article to be processed into the standard classification model to obtain a standard classification result.

[0085] In the embodiment of the present invention, the article to be processed refers to any randomly selected article for which text summarization needs to be performed. Since the standard classification model is obtained through model training, the classification ability of the standard classification model is strong and the classification accuracy is high. Inputting the article to be processed into the standard classification model obtains the label situations corresponding to different triples in the article to be processed. Different labels represent different meanings and can be used as a reference for subsequent data processing.

[0086] S5. Use the triples in the standard classification result that meet the preset screening conditions as target triples and perform triple splicing processing to obtain an input sequence. Input the input sequence into a preset bidirectional long short-term memory network to obtain an article summary.

[0087] In the embodiment of the present invention, the triples in the standard classification result that meet the preset screening conditions are used as target triples, and the preset screening conditions can be a limitation on the labels.

[0088] Specifically, referring to Figure 6 As shown, using the triples in the standard classification result that meet the preset screening conditions as target triples and performing triple splicing processing to obtain an input sequence includes the following steps S51 - S52:

[0089] S51. Identify the occurrence order of multiple target triples in the article to be processed respectively, and sort the multiple target triples according to the occurrence order to obtain sorted triples;

[0090] S52. Perform front - back splicing processing on multiple entities and the relationships between entities in the spliced triples to obtain an input sequence.

[0091] Specifically, sort the multiple target triples in the order of appearance to obtain sorted triples as triple 1, triple 2, etc. Since triple 1 contains head entity 1, relation 1, and tail entity 1, and triple 2 contains head entity 2, relation 2, and tail entity 2, perform front-to-back splicing processing on the multiple entities and the relationships between the entities in the spliced triple to obtain an input sequence as head entity 1, relation 1, tail entity 1, head entity 1, relation 1, tail entity 1,....

[0092] Further, with reference to Figure 7 as shown, inputting the input sequence into a preset bidirectional long short-term memory network to obtain an article abstract includes the following steps S501 - S504:

[0093] S501. Calculate the state value of the input sequence through the input gate in the bidirectional long short-term memory network;

[0094] S502. Calculate the activation value of the input sequence through the forget gate in the bidirectional long short-term memory network;

[0095] S503. Calculate the state update value of the input sequence according to the state value and the activation value;

[0096] S504. Calculate the article abstract corresponding to the state update value using the output gate in the bidirectional long short-term memory network.

[0097] In an optional embodiment, the calculation method of the state value includes:

[0098]

[0099] where i t represents the state value, represents the bias of the cell unit in the input gate, w i represents the activation factor of the input gate, h t-1 represents the peak value of the input sequence at the t - 1 moment of the input gate, x t represents the input sequence at the t moment, b i represents the weight of the cell unit in the input gate.

[0100] In an optional embodiment, the calculation method of the activation value includes:

[0101]

[0102] where f t represents the activation value, represents the bias of the cell unit in the forget gate, w f represents the activation factor of the forget gate, represents the peak value of the input sequence at the t - 1 moment of the forget gate, xt denotes the input sequence input at time t, b f denotes the weight of the cell unit in the forget gate.

[0103] In an optional embodiment, the calculation method of the state update value includes:

[0104]

[0105] where c t denotes the state update value, h t-1 denotes the peak value of the input sequence at time t - 1 of the input gate, denotes the peak value of the input sequence at time t - 1 of the forget gate.

[0106] In an optional embodiment, calculating the article abstract corresponding to the state update value by using the output gate in the bidirectional long short-term memory network includes:

[0107] Calculating the article abstract by using the following formula:

[0108] o t = tanh(c t )

[0109] where o t denotes the article abstract, tanh denotes the activation function of the output gate, c t denotes the state update value.

[0110] Specifically, the LSTM network (Long Short-Term Memory) is a time-recurrent neural network, including: an input gate, a forget gate, and an output gate. Among them, a bidirectional long short-term memory network with two hidden layers and 150 hidden units is adopted in this solution.

[0111] Specifically, since the input sequence already contains the most important information and relationships in the article, it makes up for the missing important information in the general text summarization method, making the accuracy of text summarization higher.

[0112] In the embodiments of the present invention, by performing sentence segmentation on historical reference articles, a plurality of reference sentences are obtained, and triple extraction and deduplication processing are performed on the plurality of reference sentences, so that the obtained standard triples are more accurate. By performing label marking on the plurality of standard triples and using the labeled data as a training data set, the standard classification model trained using the training data set is more accurate. The article to be processed is input into the standard classification model to obtain a standard classification result; the triples that meet the preset screening conditions in the standard classification result are used as target triples and triple splicing processing is performed to obtain an input sequence, and the input sequence is input into a preset bidirectional long short-term memory network to obtain an article abstract. Since the input sequence contains important information in the article, the obtained article abstract has higher accuracy. Therefore, the text abstract generation method proposed by the present invention can solve the problem of low accuracy in text abstract generation.

[0113] As Figure 8 shown, it is a functional module diagram of a text abstract generation device provided by an embodiment of the present invention.

[0114] The text abstract generation device 100 of the present invention can be installed in an electronic device. According to the functions achieved, the text abstract generation device 100 may include a triple generation module 101, a model training module 102, a standard classification module 103, and an abstract extraction module 104. The modules of the present invention may also be referred to as units, which refer to a series of computer program segments that can be executed by a processor of an electronic device and can complete fixed functions, and are stored in the memory of the electronic device.

[0115] In this embodiment, the functions of each module / unit are as follows:

[0116] The triple generation module 101 is used to obtain a historical reference article, perform sentence segmentation on the historical reference article according to a preset sentence segmentation rule to obtain a plurality of reference sentences, and perform triple extraction and deduplication processing on the plurality of reference sentences to obtain a plurality of standard triples;

[0117] The model training module 102 is used to perform label marking on the plurality of standard triples, use the labeled data as a training data set, and train a preset classification model using the training data set to obtain a standard classification model;

[0118] The standard classification module 103 is used to obtain an article to be processed, input the article to be processed into the standard classification model, and obtain a standard classification result;

[0119] The abstract extraction module 104 is configured to use the triples that meet the preset screening conditions in the standard classification result as target triples and perform triple splicing processing to obtain an input sequence, and input the input sequence into a preset bidirectional long short-term memory network to obtain an article abstract.

[0120] Specifically, the specific implementation manners of the modules of the text abstract generation device 100 are as follows:

[0121] Step 1: Obtain a historical reference article, and perform sentence segmentation processing on the historical reference article according to a preset sentence segmentation rule to obtain a plurality of reference sentences.

[0122] In an embodiment of the present invention, the historical reference article refers to an article for reference on any website or in a database. For example, it can be article A in the financial field, article B in the sports field, or article C in current affairs news.

[0123] Specifically, the performing sentence segmentation processing on the historical reference article according to a preset sentence segmentation rule to obtain a plurality of reference sentences includes:

[0124] Extract a plurality of sentence segmentation punctuation marks from a preset punctuation mark library;

[0125] Respectively use the plurality of sentence segmentation punctuation marks as division nodes to perform sentence division on the historical reference article to obtain a plurality of reference sentences.

[0126] Specifically, the punctuation mark library contains a plurality of different punctuation marks. For example, a period, a comma, a colon, an exclamation mark, a semicolon, double quotation marks, single quotation marks, etc. Extract a plurality of sentence segmentation punctuation marks from the preset punctuation mark library. The plurality of sentence segmentation punctuation marks in this solution can be a period, an exclamation mark, a question mark, and a semicolon. Search for a plurality of sentence segmentation punctuation marks in the historical reference article, and use the plurality of sentence segmentation punctuation marks as division nodes to perform sentence division to obtain a plurality of reference sentences.

[0127] Step 2: Perform triple extraction and deduplication processing on the plurality of reference sentences to obtain a plurality of standard triples.

[0128] In an embodiment of the present invention, the performing triple extraction and deduplication processing on the plurality of reference sentences to obtain a plurality of standard triples includes:

[0129] Input the plurality of reference sentences into a preset triple template extraction tool to obtain a plurality of initial triples;

[0130] Determine whether the entities and the relationships between the entities among the plurality of initial triples are consistent, and perform merging processing on the initial triples with consistent entities and relationships between the entities to obtain a plurality of standard triples.

[0131] Specifically, the triple template extraction tools are Stanford CoreNLP OpenIE and UW OpenIE. Among them, the defined triple template is <entity, relationship, entity>. When multiple reference sentences are input into these two triple template extraction tools, that is, the tool code library, if the reference sentences cover entities and the relationships between entities, they will be extracted as initial triples.

[0132] Preferably, the main purpose of using two triple template extraction tools for extraction is to identify the triples in as many reference sentences as possible. At the same time, there will also be many cases where triples are extracted repeatedly. Therefore, after extracting multiple initial triples, duplicate removal processing is required, and the same initial triples are merged. Among them, the initial triples with the same entity and entity relationship are determined as the same triples and merged.

[0133] Step 3: Label multiple standard triples, and use the labeled data as a training dataset to train a preset classification model to obtain a standard classification model.

[0134] In the embodiment of the present invention, the labeling of multiple standard triples includes:

[0135] Obtain the labeled triples in the labeled data summary, and compare the multiple standard triples with the labeled triples in the labeled data summary;

[0136] Assign positive labels to the triples in the multiple standard triples that are consistent with the labeled triples, and assign negative labels to the triples in the multiple standard triples that are inconsistent with the labeled triples;

[0137] Summarize the positive labels, negative labels and the corresponding labeled triples to obtain the labeled data.

[0138] Specifically, the labeled triples in the labeled data summary are triple2, and the multiple standard triples extracted from Article A are triple1, triple2, and triple3. Then, the triple that is consistent with the labeled triple among the multiple standard triples is triple2. Therefore, a positive label is assigned to triple2, and the triples that are inconsistent with the labeled triple among the multiple standard triples are triple1 and triple3. Thus, negative labels are assigned to triple1 and triple3. The positive label, the negative label, and the corresponding labeled triples are aggregated to obtain the labeled data, namely [Article A + triple1: negative label], [Article A + triple2: positive label], and [Article A + triple3: negative label]. The labeled data contains data with both positive and negative labels.

[0139] Specifically, training the preset classification model using the training data set to obtain the standard classification model includes:

[0140] Obtain a preset number of random vectors, and combine the random vectors with the training data in the training data set to obtain a training sequence;

[0141] Input the training sequence into the classification model to obtain an output vector;

[0142] Use a preset multi-layer perceptron to perform vector mapping processing on the output vector to obtain a final mapped vector, and input the final mapped vector into an activation function to obtain an activation value;

[0143] Obtain a predicted classification result according to the activation value and a preset reference table, and compare the predicted classification result with the labels in the training data set based on a preset cross-entropy loss function to obtain an error value;

[0144] Backpropagate the error value to train the classification model until the convergence condition is met to obtain the standard classification model.

[0145] Specifically, the preset number of random vectors are [CLS] and [SEP]. Combining the random vectors with the training data in the training data set, the obtained training sequence is [CLS] Article 1 [SEP] triple1. The classification model is a BERT model. The preset multi-layer perceptron (MLP, Multilayer Perceptron) is a feedforward artificial neural network model that maps multiple input data sets to a single output data set. The activation function is a sigmoid function, and the sigmoid function outputs values from 0 to 1.

[0146] Step 4: Obtain the article to be processed, input the article to be processed into the standard classification model, and obtain the standard classification result.

[0147] In the embodiment of the present invention, the article to be processed refers to any randomly selected article for which text summarization needs to be performed. Since the standard classification model is obtained through model training, the classification ability of the standard classification model is relatively strong and the classification accuracy is relatively high. Inputting the article to be processed into the standard classification model, the label situations corresponding to different triples in the article to be processed are obtained. Different labels represent different meanings and can be used as a reference for subsequent data processing.

[0148] Step 5: Use the triples that meet the preset screening conditions in the standard classification result as target triples and perform triple splicing processing to obtain an input sequence. Input the input sequence into a preset bidirectional long short-term memory network to obtain an article summary.

[0149] In the embodiment of the present invention, use the triples that meet the preset screening conditions in the standard classification result as target triples. The preset screening conditions can be limitations on labels.

[0150] Specifically, using the triples that meet the preset screening conditions in the standard classification result as target triples and performing triple splicing processing to obtain an input sequence includes:

[0151] Identify the occurrence order of multiple target triples in the article to be processed respectively, sort the multiple target triples according to the occurrence order to obtain sorted triples;

[0152] Perform front-back splicing processing on multiple entities and the relationships between entities in the spliced triples to obtain an input sequence.

[0153] In detail, sorting the multiple target triples according to the occurrence order to obtain sorted triples as triple 1, triple 2... Since triple 1 includes head entity 1, relationship 1, and tail entity 1, and triple 2 includes head entity 2, relationship 2, and tail entity 2, perform front-back splicing processing on multiple entities and the relationships between entities in the spliced triples to obtain an input sequence as head entity 1, relationship 1, tail entity 1, head entity 1, relationship 1, tail entity 1...

[0154] Further, inputting the input sequence into a preset bidirectional long short-term memory network to obtain an article summary includes:

[0155] Calculate the state value of the input sequence through the input gate in the bidirectional long short-term memory network;

[0156] Calculate the activation value of the input sequence through the forget gate in the bidirectional long short-term memory network;

[0157] Calculate the state update value of the input sequence according to the state value and the activation value;

[0158] Calculate the article abstract corresponding to the state update value by using the output gate in the bidirectional long short-term memory network.

[0159] In an alternative embodiment, the calculation method of the state value includes:

[0160]

[0161] where i t represents the state value, represents the bias of the cell unit in the input gate, w i represents the activation factor of the input gate, h t-1 represents the peak value of the input sequence at the t-1 moment of the input gate, x t represents the input sequence at the t moment, b i represents the weight of the cell unit in the input gate.

[0162] In an alternative embodiment, the calculation method of the activation value includes:

[0163]

[0164] where f t represents the activation value, represents the bias of the cell unit in the forget gate, w f represents the activation factor of the forget gate, represents the peak value of the input sequence at the t-1 moment of the forget gate, x t represents the input sequence input at the t moment, b f represents the weight of the cell unit in the forget gate.

[0165] In an alternative embodiment, the calculation method of the state update value includes:

[0166]

[0167] where c t represents the state update value, h t-1 represents the peak value of the input sequence at the t-1 moment of the input gate, represents the peak value of the input sequence at the t-1 moment of the forget gate.

[0168] In an alternative embodiment, the calculation of the article abstract corresponding to the state update value by using the output gate in the bidirectional long short-term memory network includes:

[0169] Calculate the article abstract using the following formula:

[0170] o t = tanh(c t )

[0171] Wherein, o t represents the article abstract, tanh represents the activation function of the output gate, and c t represents the state update value.

[0172] Specifically, the LSTM network (Long Short-Term Memory) is a time-recurrent neural network, including: an input gate, a forget gate, and an output gate. Among them, in this solution, a bidirectional long short-term memory network with two hidden layers and 150 units is adopted.

[0173] Specifically, since the input sequence already contains the most important information and relationships in the article, it makes up for the omission of important information in the article in general text summarization methods, making the generated text summary more accurate.

[0174] In the embodiments of the present invention, by performing sentence segmentation on historical reference articles, a plurality of reference sentences are obtained, and triple extraction and deduplication processing are performed on the plurality of reference sentences, so that the obtained standard triples are more accurate. By performing label marking on the plurality of standard triples and using the labeled data as a training data set, the obtained standard classification model is more accurate. Input the article to be processed into the standard classification model to obtain a standard classification result; use the triples that meet the preset screening conditions in the standard classification result as target triples and perform triple splicing processing to obtain an input sequence, and input the input sequence into a preset bidirectional long short-term memory network to obtain an article abstract. Since the input sequence contains the important information in the article, the obtained article abstract is more accurate. Therefore, the text summary generation device proposed by the present invention can solve the problem of low accuracy in text summary generation.

[0175] As Figure 9 shown, it is a schematic structural diagram of an electronic device for implementing a text summary generation method provided by an embodiment of the present invention.

[0176] The electronic device 1 may include a processor 10, a memory 11, a communication bus 12, and a communication interface 13, and may also include a computer program stored in the memory 11 and executable on the processor 10, such as a text summary generation program.

[0177] Among them, in some embodiments, the processor 10 may be composed of an integrated circuit. For example, it may be composed of a single packaged integrated circuit, or may be composed of multiple packaged integrated circuits with the same or different functions, including a combination of one or more central processing units (CPUs), microprocessors, digital processing chips, graphics processors, and various control chips. The processor 10 is the control core (Control Unit) of the electronic device, connecting various components of the entire electronic device through various interfaces and lines, and by running or executing programs or modules stored in the memory 11 (such as executing a text summary generation program, etc.), and calling data stored in the memory 11, to perform various functions of the electronic device and process data.

[0178] The memory 11 includes at least one type of readable storage medium, and the readable storage medium includes flash memory, mobile hard disks, multimedia cards, card-type memories (such as SD or DX memories, etc.), magnetic memories, magnetic disks, optical discs, etc. In some embodiments, the memory 11 may be an internal storage unit of the electronic device, such as the mobile hard disk of the electronic device. In some other embodiments, the memory 11 may also be an external storage device of the electronic device, such as a plug-in mobile hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc. equipped on the electronic device. Further, the memory 11 may also include both an internal storage unit and an external storage device of the electronic device. The memory 11 can not only be used to store application software installed on the electronic device and various types of data, such as the code of a text summary generation program, etc., but can also be used to temporarily store data that has been output or will be output.

[0179] The communication bus 12 may be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This bus can be divided into an address bus, a data bus, a control bus, etc. The bus is configured to implement the connection and communication between the memory 11 and at least one processor 10, etc.

[0180] The communication interface 13 is used for communication between the above-mentioned electronic device and other devices, including a network interface and a user interface. Optionally, the network interface may include a wired interface and / or a wireless interface (such as a WI-FI interface, a Bluetooth interface, etc.), which is generally used to establish a communication connection between this electronic device and other electronic devices. The user interface may be a display, an input unit (such as a keyboard), and optionally, the user interface may also be a standard wired interface or a wireless interface. Optionally, in some embodiments, the display may be an LED display, a liquid crystal display, a touch liquid crystal display, and an OLED (Organic Light-Emitting Diode) toucher, etc. Among them, the display may also be appropriately referred to as a display screen or a display unit, which is used to display the information processed in the electronic device and to display a visual user interface.

[0181] Figure 9 Only the electronic device with components is shown. Those skilled in the art can understand that Figure 9 the shown structure does not constitute a limitation on the electronic device 1, and it may include fewer or more components than shown, or combine certain components, or have a different component arrangement.

[0182] For example, although not shown, the electronic device may further include a power source (such as a battery) for supplying power to each component. Preferably, the power source may be logically connected to the at least one processor 10 through a power management device, so as to implement functions such as charging management, discharging management, and power consumption management through the power management device. The power source may also include any components such as one or more DC or AC power sources, a recharge device, a power failure detection circuit, a power converter or an inverter, and a power status indicator. The electronic device may also include various sensors, a Bluetooth module, a Wi-Fi module, etc., which will not be elaborated here.

[0183] It should be understood that the above embodiments are only for illustration purposes and are not limited by this structure in the scope of the patent application.

[0184] The text summary generation program stored in the memory 11 in the electronic device 1 is a combination of multiple instructions. When running in the processor 10, it can implement:

[0185] Obtain historical reference articles, perform sentence segmentation processing on the historical reference articles according to preset sentence segmentation rules, and obtain a plurality of reference sentences;

[0186] Perform triple extraction and deduplication processing on the multiple reference sentences to obtain a plurality of standard triples;

[0187] Label multiple of the standard triples, and use the labeled data as a training data set to train a preset classification model to obtain a standard classification model;

[0188] Obtain an article to be processed, and input the article to be processed into the standard classification model to obtain a standard classification result;

[0189] Use the triples that meet the preset screening conditions in the standard classification result as target triples and perform triple splicing processing to obtain an input sequence, and input the input sequence into a preset bidirectional long short-term memory network to obtain an article abstract.

[0190] Specifically, for the specific implementation method of the above instructions by the processor 10, reference can be made to the description of the relevant steps in the corresponding embodiments of the attached drawings, which will not be elaborated here.

[0191] Further, if the modules / units integrated in the electronic device 1 are implemented in the form of software functional units and sold or used as independent products, they can be stored in a storage medium. The storage medium can be volatile or non-volatile. For example, the computer-readable medium may include: any entity or device capable of carrying the computer program code, a recording medium, a USB flash drive, a mobile hard disk, a magnetic disk, an optical disk, a computer memory, a read-only memory (ROM, Read-Only Memory).

[0192] The present invention also provides a storage medium. The readable storage medium stores a computer program, and when the computer program is executed by a processor of an electronic device, it can implement:

[0193] Obtain a historical reference article, and perform sentence segmentation processing on the historical reference article according to a preset sentence segmentation rule to obtain multiple reference sentences;

[0194] Perform triple extraction and deduplication processing on multiple of the reference sentences to obtain multiple standard triples;

[0195] Label multiple of the standard triples, and use the labeled data as a training data set to train a preset classification model to obtain a standard classification model;

[0196] Obtain an article to be processed, and input the article to be processed into the standard classification model to obtain a standard classification result;

[0197] Use the triples that meet the preset screening conditions in the standard classification result as target triples and perform triple splicing processing to obtain an input sequence, and input the input sequence into a preset bidirectional long short-term memory network to obtain an article abstract.

[0198] In several embodiments provided by the present invention, it should be understood that the disclosed devices, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative. For example, the division of the modules is only a logical function division, and there can be other division methods in actual implementation.

[0199] The modules described as separate components may or may not be physically separated, and the components shown as modules may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0200] In addition, the functional modules in various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above integrated units can be implemented in the form of hardware or in the form of a combination of hardware and software functional modules.

[0201] For those skilled in the art, it is obvious that the present invention is not limited to the details of the above exemplary embodiments, and without departing from the spirit or basic characteristics of the present invention, the present invention can be implemented in other specific forms.

[0202] Therefore, from any point of view, the embodiments should be regarded as exemplary and non-limiting. The scope of the present invention is defined by the appended claims rather than the above description. Therefore, all changes falling within the meaning and scope of the equivalent elements of the claims are intended to be included in the present invention. Any associated drawing marks in the claims should not be regarded as limiting the claims involved.

[0203] The blockchain referred to in the present invention is a new application mode of computer technologies such as distributed data storage, peer-to-peer transmission, consensus mechanism, and encryption algorithms. Blockchain, in essence, is a decentralized database, a string of data blocks generated by using cryptographic methods. Each data block contains information about a batch of network transactions, which is used to verify the validity (anti-counterfeiting) of the information and generate the next block. The blockchain can include a blockchain underlying platform, a platform product service layer, an application service layer, etc.

[0204] The embodiments of the present application can acquire and process relevant data based on artificial intelligence technology. Among them, artificial intelligence (AI) is to use a digital computer or a machine controlled by a digital computer to simulate, extend, and expand human intelligence, sense the environment, acquire knowledge, and use knowledge to obtain the best results of theory, method, technology, and application systems.

[0205] In addition, it is obvious that the term "comprising" does not exclude other units or steps, and the singular does not exclude the plural. A plurality of units or devices stated in the system claims can also be implemented by one unit or device through software or hardware. Terms such as first, second, etc. are used to denote names and do not denote any particular order.

[0206] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention rather than to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present invention.

Claims

1. A method for generating a text summary, characterized in that, The method includes: Obtain historical reference articles, perform sentence segmentation processing on the historical reference articles according to preset sentence segmentation rules to obtain a plurality of reference sentences; Perform triple extraction and deduplication processing on the plurality of reference sentences to obtain a plurality of standard triples; Perform label marking on the plurality of standard triples, and use the data with label marking as a training data set, and train a preset classification model using the training data set to obtain a standard classification model; Obtain an article to be processed, input the article to be processed into the standard classification model to obtain a standard classification result; Use the triples that meet the preset screening conditions in the standard classification result as target triples and perform triple splicing processing to obtain an input sequence, and input the input sequence into a preset bidirectional long short-term memory network to obtain an article abstract; Among them, the performing label marking on the plurality of standard triples includes: obtaining the labeled triples in the labeled data abstract, and performing triple comparison between the plurality of standard triples and the labeled triples in the labeled data abstract; assigning positive labels to the triples that are the same as the labeled triples among the plurality of standard triples, and assigning negative labels to the triples that are different from the labeled triples among the plurality of standard triples; summarizing the positive labels, the negative labels and the corresponding labeled triples to obtain the data with label marking; The training the preset classification model using the training data set to obtain a standard classification model includes: obtaining a preset number of random vectors, combining the random vectors with the training data in the training data set to obtain a training sequence; inputting the training sequence into the classification model to obtain an output vector; performing vector mapping processing on the output vector using a preset multi-layer perceptron to obtain a final mapped vector, and inputting the final mapped vector into an activation function to obtain an activation value; obtaining a predicted classification result according to the activation value and a preset reference table, and comparing the predicted classification result with the label marking in the training data set based on a preset cross-entropy loss function to obtain an error value; backpropagating the error value to train the classification model until the convergence condition is met to obtain a standard classification model.

2. The method for generating a text summary according to claim 1, characterized in that, The using the triples that meet the preset screening conditions in the standard classification result as target triples and performing triple splicing processing to obtain an input sequence includes: Identify the occurrence order of the plurality of target triples in the article to be processed respectively, and sort the plurality of target triples according to the occurrence order to obtain sorted triples; Perform front-back splicing processing on the multiple entities and the relationships between the entities in the spliced triples to obtain an input sequence.

3. The method for generating a text summary according to claim 1, characterized in that, The inputting the input sequence into a preset bidirectional long short-term memory network to obtain an article abstract includes: Calculating the state value of the input sequence through the input gate in the bidirectional long short-term memory network; Calculating the activation value of the input sequence through the forget gate in the bidirectional long short-term memory network; Calculating the state update value of the input sequence according to the state value and the activation value; Calculate the article abstract corresponding to the state update value by using the output gate in the bidirectional long short-term memory network.

4. The method for generating a text summary according to claim 1, characterized in that, Performing triple extraction and deduplication processing on multiple reference sentences to obtain multiple standard triples, including: Inputting multiple reference sentences into a preset triple template extraction tool to obtain multiple initial triples; Determine whether the entities and the relationships between entities among multiple initial triples are consistent, and merge the initial triples with consistent entities and relationships between entities to obtain multiple standard triples.

5. The method for generating a text summary according to any one of claims 1 to 4, characterized in that, Performing sentence segmentation processing on the historical reference article according to preset sentence segmentation rules, including: Extract multiple sentence segmentation punctuation marks from a preset punctuation mark library; Respectively use multiple sentence segmentation punctuation marks as division nodes to segment the historical reference article to obtain multiple reference sentences.

6. A text summary generation device for implementing the text summary generation method according to any one of claims 1 to 5, characterized in that, The device includes: A triple generation module, configured to obtain a historical reference article, perform sentence segmentation processing on the historical reference article according to preset sentence segmentation rules to obtain multiple reference sentences, perform triple extraction and deduplication processing on multiple reference sentences to obtain multiple standard triples; A model training module, configured to perform label marking on multiple standard triples, use the labeled data as a training data set, and train a preset classification model with the training data set to obtain a standard classification model; A standard classification module, configured to obtain an article to be processed, input the article to be processed into the standard classification model to obtain a standard classification result; An abstract extraction module, configured to use the triples that meet the preset screening conditions in the standard classification result as target triples and perform triple splicing processing to obtain an input sequence, and input the input sequence into a preset bidirectional long short-term memory network to obtain an article abstract.

7. An electronic device, characterized in that, The electronic device includes: At least one processor; and, A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the text abstract generation method according to any one of claims 1 to 5.

8. A storage medium stores a computer program, characterized in that, When the computer program is executed by the processor, it implements the text abstract generation method according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • Abstract generation method fusing key information

    CN113111663A

  • Cboth generation method and device, equipment and storage medium

    CN114386392A