Dialogue abstract generation method and system based on key information
By extracting key knowledge of dialogue text from discourse-level semantic information, and using natural language processing tools and pre-trained language models for processing and generation, the colloquial characteristics and personal reference problems in dialogue summary generation are solved, significantly improving the quality of the abstract.
Patent Information
- Application Number
- CN202311466394.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-07
- Publication Date
- 2025-05-09
AI Technical Summary
In the generation of dialogue summary, it is difficult for the existing technology to effectively deal with the colloquial characteristics, personal reference issues and the problem of changing topic styles of dialogue texts, resulting in low quality of abstracts.
By mining key knowledge inside dialogue text from discourse-level semantic information, using natural language processing tools and pre-trained language models, extracting and filtering key sentences and relationship triplets, performing reference dissolution and language style transformation, and generating a concise dialogue summary.
Without relying on external large models and additional manual processing, the quality of conversation summary is significantly improved, helping the model better capture important information in conversation text.
Smart Images

Figure CN119961446A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of natural language processing, and more particularly to a method and system for extracting internal knowledge of a conversation text from a discourse-level semantic layer to enhance a conversation summary. Background Art
[0002] The purpose of the dialogue summary task is to extract the most important information in the dialogue into a shorter text. It can help people quickly capture the important information in the dialogue text of multiple participants without reviewing the complex dialogue context. As a very important new task, dialogue summarization has great application potential in many scenarios, such as court debates in civil trials, service conversations between customer service and users, and telephone conferences with multiple participants. However, due to the particularity of dialogue texts, there are more colloquial content in dialogue texts, the language style is inconsistent with the summary text, there are also many person reference issues in the dialogue, the topic style in the dialogue changes quickly, and the key information is relatively scattered. These characteristics have brought great challenges to the research of dialogue summarization.
[0003] In order to solve the above-mentioned problems in the generation of dialogue summaries, in recent years, there have been some studies on dialogue summary tasks, such as converting the overall language style of the dialogue text to replace the dialogue text input, or using external large models to obtain external knowledge, or using word-level structured knowledge to perform reference resolution, etc. Although the improvements in these methods do have some improvements in the quality of the summary, the above methods focus on improvements from the paragraph level or character level semantics, and some models rely on external large models, such as COMET, InstructGPT, or require a lot of manual processing. Compared with the above research, the present invention starts from the semantics of the discourse level and extracts internal key knowledge from the dialogue text in a lightweight manner, helping the model generate higher quality summaries without relying on external large models and additional manual processing. Summary of the invention
[0004] The purpose of the present invention is to provide a generative dialogue summarization system that enhances the quality of generated summaries by mining key knowledge within dialogue texts from discourse-level semantic information. The system aims to help the model learn the ability to utilize key knowledge by utilizing key knowledge in dialogue datasets, encode the input text based on the encoding module of the autoregressive pre-trained language model, and use the model's decoder to generate concise and accurate dialogue summaries based on the dialogue text and key knowledge data, so as to improve the machine's ability to generate dialogue texts.
[0005] The specific technical solution for achieving the purpose of the present invention is:
[0006] A generative dialogue summary system that enhances the quality of generated summaries by mining key knowledge inside dialogue texts, uses natural language processing tools to simultaneously extract key knowledge from dialogue datasets and reference summaries, and uses a pre-trained language model to calculate the semantic similarity at the discourse level to screen the key knowledge extracted from the dialogue dataset, and uses an autoregressive pre-trained language model to generate concise and accurate dialogue summaries. The system includes a dialogue data collection and data preprocessing module, a key knowledge extraction module, a key knowledge screening module, a key knowledge processing and dialogue content splicing module, a dialogue context representation module, and a summary generation module, wherein:
[0007] The dialogue data acquisition and preprocessing module is used to obtain the SAMSum data set and DialogSum data resources respectively, and perform format normalization processing on the dialogue data to adapt to the present summary system;
[0008] The key knowledge extraction module is used to extract key knowledge from the dialogue text and the reference summary. The key knowledge to be extracted includes two types: key sentences and relation triples. The key knowledge is used to help the summary generation module learn the ability to identify key knowledge of the dialogue during the training process;
[0009] The key knowledge screening module is used to screen the extracted key knowledge and select the key knowledge with the most appropriate semantic similarity to the reference summary;
[0010] The key knowledge processing module is used to process the above-screened key knowledge, convert the language style of the screened key sentence information into the third-person form, and perform first-person and second-person reference resolution on the above-screened relation triple information.
[0011] The key knowledge and dialogue content splicing module is used to combine the dialogue dataset and key knowledge, and convert the normalized input into an input form that matches the pre-trained language model training;
[0012] The dialogue context representation module is used to input the constructed content into BART-Large for encoding, and obtain the dialogue context representation using the above-mentioned pre-trained language model;
[0013] The summary generation module is used to effectively generate the conversation summary, update the network parameters through fine-tuning technology to obtain the optimal checkpoints, and generate the final summary using the checkpoints.
[0014] The SAMSum dataset and DialogSum dataset are public and open access research data resources, obtained from open source resources;
[0015] Specifically, the key knowledge extraction module extracts relation triples by using the OpenIE tool to extract relation triples from the dialogue text and the reference summary, and uses the sentence segmentation tool to segment the discourse in the dialogue text and the reference summary; the relation triples can be expressed as (Subject, Relation, Object) or (Subject, Prodicate, Object), where Subject and Object are two entities, respectively called the head entity (Head Entity) and the tail entity (Tail Entity), and Relation represents the relationship category;
[0016] The key knowledge screening module uses the pre-trained model Sentence-Bert to calculate the semantic similarity between the key knowledge extracted from the dialogue text and the reference summary. Specifically, the semantic similarity between the key sentence information and the flattened relationship triple information is input into the Sentence-Bert model. If the similarity is higher than the preset threshold, the key information is regarded as the key information of the final dialogue text. Among them, the key sentence information can be directly used to calculate the semantic similarity, while the relationship triple needs to be flattened before the semantic similarity is calculated. The relationship triple needs to be flattened before the semantic similarity is calculated. For example: the relationship triple is (zhangsan, want, go toBeijing) and the sentence after flattening is "zhangsan want go to Beijing".
[0017] The key knowledge processing module is used to process the selected key knowledge. Specifically, the module uses a natural language processing tool (such as spaCy) to perform named entity recognition on sentences of relation triples. By identifying the first-person and second-person pronouns in the sentences, the module replaces these pronouns with corresponding actual entities. Through such a reference resolution operation, the processed key information is obtained.
[0018] The key sentence information converts the language style from the third person to the first person as follows: for a key sentence information in the dialogue text, if it ends with a period, add "said" after the speaker in the key sentence, and put the speech content in the key sentence in quotation marks after "said"; for a key sentence information in the dialogue text that ends with a question mark, add "asked" after the speaker in the key sentence, and put the speech content in the key sentence in quotation marks after "asked";
[0019] The coreference resolution uses the natural language processing tool spaCy to perform named entity recognition on the sentence consisting of the selected relation triples, and recognizes the first-person and second-person pronouns in the sentence;
[0020] The first-person pronoun is replaced according to the category as follows: if the first-person pronoun in the sentence is "me", "I", or "myself", the first-person pronoun is replaced with the current relation triple for the speaker in the discourse; if the first-person pronoun in the sentence is "my", the first-person pronoun is replaced with the current relation triple for the speaker in the discourse, and "'s" is added; if the first-person pronoun in the sentence is "we" and "us", the first-person pronoun is replaced with the current relation triple for the speaker in the discourse and the previous speaker; if the first-person pronoun in the sentence is "our", the first-person pronoun is replaced with the current relation triple for the speaker in the discourse and the previous speaker, and "'s" is added;
[0021] The second person pronoun is replaced according to the category as follows: if the second person pronoun in the sentence is "you" or "u", the second person pronoun is replaced by the speaker of the previous round of the utterance corresponding to the current relation triple; if the second person pronoun in the sentence is "your" or "ur", the second person pronoun is replaced by the speaker of the previous round of the utterance corresponding to the current relation triple, and "'s" is added; thereby obtaining the processed key information;
[0022] The key knowledge and dialogue content splicing module splices the processed key knowledge with the dialogue text content, and uses a special token to place it at the end of the key information, so as to facilitate the dialogue summary model to distinguish and learn during training;
[0023] The conversation context representation module is used to input the constructed conversation content into the BART-Large pre-trained language model to obtain a context representation vector;
[0024] The summary generation module processes the SAMSum dataset and the DialogSum dataset of the dialogue context representation module and sends them to the BART-Large model for fine-tuning. Through reasonable hyperparameter settings and iterative optimization strategies, the validation set data is used to select the optimal network checkpoints, and the final summary of the test set is generated on the pre-trained language model BART-Large using the checkpoints.
[0025] The present invention has demonstrated superior characteristics on public datasets such as SAMSum and DialogSum, verifying that the system has good conversation summary generation capabilities; this also shows that by utilizing the key knowledge in the conversation content, it can help the summary model utilize the key knowledge in the conversation content during the training phase and improve the ability to identify key knowledge. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] Figure 1 This is the key sentence key information extraction and screening method described in the present invention.
[0027] Figure 2 This is the key knowledge extraction and screening method of relation triples described in the present invention.
[0028] Figure 3 It is the model method structure for calculating key relationship triples described in the present invention.
[0029] Figure 4 This is the first-person and second-person pronoun replacement rule for the relation triples described in the present invention.
[0030] Figure 5 This is the overall structure of the model of the present invention. DETAILED DESCRIPTION
[0031] The present invention provides a generative dialogue summarization system that enhances the quality of generated summaries by mining key knowledge within dialogue text from discourse-level semantic information. The system aims to help the model learn the ability to utilize key knowledge by utilizing key knowledge in dialogue datasets. The system is further described in detail below in conjunction with the accompanying drawings and specific embodiments.
[0032] Reference Figure 1 , extracting key sentence information from the conversation text and the reference summary:
[0033] Firstly, the dialogue text and the reference summary text are divided into sentences according to the dialogue turns to obtain the dialogue text and the reference summary sentence set;
[0034] The sentence set in the dialogue text and the summary sentence set are respectively used as inputs of Sentence-Bert to calculate semantic similarity, thereby obtaining a semantic similarity score matrix;
[0035] The sentence with the highest semantic similarity score in the sentence set in the dialogue text, that is, the sentence in the dialogue text corresponding to the sentence pair with the highest score in the similarity score matrix, is used as the key sentence;
[0036] Reference Figure 2 , extracting relation triple information from the conversation text and the reference summary:
[0037] Firstly, the dialogue text and the reference summary text are divided into sentences according to the dialogue turns to obtain the dialogue text and the reference summary sentence set;
[0038] Using OpenIE to extract relation triples from the conversation text and the reference summary text;
[0039] Performing noise removal and screening processing on the extracted relationship triples;
[0040] Flattening the relation triples into sentence form;
[0041] The processed triple sentence set of the dialogue text and the processed triple sentence set of the reference summary dialogue text are used as Sentence-Bert inputs, and the semantic similarity is calculated to obtain a semantic similarity score matrix;
[0042] The sentences in the sentence set after the relation triples in the conversation text are processed and have a semantic similarity greater than a threshold of 0.6 are used as the relation triple key information;
[0043] Reference Figure 3 , which is the model structure of the Sentence-Bert model used in the method of calculating the semantic similarity matrix:
[0044] For sentences 1 and 2, first use them as the input of the pre-trained model BERT to obtain the vector representation of sentences 1 and 2;
[0045] The vector representations of sentences 1 and 2 are then processed through a parameter-sharing pooling layer to obtain two vector representations u and v respectively;
[0046] The cosine similarity of the two vectors u and v is calculated to obtain the cosine similarity score of sentence 1 and sentence 2, and the score range is between -1 and 1.
[0047] Reference Figure 4 , for the detailed first-person and second-person pronoun replacement rules when the key relationship triple is processed in the module, refer to Figure 4 , the specific rules are as follows:
[0048] First, the natural language processing tool spaCy is used to perform named entity recognition on the sentences consisting of the selected relation triples, and the first-person and second-person pronouns in the sentences are identified;
[0049] Replace the first-person pronoun according to the category; if the first-person pronoun in the sentence is "me", "I", or "myself", replace the first-person pronoun with the current relation triple for the speaker in the discourse; if the first-person pronoun in the sentence is "my", replace the first-person pronoun with the current relation triple for the speaker in the discourse, and add "'s"; if the first-person pronoun in the sentence is "we" and "us", replace the first-person pronoun with the current relation triple for the speaker in the discourse and the previous speaker; if the first-person pronoun in the sentence is "our", replace the first-person pronoun with the current relation triple for the speaker in the discourse and the previous speaker, and add "'s";
[0050] Replace the second-person pronoun according to the category; if the second-person pronoun in the sentence is "you" or "u", replace the second-person pronoun with the speaker of the previous round of the utterance corresponding to the current relation triple; if the second-person pronoun in the sentence is "your" or "ur", replace the second-person pronoun with the speaker of the previous round of the utterance corresponding to the current relation triple, and add "'s";
[0051] Thus, the processed key information is obtained; wherein, if the current relation triple corresponds to the first round of the current dialogue text, the speaker of the previous round is changed to the speaker of the next round; through the above first-person and second-person pronoun replacement rules, the processed key information is obtained;
[0052] Reference Figure 5 , which is the overall structure of the model. The extracted key information is injected into the pre-trained language model BART-Large model during the training phase to help the model identify key knowledge during the training phase, learn to capture the key knowledge inside the dialogue text, and use the key knowledge to help the model generate higher quality summaries during the training phase. In the reasoning phase, the trained BART-Large model is used to enable the model to capture the key knowledge in the dialogue text with higher quality and generate higher quality summaries when performing new reasoning tasks.
Claims
1. A method for generating a conversation summary, characterized in that: The method for generating a dialogue summary includes: obtaining a dialogue text, and dividing the dialogue text into corresponding dialogue discourse sets; Extracting key knowledge from the conversation text content and reference summary; Screening the extracted key knowledge, that is, calculating the semantic similarity between the extracted key knowledge and the key knowledge extracted from the reference summary, and screening out the key knowledge whose semantic similarity is greater than a preset threshold; Processing the selected key knowledge, including converting the key information of key sentences from the first person to the third person language style, and performing first-person to second-person reference resolution on the key information of relation triples; Splice the processed key knowledge with the dialogue text; The conversation text after concatenating the key knowledge is used as the input of the model, and the pre-trained language model BART-Large is fine-tuned to generate the target conversation summary.
2. The method for generating a conversation summary according to claim 1, characterized in that: In the step of extracting key knowledge from the dialogue text and the reference summary, the method further comprises: The dialogue text and the reference summary text are divided and processed according to discourse; Calculating semantic similarity scores between a set of utterances in the conversation text and a set of utterances in the reference summary using Transformer-based Sentence-Bert to obtain a calculated score matrix; Selecting the utterance with the highest semantic similarity score in the score matrix from the utterance set in the dialogue text as the key sentence; Using OpenIE to extract relation triples from the conversation text and the reference summary text; Performing noise removal and screening processing on the extracted relationship triples; Flattening the relation triples into sentence form; Using the Sentence-Bert model, calculating the semantic similarity between the processed triple sentence set in the dialogue text and the processed sentence set in the reference summary, to obtain a calculated score matrix; The sentences in the sentence set after the relation triples in the dialogue text are processed and have a semantic similarity greater than a threshold of 0.6 are used as relation triple key information.
3. The method for generating a conversation summary according to claim 1, wherein: After processing the key sentence information extracted from the dialogue text, the key sentence information is used as the key sentence information of the dialogue text, and the method further includes: Processing the filtered key sentence information; For the selected key sentence information, the language style is changed from the third person to the first person; For a key sentence in the dialogue text, if it ends with a period, add "said" after the speaker in the key sentence, and put the speech content in the key sentence in quotation marks after "said"; When a key sentence in the dialogue text ends with a question mark, "asked" is added after the speaker in the key sentence, and the speech content of the key sentence is enclosed in quotation marks and placed after "asked".
4. The method for generating a conversation summary as claimed in claim 3, characterized in that: The filtered relation triple key information is subjected to first-person and second-person reference resolution as the relation triple key information of the dialogue text, and the method further includes: Using the natural language processing tool spaCy to perform named entity recognition on the flattened sentences of the selected relation triples, and identifying the first-person and second-person pronouns in the sentences; Replace the first-person pronouns according to their categories; If the first-person pronoun in the sentence is "me", "I", or "myself", replace the first-person pronoun with the current relation triple for the speaker in the utterance; If the first-person pronoun in the sentence is "my", replace the first-person pronoun with the current relation triple for the speaker in the utterance, and add "'s"; If the first-person pronouns in the sentence are "we" and "us", the first-person pronouns are replaced by the current relation triple for the speaker in the utterance and the speaker in the previous turn; The first person pronoun in the sentence is "our", and the first person pronoun is replaced by the current relation triple for the speaker in the utterance and the speaker in the previous turn, and "'s" is added; Replace the second-person pronoun according to category; If the second-person pronoun in the sentence is "you" or "u", the second-person pronoun is replaced by the speaker of the previous round of the utterance corresponding to the current relation triple; If the second-person pronoun in the sentence is "your" and "ur", replace the second-person pronoun with the speaker of the previous round of the utterance corresponding to the current relation triple, and add "'s"; Thus, the processed key information is obtained; wherein, if the current relation triplet corresponds to the first round of the current dialogue text, the speaker of the previous round is changed to the speaker of the next round.
5. A method for generating a conversation summary according to claim 1, characterized in that: The specific steps of generating the summary include: Use the pre-trained language model BART-Large as the dialogue encoder and decoder; The decoder receives the filtered conversation information and key knowledge, and uses the key knowledge to guide the generation of summary; The output of the decoder is a candidate summary, and Beam Search is used to select the best summary from multiple candidate summaries as the generated summary.
6. The system for generating a conversation summary is characterized in that: The system comprises: The key knowledge extraction module extracts key sentences and relation triples from the dialogue text and reference summary respectively; A key knowledge screening module is used to screen the extracted key knowledge based on semantic similarity to obtain the most semantically important key knowledge; The key knowledge processing module processes the key knowledge respectively, transforms the language style from the first person to the third person for the key sentence information, and performs the first-person and second-person reference resolution for the flattened sentences for the relation triple information; The key knowledge and dialogue text splicing module inserts special symbols between input text sequences and changes the format of dialogue sequences; The dialogue context representation module uses the pre-trained model BART-Large to encode the input text and obtain the vector of the dialogue sentence; The summary generation module generates a dialogue summary based on the dialogue text and key knowledge.