Conference summary generation method and device, electronic equipment and storage medium

By using the summary generation model in the conference summary generation method, processing meeting records and calculating segmented scores, the problem of insufficient accuracy, completeness and coherence of meeting summary content is solved, and a higher quality summary generation is achieved.

CN119938904APending Publication Date: 2025-05-06CHINA TELECOM CORP LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411814826.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-10
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

The generation of conference summary has problems with insufficient content accuracy, completeness and coherence.

Method used

A meeting summary generation method is adopted to obtain the conference record and summary generation model, including an encoder, decoder and output layer, perform content processing, calculate the segmentation score of text segments, extract the target text segments, perform encoding and decoding processing, and finally output the conference summary.

Benefits of technology

Improves the accuracy, completeness and coherence of generated conference summary, ensuring the quality of the summary through the integration of segmented scores and knowledge triplets.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119938904A_ABST
    Figure CN119938904A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a conference summary generation method and device, electronic equipment and a storage medium, and relates to the technical field of artificial intelligence, and the method comprises the steps: obtaining a conference record and a summary generation model, the summary generation model comprises an encoder, a decoder and an output layer; in response to a query instruction for the conference record, determining query information corresponding to the query instruction, and inputting the conference record into an abstract generation model to perform content processing on the conference record to obtain a text segment corresponding to the query information and a target knowledge triple corresponding to the text segment; calculating segmentation scores corresponding to the text segments, and based on the segmentation scores, extracting a plurality of target text segments of which the scores are ranked in the front from the text segments; inputting the target text segments into an encoder for encoding, and outputting corresponding encoding features; and inputting the query information, the coding feature and the target knowledge triple into a decoder for decoding processing, and outputting a conference abstract corresponding to the query information through an output layer.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a method for generating a conference summary, a device for generating a conference summary, an electronic device, and a computer-readable storage medium. Background Art

[0002] With the rapid advancement of Internet technology, the surge in network data has made it difficult for people to efficiently obtain information resources. Therefore, automated text summarization technology has gradually attracted attention. Using natural language processing technology, people can more effectively extract important information from massive data, and achieve content compression and information filtering through automatic summarization to meet users' different information needs. This technological advancement not only improves information processing efficiency, but also brings users a more convenient and personalized information acquisition experience.

[0003] The significance of internal meetings is to provide a platform for team members to exchange ideas, discuss issues, develop strategies, and ensure that team members are aligned on company goals and work priorities. Meetings can enhance team cohesion, promote information sharing and communication, and drive work efficiency and overall business development.

[0004] However, automatically generating conference summaries is not an easy task for computers. It faces several key challenges: First, data resources are scarce. Obtaining sufficient and high-quality conference data is key. Second, effective conference data modeling is required to deeply understand conference texts and generate accurate and rich summaries. Third, conference summaries need to address challenges in specific fields, because different conference data have their own forms and characteristics. How to effectively capture these characteristics is crucial to generating high-quality summaries. Summary of the invention

[0005] The embodiments of the present invention provide a method, device, electronic device and computer-readable storage medium for generating a conference summary, so as to solve or partially solve the problem of insufficient content accuracy, completeness and coherence in the generation of conference summaries.

[0006] The embodiment of the present invention discloses a method for generating a conference summary, comprising:

[0007] Obtaining a meeting record and a summary generation model for generating a meeting summary, wherein the summary generation model includes at least an encoder, a decoder, and an output layer;

[0008] In response to a query instruction for the conference record, query information corresponding to the query instruction is determined, and the conference record is input into the summary generation model to perform content processing on the conference record to obtain a text segment corresponding to the query information and a target knowledge triple corresponding to the text segment;

[0009] Calculating the segmentation scores corresponding to each of the text segments, and extracting a plurality of target text segments ranked first in score from the text segments based on the segmentation scores;

[0010] Input the target text into the encoder for encoding, and output corresponding encoding features;

[0011] The query information, the encoding feature and the target knowledge triple are input into the decoder for decoding processing, and the conference summary corresponding to the query information is output through the output layer.

[0012] In some feasible implementations, the summary generation model further includes a knowledge-aware summarizer, and the step of inputting the meeting minutes into the summary generation model to perform content processing on the meeting minutes to obtain text segments corresponding to the query information and target knowledge triples corresponding to the text segments includes:

[0013] Inputting the meeting record into the knowledge-aware summarizer for segmentation to obtain a plurality of text segments;

[0014] Extracting information from the text segment to obtain knowledge triples corresponding to the text segment;

[0015] The knowledge triples are filtered to extract target knowledge triples that overlap with the query information once.

[0016] In some feasible implementations, the calculating the segmentation scores corresponding to the respective text segments includes:

[0017] Using the query information and the text segment to perform calculations to obtain a semantic search score corresponding to the text segment;

[0018] Obtaining the total number of paragraphs corresponding to the text segment, performing normalization calculation using the total number of paragraphs and the number of target knowledge triples, and obtaining a knowledge perception score corresponding to the text segment;

[0019] The sum of the semantic search score and the knowledge-aware score is used as the segmentation score of the text segment.

[0020] In some feasible implementations, the using the query information and the text segment to perform calculations to obtain a semantic search score corresponding to the text segment includes:

[0021] Calculate the semantic search score corresponding to each of the text segments based on cosine similarity:

[0022]

[0023] Among them, Score se Represents query information Q and text segment S i The similarity score between them, MPNet(Q) represents the encoding result of the multimodal network MPNet for the query information Q, that is, the vector representation of Q, MPNet(S i ) represents the multimodal network MPNet for text segmentation S i The encoding result, namely S i The vector representation of .

[0024] In some feasible implementations, extracting a plurality of target text segments ranked first in score from the text segments based on the segmentation scores includes:

[0025] Sorting the text segments according to the values ​​of the segmentation scores, and extracting top k text segments from the sorted text segments as target text segments;

[0026] Wherein, k is a positive integer.

[0027] In some feasible implementations, the encoder includes at least a vector mapping layer, a multi-head attention layer, and a residual connection and normalization layer, and the target text is segmented and input into the encoder for encoding, and the corresponding encoding features are output, including:

[0028] Inputting the target text segment into the encoder, extracting words from the target text segment through the vector mapping layer, and obtaining word vectors corresponding to the target text segment;

[0029] Adding position coding through the vector mapping layer to obtain a corresponding position coding vector, and adding the word vector to the position coding vector to obtain an input sequence corresponding to the target text segment;

[0030] Input the input sequence into the multi-head attention layer for encoding, and calculate the attention information corresponding to each word in the input sequence;

[0031] The attention information is input into the residual connection and normalization layer for processing, and the encoding features corresponding to the target segment record are output.

[0032] In some feasible implementations, the step of inputting the query information, the encoding feature, and the target knowledge triple into the decoder for decoding, and outputting the meeting summary corresponding to the query information through the output layer includes:

[0033] removing stop words in the target knowledge triples and merging the remaining words into knowledge terms;

[0034] Concatenate the query information, the coding feature and the knowledge term to obtain a corresponding concatenation result;

[0035] The splicing result is input into the decoder for decoding processing, and the conference summary corresponding to the query information is output through the output layer.

[0036] The embodiment of the present invention also discloses a device for generating a conference summary, including:

[0037] A record acquisition module, used to acquire meeting records, and a summary generation model for generating a meeting summary, wherein the summary generation model includes at least an encoder, a decoder, and an output layer;

[0038] A record processing module, for responding to a query instruction for the conference record, determining query information corresponding to the query instruction, and inputting the conference record into the summary generation model to perform content processing on the conference record, and obtaining a text segment corresponding to the query information and a target knowledge triple corresponding to the text segment;

[0039] A score calculation module, used to calculate the segmentation scores corresponding to each of the text segments, and extract a plurality of target text segments ranked first in the score order from the text segments based on the segmentation scores;

[0040] An encoding module, used for inputting the target text into the encoder for encoding in segments, and outputting corresponding encoding features;

[0041] The summary generation module is used to input the query information, the encoding features and the target knowledge triples into the decoder for decoding processing, and output the meeting summary corresponding to the query information through the output layer.

[0042] In some feasible implementations, the summary generation model further includes a knowledge-aware summarizer, and the record processing module is specifically used to:

[0043] Inputting the meeting record into the knowledge-aware summarizer for segmentation to obtain a plurality of text segments;

[0044] Extracting information from the text segment to obtain knowledge triples corresponding to the text segment;

[0045] The knowledge triples are filtered to extract target knowledge triples that overlap with the query information once.

[0046] In some feasible implementations, the score calculation module is specifically used to:

[0047] Using the query information and the text segment to perform calculations to obtain a semantic search score corresponding to the text segment;

[0048] Obtaining the total number of paragraphs corresponding to the text segment, performing normalization calculation using the total number of paragraphs and the number of target knowledge triples, and obtaining a knowledge perception score corresponding to the text segment;

[0049] The sum of the semantic search score and the knowledge-aware score is used as the segmentation score of the text segment.

[0050] In some feasible implementations, the score calculation module is specifically used to:

[0051] Calculate the semantic search score corresponding to each of the text segments based on cosine similarity:

[0052]

[0053] Among them, Score se Represents query information Q and text segment S i The similarity score between them, MPNet(Q) represents the encoding result of the multimodal network MPNet for the query information Q, that is, the vector representation of Q, MPNet(S i ) represents the multimodal network MPNet for text segmentation S i The encoding result, namely S i The vector representation of .

[0054] In some feasible implementations, the score calculation module is specifically used to:

[0055] Sorting the text segments according to the values ​​of the segmentation scores, and extracting top k text segments from the sorted text segments as target text segments;

[0056] Wherein, k is a positive integer.

[0057] In some feasible implementations, the encoder includes at least a vector mapping layer, a multi-head attention layer, and a residual connection and normalization layer, and the encoding module is specifically used for:

[0058] Inputting the target text segment into the encoder, extracting words from the target text segment through the vector mapping layer, and obtaining word vectors corresponding to the target text segment;

[0059] Adding position coding through the vector mapping layer to obtain a corresponding position coding vector, and adding the word vector to the position coding vector to obtain an input sequence corresponding to the target text segment;

[0060] Input the input sequence into the multi-head attention layer for encoding, and calculate the attention information corresponding to each word in the input sequence;

[0061] The attention information is input into the residual connection and normalization layer for processing, and the encoding features corresponding to the target segment record are output.

[0062] In some feasible implementations, the summary generation module is specifically used to:

[0063] removing stop words in the target knowledge triples and merging the remaining words into knowledge terms;

[0064] Concatenate the query information, the coding feature and the knowledge term to obtain a corresponding concatenation result;

[0065] The splicing result is input into the decoder for decoding processing, and the conference summary corresponding to the query information is output through the output layer.

[0066] The embodiment of the present invention further discloses an electronic device, comprising a processor, a communication interface, a memory and a communication bus, wherein the processor, the communication interface and the memory communicate with each other via the communication bus;

[0067] The memory is used to store computer programs;

[0068] The processor is used to implement the method described in the embodiment of the present invention when executing the program stored in the memory.

[0069] The embodiment of the present invention further discloses a computer-readable storage medium having instructions stored thereon, which, when executed by one or more processors, enables the processors to execute the method described in the embodiment of the present invention.

[0070] The embodiments of the present invention include the following advantages:

[0071] In an embodiment of the present invention, when a user wants to generate a corresponding meeting summary for a meeting record, a query instruction for the meeting record can be output. The terminal can respond to the query instruction for the meeting record, determine the query information corresponding to the query instruction, and input the meeting record into the summary generation model to process the content of the meeting record, obtain the text segment corresponding to the query information and the target knowledge triple corresponding to the text segment, then calculate the segmentation score corresponding to each text segment, and extract several target text segments with the highest score from the text segment based on the segmentation score, and then input the target text segment into the encoder for encoding, output the corresponding encoding feature, and finally input the query information, the encoding feature and the target knowledge triple into the decoder for decoding processing, and output the meeting summary corresponding to the query information through the output layer, so that in the process of generating the meeting summary, the fragments related to the query information are compared through the segmentation score, and the knowledge triples related to the query information are integrated into the summary generation process, thereby improving the accuracy, completeness and coherence of the generated meeting summary. BRIEF DESCRIPTION OF THE DRAWINGS

[0072] Figure 1 is a flowchart of a method for generating a conference summary provided in an embodiment of the present invention;

[0073] Figure 2 is a schematic diagram of processing an input sequence provided in an embodiment of the present invention;

[0074] Figure 3 is a schematic diagram of the generation process of a conference summary provided in an embodiment of the present invention;

[0075] Figure 4 It is a structural block diagram of a device for generating a conference summary provided in an embodiment of the present invention. DETAILED DESCRIPTION

[0076] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the present invention is further described in detail below with reference to the accompanying drawings and specific embodiments.

[0077] As an example, automatically generating conference summaries is not an easy task for computers and faces several key challenges: First, data resources are scarce, and obtaining sufficient and high-quality conference data is key. Second, effective conference data modeling is required to deeply understand conference texts and generate accurate and rich summaries. Third, conference summaries need to address challenges in specific fields, because different conference data have their own forms and characteristics, and how to effectively capture these characteristics is crucial to generating high-quality summaries.

[0078] In this regard, in the present invention, when a user wants to generate a corresponding meeting summary for a meeting record, a query instruction for the meeting record can be output. The terminal can respond to the query instruction for the meeting record, determine the query information corresponding to the query instruction, and input the meeting record into the summary generation model to process the content of the meeting record, obtain the text segment corresponding to the query information and the target knowledge triple corresponding to the text segment, then calculate the segmentation score corresponding to each text segment, and extract several target text segments with the highest score from the text segment based on the segmentation score, and then input the target text segment into the encoder for encoding, output the corresponding encoding feature, and finally input the query information, the encoding feature and the target knowledge triple into the decoder for decoding processing, and output the meeting summary corresponding to the query information through the output layer, so that in the process of generating the meeting summary, the fragments related to the query information are matched through the segmentation score, and the knowledge triples related to the query information are integrated into the summary generation process, thereby improving the accuracy, completeness and coherence of the generated meeting summary.

[0079] Reference Figure 1 , shows a flowchart of a method for generating a conference summary provided in an embodiment of the present invention, which may specifically include the following steps:

[0080] Step 101, obtaining a meeting record and a summary generation model for generating a meeting summary, wherein the summary generation model at least includes an encoder, a decoder, and an output layer;

[0081] When a user wants to extract key information from a meeting record, he or she can first select the corresponding meeting record and determine the summary generation model that has been trained to generate a meeting summary, so as to generate a summary of the meeting record through the summary generation model to obtain the corresponding meeting summary.

[0082] Among them, the summary generation model can be a model trained according to the corresponding training data, which can at least include an encoder, a decoder and an output layer. The encoder can be used to extract features from the input data to obtain information that can effectively express the data features; the decoder can be used to decode the features extracted by the encoder and restore them to corresponding text information to construct a corresponding meeting summary; the output layer can be used to output the constructed meeting summary, and the present invention is not limited to this.

[0083] In some feasible implementation methods, the collection of training data can be done by building corresponding applications. When users apply for meeting rooms for meetings through the applications on a daily basis, the applications can collect corresponding scheduled meeting data, such as meeting topics, estimated duration, and meeting requirements, and build corresponding meeting content data sets to provide valuable data support for in-depth analysis of meeting efficiency, optimization of meeting processes, and improvement of team collaboration levels. It also provides a good data foundation for model training.

[0084] For example, you can select several meetings that have ended, and these meetings can cover as many work departments and employees in different positions as possible to improve the comprehensiveness and completeness of the training data.

[0085] Among them, the meeting topics may involve various types such as daily communication meetings, type one meetings, type two meetings, type three meetings and other meetings, and the number of meetings corresponding to each meeting topic and the number of meetings corresponding to each work department shall be counted.

[0086] In addition, the minutes recorded by the user after the meeting can also be processed, for example, emphasis sentences, topic classification, segment titles, segment summaries, and overall summaries can be marked in the meeting minutes. Among them, emphasis sentences refer to some sentences with rich information in the original meeting minutes. The marked sentences must be fluent and complete. At the same time, to avoid excessive annotations, the annotated sentences should not exceed 10% of the entire manuscript, and the corresponding colors should be used for emphasis. For topic classification, since meetings usually involve multiple topics, different topic segments can be identified by inserting a special token [EOS] at the end of each segment. It should be noted that topic segmentation is annotated at the discourse level. After obtaining the topic segment, each meeting should be identified with a topic, and the number of words for each title should be between 5-20 words. For segment titles, after obtaining the topic segment, the annotator should provide a title to identify the topic of each segment, and the number of words for each title should be between 5-20 words. For segment summaries, a summary should be written for each segment. Unlike titles, segment summaries pay more attention to the details of the content. Each segment summary should contain 100-150 words; for the overall summary, it can be an overall summary covering the most prominent and informative content of the meeting, and the number of words in each overall summary should be between 200 and 250 words, etc., so as to improve the quality and availability of the training data by marking the corresponding content in the original meeting records and processing the training data.

[0087] After the processed training data is obtained through the above process, the training data can be converted into structured data so as to produce corresponding training data sets and test data sets, etc., for training the summary generation model.

[0088] In addition, for the trained summary generation model, the Rouge indicator can be used to evaluate the conference summary generation results. The Rouge indicator is an indicator used in the field of natural language processing to evaluate the quality of automatically generated text. The Rouge indicator can calculate the overlap between the reference summary and the generated summary to evaluate their similarity. Rouge-1 represents the word overlap ratio between the system summary and the reference summary. Rouge-2 represents the double-word group overlap ratio between the system summary and the reference summary. Rouge-L: represents the similarity between the longest common subsequence of the generated summary and the reference summary. The higher the score, the more accurate the generated summary is. The calculation formula is as follows.

[0089]

[0090] Among them, the numerator refers to the number of N-grams that are the same in the generated summary and the reference summary, and the denominator is the number of N-grams in the reference summary. The model effect of the summary generation model is verified through experiments to ensure the predictive effect of the summary generation model in applications.

[0091] Step 102, in response to a query instruction for the conference record, determining query information corresponding to the query instruction, and inputting the conference record into the summary generation model to perform content processing on the conference record to obtain a text segment corresponding to the query information and a target knowledge triple corresponding to the text segment;

[0092] In an embodiment of the present invention, for a meeting record for which a meeting summary needs to be generated, a user can input corresponding query information so as to guide the summary generation model to generate a corresponding meeting summary through the query information. The query information can be information for guiding the model to extract relevant information from the meeting record to generate a meeting summary. The query information is used to clarify the goal and scope of the summary and ensure that the generated meeting summary is relevant to the user's needs or topics of interest.

[0093] For example, suppose there is a meeting record that discusses multiple topics, such as "marketing strategy", "product development", and "customer feedback". If the user is interested in "marketing strategy", the query information can be set to "marketing strategy". The system will extract relevant fragments from the meeting record based on this query information to generate a summary about marketing strategy.

[0094] In a specific implementation, based on the query information input by the user, the meeting minutes are input into the summary generation model, and the summary generation model processes the content of the meeting minutes to obtain the text segmentation corresponding to the query information and the target knowledge triples corresponding to the text segmentation. Among them, the summary generation model also includes a knowledge-aware summarizer, and the meeting minutes are input into the knowledge-aware summarizer for segmentation to obtain a number of text segments, and then information is extracted from the text segments to obtain the knowledge triples corresponding to the text segments, and then the knowledge triples are filtered to extract the target knowledge triples that overlap with the query information once.

[0095] It should be noted that in order to preserve the contextual information between utterances, the summary generation model can divide the meeting record into several text segments, each of which can contain multiple sentences. Specifically, given an input meeting record T, the knowledge-aware summarizer in the summary generation model can divide the meeting record into n segments S = {S1, S2, ..., S n}, each S i The length of is less than l tokens. For example, by feeding the meeting minutes into the segment one by one until it reaches l.

[0096] For knowledge triples, after the corresponding text fragments are extracted, the corresponding knowledge triples can be extracted from the text fragments through a two-stage framework. Specifically, in the first stage, knowledge triples can be extracted from the text fragments. During the extraction process, user-annotated sentences can be emphasized; in the second stage, the extracted knowledge triples related to the query can be used as additional input for generating summaries. Related content can be seen in the relevant description below.

[0097] For example, the following Table 1 shows an example of extracting knowledge triples related to query information from conference records:

[0098]

[0099] Table 1

[0100] Specifically, in the above example, the user can input a corresponding question (query information), and based on the input question, relevant knowledge triples are extracted from the conference record. The knowledge triples can be used for query-related segment extraction and summary generation.

[0101] Step 103, calculating the segmentation scores corresponding to the respective text segments, and extracting a plurality of target text segments ranked first in the score order from the text segments based on the segmentation scores;

[0102] After extracting the corresponding text segments and knowledge triples, the text segments can be processed based on the knowledge-aware sorting method to calculate the segmentation scores corresponding to each text segment. Then, based on the segmentation scores, several target text segments with higher scores can be extracted from the text segments so as to generate summaries based on the target text segments.

[0103] In some feasible implementations, the query information and the text segment can be used to perform calculations to obtain a semantic search score corresponding to the text segment, and the total number of paragraphs corresponding to the text segment can be obtained. The total number of paragraphs and the number of target knowledge triples can be used to perform normalization calculations to obtain a knowledge perception score corresponding to the text segment, and then the sum of the semantic search score and the knowledge perception score can be used as the segmentation score of the text segment.

[0104] Among them, for the semantic search score, the semantic search score corresponding to each of the text segments can be calculated according to the cosine similarity:

[0105]

[0106] Among them, Score se Represents query information Q and text segment S i The similarity score between them (i.e., semantic search score), MPNet(Q) represents the encoding result of the multimodal network MPNet for the query information Q, i.e., the vector representation of Q, and MPNet(S i ) represents the multimodal network MPNet for text segmentation S i The encoding result, namely S i The vector representation of .

[0107] For the knowledge-aware score, after extracting the corresponding knowledge triples, the knowledge triples irrelevant to the query can be filtered out, and only the target knowledge triples containing overlapping words with can be retained. Then the knowledge-aware score can be obtained by L2 summing the remaining knowledge triples m in each text segment. i Normalizing the number of , we get:

[0108] Score rank =Score se +Score ka

[0109] Among them, Score rank Score is the segment score. se Score is the semantic search score. kais the knowledge perception score. After the corresponding segmentation scores are calculated through the above process, the text segments can be sorted according to the numerical values ​​of the segmentation scores, and the top k text segments can be extracted from the sorted text segments as the target text segments. For example, S = {S1, S2, ..., S n} represents the target text segment. Wherein, k is a positive integer.

[0110] Step 104, inputting the target text into the encoder for encoding, and outputting corresponding encoding features;

[0111] After obtaining the target text segmentation, the target text segmentation can be input into the encoder for encoding to extract the encoding features that can effectively characterize the text segmentation. Among them, the encoder includes a vector mapping layer, a multi-head attention layer, a residual connection and a normalization layer. Then, the target text segmentation can be input into the encoder first, and the target text segmentation is extracted through the vector mapping layer to obtain the word vector corresponding to the target text segmentation. Then, the position encoding is added through the vector mapping layer to obtain the corresponding position encoding vector, and the word vector is added to the position encoding vector to obtain the input sequence corresponding to the target text segmentation. The input sequence is input into the multi-head attention layer for encoding, and the attention information corresponding to each word in the input sequence is calculated. Then, the attention information is input into the residual connection and normalization layer for processing, and the encoding features corresponding to the target segmentation record are output.

[0112] In one example, when encoding the words in the original article using the Transformer model, each word is first embedded, and then the position information of the word in the input sequence is considered by adding position encoding. In the input sequence, the position of each word is marked starting from 1. If the length of the input sequence is less than the length set by the encoder, a special marker with a position marker of 0 is added from the end of the text, and the corresponding position vector is a zero vector. The position vectors of other words are calculated using sine and cosine functions. Next, the input sequence is encoded through a multi-head attention layer to calculate the attention of all words in the input sequence. This process enables the model to consider both the global information and local relations of the input sequence while encoding the input sequence. Afterwards, residual connections and normalization layers are used to process the output, making the model more stable and easier to train. Finally, the input sequence passes through a position feedforward neural network layer to capture deeper input sequence features. This process utilizes the features encoded by the previous multi-head self-attention mechanism to improve the model's ability to understand the input sequence.

[0113] For example, refer to Figure 2, shows a schematic diagram of processing an input sequence provided in an embodiment of the present invention. In the Transformer model, the encoding process of the input sequence involves multiple steps, including word embedding, position encoding, and paragraph encoding. The following is a detailed description of the input sequence encoding process:

[0114] 1. Composition of the input sequence

[0115] The input sequence consists of the following:

[0116] [CLS]: Classification marker, used to indicate the start of a sequence.

[0117] Query: Query, used to guide summary generation.

[0118] [SEP]: Separator, used to separate different parts.

[0119] Meeting: Meeting minutes text.

[0120] [SEP]: Separator, used to separate different parts.

[0121] 2. Token Embeddings

[0122] Each word in the input sequence is first converted into a word embedding vector. Assuming there are N words in the input sequence, the word embedding vector of each word is represented as Ei, where i is the index of the word.

[0123] Query part:

[0124] E C2S : The word embedding vector of the first word of the query.

[0125] E1 to E N : The word embedding vectors of other words in the query.

[0126] Meeting minutes section:

[0127] E SEY : Word embedding vector of the first word of the conference transcript.

[0128] E1 to E M : Word embedding vectors for other words in conference proceedings.

[0129] 3. Segment Embeddings

[0130] To distinguish between query and transcript parts, we use paragraph embeddings. Paragraph embeddings are used to identify which part each word belongs to.

[0131] Query part:

[0132] EA : The paragraph embedding vector of the query part.

[0133] Meeting minutes section:

[0134] E B : Paragraph embedding vectors for conference proceedings.

[0135] 4. Position Embdings

[0136] In order to consider the position information of the word in the input sequence, we add position encoding. The position encoding is calculated by sine and cosine functions, and each position corresponds to a unique vector.

[0137] Query part:

[0138] E0: The positional encoding vector of the first word in the query.

[0139] E1 to E n+1 : The positional encoding vectors of other words in the query part.

[0140] Meeting minutes section:

[0141] E n : Positional encoding vector of the first word of a conference transcript section.

[0142] E n+1 To E m+1 : Positional encoding vectors for other words in the conference transcript section.

[0143] 5. Sentence Embeddings

[0144] To further distinguish different sentences, we use sentence embeddings. Sentence embeddings are used to identify which sentence each word belongs to.

[0145] Query part:

[0146] E0: Sentence embedding vector of the first word in the query.

[0147] E1: Sentence embedding vectors of other words in the query part.

[0148] Meeting minutes section:

[0149] E0: Sentence embedding vector of the first word of the conference transcript section.

[0150] Through word embedding, paragraph embedding, position encoding and sentence embedding, the input sequence is converted into a vector representation containing semantic information, position information and structural information. These vector representations will be input into the encoder of the Transformer model for further processing and encoding.

[0151] Step 105: input the query information, the encoding feature and the target knowledge triple into the decoder for decoding, and output the meeting summary corresponding to the query information through the output layer.

[0152] For the decoder of the summary generation model, it can be a pre-trained language model generated based on the transformer. Using it as the backbone model of the decoder can effectively ensure the performance of the text summary benchmark test.

[0153] In a specific implementation, during the decoding process, the extracted knowledge triples and query information can be used as additional inputs and spliced ​​with the encoded features output by the encoder, and then decoding is performed based on the spliced ​​data to output the meeting summary corresponding to the query information through the output layer. In the process of generating the meeting summary, the fragments related to the query information are matched through segmentation scores, and the knowledge triples related to the query information are integrated into the summary generation process, thereby improving the accuracy, completeness and coherence of the generated meeting summary.

[0154] Specifically, the stop words in the target knowledge triples can be removed first, and the remaining words can be merged into knowledge words. Then, the query information, the encoded features and the knowledge words can be concatenated to obtain the corresponding concatenated results. Finally, the concatenated results can be input into the decoder for decoding processing, and the conference summary corresponding to the query information can be output through the output layer. In the process of generating the conference summary, the fragments related to the query information are matched through the segmentation scores, and the knowledge triples related to the query information are integrated into the summary generation process, thereby improving the accuracy, completeness and coherence of the generated conference summary.

[0155] It should be noted that in order to incorporate the extracted knowledge into summary generation, we use this knowledge as additional input in addition to the query and fragments. Specifically, we remove the stop words in the knowledge triples and then merge all the remaining words in the knowledge triples in each fragment into a set of knowledge phrases. The concatenation of each fragment Si, knowledge phrase Ki and query Q will be processed by the FiD-BART encoder:

[0156]

[0157] Among them, Encoder represents the encoding method, Q is the query, K is i is a knowledge phrase, S iis the encoded feature. Finally, the decoder performs an encoder-decoder on the concatenation of the encoder outputs of all segments, so that the computational complexity grows linearly with the number of segments, rather than quadratically. At the same time, all segments are processed jointly in the decoder, allowing the model to aggregate information from multiple segments and improve the accuracy, completeness, and coherence of the generated conference summary.

[0158] It should be noted that the embodiments of the present invention include but are not limited to the above examples. It is understandable that those skilled in the art can also make settings according to actual needs under the guidance of the ideas of the embodiments of the present invention, and the present invention is not limited to this.

[0159] In an embodiment of the present invention, when a user wants to generate a corresponding meeting summary for a meeting record, a query instruction for the meeting record can be output. The terminal can respond to the query instruction for the meeting record, determine the query information corresponding to the query instruction, and input the meeting record into the summary generation model to process the content of the meeting record, obtain the text segment corresponding to the query information and the target knowledge triple corresponding to the text segment, then calculate the segmentation score corresponding to each text segment, and extract several target text segments with the highest score from the text segment based on the segmentation score, and then input the target text segment into the encoder for encoding, output the corresponding encoding feature, and finally input the query information, the encoding feature and the target knowledge triple into the decoder for decoding processing, and output the meeting summary corresponding to the query information through the output layer, so that in the process of generating the meeting summary, the fragments related to the query information are compared through the segmentation score, and the knowledge triples related to the query information are integrated into the summary generation process, thereby improving the accuracy, completeness and coherence of the generated meeting summary.

[0160] In order to enable those skilled in the art to better understand the technical solutions in the embodiments of the present invention, the following examples are used for exemplary description:

[0161] As an example, see Figure 3 , shows a schematic diagram of the generation process of the conference summary provided in an embodiment of the present invention, which may include the processes of cleaning and building the original conference data, loading the conference model, initially building the conference summary, and optimizing the conference summary, wherein:

[0162] 1. Cleaning and construction of original conference data

[0163] In the process of cleaning and constructing the original meeting data, the first task is to ensure the quality and availability of the data. This involves correcting typos in the original meeting records and removing or modifying parts of the sentences that are not fluent to ensure the accurate communication of information. In addition, irrelevant information, such as repeated speeches and informal exchanges, needs to be eliminated to ensure the refinement and efficiency of the input data. By collecting and organizing a large amount of real meeting record data, a dataset specifically for meeting summary generation was constructed. This dataset covers the content of various meeting types and fields, ensuring the diversity and representativeness of the training data for the subsequent summary generation model.

[0164] 2. Conference model loading

[0165] Conference model loading refers to the process of applying a pre-trained model to conference summary generation. The model file is read from storage and loaded into memory for immediate use. The loaded model can quickly process new conference data and automatically generate summaries. In order to ensure that the model can run efficiently, it is necessary to continuously optimize the model to ensure that the model can work stably on different devices and provide users with a smoother service experience.

[0166] 3. Preliminary construction of conference summary

[0167] The model can understand the text content and automatically identify key information. The loaded model is used to analyze the text structure, keywords and sentence importance through algorithms; then, the model selects the sentences or paragraphs that best represent the main theme of the meeting to generate a summary based on the preset summary length requirements. In addition, the level of detail of the summary can be adjusted by setting different parameters to meet the needs of different users.

[0168] 4. Meeting summary optimization

[0169] The purpose of conference summary optimization is to improve the quality and practicality of the summary, ensure that it can accurately and concisely convey the core content of the meeting, delete unnecessary repetitive information, and maintain the simplicity of the summary. At the same time, the model is continuously optimized to make the conference summary more accurate.

[0170] Through the above process, in the process of generating conference summaries, the fragments related to the query information are matched by segmentation scores, and the knowledge triples related to the query information are integrated into the summary generation process, thereby improving the accuracy, completeness and coherence of the generated conference summary.

[0171] It should be noted that, for the sake of simplicity, the method embodiments are described as a series of action combinations, but those skilled in the art should be aware that the embodiments of the present invention are not limited by the order of the actions described, because according to the embodiments of the present invention, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions involved are not necessarily required by the embodiments of the present invention.

[0172] Reference Figure 4 , shows a structural block diagram of a device for generating a conference summary provided in an embodiment of the present invention, which may specifically include the following modules:

[0173] A record acquisition module 401, used to acquire meeting records and a summary generation model for generating a meeting summary, wherein the summary generation model includes at least an encoder, a decoder, and an output layer;

[0174] A record processing module 402 is used to respond to a query instruction for the conference record, determine query information corresponding to the query instruction, and input the conference record into the summary generation model to perform content processing on the conference record to obtain a text segment corresponding to the query information and a target knowledge triple corresponding to the text segment;

[0175] A score calculation module 403 is used to calculate the segmentation scores corresponding to each of the text segments, and extract a plurality of target text segments ranked first in the score order from the text segments based on the segmentation scores;

[0176] An encoding module 404, configured to input the target text segments into the encoder for encoding and output corresponding encoding features;

[0177] The summary generation module 405 is used to input the query information, the encoding feature and the target knowledge triple into the decoder for decoding processing, and output the meeting summary corresponding to the query information through the output layer.

[0178] In some feasible implementations, the summary generation model further includes a knowledge-aware summarizer, and the record processing module 402 is specifically used to:

[0179] Inputting the meeting record into the knowledge-aware summarizer for segmentation to obtain a plurality of text segments;

[0180] Extracting information from the text segment to obtain knowledge triples corresponding to the text segment;

[0181] The knowledge triples are filtered to extract target knowledge triples that overlap with the query information once.

[0182] In some feasible implementations, the score calculation module 403 is specifically used to:

[0183] Using the query information and the text segment to perform calculations to obtain a semantic search score corresponding to the text segment;

[0184] Obtaining the total number of paragraphs corresponding to the text segment, performing normalization calculation using the total number of paragraphs and the number of target knowledge triples, and obtaining a knowledge perception score corresponding to the text segment;

[0185] The sum of the semantic search score and the knowledge-aware score is used as the segmentation score of the text segment.

[0186] In some feasible implementations, the score calculation module 403 is specifically used to:

[0187] Calculate the semantic search score corresponding to each of the text segments based on cosine similarity:

[0188]

[0189] Among them, Score se Represents query information Q and text segment S i The similarity score between them, MPNet(Q) represents the encoding result of the multimodal network MPNet for the query information Q, that is, the vector representation of Q, MPNet(S i ) represents the multimodal network MPNet for text segmentation S i The encoding result, namely S i The vector representation of .

[0190] In some feasible implementations, the score calculation module 403 is specifically used to:

[0191] Sorting the text segments according to the values ​​of the segmentation scores, and extracting top k text segments from the sorted text segments as target text segments;

[0192] Wherein, k is a positive integer.

[0193] In some feasible implementations, the encoder includes at least a vector mapping layer, a multi-head attention layer, and a residual connection and normalization layer, and the encoding module 404 is specifically used to:

[0194] Inputting the target text segment into the encoder, extracting words from the target text segment through the vector mapping layer, and obtaining word vectors corresponding to the target text segment;

[0195] Adding position coding through the vector mapping layer to obtain a corresponding position coding vector, and adding the word vector to the position coding vector to obtain an input sequence corresponding to the target text segment;

[0196] Input the input sequence into the multi-head attention layer for encoding, and calculate the attention information corresponding to each word in the input sequence;

[0197] The attention information is input into the residual connection and normalization layer for processing, and the encoding features corresponding to the target segment record are output.

[0198] In some feasible implementations, the summary generation module 405 is specifically used to:

[0199] removing stop words in the target knowledge triples and merging the remaining words into knowledge terms;

[0200] Concatenate the query information, the coding feature and the knowledge term to obtain a corresponding concatenation result;

[0201] The splicing result is input into the decoder for decoding processing, and the conference summary corresponding to the query information is output through the output layer.

[0202] As for the device embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.

[0203] In addition, an embodiment of the present invention further provides an electronic device, including: a processor, a memory, and a computer program stored in the memory and executable on the processor. When the computer program is executed by the processor, the various processes of the above-mentioned embodiment of the method for generating a conference summary are implemented, and the same technical effect can be achieved. To avoid repetition, it will not be described here.

[0204] The embodiment of the present invention further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, each process of the above-mentioned method for generating a conference summary is implemented, and the same technical effect can be achieved. To avoid repetition, it is not repeated here. The computer-readable storage medium is, for example, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.

[0205] The various embodiments in this specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the various embodiments can be referenced to each other.

[0206] It should be understood by those skilled in the art that the embodiments of the embodiments of the present invention may be provided as methods, devices, or computer program products. Therefore, the embodiments of the present invention may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, the embodiments of the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, EEPROM, Flash, and eMMC, etc.) containing computer-usable program codes.

[0207] The embodiments of the present invention are described with reference to the flowcharts and / or block diagrams of the methods, terminal devices (systems), and computer program products according to the embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of the processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing terminal device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing terminal device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0208] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing terminal device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce a manufactured product including an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.

[0209] These computer program instructions can also be loaded onto a computer or other programmable data processing terminal device so that a series of operating steps are executed on the computer or other programmable terminal device to produce a computer-implemented process, thereby providing instructions for executing on the computer or other programmable terminal device to implement the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.

[0210] Although the preferred embodiments of the present invention have been described, those skilled in the art may make additional changes and modifications to these embodiments once they have learned the basic creative concept. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the embodiments of the present invention.

[0211] Finally, it should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or terminal device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or terminal device. In the absence of further restrictions, the elements defined by the sentence "comprise a ..." do not exclude the existence of other identical elements in the process, method, article or terminal device including the elements.

[0212] The above is a detailed introduction to a method for generating a conference summary and a device for generating a conference summary provided by the present invention. Specific examples are used in this article to illustrate the principle and implementation mode of the present invention. The description of the above embodiments is only used to help understand the method of the present invention and its core idea. At the same time, for those skilled in the art, according to the idea of ​​the present invention, there will be changes in the specific implementation mode and application scope. In summary, the content of this specification should not be understood as a limitation on the present invention.

Claims

1. A method for generating a conference summary, characterized in that: include: Obtaining a meeting record and a summary generation model for generating a meeting summary, wherein the summary generation model includes at least an encoder, a decoder, and an output layer; In response to a query instruction for the conference record, query information corresponding to the query instruction is determined, and the conference record is input into the summary generation model to perform content processing on the conference record to obtain a text segment corresponding to the query information and a target knowledge triple corresponding to the text segment; Calculating the segmentation scores corresponding to each of the text segments, and extracting a plurality of target text segments ranked first in score from the text segments based on the segmentation scores; Input the target text into the encoder for encoding, and output corresponding encoding features; The query information, the encoding feature and the target knowledge triple are input into the decoder for decoding processing, and the conference summary corresponding to the query information is output through the output layer.

2. The method according to claim 1, characterized in that The summary generation model further includes a knowledge-aware summarizer, wherein the step of inputting the conference record into the summary generation model to perform content processing on the conference record to obtain text segments corresponding to the query information and target knowledge triples corresponding to the text segments includes: Inputting the meeting record into the knowledge-aware summarizer for segmentation to obtain a plurality of text segments; Extracting information from the text segment to obtain knowledge triples corresponding to the text segment; The knowledge triples are filtered to extract target knowledge triples that overlap with the query information once.

3. The method according to claim 1 or 2, characterized in that: The calculating the segmentation scores corresponding to each of the text segments includes: Using the query information and the text segment to perform calculations to obtain a semantic search score corresponding to the text segment; Obtaining the total number of paragraphs corresponding to the text segment, performing normalization calculation using the total number of paragraphs and the number of target knowledge triples, and obtaining a knowledge perception score corresponding to the text segment; The sum of the semantic search score and the knowledge-aware score is used as the segmentation score of the text segment.

4. The method according to claim 3, characterized in that The using the query information and the text segment to perform calculations to obtain a semantic search score corresponding to the text segment includes: Calculate the semantic search score corresponding to each of the text segments based on cosine similarity: Among them, Score se Represents query information Q and text segment S i The similarity score between them, MPNet(Q) represents the encoding result of the multimodal network MPNet for the query information Q, that is, the vector representation of Q, MPNet(S i ) represents the multimodal network MPNet for text segmentation S i The encoding result, namely S i The vector representation of .

5. The method according to claim 1, characterized in that The step of extracting a plurality of target text segments ranked first in score from the text segments based on the segmentation scores includes: Sorting the text segments according to the values ​​of the segmentation scores, and extracting top k text segments from the sorted text segments as target text segments; Wherein, k is a positive integer.

6. The method according to claim 1, characterized in that The encoder includes at least a vector mapping layer, a multi-head attention layer, and a residual connection and normalization layer. The target text is segmented and input into the encoder for encoding, and the corresponding encoding features are output, including: Inputting the target text segment into the encoder, extracting words from the target text segment through the vector mapping layer, and obtaining word vectors corresponding to the target text segment; Adding position coding through the vector mapping layer to obtain a corresponding position coding vector, and adding the word vector to the position coding vector to obtain an input sequence corresponding to the target text segment; Input the input sequence into the multi-head attention layer for encoding, and calculate the attention information corresponding to each word in the input sequence; The attention information is input into the residual connection and normalization layer for processing, and the encoding features corresponding to the target segment record are output.

7. The method according to claim 1, characterized in that The step of inputting the query information, the encoding feature, and the target knowledge triple into the decoder for decoding, and outputting the conference summary corresponding to the query information through the output layer includes: removing stop words in the target knowledge triples and merging the remaining words into knowledge terms; Concatenate the query information, the coding feature and the knowledge term to obtain a corresponding concatenation result; The splicing result is input into the decoder for decoding processing, and the conference summary corresponding to the query information is output through the output layer.

8. A device for generating a conference summary, characterized in that: include: A record acquisition module, used to acquire meeting records, and a summary generation model for generating a meeting summary, wherein the summary generation model includes at least an encoder, a decoder, and an output layer; A record processing module, for responding to a query instruction for the conference record, determining query information corresponding to the query instruction, and inputting the conference record into the summary generation model to perform content processing on the conference record, and obtaining a text segment corresponding to the query information and a target knowledge triple corresponding to the text segment; A score calculation module, used to calculate the segmentation scores corresponding to each of the text segments, and extract a plurality of target text segments ranked first in the score order from the text segments based on the segmentation scores; An encoding module, used for inputting the target text into the encoder for encoding in segments, and outputting corresponding encoding features; The summary generation module is used to input the query information, the encoding features and the target knowledge triples into the decoder for decoding processing, and output the meeting summary corresponding to the query information through the output layer.

9. An electronic device, characterized in that: It includes a processor, a communication interface, a memory and a communication bus, wherein the processor, the communication interface and the memory communicate with each other through the communication bus; The memory is used to store computer programs; The processor is used to implement the method according to any one of claims 1 to 7 when executing the program stored in the memory.

10. A computer-readable storage medium having instructions stored thereon, which, when executed by one or more processors, cause the processors to perform the method according to any one of claims 1 to 7.