Method and System for Extracting Multi-Level Events from Telephone Audio Data and Visualizing Them

Through an agent system combining a multi-level event extraction method and a large language model, the problem of event extraction and visualization in telephone audio data is solved, accurate event extraction and visualization in complex scenarios is realized, and rapid analysis of multiple application scenarios is supported.

CN120087364BActive Publication Date: 2025-07-18ZHEJIANG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510567129.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-30
Publication Date
2025-07-18
Estimated Expiration
2045-04-30

AI Technical Summary

Technical Problem

The prior art is difficult to accurately extract events from telephone audio data and visualize them, especially when facing the problems of language colloquialization, logical jumps and unclear grammar, and lack of effective integration and display methods.

Method used

A multi-level event extraction method is adopted to generate an event map and visualize it through sub-task processes such as semantic chunking, abstract generation, event extraction and relationship extraction, etc., combined with large language models and feedback optimization agents.

Benefits of technology

It realizes the accurate extraction and visualization of event information in telephone audio in various scenarios without manual intervention, and supports the rapid understanding and analysis of complex dialogue scenarios such as customer service quality inspection, meeting minutes analysis and medical consultation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120087364B_ABST
    Figure CN120087364B_ABST
Patent Text Reader

Abstract

The present invention discloses a method and system for extracting multi-level events from telephone audio data and visualizing them. First, the target telephone audio data is converted into a target text for event extraction. Then, the dialogue type and scenario of the target text are identified to generate event extraction targets, and the event extraction task is decomposed into a sub-task process composed of semantic chunking sub-tasks, summary generation sub-tasks, event extraction sub-tasks, event relationship extraction sub-tasks, and visualization sub-tasks. At the same time, the execution tools to be scheduled for each sub-task are specified. Finally, according to the planned sub-task process, each sub-task is executed in sequence by scheduling the pre-specified execution tools and the knowledge base corresponding to the current dialogue type and scenario. The method of the present invention makes full use of the autonomy of the intelligent agent and the generalized learning and reasoning ability, and can complete the event extraction without a priori structure without manual intervention, and can be used as a general event extraction and visualization solution applicable to various scenarios.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of audio data processing, and particularly relates to a method and system for extracting multi-level events from telephone audio data and visualizing them. Background Art

[0002] In modern communication, telephone calls have become an important way to obtain real-time information, especially in the fields of business, customer service, and emergency response. Traditional telephone records usually exist in audio form. However, extracting valuable information from them and performing visual analysis faces many challenges. The rapid development of speech recognition technology has made it possible to convert audio into text, but there are still difficulties in extracting event information from the text and visualizing it effectively.

[0003] In recent years, event extraction technology has received extensive attention, which aims to identify and extract specific events and their related information from text. This process usually involves natural language processing (NLP) and machine learning technologies, especially the application of deep learning models. By combining telephone audio data with advanced speech recognition and event extraction models, structured event lists can be automatically generated and further transformed into graph structures to achieve more intuitive data visualization.

[0004] Existing research mainly focuses on the single application of speech recognition and event extraction, lacking effective integration of the two. In addition, how to display the extracted events in a graphical way to help users quickly understand and analyze the relationships between events is still an unsolved problem. Moreover, in fact, when extracting events from telephone audio data, the transcribed text has problems such as colloquial language, easy logical jumps, and unclear grammatical structures. Whether using conventional natural language processing tools or large language models, it is not possible to accurately extract the event information in them. Therefore, there is an urgent need to propose an event extraction and visualization method based on telephone audio data to provide a new solution for information processing and decision support. Summary of the Invention

[0005] The purpose of the present invention is to solve the problem in the prior art that it is difficult to accurately extract events from voice calls and visualize them, and to provide a method and system for extracting multi-level events from telephone audio data and visualizing them.

[0006] The specific technical solutions adopted by the present invention are as follows:

[0007] In the first aspect, the present invention provides a method for extracting multi-level events from telephone audio data and visualizing them, which includes:

[0008] S1. Convert the target telephone audio data into text dialogue data as the target text for event extraction;

[0009] S2. Identify the dialogue type and scenario of the target text, generate the target of event extraction, and decompose the event extraction task into a sub-task process consisting of semantic chunking sub-tasks, summary generation sub-tasks, event extraction sub-tasks, event relationship extraction sub-tasks, and visualization sub-tasks. At the same time, specify the execution tools to be scheduled for each sub-task;

[0010] S3. According to the planned sub-task process, sequentially execute each sub-task by scheduling the pre-specified execution tools and the knowledge base corresponding to the current dialogue type and scenario, and perform semantic chunking, summary generation, event extraction, event relationship extraction, and visualization in sequence. Among the semantic chunking sub-task, summary generation sub-task, and event extraction sub-task, the execution tools need to call the large language model and the benchmark method corresponding to the current sub-task in a double-branch form to generate the execution results of the current sub-task respectively. Then, the feedback optimization agent built based on the large language model performs consistency verification on the execution results of the two branches, and fuses and aligns the execution results of the two branches through a multi-round self-check feedback strategy to generate the final execution result of each sub-task. In the event relationship extraction sub-task, it is necessary to construct an event graph with all the events extracted by the event extraction sub-task as nodes through the execution tool, and classify the event relationships between the events through a graph classification algorithm. Finally, the execution tool of the visualization sub-task visualizes the event graph.

[0011] As a preference of the first aspect above, the execution tool specified to be scheduled for the semantic chunking sub-task is a semantic chunking agent, and the internal processing process of the semantic chunking agent is as follows:

[0012] A1. Convert each statement in the target text into a vector representation through encoding, extract the average vector representation of multiple statements within the window through a sliding window of a fixed size, calculate the similarity of the average vector representations of adjacent sliding windows as the similarity between windows, select the local minimum value on the similarity sequence between windows as the candidate semantic deep valley point, and for the statement sequence composed of all the statements extracted from two sliding windows within each candidate semantic deep valley point, calculate the similarity of the vector representations of all adjacent statements and take the minimum value as the mutation depth. If the mutation depth of a candidate semantic deep valley point exceeds the threshold, take the adjacent statement with the largest vector representation similarity as the semantic segmentation point, otherwise discard the candidate semantic deep valley point; divide the target text at all semantic segmentation points to obtain a series of semantic text blocks representing different independent topics, and obtain the first set of semantic text blocks;

[0013] A2. Guide the large language model to divide the target text into semantic text blocks through the first prompt word to obtain the second set of semantic text blocks;

[0014] A3. The feedback optimization agent built based on the large language model performs consistency verification on the first semantic text block set and the second semantic text block set, and fuses and aligns the two sets through a multi-round self-check feedback strategy to generate a third semantic text block set as the execution result of the final output of the semantic chunking subtask.

[0015] As a preference of the first aspect above, the execution tool designated and scheduled by the abstract generation subtask is an abstract generation agent, and the internal processing process of the abstract generation agent is as follows:

[0016] B1. Generate an abstract sentence for each semantic text block in the execution result output by the semantic chunking subtask through a pre-trained language model with an encoder-decoder architecture to obtain a first abstract sentence set;

[0017] B2. Guide the large language model through a second prompt word to generate an abstract sentence for each semantic text block in the execution result output by the semantic chunking subtask to obtain a second abstract sentence set;

[0018] B3. The feedback optimization agent built based on the large language model performs consistency verification on the first abstract sentence set and the second abstract sentence set, and fuses and aligns the two sets through a multi-round self-check feedback strategy to generate a third abstract sentence set as the execution result of the final output of the abstract generation subtask.

[0019] As a preference of the first aspect above, the execution tool designated and scheduled by the event extraction subtask is an event extraction agent, and the internal processing process of the event extraction agent is as follows:

[0020] C1. Use a semantic role labeling tool to identify each abstract sentence in the execution result output by the abstract generation subtask, determine the predicate and the corresponding semantic arguments in the abstract sentence, use the predicate as the trigger word, and the corresponding semantic arguments as the event arguments, and output a first set of event elements composed of the trigger words and event arguments in all abstract sentences;

[0021] C2. Guide the large language model through a third prompt word to identify each abstract sentence in the execution result output by the abstract generation subtask to obtain a second set of event elements composed of the event type, trigger word, event role, and event argument in all abstract sentences;

[0022] C3. The feedback optimization agent built based on the large language model performs consistency verification on the trigger words and event arguments in the first set of event elements and the second set of event elements, and fuses and aligns the two sets through a multi-round self-check feedback strategy to generate a third set of event elements composed of the event type, trigger word, event role, and event argument in all abstract sentences as the execution result of the final output of the event extraction subtask.

[0023] Preferably, for the first aspect described above, the execution tool designated for scheduling by the event relationship extraction subtask is an event relationship extraction agent, and the internal processing procedure of the event relationship extraction agent is as follows:

[0024] D1. Each semantic text block in the execution result output by the semantic chunking subtask is regarded as an independent event, and the summary sentence corresponding to this semantic text block in the execution result output by the summary generation subtask is used as the event backbone sentence. The event backbone sentence is encoded into an event vector representation through a language representation model. All events are used as nodes to construct an event graph in the form of an undirected fully connected graph, and the original event vector representations of each event node are propagated and updated among nodes through a graph attention network to obtain the final event vector representation of each event node;

[0025] D2. The final event vector representations of each pair of event nodes in the event graph are fused into a combined vector, and the combined vector is input into an event relationship classifier to obtain the event relationship between this pair of event nodes.

[0026] Preferably, for the first aspect described above, the execution tool designated for scheduling by the visualization subtask is a triple data visualization tool, and the internal processing procedure is as follows:

[0027] First, the format of the event extraction results obtained from other subtasks is normalized to form structured triple data, and then a graphical visualization tool is called to render the triple data to form a rendered event graph for display on the user interface.

[0028] In a second aspect, the present invention provides a system for extracting and visualizing multi-level events from telephone audio data, which includes:

[0029] An audio text conversion module, configured to convert target telephone audio data into text dialogue data as the target text for event extraction;

[0030] A task decomposition module, configured to identify the dialogue type and scenario of the target text, generate event extraction targets, decompose the event extraction task into a subtask process composed of a semantic chunking subtask, a summary generation subtask, an event extraction subtask, an event relationship extraction subtask, and a visualization subtask, and simultaneously specify the execution tools required for scheduling each subtask;

[0031] A task execution module, which is used to execute each subtask in sequence according to the planned subtask process, by scheduling pre-specified execution tools and the knowledge base corresponding to the current dialogue type and scenario, and successively perform semantic chunking, summary generation, event extraction, event relationship extraction, and visualization; among them, in the semantic chunking subtask, summary generation subtask, and event extraction subtask, the execution tool needs to call the large language model and the benchmark method corresponding to the current subtask in a two-branch form to generate the execution results of the current subtask respectively. Then, the feedback optimization agent constructed based on the large language model performs consistency verification on the execution results of the two branches, and fuses and aligns the execution results of the two branches through a multi-round self-check feedback strategy to generate the final execution result of each subtask; in the event relationship extraction subtask, it is necessary to construct an event graph with all the events extracted by the event extraction subtask as nodes through the execution tool, and classify the event relationships between the events through a graph classification algorithm. Finally, the execution tool of the visualization subtask visualizes the event graph.

[0032] In a third aspect, the present invention provides a computer program product, including a computer program / instructions, which when executed by a processor, can implement the method of extracting multi-level events from telephone audio data and visualizing as described in any one of the above first aspect solutions.

[0033] In a fourth aspect, the present invention provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it can implement the method of extracting multi-level events from telephone audio data and visualizing as described in any one of the above first aspect solutions.

[0034] In a fifth aspect, the present invention provides a computer electronic device, which includes a memory and a processor;

[0035] The memory is used to store a computer program;

[0036] The processor is used to, when executing the computer program, be able to implement the method of extracting multi-level events from telephone audio data and visualizing as described in any one of the above first aspect solutions.

[0037] Compared with the prior art, the present invention has the following beneficial effects:

[0038] The present invention provides a method for extracting multi-level events from telephone audio data and visualizing. This method makes full use of the autonomy, generalization adaptability, and learning and reasoning ability characteristics of the intelligent agent, and can complete the extraction of events without a priori structure without manual intervention. It can be used as a general event extraction and visualization solution applicable to various scenarios. The method of the present invention can be used in customer service quality inspection, meeting minutes analysis, medical consultations, etc., which require rapid and complex dialogue scenarios, and has important application value. Brief Description of the Drawings

[0039] Figure 1 Schematic diagram of the method steps for extracting multi-level events from telephone audio data and visualizing them;

[0040] Figure 2 Schematic diagram of the key steps inside the semantic chunking agent;

[0041] Figure 3 Schematic diagram of the key steps inside the abstract generation agent;

[0042] Figure 4 Schematic diagram of the key steps inside the event extraction agent;

[0043] Figure 5 Schematic diagram of the system module composition for extracting multi-level events from telephone audio data and visualizing them;

[0044] Figure 6 Schematic diagram of the structure of a computer electronic device. Detailed Description of the Invention

[0045] To make the above objects, features, and advantages of the present invention more obvious and understandable, the following detailed description of the specific embodiments of the present invention will be given in conjunction with the accompanying drawings. Many specific details are set forth in the following description to facilitate a full understanding of the present invention. However, the present invention can be implemented in many other ways different from those described herein, and those skilled in the art can make similar improvements without departing from the connotation of the present invention. Therefore, the present invention is not limited by the specific embodiments disclosed below. The technical features in the various embodiments of the present invention can be combined correspondingly without conflict.

[0046] In the description of the present invention, it should be understood that the terms "first" and "second" are only used for the purpose of distinguishing descriptions, and cannot be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one of such features.

[0047] As Figure 1 shown, in a preferred embodiment of the present invention, a method for extracting multi-level events from telephone audio data and visualizing them is provided, which includes:

[0048] S1. Convert the target telephone audio data into text conversation data as the target text for the events to be extracted.

[0049] It should be noted that the text transcription technology for telephone audio data belongs to the prior art and can be achieved through operations such as speech-to-text conversion, speaker separation processing, and timestamp alignment for telephone audio data. Specifically, the speech stream can be first converted into a common audio format file, such as WAV or MP3; then the above audio format file is successively subjected to background noise reduction, voice activity detection (VAD), automatic speech recognition (ASR), speaker diarization, contact matching, and timestamp alignment operations to be converted into text conversation data. The text conversation data contains information such as call contacts, timestamps, and call content text, and records the speech content of different speakers at different times in text form. In order to reduce the redundant information in the directly transcribed text and standardize its language expression, the call content text directly transcribed in the text conversation data can be processed through a large language model to adjust the redundant information (such as filler words and repeated sentences) and language errors in the sentences while restricting the information expressed to be strictly consistent.

[0050] S2. Identify the dialogue type and scenario of the target text, generate an event extraction target, decompose the event extraction task into a sub-task process composed of semantic chunking sub-tasks, summary generation sub-tasks, event extraction sub-tasks, event relationship extraction sub-tasks, and visualization sub-tasks, and simultaneously specify the execution tools to be scheduled for each sub-task.

[0051] In an embodiment of the present invention, the above-mentioned step S2 can be executed by a large language model. By means of prompt words, it is driven to identify the dialogue type and scenario of the target text, and then generate the target of event extraction. However, it does not directly perform event extraction. Instead, the event extraction task is decomposed into five-step processes: semantic chunking subtask, summary generation subtask, event extraction subtask, event relationship extraction subtask, and visualization subtask, so as to perform quality control on the event extraction task step by step. Each subtask needs to specify corresponding execution tools according to the task characteristics to facilitate the completion of the subtask and obtain the execution result. The relevant execution tools called according to the task include but are not limited to calling external APIs, neural network models, large language models, and agents. At the same time, the database can also be accessed, and scripts or codes can be executed. In an embodiment of the present invention, it is possible to support calling one or more of the following tool libraries: semantic chunking agent, summary generation agent, event extraction agent, event relationship extraction agent, knowledge base call module (supporting RAG or SPARQL interfaces), triple data visualization tool (such as Graphviz or Neo4j), code execution environment (for running auxiliary scripts and graph construction logic). When executing, the Model Context Protocol (MCP) can be used to implement the integration of the large language model with external data sources and tools. Subsequently, the specific internal processing processes of the execution tools in each subtask will be described in combination with the specific subtask execution process.

[0052] S3. According to the planned subtask process, each subtask is sequentially executed by scheduling the pre-specified execution tools and the knowledge base corresponding to the current dialogue type and scenario, and semantic chunking, summary generation, event extraction, event relationship extraction, and visualization are carried out in sequence; among them, in the semantic chunking subtask, summary generation subtask, and event extraction subtask, the execution tools are required to call the large language model and the benchmark method corresponding to the current subtask in a two-branch form to generate the execution results of the current subtask respectively. Then, the feedback optimization agent constructed based on the large language model performs consistency verification on the execution results of the two branches, and fuses and aligns the execution results of the two branches through a multi-round self-check feedback strategy to generate the final execution result of each subtask; in the event relationship extraction subtask, it is necessary to construct an event graph with all the events extracted by the event extraction subtask as nodes through the execution tools, and classify the event relationships between the events through graph classification algorithms. Finally, the execution tool of the visualization subtask visualizes the event graph.

[0053] In an embodiment of the present invention, the execution tool specified for scheduling in the above-mentioned semantic chunking subtask is a semantic chunking agent, such as Figure 2 shown, the internal processing process of the semantic chunking agent is as follows:

[0054] A1. Convert each statement in the target text into a vector representation through encoding. Extract the average vector representation of multiple statements within the window using a sliding window of a fixed size. Calculate the similarity between the average vector representations of adjacent sliding windows as the similarity between windows. Select the local minimum value on the sequence of window similarities as the candidate semantic deep valley point. For the sequence of statements formed by all the statements extracted from two sliding windows within each candidate semantic deep valley point, calculate the similarity between the vector representations of all adjacent statements and take the maximum value as the mutation depth. If the mutation depth of a candidate semantic deep valley point exceeds the threshold, take the adjacent statement with the minimum vector representation similarity as the semantic segmentation point; otherwise, discard the candidate semantic deep valley point. Divide the target text at all semantic segmentation points into a series of semantic text blocks representing different independent topics to obtain the first set of semantic text blocks.

[0055] The core of the above step A1 lies in detecting the semantic mutation points in the target text, so as to divide the target text into several semantic text blocks. Each semantic text block is composed of one or more dialogue statements recorded during the conversation. Each semantic text block represents a relatively independent topic, and an event will be extracted from each semantic text block later. In the embodiments of the present invention, the specific implementation process of the above step A1 is as follows:

[0056] 1) Use the pre-trained language model Sentence - BERT to encode each dialogue statement to generate the vector representation of this dialogue statement. A sequence of sentence vector representations S is composed of the vector representations of all dialogue statements.

[0057] 2) Apply a sliding window with a size of 3 and a step size of 1 (both the size and the step size are in units of the vector representation of dialogue statements) on the sequence of dialogue statement vector representations for sliding extraction. Each time of sliding can extract the vector representations of 3 dialogue statements, and calculate the average vector of the vector representations of the 3 extracted dialogue statements as the window average vector representation. Assume that the sliding window extracts n times in total, and the i-th sliding window is denoted as , The average vector of the 3 dialogue statements extracted from it is denoted as ,i = 1, 2, ……, n.

[0058] 3) Calculate the cosine similarity between adjacent window average vector representations on the sequence of window average vector representations extracted and calculated in order to obtain the block similarity sequence 。

[0059] 4) Find all local minimum values on the block similarity sequence. The two windows corresponding to the local minimum values are candidate semantic valleys. The local minimum value satisfies:

[0060]

[0061] If the local minimum satisfies the above two formulas, then and constitute a candidate semantic valley bottom, indicating that there may be a semantic transition point at the position of the sliding window and .

[0062] 5) Select all 4 dialogue sentences in two sliding windows within the candidate semantic valley bottom (the sentences in the two sliding windows are repeated), calculate the cosine similarity between the vector representations of adjacent two dialogue sentences, and obtain a cosine similarity sequence composed of 3 cosine similarities in total. Take the minimum cosine similarity as the mutation depth . If is greater than the preset threshold , then find the two dialogue sentences corresponding to the minimum cosine similarity, use them as semantic segmentation points to segment the dialogue text, and divide the previous sentence and the next sentence of these two dialogue sentences into different semantic text blocks.

[0063] 6) Segment all the dialogue sentences in the target text into several semantic text blocks according to the positions of all semantic segmentation points, where the two dialogue sentences corresponding to the semantic segmentation point need to be respectively segmented into the previous and the subsequent semantic text blocks.

[0064] A2. Guide the large language model to divide the target text into semantic text blocks through the first prompt word, and obtain the second semantic text block set.

[0065] It should be noted that the above first prompt word can be constructed according to the prompt engineering construction instructions of the large language model to guide the large model to block the dialogue text. For example, the first prompt word template can be constructed as: "The current input target text is a transcription text of a telephone voice data. Please detect whether there are different topics in this text. If so, please block the entire text to form a series of semantic text blocks, ensuring that each block corresponds to an independent topic. Next, I will give the target text as follows: [Target text placeholder]".

[0066] In addition, due to the respective characteristics of call records in different dialogue types and scenarios, it is also possible to, based on the few-shot learning concept, based on the dialogue types and scenarios identified in advance from the target text, call the inference examples set for different dialogue types and scenarios in the memory bank, and add multiple exemplary semantic text block segmentation cases to the above first prompt word to facilitate the large model to understand how to segment the target text into semantic text blocks.

[0067] A3. The feedback optimization agent built based on the large language model performs consistency verification on the first semantic text block set and the second semantic text block set, and fuses and aligns the two sets through a multi-round self-check feedback strategy to generate a third semantic text block set as the execution result of the final output of the semantic chunking subtask.

[0068] It should be noted that the above feedback optimization agent is built based on the large language model. It can directly adopt a large language model with high generalization ability, or a large language model fine-tuned in the scenario of this subtask. The specific execution of the feedback optimization agent for performing consistency verification on the first semantic text block set and the second semantic text block set, and fusing and aligning the two sets through a multi-round self-check feedback strategy can be driven by input prompts.

[0069] The multi-round self-check feedback strategy in the present invention means that the large language model self-checks its verification results, examines its own thinking chain and answers, finds possible errors and corrects them, then performs consistency verification again according to the correction results, iterates multiple times, and finally outputs the results, thereby optimizing the model output and improving the accuracy and reliability of the model.

[0070] In the embodiment of the present invention, the execution tool specified and scheduled by the above abstract generation subtask is the abstract generation agent, as Figure 3 shown, the internal processing process of the abstract generation agent is as follows:

[0071] B1. Generate an abstract sentence for each semantic text block in the execution result (i.e., the third semantic text block set) output by the semantic chunking subtask through a pre-trained language model with an encoder-decoder architecture to obtain a first abstract sentence set.

[0072] It should be noted that the above pre-trained language model with an encoder-decoder architecture can adopt any model that can accurately generate abstracts, such as the T5 (Text-to-Text Transfer Transformer) model. The T5 model is a pre-trained language model based on the Transformer architecture, and its core idea is to unify various natural language processing (NLP) tasks into the "text-to-text" format, that is, the input and output of all tasks are in text form.

[0073] B2. Guide the large language model to generate an abstract sentence for each semantic text block in the execution result output by the semantic chunking subtask through a second prompt to obtain a second abstract sentence set.

[0074] It should be noted that the above second prompt can, according to the prompting engineering construction instructions of the large language model, guide the large model to convert semantic text blocks into summary sentences. For example, the second prompt template can be constructed as: "The current input semantic text block is the text describing a certain topic in a telephone voice data. Please extract a summary sentence for this text. Next, I will give the semantic text block as follows: [Semantic text block placeholder]".

[0075] Similarly, due to the respective characteristics of call records in different dialogue types and scenarios, it is also possible to, based on the few-shot learning concept, based on the dialogue types and scenarios identified in advance from the target text, call the inference examples set for different dialogue types and scenarios in the memory bank, and add multiple exemplary cases of converting semantic text blocks into summary sentences in the above second prompt, so as to assist the large model in understanding how to convert the semantic text blocks in this dialogue type and scenario into summary sentences.

[0076] B3. The feedback optimization agent constructed based on the large language model performs consistency verification on the first summary sentence set and the second summary sentence set, and fuses and aligns the two sets through a multi-round self-check feedback strategy to generate a third summary sentence set as the execution result of the final output of the summary generation subtask.

[0077] It should be noted that the feedback optimization agent used in the above B3 step and the feedback optimization agent used in the above A3 step can be the same or different, and there is no limitation on this.

[0078] In the embodiments of the present invention, the execution tool specified and scheduled by the above event extraction subtask is an event extraction agent, such as Figure 4 shown, the internal processing process of the event extraction agent is as follows:

[0079] C1. Use a semantic role labeling tool to identify each summary sentence in the execution result (i.e., the third summary sentence set) output by the summary generation subtask, determine the predicate and the corresponding semantic arguments in the summary sentence, use the predicate as the trigger word, and the corresponding semantic arguments as event arguments, and output a first set of event elements composed of the trigger words and event arguments in all summary sentences.

[0080] It should be noted that the above semantic role labeling tool can be any tool or model that can label event elements such as predicates and semantic arguments. In the embodiments of the present invention, the above semantic role labeling tool can be implemented using AllenNLP, which can automatically identify the predicate of a sentence and label its corresponding semantic arguments. Thus, using the predicate as the trigger word and its corresponding semantic arguments as event arguments, event elements can be output.

[0081] C2. Guide the large language model to identify each summary sentence in the execution result output by the summary generation subtask through the third prompt word, and obtain a second set of event elements composed of event types, trigger words, event roles, and event arguments in all summary sentences.

[0082] It should be noted that the above third prompt word can guide the large model to chunk the dialogue text according to the prompt engineering construction instructions of the large language model. For example, the third prompt word template can be constructed as: "The current input summary sentence is a sentence describing a certain event. Please detect the event elements existing in this sentence. The event elements should include event types, trigger words, event roles, and argument topics. Next, I will give the following summary sentence: [summary sentence placeholder]".

[0083] In addition, due to the respective characteristics of call records in different dialogue types and scenarios, based on the few-shot learning concept, based on the dialogue types and scenarios identified in advance from the target text, call the inference examples set for different dialogue types and scenarios in the memory bank, and add multiple exemplary cases of extracting event elements from the summary sentence in the above third prompt word, so as to assist the large model to understand how to extract the required event elements from the summary sentence in this dialogue type and scenario.

[0084] C3. The feedback optimization agent constructed based on the large language model performs consistency verification on the trigger words and event arguments in the first set of event elements and the second set of event elements, and fuses and aligns the two sets through a multi-round self-check feedback strategy to generate a third set of event elements composed of event types, trigger words, event roles, and event arguments in all summary sentences as the execution result of the final output of the event extraction subtask.

[0085] It should be noted that the event elements extracted by the semantic role annotation tool in step C1 only include trigger words and arguments, and cannot generate event types and event roles. Although the large language model can extract complete event elements composed of event types, trigger words, event roles, and event arguments in step C2, it may produce "hallucination" results that are not faithful to the original text. Therefore, in this step, it is necessary to perform verification and alignment on it through the feedback optimization agent. The feedback optimization agent used in the above C3 step and the feedback optimization agent used in the above A3 step can be the same or different, and there is no limit to this.

[0086] In the embodiment of the present invention, the execution tool specified and scheduled by the above event relationship extraction subtask is an event relationship extraction agent, and the internal processing process of the event relationship extraction agent is as follows:

[0087] D1. Take each semantic text block in the execution result of the semantic chunking subtask (i.e., the third set of event elements) as an independent event, take the summary sentence corresponding to this semantic text block in the execution result of the summary generation subtask as the main sentence of the event, encode the main sentence of the event into an event vector representation through a language representation model, construct an event graph with all events as nodes in the form of an undirected fully connected graph, and perform node-to-node propagation and update on the original event vector representation of each event node through a graph attention network to obtain the final event vector representation of each event node.

[0088] In an embodiment of the present invention, since the events in each event node are extracted from a semantic text block, the encoded vector of its summary sentence can be used as the node representation of the event node, that is, the summary sentence is input into a pre-trained BERT network, and the [CLS] vector output is taken as the vector representation of the corresponding event vector representation . After obtaining the vector representations of all events, an undirected fully connected graph can be constructed, where the nodes are events , and an edge connection needs to be preset for each pair of events . The initially constructed event graph is input into a graph attention network (GAT) to update the node features in multiple rounds. In each round of attention calculation, the system splices the feature vectors of each pair of connected nodes and sends them into a scoring function with an activation function to generate attention coefficients, and normalizes the attention scores of the neighbors of the same node through the softmax function. Further, a multi-head attention mechanism is used for parallel propagation of multiple feature channels to improve the diversity and robustness of event semantic representations.

[0089] D2. Fuse the final event vector representations of each pair of event nodes in the event graph into a combined vector, and input the combined vector into an event relationship classifier to obtain the event relationship between this pair of event nodes.

[0090] In an embodiment of the present invention, for any pair of event nodes , the vector representation finally output after passing through the graph attention network (i.e., the output of the graph attention network for and output and ) can be combined into an edge representation vector , and the edge representation vector can adopt the following combination strategy:

[0091]

[0092] The above combined vector is input into a relationship classifier (such as a multi-layer perceptron MLP), and the output is the pair of event nodes The event relationships among them. In the embodiments of the present invention, relationship modeling can be pre - carried out for all abstract events, and the identified event relationship tags can be divided into the following four categories: temporal relationship (occurring successively), causal relationship (cause → effect), parallel relationship (parallel events), and inclusion relationship (sub - events).

[0093] In the embodiments of the present invention, the execution tool specified for the above - mentioned visual subtask scheduling is a triple data visualization tool, and the internal processing process is as follows: First, the format of the event extraction results obtained from other subtasks is normalized to form structured triple data, and then a graphical visualization tool is called to render the triple data to form a rendered event graph for display on the user interface. Specifically, the visualization process of the graph in the embodiments of the present invention is as follows:

[0094] 1) Uniformly organize the format of the event results after feedback optimization.

[0095] 2) Construct a directed graph or a graph database format (such as Neo4j, Graphviz data format). Among them, the nodes are event nodes containing event types and trigger words and independent argument nodes, and the edges are role edges of event nodes → argument nodes and event - to - event relationship edges.

[0096] 3) Call a graphical visualization tool (such as D3.js, Cytoscape.js, ECharts, Graphviz, etc.) to render the event graph and display it on the user interface.

[0097] In addition, when specifically rendering the graph in the embodiments of the present invention, a dynamic color - coding technology can be introduced to display the specific information of each event, which specifically includes the following three aspects:

[0098] Map to a gradient color based on event confidence, and trigger an alarm mechanism when the confidence is lower than the threshold;

[0099] Topological layout optimization: Adopt a force - directed layout algorithm to avoid node overlap and maintain the visibility of the critical path;

[0100] Interactive focused browsing: Support zooming, dragging, and right - clicking to bring up the event details panel, and the details panel includes a timeline, semantic roles, and an associated evidence chain.

[0101] It should be noted that in the above implementation method of the present invention, each module is equivalent to constructing a general event extraction intelligent agent, which includes a task decomposition module, a task execution module, a memory module, a feedback and optimization module. Among them, the memory module stores the aforementioned memory bank, which can be divided into long-term memory and short-term memory. The long-term memory is a knowledge base for different conversation types and scenarios, providing necessary background knowledge for event extraction in different conversation types and scenarios. The short-term memory is used to cache task status management, task process results, conversation context, and abnormal situations. The feedback and optimization module stores the feedback optimization intelligent agents used in each step.

[0102] It should be noted that the method steps shown in S1~S3 above can essentially be implemented in the form of a computer program.

[0103] Therefore, based on the same inventive concept, a system for extracting and visualizing multi-level events from telephone audio data is also provided, as Figure 5 shown. The system includes:

[0104] An audio text conversion module, which is used to convert the target telephone audio data into text conversation data as the target text for event extraction.

[0105] A task decomposition module, which is used to identify the conversation type and scenario of the target text, generate event extraction targets, and decompose the event extraction task into a sub-task process composed of semantic chunking sub-tasks, summary generation sub-tasks, event extraction sub-tasks, event relationship extraction sub-tasks, and visualization sub-tasks. At the same time, it specifies the execution tools to be scheduled for each sub-task.

[0106] A task execution module, which is used to sequentially execute each sub-task according to the planned sub-task process by scheduling the pre-specified execution tools and the knowledge base corresponding to the current conversation type and scenario, and perform semantic chunking, summary generation, event extraction, event relationship extraction, and visualization in sequence. Among them, in the semantic chunking sub-task, summary generation sub-task, and event extraction sub-task, the execution tool needs to call the large language model and the benchmark method corresponding to the current sub-task in a double-branch form to generate the execution results of the current sub-task respectively. Then, the feedback optimization intelligent agent based on the large language model performs consistency verification on the execution results of the two branches, and fuses and aligns the execution results of the two branches through a multi-round self-check feedback strategy to generate the final execution result of each sub-task. In the event relationship extraction sub-task, the execution tool needs to construct an event graph with all the events extracted by the event extraction sub-task as nodes, and classify the event relationships between the events through a graph classification algorithm. Finally, the execution tool of the visualization sub-task visualizes the event graph.

[0107] Thus, based on the same inventive concept, asFigure 6 As shown in Figure 6 , the present invention also provides a computer electronic device corresponding to the method for extracting multi-level events from telephone audio data and visualizing provided in the above embodiments. It includes a memory and a processor;

[0108] The memory is used to store computer programs;

[0109] The processor is used to implement the method for extracting multi-level events from telephone audio data and visualizing as described above when executing the computer program;

[0110] In addition, when the logical instructions in the above memory are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention.

[0111] In addition, based on the same inventive concept, the present invention provides a computer-readable storage medium corresponding to the method for extracting multi-level events from telephone audio data and visualizing. A computer program is stored on the storage medium. When the computer program is executed by a processor, it can implement the method for extracting multi-level events from telephone audio data and visualizing as described above.

[0112] Thus, based on the same inventive concept, the present invention provides a computer program product, including computer programs / instructions. When the computer programs / instructions are executed by a processor, they can implement the method for extracting multi-level events from telephone audio data and visualizing as described above.

[0113] Specifically, in the computer-readable storage media of the above three embodiments, the stored computer programs are executed by a processor, and the steps of S1 to S3 described above can be executed.

[0114] It can be understood that the above storage medium may include a Random Access Memory (RAM), and may also include a Non-Volatile Memory (NVM), such as at least one disk memory. At the same time, the storage medium may also be various media such as a USB flash drive, a mobile hard disk, a magnetic disk, or an optical disc that can store program codes.

[0115] It can be understood that the above-mentioned processor can be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.

[0116] In addition, it should be noted that those skilled in the art can clearly understand that for the convenience and simplicity of description, the specific working process of the above-described system can refer to the corresponding process in the foregoing method embodiments, and will not be elaborated herein. In the embodiments provided in the present application, the division of steps or modules in the system and method is only a logical function division, and there may be other division methods in actual implementation. For example, multiple modules or steps can be combined or integrated together, and a module or step can also be split.

[0117] The above-described embodiments are only some preferred implementation solutions of the present invention, but are not intended to limit the present invention. Those of ordinary skill in the relevant art can make various changes and modifications without departing from the spirit and scope of the present invention. Therefore, all technical solutions obtained by adopting equivalent replacement or equivalent transformation fall within the protection scope of the present invention.

Claims

1. A method for extracting multi-level events from telephone audio data and visualizing them, characterized in that, Including: S1. Convert the target telephone audio data into text conversation data as the target text of the event to be extracted; S2. Identify the conversation type and scenario of the target text, generate the event extraction target, and decompose the event extraction task into a sub-task process composed of semantic chunking sub-tasks, summary generation sub-tasks, event extraction sub-tasks, event relationship extraction sub-tasks, and visualization sub-tasks. At the same time, specify the execution tools to be scheduled for each sub-task; S3. According to the planned sub-task process, sequentially execute each sub-task by scheduling the pre-specified execution tools and the knowledge base corresponding to the current conversation type and scenario, and perform semantic chunking, summary generation, event extraction, event relationship extraction, and visualization in sequence; Among them, in the semantic chunking sub-task, summary generation sub-task, and event extraction sub-task, the execution tools are required to call the large language model and the benchmark method corresponding to the current sub-task in a two-branch form to generate the execution results of the current sub-task respectively. Then, the feedback optimization agent constructed based on the large language model performs consistency verification on the execution results of the two branches, and fuses and aligns the execution results of the two branches through a multi-round self-check feedback strategy to generate the final execution result of each sub-task; In the event relationship extraction sub-task, it is necessary to construct an event graph with all the events extracted by the event extraction sub-task as nodes through the execution tool, and classify the event relationships between the events through the graph classification algorithm. Finally, the execution tool of the visualization sub-task visualizes the event graph.

2. The method for extracting multi-level events from telephone audio data and visualizing them as claimed in claim 1, wherein The execution tool specified for scheduling in the semantic chunking sub-task is the semantic chunking agent, and the internal processing process of the semantic chunking agent is as follows: A1. Encode each statement in the target text into a vector representation, extract the average vector representation of multiple statements within the window through a sliding window of a fixed size, calculate the similarity of the average vector representations of adjacent sliding windows as the similarity between windows, select the local minimum value on the similarity sequence between windows as the candidate semantic deep valley point, and calculate the vector representation similarity of all adjacent statements in the statement sequence composed of all statements extracted from the two sliding windows within each candidate semantic deep valley point and take the maximum value as the mutation depth. If the mutation depth of a candidate semantic deep valley point exceeds the threshold, take the adjacent statement with the smallest vector representation similarity as the semantic segmentation point, otherwise discard the candidate semantic deep valley point; Divide the target text at all semantic segmentation points to obtain a series of semantic text blocks representing different independent topics, and obtain the first set of semantic text blocks; A2. Guide the large language model to divide the target text into semantic text blocks through the first prompt word to obtain the second set of semantic text blocks; A3. The feedback optimization agent constructed based on the large language model performs consistency verification on the first set of semantic text blocks and the second set of semantic text blocks, and fuses and aligns the two sets through a multi-round self-check feedback strategy to generate the third set of semantic text blocks as the final execution result output by the semantic chunking sub-task.

3. The method for extracting multi-level events from telephone audio data and visualizing them according to claim 1, characterized in that The execution tool specified for the abstract generation subtask scheduling is the abstract generation agent, and the internal processing process of the abstract generation agent is as follows: B1. Generate an abstract sentence for each semantic text block in the execution result output by the semantic chunking subtask through a pre-trained language model with an encoder-decoder architecture, and obtain a first set of abstract sentences; B2. Guide the large language model to generate an abstract sentence for each semantic text block in the execution result output by the semantic chunking subtask through a second prompt, and obtain a second set of abstract sentences; B3. The feedback optimization agent built based on the large language model performs consistency verification on the first set of abstract sentences and the second set of abstract sentences, and fuses and aligns the two sets through a multi-round self-check feedback strategy to generate a third set of abstract sentences as the execution result finally output by the abstract generation subtask.

4. The method for extracting multi-level events from telephone audio data and visualizing as claimed in claim 1, wherein The execution tool specified for the event extraction subtask scheduling is the event extraction agent, and the internal processing process of the event extraction agent is as follows: C1. Use a semantic role labeling tool to identify each abstract sentence in the execution result output by the abstract generation subtask, determine the predicate and the corresponding semantic arguments in the abstract sentence, use the predicate as the trigger word, and the corresponding semantic arguments as the event arguments, and output a first set of event elements composed of the trigger words and event arguments in all abstract sentences; C2. Guide the large language model to identify each abstract sentence in the execution result output by the abstract generation subtask through a third prompt, and obtain a second set of event elements composed of the event type, trigger word, event role, and event arguments in all abstract sentences; C3. The feedback optimization agent built based on the large language model performs consistency verification on the trigger words and event arguments in the first set of event elements and the second set of event elements, and fuses and aligns the two sets through a multi-round self-check feedback strategy to generate a third set of event elements composed of the event type, trigger word, event role, and event arguments in all abstract sentences as the execution result finally output by the event extraction subtask.

5. The method for extracting multi-level events from telephone audio data and visualizing them as claimed in claim 1, wherein The execution tool specified for the event relationship extraction subtask scheduling is the event relationship extraction agent, and the internal processing process of the event relationship extraction agent is as follows: D1. Take each semantic text block in the execution result output by the semantic chunking subtask as an independent event, take the abstract sentence corresponding to the semantic text block in the execution result output by the abstract generation subtask as the event main sentence, encode the event main sentence into an event vector representation through a language representation model, construct an event graph in the form of an undirected fully connected graph with all events as nodes, and perform node-to-node propagation and update on the original event vector representation of each event node through a graph attention network to obtain the final event vector representation of each event node; D2. Fuse the final event vector representations of each pair of event nodes in the event graph into a combined vector, and input the combined vector into an event relationship classifier to obtain the event relationship between this pair of event nodes.

6. The method for extracting multi-level events from telephone audio data and visualizing as claimed in claim 1, wherein, The execution tool specified for the visualization subtask scheduling is the triple data visualization tool, and the internal processing process is as follows: First, format the event extraction results obtained in other subtasks to form structured triple data, and then call a graphical visualization tool to render the triple data to form a rendered event graph for display on the user interface.

7. A system for extracting multi-level events from telephone audio data and visualizing them, characterized in that, Including: An audio-text conversion module for converting target telephone audio data into text dialogue data as the target text of the event to be extracted; A task decomposition module for identifying the dialogue type and scenario of the target text, generating event extraction targets, and decomposing the event extraction task into a subtask process consisting of semantic chunking subtasks, summary generation subtasks, event extraction subtasks, event relationship extraction subtasks, and visualization subtasks, and specifying the execution tools to be scheduled for each subtask; A task execution module for sequentially executing each subtask according to the planned subtask process by scheduling the pre-specified execution tools and the knowledge base corresponding to the current dialogue type and scenario, and performing semantic chunking, summary generation, event extraction, event relationship extraction, and visualization in turn; Among them, in the semantic chunking subtask, summary generation subtask, and event extraction subtask, the execution tool needs to call the large language model and the benchmark method corresponding to the current subtask in a two-branch form to generate the execution results of the current subtask respectively. Then, the feedback optimization agent constructed based on the large language model performs consistency verification on the execution results of the two branches, and fuses and aligns the execution results of the two branches through a multi-round self-check feedback strategy to generate the final execution result of each subtask; In the event relationship extraction subtask, the execution tool needs to construct an event graph with all the events extracted by the event extraction subtask as nodes, classify the event relationships between the events through a graph classification algorithm, and finally visualize the event graph by the execution tool of the visualization subtask.

8. A computer program product comprising a computer program / instructions, characterized in that, When the computer program / instructions are executed by the processor, it can implement the method of extracting multi-level events from telephone audio data and visualizing as described in any one of claims 1 to 6.

9. A computer-readable storage medium, characterized in that, A computer program is stored on the storage medium, and when the computer program is executed by the processor, it implements the method of extracting multi-level events from telephone audio data and visualizing as described in any one of claims 1 to 6.

10. A computer electronic device, characterized in that, Including a memory and a processor; The memory is used to store computer programs; The processor is used to implement the method of extracting multi-level events from telephone audio data and visualizing as described in any one of claims 1 to 6 when executing the computer program.

Citation Information

Patent Citations

  • Generation type and extraction type combined text abstract generation method

    CN115757762A

  • Smart home quality safety determination method based on affair knowledge graph

    CN119089065A