Intelligent extraction and abstract generation method and system for call content
By using automatic recording, speech recognition, and multi-model collaborative processing, the problem of low efficiency and insufficient accuracy in obtaining key points of a call has been solved in existing technologies, enabling the rapid and accurate generation of call summaries and customized meeting minutes.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-19
- Publication Date
- 2026-03-13
AI Technical Summary
Existing technologies are inefficient at capturing key points of a call and are prone to missing important information, resulting in insufficient accuracy.
By automatically recording and storing audio during calls, the system utilizes the Whisper model for speech recognition and Pyannote pipeline processing, combined with the BERT model for entity recognition, sentiment analysis, and intent recognition, to generate timestamped text and extract key information. A multi-channel action item extraction model is then used to generate a summary.
It enables the rapid and accurate acquisition of key points from calls, generating meeting minutes in a custom format, thus improving the efficiency and accuracy of information acquisition.
Smart Images

Figure CN121658644A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of information processing technology, and in particular to a method and system for intelligent extraction and summarization of call content. Background Technology
[0002] With the rapid development of smart devices, information is flooding in, and more and more phone calls are being made through them. Traditional methods for obtaining call information typically involve first recording the call and saving the audio file, then having someone listen to the recording and transcribe it. However, this method is prone to missing important information during long calls or when there are many calls.
[0003] Processing voice messages to quickly and accurately extract key information from conversations plays a crucial role in optimizing corporate services and enhancing corporate competitiveness. Therefore, there is an urgent need for technology that can quickly and accurately extract key information from conversations. Summary of the Invention
[0004] The present invention aims to solve the problems of slow time, easy omission and low accuracy in the existing methods of obtaining key points of call content, and provides a method and system for intelligent extraction and summary generation of call content.
[0005] This invention provides a method for intelligent extraction and summary generation of call content, comprising the following steps: Automatically trigger recording during a call to obtain the call audio; The call audio is stored in a preset user identification card, and a traceable index directory is generated; The call audio is obtained and extracted to a specified page through directory location; The call audio is input into a preset Whisper model for speech recognition within the designated page to obtain timestamped text; simultaneously, the call audio is processed through the Pyannote pipeline to obtain caller and caller segmentation information; wherein, the Pyannote pipeline includes VAD and speaker embedding model; The timestamped text, the caller and the caller segment information are merged, and key information is sorted and extracted to generate text with caller tags. The text with caller tags is sent to the adjusted BERT model, where entity recognition, sentiment analysis, and intent recognition are performed to obtain sentiment tendencies and target intent sentences. The target intent sentence is input into a multi-channel action item extraction model, and the action items and triples of the text with caller tags are extracted by the multi-channel action item extraction model; wherein, the multi-channel action item extraction model includes an action item extraction model, the BERT model and a triple extraction model; The triples and the timestamped text are processed in parallel using multiple preset algorithm models to generate a summary or structured document; The meeting minutes template is dynamically selected and adjusted, and the action items, the summary, the sentiment, and the target intent sentence are filled into the meeting minutes template to obtain the final meeting minutes, which are then saved and displayed on the designated page.
[0006] Furthermore, the step of inputting the call audio into a preset Whisper model for speech recognition within the designated page to obtain timestamped text, and simultaneously processing the call audio through the Pyannote pipeline to obtain caller and caller segmentation information, includes: Caller identification: Converts continuous sound waves into discrete text and distinguishes different callers through voiceprint features; Call content recognition: Based on the text, the caller is divided and organized into segments to obtain segment information of the caller; wherein, the segment information of the caller includes an opening segment, a statement segment, and a closing statement.
[0007] Furthermore, the algorithm model includes a speech recognition model, a caller separation and recognition model, an understanding model, and an extraction model; The speech recognition model includes a deep neural network model based on a recurrent neural network or a Transformer architecture; The caller separation and identification model includes a VAD model and a speaker embedding model; The understanding model includes a basic natural language understanding model and a pre-trained language model based on the Transformer architecture; The extraction models include Transformer-based sequence labeling models and Seq2Seq models.
[0008] Furthermore, sentiment analysis and intent recognition are performed on the opening paragraph, the statement paragraph, and the closing remarks, respectively. The steps of performing sentiment analysis and intent recognition on the opening paragraph, the statement paragraph, and the closing remarks respectively include: The meaning of the opening paragraph, the statement paragraph, and the conclusion is understood using the aforementioned understanding model; Analyze the grammatical structure and semantic roles of the opening paragraph, the statement paragraph, and the conclusion; Determine the attitude attributes of the caller and identify their intent to obtain the target intent sentence.
[0009] Furthermore, the step of sending the text with caller tags to the adjusted BERT model and performing entity recognition, sentiment analysis, and intent recognition to obtain the sentiment tendency and target intent sentence includes: The intent recognition and classification are performed on the text containing the caller's tag; Extract the corresponding text based on the preset target intent to obtain the target intent sentence; The entity recognition includes person recognition, time recognition, and item recognition.
[0010] Furthermore, the step of processing the triples and the timestamped text in parallel using multiple preset algorithm models to generate a summary or structured document includes: The system outputs contextual semantics through a pre-defined semantic understanding model and infers the missing elements in the triples; wherein the triples include the task, the person in charge, and the deadline. Complete the missing elements in the triplet.
[0011] Furthermore, the step of processing the triples and the timestamped text in parallel using multiple preset algorithm models and generating a summary further includes: The content of the summary is determined to be consistent with the semantics of the triple by a preset alignment loss function.
[0012] This invention also provides an intelligent system for extracting and summarizing call content, comprising: The recording module is used to automatically trigger recording during a call to obtain the call audio; The storage module is used to store the call audio into a preset user identification card and generate a traceable index directory; The call audio extraction module is used to acquire the call audio and extract the call audio to a specified page through directory location; The speech recognition module is used to input the call audio into a preset Whisper model for speech recognition within the specified page, and obtain text with timestamps. The audio processing module is used to process the call audio through the Pyannote pipeline to obtain the caller and caller segmentation information; The information merging module is used to merge the timestamped text, the caller and the segmented information of the caller, and to organize and extract key information to generate text with caller tags; The analysis module is used to send the text with caller tags to the adjusted BERT model and perform entity recognition, sentiment analysis and intent recognition to obtain sentiment tendency results and target intent sentences; The action item extraction module is used to input the target intent sentence into the multi-channel action item extraction model, and extract the action items and triples of the text with caller tags through the multi-channel action item extraction model; The parallel processing module is used to process the triples and the timestamped text in parallel using multiple preset algorithm models, and generate a summary or structured document. The summary module is used to dynamically select and adjust the meeting minutes template, and fill the action items, the summary, the sentiment, and the target intent into the meeting minutes template to obtain the final meeting minutes, which are then saved and displayed on a designated page.
[0013] The present invention also provides a computer device, including a memory and a processor, wherein the memory stores a computer program and the processor executes the computer program to implement the steps in any of the above methods.
[0014] The present invention also provides a computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the steps of any of the above methods.
[0015] This invention provides a method and system for intelligent extraction and summarization of call content, which has the following beneficial effects: This application processes call audio from smart devices, converting speech into timestamped text, segmented information about the caller and the caller, and then organizing and extracting key information to generate text tagged with the caller. This tagged text is then intelligently recognized to extract important information, obtaining sentiment and target intent sentences, ultimately resulting in a final meeting summary. Through multi-model collaboration and parallel processing, call audio can be processed quickly, rapidly extracting key points and saving time. Furthermore, based on the results of "understanding," key information is extracted—that is, after entity recognition, sentiment analysis, and intent recognition to obtain sentiment and target intent sentences, action items and triples are extracted, resulting in higher accuracy. After obtaining the summary, the meeting summary template can be dynamically selected and adjusted to generate a final meeting summary in a custom format. Attached Figure Description
[0016] Figure 1 This is a schematic diagram of the method steps for intelligent extraction and summary generation of call content in this invention; Figure 2 This is a schematic diagram illustrating the steps of an embodiment of a method for intelligent extraction and summary generation of call content according to the present invention. Figure 3 This is a structural block diagram of a call content intelligent extraction and summary generation system according to the present invention; Figure 4 This is a structural block diagram of a computer device according to the present invention.
[0017] Labeling description: Recording module 10, storage module 20, call audio extraction module 30, speech recognition module 40, audio processing module 50, caller identification unit 501, call content recognition unit 502, emotion analysis unit 503, understanding unit 504, syntax analysis unit 505, first judgment unit 506, information merging module 60, analysis module 70, intent processing unit 701, extraction unit 702, action item extraction module 80, parallel processing module 90, reasoning unit 901, completion unit 902, second judgment unit 903, summary module 110. Detailed Implementation
[0018] It should be understood that the specific embodiments described herein are merely illustrative of the invention and are not intended to limit the invention.
[0019] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, and not all of them. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0020] Reference Appendix Figure 1 The present invention provides a method for intelligent extraction and summary generation of call content, comprising the following steps: S1, automatically triggers recording during the call to obtain the call audio; S2, store the call audio into a preset user identification card and generate a traceable index directory; S3, obtain the call audio and extract it to the specified page by locating it in the directory; S4. Input the call audio into the preset Whisper model for speech recognition within the specified page to obtain time-stamped text; at the same time, process the call audio through the Pyannote pipeline to obtain the caller and caller segmentation information; the Pyannote pipeline includes VAD and speaker embedding model. S5 merges the timestamped text, caller and caller segment information, organizes and extracts key information, and generates text with caller tags; S6, send the text with caller tags to the adjusted BERT model, and perform entity recognition, sentiment analysis and intent recognition to obtain sentiment tendency and target intent sentence; S7, the target intent sentence is input into the multi-channel action item extraction model, and the action items and triples of the text with caller tags are extracted by the multi-channel action item extraction model; the multi-channel action item extraction model includes an action item extraction model, a BERT model and a triple extraction model; S8 processes triples and timestamped text in parallel using multiple preset algorithm models and generates summaries or structured documents. S9 allows for dynamic selection and adjustment of meeting minutes templates, and the filling of action items, summaries, sentiments, and target intent sentences into the meeting minutes templates to obtain the final meeting minutes, which are then saved and displayed on the specified page.
[0021] In the above steps, the recording module 10 of the smart device first automatically triggers recording during the call to obtain the call audio. After obtaining the call audio, it is stored in a preset user identification card, and a traceable index directory is generated. The call audio is stored in the index directory. The user identification card is the operator's SIM card (Subscriber Identity Module). The index directory is the storage location of the call audio and is automatically generated. In a specific embodiment, the index directory is C: / recorder / ai / . Next, the call audio is obtained and extracted to a specified page through directory location. Specifically, after obtaining the call audio, the storage location of the call audio is located through directory location, and then the call audio is extracted to a specified page, such as the display page. Then, the call audio is processed simultaneously through two paths: on the specified page, the call audio is input into a preset Whisper model for speech recognition to obtain time-stamped text; at the same time, the call audio is processed through the Pyannote pipeline to obtain the caller and caller segment information; the Whisper model is a speech recognition model capable of multilingual recognition, speech translation, and language recognition. After speech recognition of the call audio using the Whisper model, timestamped text is obtained. Pyannote is a sound processing tool, and its pipeline includes VAD (Voice Activity Detection) and a speaker embedding model. Specifically, after processing the call audio using VAD and the speaker embedding model, caller and caller segment information are obtained. Then, the timestamped text, caller and caller segment information are merged, organized, and key information is extracted to generate text tagged with the caller. This text tagged with the caller is sent to an adjusted BERT (Bidirectional Encoder Representations from Transformers) model for entity recognition, sentiment analysis, and intent recognition to obtain sentiment tendency and target intent sentence. Entity recognition includes identifying the caller, time, and item; sentiment analysis determines the caller's attitude; and intent recognition identifies the intent of the text tagged with the caller and categorizes the text according to intent, ultimately obtaining the target intent sentence based on a preset target intent. In one specific embodiment, the preset target intent may include a question, commitment, objection, suggestion, and summary.After obtaining the target intent sentence, such as "commitment," it is input into a multi-channel action item extraction model to extract action items and triples from the text labeled with the caller's name. The multi-channel action item extraction model includes an action extraction model, a BERT (Bidirectional Encoder Representations from Transformers) model, and a relation extraction model. Specifically, the action extraction model identifies task instructions from the text using Seq2Seq or Token Classification. The BERT model performs intent and semantic understanding, providing contextual understanding and semantic support for action items. Triples include the task, the responsible person, and the deadline. The relation extraction model clarifies the correspondence between the task, the responsible person, and the deadline. The action extraction model, the BERT model, and the relation extraction model together constitute the semantic channel, each responsible for extracting and completing different dimensions, and are finally combined for output.
[0022] More specifically, traditional action item extraction models can only extract explicit information, such as "submit the report tomorrow," but they cannot extract implicit information, such as "the report is almost finished," where the processing capability for the implicit action item "complete the report as soon as possible" is weak. This application integrates the BERT model for deep semantic understanding on the basis of action item extraction models, enabling the model to complete based on context.
[0023] Next, multiple pre-defined algorithm models are used to process triples and timestamped text in parallel, generating summaries or other structured documents, such as reports. Specifically, intelligent recognition, core understanding, and task extraction are performed through multiple algorithm models. These models include speech recognition, speaker separation and recognition, understanding, and extraction models. Speech recognition models include deep neural network models based on recurrent neural networks (RNN) or Transformer architectures. Speaker separation and recognition models include VAD (Voice Activity Detection) models and speaker embedding models. Understanding models include basic natural language understanding models and pre-trained language models based on Transformer architectures. The basic natural language understanding models can be CLIP (Contrastive Language-Image Pre-training) or SAM (SegmentAnything Model). Extraction models include Transformer-based sequence labeling models and Seq2Seq models. Sequence labeling models include HMM and Hidden Markov Models. After intelligent recognition, core understanding, and task extraction by multiple algorithm models, a summary is generated. Next, dynamically select and adjust the meeting minutes template to suit different call types, such as sales calls, customer support, and project meetings. Then, fill the meeting minutes template with action items, summaries, sentiments, and target intent sentences to obtain the final meeting minutes, which are then saved and displayed on the designated page.
[0024] In one specific embodiment, such as Figure 2As shown, phone calls are made using a carrier SIM card installed on the phone, and the call is recorded during the call. After recording ends, the call audio is stored on the SIM card, and a traceable index directory is generated. Next, the call audio is retrieved and extracted to a designated page using the directory location. Intelligent recognition is performed on the caller and the recording content to identify the caller and extract key information. Specifically, the call audio is input into a preset Whisper model for speech recognition within the designated page, resulting in timestamped text. Simultaneously, the call audio is processed through a Pyannote pipeline to obtain the caller and caller segmentation information. The timestamped text, caller and caller segment information are merged, and key information is extracted to generate caller-tagged text. This caller-tagged text is then sent to an adjusted BERT model for entity recognition, sentiment analysis, and intent recognition to obtain sentiment tendency and target intent sentence. The target intent sentence is input into a multi-channel action item extraction model, which extracts action items and triples from the caller-tagged text. The multi-channel action item extraction model includes an action item extraction model, a BERT model, and a triple extraction model. Multiple pre-defined algorithm models process the triples and timestamped text in parallel and output a textual summary. The action items, summary, sentiment tendency, and target intent sentence are then filled into a meeting minutes template to obtain the final meeting minutes, which are then saved.
[0025] In one embodiment, the step of inputting call audio into a preset Whisper model for speech recognition within a designated page to obtain timestamped text, and simultaneously processing the call audio through a Pyannote pipeline to obtain caller and caller segmentation information, includes: Caller identification: Converts continuous sound waves into discrete text and distinguishes different callers through voiceprint features; Call content recognition: Based on the text, the caller is segmented and organized to obtain segment information of the caller; the segment information of the caller includes the opening paragraph, the statement paragraph and the closing remarks.
[0026] In this embodiment, caller identification and call content identification are performed on the call audio. Caller identification specifically involves converting continuous sound waves into discrete text and distinguishing different callers using voiceprint features. Voiceprint refers to the sound wave spectrum carrying speech information, displayed using electroacoustic instruments. Call content identification specifically involves segmenting and organizing the text into paragraphs to obtain caller segment information; this segment information includes an opening paragraph, a statement paragraph, and a closing remark. The text is structurally segmented and organized, removing redundancy to prevent it from becoming a continuous, disorganized stream of text.
[0027] In one embodiment, the algorithm model includes a speech recognition model, a caller separation and recognition model, an understanding model, and an extraction model; Speech recognition models include deep neural network models based on recurrent neural networks or Transformer architectures; Caller separation and identification models include VAD models and speaker embedding models; Understanding models include basic natural language understanding models and pre-trained language models based on the Transformer architecture; The extraction models include Transformer-based sequence labeling models and Seq2Seq models.
[0028] In this embodiment, the algorithm model includes a speech recognition model, a speaker separation and recognition model, an understanding model, and an extraction model. The speech recognition model includes a deep neural network model based on a recurrent neural network or a Transformer architecture. The speaker separation and recognition model includes a VAD model and a speaker embedding model. The VAD model can extract features and distinguish between speech and non-speech parts. The speaker embedding model can perform speaker recognition, i.e., determine which speaker is making an audio clip. The understanding model includes a basic natural language understanding model and a pre-trained language model based on a Transformer architecture. The extraction model includes a Transformer-based sequence labeling model and a Seq2Seq model.
[0029] In one embodiment, sentiment analysis and intent recognition are performed on the opening paragraph, the statement paragraph, and the closing paragraph, respectively.
[0030] In one specific embodiment, after segmenting the call audio, the opening segment is a greeting between the merchant and the consumer; the statement segment is the consumer describing their problem or need, followed by the merchant's explanation and coordination; and the closing segment is the merchant's expression of gratitude. A preset sentiment analysis model, such as the BERT model, is used to perform sentiment analysis on the opening segment, statement segment, and closing segment respectively. In another specific embodiment, if the opening segment is "Hello, it's a pleasure to serve you," the sentiment analysis will determine it to be "positive." However, if the statement segment is "Why haven't you replied for so long? I'm very unhappy," the sentiment analysis will determine it to be "negative."
[0031] In one embodiment, the steps of performing sentiment analysis and intent recognition on the opening paragraph, the statement paragraph, and the closing remarks respectively include: Understand the meaning of the opening paragraph, the statement paragraph, and the conclusion by understanding the model; Analyze the grammatical structure and semantic roles of the opening paragraph, the declarative paragraph, and the concluding remarks; Determine the speaker's attitude and identify their intent to obtain the target intent sentence.
[0032] In one specific embodiment, the opening paragraph is "Hello, what do you need?"; the closing paragraph is the customer's response, "I need to know the status of my order; I've been waiting for several days"; and the closing paragraph is "Okay, thank you for your call." Next, the meaning of the opening, closing, and closing paragraphs is understood using an understanding model. Then, the grammatical structure and semantic roles of the opening, closing, and closing paragraphs are analyzed to determine the caller's attitude and identify the intent, resulting in the target intent sentence. For example, "I need to know the status of my order; I've been waiting for several days" indicates a negative attitude and an intent to check the status, resulting in the target intent sentence "Know the order status."
[0033] In one embodiment, the timestamped text, caller and caller segment information are merged, sorted and key information is extracted to generate text with caller tags.
[0034] In this embodiment, key information is extracted from the timestamped text, the caller, and the caller segmentation information based on the target intent sentence. Specifically, key information such as time, question, task, or commitment can be extracted using TF-IDF and TextRank algorithms. After processing, the text with caller tags is obtained. For example, the speech content is segmented using timestamps, and each key segment is listed in chronological order, with the caller identified.
[0035] In one embodiment, the step of sending text with caller tags to an adjusted BERT model and performing entity recognition, sentiment analysis, and intent recognition to obtain sentiment tendency and target intent sentence includes: Perform intent recognition and classification on text with caller tags; Extract the corresponding text based on the preset target intent to obtain the target intent sentence; Entity recognition includes person recognition, time recognition, and item recognition.
[0036] In this embodiment, the text with caller tags is sent to the adjusted BERT model for entity recognition. Entity recognition includes caller identification, time identification, and item identification. Specifically, all callers are identified from the caller tags and named accordingly. Time identification identifies time expressions, such as "next Friday" or "this afternoon." Item identification identifies event information, such as "meeting" or "submitting a report." The BERT model performs intent recognition and classification on the text with caller tags, such as "check logistics progress," "provide solutions," or "fulfill exchange commitments." Classification can be categorized into after-sales categories, for example, "provide solutions" and "fulfill exchange commitments" are classified as after-sales. The corresponding text is extracted based on the preset target intent to obtain the target intent sentence. In a specific embodiment, if the target intent is a commitment, the extracted target intent sentence is "fulfill exchange commitments."
[0037] In one embodiment, the step of processing triples and timestamped text in parallel using multiple preset algorithm models to generate a summary or structured document includes: The system outputs contextual semantics through a pre-defined semantic understanding model and infers the missing elements in the triples; the triples include the task, the person in charge, and the deadline. Complete the missing elements in the triplet.
[0038] In this embodiment, when the extracted triple is missing "responsible person" or "deadline," the semantic understanding model outputs contextual semantics and infers the missing element in the triple. In one specific embodiment, the default responsible person is inferred based on the task allocation mentioned above. In another specific embodiment, the deadline is inferred based on the timeline of the discussion, such as "before the meeting next week," combined with the time entity identified by entity recognition. For the missing element that is completed, it can be marked in the generated final meeting minutes, such as "Inferred responsible person: Zhang San," to prompt the user for review.
[0039] In one embodiment, the step of processing triples and timestamped text in parallel using multiple preset algorithm models to generate a summary further includes: The content of the summary is determined to be consistent with the semantics of the triple by a preset alignment loss function.
[0040] In this embodiment, the alignment loss function is used to evaluate the semantic consistency between the generated summary and the triples. Specifically, by calculating the cosine similarity between the generated summary and the extracted triples, it is determined whether the generated summary accurately reflects the task, responsible person, and deadline in the triples, thereby improving the accuracy of the generated summary.
[0041] Reference Appendix Figure 3 A system for intelligent extraction and summarization of call content, comprising: The recording module 10 is used to automatically trigger recording during a call to obtain the call audio; Storage module 20 is used to store call audio into a preset user identification card and generate a traceable index directory; The call audio extraction module 30 is used to acquire call audio and extract the call audio to a specified page through directory location; The speech recognition module 40 is used to input the call audio into the preset Whisper model for speech recognition within the specified page, and obtain the text with timestamps. The audio processing module 50 is used to process the call audio through the Pyannote pipeline to obtain the caller and caller segmentation information. The caller identification unit 501 is used to convert continuous sound waves into discrete text and distinguish different callers through voiceprint features. The call content recognition unit 502 is used to divide and organize the text into segments to obtain the segment information of the caller. The sentiment analysis unit 503 is used to perform sentiment analysis and intent recognition on the opening paragraph, the statement paragraph and the closing remarks respectively; Unit 504 is used to understand the meaning of the opening paragraph, the statement paragraph, and the conclusion through the understanding model; Syntax analysis unit 505 is used to analyze the grammatical structure and semantic roles of the opening paragraph, the statement paragraph, and the closing paragraph; The first judgment unit 506 is used to judge the attitude attribute of the caller and identify the intention to obtain the target intention sentence; The information merging module 60 is used to merge the timestamped text, caller and caller segment information, and to organize and extract key information to generate text with caller tags. Analysis module 70 is used to send text with caller tags to the adjusted BERT model and perform entity recognition, sentiment analysis and intent recognition to obtain sentiment tendency results and target intent sentences; The intent processing unit 701 is used to perform intent recognition and classification on text with caller tags; The extraction unit 702 is used to extract the corresponding text according to the preset target intent to obtain the target intent sentence; The action item extraction module 80 is used to input the target intent sentence into the multi-channel action item extraction model, and extract the action items and triples with caller-tagged text through the multi-channel action item extraction model; Parallel processing module 90 is used to process triples and timestamped text in parallel using multiple preset algorithm models, and generate summaries or structured documents. The reasoning unit 901 is used to output contextual semantics through a preset semantic understanding model and infer the missing elements in the triples; Completion unit 902 is used to complete missing elements in triples; The second judgment unit 903 is used to determine whether the content of the summary is consistent with the semantics of the triples through a preset alignment loss function; The summary module 110 is used to dynamically select and adjust the meeting minutes template, and fill the meeting minutes template with action items, summaries, sentiments and target intent sentences to obtain the final meeting minutes, which are then saved and displayed on the specified page.
[0042] In this embodiment, the recording module 10 first automatically triggers recording during the call to obtain the call audio; then, the storage module 20 stores the call audio in a preset user identification card and generates a traceable index directory; next, the call audio extraction module 30 obtains the call audio and extracts it to a specified page through directory location; then, the speech recognition module 40 inputs the call audio into a preset Whisper model for speech recognition within the specified page to obtain timestamped text; the audio processing module 50 processes the call audio through the Pyannote pipeline to obtain the caller and caller segmentation information; the caller identification unit 501 connects... The continuous sound waves are converted into discrete text, and different callers are distinguished by voiceprint features. During this process, the call content recognition unit 502 divides and organizes the text into segments to obtain segmented information of the callers. The sentiment analysis unit 503 performs sentiment analysis and intent recognition on the opening paragraph, statement paragraph, and closing remarks respectively. The understanding unit 504 understands the meaning of the opening paragraph, statement paragraph, and closing remarks through an understanding model. The grammar analysis unit 505 analyzes the grammatical structure and semantic roles of the opening paragraph, statement paragraph, and closing remarks. The first judgment unit 506 judges the caller's attitude attributes and identifies the intent to obtain the target intent sentence. The information merging module 60 combines the time-bound... The interstamped text, caller and caller segment information are merged, sorted and key information is extracted to generate text with caller tags; the analysis module 70 sends the text with caller tags to the adjusted BERT model and performs entity recognition, sentiment analysis and intent recognition to obtain sentiment tendency results and target intent sentences; the intent processing unit 701 performs intent recognition and classification on the text with caller tags; the extraction unit 702 extracts the corresponding text according to the preset target intent to obtain the target intent sentence; the action item extraction module 80 inputs the target intent sentence into the multi-channel action item extraction model, and extracts the action items of the text with caller tags through the multi-channel action item extraction model. The triplet; the parallel processing module 90 processes the triplet and timestamped text in parallel using multiple preset algorithm models, and generates a summary or structured document; the inference unit 901 outputs contextual semantics through a preset semantic understanding model and infers the missing elements in the triplet; the completion unit 902 completes the missing elements in the triplet; the second judgment unit 903 judges whether the content of the summary is consistent with the semantics of the triplet through a preset alignment loss function; the summary module 110 dynamically selects and adjusts the meeting minutes template, and fills the action items, summary, sentiment and target intent sentences into the meeting minutes template to obtain the final meeting minutes, which are then saved and presented on the specified page.
[0043] This application also provides a computer device, which may be a server, and its internal structure may be as follows: Figure 4As shown. The computer device includes a processor, memory, network interface, and database. The processor provides computing and control capabilities. The memory includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores the operating system, computer programs, and database. The internal memory stores the operating system, computer programs, and database in the non-volatile storage medium. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The database stores templates, tables, preset fields, and other data. The network interface is used to communicate with external terminals via a network connection. When the computer program is executed by the processor, it implements a method for intelligent extraction and summary generation of call content, including the following steps: Automatically trigger recording during a call to obtain the call audio; The call audio is stored in a preset user identification card, and a traceable index directory is generated; Get the call audio and extract it to a specified page by locating it in the directory; The call audio is input into a preset Whisper model for speech recognition within a specified page to obtain timestamped text; at the same time, the call audio is processed through the Pyannote pipeline to obtain caller and caller segmentation information; the Pyannote pipeline includes VAD and speaker embedding model; The timestamped text, caller and caller segment information are merged, and key information is sorted and extracted to generate text with caller tags. The text with caller tags is sent to the adjusted BERT model, where entity recognition, sentiment analysis, and intent recognition are performed to obtain the sentiment tendency and target intent sentence. The target intent sentence is input into a multi-channel action item extraction model, and the action items and triples of the caller-tagged text are extracted by the multi-channel action item extraction model; the multi-channel action item extraction model includes an action item extraction model, a BERT model and a triple extraction model; Multiple pre-defined algorithm models are used to process triples and timestamped text in parallel and generate summaries. Dynamically select and adjust the meeting minutes template, and fill in the action items, summary, sentiment, and target intent sentences into the meeting minutes template to obtain the final meeting minutes, which are then saved and displayed on the specified page.
[0044] In one embodiment, the step of inputting call audio into a preset Whisper model for speech recognition within a designated page to obtain timestamped text, and simultaneously processing the call audio through a Pyannote pipeline to obtain caller and caller segmentation information, includes: Caller identification: Converts continuous sound waves into discrete text and distinguishes different callers through voiceprint features; Call content recognition: Based on the text, the caller is segmented and organized to obtain segment information of the caller; the segment information of the caller includes the opening paragraph, the statement paragraph and the closing remarks.
[0045] In one embodiment, the algorithm model includes a speech recognition model, a caller separation and recognition model, an understanding model, and an extraction model; Speech recognition models include deep neural network models based on recurrent neural networks or Transformer architectures; Caller separation and identification models include VAD models and speaker embedding models; Understanding models include basic natural language understanding models and pre-trained language models based on the Transformer architecture; The extraction models include Transformer-based sequence labeling models and Seq2Seq models.
[0046] In one embodiment, sentiment analysis and intent recognition are performed on the opening paragraph, the statement paragraph, and the closing paragraph, respectively.
[0047] In one embodiment, the steps of performing sentiment analysis and intent recognition on the opening paragraph, the statement paragraph, and the closing remarks respectively include: Understand the meaning of the opening paragraph, the statement paragraph, and the conclusion by understanding the model; Analyze the grammatical structure and semantic roles of the opening paragraph, the declarative paragraph, and the concluding remarks; Determine the speaker's attitude and identify their intent to obtain the target intent sentence.
[0048] In one embodiment, the step of sending text with caller tags to an adjusted BERT model and performing entity recognition, sentiment analysis, and intent recognition to obtain sentiment tendency and target intent sentence includes: Perform intent recognition and classification on text with caller tags; Extract the corresponding text based on the preset target intent to obtain the target intent sentence; Entity recognition includes person recognition, time recognition, and item recognition.
[0049] In one embodiment, the step of processing triples and timestamped text in parallel using multiple preset algorithm models to generate a summary or structured document includes: The system outputs contextual semantics through a pre-defined semantic understanding model and infers the missing elements in the triples; the triples include the task, the person in charge, and the deadline. Complete the missing elements in the triplet.
[0050] In one embodiment, the step of processing triples and timestamped text in parallel using multiple preset algorithm models to generate a summary further includes: The content of the summary is determined to be consistent with the semantics of the triple by a preset alignment loss function.
[0051] Those skilled in the art will understand that the structure shown in the figure is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer equipment on which the present application is applied.
[0052] One embodiment of this application also provides a computer storage medium storing a computer program thereon. When the computer program is executed by a processor, it implements a method for intelligent extraction and summary generation of call content, including the following steps: Automatically trigger recording during a call to obtain the call audio; The call audio is stored in a preset user identification card, and a traceable index directory is generated; Get the call audio and extract it to a specified page by locating it in the directory; The call audio is input into a preset Whisper model for speech recognition within a specified page to obtain timestamped text; at the same time, the call audio is processed through the Pyannote pipeline to obtain caller and caller segmentation information; the Pyannote pipeline includes VAD and speaker embedding model; The timestamped text, caller and caller segment information are merged, and key information is sorted and extracted to generate text with caller tags. The text with caller tags is sent to the adjusted BERT model, where entity recognition, sentiment analysis, and intent recognition are performed to obtain the sentiment tendency and target intent sentence. The target intent sentence is input into a multi-channel action item extraction model, and the action items and triples of the caller-tagged text are extracted by the multi-channel action item extraction model; the multi-channel action item extraction model includes an action item extraction model, a BERT model and a triple extraction model; Multiple pre-defined algorithm models are used to process triples and timestamped text in parallel and generate summaries. Dynamically select and adjust the meeting minutes template, and fill in the action items, summary, sentiment, and target intent sentences into the meeting minutes template to obtain the final meeting minutes, which are then saved and displayed on the specified page.
[0053] In one embodiment, the step of inputting call audio into a preset Whisper model for speech recognition within a designated page to obtain timestamped text, and simultaneously processing the call audio through a Pyannote pipeline to obtain caller and caller segmentation information, includes: Caller identification: Converts continuous sound waves into discrete text and distinguishes different callers through voiceprint features; Call content recognition: Based on the text, the caller is segmented and organized to obtain segment information of the caller; the segment information of the caller includes the opening paragraph, the statement paragraph and the closing remarks.
[0054] In one embodiment, the algorithm model includes a speech recognition model, a caller separation and recognition model, an understanding model, and an extraction model; Speech recognition models include deep neural network models based on recurrent neural networks or Transformer architectures; Caller separation and identification models include VAD models and speaker embedding models; Understanding models include basic natural language understanding models and pre-trained language models based on the Transformer architecture; The extraction models include Transformer-based sequence labeling models and Seq2Seq models.
[0055] In one embodiment, sentiment analysis and intent recognition are performed on the opening paragraph, the statement paragraph, and the closing paragraph, respectively.
[0056] In one embodiment, the steps of performing sentiment analysis and intent recognition on the opening paragraph, the statement paragraph, and the closing remarks respectively include: Understand the meaning of the opening paragraph, the statement paragraph, and the conclusion by understanding the model; Analyze the grammatical structure and semantic roles of the opening paragraph, the declarative paragraph, and the concluding remarks; Determine the speaker's attitude and identify their intent to obtain the target intent sentence.
[0057] In one embodiment, the step of sending text with caller tags to an adjusted BERT model and performing entity recognition, sentiment analysis, and intent recognition to obtain sentiment tendency and target intent sentence includes: Perform intent recognition and classification on text with caller tags; Extract the corresponding text based on the preset target intent to obtain the target intent sentence; Entity recognition includes person recognition, time recognition, and item recognition.
[0058] In one embodiment, the step of processing triples and timestamped text in parallel using multiple preset algorithm models to generate a summary or structured document includes: The system outputs contextual semantics through a pre-defined semantic understanding model and infers the missing elements in the triples; the triples include the task, the person in charge, and the deadline. Complete the missing elements in the triplet.
[0059] In one embodiment, the step of processing triples and timestamped text in parallel using multiple preset algorithm models to generate a summary further includes: The content of the summary is determined to be consistent with the semantics of the triple by a preset alignment loss function.
[0060] In summary, this application provides a method and system for intelligent extraction and summarization of call content. By processing call audio from a smart device, the speech is converted into timestamped text, caller and caller segmentation information. Key information is then organized and extracted to generate text tagged with the caller. This tagged text undergoes intelligent recognition processing to extract important information, obtaining sentiment and target intent sentences, ultimately resulting in a final meeting summary. Through multi-model collaboration and parallel processing, call audio can be processed quickly, rapidly extracting key points and saving time. Furthermore, based on the "understanding" results, key information is extracted after entity recognition, sentiment analysis, and intent recognition to obtain sentiment and target intent sentences, followed by the extraction of action items and triples, resulting in higher accuracy. After obtaining the summary, the meeting summary template can be dynamically selected and adjusted to generate a final meeting summary in a custom format.
[0061] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in this application and in the embodiments can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in a variety of forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual-speed SDRAM (SSRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), RAMbus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).
[0062] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, apparatus, article, or method that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, apparatus, article, or method. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, apparatus, article, or method that includes that element.
[0063] The above description is only a preferred embodiment of this application and does not limit the patent scope of this application. Any equivalent structural or procedural changes made based on the content of this application's specification and drawings, or direct or indirect applications in other related technical fields, are similarly included within the patent protection scope of this application.
Claims
1. A method for intelligent extraction and summary generation of call content, characterized in that, Includes the following steps: Automatically trigger recording during a call to obtain the call audio; The call audio is stored in a preset user identification card, and a traceable index directory is generated; The call audio is obtained and extracted to a specified page through directory location; The call audio is input into a preset Whisper model for speech recognition within the designated page to obtain timestamped text; simultaneously, the call audio is processed through the Pyannote pipeline to obtain caller and caller segmentation information; wherein, the Pyannote pipeline includes VAD and speaker embedding model; The timestamped text, the caller and the caller segment information are merged, and key information is sorted and extracted to generate text with caller tags. The text with caller tags is sent to the adjusted BERT model, where entity recognition, sentiment analysis, and intent recognition are performed to obtain sentiment tendencies and target intent sentences. The target intent sentence is input into a multi-channel action item extraction model, and the action items and triples of the text with caller tags are extracted by the multi-channel action item extraction model; wherein, the multi-channel action item extraction model includes an action item extraction model, the BERT model and a triple extraction model; The triples and the timestamped text are processed in parallel using multiple preset algorithm models to generate a summary or structured document; The meeting minutes template is dynamically selected and adjusted, and the action items, the summary, the sentiment, and the target intent sentence are filled into the meeting minutes template to obtain the final meeting minutes, which are then saved and displayed on the designated page.
2. The method for intelligent extraction and summary generation of call content according to claim 1, characterized in that, The call audio is input into a preset Whisper model for speech recognition within the designated page to obtain timestamped text; Simultaneously, the step of processing the call audio through the Pyannote pipeline to obtain the caller and caller segmentation information includes: Caller identification: Converts continuous sound waves into discrete text and distinguishes different callers through voiceprint features; Call content recognition: Based on the text, the caller is divided and organized into segments to obtain segment information of the caller; wherein, the segment information of the caller includes an opening segment, a statement segment, and a closing statement.
3. The method for intelligent extraction and summary generation of call content according to claim 2, characterized in that, The algorithm model includes a speech recognition model, a caller separation and recognition model, an understanding model, and an extraction model; The speech recognition model includes a deep neural network model based on a recurrent neural network or a Transformer architecture; The caller separation and identification model includes a VAD model and a speaker embedding model; The understanding model includes a basic natural language understanding model and a pre-trained language model based on the Transformer architecture; The extraction models include Transformer-based sequence labeling models and Seq2Seq models.
4. The method for intelligent extraction and summary generation of call content according to claim 3, characterized in that, Sentiment analysis and intent recognition were performed on the opening paragraph, the statement paragraph, and the closing remarks, respectively. The steps of performing sentiment analysis and intent recognition on the opening paragraph, the statement paragraph, and the closing remarks respectively include: The meaning of the opening paragraph, the statement paragraph, and the conclusion is understood using the aforementioned understanding model; Analyze the grammatical structure and semantic roles of the opening paragraph, the statement paragraph, and the conclusion; Determine the attitude attributes of the caller and identify their intent to obtain the target intent sentence.
5. The method for intelligent extraction and summary generation of call content according to claim 1, characterized in that, The step of sending the text with caller tags to the adjusted BERT model and performing entity recognition, sentiment analysis, and intent recognition to obtain the sentiment tendency and target intent sentence includes: The intent recognition and classification are performed on the text containing the caller's tag; Extract the corresponding text based on the preset target intent to obtain the target intent sentence; The entity recognition includes person recognition, time recognition, and item recognition.
6. The method for intelligent extraction and summary generation of call content according to claim 1, characterized in that, The step of processing the triples and the timestamped text in parallel using multiple preset algorithm models to generate a summary or structured document includes: The system outputs contextual semantics through a pre-defined semantic understanding model and infers the missing elements in the triples; wherein the triples include the task, the person in charge, and the deadline. Complete the missing elements in the triplet.
7. The method for intelligent extraction and summary generation of call content according to claim 1, characterized in that, The step of processing the triples and the timestamped text in parallel using multiple preset algorithm models to generate a summary further includes: The content of the summary is determined to be consistent with the semantics of the triple by a preset alignment loss function.
8. A system for intelligent extraction and summarization of call content, characterized in that, include: The recording module is used to automatically trigger recording during a call to obtain the call audio; The storage module is used to store the call audio into a preset user identification card and generate a traceable index directory; The call audio extraction module is used to acquire the call audio and extract the call audio to a specified page through directory location; The speech recognition module is used to input the call audio into a preset Whisper model for speech recognition within the specified page, and obtain text with timestamps. The audio processing module is used to process the call audio through the Pyannote pipeline to obtain the caller and caller segmentation information; The information merging module is used to merge the timestamped text, the caller and the segmented information of the caller, and to organize and extract key information to generate text with caller tags; The analysis module is used to send the text with caller tags to the adjusted BERT model and perform entity recognition, sentiment analysis and intent recognition to obtain sentiment tendency results and target intent sentences; The action item extraction module is used to input the target intent sentence into the multi-channel action item extraction model, and extract the action items and triples of the text with caller tags through the multi-channel action item extraction model; The parallel processing module is used to process the triples and the timestamped text in parallel using multiple preset algorithm models, and generate a summary or structured document. The summary module is used to dynamically select and adjust the meeting minutes template, and fill the action items, the summary, the sentiment, and the target intent into the meeting minutes template to obtain the final meeting minutes, which are then saved and displayed on a designated page.
9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method for intelligent extraction and summary generation of call content according to any one of claims 1 to 8.
10. A computer storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the steps of the method for intelligent extraction and summary generation of call content according to any one of claims 1 to 8.