Generate summaries and next actions for users in real time from interaction records in natural language

By analyzing meeting interaction records through natural language processing and machine learning models, identifying action items and solutions, and generating real-time meeting summaries, the problem of difficulty in extracting action items in existing technologies is solved, and users can achieve real-time understanding and improve collaboration efficiency.

CN114691857BActive Publication Date: 2025-09-09INTERNATIONAL BUSINESS MACHINE CORPORATION
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111441651.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-12-31
Filing Date
2021-11-30
Publication Date
2025-09-09
Estimated Expiration
2041-11-30

AI Technical Summary

Technical Problem

Existing technologies make it difficult to quickly and automatically extract action items from meeting sessions and generate meeting summaries in real time, resulting in people who arrive late to the meeting being unable to understand the meeting discussion content and task allocation in a timely manner.

Method used

Use natural language processing and machine learning models to analyze meeting interaction records, identify action items, relationships, and solutions, and generate real-time meeting summaries to provide users with the tasks and information they need.

Benefits of technology

It enables fast and automatic extraction of action items and their solutions from meetings, ensuring that users understand meeting progress and task allocation, and promoting continuous and effective collaboration among distributed users.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114691857B_ABST
    Figure CN114691857B_ABST
Patent Text Reader

Abstract

The present disclosure relates to generating summaries and next actions for users in real time from interaction records in natural language. The system receives a message transmission and reception, video and / or audio input stream comprising conversations spoken by users at a group conference. The system obtains a single or multiple interaction records (including natural language texts that memorize the content spoken by each speaker at the conference) from these inputs, analyzes the content, and identifies a single or multiple action item tasks in the interaction record. The system then generates a summary indicating the action item tasks for the user. The system further detects whether each action item is addressed and whether the user's action item has a solution based on the content of the conversation. The system further detects whether an action item is a prerequisite for solving another action item by the user or in combination with another user. Using pre-configured templates, the system generates an action item summary, any associated solutions, and any relationship or prerequisite between the action items, and presents the summary to the user.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the application of natural language processing (NLP) and machine learning models (ML models) for managing meetings, and more particularly, to detecting action items of meeting transcripts and determining their status (e.g., completed or open, etc.) to facilitate continuous and effective collaboration among distributed user groups. Background Art

[0002] Typically, participants in a meeting consider "action items" (i.e., public commitments to perform a given task) as one of the most important outputs of a meeting. However, spoken action item detection is a relatively new task.

[0003] Someone joining the meeting a few minutes late should be able to learn what was discussed before joining, such as what tasks (action items) have been assigned, what types of tasks have been answered by others, and what types of action items are waiting unanswered for a particular person. Summary of the Invention

[0004] The following summary is intended to be illustrative only. It is not intended to limit the scope of the claims.

[0005] According to one aspect, the present disclosure provides a system and method for quickly and automatically extracting action items from a conference session and summarizing the conversation of the meeting and presenting required actions in real time so that users can address unresolved items in the meeting.

[0006] The system and method described herein are intended to summarize action items in a meeting, including: 1) action items; 2) relationships between different action items; and 3) solutions for each action item.

[0007] According to one aspect, the system and method invoke operations to determine tasks for which users are responsible based on interactions among attendees of a meeting, and provide a meeting summary to inform users of tasks and other information required of users attending the meeting.

[0008] According to one aspect, the system and method invoke operations that provide functionality to notify a team member or other user of tasks and other information required of the team member or user participating in a meeting, wherein the team member / user participates in a meeting after the meeting has begun.

[0009] In an embodiment, the meeting summary provides a list of action items, identified relationships between different action items, and further provides a recommended solution for each action item.

[0010] According to one aspect, a computer-implemented method is provided. The method includes: receiving, at one or more processors, a plurality of interaction records of natural language text associated with a group of users participating in a conference; analyzing the text of the plurality of interaction records using one or more classifier models to identify one or more action items from the text, the action items being related to respective tasks to be performed; detecting, from the interaction records, respective users assigned to perform the respective tasks; detecting, from the interaction records, whether each action item is addressed using a solution for performing the task; and generating a text summary of a solution addressing each action item for the user.

[0011] According to one aspect, a computer-implemented system is provided. The system includes: a memory storage device for storing a computer-readable program; and at least one processor adapted to execute the computer-readable program to configure the at least one processor to: receive a plurality of interaction records of natural language text associated with a group of users participating in a conference; analyze the text of the plurality of interaction records using one or more classifier models to identify one or more action items from the text, the action items being associated with respective tasks to be performed; detect, from the interaction records, respective users assigned to perform respective tasks; detect, from the interaction records, whether each action item is addressed using a solution for performing the task; and generate a text summary of a solution for addressing each action item for the user.

[0012] In a further aspect, a computer program product for performing operations is provided. The computer program product includes a storage medium readable by a processing circuit and storing instructions executed by the processing circuit for executing a method. The method is the same as listed above. BRIEF DESCRIPTION OF THE DRAWINGS

[0013] The foregoing various aspects and other features are explained in the following description taken in conjunction with the accompanying drawings, in which:

[0014] Figure 1 depicts a system block diagram of an embodiment of a computer system tool for automatically generating action item summaries and next actions for multiple users in real time from interaction records according to an embodiment of the present invention;

[0015] Figure 2 Depicted by Figure 1 The flow of operations performed by computing system tools;

[0016] Figures 3A-3B Described are example automated methods that implement operations to generate action item summaries, including next actions, for users in real time, in accordance with embodiments;

[0017] Figure 4illustrates a schematic diagram of an exemplary computer or processing system according to an embodiment of the present invention;

[0018] Figure 5 A cloud computing environment according to an embodiment of the present invention is described; and

[0019] Figure 6 Abstract model layers according to an embodiment of the present invention are depicted. DETAILED DESCRIPTION

[0020] According to one aspect, the present disclosure generally provides systems and methods for quickly and automatically facilitating manual coordination tasks for continuous and efficient collaboration of distributed users. In this effort, the systems and methods invoke operations to summarize a group meeting in real time and present action items that need to be taken, so that each user can address unresolved action items in the meeting.

[0021] In such efforts, the system enables a person who joins a meeting a few minutes late to learn what was discussed in the meeting before joining the meeting, i.e., to be informed about what types of new tasks have been assigned, what types of tasks have been answered by others, and what types of questions are waiting for this person to answer.

[0022] For users who need to handle multiple tasks while participating in a meeting (e.g., employees of a business), and users whose names are called in a meeting, the system and method ensure that users can: understand what the current status of the meeting is and what types of questions have been asked by receiving an interactive summary of what has been discussed in the meeting.

[0023] Thus, according to further aspects of the present disclosure, systems and methods are described herein that allow multiple users to interact on tasks based on input interaction records in natural language text or audio shared by multiple users (such as logs of conversations obtained from meetings, threads of emails, instant messages, messages communicated via Slack Technologies Inc. messaging platform, Twitter messaging platform, etc.). The systems and methods described herein further continuously summarize content that has been discussed, mentioned, or summarized in an expected or implicit manner. Further, the systems and methods described herein present a summary containing one or more pending action items for each user that appears in a non-existent form. The systems and methods described herein further present action items addressed by other groups of users (e.g., summaries with actions, solutions, etc.).

[0024] According to one aspect, in addition to the meeting summary, the system and method analyzes the relationship between different action items and further provides a solution for each action item.

[0025] The systems and methods described herein can be applied to any meeting to detect relationships between action items and solutions for each action item in real time from a meeting transcript.

[0026] The system and method described herein invoke operations to summarize action items in a meeting, including: 1) action items; 2) relationships between different action items; and 3) solutions for each action item. In one aspect, the system summarizes the meeting in real time and presents actions that need to be taken so that unresolved items in the meeting can be addressed.

[0027] In an embodiment, the method and system provide a meeting summary with an action list, analyze the relationship between different action items and further provide a solution for each action item. The method and system manages the conversation by identifying action items, detecting the relationship between the action items, and suggesting solutions.

[0028] Now see Figure 1 , a system 100 (e.g., a collaborative tool embodied as one or more computer systems implementing a method for automatically making action items available to meeting attendees who arrive late to a meeting) is provided. As a non-limiting example scenario, a collaborative user group meets via an online web conference or messaging platform that enables remote conferencing based on VoIP, online video, instant messaging, file sharing, and screen sharing, and the tool is operated when a user is late or cannot join the meeting until after the user has already started. According to an embodiment, a software application analyzes a meeting dialog box by invoking NLP operations to determine the tasks (i.e., action items) for which the user is responsible based on interactions between the attendees of the meeting. The system notifies the user of the tasks and other information required for the user to attend the meeting (i.e., before the user joins the meeting).

[0029] like Figure 1 As shown, system 100 is a computing system that provides a tool that is configured to generate a summary of action items including a user or a group of multiple users to promote communication and collaboration between user groups. The tool is embodied as a computing system having at least the components shown. In one embodiment, the action item and summary generation tool 100 implements an NLP analysis method for detecting action items based on one or more natural language text, video, audio and / or audio / video (A / V) inputs. The system tool 100 implements NLP analysis technology and a machine learning model 190, which includes a learning algorithm for training a model to detect action items including the next action item (for example, to enhance the collaborative efforts of the user group). In an embodiment, the tool 100 calls an operation that provides functionality to notify team members or other users of tasks and other information required by the team members or users to attend a meeting, where the team members / users attend the meeting after the meeting has already begun.

[0030] like Figure 1 As shown, in the context of automatically generating summaries and action items for a user or a group of multiple users according to one embodiment, the tool 100 is a computer system, computing device, mobile device, or server. In some aspects, the computing device 100 may include, for example, a personal computer, a laptop computer, a tablet computer, a smart device, a smartphone, or any other similar computing device.

[0031] The computing system 100 includes one or more hardware processors (e.g., hardware processor 152), memory 150, such as for storing an operating system, an application program interface (API) and program instructions, a network interface 156, a display device 158, an input device 159, and any other features common to computing devices. In some aspects, the computing system 100 can be, for example, any computing device configured to communicate with one or more websites 125, including web-based or cloud-based servers 120, via a public or private communication network 99. For example, the website can host web conferencing software or store a record / transcription of a current conference video conference or other meeting or interaction record. Such structured / unstructured electronic text-based interaction records can be stored in a database 130.

[0032] Further, as shown as part of the system 100, local storage is provided for the data processing framework, which may include an additional memory storage device 160, or a remote memory storage device (e.g., a database, a vocabulary database, an ontology) that can be accessed via a remote network connection for input to or use by the system 100.

[0033] exist Figure 1 In the depicted embodiment, the processor 152 may include, for example, a microcontroller, a field programmable gate array (FPGA), or any other processor configured to perform various operations. Also shown is a communication channel 140 (e.g., a wired connection, such as a data bus line, an address bus line, an input / output (I / O) data line, a video bus, an expansion bus, etc.) for routing signals between the various components of the system 100. The processor 152 is configured to execute method instructions as described below. These instructions may be stored, for example, as programming modules in another associated memory storage device 150.

[0034] The memory 150 may include, for example, non-transitory computer-readable media in the form of volatile memory (such as random access memory (RAM) and / or cache memory or other). The memory 150 may include, for example, other removable / non-removable, volatile / non-volatile storage media. By way of non-limiting example only, the memory 150 may include a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0035] The network interface 156 is configured to transmit data or information to and receive data or information from the website server 120, for example, via a wired or wireless connection. For example, the network interface 156 may utilize wireless technologies and communication protocols such as , WIFI (e.g., 802.11a / b / g / n), cellular networks (e.g., CDMA, GSM, M2M, and 3G / 4G / 4G LTE, 5G), near field communication systems, satellite communications, via a local area network (LAN), via a wide area network (WAN), or any other form of communication that allows the computing device 100 to transmit information to or receive information from the server 120).

[0036] Display 158 may include, for example, a computer monitor, a television, a smart TV, a display screen integrated into a personal computing device (such as, for example, a laptop, a smartphone, a smartwatch, a virtual reality headset, a smart wearable device, or any other mechanism for displaying information to a user). In some aspects, display 158 may include a liquid crystal display (LCD), an electronic paper / electronic ink display, an organic LED (OLED) display, or other similar display technology. In some aspects, display 158 may be touch-sensitive and may also serve as an input device.

[0037] The input device 159 may include, for example, a keyboard, a mouse, a touch-sensitive display, a keypad, a microphone, or other similar input devices or any other input devices that may be used alone or together to provide a user with the ability to interact with the computing device 100. In embodiments, through the user interface, a user may enter commands or utterances to activate or deactivate system modules and invoke operations at the modules for real-time system processing.

[0038] In one embodiment, an audio / video data capture unit 165 is provided to capture real-time A / V and conference streams from one or more users at a conference (e.g., a virtual or online conference such as ZOOM™ (Zoom Video Communications, Inc.), Webex™ (Cisco Systems, Inc.), GoToMeeting™ (LogMeIn USA, Inc.), etc.) for analysis and action item extraction processing according to embodiments of the present invention.

[0039] With respect to configuring a computer system as an analysis tool for summarizing action items in a meeting, including: 1) action items; 2) relationships between different action items; and 3) solutions for each action item, local or remote memory 160 may be configured to temporarily store or use text data or information 162, including but not limited to: any captured message threads or messaging conversations to be analyzed by NLP, action items, extracted action item-solution pairs, identified relationships between action items, or templates for automatically generating summaries.

[0040] The data may alternatively be stored in a separate local memory storage device attached to the computer system 100 .

[0041] like Figure 1 As shown, the memory 150 of the computer system 100 also stores processing software or application modules including programming instructions suitable for configuring the tool to provide the user with an analysis of the meeting and summary of detected action items, their resolutions and any next action items.

[0042] In one embodiment, one of the programmed processing modules stored in the associated memory 150 includes an audio / video control interface module 166 that provides instructions and logic for operating circuitry to initiate capture of audio or A / V conversation or similar messaging media (e.g., text) input for a real-time conference session or online meeting. In one embodiment, the system receives input activation or deactivation signals from a user device via at least one application program interface 166 for controlling A / V recording, for example, to responsively trigger or terminate system module operations.

[0043] In one embodiment, the speech-to-text conversion module 170 provides instructions and logic for operating circuits to, for example, use IBM's Speech to Text (a trademark of International Business Machines Corporation), Speech recognition software (Nuance Communications, Inc.) or the like is used to receive the input audio speech captured by the video capture device 165 and convert it into a text format. Module 170 initiates the storage / buffering of the online conference session and converts it into natural language text data for NLP processing and for use and analysis by other modules that process data according to various aspects of the present disclosure. The text output of the converted speech input forms a record of the interaction that is processed by other system components.

[0044] In one embodiment, another programmed processing module stored at the associated memory 150 of the system tool 100 includes a natural language processing (NLP) tool 175 to implement operations and techniques for analyzing text, natural language data for use by other system components that perform language summarization, translation, named entity recognition, relationship extraction, speech recognition, and topic segmentation. In one embodiment, the NLP processing tool utilizes a neural network that enables the capture of semantic properties of words and augments learning to perform advanced tasks such as question answering.

[0045] Another programming processing module stored at the associated memory 150 of the system tool 100 includes an action item identification and extraction module 180, which uses logic and instructions for operating circuits to identify action items from the interaction record or transcription of the meeting. Action items are tasks discussed during the meeting and need to be addressed by a person or a small group of people. The action item extraction module uses an NLP information extraction system for extracting action items. Reference is made to M. Purvery, J. Dowding, J. Niekrasz, P. Ehlen, S. Noorbaloochi and S. Peters, entitled "Detecting and Summarizing Action Items in Multi-Party Dialogue", SIGDIAL 2007, for a description of the exemplary NLP information extraction system used, the entire contents and disclosure of which are incorporated herein by reference.

[0046] Another programmed processing module stored at the associated memory 150 of the system tool 100 includes an action item solution extraction module 184 that employs instructions to configure the system to employ NLP tools and ML models to detect the existence of any solutions related to the user's action items discussed during the meeting.

[0047] Further programmed processing modules stored at the associated memory 150 of the system tool 100 include an action item relationship extraction module 188 that executes instructions to configure the system to employ NLP tools and ML models to detect the existence of any relationships between detected action items (e.g., action items that must be performed as a prerequisite to resolving another action item of a user being discussed during the meeting).

[0048] Another programming processing module 190 stored at the associated memory 150 of the system tool 100 uses instructions to configure the system to build, access and / or run supervised (or unsupervised) machine learning (ML) models for use by other system components. For example, the ML model module 190 calls operations to train a machine learning model for learning the next item related to the currently identified action item. The machine learning model can include, but is not limited to: a classifier, a regression model, a knowledge graph, or a neural network. For example, supervised learning (classification) techniques using regression models are used to predict the causality of an action item or the next action item recommendation associated with the occurrence of that particular action item, and / or unsupervised learning (clustering) techniques.

[0049] Further programmed processing modules stored at associated memory 150 of the system tool 100 include a summary generator 195 that employs instructions to configure the system to access a preconfigured template, populate the template with extracted information about action items associated with the user, including any detected solutions for resolving the action item and any detected related action items or preconditions that must be performed in order to resolve another action item of the user being discussed during the meeting, and present the generated summary to the user(s) of the meeting when requested.

[0050] like Figure 1 As further shown, the memory 150 includes a monitoring program 110 having instructions for configuring the computing system 100 to call each program module and invoke operations as described herein. In one embodiment, such monitoring program 110 calls methods and provides an application program interface for real-time control of receiving real-time input A / V signals, interaction records, text, etc., and determining action items from text, determining action item pairs with users, determining action item relationships, and determining action item solutions for late users or meeting attendees, and generating summaries.

[0051] Figure 2 Depicts the process of monitoring the Figure 1 An exemplary flow of operations 200 performed by the computing system tool 100 of FIG. 1 for controlling recording, processing, and analyzing spoken conference data to generate action item summaries for users in real time.

[0052] like Figure 2As shown in FIG, a video camera or similar recording device is physically located at the location where the user is conducting the meeting and is activated to generate an input stream 202 (such as an audio or A / V stream including content spoken by the attendees (users) of the meeting). The A / V or audio capture device 165 is configured to receive / capture the audio or A / V stream. In one embodiment, the user or attendee can dynamically adjust the recording of the conference audio or A / V stream through the audio / video stream capture processor 165 based on the user's information needs.

[0053] That is, via at least one application interface 166 for controlling A / V recording, the system receives an input activation or deactivation signal 205 transmitted from a user device, for example to trigger an operation for activating a system module and calling an operation at the module for real-time A / V signal capture, text capture action item determination, etc. when the user leaves a meeting, or to deactivate such an action item determination operation when the user appears in a meeting.

[0054] For example, in the case of an action item where the user only wants to summarize a portion of the meeting time, the input can be adjusted accordingly by the user invoking an activation / deactivation or toggle function so that the user can keep themselves updated about the meeting if they need to leave for a few minutes, cannot arrive in time, or must leave before the meeting ends.

[0055] In such an embodiment, the switch 210 initiates the generation of multimedia recording and interaction recording of the meeting in response to input activation or deactivation signal 205. For example, a user or attendee can select an "activate" button via the user device interface to initiate the recording of the A / V media stream of the meeting before the user arrives at the meeting or during the duration when the user must leave the meeting. When the user or attendee leaves the meeting, the meeting can be recorded continuously. Upon arriving at the meeting, the user or attendee can select a "deactivate" button to terminate the recording of the meeting. At this point, the system 100 will then automatically call an operation to process the recording of the meeting and generate a summary of action items for the corresponding duration for the user or attendee to leave. The action item summary generation system can be activated and deactivated by issuing a "wake-up" word via a microphone or similar audio interface device, rather than selecting an activation / deactivation button via the interface.

[0056] Alternatively, in various embodiments, the A / V stream 202 of the conference is recorded by default or upon consent of the participants, regardless of the presence / status of any attendee or conference participant. Then, after the conference is terminated, the system will generate a summary of each user's action items for the corresponding time period.

[0057] Whether a portion of the conference time is recorded for a user by selectively activating / deactivating it, or the entire conference A / V stream content is recorded, the recorded voice data is input into a real-time speech-to-text converter unit 170. The speech-to-text converter unit 170 is trained to recognize the voice of each speaker, transcribe the user's voice into text, and associate the text of the transcribed voice content spoken at the conference with its corresponding user. The system can generate and / or store a transcript of each speaker's spoken content (e.g., Speaker AD) in real time in the order in which the corresponding speaker spoke, as follows:

[0058] Speaker A: Content A

[0059] Speaker B: Content B

[0060] Speaker A: Content A"

[0061] Speaker C: Content C

[0062] Speaker D: Content D

[0063]

[0064] Then, in an embodiment, the action item identifier module 180 is then called to identify action items from the stored transcript or interaction record of the meeting. In an embodiment, the module 180 can call the operation of a hierarchical classifier including several sub-classifier models in a two-stage process, which includes: 1) detecting sub-dialogues (i.e., short paragraphs of conversation) of the meeting in which action items are proposed, discussed, and submitted; and 2) extracting phrases that accurately capture or summarize the tasks involved. In an embodiment, the detection of sub-dialogues is performed by using a hierarchical classifier that includes multiple independent sub-classifiers for detecting utterances in the conversation that express action items and user suggestions, agreements, or commitments. Then, a semantic parser is used to extract phrases from the utterances that describe the tasks and the time frame in which the tasks are to be performed. For example, to extract the time frame and task description, a rule-based NLP semantic parser is used that can find the basic predicate-argument structure of the main phrase types (e.g., symbols, noun phrases, verb phrases, prepositional phrases). In an embodiment, four independent classifier models 190 (e.g., linear kernel support vector machines) are run for task detection based on a set of different action item-specific dialogue action (AIDA) utterance classifications (description of the task to be performed, discussion of timeframes, ownership or assignment of responsibilities, and any explicit agreements). Four independent AIDA sub-classifiers can be trained based on features derived from various attributes of the utterances in context (e.g., n-grams, persistence and positional features from transcriptions, prosodic features, etc.). A super-classifier is then used to detect a window encompassing the captured sub-dialogue sequence based on the task utterances detected by the various AIDA sub-classifiers.

[0065] For each user, the system generates and records in real time each of the action items identified by the action item identifier unit 180, including the identity of the owner of the action item, a description of the task, and a designation of the time frame. For example, the module 180 determines a list of action items for the identified speaker A, such as:

[0066] Action Item 1

[0067] Action Item 2

[0068] Action Item 3

[0069]

[0070] In an embodiment, the urgency of the action items may be detected, for example, using a determined sentiment, semantic, or lexical analysis of the conference conversation, and the list of action items may be sorted according to the determined urgency.

[0071] Once the action item identifier unit 180 identifies each action item for a specific user / speaker, the system invokes an operation to determine the existence of a solution for each action item by the subject user / speaker, and extracts the solution for each action from the meeting transcript. For action item solution extraction, the action item solution extraction unit 184 converts the action item into a question-answering (QA) task. That is, given that the extracted action item is converted into a question, and based on the meeting transcript as context, the task is to extract a span from the context as an answer. Further identified is the specific user / speaker who is assigned the task of providing a solution to a specific action item. In an embodiment, a QA model trained on the Stanford Question Answering Dataset (SQuAD) dataset (https: / / rajpurkar.github.io / SQuAD-explorer / ) can be transmitted to perform the QA task. The SQuAD dataset has a large number of answers associated with its corresponding questions. One embodiment is to train a word embedding model on text related to the question, vectorize the questions and answers using the word embedding model, and finally use the vectorized questions and answers to train a machine learning model for the QA task. That is, the SQuAD dataset can be used to train a QA model or tool. An exemplary output of the action item solution extraction unit 184 is as follows:

[0072] Action 1: Solution 1 for speaker A / B, addressed

[0073] Action Item 2: Speaker C’s Solution 2, to be executed by Speaker D

[0074] Action Item 3: Open...

[0075] To improve performance, user-annotated datasets (e.g., transcription-action item solutions) serve as feedback to further fine-tune the QA model and improve its performance. For example, suppose an action item list indicates "This action item is open and assigned to speaker A." If this is incorrect, speaker A provides feedback by explicitly stating "Speaker B proposed a solution, and this action item has been resolved." This feedback is used to retrain the QA model.

[0076] In an embodiment, while identifying the solution (or not) for each of the action items to be assigned to a particular user / speaker via the action item solution extraction unit 184, the system invokes an operation to determine the existence of any related action items for the subject action item. That is, the system tool 100 simultaneously invokes the action item relationship extraction unit 188 to identify dependencies between the action items. Such dependencies, for example, are a prerequisite for determining whether one action item is a prerequisite for another action item.

[0077] In an embodiment, the NLP text entailment system detects whether a first action item (e.g., action 1) "entails" a second action (e.g., action 2). If such entailment is found, the solution to action 2 is determined by first solving action 1, i.e.:

[0078] Action 1←Action 2

[0079] In an embodiment, once the action item solution extraction unit 184 and the action item relationship extraction unit 188 identify and extract action item solution pairs for a user / speaker, and identify any relationships between action items for a particular user / speaker, the system calls an operation to generate a summary report for the user. A summary generator module 195 generates a human-readable summary of meeting action items based on a predefined template with selectable statements having fields (e.g., ***) that can be filled with corresponding information. In an embodiment, based on the relationship between the extracted action item solution pairs and the identified action items, the module fills the predefined template fields with the information and generates a summary output 198 for presentation to the user, for example, via the user's device display. In a non-limiting example, the predefined template can be filled with detailed information extracted from the action item solution / relationship unit. Such selectable exemplary statements with fields to be filled include, but are not limited to:

[0080] "*** action items have been discussed so far."

[0081] "The first action item is ***, and the speaker is ***"

[0082] "And speaker *** proposed a solution, and this action item has been addressed."

[0083] "The second action item is ***, which depends on the first action item."

[0084] "Speaker *** proposed a solution"

[0085] "And it will be accomplished or will be accomplished by the speaker***."

[0086] "The third action item is ***"

[0087] “There are no current discussions about a solution.”

[0088] In an embodiment, the system tool 100 automatically selects from these statements, populates these statements with the relevant information found, and provides a summary output 198 to the user, which explains, for example, the action items and corresponding solutions and related items discussed in the meeting during the time period responsive to the user's activation / deactivation control input. Thus, when the user enters the meeting or returns to the meeting after having been away from the meeting for a period of time, the user can be immediately made aware of the user's action items and, therefore, the responsibilities discussed during that time period.

[0089] Figures 3A-3B An example automated method 300 executed by monitoring program 110 for generating an action item summary for a group of multiple users is shown.

[0090] In the first step 303, when a meeting is being held among a group of users / attendees, it is determined that a user input deactivation control signal has not yet been received, or alternatively, an activation signal has been received. In response, at 305, the system initiates real-time recording / capturing of spoken content at the group conference, and initiates real-time speech-to-text conversion and storage of the text at 310. Then, at 315, it is continuously determined whether an interrupt deactivation signal input by the user has been received, which will initiate the stopping of the reception, recording, and speech conversion of the A / V stream. If an interrupt signal is not received at 315, the method continues to record the conference, as shown by the continuous return to step 305. Otherwise, at 315, once a deactivation signal is received, or by default, if the group conference has ended, the process continues to 320 to call NLP processing and hierarchical classifier processing to identify the action items of each of the users at the action item detector module 170.

[0091] In alternative embodiments, the received input may include a transcript of interactions in natural language text from input including, but not limited to: instant message exchanges, messages in a message board, messages received in a chat room or message board, and / or text converted from audio from a meeting scenario or other meeting transcription, email threads, social media messages / text, video streams from which text data is extracted and NLP applied.

[0092] from Figure 3A In step 320, the process continues with the simultaneous path to Figure 3B Step 330 and Figure 3B Step 333. Figure 3B At 330, for each action item owned by the determined user, the action item solution extraction unit 184 generates a question based on the action item at 335. For example, when extracting "Action Item 1", the module will rephrase it into a corresponding question, such as, "What is the solution for Action Item 1". Then, at 340, the question-answering model is run to obtain the action item solution from the context determined based on the meeting conversation transcription. For example, using any question-answering system (e.g., IBM's Watson deepQA architecture), the formed question can be parsed into different keywords and sentence fragments to find statistically relevant phrases from the meeting transcription history. Continue to Figure 3B At 345, it is determined whether there are any more action items for the user. If there are more action items for the user, the process proceeds to 350, where the next action item is retrieved and the process returns to Figure 3B 335, wherein the step of forming a related question from the action item and QA processing is performed to again find an answer to the question (i.e., a solution to the action item). Step 345 is repeated for each action item found for the user until such time as when it is determined at 345 that there are no more action items for the user, in which case the process proceeds to a summary generator step at 375, in which a template is obtained and a summary is generated using the template to populate the template with information such as listing the action items, the identified speakers, and the corresponding solutions (i.e., how the user would solve the action item if it were indeed to be solved).

[0093] return Figure 3A In step 333, the process simultaneously performs steps to obtain a combination of action item pairs. For three (3) action items, such pairs may include, for example:

[0094] Action Item 1 → Action Item 2

[0095] Action Item 1 → Action Item 3

[0096] Action Item 2 → Action Item 3

[0097]

[0098] Then, at 338, for each action item pair associated with the user, the method proceeds to Figure 3B 343, wherein the method runs an NLP entailment model to determine whether there is a relationship between the action item pairs or whether there is no relationship (no correlation or neutrality). Specifically, at 343, the NLP tool finds the directional relationship between the text fragments, for example, whether one action item pair can be inferred from another. Whenever the truth value of one text fragment follows from another text, the relationship holds. In the NLP entailment framework, the entailment text is the action item text, and the entailment text is another action item (hypothesis), and the NLP model determines whether the hypothesis is entailed by the text. Using a dataset of examples marked for entailment, a statistical classifier (e.g., a logistic regression classifier) ​​is learned using techniques from supervised machine learning. The input to the classifier is two sentences or phrases (e.g., S1 - submit a claim for a refund, S2 - submit a spending request), and the output of the classifier will be the relationship between S1 and S2 in terms of entailment. If the classifier predicts that S1 entails S2, it means that action S2 will occur before action S1.

[0099] Continue to Figure 3B At 348, a determination is made as to whether there are any more action item pairs for the user. If there are more action item pairs for the user, the process proceeds to 360 where the next action item pair is retrieved, and the process returns to Figure 3B 343, wherein a step of running an NLP entailment model based on the text snippet of the next action-term pair is performed to again determine the relationship of the action-term pair. Steps 345-360 are repeated for each action-term pair found for the user until such time as it is determined at 348 that there are no more action-term pairs to be processed for entailment for the user, in which case the process proceeds to a summary generator step at 375, wherein a template is obtained at 375 and the relevant action terms are added to the user's summary.

[0100] Finally, in Figure 3B At 375, the summary generated by the system 100 is output to the user's designated device for presentation to the user. For a particular user, the output may include: a summary of the interaction record, including a summary of the action items and specifications of to whom the action items are assigned, which action items are closed (completed) or open, and the relationships between the action items (e.g., dependencies, who needs to handle open items, and who handled closed items and how they were completed).

[0101] In further embodiments, the system can receive feedback manually input into the system by users and respond to other users and their actions. For example, to improve the performance of a QA task, users can annotate a dataset (transcription-action item-solution) to further fine-tune and improve the QA model. One embodiment uses action items from actual meeting transcripts that appear in the system. Based on user feedback on the action items, an additional training set is generated. This training set is used to further train the QA model.

[0102] Figure 4 An example computing system according to the present invention is described. It should be understood that the depicted computer system is only one example of a suitable processing system and is not intended to suggest any limitation as to the scope of use or functionality of embodiments of the invention. For example, the illustrated system is operational with numerous other general purpose or special purpose computing system environments or configurations. Figure 4 Examples of well-known computing systems, environments, and / or configurations for the illustrated system may include, but are not limited to, personal computer systems, server computer systems, thin clients, thick clients, handheld or laptop devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputer systems, mainframe computer systems, and distributed cloud computing environments that include any of the above systems or devices, etc.

[0103] In some embodiments, the computer system may be described in the general context of computer system executable instructions embodied as program modules stored in memory 16 and executed by the computer system. Generally speaking, program modules may include routines, programs, objects, components, logic, data structures, etc. that perform particular tasks and / or implement particular input data and / or data types according to the present invention (see, for example, Figures 3A-3B ).

[0104] The components of the computer system may include, but are not limited to, one or more processors or processing units 12, memory 16, and a bus 14 that operably couples various system components, including memory 16, to the processor 12. In some embodiments, the processor 12 may execute one or more modules 11 loaded from the memory 16, wherein the program module(s) are embodied as software (program instructions) that causes the processor to perform one or more method embodiments of the present invention. In some embodiments, the modules 11 may be programmed into the integrated circuits of the processor 12 and loaded from the memory 16, the storage device 18, the network 24, and / or a combination thereof.

[0105] Bus 14 may represent one or more of any of several types of bus structures, including a memory bus or memory controller, a peripheral bus, an accelerated graphics port, and a processor or local bus using any of a variety of bus architectures. By way of example and not limitation, such architectures include an Industry Standard Architecture (ISA) bus, a Micro Channel Architecture (MCA) bus, an Enhanced ISA (EISA) bus, a Video Electronics Standards Association (VESA) local bus, and a Peripheral Component Interconnect (PCI) bus.

[0106] The computer system may include a variety of computer system readable media. Such media can be any available media that can be accessed by the computer system, and it can include both volatile and nonvolatile media, removable and non-removable media.

[0107] Memory 16 (sometimes referred to as system memory) may include computer-readable media in the form of volatile memory (such as random access memory (RAM), cache memory, and / or other forms). The computer system may also include other removable / non-removable, volatile / non-volatile computer system storage media. By way of example only, a storage system 18 may be provided for reading from and writing to a non-removable, non-volatile magnetic medium (e.g., a "hard drive"). Although not shown, a magnetic disk drive may be provided for reading from and writing to a removable non-volatile magnetic disk (e.g., a "floppy disk") and an optical disk drive may be provided for reading from and writing to a removable non-volatile optical disk (such as a CD-ROM, DVD-ROM, or other optical media). In such instances, each may be connected to bus 14 via one or more data media interfaces.

[0108] The computer system may also communicate with one or more external devices 26, such as a keyboard, pointing device, display 28, etc.; one or more devices that enable a user to interact with the computer system; and / or any device that enables the computer system to communicate with one or more other computing devices (e.g., a network card, a modem, etc.). Such communication may occur via input / output (I / O) interface 20.

[0109] In addition, the computer system can communicate with one or more networks 24, such as a local area network (LAN), a general wide area network (WAN), and / or a public network (e.g., the Internet), via a network adapter 22. As depicted, the network adapter 22 communicates with the other components of the computer system via bus 14. It should be understood that, although not shown, other hardware and / or software components may be used in conjunction with the computer system. Examples include, but are not limited to, microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data archival storage systems.

[0110] The present invention may be a system, method and / or computer program product of any possible level of integrated technical detail. The computer program product may include a computer-readable storage medium (or medium) having computer-readable program instructions thereon, the computer-readable program instructions being used to cause a processor to perform various aspects of the present invention.

[0111] Computer readable storage medium can be a tangible device that can retain and store instructions for use by instruction execution equipment.Computer readable storage medium can be, for example, but not limited to, electronic storage device, magnetic storage device, optical storage device, electromagnetic storage device, semiconductor storage device or any suitable combination of the foregoing.A non-exhaustive list of more specific examples of computer readable storage medium includes the following: portable computer disk, hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), portable compact disc read-only memory (CD-ROM), digital versatile disk (DVD), memory stick, floppy disk, mechanical encoding device (such as punch card or the convex structure in the groove with the instruction recorded thereon) and any suitable combination of the foregoing.Computer readable storage medium as used herein should not be interpreted as transient signal itself (such as radio wave or other free propagating electromagnetic wave, electromagnetic wave propagated by waveguide or other transmission medium (for example, light pulse by fiber optic cable) or electrical signal transmitted by wire).

[0112] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to a corresponding computing / processing device, or downloaded to an external computer or external storage device via a network (e.g., the Internet, a local area network, a wide area network, and / or a wireless network). The network can include copper transmission cables, optical transmission fibers, wireless transmissions, routers, firewalls, switches, gateway computers, and / or edge servers. The network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions for storage in a computer-readable storage medium within the corresponding computing / processing device.

[0113] The computer-readable program instructions for performing the operations of the present invention may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state setting data, configuration data for an integrated circuit, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages ​​(such as Smalltalk, C++, etc.) and procedural programming languages ​​(such as the "C" programming language or similar programming languages). The computer-readable program instructions may be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the latter case, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet service provider). In some embodiments, electronic circuits (including, for example, programmable logic circuits, field programmable gate arrays (FPGAs), or programmable logic arrays (PLAs)) may execute computer-readable program instructions by utilizing state information of the computer-readable program instructions to personalize the electronic circuits so as to perform various aspects of the present invention.

[0114] Aspects of the present invention are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the present invention. It should be understood that each block of the flowchart and / or block diagram, and combinations of blocks in the flowchart and / or block diagram, can be implemented by computer-readable program instructions.

[0115] These computer-readable program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine, such that the instructions executed by the processor of the computer or other programmable data processing apparatus create a device for implementing the functions / actions specified in one or more blocks of the flowchart illustrations and / or block diagrams. These computer-readable program instructions can also be stored in a computer-readable storage medium that can direct the computer, programmable data processing apparatus, and / or other equipment to operate in a particular manner, such that the computer-readable storage medium having the instructions stored therein comprises an article of manufacture that includes instructions for implementing various aspects of the functions / actions specified in one or more blocks of the flowchart illustrations and / or block diagrams.

[0116] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device, so that a series of operational steps are performed on the computer, other programmable apparatus, or other device to produce a computer-implemented process, so that the instructions executed on the computer, other programmable apparatus, or other device implement the functions / actions specified in one or more boxes of the flowchart and / or block diagram.

[0117] The flow charts and block diagrams in the accompanying drawings illustrate the architecture, functions and operations of the possible implementations of the systems, methods and computer program products according to various embodiments of the present invention. To this end, each frame in the flow chart or block diagram may represent a part for a module, segment or instruction, which includes one or more executable instructions for realizing the logical function of the specification. In some alternative embodiments, the functions marked in the frame may not occur in the order marked in the figure. For example, depending on the functions involved, the two frames shown in succession may actually be performed substantially simultaneously, or these frames may sometimes be performed in the opposite order. It will also be noted that each frame in the block diagram and / or flow chart, and the combination of the frames in the block diagram and / or flow chart can be implemented by a system based on special-purpose hardware, which performs a specified function or action or performs a combination of special-purpose hardware and computer instructions.

[0118] The terms used herein are only used to describe the purpose of specific embodiments and are not intended to limit the present invention. As used herein, the singular forms "a", "an" and "the" are intended to also include the plural forms, unless the context clearly indicates otherwise. It will be further understood that when the terms "comprise" and / or "include" are used in this specification, they specify the presence of the stated features, wholes, steps, operations, elements and / or components, but do not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components and / or their groups. The corresponding structures, materials, actions and equivalents of all elements in the following claims are intended to include any structure, material or action for performing a function in combination with other claimed elements as specifically claimed. The description of the present invention is presented for the purpose of illustration and description, but is not intended to be exhaustive or limited to the present invention in the disclosed form. Without departing from the scope and spirit of the present invention, many modifications and variations will be apparent to those of ordinary skill in the art. The embodiments are selected and described in order to best explain the principles and practical applications of the present invention, and to enable those of ordinary skill in the art to understand the present invention for various embodiments with various modifications suitable for the specific purposes under consideration.

[0119] It should be understood that although the present disclosure includes detailed descriptions about cloud computing, the implementation of the teachings cited herein is not limited to cloud computing environments. Instead, embodiments of the present invention can be implemented in conjunction with any other type of computing environment now known or later developed.

[0120] Cloud computing is a service delivery model for enabling convenient, on-demand network access to a shared pool of configurable computing resources (e.g., networks, network bandwidth, servers, processing, memory, storage, applications, virtual machines, and services) that can be quickly provisioned and released with minimal management effort or interaction with the service provider. The cloud model can include at least five characteristics, at least three service models, and at least four deployment models.

[0121] Features are as follows:

[0122] On-demand self-service: A cloud consumer can unilaterally provision computing capacity (such as server time and network storage) automatically on demand without requiring human interaction with the provider of the service.

[0123] Broad Network Access: Capabilities are available over the network and accessed through standard mechanisms that facilitate use by heterogeneous thin or thick client platforms (e.g., mobile phones, laptops, and PDAs).

[0124] Resource pooling: A provider's computing resources are pooled to serve multiple consumers using a multi-tenant model, where different physical and virtual resources are dynamically assigned and reassigned as needed. Location independence is a key consideration, as consumers typically have no control or knowledge of the exact location of the provided resources, but may be able to specify a location at a higher level of abstraction (e.g., country, state, or data center).

[0125] Rapid elasticity: Capacity can be quickly and elastically provisioned (in some cases, automatically) to quickly scale down and quickly released to quickly scale up. To the consumer, the capacity available for provisioning generally appears unlimited and can be purchased at any time and in any quantity.

[0126] Measured services: Cloud systems automatically control and optimize resource usage by leveraging metering capabilities at some level of abstraction appropriate to the type of service (e.g., storage, processing, bandwidth, and active user accounts). Resource usage can be monitored, controlled, and reported, providing transparency to both the provider and consumer of the utilized service.

[0127] The service model is as follows:

[0128] Software as a Service (SaaS): The ability provided to consumers is to use the provider's applications running on a cloud infrastructure. The applications are accessible from individual client devices through a thin client interface such as a web browser (e.g., web-based email). The consumer does not manage or control the underlying cloud infrastructure, including the network, servers, operating systems, storage, or even individual application capabilities, with the possible exception of limited user-specific application configuration settings.

[0129] Platform as a Service (PaaS): The capability provided to consumers is to deploy applications created or acquired by the consumer on cloud infrastructure. The applications are built using programming languages ​​and tools supported by the provider. Consumers do not manage or control the underlying cloud infrastructure, including networks, servers, operating systems, or storage, but do have control over the deployed applications and possibly the configuration of the application hosting environment.

[0130] Infrastructure as a Service (IaaS): The capability provided to consumers is the provisioning of processing, storage, networking, and other basic computing resources on which consumers can deploy and run any software, including operating systems and applications. Consumers do not manage or control the underlying cloud infrastructure, but rather have control over the operating system, storage, deployed applications, and possibly limited control over selected networking components (e.g., host firewalls).

[0131] The deployment model is as follows:

[0132] Private cloud: A cloud infrastructure operated solely for an organization. It can be managed by the organization or a third party and can exist on or off-premises.

[0133] Community cloud: Cloud infrastructure is shared by several organizations and supports a specific community with shared concerns (e.g., mission, security requirements, policies, and compliance considerations). It can be managed by the organization or a third party and can exist on-premises or off-premises.

[0134] Public cloud: Cloud infrastructure is made available to the public or a large industry group and is owned by an organization that sells cloud services.

[0135] Hybrid cloud: A cloud infrastructure that is composed of two or more clouds (private, community, or public) that remain unique entities but are bound together by standardized or proprietary technologies that enable data and application portability (e.g., cloud bursting for load balancing between clouds).

[0136] The cloud computing environment is service-oriented and focuses on statelessness, low coupling, modularity, and semantic interoperability. The core of cloud computing is the infrastructure consisting of a network of interconnected nodes.

[0137] Now see Figure 5, depicts an illustrative cloud computing environment 50. As shown, the cloud computing environment 50 includes one or more cloud computing nodes 11, and local computing devices used by cloud consumers (such as, for example, personal digital assistants (PDAs) or mobile phones 54A, desktop computers 54B, laptop computers 54C and / or automobile computer systems 54N) can communicate with the cloud computing nodes 11. The nodes 10 can communicate with each other. They can be physically or virtually grouped (not shown) in one or more networks (such as the private clouds, community clouds, public clouds, or hybrid clouds described above, or a combination thereof). This allows the cloud computing environment 50 to provide infrastructure, platforms, and / or software as services without the cloud consumer needing to maintain resources on local computing devices. It should be understood that Figure 5 The types of computing devices 54A-N shown are intended to be illustrative only, and computing node 10 and cloud computing environment 50 may communicate with any type of computerized device over any type of network and / or network-addressable connection (eg, using a web browser).

[0138] Now see Figure 6 , showing the cloud computing environment 50 ( Figure 5 ) provides a set of functional abstraction layers. It should be understood in advance that Figure 6 The components, layers, and functions shown in are intended to be illustrative only, and embodiments of the present invention are not limited thereto. As depicted, the following layers and corresponding functions are provided:

[0139] The hardware and software layer 60 includes hardware and software components. Examples of hardware components include: host computers 61; servers based on RISC (Reduced Instruction Set Computer) architecture 62; servers 63; blade servers 64; storage devices 65; and network and networking components 66. In some embodiments, software components include network application server software 67 and database software 68.

[0140] Virtualization layer 70 provides an abstraction layer from which the following examples of virtual entities can be provided: virtual servers 71 ; virtual storage 72 ; virtual networks 73 , including virtual private networks; virtual applications and operating systems 74 ; and virtual clients 75 .

[0141] In one example, the management layer 80 may provide the functionality described below. Resource provisioning 81 provides dynamic acquisition of computing resources and other resources for performing tasks within the cloud computing environment. Metering and pricing 82 provides cost tracking when resources are utilized within the cloud computing environment and bills or invoices for the consumption of these resources. In one example, these resources may include application software licenses. Security provides identity verification for cloud consumers and tasks, as well as protection of data and other resources. User portal 83 provides access to the cloud computing environment for consumers and system administrators. Service level management 84 provides cloud computing resource allocation and management so that required service levels are met. Service level agreement (SLA) planning and fulfillment 85 provides pre-arrangement and procurement of cloud computing resources in anticipation of future requirements for the cloud computing resources according to the SLA.

[0142] The workload layer 90 provides examples of functionality that can utilize a cloud computing environment. Examples of workloads and functionality that can be provided from this layer include: mapping and navigation 91; software development and lifecycle management 92; virtual classroom education delivery 93; data analytics processing 94; transaction processing 95; and processing 96 to automatically generate summaries and next actions from natural language interaction records in real time according to aspects of the present disclosure.

Claims

1. A computer-implemented method for processing an interaction record, comprising: receiving, at one or more processors, a plurality of interaction records of natural language text associated with a group of users participating in a conference; analyzing the text of the plurality of interaction records using one or more classifier models to identify one or more action items from the text, the action items being related to corresponding tasks to be performed; detecting, from the interaction record, a corresponding user assigned to perform the corresponding task; Detecting from the interaction record whether each action item is addressed using a solution for performing the task, wherein detecting from the interaction record whether each action item is addressed using a solution for performing the task comprises: Convert the detected action items into question form; Forming question-answering tasks; and Running the trained question-answering model to extract spans from the interaction record and determine whether there is an answer to the question from the extracted spans; and Generate a text summary of the solution to each action item for the user.

2. The computer-implemented method of claim 1 , further comprising: receiving a real-time audio signal stream or an audio / video signal stream capturing user spoken content during the conference; and The real-time audio signal stream or audio / video signal stream is converted into a text representation to form the plurality of interaction records.

3. The computer-implemented method of claim 1 , further comprising: One of an activation signal or a deactivation signal is received at the one or more processors, the activation signal being used to initiate the reception, analysis, action item detection, and summary generation of the interaction record at the beginning of the meeting or during the meeting, and the deactivation signal being used to terminate the reception, analysis, action item detection, and summary generation of the interaction record during the meeting.

4. The computer-implemented method of claim 1 , further comprising: Use NLP entailment models to identify whether any action item depends on another action item or whether a task-related action item is a prerequisite for performing another action item task.

5. The computer-implemented method of claim 4, wherein: Generating a text summary of a solution to each action item for the user further includes: obtaining a template having predefined fields for receiving information related to the action item detected in the interaction record; and The template is populated with action item information including the corresponding user who owns the action item, and for each action item: a solution for the action item and / or any prerequisite action items on which the action item depends are specified in the template information.

6. The computer-implemented method of claim 5 , further comprising: receiving feedback information from a user regarding an action item task, its detected corresponding user, and a solution to the action item based on the populated template summary information; and The feedback information is used to refine and adjust one or more of the following: the one or more classifier models, the trained question-answering model, and the NLP entailment model.

7. A computer-implemented system for processing interaction records, comprising: a memory storage device for storing a computer readable program, and At least one processor is adapted to run the computer-readable program to configure the at least one processor to perform the method according to any one of claims 1 to 6.

8. A computer program product comprising a computer readable program, wherein: When executed on a computer comprising at least one processor, the computer readable program causes the at least one processor to perform the method according to any one of claims 1 to 6.

9. A computer-implemented system for processing interaction records, comprising means for performing the method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Computerized intelligent assistant for conferences

    CN112075075A

  • Automatic generation of action items from a meeting transcript

    US20170161258A1