Computer-implemented method, computer-implemented system, and computer program product for generating summaries and next actions for multiple users in real time from natural language dialogue recordings

The system uses NLP and ML to automatically summarize meeting action items and their resolutions, ensuring participants joining late can understand and address outstanding tasks effectively.

JP7754590B2Active Publication Date: 2025-10-15INTERNATIONAL BUSINESS MACHINE CORPORATION
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2021201458
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2020-12-31
Filing Date
2021-12-13
Publication Date
2025-10-15
Estimated Expiration
2041-12-13

AI Technical Summary

Technical Problem

Participants in meetings often join late and struggle to understand the action items discussed and their status, leading to inefficiencies in collaboration.

Method used

A system and method that uses NLP and ML models to automatically extract action items, their relationships, and resolutions from meeting transcripts, providing real-time summaries and notifications to attendees.

Benefits of technology

Enables participants to quickly grasp meeting outcomes, including assigned tasks, their status, and necessary actions, facilitating effective collaboration among distributed groups.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007754590000001
    Figure 0007754590000001
  • Figure 0007754590000002
    Figure 0007754590000002
  • Figure 0007754590000003
    Figure 0007754590000003
Patent Text Reader

Abstract

To provide a computer-implemented method, computer-implemented system and computer program product for generating summaries and next actions in real time for multiple users from interaction records in a natural language.SOLUTION: A system receives input streams of messaging, video or audio, or a combination thereof, including dialogue spoken by users at a group meeting. From these inputs, the system obtains single or multiple interaction records including natural language text recording contents spoken by each speaker at a meeting, analyzes the contents, and identifies single or multiple action item tasks in the interaction records. The system then generates summaries indicating the action item tasks for the users. Using a pre-configured template, the system generates action item summaries, any associated solution, and any relationship or precondition between the action items and presents the summaries to the users.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to the application of Natural Language Processing (NLP) and Machine Learned models (ML models) to manage meetings, and more particularly to detecting action items in meeting transcripts and determining their status (e.g., completed or unresolved) to facilitate continuous and effective collaboration among a distributed group of users. [Background technology]

[0002] Meeting participants generally consider "action items," i.e., declarations of commitment to perform a given task, to be one of the most important outcomes of a meeting. However, detecting stated action items is a relatively new task. [Prior art documents] [Non-patent literature]

[0003] [Non-Patent Document 1] M. Purvery, J. Dowding, J. Niekrasz, P. Ehlen, S. Noorbaloochi and S. Peters entitled “Detecting and Summarizing Action Items in Multi-Party Dialogue”, SIGDIAL 2007 Summary of the Invention [Problem to be solved by the invention]

[0004] A person who joins a meeting a few minutes late should be able to understand what was discussed before joining, including which tasks (action items) have been assigned, what tasks have already been answered by others, and what action items remain unanswered for a particular person. Embodiments of the present disclosure aim to generate summaries and next actions for multiple users in real time from natural language dialogue recordings. [Means for solving the problem]

[0005] The following abstract is intended to be illustrative only and is not intended to limit the scope of the claims.

[0006] According to one aspect, the present disclosure provides a system and method for quickly and automatically extracting action items from a meeting conversation, summarizing the meeting dialogue in real time and suggesting actions that need to be taken so that a user can address outstanding action items in the meeting.

[0007] The systems and methods described herein aim to summarize action items in a meeting, including 1) the action items, 2) the relationships between different action items, and 3) the resolution of each action item.

[0008] According to one aspect, the system and method determine tasks for which users are responsible based on interactions between conference attendees and invokes operations to provide a conference summary to inform users of tasks and other information required for attending the conference.

[0009] According to one aspect, the system and method invokes operations that provide functionality for notifying team members or users attending a meeting of required tasks and other information when the team members / users join a meeting after the meeting has started.

[0010] In one embodiment, the meeting summary provides a list of action items, identified relationships between different action items, and further provides recommended solutions for each action item.

[0011] According to one aspect, a computer-implemented method is provided that includes, at one or more processors, receiving a plurality of interaction records of natural language text related to a group of users attending a conference, analyzing the text of the plurality of interaction records using one or more classifier models to identify from the text one or more action items related to respective tasks to be performed, detecting from the interaction records respective users assigned to perform the respective tasks, detecting from the interaction records whether each action item is addressed with a solution for performing the task, and generating a text summary of solutions addressing each action item for the users.

[0012] According to one aspect, a computer-implemented system is provided, the system including: a memory storage device for storing a computer-readable program; and at least one processor adapted to execute the computer-readable program, the computer-readable program configuring the one or more processors to: receive a plurality of interaction records of natural language text related to a group of users attending a conference; analyze the text of the plurality of interaction records using one or more classifier models to identify from the text one or more action items associated with respective tasks to be performed; detect from the interaction records respective users assigned to perform the respective tasks; detect from the interaction records whether each action item is addressed with a solution for performing the task; and generate a text summary of solutions addressing each action item for the users.

[0013] In a further aspect, a computer program product for performing operations is provided, the computer program product including a storage medium readable by a processing circuit and storing instructions for execution by the processing circuit to perform a method, the method being the same as recited above.

[0014] The foregoing aspects and other features are explained in the following description in conjunction with the accompanying drawings. [Brief explanation of the drawings]

[0015] [Figure 1] FIG. 1 is a system block diagram of an embodiment of a computer system tool for automatically generating action item summaries and next actions for multiple users in real time from interaction records, according to an embodiment of the present invention. [Figure 2] Figure 1 is a flow diagram of operations performed by the computing system tool. [Figure 3A] FIG. 1 illustrates an exemplary automated method that performs operations to generate a real-time summary of action items, including a user's next actions, in one embodiment. [Figure 3B] FIG. 1 illustrates an exemplary automated method that performs operations to generate a real-time summary of action items, including a user's next actions, in one embodiment. [Figure 4] FIG. 1 is a schematic diagram of an exemplary computer or processing system according to an embodiment of the present invention. [Figure 5] FIG. 1 illustrates a cloud computing environment according to one embodiment of the present invention. [Figure 6] FIG. 2 illustrates abstraction model layers according to one embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0016] According to one aspect, the present disclosure generally provides systems and methods for quickly and automatically facilitating manual coordination tasks for continuous, effective collaboration of distributed users. In this approach, the systems and methods invoke operations to summarize group meetings in real time and present action items that need to be performed so that each user can address outstanding action items in the meetings.

[0017] In such an approach, the system allows a person who joins a meeting a few minutes late to know what was discussed in the meeting before joining, i.e., what new tasks have been assigned, what tasks have already been answered by others, and what questions remain unanswered for this person.

[0018] For users who need to handle multiple tasks while participating in a meeting, and for users designated at a meeting, such as employees of a company, the system and method ensures that the user understands what the current status of the meeting is and what questions have been asked by receiving a dialogue summary of what has already been discussed in the meeting.

[0019] Thus, according to further aspects of the present disclosure, systems and methods are described herein that enable multiple users to interact regarding a task based on input interaction records of natural language text or voice shared by the multiple users, such as conversation logs obtained from meetings, email threads, instant messages, messages communicated via Slack Technologies Inc. messaging platforms, Twitter messaging platforms, etc. The systems and methods described herein further provide predictive or implicit ongoing summaries of what has been discussed, mentioned, or concluded. Furthermore, the systems and methods described herein present summaries including one or more pending action items for each user that occurred in the user's absence. The systems and methods described herein also present action items addressed by other group users, including, for example, action summaries, solutions, etc.

[0020] According to one aspect, in addition to summarizing the meeting, the system and method analyzes the relationships between different action items and further provides a solution for each action item.

[0021] The systems and methods described herein can be applied to any meeting to detect, in real time, relationships between action items and their solutions from an audio transcript of the meeting.

[0022] The systems and methods described herein invoke operations to summarize action items in a meeting, including 1) the action items, 2) the relationships between different action items, and 3) the resolution of each action item. In one aspect, the system summarizes the meeting in real time and presents actions that need to be taken so that outstanding items in the meeting can be addressed.

[0023] In one embodiment, the method and system provides a summary of the meeting with a list of actions, analyzes the relationships between different action items, and provides solutions for each action item. The method and system manages the conversation by identifying action items and detecting relationships between them, and proposes solutions.

[0024] Referring now to FIG. 1 , a collaboration tool is provided, embodied as system 100, e.g., one or more computer systems, implementing a method for automatically making action items available to meeting attendees who arrive late. In one non-limiting example scenario, users in a collaboration group meet via an online web conferencing or messaging platform that enables remote conferencing based on VoIP, online video, instant messaging, file sharing, and screen sharing, and the tool operates when a user is late or unable to join before the meeting begins. According to one embodiment, a software application analyzes the meeting conversation by invoking NLP operations to determine tasks, i.e., action items, for which users are responsible based on the interactions between meeting attendees. The system notifies users of the tasks and other information required for attending the meeting, i.e., before they join the meeting.

[0025] As shown in FIG. 1 , system 100 is a computing system that provides a tool configured to generate summaries containing action items for a user or a group of users to facilitate communication and collaboration among the group of users. The tool is embodied as a computing system having at least the components shown. In one embodiment, action item and summary generation tool 100 implements NLP analysis methods for detecting action items based on one or more natural language text inputs, video inputs, audio inputs, and / or audio / video (A / V) inputs. System tool 100 implements NLP analysis techniques and machine-learned models 190, including learning algorithms for training the models to detect action items, including upcoming action items, to enhance the collaborative efforts of a group of users, for example. In one embodiment, tool 100 invokes operations that provide functionality to notify team members or other users, when a team member / user joins a meeting after the meeting has started, of tasks and other information required for the team member or user attending the meeting.

[0026] 1, in the context of automatically generating summaries and action items for a user or group of users, according to one embodiment, tool 100 is a computer system, computing device, mobile device, or server. In some aspects, computing device 100 may include, for example, a personal computer, a laptop, a tablet, a smart device, a smartphone, or any other similar computing device.

[0027] Computing system 100 includes one or more hardware processors, such as hardware processor 152; memory 150 for storing, for example, an operating system, an application program interface (API), and program instructions; a network interface 156; a display device 158; an input device 159; and any other functionality common to computing devices. In some embodiments, computing system 100 may be any computing device configured to communicate with one or more websites 125, including web-based or cloud-based servers 120, via, for example, a public or private communications network 99. For example, the website may host web conferencing software or store video conferences of current meetings, or recordings / audio transcripts or interaction records of those meetings. Such structured or unstructured electronic text interaction records may be stored in database 130.

[0028] Additionally, as shown as part of system 100, there is provided local memory useful to the data processing framework, which may include attached memory storage device 160, or a remote memory storage device accessible via a remote network connection for input to or use by system 100, e.g., databases, vocabulary databases, ontologies, etc.

[0029] 1, processor 152 may include, for example, a microcontroller, a field programmable gate array (FPGA), or any other processor configured to perform various operations. Additionally, communication channels 140, e.g., wired connections such as data bus lines, address bus lines, input / output (I / O) data lines, a video bus, and an expansion bus, are shown for routing signals between various components of system 100. Processor 152 is configured to execute method instructions, as described below. These instructions may be stored, for example, as programmed modules, in associated memory storage device 150.

[0030] Memory 150 may include, for example, non-transitory computer-readable media in the form of volatile memory, such as random access memory (RAM) and / or cache memory. Memory 150 may also include, for example, other removable / non-removable, volatile / non-volatile storage media. By way of non-limiting example only, memory 150 may include a portable computer diskette, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), portable compact disc read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination of the foregoing.

[0031] The network interface 156 is configured to send and receive data or information to and from the website server 120, for example, via a wired or wireless connection. For example, the network interface 156 may utilize wireless technologies and communication protocols such as Bluetooth, WIFI (e.g., 802.11a / b / g / n), cellular networks (e.g., CDMA, GSM, M2M, and 3G / 4G / 4G LTE, 5G), near field communication systems, satellite communications, or any other form of communication that enables the computing device 100 to send and receive information to and from the server 120 over a local area network (LAN), a wide area network (WAN).

[0032] Display 158 may include, for example, a computer monitor, a television, a smart television, a display screen integrated into a personal computing device such as a laptop, a smartphone, a smart watch, a virtual reality headset, a smart wearable device, or any other mechanism for displaying information to a user. In some embodiments, display 158 may include a liquid crystal display (LCD), an e-paper / e-ink display, an organic light-emitting diode (OLED) display, or other similar display technology. In some embodiments, display 158 may be touch-sensitive and may also function as an input device.

[0033] Input device 159 may include, for example, a keyboard, a mouse, a touch-sensitive display, a keypad, a microphone, or other similar input device, or any other input device that may be used alone or together to provide a user with the ability to interact with computing device 100. In one embodiment, a user may enter commands or utterances through the user interface to start or stop system modules and invoke actions on modules for real-time system processing.

[0034] In one embodiment, an audio / video data capture unit 165 is provided for capturing real-time A / V and conference streams from one or more users in a conference, e.g., a virtual or online conference such as ZOOM™ (Zoom Video Communications, Inc.), Webex™ (Cisco Systems, Inc.), GoToMeeting™ (LogMeIn USA, Inc.), etc., for analysis and action item extraction processing according to embodiments herein.

[0035] With regard to configuring the computer system as an analytical tool for summarizing action items in a meeting, including 1) the action items, 2) the relationships between different action items, and 3) the resolutions for each action item, the local or remote memory 160 may be configured to temporarily store or use text data or information 162, including, but not limited to, any captured message threads or message interactions to be NLP analyzed, action items, extracted action item-solution pairs, identified relationships between action items, and templates used to automatically generate summaries.

[0036] Alternatively, this data may be stored on a separate local memory storage device attached to computer system 100.

[0037] As shown in FIG. 1, the memory 150 of the computer system 100 further stores processing software or application modules including programmed instructions adapted to configure the tool to provide an analysis of the meeting and a summary of the user's detected action items, their resolutions, and next action items.

[0038] In one embodiment, one of the programmed processing modules stored in the associated memory 150 includes an audio / video control interface module 166 that provides instructions and logic for operational circuitry to initiate capture of real-time conference conversation or online meeting voice or A / V dialogue or similar message medium (e.g., text) input. In one embodiment, the system receives input activation or deactivation signals from a user device via at least one application program interface 166 for controlling A / V recording, e.g., triggering or terminating operation of system modules in response thereto.

[0039] In one embodiment, speech-to-text module 170 provides instructions and logic for operational circuitry to receive input audio speech captured by video capture device 165 and convert the input audio speech into text format, using, for example, IBM's Watson® Speech to Text (a trademark of International Business Machines Corporation), Dragon® speech recognition software (Nuance Communications, Inc.), etc. Module 170 initiates the storage / buffering and conversion of online conference conversations into natural language text data for NLP processing and for use and analysis by other modules that process data according to aspects of the present disclosure. The text output of the dialogue record from the converted speech input format is processed by other system components.

[0040] In one embodiment, another programmed processing module stored in the associated memory 150 of the system tool 100 includes a natural language processing (NLP) tool 175 for implementing operations and techniques for analyzing textual natural language data used by other system components to perform language summarization, translation, named entity recognition, relationship extraction, speech recognition, and topic segmentation, etc. In one embodiment, the NLP processing tool utilizes neural networks, which enable capturing semantic properties of words and increased learning to perform higher level tasks such as question answering.

[0041] Other programmed processing modules stored in the associated memory 150 of the system tool 100 include an action item identification and extraction module 180 whose operating circuitry utilizes logic and instructions to identify action items from a meeting dialogue record or audio transcript. Action items are tasks discussed during a meeting that need to be addressed by individuals or small groups. The action item extraction module utilizes an NLP information extraction system that is applied to the extracted action items. An exemplary NLP information extraction system that may be utilized is described in the reference entitled “Detecting and Summarizing Action Items in Multi-Party Dialogue,” by M. Purvery, J. Dowding, J. Niekrasz, P. Ehlen, S. Noorbaloochi, and S. Peters, SIGDIAL 2007, the entire contents and disclosure of which are incorporated herein by reference.

[0042] Another programmed processing module stored in the associated memory 150 of the system tool 100 includes an action item solution extraction module 184 that utilizes instructions to configure the system to utilize NLP tools and ML models to detect the presence of solutions related to a user's action items discussed during a meeting.

[0043] Further programmed processing modules stored in the associated memory 150 of the system tool 100 include an action item relationship extraction module 188 that executes instructions to configure the system to utilize NLP tools and ML models to detect the existence of relationships between detected action items, for example, an action item that must be performed as a prerequisite for resolving another action item of the user that was discussed during a meeting.

[0044] An additional programmed processing module 190 stored in the associated memory 150 of the system tool 100 utilizes instructions to configure the system to access and / or execute supervised (or unsupervised) machine learning (ML) models used by other system components. For example, the ML model module 190 invokes operations to train a machine learning model to learn next items related to currently identified action items. Machine learning models may include, but are not limited to, classifiers, regression models, knowledge graphs, or neural networks. For example, supervised learning (classification) techniques using regression models to predict causal relationships between action items or suggested next action items related to the occurrence of a particular action item, and / or unsupervised learning (clustering) techniques may be used.

[0045] Further programmed processing modules stored in the associated memory 150 of the system tool 100 include a summary generator 195 that utilizes instructions to access a pre-configured template, populate the template with extracted information related to the action item associated with the user, including any detected solutions for resolving the action item and any detected related action items or prerequisites that must be performed to resolve another action item of the user that was discussed during the meeting, and configure the system to present the generated summary to the user in the meeting when requested.

[0046] 1, memory 150 includes a monitor program 110 having instructions for configuring computing system 100 to invoke each program module and perform the operations described herein. In one embodiment, such monitor program 110 invokes methods and provides an application program interface for controlling, in real time, receiving real-time input A / V signals, interaction records, text, etc., determining action items from the text, determining action item-user pairings, determining action item relationships, determining resolutions for delayed user or meeting attendee action items, and generating summaries.

[0047] FIG. 2 illustrates an exemplary flow of operations 200 performed by the computing system tool 100 of FIG. 1 according to a monitoring program 110 for controlling the recording, processing, and analysis of meeting speech data to generate real-time action item summaries for users.

[0048] 2, a video camera or similar recording device is physically located where a user is conducting a conference and activated to generate an input stream 202, such as an audio or A / V stream, containing what is being said by the attendees (users) at the conference. An A / V or audio capture device 165 is configured to receive / capture the audio or A / V stream. In one embodiment, the user or attendee can dynamically adjust the recording of the conference audio or A / V stream via the audio / video stream capture processor 165 depending on the user's information needs.

[0049] That is, the system receives an input activate or stop signal 205 communicated from a user device via at least one application program interface 166 for controlling A / V recording to trigger an operation to activate a system module and invoke an operation in that module, for example, for real-time A / V signal capture, text capture, action item determination, etc., while the user is absent from the conference, or to trigger an operation to stop such action item determination operation while the user is present in the conference.

[0050] For example, if a user wants to summarize action items for only a small portion of a meeting, if the user needs to be away for a few minutes, cannot arrive on time, or has to leave before the meeting ends, the input can be adjusted accordingly by the user invoking a start / stop or switch function that keeps the user up to date with the meeting.

[0051] In such an embodiment, the switch 210 initiates generation of a multimedia recording and interaction record of the conference in response to an input activation or stop signal 205. For example, a user or attendee can select a “activate” button via a user device interface to begin recording the conference A / V media stream before the user arrives at the conference or during a period when the user must leave the conference. The conference can be continuously recorded while the user or attendee is absent from the conference. Upon arriving at the conference, the user or attendee may select a “stop” button to terminate recording of the conference. At such time, the system 100 then automatically invokes operations to process the conference recording and generate a summary of action items for the corresponding period when the user or attendee was absent. The action item summary generation system may be activated and deactivated by issuing a “wake” word via a microphone or similar audio interface device, rather than by selecting an activation / stop button via an interface.

[0052] Alternatively, in an embodiment, by default or based on the consent of the participants, the A / V stream 202 of the conference is recorded regardless of the presence / status of either the attendees or conference participants. Then, after the conference ends, the system generates a summary of each user's action items for the corresponding period.

[0053] Whether recording is performed for only a small portion of the conference, selectively activated / deactivated by the user, or the entire A / V stream content of the conference, the recorded speech data is input to a real-time speech-to-text converter unit 170. This speech-to-text converter unit 170 is trained to recognize each speaker's voice, transcribe the user's speech into text, and associate the text of the transcribed speech content spoken in the conference with each user. The system can generate and / or store speech transcriptions of each speaker's speech content, e.g., Speakers A-D, in real time, in the order in which they were spoken by each speaker, as follows: Speaker A: Content A Speaker B: Content B Speaker A: Content A' Speaker C: Content C Speaker D: Content D ...

[0054] In one embodiment, the action item identifier module 180 is then invoked to identify action items from the stored meeting audio transcript or dialogue record. In one embodiment, the module 180 can invoke the operation of a hierarchical classifier including several sub-classifier models in a two-stage process that includes: 1) detecting sub-dialogues (i.e., short sentences of conversation) of the meeting in which action items are proposed, discussed, and committed; and 2) extracting phrases that accurately capture or summarize the tasks contained in the sub-dialogues. In one embodiment, the detection of sub-dialogues is performed by using a hierarchical classifier including multiple independent sub-classifiers used to detect utterances within the dialogue that express action items and user proposals, agreements, or commitments. A semantic parser is then used to extract phrases from the utterances that describe tasks and the timeframes within which the tasks should be performed. For example, to extract the timeframes and task descriptions, a rule-based NLP semantic parser is used that can find basic predicate-argument structures of key phrase types (e.g., symbols, noun phrases, verb phrases, prepositional phrases). In one embodiment, four independent classifier models 190, e.g., linear kernel support vector machines, are implemented for task detection based on a set of distinct action item specific dialog act (AIDA) utterance classes: description of the task to be performed, discussion of a time frame, ownership or assignment of duties, and any explicit agreement. The four independent AIDA subclassifiers can be trained based on features derived from various characteristics of the contextual utterances (e.g., n-grams, durational and positional features from speech transcriptions, prosodic features, etc.). A superclassifier is then used to detect time periods containing the captured subdialogue sequences based on the task utterances detected by the individual AIDA subclassifiers.

[0055] For each user, the system generates and records in real time each action item identified by the action item identifier unit 180, including the action item owner's identification information, task description, and time frame specification. For example, module 180 determines, for identified speaker A, a list of action items, for example: Action item 1 Action item 2 Action item 3 ...

[0056] In one embodiment, for example, determined sentiment, semantic, or lexical analysis of the meeting dialogue can be used to detect the urgency of action items, which can be ordered according to the determined urgency.

[0057] Once the action item identifier unit 180 identifies each action item of a specific user / speaker, the system invokes an operation to determine the existence of a solution for each action item by the target user / speaker and extract a solution for each action item from the meeting audio transcription. For action item solution extraction, the action item solution extraction unit 184 converts the action item into a question answering (QA) task. That is, if the extracted action item is converted as a question and based on the meeting audio transcription as the context, the task is to extract a span from the context as the answer. Furthermore, a specific user / speaker responsible for realizing the solution for the specific action item is identified. In one embodiment, a QA model trained on the Stanford Question Answering Dataset (SQuAD) dataset (https: / / rajpurkar.github.io / SQuAD-explorer / ) can be transferred to perform the QA task. The SQuAD dataset has a large number of answers associated with the corresponding question. One embodiment is to train a word embedding model on the text to which the question is related, vectorize the questions and answers using the word embedding model, and finally train a machine learning model for the QA task using the vectorized questions and answers. That is, the SQuAD dataset can be used to train a QA model or QA tool. An example output of the action item solution extraction unit 184 is as follows: Action Item 1: Solution 1 by Speaker A / B, addressed Action Item 2: Solution 2 by Speaker C, to be implemented by Speaker D Action item 3: Unresolved ...

[0058] To improve performance, users annotate the dataset (e.g., speech transcript-action item-solution) as feedback to further fine-tune the QA model and improve the model's performance. As an example, assume that a list of action items indicates "This action item is unresolved and assigned to speaker A." However, if that is incorrect, speaker A provides feedback by stating "Speaker B has proposed a solution and this action item has already been resolved." This feedback is used to retrain the QA model.

[0059] In one embodiment, in parallel with (or without) identifying solutions for each action item assigned to a particular user / actor by action item solution extraction unit 184, the system invokes an operation to determine the existence of action items related to the action item of interest. That is, system tool 100 simultaneously invokes action item relationship extraction unit 188 to identify dependencies between action items. Such dependencies may, for example, determine that one action item is a prerequisite for another action item.

[0060] In one embodiment, the NLP textual entailment system detects whether a first action item, e.g., Action 1, "implies" a second action, e.g., Action 2. If such an entailment is found to exist, the resolution of Action 2 is determined dependent on first resolving Action 1. That is, Action item 1←Action item 2

[0061] In one embodiment, once the action item / solution extraction unit 184 and the action item relationship extraction unit 188 identify and extract user / speaker action item / solution pairs and identify relationships between specific user / speaker action items, the system invokes an operation to generate a summary report for the user. The summary generation module 195 generates a human-readable summary of the meeting action items based on a predefined template having selectable statements with fields (e.g., ***) that can be populated with corresponding information. In one embodiment, based on the extracted action item / solution pairs and the identified relationships between the action items, the module populates predefined template fields with information and generates a summary output 198 for presentation to the user, for example, via the user's device display. In a non-limiting example, the predefined template can populate statements with detailed information extracted from the action item / solution / relationship units. Examples of such selectable statements with populating fields include, but are not limited to, the following: "There are *** action items that have been discussed so far." "The first action item is ***, spokesperson ***" "And Speaker *** proposed a solution and this action item has been addressed." "The second action item is *** and depends on the first action item." "Speaker *** proposes a solution," "And that solution is planned to be implemented, or that solution is planned to be implemented by the speaker ***." "The third action item is ***," "Currently, no solutions are being discussed."

[0062] In one embodiment, system tool 100 automatically selects among these statements, populates the statements with any relevant information found, and provides the user with a summary output 198 describing the action items discussed in the meeting during the time period, as well as their respective solutions and related items, for example, in response to the user's start / stop control input. Thus, when a user joins a meeting, or returns to the meeting after being away from the meeting for a period of time, the user can immediately see the user's action items, and therefore responsibilities, that were discussed during that time period.

[0063] 3A-3B illustrate an exemplary automated method 300 performed by the monitoring program 110 to generate a summary of action items for a group of multiple users.

[0064] In a first step 303, during a conference between a group of users / attendees, a determination is made that a user-inputted stop control signal has not been received or that an activation signal has been received. In response, in 305, the system begins real-time recording / capture of what is being said in the group conference and, in 310, begins real-time speech-to-text conversion and text storage. A determination is then made continuously in 315 whether a user-inputted interrupt stop signal has been received to initiate cessation of A / V stream reception, recording, and speech conversion. If an interrupt signal is not received in 315, the method continues recording the conference, as indicated by a continuous return to step 305. Otherwise, if a stop signal is received in 315, or by default, the group conference has ended, the process proceeds to 320, where NLP and hierarchical classifier processing is invoked to identify action items for each user in the action item detector module 170.

[0065] In alternative embodiments, the received input may include natural language text dialogue recordings from inputs including, but not limited to, instant message exchanges, messages in a message board, messages received in a chat room or message board, or text converted from speech from a meeting scenario, or a combination thereof, or other meetings where text data is extracted and NLP is applied, including, but not limited to, audio transcriptions, email threads, social media messages / texts, video streams.

[0066] The process continues from step 320 of FIG. 3A along simultaneous paths to steps 330 of FIG. 3B and 333 of FIG. 3B. For each action item determined to be owned by the user at step 330 of FIG. 3B, the action item solution extraction unit 184 generates a question based on the action item at step 335. For example, the module extracts "action item 1" and rephrases it into a corresponding question, such as "What is the solution to action item 1?" At step 340, a question-answering model is then executed to obtain a solution for the action item from the context determined from the audio transcription of the meeting dialogue. The formulated question can be parsed into various keywords and sentence fragments using any question-answering system, such as IBM's Watson deep QA architecture, to find statistically relevant phrases from the meeting audio transcription history record. Continuing at step 345 of FIG. 3B, a determination is made as to whether there are any more action items for the user. If there are more action items for the user, the process proceeds to 350, where the next action item is retrieved, and the process returns to 335 in Figure 3B, where steps are performed to form an associated question from the action item and a QA process to again find an answer to that question, i.e., a solution for the action item. The steps up to 345 are repeated for each action item found for the user until it is determined at 345 that there are no more action items for the user, at which point the process proceeds to a summary generator step at 375, where a template is obtained and a summary is generated using the template by populating the template with information such as the list of action items, the identified authors, and the corresponding solutions, i.e., how the user would resolve the action items if they were unresolved.

[0067] 3A, the process simultaneously performs steps to obtain combinations of pairs of action items. For three action items, such pairs may include, for example: Action item 1 → Action item 2 Action item 1 → Action item 3 Action item 2 → Action item 3 ...

[0068] Then, at 338, for each pair of action items relevant to the user, the method proceeds to 343 of FIG. 3B, where the method executes an NLP entailment model to determine whether a relationship exists between the pair of action items or whether no relationship exists (uncorrelated or neutral). Specifically, at 343, the NLP tool finds directional relationships between text fragments, e.g., whether one pair of action items can be inferred from another pair of action items. This relationship holds whenever the truth of one text fragment is derived from another text. In an NLP entailment framework, the entailing text is the action item text, the entailed text is another action item (hypothesis), and the NLP model determines whether the hypothesis is entailed by the text. Using a dataset of entailment-labeled examples, a statistical classifier, e.g., a logistic regression classifier, is trained using techniques from supervised machine learning. The input to the classifier is two sentences or phrases (e.g., S1—file a reimbursement claim, S2—file an expense claim), and the output of the classifier is the entailment relationship between S1 and S2. If the classifier predicts that S1 entails S2, it means that action S2 occurs before action S1.

[0069] Continuing at 348 of Figure 3B, a determination is made as to whether there are more action item pairs for the user. If there are more action item pairs for the user, the process proceeds to 360, where the next action item pair is retrieved, and the process returns to 343 of Figure 3B, where the step of running an NLP entailment model based on the text fragment of the next action item pair is again performed to determine the relationship of the action item pair. Steps 345-360 are repeated for each action item pair found for the user, until it is determined at 348 that there are no more action item pairs to process for the user's entailments, at which point the process proceeds to the summary generator step at 375, where a template is obtained and the associated action items are added to the user's summary.

[0070] 3B 375, the summary generated by system 100 is output to the user's designated device for presentation to the user. This output may include a summary of the interaction record for the particular user, including a summary of action items with details about who action items are assigned to, which action items are completed (completed) or unresolved, and the relationships between these action items (e.g., dependencies, who needs to work on unresolved items, and who worked on completed items and how they were addressed).

[0071] In further embodiments, the system can receive feedback manually entered by users into the system to respond to other users and their actions. For example, to improve performance of a QA task, users can annotate the dataset (audio transcription-action items-solutions) to further fine-tune and improve the QA model. One embodiment is to use action items that appear in audio transcriptions of actual meetings in the system. Based on user feedback on the action items, an additional training set is generated. This training set is used to further train the QA model.

[0072] FIG. 4 illustrates an exemplary computing system according to the present invention. It should be understood that the illustrated computer system is merely one example of a suitable processing system and is not intended to suggest any limitation as to the scope of use or functionality of embodiments of the present invention. For example, the illustrated system may be operational with numerous other general-purpose or special-purpose computing system environments or configurations. Examples of well-known computing systems, environments, or configurations, or combinations thereof, that may be suitable for use with the system illustrated in FIG. 4 may include, but are not limited to, personal computer systems, server computer systems, thin clients, thick clients, handheld or laptop devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputer systems, mainframe computer systems, and distributed cloud computing environments that include any of the above systems or devices.

[0073] In some embodiments, a computer system may be described in the general context of computer system-executable instructions embodied as program modules stored in memory 16 that are executed by the computer system. Generally, program modules may include routines, programs, objects, components, logic, data structures, etc. that perform particular tasks and / or implement particular input data and / or data types in accordance with the present invention (see, e.g., FIGS. 3A-3B).

[0074] Components of a computer system may include, but are not limited to, one or more processors or processing units 12, memory 16, and a bus 14 that operatively couples various system components, including memory 16, to the processor 12. In some embodiments, the processor 12 may execute one or more modules 11 that are loaded from memory 16, the program modules embodying software (program instructions) that cause the processor to perform one or more method embodiments of the present invention. In some embodiments, the modules 11 may be programmed within integrated circuits of the processor 12 and loaded from memory 16, a storage device 18, or a network 24, or a combination thereof.

[0075] Bus 14 may represent any one or more of several types of bus structures, including a memory bus or memory controller, a peripheral bus, an accelerated graphics port, and a processor or local bus using any of a variety of bus architectures, including, by way of example and not limitation, an Industry Standard Architecture (ISA) bus, a Micro Channel Architecture (MCA) bus, an Enhanced ISA (EISA) bus, a Video Electronics Standards Association (VESA) local bus, and a Peripheral Component Interconnect (PCI) bus.

[0076] The computer system may include a variety of computer system readable media, which can be any available media that can be accessed by the computer system and can include both volatile and nonvolatile media, removable and non-removable media.

[0077] Memory 16 (sometimes referred to as system memory) may include computer-readable media in the form of volatile memory, such as random access memory (RAM), cache memory, or other forms or combinations thereof. The computer system may also include other removable / non-removable, volatile / non-volatile computer system storage media. By way of example only, storage system 18 may be provided for reading from and writing to non-removable, non-volatile magnetic media (e.g., a "hard drive"). Although not shown, a magnetic disk drive may be provided for reading from and writing to removable, non-volatile magnetic disks (e.g., a "floppy disk"), and an optical disk drive may be provided for reading from or writing to removable, non-volatile optical disks, such as CD-ROMs, DVD-ROMs, or other optical media. In such cases, each may be connected to bus 14 by one or more data media interfaces.

[0078] The computer system may also communicate with one or more external devices 26, such as a keyboard, pointing device, display 28, one or more devices that allow a user to interact with the computer system, or any device that allows the computer system to communicate with one or more other computing devices (e.g., network card, modem, etc.), or a combination thereof. Such communication may occur through input / output (I / O) interface 20.

[0079] Additionally, the computer system may communicate with one or more networks 24, such as a local area network (LAN), a general wide area network (WAN), and / or a public network (e.g., the Internet), via a network adapter 22. As shown, the network adapter 22 communicates with other components of the computer system via a bus 14. Although not shown, it should be understood that other hardware and / or software components may be used in conjunction with the computer system. Examples include, but are not limited to, microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data archive storage systems.

[0080] The present invention may be a system, method, and / or computer program product at any possible level of integration of technical details, and may include a computer-readable storage medium (or multiple computer-readable storage media) having computer-readable program instructions for causing a processor to implement aspects of the present invention.

[0081] A computer-readable storage medium may be a tangible device capable of retaining and storing instructions for use by an instruction-execution device. A computer-readable storage medium may be, for example, but not limited to, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the above. A non-exhaustive list of more specific examples of computer-readable storage media includes the following: portable computer diskettes, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), portable compact disk read-only memory (CD-ROM), digital versatile disks (DVDs), memory sticks, floppy disks, mechanically encoded devices such as punch cards or ridge structures in grooves with instructions recorded on them, and any suitable combination of the above. As used herein, computer-readable storage media should not be construed as being ephemeral signals per se, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through a waveguide or other transmission medium (e.g., light pulses through fiber optic cable), or electrical signals transmitted over electrical wires.

[0082] The computer-readable program instructions described herein may be downloaded from a computer-readable storage medium to each computing / processing device or to an external computer or storage device via a network, such as the Internet, a local area network, a wide area network, or a wireless network, or a combination thereof. The network may include copper transmission cables, optical fiber transmissions, wireless transmissions, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions for storage in a computer-readable storage medium within the respective computing / processing device.

[0083] The computer-readable program instructions for carrying out the operations of the present invention may be source or object code written in any combination of one or more programming languages, including assembler instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state setting data, configuration data for integrated circuits, or object code written in any combination of one or more programming languages, including object-oriented programming languages ​​such as Smalltalk, C++, and procedural programming languages ​​such as the "C" programming language or similar programming languages. The computer-readable program instructions may execute entirely on the user's computer as a standalone software package, partially on the user's computer, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the latter scenario, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet service provider). In some embodiments, electronic circuitry including, for example, a programmable logic circuit, a field programmable gate array (FPGA), or a programmable logic array (PLA) may execute computer readable program instructions by utilizing state information of the computer readable program instructions to personalize the electronic circuitry to perform aspects of the present invention.

[0084] Aspects of the present invention are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.

[0085] These computer-readable program instructions may be provided to a processor of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute on the processor of the computer or other programmable data processing apparatus, produce means for performing the functions / acts specified in one or more blocks of the flowcharts and / or block diagrams.The computer-readable program instructions may also be stored on a computer-readable medium, such that the computer-readable storage medium on which the instructions are stored comprises an article of manufacture containing instructions implementing aspects of the functions / acts specified in one or more blocks of the flowcharts and / or block diagrams, and can direct a computer, programmable data processing apparatus, or other device, or combination thereof, to function in a particular manner.

[0086] The computer-readable program instructions may also be loaded into a computer, other programmable data processing apparatus, or other device to create a computer-implemented process that causes the computer, other programmable apparatus, or other device to perform a series of operational steps, such that the instructions, which execute on the computer, other programmable apparatus, or other device, perform the functions / acts specified in one or more blocks of the flowcharts and / or block diagrams.

[0087] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in the flowcharts or block diagrams may represent a module, segment, or portion of instructions, including one or more executable instructions for implementing a specified logical function. In some alternative implementations, the functions noted in the blocks may occur out of the order noted in the figures. For example, two blocks shown in succession may, in fact, be executed substantially concurrently, or the blocks may even be executed in the reverse order, depending on the functionality involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented by a dedicated hardware-based system that performs the specified functions or actions or a combination of dedicated hardware and computer instructions.

[0088] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the invention. As used herein, the singular forms "a," "an," and "the" are intended to include the plural forms unless the context clearly dictates otherwise. It will be further understood that the terms "comprises" and / or "comprising," when used herein, indicate the presence of stated features, integers, steps, operations, elements, or components, or combinations thereof, but do not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components, or groups or combinations thereof. Corresponding structure, materials, acts, and equivalents of all elements within the scope of the appended claims are intended to include any structure, material, or act for performing a function in combination with other claimed elements when specifically claimed. The description of the invention has been presented for purposes of illustration and description, but is not intended to be exhaustive or to limit the invention to the form disclosed. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the invention. The embodiments have been chosen and described in order to best explain the principles and practical applications of the invention and to enable those skilled in the art to understand the invention in various embodiments with various modifications suited to the particular uses contemplated.

[0089] Although this disclosure includes detailed descriptions of cloud computing, it should be understood that implementation of the teachings described herein is not limited to cloud computing environments. Rather, embodiments of the present invention may be implemented in conjunction with any other type of computing environment now known or later developed.

[0090] Cloud computing is a service delivery model for enabling convenient, on-demand network access to a shared pool of configurable computing resources (e.g., networks, network bandwidth, servers, processing, memory, storage, applications, virtual machines, and services) that can be rapidly provisioned and released with minimal administrative effort or interaction with a service provider. This cloud model may include at least five characteristics, at least three service models, and at least four deployment models.

[0091] The features are as follows:

[0092] On-Demand Self-Service: Cloud consumers can unilaterally provision computing capacity, such as server time and network storage, automatically as needed, without the need for human interaction with the service provider.

[0093] Broad network access: Functionality is available over the network and accessed through standard mechanisms that facilitate use by heterogeneous thin or thick client platforms (e.g., mobile phones, laptops, and personal digital assistants (PDAs)).

[0094] Resource Pooling: Provider computing resources are pooled to serve multiple consumers using a multi-tenant model, with different physical and virtual resources dynamically allocated and reallocated according to demand. Location independence exists in that consumers generally have no control or knowledge of the exact location of the resources provided, although at a higher level of abstraction they may specify a location (e.g., country, state, or data center).

[0095] Rapid Flexibility: Capabilities can be rapidly and flexibly, sometimes automatically, provisioned to quickly scale out and rapidly released to quickly scale in. To the consumer, the capabilities available for provisioning often appear unlimited and can be purchased at any time and in any quantity.

[0096] Service Metering: Cloud systems automatically control and optimize resource usage by leveraging metering capabilities at a certain level of abstraction appropriate to the type of service (e.g., storage, processing, bandwidth, and active user accounts). Resource usage can be monitored, controlled, and reported, providing transparency to both providers and consumers of the services they utilize.

[0097] The service model is as follows:

[0098] Software as a Service (SaaS): The ability to offer consumers the ability to use a provider's applications running on a cloud infrastructure. The applications are accessible from a variety of client devices through thin-client interfaces such as web browsers (e.g., web-based email). With the expected exception of limited user-specific application configuration settings, the consumer does not manage or control the underlying cloud infrastructure, including the network, servers, operating systems, storage, or individual application features.

[0099] Platform as a Service (PaaS): The capability offered to consumers to deploy consumer-generated or acquired applications generated using programming languages ​​and tools supported by the provider onto a cloud infrastructure. The consumer does not manage or control the underlying cloud infrastructure, such as the network, servers, operating systems, or storage, but does control the deployed applications and, in some cases, the configuration of the environment that hosts the applications.

[0100] Infrastructure as a Service (IaaS): The capability offered to consumers to provision processing, storage, network, and other basic computing resources on which they can deploy and run any software, which may include operating systems and applications. Consumers do not manage or control the underlying cloud infrastructure, but rather control the operating systems, storage, deployed applications, and possibly limited control over selected networking components (e.g., host firewalls).

[0101] The deployment model is as follows:

[0102] Private Cloud: Cloud infrastructure is operated exclusively for an organization. This cloud infrastructure may be managed by the organization or a third party and may exist on-premise or off-premise.

[0103] Community Cloud: Cloud infrastructure is shared among multiple organizations to support a specific community of shared interests (e.g., mission, security requirements, policies, and compliance considerations). This cloud infrastructure may be managed by the organization or a third party and may exist on-premises or off-premises.

[0104] Public Cloud: Cloud infrastructure is available to the general public or large industry groups and is owned by organizations that sell cloud services.

[0105] Hybrid Cloud: A cloud infrastructure is a composition of two or more clouds (private, community, or public) that remain unique entities but are tied together by standardized or proprietary technologies that allow for data and application portability (e.g., cloud bursting for load balancing between clouds).

[0106] A cloud computing environment is a service that emphasizes statelessness, low coupling, modularity, and semantic interoperability. At the heart of cloud computing is an infrastructure that includes a network of interconnected nodes.

[0107] Referring now to FIG. 5, an exemplary cloud computing environment 50 is shown. As shown, the cloud computing environment 50 includes one or more cloud computing nodes 10 with which local computing devices used by cloud consumers, such as, for example, a personal digital assistant (PDA) or mobile phone 54A, a desktop computer 54B, a laptop computer 54C, or an automobile computer system 54N, or combinations thereof, can communicate. The nodes 10 may also communicate with each other. The nodes 10 may be physically or virtually grouped in one or more networks, such as a private cloud, a community cloud, a public cloud, or a hybrid cloud (not shown), as described herein above. This enables the cloud computing environment 50 to provide infrastructure, a platform, and / or software as a service without requiring the cloud consumer to maintain resources on their local computing device. It should be understood that the types of computing devices 54A-54N shown in FIG. 5 are intended to be illustrative only, and that computing node 10 and cloud computing environment 50 can communicate with any type of computerized device over any type of network and / or network-addressable connection (e.g., using a web browser).

[0108] Referring now to Figure 6, a set of functional abstraction layers provided by cloud computing environment 50 (Figure 5) is shown. It should be understood that the components, layers, and functions shown in Figure 6 are intended to be illustrative only, and embodiments of the present invention are not limited thereto. As shown, the following layers and corresponding functions are provided:

[0109] Hardware and software layer 60 includes hardware and software components. Examples of hardware components include mainframe 61, RISC (reduced instruction set computer) architecture-based servers 62, servers 63, blade servers 64, storage devices 65, and networks and networking components 66. In some embodiments, software components include network application server software 67 and database software 68.

[0110] The virtualization layer 70 provides an abstraction layer from which the following examples of virtual entities may be provided: virtual servers 71, virtual storage 72, virtual networks including virtual private networks 73, virtual applications and operating systems 74, and virtual clients 75.

[0111] In one example, the management layer 80 may provide the following functions: Resource provisioning 81 provides dynamic procurement of computing and other resources utilized to execute tasks within the cloud computing environment. Metering and pricing 82 provides cost tracking as resources are utilized within the cloud computing environment and billing or invoicing for the consumption of these resources. In one example, these resources may include application software licenses. Security provides identity verification for cloud consumers and tasks, as well as protection for data and other resources. User portal 83 provides consumers and system administrators with access to the cloud computing environment. Service level management 84 provides allocation and management of cloud computing resources to ensure requested service levels are met. Service level agreement (SLA) planning and fulfillment 85 provides pre-allocation and procurement of cloud computing resources anticipated to be required in the future according to SLAs.

[0112] Workload tier 90 provides examples of functions for which a cloud computing environment may be utilized. Examples of workloads and functions that may be provided from this tier include mapping and navigation 91, software development and lifecycle management 92, virtual classroom instruction delivery 93, data analytics processing 94, transaction processing 95, and processing for automatically generating summaries and next actions in real time from natural language interaction recordings 96, according to aspects of the present disclosure. [Explanation of symbols]

[0113] 100 systems 100 Action Items and Summary Generator 110 Surveillance Program 130 databases 150 memory 156 Network Interfaces 159 Input Devices 160 Memory Storage Devices 165 Audio / Video Data Acquisition Unit 166 Audio / Video Control Interface Module 170 Speech-to-Text Module 175 Natural Language Processing (NLP) Tools 180 Action Item Identification and Extraction Module 184 Action Item Solution Extraction Module 188 Action Item Relationship Extraction Module 190 machine learning models 195 Summary Generation Module

Claims

1. 1. A computer-implemented method comprising: receiving, at one or more processors, a plurality of interaction records of natural language text associated with a group of users attending a conference; analyzing the natural language text of the plurality of interaction records using one or more classifier models to identify from the natural language text one or more action items associated with each task to be performed; detecting from the interaction record each user assigned to perform each of the tasks; Detecting from the interaction record whether each action item is addressed with a solution for performing the task; generating a text summary of solutions that address each action item of the user; 20. A computer-implemented method comprising:

2. receiving a real-time audio or audio / video signal stream capturing what is being said by users in the conference; converting said real-time audio signal stream or audio / video signal stream into a text representation to form said plurality of dialogue records; The computer-implemented method of claim 1 further comprising:

3. 3. The computer-implemented method of claim 1, further comprising receiving, in the one or more processors, one of a start signal to start the receiving, analyzing, identifying action items, and generating text summaries at the beginning or during the conference, or a stop signal to end the receiving, analyzing, identifying action items, and generating text summaries during the conference.

4. detecting whether each action item from the interaction record is addressed with a solution for performing the task; Converting the detected action items into a question format; forming a question-answering task; extracting ranges from the dialogue recording and executing a trained question-answering model to determine whether an answer to the question exists from the extracted ranges; The computer-implemented method of any one of claims 1 to 3, comprising:

5. 5. The computer-implemented method of claim 4, further comprising using an NLP implication model to identify whether any action item is dependent on another action item or whether an action item associated with a task is a prerequisite for performing another action item task.

6. said generating a text summary of a solution that addresses each action item of said user; obtaining a template having predefined fields for receiving information related to the action items detected in the interaction record; populating said template with action item information including corresponding users who own said action items, and for each action item, specifying in said template information including a solution for said action item, or solutions for any prerequisite action items on which the action item depends, or both; The computer-implemented method of claim 5 further comprising:

7. receiving feedback information from a user based on the input action item information, the feedback information relating to the action item task, the detected corresponding user, and a solution for the action item; using the feedback information to improve tuning of one or more of the one or more classifier models, the trained question-answering model, and the NLP implication model; The computer-implemented method of claim 6 further comprising:

8. 1. A computer-implemented system comprising: a memory storage device for storing a computer readable program; at least one processor adapted to execute the computer-readable program; the computer readable program causing the at least one processor to: receiving a plurality of natural language text interaction records relating to a group of users attending a conference; analyzing the natural language text of the plurality of interaction records using one or more classifier models to identify one or more action items from the natural language text associated with each task to be performed; Detecting from the interaction record each user assigned to perform each of the tasks; Detecting from the interaction record whether each action item is addressed with a solution for performing the task; Generate a text summary of a solution that addresses each of the user's action items. A computer-implemented system configured to:

9. the at least one processor: receiving a real-time audio or audio / video signal stream capturing what is being said by users in the conference; converting the real-time audio or audio / video signal stream into a text representation to form the plurality of dialogue records; 10. The computer-implemented system of claim 8, further configured to:

10. the at least one processor:

10. The computer-implemented system of claim 8 or 9, further configured to receive one of a start signal to start the receiving, analyzing, identifying action items, and generating text summaries of interaction records at the beginning or during the conference, or a stop signal to end the receiving, analyzing, identifying action items, and generating text summaries of interaction records during the conference.

11. to detect from the interaction record whether each action item is addressed with a solution for performing the task, the at least one processor: Convert the detected action items into questions, Form a question-answering task, extracting ranges from the interaction record and running a trained question-answering model to determine whether an answer to the question exists from the extracted ranges; The computer-implemented system of any one of claims 8 to 10, further configured to:

12. the at least one processor:

12. The computer-implemented system of claim 11, further configured to use an NLP implication model to identify whether any action item depends on another action item or whether an action item associated with a task is a prerequisite for performing another action item task.

13. To generate the text summary of solutions that address each action item of the user, the at least one processor: obtaining a template having predefined fields for receiving information related to the action items detected in the interaction record; Populating the template with action item information including corresponding users who own the action items, and for each action item, specifying in the template information including a solution for the action item, solutions for any prerequisite action items on which the action item depends, or both.

13. The computer-implemented system of claim 12, further configured to:

14. the at least one processor: receiving feedback information from a user based on the input action item information, related to the action item task, its detected corresponding user, and a solution to the action item; Using the feedback information to improve tuning of one or more of the one or more classifier models, the trained question-answering model, and the NLP implication model.

14. The computer-implemented system of claim 13, further configured to:

15. A computer program for causing a computer to execute the computer-implemented method according to any one of claims 1 to 7.

16. A computer-readable recording medium on which the computer program according to claim 15 is recorded.

Citation Information

Patent Citations

  • Computerized intelligent assistant for conferences

    CN112075075A

  • Information processor and information processing method

    JP2012053628A

  • Conference support system and conference support program

    JP2019061594A

  • Extraction program, extraction method and extraction device

    JP2019191276A

  • Automatic generation of action items from a meeting transcript

    US20170161258A1