Video Conference Speech Recognition for Action Item Tracking

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Participants in video conferences often struggle to keep track of action items and information discussed during meetings, leading to forgotten tasks and inefficiencies due to manual note-taking and multitasking, which can distract from the meeting itself.

Innovation Solution

A software client that performs speech recognition during video conferences, identifies keywords and contexts, and suggests actions by matching them with predefined rules, allowing users to confirm and execute relevant applications or functionalities directly.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Loss of information

If participants manually take notes and multitask during video conferences, then they can track action items and information, but they become distracted from the meeting itself and experience reduced productivity

Engineering Contradiction:
Improvetracking action items and informationVSAvoidproductivity
Core Design Contradiction:
Loss of informationVSProductivity

Solution Approach 1:

The system enables self-service by automatically transcribing speech, identifying keywords, and suggesting actions without requiring manual note-taking. The computing device autonomously processes audio input and generates action suggestions, freeing participants from manual documentation tasks while maintaining information tracking.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system performs preliminary action by proactively analyzing speech during the meeting and suggesting actions before participants need to manually document them. The system continuously monitors audio input, identifies relevant keywords, and presents action suggestions in real-time, allowing participants to respond without breaking their focus on the meeting.

Inventive Principle:
Principle #10Preliminary action

2Extent of automation

If the system analyzes speech to suggest actions, then automation and productivity improve, but user privacy concerns may arise

Engineering Contradiction:
ImproveautomationVSAvoidprivacy concerns
Core Design Contradiction:
Extent of automationVSObject-affected harmful factors

Solution Approach 1:

The system applies local quality by processing and analyzing only the speech of the local user rather than all participants. The computing device receives audio input, identifies keywords from the user's speech specifically, and generates personalized action suggestions. This localized approach minimizes privacy intrusion while maintaining automation benefits.

Inventive Principle:
Principle #3Local quality

3Reliability

If the system requires user confirmation before executing actions, then privacy and control are maintained, but response time and efficiency may be reduced

Engineering Contradiction:
Improveuser controlVSAvoidresponse time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The system applies partial action by requiring confirmation only for suggested actions rather than executing all possible actions automatically. This selective confirmation approach balances automation efficiency with user control, allowing rapid processing of obvious actions while seeking validation for more significant or ambiguous actions.

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS12014731B2Suggesting user actions during a video conference
Publication Date: 2024.06.18 ZOOM VIDEO COMM INC
  • US12014731B2 patent drawing
  • US12014731B2 patent drawing
  • US12014731B2 patent drawing

AI summary

One example method includes receiving, by a computing device, audio during a video conference having a plurality of participants, the audio comprising spoken words by a user of the computing device; recognizing one or more words from the spoken words; identifying one or more keywords within the one or more recognized words; accessing a set of rules comprising one or more rules, each rule of the one or more rules associated with an application of a set of applications, and at least one rule of the one or more rules associated with a functionality of a respective application; determining a context associated with the one or more keywords; determining an application to execute based on the one or more keywords, the context, and the one or more rules, wherein determining the application comprises determining a functionality of the application to invoke; and in response to receiving user confirmation of the functionality of the application to invoke, executing the application and invoking the functionality.