Collaborative Speech Processing for Meeting Tracking
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Tracking progress and contributions during project tasks in scrum meetings is time-consuming and can disrupt collaborative discussions among team members, hindering effective project management.
Innovation Solution
A collaborative speech processing computer system that uses speech-to-text conversion to identify speakers and associate them with project tasks, allowing for real-time tracking of progress and contributions by forwarding audio streams to a speech-to-text server, comparing spectral characteristics to select the correct speaker, and controlling microphone settings to manage audio inputs.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of information
If manual tracking of progress and contributions is performed during scrum meetings, then project management tracking is achieved, but collaborative discussions are disrupted and time is consumed
Solution Approach 1:
The patent replaces manual mechanical tracking methods with an automated speech processing system that uses audio capture, speech-to-text conversion, and natural language processing to automatically track progress and contributions during scrum meetings, eliminating the need for manual intervention and preserving collaborative discussion flow
Solution Approach 2:
The system enables self-service tracking by automatically capturing, processing, and analyzing meeting discussions without requiring team members to manually track their own progress or contributions, with the system autonomously identifying speakers, converting speech to text, and extracting relevant project information
2Loss of information
If speech-to-text conversion is performed on all audio streams, then complete transcription is achieved, but processing time and computational resources increase
Solution Approach 1:
The patent extracts and processes only relevant audio segments containing actual speech content, using voice activity detection and speaker identification to isolate meaningful contributions from background noise and non-speech audio, thereby reducing overall processing time while maintaining conversion completeness for relevant content
Solution Approach 2:
The system performs partial processing by focusing computational resources on identifying and transcribing only the portions of audio that contain relevant speech content, rather than processing entire audio streams uniformly, achieving efficient balance between completeness and processing speed
3Loss of information
If multiple microphones are used to capture audio from different speakers, then audio coverage is improved, but difficulty in identifying which speaker is talking increases
Solution Approach 1:
The patent applies local quality by assigning unique spectral characteristics to each speaker and using these characteristics to identify which speaker is talking at any given time, allowing the system to distinguish between multiple speakers across different microphones by analyzing the local spectral fingerprint of each speaker's voice
Solution Approach 2:
The system introduces spectral characteristic analysis as an intermediary process between audio capture and speaker identification, using voice biometrics to bridge the gap between multiple audio sources and their corresponding speakers, thereby solving the ambiguity of identifying which speaker is talking
Applied Scientific Principles
This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.
Function Achieved in This Case
Enhances project management by efficiently tracking progress and contributions, reducing distractions during meetings, and improving speech recognition through targeted audio processing and speaker identification.
Implementation Method 1
The sampled audio streams are forwarded to a speech-to-text conversion server via a data network. Data packets are received, via the data network, which contain text strings converted from the sampled audio steams by the speech-to-text conversion server.
Implementation Method 2
Spectral characteristics of a voice contained in the sampled audio stream that was converted to the one of the text strings is compared to known spectral characteristics that are defined for the candidate speakers in the group. One person is selected as the speaker from among the candidate speakers in the group, based on a relatively closeness of the comparisons of spectral characteristics.
Data Source
AI summary
A collaborative speech processing computer receives packets of sampled audio streams. The sampled audio streams are forwarded to a speech-to-text conversion server via a data network. Packets are received via the data network that contain text strings converted from the sampled audio steams by the speech-to-text conversion server. Speakers are identified who are associated with the text strings contained in the data packets. The text strings and the identifiers of the associated speakers are added to a dialog data structure in a repository memory. Content of at least a portion of the dialog data structure is displayed on a display device.


