Voice Stream Augmented Note Taking with Speech Recognition

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional note-taking methods, such as shorthand and rapid typing, are difficult to learn and impractical for capturing detailed information during casual conversations or presentations, leading to missed details when reviewing notes later.

Innovation Solution

A voice stream augmented note-taking system that records and converts audio into text chunks, allowing users to match and select relevant phrases while typing, using speech recognition algorithms like Hidden Markov Models to provide autocomplete suggestions based on context and user input.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If conventional note-taking methods (shorthand, rapid typing) are used, then note-taking speed may be improved, but the difficulty of learning and practicality for casual conversations increases

Engineering Contradiction:
Improvenote-taking speedVSAvoidease of learning and practicality
Core Design Contradiction:
ProductivityVSEase of operation

Solution Approach 1:

The patent replaces manual note-taking methods (shorthand, rapid typing) with an automated speech recognition system that converts spoken words into text automatically. This substitution eliminates the need for users to learn difficult note-taking techniques while maintaining high information capture speed during casual conversations and presentations.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Loss of information

If users attempt to include all details while listening to presentations, then information completeness may be improved, but the risk of missing later details due to keeping up with typing increases

Engineering Contradiction:
Improveinformation completenessVSAvoidtime to keep up with typing
Core Design Contradiction:
Loss of informationVSLoss of time

Solution Approach 1:

The speech recognition system performs the note-taking function automatically without requiring user intervention for typing. The system captures and transcribes speech in real-time, allowing users to focus on listening and understanding presentations rather than competing to keep up with transcription speed, thereby preventing information loss.

Inventive Principle:
Principle #25Self-service

3Productivity

If speech recognition algorithms are used to convert audio to text, then note-taking efficiency is improved, but system complexity increases

Engineering Contradiction:
Improvenote-taking efficiencyVSAvoidsystem complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent employs speech recognition algorithms as an intermediary component that bridges the gap between audio input and text output. This intermediary automatically handles the complex task of converting spoken language into written form, improving note-taking efficiency while shielding users from the underlying system complexity through an intuitive interface.

Inventive Principle:
Principle #24Intermediary (Mediator)

Applied Scientific Principles

This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.

Function Achieved in This Case

Enhances note-taking efficiency by providing users with accurate and relevant suggestions from the audio stream, reducing the likelihood of missing important details and improving the accuracy of recorded information during note review.

Implementation Method 1

converts the audio stream into text chunks

Methodology Applied
Scientific EffectSpeech recognition:

Data Source

PatentEP2572355B1Voice stream augmented note taking
Publication Date: 2018.06.27 MICROSOFT TECHNOLOGY LICENSING LLC
  • EP2572355B1 patent drawingFigure 1
  • EP2572355B1 patent drawingFigure 2
  • EP2572355B1 patent drawingFigure 3

AI summary

Voice stream augmented note taking may be provided. An audio stream associated with at least one speaker may be recorded and converted into text chunks. A text entry may be received from a user, such as in an electronic document. The text entry may be compared to the text chunks to identify matches, and the matching text chunks may be displayed to the user for selection.