Voice-Controlled Presentation Navigation via Speech Recognition

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current methods for controlling presentations, such as using pointers, mice, or keyboards, are inefficient and distract presenters from focusing on their audience and content, as they require manual navigation through slides during a presentation.

Innovation Solution

A computer-implemented method and system that utilizes Automated Speech Recognition (ASR) and Natural Language Processing (NLP) to continuously transcribe voice inputs and detect gestures, allowing for the selection and execution of presentation item controls, such as moving to specific slides, through voice commands and gestures, thereby streamlining the navigation process.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If manual control devices (pointer, mouse, keyboard) are used to navigate presentation slides, then presentation control functionality is achieved, but presenter attention is diverted from the audience and time is lost during navigation

Engineering Contradiction:
Improvepresentation control efficiencyVSAvoidtime lost during slide navigation
Core Design Contradiction:
Ease of operationVSLoss of time

Solution Approach 1:

The patent replaces mechanical control devices (mouse, keyboard, pointer) with voice-based control systems. The system captures voice input, transcribes it to text using speech-to-text technology, and processes the transcription to execute presentation control commands. This substitution eliminates the need for manual device operation, allowing presenters to control slides through natural speech while maintaining eye contact with the audience and eliminating navigation time delays.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Adaptability or versatility

If voice-to-text transformation and speech processing are implemented, then hands-free navigation is enabled, but system complexity increases

Engineering Contradiction:
Improvehands-free control capabilityVSAvoidsystem architecture complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent implements a multi-functional system that combines voice capture, speech-to-text transformation, transcription processing, and presentation control execution within a single integrated architecture. The voice input processing system serves multiple purposes: it transcribes speech for real-time display, processes the transcription to identify control commands, and executes appropriate presentation actions. This universal approach enables hands-free navigation while managing system complexity through functional integration rather than separate dedicated components for each task.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentEP3910626A1Presentation control
Publication Date: 2021.11.17 DEUTSCHE TELEKOM AG
  • EP3910626A1 patent drawingFigure 1
  • EP3910626A1 patent drawingFigure 2
  • EP3910626A1 patent drawingFigure 3

AI summary

A presentation control method and system, wherein the method comprises the steps of: continuously running a voice to text transformation of a voice input from a user; detecting at least one spoken expression in the voice input; selecting a specific presentation item out of a set of presentation items based on the detected spoken expression; and executing an animation and/or change to the presentation item.