Voice Search Annotation Breakpoint Navigation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing voice-activated systems struggle to efficiently navigate and select specific portions of digital content, such as videos, leading to wasteful network transmissions and computational resources.

Innovation Solution

A multi-modal interface system that enables users to interact with digital content through both touch interfaces and voice commands, utilizing semantic processing and annotation techniques to identify break points within content, allowing for precise selection and playback of specific video segments.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Loss of energy

If voice-based interfaces transmit the entire digital content to the client device, then the user can access the content, but network resources and computational resources are wasted

Engineering Contradiction:
Improvenetwork resource wasteVSAvoidvoice-based navigation capability
Core Design Contradiction:
Loss of energyVSEase of operation

Solution Approach 1:

The patent segments digital content (videos, images, audio) into smaller portions using detected break points. Instead of transmitting entire content files, only relevant segments are transmitted to client devices, reducing network resource consumption while maintaining user access capability through voice-based navigation.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system extracts and transmits only the necessary portions of digital content based on voice commands and detected break points. This extraction approach removes unnecessary data transmission, directly addressing the network resource waste problem while preserving essential content delivery.

Inventive Principle:
Principle #2Taking out (Extraction)

2Loss of energy

If the system processes and transmits only specific video segments, then network resources are saved, but the system complexity increases

Engineering Contradiction:
Improvenetwork resource wasteVSAvoidsystem complexity
Core Design Contradiction:
Loss of energyVSDevice complexity

Solution Approach 1:

The system performs preliminary processing to detect break points and segment content before user requests. By pre-processing and organizing content into segments with identified break points, the system reduces the complexity of real-time processing while enabling efficient segment-based transmission that saves network resources.

Inventive Principle:
Principle #10Preliminary action

3Ease of operation

If voice commands are used to navigate digital content, then user interaction improves, but the ability to precisely locate and jump to specific content portions is limited

Engineering Contradiction:
Improveuser interactionVSAvoidcontent location precision
Core Design Contradiction:
Ease of operationVSMeasurement precision

Solution Approach 1:

The patent introduces break point detection as an intermediary mechanism between voice commands and content navigation. The system detects break points within content and uses these as precise navigation targets, enabling voice-based interfaces to accurately jump to specific content portions rather than merely playing sequential content.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentEP3685280B1Voice based search for digital content in a network
Publication Date: 2025.03.05 GOOGLE LLC
  • EP3685280B1 patent drawingFigure 1
  • EP3685280B1 patent drawingFigure 2
  • EP3685280B1 patent drawingFigure 3

AI summary

Systems and methods of the present technical solution enable a multi-modal interface for voice-based devices, such as digital assistants. The solution can enable a user to interact with video and other content through a touch interface and through voice commands. In addition to inputs such as stop and play, the present solution can also automatically generate annotations for displayed video files. From the annotations, the solution can identify one or more break points that are associated with different scenes, video portions, or how-to steps in the video. The digital assistant can receive input audio signal and parse the input audio signal to identify semantic entities within the input audio signal. The digital assistant can map the identified semantic entities to the annotations to select a portion of the video that corresponds to the users request in the input audio signal.