Meeting Transcript Display Control for Precise Audio Playback

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Users find it cumbersome to start playing recorded voice data from a specific point in time during meetings or events, especially when text data is displayed, as selecting the correct playback start time can be difficult without knowing the exact timing of desired content.

Innovation Solution

A display control apparatus and method that utilizes graphical control regions to set playback positions based on the generation time of selected text data, allowing users to easily select and play back voice data from specific points by interacting with displayed text data or screenshot images.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Loss of information

If text data is displayed during meeting playback, then users can reference content while listening, but it becomes cumbersome for users to select specific playback start times

Engineering Contradiction:
Improveinformation retention during playbackVSAvoidplayback position selection
Core Design Contradiction:
Loss of informationVSEase of operation

Solution Approach 1:

The patent introduces text data as an intermediary between the user and the voice recording. By displaying transcribed text with temporal information, users can locate desired content through text search or scanning rather than manually scrubbing through time markers. The text acts as a mediator that bridges the gap between audio content and user navigation needs.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent adds a textual dimension to the audio playback system. Instead of navigating solely through temporal progression of audio, users gain access to a parallel textual representation where they can search, scan, and select content based on semantic meaning rather than temporal position, effectively adding another dimension to the navigation space.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Measurement precision

If manual time selection is required for playback, then precise control is achieved, but user time and effort increase significantly

Engineering Contradiction:
Improveplayback position accuracyVSAvoidtime to locate playback point
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent performs preliminary transcription of voice data into text data before playback occurs. This advance preparation creates a searchable textual index of the audio content, allowing users to quickly locate desired segments through text search or scanning without having to manually navigate through the entire recording or guess time positions.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent provides visual feedback by displaying text data with temporal markers that correspond to specific playback positions. As users interact with the text (selecting, hovering, or searching), the system provides feedback by highlighting corresponding time positions and playing back the associated audio segment, creating a feedback loop that accelerates location of desired content.

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS20250329331A1Apparatus, system, and method of display control, and recording medium
Publication Date: 2025.10.23 RICOH CO LTD
  • US20250329331A1 patent drawing
  • US20250329331A1 patent drawing
  • US20250329331A1 patent drawing

AI summary

A system includes: a server including first circuitry and a memory that stores, for each event, voice data recorded during the event, text data converted from the voice data, and time information indicating a time when the text data was generated; and a display control apparatus communicably connected with the server, including second circuitry to based on information on the event stored in the memory, control a display to display text data in an order according to the time when the text data was generated, and a graphical control region that sets playback position in a total playback time of the voice data, and in response to selection of particular text data from the text data being displayed, control the display to display the graphical control region at a location determined based on a time when the particular text data was generated.