Text Voice Synchronization via Phoneme Accumulation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing technologies face challenges in synchronously processing text and voice data, especially when the reading speed of voice data varies or silent data is interposed, leading to low quality synchronization.

Innovation Solution

An apparatus that divides and phonemically converts text and voice data, calculates accumulated phoneme conversion values, and produces corresponding data to synchronize text and voice data accurately, even with varying reading speeds and silent segments, using a storing unit, text data dividing section, phoneme converting section, voice data dividing section, and output section.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If text data and voice data are manually synchronized, then synchronization can be performed, but working efficiency is very low

Engineering Contradiction:
Improveworking efficiency of synchronous processingVSAvoidtime for synchronous processing
Core Design Contradiction:
ProductivityVSLoss of time

Solution Approach 1:

The patent replaces manual synchronization (mechanical human operation) with an automatic synchronization system that uses phoneme conversion and time stamp matching algorithms. The synchronization section automatically associates text data with voice data by comparing phoneme conversion results and time stamps, eliminating the need for manual intervention and significantly improving working efficiency.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Reliability

If text data are displayed and voice data are outputted simultaneously, then both data can be presented, but synchronization quality is low when reading speed varies or silent data are interposed

Engineering Contradiction:
Improvesynchronization qualityVSAvoidadaptability to varying reading speeds and silent data
Core Design Contradiction:
ReliabilityVSAdaptability or versatility

Solution Approach 1:

The patent performs phoneme conversion on both text data and voice data in advance, before the actual synchronization and playback. The phoneme conversion results and time stamps are calculated and stored beforehand, allowing the synchronization section to quickly and accurately match text with voice during playback, even when reading speeds vary or silent data are present, thereby maintaining high synchronization quality.

Inventive Principle:
Principle #10Preliminary action

3Measurement precision

If phoneme conversion and accumulated value calculation are performed for both text and voice data, then accurate synchronization can be achieved, but processing complexity increases

Engineering Contradiction:
Improvesynchronization accuracyVSAvoidcomplexity of phoneme conversion and calculation sections
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent employs a unified phoneme conversion approach for both text data and voice data. The phoneme conversion section converts text to phonemes, while the voice data undergoes phoneme recognition to also produce phoneme sequences. This universal phoneme-based methodology allows the synchronization section to use the same comparison and matching logic for both data types, improving synchronization accuracy while managing complexity through methodological consistency.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS9679566B2Apparatus for synchronously processing text data and voice data
Publication Date: 2017.06.13 SHINANO KENSHI CO LTD
  • US9679566B2 patent drawing
  • US9679566B2 patent drawing
  • US9679566B2 patent drawing

AI summary

The apparatus for synchronously processing text data and voice data, comprises: a storing unit for storing text data and voice data; a text data dividing section for dividing the text data; a text data phoneme converting section for phonemically converting the divided text data; a text data phoneme conversion accumulated value calculating section for calculating accumulated values of text data phoneme conversion values; a voice data dividing section for dividing the voice data; a reading data phoneme converting section for phonemically converting the divided voice data; a voice data phoneme conversion accumulated value calculating section for calculating accumulated values of voice data phoneme conversion values; a phrase corresponding data producing section for producing phrase corresponding data; and an output section for synchronously outputting the text data and the divided voice data.