Sound Text Synchronization via Transcription Delay Correction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Manual transcription of sound data into text is time-consuming and costly due to the lack of automatic synchronization between sound and text, which is essential for features like synchronous playback, especially when dealing with poor-quality sound or specialized jargon.

Innovation Solution

A method and system that repeatedly query sound and text data to obtain current time positions, apply a time correction value based on transcription delay, and generate association data to synchronize sound and text, allowing for automatic synchronization during manual transcription, even when using speech recognition systems.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If manual transcription is used to obtain text data from sound data, then transcription accuracy can be maintained (especially for poor-quality sound or specialized jargon), but synchronization between sound and text becomes time-consuming and costly

Engineering Contradiction:
Improvetranscription accuracyVSAvoidsynchronization time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system performs preliminary actions by automatically querying sound data to obtain time positions and generating association data during the transcription process itself. The transcription delay is measured and stored in advance, allowing synchronization to be established automatically without requiring manual time-marking operations after transcription is complete.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system makes the transcription process self-servicing by automatically creating synchronization associations between sound and text data. The transcription system itself queries the sound data, measures the transcription delay, generates association data, and establishes time positions without requiring external manual intervention, thereby eliminating the need for separate manual synchronization operations.

Inventive Principle:
Principle #25Self-service

2Loss of time

If automatic speech recognition is used to produce text data, then synchronization can be easily established with timing data, but transcription accuracy deteriorates when dealing with poor-quality sound or specialized jargon

Engineering Contradiction:
Improvesynchronization timeVSAvoidtranscription accuracy
Core Design Contradiction:
Loss of timeVSMeasurement precision

Solution Approach 1:

The system introduces an intermediary approach where association data serves as a mediator between manually transcribed text and sound data. Instead of relying solely on automatic speech recognition (which fails for specialized content) or manual synchronization (which is time-consuming), the association data automatically links text datums to sound segments based on measured transcription delays, combining the accuracy of manual transcription with the efficiency of automatic synchronization.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Measurement precision

If manual synchronization is performed by marking sound segments with precision of a few milliseconds, then synchronization accuracy is improved, but the process becomes very time-consuming and expensive

Engineering Contradiction:
Improvesynchronization precisionVSAvoidtranscription efficiency
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The system makes the synchronization process self-servicing by automatically querying sound data during transcription to obtain time positions and generating association data. The transcription system itself performs the synchronization function without requiring separate manual marking operations, thereby achieving both precision and high productivity simultaneously.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system merges the transcription process with the synchronization process into a single integrated operation. Instead of performing transcription and synchronization as separate sequential steps (which would double the time required), the system combines both functions so that association data is generated automatically during the transcription process itself, maintaining precision while doubling productivity.

Inventive Principle:
Principle #5Merging (Combining)

Data Source

PatentUS8924216B2System and method for synchronizing sound and manually transcribed text
Publication Date: 2014.12.30 MICROSOFT TECHNOLOGY LICENSING LLC
  • US8924216B2 patent drawing
  • US8924216B2 patent drawing
  • US8924216B2 patent drawing

AI summary

A method for synchronizing sound data and text data, said text data being obtained by manual transcription of said sound data during playback of the latter. The proposed method comprises the steps of repeatedly querying said sound data and said text data to obtain a current time position corresponding to a currently played sound datum and a currently transcribed text datum, respectively, correcting said current time position by applying a time correction value in accordance with a transcription delay, and generating at least one association datum indicative of a synchronization association between said corrected time position and said currently transcribed text datum. Thus, the proposed method achieves cost-effective synchronization of sound and text in connection with the manual transcription of sound data.