Speaker-Adaptive Symbol Insertion for Voice Transcription
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing symbol insertion techniques for transcribed voice data fail to accurately insert punctuation marks, such as periods and commas, due to their inability to consider the unique speaking styles and pause patterns of individual speakers, leading to inconsistent evaluation and insertion of symbols.
Innovation Solution
A symbol insertion apparatus that evaluates symbol insertion likelihood using multiple models tailored to specific speaking style features, including linguistic and acoustic characteristics, to determine the appropriate placement of punctuation marks in transcribed voice data.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If a single symbol insertion model is used for all speakers, then the device complexity is reduced, but the measurement precision of symbol insertion accuracy deteriorates due to inability to account for individual speaking styles
Solution Approach 1:
The patent applies local quality by creating separate symbol insertion models for different speakers or speaking styles. Instead of using a single universal model, the system identifies and models the specific characteristics of each speaker's pause patterns and sentence boundary markers, thereby achieving high accuracy for each individual while maintaining the overall system's adaptability.
Solution Approach 2:
The patent utilizes parameter changes by adjusting the model parameters based on speaker-specific characteristics. The system modifies pause length thresholds, sentence boundary criteria, and weighting factors according to the identified speaking style, allowing the same basic model framework to adapt to different speakers without requiring complete model redesign.
2Measurement precision
If pause length thresholds are set to detect sentence boundaries, then symbol insertion accuracy improves, but the adaptability to different speaking styles deteriorates because fixed thresholds cannot accommodate individual variations
Solution Approach 1:
The patent implements dynamics by making the pause length thresholds and detection parameters adaptive rather than fixed. The system dynamically adjusts the thresholds based on the identified speaking style of each speaker, allowing the same detection mechanism to work effectively across different speaking rates and styles while maintaining high accuracy for each individual.
Solution Approach 2:
The patent applies preliminary action by pre-identifying and characterizing each speaker's speaking style before performing symbol insertion. The system analyzes pause patterns and sentence boundary characteristics in advance, storing these profiles for use during the actual symbol insertion process, thereby enabling personalized detection parameters to be applied automatically.
3Measurement precision
If multiple symbol insertion models are created for different speakers, then the measurement precision of symbol insertion improves, but the device complexity increases due to management of multiple models
Solution Approach 1:
The patent applies universality by designing a multi-functional symbol insertion system that can handle multiple speakers through a unified framework. The system uses a common base model that can be adapted to different speakers through parameter adjustment and profile storage, rather than requiring completely separate systems for each speaker, thereby reducing overall complexity while maintaining speaker-specific accuracy.
Data Source
AI summary
Enables symbol insertion evaluation in consideration of a difference in speaking style features between speakers. For a word sequence transcribing voice information, the symbol insertion likelihood calculation means 113 obtains a symbol insertion likelihood for each of a plurality of symbol insertion models supplied for different speaking style features. The speaking style feature similarity calculation means 112 obtains a similarity between the speaking style feature of the word sequence and the plurality of speaking style feature models. The symbol insertion evaluation means 114 weights the symbol insertion likelihood obtained for the word sequence by each of the plurality of symbol insertion models according to the similarity between the speaking style feature of the word sequence and the plurality of speaking style feature models and the relevance between the symbol insertion model and the speaking style feature model, and performs symbol insertion evaluation to the word sequence.


