Text Presentation Apparatus for Voice Recording Script Adaptation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional text speech synthesis technologies face challenges in collecting high-quality voices due to difficulties in preparing recording scripts that account for individual speakers' pronunciation and intonation preferences, leading to increased recording costs and reduced voice quality.
Innovation Solution
A text presentation apparatus that stores text with associated attribute information, allowing for dynamic replacement of difficult-to-pronounce text with alternative phrases based on attribute matching and importance scoring, reducing the need for retakes and improving voice collection.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Manufacturing precision
If a recording script is prepared in advance with consideration of phoneme and intonation selection, then the quality of synthesized speech can be improved, but it becomes difficult to accommodate individual speakers' pronunciation preferences and increases recording complexity
Solution Approach 1:
The system performs preliminary analysis of the recording script to identify difficult-to-pronounce portions before the actual recording session. By pre-marking problematic text segments based on linguistic analysis, the system prepares advance guidance without requiring complex manual script preparation, thus improving speech quality while reducing preparation complexity
Solution Approach 2:
The system introduces an intermediary linguistic analysis component that acts as a mediator between the recording script and the speaker. This intermediary automatically identifies and marks difficult portions, serving as a bridge that translates complex linguistic requirements into simple visual cues for the speaker, thereby reducing the complexity of script preparation while maintaining high speech quality
2Manufacturing precision
If the speaker encounters difficult-to-pronounce text and the recording is rejected, then voice quality standards are maintained, but recording time and costs increase due to repeated retakes
Solution Approach 1:
The system performs preliminary marking of difficult-to-pronounce portions before the speaker attempts recording. By identifying problematic text in advance and providing visual cues, the speaker can prepare mentally and pronounce difficult portions correctly on the first attempt, thereby maintaining voice quality standards while significantly reducing retakes and recording time
Solution Approach 2:
The system provides real-time visual feedback by highlighting difficult-to-pronounce portions as the speaker moves through the script. This immediate feedback mechanism allows the speaker to adjust their pronunciation on the fly, maintaining quality standards while minimizing the need for retakes and reducing overall recording time
3Ease of manufacture
If a standardized recording script is used for all speakers, then script preparation is simplified, but it cannot accommodate individual variations in pronunciation and intonation preferences
Solution Approach 1:
The system applies local quality enhancement by selectively marking only the specific portions of the script that are difficult to pronounce, rather than uniformly complicating the entire script. This localized approach maintains the overall simplicity and standardization of the script while providing targeted assistance for individual speaker needs, thus preserving both ease of preparation and adaptability
Solution Approach 2:
The system introduces dynamic adaptability by adjusting the marking and guidance provided in the script based on the specific speaker's characteristics and performance. The script remains fundamentally standardized but dynamically adapts to individual speaker variations through automated linguistic analysis and real-time feedback, thereby maintaining both ease of preparation and speaker-specific adaptability
Data Source
AI summary
According to an embodiment, a text presentation apparatus presenting text for a speaker to read aloud for voice recording includes: a text storing unit for storing first text; a presenting unit for presenting the first text; a determination unit for determining whether or not the first text needs to be replaced, on the basis of a speaker's input for the first text presented; a preliminary text storing unit for storing preliminary text; a select unit configured to select, if it is determined that the first text needs to be replaced, second text to replace the first text from among the preliminary text, the selecting being performed on the basis of attribute information describing an attribute of the first text and on the basis of at least one of attribute information describing pronunciation of the first text and attribute information describing a stress type of the first text; and a control unit configured to control the presenting unit so that the presenting unit presents the second text.


