Script Data Alignment for Performance Voice Synthesis
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current technologies lack the ability to output performance voices in accordance with the intention of a script, as they do not provide data that associates dialogue with the intended utterer and stage directions, requiring manual preparation and analysis.
Innovation Solution
An information processing device that includes a hardware processor to output second script data by associating dialogue data with utterer data, using a communication unit, user interface, storage unit, and processing unit to analyze and generate performance voice data, incorporating optical character recognition and learning models to optimize dialogue and utterer data alignment.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If manual preparation and analysis of script data is performed, then accuracy of performance voice synthesis is improved, but productivity and ease of operation deteriorate
Solution Approach 1:
The system automatically analyzes script data and generates performance voice synthesis parameters without requiring manual preparation. The processing unit automatically extracts dialogue data, utterer data, and stage direction data from the script, and automatically generates synthesis parameters based on this extracted information, enabling the system to serve itself rather than requiring human intervention for data preparation.
Solution Approach 2:
The patent replaces manual mechanical analysis of script data with automated information processing. Instead of manually analyzing and preparing script data, the system uses computational algorithms to automatically extract, associate, and process script information, substituting human cognitive operations with automated digital processing mechanisms.
2Measurement precision
If manual preparation and analysis of script data is performed, then accuracy of performance voice synthesis is improved, but ease of operation deteriorates
Solution Approach 1:
The system automatically analyzes script data and generates performance voice synthesis parameters without requiring manual preparation. The processing unit automatically extracts dialogue data, utterer data, and stage direction data from the script, and automatically generates synthesis parameters based on this extracted information, enabling the system to serve itself rather than requiring human intervention for data preparation.
3Productivity
If automated voice synthesis is used, then productivity is improved, but ability to reflect script intentions deteriorates
Solution Approach 1:
The patent segments the script data into distinct components: dialogue data, utterer data, and stage direction data. By separating these elements, the system can process each component independently and then combine them to generate comprehensive performance voice synthesis parameters that accurately reflect the original script intentions while maintaining automated processing efficiency.
Solution Approach 2:
The system introduces an intermediary processing layer that automatically analyzes and interprets script data. This intermediary processing unit extracts meaningful information from the raw script, associates dialogue with utterers and stage directions, and transforms this information into synthesis parameters, thereby bridging the gap between automated processing and accurate script intention reflection.
Data Source
AI summary
An information processing device (10) includes a hardware processor configured to function as an output unit (24). The output unit (24) outputs second script data in which dialogue data of a dialogue included in first script data is associated with utterer data of an utterer of the dialogue from the first script data as a basis for performance.


