Hierarchical Voice Synthesis Interface for Singing Expression Selection
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Users face difficulty in selecting a desired singing expression from a large number of options in existing voice synthesis techniques, leading to a cumbersome and inefficient editing process.
Innovation Solution
A display control method that utilizes a hierarchical structure to display singing expressions, allowing users to select options in a layer-by-layer manner, with the indicator moving over options to reveal subsequent layers, simplifying the selection process and reducing the complexity of choosing a desired singing expression.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of information
If all types of singing expressions are displayed in parallel form in a list, then the user can see all available options at once, but the user finds it difficult to find and select a desired type of singing expression due to the large number of options
Solution Approach 1:
The list of singing expressions is segmented into multiple pages, with each page displaying a limited number of options (e.g., 5-10 expressions per page). Users can navigate through pages to access all expressions, but only a manageable subset is visible at any given time, reducing cognitive load and improving selection ease.
Solution Approach 2:
The interface transitions from a single-dimension flat list to a multi-dimensional hierarchical structure with categories (e.g., emotional expressions, technical expressions, stylistic expressions). This adds a categorical dimension that organizes expressions logically, making it easier for users to locate desired expressions by browsing relevant categories rather than scanning the entire list.
2Ease of operation
If a hierarchical structure with multiple layers is implemented, then the selection process becomes more intuitive and visual clutter is reduced, but the interface complexity increases
Solution Approach 1:
The system pre-organizes singing expressions into hierarchical categories and subcategories before presentation to the user. This preliminary structuring allows the interface to display only relevant subsets of expressions based on the current selection level, reducing visual clutter while maintaining intuitive navigation without requiring complex real-time processing.
Solution Approach 2:
The interface implements a nested hierarchical structure where main categories contain subcategories, which in turn contain specific singing expressions. Each level is revealed progressively as users navigate, with parent categories expanding to show child categories. This nesting approach organizes complex information hierarchically while maintaining a clean, simple interface at each viewing level.
Data Source
AI summary
A display control method executed by a processor, the method includes the steps of: displaying, on a display device, a note icon that represents a note of a voice to be synthesized and an indicator that is moved in accordance with an operation received from a user; displaying, on the display device, first options that belong to a first layer among layers in a hierarchical structure, for the user to select a singing expression to be applied to the note from among a plurality of singing expressions; and displaying, on the display device, when the indicator is moved into an area corresponding to a particular option selected from among the first options, second options that correspond to the particular option and belong to a second layer that is below the first layer in the hierarchical structure.


