Singing Voice Synthesis Through Breath-Aware Score Segmentation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional singing voice synthesizing techniques fail to adequately account for breath-taking voices in musical scores, leading to unnatural phenomena such as prolonged voices or breath-holding during synthesis, which negatively impact the auditory perception of the synthesized singing voice.
Innovation Solution
Segment a musical score file into multiple segments based on breath-taking identifiers, generate audio segments corresponding to these segments, and synthesize a singing voice that mimics the natural breathing rhythm of a real singer by adding breath-taking symbols and silence symbols to ensure smooth transitions.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If conventional singing voice synthesizing techniques are used, then the singing voice can be generated, but unnatural phenomena such as prolonged voices or breath-holding occur during synthesis
Solution Approach 1:
The musical score file is segmented into multiple segments based on breath-taking identifiers. Each segment corresponds to a specific breath-taking interval, allowing the synthesis system to process and generate audio segments that naturally incorporate breath pauses, thereby eliminating unnatural breath-holding effects in the final synthesized singing voice.
2Reliability
If breath-taking identifiers are used to segment the musical score file, then natural breathing rhythm is simulated, but the complexity of the synthesis process increases
Solution Approach 1:
Breath-taking identifiers are pre-inserted into the musical score file at appropriate positions before the synthesis process begins. This preliminary action marks the breath-taking intervals in advance, allowing the segmentation and audio generation to proceed systematically without requiring complex real-time analysis of breathing patterns during synthesis.
Data Source
AI summary
Embodiments of the present disclosure relate to a method and apparatus for synthesizing a singing voice, an electronic device and a program product. The method comprises obtaining a musical score file with breath-taking identifiers, and segmenting the musical score file into a plurality of musical score segments based on the breath-taking identifiers. The method further comprises generating a plurality of audio segments corresponding to the plurality of musical score segments, and synthesizing a singing voice corresponding to the musical score file based on the plurality of audio segments. In an embodiment of the present disclosure, the musical score file is segmented into a plurality of musical score segments according to the breath-taking identifier, a plurality of audio segments corresponding to the plurality of musical score segments are generated, and finally, a singing voice corresponding to the musical score file is synthesized based on the plurality of audio segments.


