Singing Voice Synthesis Through Breath-Aware Score Segmentation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional singing voice synthesizing techniques fail to adequately account for breath-taking voices in musical scores, leading to unnatural phenomena such as prolonged voices or breath-holding during synthesis, which negatively impact the auditory perception of the synthesized singing voice.

Innovation Solution

Segment a musical score file into multiple segments based on breath-taking identifiers, generate audio segments corresponding to these segments, and synthesize a singing voice that mimics the natural breathing rhythm of a real singer by adding breath-taking symbols and silence symbols to ensure smooth transitions.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If conventional singing voice synthesizing techniques are used, then the singing voice can be generated, but unnatural phenomena such as prolonged voices or breath-holding occur during synthesis

Engineering Contradiction:
Improvenaturalness of synthesized singing voiceVSAvoidunnatural breath-holding and tone dragging
Core Design Contradiction:
ReliabilityVSObject-generated harmful factors

Solution Approach 1:

The musical score file is segmented into multiple segments based on breath-taking identifiers. Each segment corresponds to a specific breath-taking interval, allowing the synthesis system to process and generate audio segments that naturally incorporate breath pauses, thereby eliminating unnatural breath-holding effects in the final synthesized singing voice.

Inventive Principle:
Principle #1Segmentation

2Reliability

If breath-taking identifiers are used to segment the musical score file, then natural breathing rhythm is simulated, but the complexity of the synthesis process increases

Engineering Contradiction:
Improvesimulation of natural breathing rhythmVSAvoidcomplexity of synthesis process
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

Breath-taking identifiers are pre-inserted into the musical score file at appropriate positions before the synthesis process begins. This preliminary action marks the breath-taking intervals in advance, allowing the segmentation and audio generation to proceed systematically without requiring complex real-time analysis of breathing patterns during synthesis.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS20250239241A1Method and apparatus for synthesizing a singing voice, electronic device and program product
Publication Date: 2025.07.24 BEIJING ZITIAO NETWORK TECH CO LTD
  • US20250239241A1 patent drawing
  • US20250239241A1 patent drawing
  • US20250239241A1 patent drawing

AI summary

Embodiments of the present disclosure relate to a method and apparatus for synthesizing a singing voice, an electronic device and a program product. The method comprises obtaining a musical score file with breath-taking identifiers, and segmenting the musical score file into a plurality of musical score segments based on the breath-taking identifiers. The method further comprises generating a plurality of audio segments corresponding to the plurality of musical score segments, and synthesizing a singing voice corresponding to the musical score file based on the plurality of audio segments. In an embodiment of the present disclosure, the musical score file is segmented into a plurality of musical score segments according to the breath-taking identifier, a plurality of audio segments corresponding to the plurality of musical score segments are generated, and finally, a singing voice corresponding to the musical score file is synthesized based on the plurality of audio segments.