Stream matching singing beautifying method based on feature decoupling and masking reconstruction
The stream matching vocal enhancement method based on feature decoupling and masking reconstruction achieves independent control and self-supervised training of timbre, pitch, and content. It solves the problems of comprehensive expressiveness, data dependence, and style loss in existing vocal enhancement methods, improves the quality and efficiency of vocal enhancement, and is suitable for intelligent and personalized applications in the music industry.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- HARBIN INST OF TECH
- Filing Date
- 2026-01-29
- Publication Date
- 2026-05-05
AI Technical Summary
Existing vocal enhancement methods have shortcomings in terms of lack of comprehensive expressiveness, reliance on paired data for model training, and easy loss of personalized singing style, resulting in unnatural enhancement effects. They are difficult to achieve effective separation of timbre, pitch and content, and the reliance on parallel data training limits the model's generalization ability and the accuracy of pitch and rhythm correction.
We employ a stream matching vocal enhancement method based on feature decoupling and masking reconstruction. By introducing a CAM++ timbre encoder, an RMVPE pitch extractor, and a Conformer content encoder to decouple multi-dimensional acoustic features, and combining masking reconstruction and stream matching generation framework, we achieve independent control and self-supervised training of timbre, pitch, and content, thus preserving the singer's personalized style.
Without relying on parallel data, it achieves high-quality, high-fidelity vocal enhancement, significantly improving timbre consistency, pitch accuracy, rhythm stability, and artistic expression. This lowers the application threshold and provides a practical technical path for intelligent and personalized musical expression in the music industry.
Smart Images

Figure CN121983071A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to a stream matching vocal enhancement method based on feature decoupling and masking reconstruction. Background Technology
[0002] Singing, as a vital vehicle for human emotional expression and artistic creation, carries rich cultural connotations. However, for non-professional singers, the lack of systematic vocal training often results in significant deficiencies in breath control and vocal stability, leading to difficulties in accurately controlling pitch and a discrepancy between the actual sound quality and the desired effect. Therefore, despite the large number of singing enthusiasts, many still face the frustration of unsatisfactory performances. Against this backdrop, vocal enhancement technology has demonstrated continuously growing market potential in music production and multimedia entertainment. Whether it's vocal fine-tuning in professional music production, dubbing and singing in film and television dramas, real-time audio editing in online karaoke applications, or demo generation assisted by intelligent composition, vocal enhancement technology plays a crucial role, becoming a key bridge connecting the initial creative intention with the final artistic presentation. Traditional vocal enhancement processes heavily rely on professional sound engineers, requiring expensive hardware and software for tedious manual editing, pitch correction, and effects tuning—time-consuming, labor-intensive, and technically demanding. Therefore, how to achieve automated, intelligent vocal enhancement without sacrificing artistic expression has become an important issue of common concern for both academia and industry.
[0003] The core goal of vocal enhancement is to improve singing quality on two levels without altering the semantic content of the original vocals or the singer's timbre: first, correcting basic singing indicators, including the accuracy of pitch and rhythm; and second, enhancing the artistic expressiveness of the voice, such as optimizing sound quality, improving breath control, and enhancing singing techniques. Current vocal enhancement methods primarily focus on the sub-task of automatic pitch correction, aiming to correct unsatisfactory pitches in amateur singers' performances, making them more consistent with expected pitch curves, such as target MIDI melodies or reference pitches for professional singing. However, these methods typically only address the single dimension of pitch accuracy, without further considering and optimizing other key artistic attributes of the vocals, such as the fullness and purity of timbre, singing techniques like vibrato, breath control, rhythmic stability, emotional coherence, and expressiveness. A vocal performance that is merely pitch-accurate but lacks expressiveness often sounds mechanical and rigid, failing to meet the demands of professional artistic creation or a high-quality auditory experience.
[0004] With the rapid advancement of generative AI technology, vocal enhancement research is shifting from single-dimensional correction to a multi-dimensional, generative, comprehensive enhancement paradigm. Existing methods primarily rely on variational autoencoders or diffusion models for both pitch correction and sound quality enhancement. However, existing generative enhancement methods still have several limitations: First, they depend on paired data, requiring strictly supervised training with amateur-professional parallel data pairs. Collecting such data is extremely costly, requiring professional singers to record both amateur and professional versions of the same song, which is often difficult to obtain in large quantities in the real world, thus limiting the model's generalization ability. Second, there is the issue of lost singing style. While pursuing professionalism or idealization, existing methods tend to over-modify or erase the singer's original individual characteristics, resulting in the enhanced voice losing its distinctiveness and emotional authenticity. Effectively preserving the original singer's personal singing characteristics and style during the enhancement process remains a significant challenge. Summary of the Invention
[0005] The purpose of this invention is to address the shortcomings of existing vocal enhancement methods, such as a lack of comprehensive expressiveness, reliance on paired data for model training, and the easy loss of personalized singing styles leading to unnatural enhancement effects. Therefore, this invention proposes a stream matching vocal enhancement method based on feature decoupling and masking reconstruction.
[0006] The specific process of a stream matching vocal enhancement method based on feature decoupling and masking reconstruction is as follows:
[0007] Step S1: Extract the Mel spectrogram of the singing audio and perform preprocessing;
[0008] Step S2: Decouple and extract multi-dimensional acoustic features from the original audio waveform and Mel spectrogram, including timbre, pitch, and content features;
[0009] Step S3: Fuse and encode the extracted conditional features;
[0010] Step S4: Randomly mask the original Mel spectrogram to obtain the masking features;
[0011] Step S5: Construct a flow matching generative model by inputting conditional features, masking features, and the original Mel spectrogram into the model for training to obtain a trained flow matching model.
[0012] Step S6: Construct hybrid inference conditions, beautify amateur singing based on the trained stream matching model, and input the beautified Mel spectrogram into the vocoder to obtain the beautified singing;
[0013] The beneficial effects of this invention are as follows:
[0014] To address the shortcomings of existing vocal enhancement methods, such as insufficient overall expressiveness, reliance on paired data for model training, and the tendency to lose personalized singing styles leading to unnatural enhancement effects, this invention aims to overcome existing technical bottlenecks and construct an accurate, comprehensive, and highly practical vocal enhancement method.
[0015] First, the primary technical problem this invention addresses is how to efficiently decouple information from the three dimensions of timbre, pitch, and content. There is often a deep coupling between timbre, pitch, and content information in singing, and this complex interrelationship makes effective feature separation difficult using traditional methods. Therefore, the primary technical problem this invention aims to solve is how to establish an effective multi-dimensional feature decoupling mechanism, and through the design of an independent feature extraction module, achieve accurate separation and representation learning of the three dimensions of timbre, pitch, and content, thus providing a foundation for subsequent vocal enhancement generation.
[0016] Secondly, the key technical problem this invention addresses is how to correct pitch and rhythm and achieve self-supervised training, freeing the model from dependence on parallel data. Existing methods for automatic pitch and rhythm correction often suffer from insufficient accuracy, poor naturalness, or weak adaptability to complex musical structures. Furthermore, mainstream generative vocal enhancement models rely on large-scale amateur-professional parallel datasets for supervised training. Acquiring such data is difficult, costly, and has limited coverage, severely restricting the model's scalability and generalization ability in practical applications. Therefore, designing a self-supervised training mechanism that does not rely on parallel data, enabling the model to learn acoustic feature representations and generation rules solely from unpaired vocal data, and achieving more practical pitch and rhythm correction capabilities, is a crucial technical problem this invention aims to solve.
[0017] Furthermore, this invention also needs to address the technical problem of how to maintain the original singer's stylistic characteristics during the enhancement process. Most current methods, while improving pitch accuracy and rhythm regularity, often excessively affect the singer's original timbre, enunciation habits, emotional expression, and other personalized stylistic features, resulting in the enhancement result losing the singer's distinctiveness and artistic appeal. Therefore, how to effectively preserve the singing style characteristics during professional pitch and rhythm correction, so that the enhanced singing both meets professional singing technical standards and authentically restores the singer's personal characteristics, is the key technical problem this invention aims to solve.
[0018] In summary, this invention aims to overcome the bottlenecks of existing vocal enhancement technologies by proposing a stream-matching vocal enhancement method based on feature decoupling and masking reconstruction. This method effectively separates timbre, pitch, and content information by constructing a multi-dimensional acoustic feature decoupling mechanism. Combined with a generative framework of masking reconstruction and stream matching, it achieves high-quality, high-fidelity vocal enhancement without requiring parallel data. This invention significantly lowers the application threshold for high-quality vocal enhancement technologies, providing a practical technical path for cutting-edge application scenarios such as intelligent music production and personalized music expression, and has significant industrial application value.
[0019] This invention proposes a stream-matching vocal enhancement method based on feature decoupling and masking reconstruction, achieving a significant technological breakthrough in the field of vocal enhancement. Compared with existing methods based on variational autoencoders and diffusion models, this invention achieves significant breakthroughs in pitch accuracy, rhythm regularity, timbre consistency, and artistic expressiveness, while greatly reducing the requirements for training data. It provides a complete, efficient, and high-fidelity technical solution for intelligent vocal enhancement.
[0020] Specifically, this invention achieves precise decoupling of the three core attributes of timbre, pitch, and content in the Mel spectrum by introducing three dedicated acoustic coding modules—CAM++ timbre encoder, RMVPE pitch extractor, and Conformer content encoder. The extracted features respectively characterize the singer's identity traits, melody outline, and temporal semantics, laying a solid foundation for subsequent enhancement generation. Furthermore, this invention combines the training objective of masking reconstruction with a stream matching generation framework. During the training phase, by randomly masking continuous temporal segments, the model is forced to learn the ability to reconstruct missing regions based on visible context and multi-dimensional conditional signals, thereby significantly enhancing its ability to model long-term dependencies and local details. Crucially, one of the core challenges of vocal enhancement is how to retain the personal characteristics of amateur singers (such as unique timbre and singing habits) while introducing professional standards (such as accurate pitch and regular rhythm). This masking strategy naturally preserves the stylistic cues of the original audio, avoiding the loss of individuality caused by global reconstruction. During inference, this invention employs a hybrid conditional construction and temporal concatenation strategy, using the first half of an amateur singer's voice as the real context to generate the beautified second half. Significant improvements are achieved in pitch accuracy, timbre naturalness, and consistency. Compared to traditional diffusion models, the flow matching model significantly reduces the sampling steps, providing a feasible guarantee for real-time vocal enhancement applications while maintaining or even surpassing the audio quality of existing methods. Furthermore, the entire training process does not rely on parallel datasets with amateur-professional pairings; model learning can be completed with only a large number of unpaired ordinary vocal samples, greatly reducing the dependence on high-quality paired data and significantly improving the scalability and practical deployment feasibility of the method.
[0021] In summary, this invention proposes a stream-matching vocal enhancement method based on feature decoupling and masking reconstruction. A complete technical system of "feature decoupling—conditional fusion—efficient generation—high-fidelity synthesis" is constructed, providing an efficient, accurate, and natural vocal enhancement solution. First, a multi-dimensional feature decoupling mechanism is designed to achieve independent and fine-grained control over the multi-dimensional attributes of the vocal performance. Simultaneously, a masking reconstruction strategy and a stream-matching generation framework are designed, enabling the model to achieve high-quality reconstruction from noise to a complete Mel spectrum even when some Mel spectrum information is missing, by combining decoupling features as control conditions. This significantly improves efficiency while ensuring high-quality generation. This design gives the model a deep understanding of the spectral context, effectively enhancing the naturalness and stylistic consistency of the generated audio. Furthermore, the training process employs a self-supervised strategy, completely eliminating dependence on parallel training data. Ultimately, this method can generate enhanced vocals that achieve professional-level performance in multiple dimensions such as pitch, rhythm, and expressiveness while maintaining the original singer's timbre and singing style. This invention enhances the practicality and generalizability of the method, providing an accurate and comprehensive end-to-end solution for intelligent vocal enhancement. It significantly lowers the application threshold for high-quality vocal enhancement technology, offering a feasible technical path for cutting-edge application scenarios such as intelligent music production and personalized music expression. It also provides new ideas and strong technical support for achieving high-fidelity and highly controllable processing in a wider range of music generation and editing tasks, demonstrating significant industrial application value. Attached Figure Description
[0022] Figure 1 This is a structural diagram of the model of the present invention;
[0023] Figure 2 This is a flowchart of the present invention. Detailed Implementation
[0024] Specific implementation method one: Combining Figure 1 , Figure 2 This embodiment describes a stream matching vocal enhancement method based on feature decoupling and masking reconstruction, which specifically includes the following steps:
[0025] This invention proposes a stream-matching vocal enhancement method based on feature decoupling and masking reconstruction, aiming to overcome the problems of strong acoustic feature coupling that is difficult to control independently, low generation efficiency, and low timbre similarity of the enhancement results in existing technologies. The following are the key points and areas to be protected in this invention:
[0026] 1. Decoupling and Fusion Mechanism of Multi-Dimensional Acoustic Features. This invention achieves precise decoupling of the three core acoustic attributes: timbre, pitch, and content. Specifically, by introducing three dedicated modules for extracting pure timbre embeddings related to the singer's identity, high-precision continuous and robust fundamental frequency sequences, and phoneme posterior maps of language content, timbre, pitch, and content features are separated from the original audio waveform and Mel spectrum, respectively. This decoupling mechanism allows each dimension of features to be independently manipulated and recombined, laying the foundation for subsequent conditional guided generation and effectively avoiding the timbre distortion or rhythmic misalignment problems caused by feature mixing in traditional methods.
[0027] 2. A Generative Framework Combining Masking Reconstruction and Flow Matching: This invention innovatively combines the masking retraining objective with a flow matching generation paradigm. During the training phase, a dynamic masking strategy is employed to randomly mask local regions of the input Mel spectrum, forcing the model to reconstruct missing parts based on visible context and multi-dimensional decoupling conditions. This significantly enhances the model's ability to model long-term dependencies and local details. In the generation phase, a flow matching framework replaces the traditional diffusion model, drastically reducing the sampling steps from hundreds to thousands to 20–50 steps. While maintaining or even surpassing the audio quality of existing methods, this greatly improves inference efficiency, providing a feasible guarantee for real-time vocal enhancement applications.
[0028] 3. A Hybrid Conditional Inference Strategy for Amateur-to-Enhanced Singing: This invention designs a novel hybrid conditional construction and temporal splicing inference strategy. During the inference process, unprocessed amateur singing is used as the real context input, while a hybrid guiding condition containing three elements is constructed: first, a timbre embedding from the original singer to ensure timbre consistency; second, a target pitch sequence obtained from professional reference audio to improve accuracy; and third, temporally aligned content features to ensure semantic consistency. The resulting singing strictly adheres to the pitch and rhythm specifications of professional singing while maintaining the original singer's timbre characteristics.
[0029] In summary, this invention significantly improves the overall performance of the beautified singing voice in terms of timbre consistency, pitch accuracy, rhythm stability, and artistic expression by constructing a decoupled representation of multi-dimensional acoustic features, fusing masking reconstruction and flow matching into an efficient generation architecture, and introducing an inference strategy guided by real context and mixed conditions. At the same time, it greatly reduces the complexity of inference and has good potential for real-time applications.
[0030] Step S1: Extract the Mel spectrogram of the singing audio and perform preprocessing;
[0031] Step S2: Decouple and extract multi-dimensional acoustic features from the original audio waveform and Mel spectrogram, including timbre, pitch, and content features;
[0032] Step S3: Fuse and encode the extracted conditional features;
[0033] Step S4: Randomly mask the original Mel spectrogram to obtain the masking features;
[0034] Step S5: Construct a flow matching generative model by inputting conditional features, masking features, and the original Mel spectrogram into the model for training to obtain a trained flow matching model.
[0035] Step S6: Construct hybrid inference conditions, beautify amateur singing based on the trained stream matching model, and input the beautified Mel spectrogram into the vocoder to obtain the beautified singing;
[0036] Specific Implementation Method Two: This implementation method differs from Specific Implementation Method One in that: in step S1, the Mel-spectrum of the singing audio is extracted and preprocessed; the specific process is as follows:
[0037] S11, Mel spectrum extraction:
[0038] In the acoustic feature extraction stage, this invention employs a Mel spectrum extraction and preprocessing workflow adapted to the HiFiGAN vocoder.
[0039] (1) Parameter configuration and audio loading: Configure the parameters required for Mel spectrum extraction according to the predefined acoustic hyperparameters, including sampling rate, number of FFT points, number of Mel bands, frame shift, window length, and minimum and maximum frequencies of the Mel filter bank; then, load the audio file from the specified path and resample it to the target sampling rate to obtain mono audio waveform data.
[0040] (2) Audio waveform normalization: Perform dynamic range analysis on the resampled audio waveform and perform amplitude normalization to strictly constrain the waveform amplitude to a preset reasonable range of values to prevent numerical overflow or calculation instability caused by excessive amplitude during subsequent spectrum calculation.
[0041] (3) Mel filter bank and window function caching: Based on the above acoustic hyperparameter configuration, the corresponding Mel filter bank matrix and window function are generated or obtained from the cache. A caching mechanism is adopted so that the same parameter configuration and computing device are generated only once to improve the efficiency of repeated calculations.
[0042] (4) Short-time Fourier Transform (STFT): The normalized audio waveform is framed, windowed and zero-filled using the buffer window function. Then, the short-time Fourier transform is performed to obtain the frequency domain representation in complex form, and its linear amplitude spectrum is taken as the STFT amplitude spectrum.
[0043] (5) Mel spectrum mapping and normalization: The STFT amplitude spectrum and the Mel filter bank matrix are multiplied to achieve the mapping from the linear frequency domain to the Mel frequency domain, and the amplitude spectrum at the Mel scale is obtained. Then, a logarithmic transformation is applied to the Mel amplitude spectrum to compress the dynamic range and enhance the perceptual importance of the low-energy frequency band, and finally the logarithmic Mel spectrum feature is output.
[0044] S12, Feature Normalization:
[0045] To further improve the convergence speed and numerical stability of model training, this invention performs global normalization on the Mel spectrum before inputting it into the neural network.
[0046] (1) Statistical training set extrema: Traverse the log-Mel spectrum of all samples in the current training set, and statistically analyze the global minimum and maximum values as reference boundaries for normalization.
[0047] (2) Linear mapping and pruning: Each Mel spectrum feature value is linearly mapped from the original interval to the target interval [-1, 1]. Specifically, it is first mapped to [0,1], and then extended to [-1, 1] through affine transformation. For feature points that exceed the interval due to outliers or other reasons, boundary pruning is performed to ensure that all feature values are strictly within the interval, thereby ensuring the consistency and robustness of the model input.
[0048] Specific Implementation Method Three: This implementation method differs from Specific Implementation Method One or Two in that: in step S2, multi-dimensional acoustic features, including timbre, pitch, and content features, are decoupled and extracted from the original audio waveform and Mel spectrogram; the specific process is as follows:
[0049] S21. Timbre Feature Extraction:
[0050] This invention employs the CAM++ model as a timbre encoder to extract timbre embedding vectors strongly correlated with the singer's identity from the input Mel spectrogram. CAM++, based on an improved convolutional attention mechanism, captures local textures and formant structures in the spectrum through a multi-layered, deeply separable convolutional network. It also introduces a channel attention module to adaptively weight the contributions of different Mel frequency bands, thereby focusing on the most discriminative acoustic cues for timbre discrimination. The extracted timbre embedding is a 192-dimensional fixed-length vector that effectively summarizes the singer's vocal texture, vocal tract resonance characteristics, and individual vocal habits, and is designed to be independent of specific lyrics, melody direction, and performance duration. In subsequent vocal enhancement processes, this feature is injected as a key conditional signal into the generation model, ensuring that the synthesized voice strictly maintains consistency with the original singer in the timbre dimension, avoiding identity confusion or timbre drift.
[0051] S22, Pitch Feature Extraction:
[0052] This invention selects the RMVPE model as the core component for pitch extraction. RMVPE employs a multi-scale convolutional neural network architecture, enabling robust fundamental frequency estimation of noisy or complex vocal signals (such as vibrato, breathy sounds, weak onsets, etc.), outputting a high-precision frame-level fundamental frequency sequence. To accommodate the timing alignment requirements of subsequent processing modules, the original fundamental frequency sequence undergoes two post-processing steps.
[0053] (1) The time resolution is precisely aligned to the frame rate of the target Mel spectrum by linear interpolation or logarithmic interpolation;
[0054] (2) Map the continuous fundamental frequency values to normalized discrete pitch level codes to form pitch feature sequences. .
[0055] This sequence precisely depicts the temporal contours of the melody, including not only the absolute pitch of the notes but also micro-expression details such as glissando and vibrato. It serves as the core basis for performing vocal enhancement operations such as pitch correction, tonality adjustment, and rhythm alignment.
[0056] S23. Content Feature Extraction:
[0057] This invention constructs a content encoder based on the Conformer architecture to extract content representations related to speech semantics. Conformer combines the modeling ability of convolutional neural networks for local context with the advantage of self-attention mechanisms in capturing long-distance dependencies, achieving both efficiency and expressiveness. This encoder is pre-trained on large-scale multi-speaker, bilingual singing corpora, learning the mapping relationship from Mel-spectrum to phoneme-level semantic units. Its output is a frame-level content feature sequence. It encodes linguistic information such as "what is sung" (lyric content) and "how it is pronounced" (articulation, syllable duration, stress distribution), while being independent of variables such as "who is singing" (timbre) and "at what pitch" (melody). During the vocal enhancement process, content features serve as a content fidelity constraint, ensuring that even with significant adjustments to pitch or timbre, the generated vocals maintain clear articulation, natural rhythm, and accurate semantics, effectively preventing lyric distortion or unclear pronunciation caused by modifications to acoustic parameters.
[0058] Other steps and parameters are the same as in specific implementation method one or two.
[0059] Specific Implementation Method Four: This implementation method differs from Specific Implementation Methods One to Three in that: in step S3, the extracted conditional features are fused and encoded for projection; the specific process is as follows:
[0060] To effectively integrate heterogeneous control signals from multiple dimensions in the singing enhancement task, this invention designs a conditional feature fusion and encoding projection module. This module can map multi-source conditional features such as timbre, content, pitch, and effective speech region masks to a unified latent representation space, and generate a structured, time-aligned fusion conditional sequence, providing high-fidelity and highly guided conditional input for subsequent generative models.
[0061] S31. Independent encoding and time alignment of multi-dimensional features:
[0062] This module first encodes and aligns the raw features from different sources in terms of time dimension to ensure that all conditions are strictly synchronized at the frame level. Specifically, it includes the following steps:
[0063] (1) Timbre feature extension: Receive the 192-dimensional global voiceprint embedding vector pre-extracted by the CAM++ encoder. It maps to the target hidden dimension through a learnable linear projection layer. Subsequently, the projection result is replicated and expanded in the time dimension to generate a temporal timbre feature sequence consistent with the target Mel spectrum frame number. .
[0064] (2) Pitch feature encoding: Receive discretized frame-level pitch sequences First, it is mapped to dense vectors using a learnable embedding lookup table; then, context modeling is performed via a multi-layer convolutional stack to extract pitch representations with local smoothness and dynamic continuity. This ensures the smoothness and local relevance of the pitch trajectory.
[0065] (3) Content feature upsampling: Receive frame-level content features output by the pre-trained Conformer content encoder. Since the original temporal resolution is usually lower than the target Mel spectrum, this invention introduces an upsampling network composed of multiple transposed convolutional modules cascaded together to accurately upsample low-frequency content features to the target frame length. The aligned content feature sequence is obtained. .
[0066] (4) Effective speech region mask generation: In order to enhance the model's ability to distinguish between spoken and silent segments, this invention introduces a non-silent mask. As an auxiliary condition, this mask is generated based on pitch features. Frames with a valid (non-zero) fundamental frequency value are marked as "valid," and the rest are marked as "silent." This binary mask is expanded to form a temporal signal of the same dimension as other features, which is used to guide the generative model to suppress artifacts in silent regions and focus on detail reconstruction in sound regions.
[0067] S32. Multimodal feature splicing and fusion:
[0068] After completing the independent encoding and time alignment of features in each dimension, this module performs cross-modal fusion, specifically including:
[0069] (1) Feature splicing and pre-normalization: The aligned timbre features, content features, pitch features and effective region masks are spliced together in the feature channel dimension to form combined conditional features. Subsequently, a pre-layer normalization is applied to the spliced features to eliminate dimensional differences and distribution shifts between different feature sources, thereby improving training stability and accelerating convergence.
[0070] (2) Unified Spatial Projection and Post-Normalization: A linear projection is used to map high-dimensional combined conditional features to a preset hidden layer dimension. This achieves deep fusion of heterogeneous features, enabling information from different modalities to interact with each other and adapting the dimensions to the input requirements of subsequent generative models. The projected features are then normalized in a second layer to further stabilize the feature distribution. The final output fused conditional sequence is directly used as the conditional input of the flow matching model.
[0071] The other steps and parameters are the same as those in one of the specific implementation methods one to three.
[0072] Specific Implementation Method Five: This implementation method differs from Specific Implementation Methods One to Four in that: in step S4, the original Mel-spectrum image is randomly masked to obtain masking features; the specific process is as follows:
[0073] This invention proposes a joint training method based on a dynamic masking strategy and a continuous flow matching framework. This method simulates scenarios with arbitrary missing regions during the training phase, enabling the model to possess strong context awareness and local completion capabilities. Simultaneously, it employs a flow matching generation framework, allowing it to reconstruct high-quality target Mel spectrum from random noise based on input conditional features, significantly improving training stability and inference efficiency.
[0074] To enable the model to complete any missing regions, this invention designs a dynamic masking training strategy.
[0075] (1) Continuous segment masking strategy: For an input with a length of Complete Mel spectrogram A random start time is selected, and a continuous time frame is masked. The length of the masked area is determined proportionally to the total duration, i.e. The proportionality coefficient Samples are taken within a preset interval (e.g., [0.3, 0.5]). This yields the masked Mel spectrum. .
[0076] (2) Masking signal construction: Synchronously generate a binary masking marker vector with the same time dimension as the Mel spectrum. , among which, if the first If the frame is not masked, then If the first If the frame is masked, then This marker clearly indicates the area that the model needs to reconstruct, enhancing its ability to detect missing locations. Subsequently, and The masking features of the splicing structure model .
[0077] The other steps and parameters are the same as those in one of the specific implementation methods one to four.
[0078] Specific Implementation Method Six: This implementation method differs from Specific Implementation Methods One through Five in that: in S5, a flow matching generative model is constructed, and conditional features, masking features, and the original Mel-spectrum are input into the model for training to obtain a trained flow matching model; the specific process is as follows:
[0079] This invention adopts a flow matching framework based on continuous normalization. Compared with the traditional diffusion model, this framework does not require a multi-step denoising process. It can complete the mapping from noise to data with only a single forward propagation, which has the advantages of more stable training, faster sampling speed and more efficient gradient propagation.
[0080] (1) Definition of probabilistic path: The definition is derived from the noise distribution Data distribution Probability path. For a uniform distribution Mid-sampling time step Random noise with the same shape as the target spectrum is sampled from the standard normal distribution, and the intermediate noisy sample is obtained by linear interpolation. :
[0081]
[0082] in It is Gaussian noise. When hour, For real data; when hour, This is pure noise.
[0083] (2) Velocity field prediction: The core of the model is to train a parameterized velocity field. This function describes the instantaneous velocity of a sample evolving along a probability path, guiding the sample from a straight line from the data distribution to the noise distribution. During training, the model needs to mask each sample. and conditional features Given the input, predict the state from the current state. Reaching complete data The velocity field satisfies the ordinary differential equation (ODE):
[0084]
[0085] To achieve this goal, this invention employs a U-Net structure as the velocity field predictor, comprising a downsampling encoder, an intermediate bottleneck layer, and an upsampling decoder. Temporal embedding. It is broadcast as a global scalar condition to all network layers.
[0086] (3) Loss function design: The goal of training is to minimize the mean square error (MSE) between the predicted velocity field and the target velocity field. The loss function is defined as:
[0087]
[0088] The loss is calculated element-wise across the entire Mel spectrogram, forcing the model to accurately infer the evolution direction required to regress from the current noise state to the complete spectrum at any time step, based on known context, masking features, and multidimensional control conditions. Through this training strategy, the model not only learns to generate high-quality Mel spectrograms from pure noise, but also possesses the ability to perform context-aware intelligent completion under partial observation conditions.
[0089] The other steps and parameters are the same as those in one of the specific implementation methods one to five.
[0090] Specific Implementation Method Seven: This implementation method differs from Specific Implementation Methods One through Six in that: in S6, a hybrid inference condition is constructed, the amateur singing voice is beautified based on the trained stream matching model, and the beautified Mel spectrogram is input into the vocoder to obtain the beautified singing voice; the specific process is as follows:
[0091] S61. Construction of Mixed Conditions and Temporal Separation:
[0092] To enhance amateur singing while maintaining timbre consistency and professional melody and rhythm, this invention employs a temporal splicing strategy to construct a composite input sequence containing a reference context and the region to be generated:
[0093] (1) Masking Mel spectrum construction: Construct a spliced Mel spectrum of length. The first half is the Mel spectrum of a real amateur singing voice, serving as a reference for the known acoustic context. The second half is the region of the enhanced singing voice to be generated.
[0094] (2) Masking mark construction: Construct the corresponding masking mark sequence Setting the first part to 1 indicates that the original model should be retained, while setting the second part to 0 indicates that the model needs to be reconstructed in this area. and The masking features are obtained by splicing. .
[0095] (3) Construction of Multidimensional Hybrid Control Conditions: In order to preserve the personal singing characteristics of amateur singers and achieve the accurate pitch and rhythm of professional singers, this invention splices the conditional features of each dimension: For timbre conditions, the same singer's voiceprint is embedded in both the front and back sections to ensure that the timbre identity is strictly consistent before and after beautification; for pitch conditions, the original pitch of the amateur singer is used in the first half and the professional standard pitch is used in the second half to guide the model to output an accurate and stable melody outline in the generation area; for content conditions, the content features in the second half are obtained by mapping the content features of amateur audio to the time axis of professional audio through a time warping algorithm to ensure that the generated lyrics content is consistent with the original input, but the rhythm and rhyme are aligned with the professional singing; the validity mask marks the valid vocal areas of amateur and professional sections respectively, to help the model suppress artifacts in silent sections and focus on details in vocal sections. The above four types of hybrid conditional features are then integrated according to the fusion process described in S32 to generate the final inference control conditions. , which serves as the core guiding signal for the flow matching model.
[0096] S62, Stream Matching Sampling Generation:
[0097] This invention uses the Euler method to solve the ordinary differential equations of flow matching, thereby achieving deterministic generation of the target Mel spectrum from random noise.
[0098] (1) Initialization: Sample initial noise from a standard normal distribution. Discretize the continuous time interval [0,1]. The time intervals are equal, and the number of steps is usually set to 20-50.
[0099] (2) Euler Iterative Update: From arrive Iteration, or evolution from noise to data, involves continuously updating the state using the Euler method.
[0100]
[0101] (3) Extraction: Completed After one iteration, the complete reconstructed Mel spectrum is obtained, and only the latter half is truncated as the final beautified Mel spectrum output.
[0102] S63, High-quality enhanced audio synthesis:
[0103] To obtain audible high-fidelity audio, this invention performs the following vocoder synthesis process:
[0104] (1) Inverse normalization: The beautified Mel spectrum generated in step S52 is inverse normalized, that is, the inverse transformation of the normalization operation in S12 is performed to restore its numerical range to the input domain expected by the HiFiGAN vocoder.
[0105] (2) Vocoder synthesis: The denormalized Mel spectrum is input into the pre-trained HiFiGAN generator to synthesize the final beautified singing audio with natural sound quality, rich details and high fidelity.
[0106] The other steps and parameters are the same as those in one of the specific implementation methods one to six.
[0107] Although there are theoretical alternatives to its core steps, these alternatives all have significant shortcomings in different key dimensions and cannot achieve the invention's overall objectives of high quality, high efficiency, and low data dependence at the same time.
[0108] The key step in this invention is the decoupling of multi-dimensional acoustic features. An alternative to step S2 could be to use a single, non-targeted neural network to extract all features. However, timbre, pitch, and content are acoustic attributes of different natures, making it difficult for general-purpose modules to achieve optimal performance across various tasks. The combination of CAM++, RMVPE, and Conformer used in this invention comprises highly advanced specialized modules in their respective fields. This combination ensures the purity and accuracy of feature decoupling, which is unmatched by general-purpose modules.
[0109] The key step of this invention is a generative framework based on a combination of masking reconstruction and flow matching. Considering alternatives to step S4, one could use only masking reconstruction for self-supervised pre-training, followed by other generative models. One possible alternative is to use traditional adversarial networks or variational autoencoder frameworks. Adversarial networks suffer from training instability and pattern collapse, making it difficult to generate stable, high-quality audio over long periods; variational autoencoders typically face challenges of overly smooth and insufficiently detailed generated results. Both generally have lower audio fidelity and stability than flow matching models in complex vocal generation tasks. Another possible alternative is to use a diffusion model, learning how to progressively denoise Gaussian noise. However, diffusion models require hundreds to thousands of iterative denoising steps, resulting in slow inference speed and low efficiency, making them unsuitable for real-time applications. The flow matching framework used in this invention mathematically provides a more efficient sampling trajectory, achieving equal or even better quality with very few steps, representing a better solution for balancing efficiency and quality. Furthermore, another possible alternative is to use an autoregressive model for sequential generation, but the inference speed is limited by the sequence length and is prone to error accumulation. Flow matching, being non-autoregressive, supports flexible generation from one to multiple steps, offering advantages in efficiency and long-term consistency. Furthermore, jointly training the masked reconstruction target with flow matching allows the model to learn its generation capabilities within the masking context from the outset, resulting in superior integration and performance compared to two-stage concatenation schemes.
[0110] In summary, although there are some alternatives to certain technical solutions of this invention, these alternatives have significant shortcomings in terms of generation quality and controllability, inference efficiency, etc., and are difficult to achieve the technical effects of this invention. The stream matching vocal enhancement method based on feature decoupling and masking reconstruction proposed in this invention, along with its specific module design and training strategy, is designed to more effectively solve the specific problem of vocal enhancement, and is therefore necessary and innovative.
[0111] What are the advantages of this invention compared to the closest existing technology?
[0112] Compared with current vocal enhancement techniques, the flow matching vocal enhancement method based on feature decoupling and masking reconstruction proposed in this invention exhibits significant advantages in several aspects and effectively overcomes the limitations of existing technologies. Specific comparisons are as follows:
[0113] 1. Precise decoupling and independent control of acoustic features: Existing methods, such as directly inputting Mel spectrograms into diffusion models, struggle to independently and accurately control core attributes of vocal qualities such as timbre, pitch, and rhythm. This invention designs a dedicated multi-dimensional feature decoupling and extraction pipeline. By introducing and integrating targeted modules such as CAM++, RMVPE, and Conformer, it extracts clean, decoupled feature representations from the source spectrum. This allows for the independent injection of professional-standard pitches during enhancement, while strictly locking in and preserving the original singer's timbre and vocal semantics.
[0114] 2. Consistency of Singer's Personal Style: Most existing methods focus only on improving a single metric, such as optimizing pitch. Some comprehensive vocal enhancement methods either over-correct, causing the output to lose the singer's original characteristics, or are insufficiently enhanced, resulting in limited improvement in professionalism. This invention introduces masked reconstruction as a training objective and combines it with flow matching for training. The model is forced to learn to reconstruct the complete vocal performance based on large spectral regions that have been randomly and dynamically masked, as well as decoupled control conditions. This pre-training task greatly enhances the model's long-term context awareness and dependency modeling capabilities, making the generated vocal performance more holistic and natural in terms of melodic lines, rhythmic flow, and emotional expression. Furthermore, this invention proposes a clever hybrid conditional inference strategy. During inference, the model uses the first half of the original vocal performance as the real context, while the generation of the second half is guided by a hybrid condition. This strategy ensures at the algorithmic level that the generated vocal performance possesses both professional-level pitch and rhythm while maximally preserving the original singer's unique timbre and singing habits, perfectly balancing the core contradiction of introducing professional standards and retaining personal characteristics.
[0115] 3. Improved Generation Efficiency: To achieve high audio quality, diffusion-based methods typically require hundreds or even thousands of iterative sampling steps, resulting in excessively long inference times per iteration, making it difficult to meet the demands of real-time or interactive applications. This invention innovatively introduces a streaming generation framework into the vocal enhancement task. This framework mathematically possesses superior sampling path planning, achieving audio quality comparable to or even better than diffusion models with hundreds of steps in just 20-50 sampling steps. This represents an order-of-magnitude improvement in inference speed, laying the technical foundation for real-time, efficient vocal enhancement applications.
[0116] 4. Significantly reduced requirements for training data, eliminating the need for parallel data: Existing techniques are typically based on supervised learning methods, heavily reliant on large-scale parallel datasets for mapping learning—that is, strictly aligned data pairs of the same lyrics from the same song, sung by amateur and professional singers respectively. Acquiring such data is extremely costly and difficult, severely limiting the training scale and generalization ability of the model. This invention benefits from a generative framework combining masking reconstruction and flow matching, a self-supervised training paradigm that does not require parallel data during training. The model only needs to be trained on a large amount of non-parallel amateur or professional vocal data to learn the inherent structure and feature distribution of the vocals. This greatly reduces the data acquisition threshold and cost, improving the practicality and scalability of the method.
[0117] This invention may have other embodiments. Without departing from the spirit and essence of this invention, those skilled in the art can make various corresponding changes and modifications according to this invention, but these corresponding changes and modifications should all fall within the protection scope of the appended claims.
Claims
1. A stream matching vocal enhancement method based on feature decoupling and masking reconstruction, characterized in that: The specific process of the method is as follows: Step S1: Extract the Mel spectrogram of the singing audio and perform preprocessing; Step S2: Decouple and extract multi-dimensional acoustic features from the original audio waveform and Mel spectrogram, including timbre, pitch, and content features; Step S3: Fuse and encode the extracted conditional features; Step S4: Randomly mask the original Mel spectrogram to obtain the masking features; Step S5: Construct a flow matching generative model by inputting conditional features, masking features, and the original Mel spectrogram into the model for training to obtain a trained flow matching model. Step S6: Construct hybrid inference conditions, beautify amateur singing based on the trained stream matching model, and input the beautified Mel spectrogram into the vocoder to obtain the beautified singing.
2. The method for stream matching vocal enhancement based on feature decoupling and masking reconstruction according to claim 1, characterized in that: In step S1, the Mel-spectrum of the singing audio is extracted and preprocessed; the specific process is as follows: S11, Mel spectrum extraction: (1) Parameter configuration and audio loading: Configure the parameters required for Mel spectrum extraction according to the predefined acoustic hyperparameters, including sampling rate, number of FFT points, number of Mel frequency bands, frame shift, window length, and minimum and maximum frequencies of the Mel filter bank; then, load the audio file from the specified path and resample it to the target sampling rate to obtain mono audio waveform data; (2) Audio waveform normalization: Perform dynamic range analysis on the resampled audio waveform and perform amplitude normalization to strictly constrain the waveform amplitude to a preset reasonable range of values to prevent numerical overflow or calculation instability due to excessive amplitude in the subsequent spectrum calculation process. (3) Mel filter bank and window function caching: Based on the above acoustic hyperparameter configuration, generate or obtain the corresponding Mel filter bank matrix and window function from the cache. The caching mechanism is adopted so that the same parameter configuration and computing device are generated only once to improve the efficiency of repeated calculations. (4) Short-time Fourier Transform (STFT): The normalized audio waveform is framed, windowed and zero-filled using the buffer window function. Then, the short-time Fourier transform is performed to obtain the frequency domain representation in complex form, and its linear amplitude spectrum is taken as the STFT amplitude spectrum. (5) Mel spectrum mapping and normalization: Perform matrix multiplication on the STFT amplitude spectrum and the Mel filter bank matrix to realize the mapping from the linear frequency domain to the Mel frequency domain and obtain the amplitude spectrum under the Mel scale; S12, Feature Normalization: (1) Statistical training set extrema: Traverse the log-Mel spectrum of all samples in the current training set, and statistically analyze the global minimum and maximum values as reference boundaries for normalization; (2) Linear mapping and pruning: Each Mel spectrum feature value is linearly mapped from the original interval to the target interval [-1,1]. For feature points that exceed the interval due to outliers or other reasons, boundary pruning is performed to ensure that all feature values are strictly within the interval, thereby ensuring the consistency and robustness of the model input.
3. The method for stream matching vocal enhancement based on feature decoupling and masking reconstruction according to claim 2, characterized in that: In S2, multi-dimensional acoustic features are extracted from the original audio waveform and Mel spectrogram, including timbre, pitch, and content features; The specific process is as follows: S21. Timbre Feature Extraction: The CAM++ model is used as a timbre encoder to extract timbre embedding vectors that are strongly correlated with the singer's identity from the input Mel spectrogram. The extracted timbre is embedded as a 192-dimensional fixed-length vector, which effectively summarizes the singer's vocal texture, vocal tract resonance characteristics and individual vocal habits. In its design, it is independent of the specific lyrics, melody direction and singing duration. In the subsequent vocal enhancement process, this feature is injected into the generation model as a key condition signal to ensure that the synthesized vocals strictly maintain consistency with the original singer in the timbre dimension, avoiding identity confusion or timbre drift. S22, Pitch Feature Extraction: The RMVPE model is selected as the core component for pitch extraction. To meet the timing alignment requirements of subsequent processing modules, the original fundamental frequency sequence needs to undergo two post-processing steps: (1) The time resolution is precisely aligned to the frame rate of the target Mel spectrum by linear interpolation or logarithmic interpolation; (2) Map the continuous fundamental frequency values to normalized discrete pitch level codes to form pitch feature sequences. This sequence precisely depicts the temporal contours of the melody, including not only the absolute pitch of the notes but also the micro-expression details such as glissando and vibrato. It is the core basis for performing vocal enhancement operations such as pitch correction, tonality adjustment, and rhythm alignment. S23. Content Feature Extraction: Construct a content encoder based on the Conformer architecture to extract content representations and content feature sequences related to speech and semantics. It encodes linguistic information such as "what to sing" (lyrics content) and "how to pronounce" (articulation, syllable duration, stress distribution), while being independent of variables such as "who is singing" (timbre) and "at what pitch" (melody). In the process of vocal enhancement, content features serve as content fidelity constraints, ensuring that even when pitch or timbre is significantly adjusted, the generated vocals still maintain clear articulation, natural speech rhythm, and accurate linguistic semantics, effectively preventing lyric distortion or unclear pronunciation caused by modifications to acoustic parameters.
4. The method for stream matching vocal enhancement based on feature decoupling and masking reconstruction according to claim 3, characterized in that: In step S3, the extracted conditional features are fused and encoded for projection; the specific process is as follows: S31. Independent encoding and time alignment of multi-dimensional features: First, the raw features from different sources are encoded and aligned in the time dimension to ensure that all conditions are strictly synchronized at the frame level. This includes the following steps: (1) Timbre feature extension: Receive the 192-dimensional global voiceprint embedding vector pre-extracted by the CAM++ encoder. It maps to the target hidden dimension through a learnable linear projection layer. The projection result is replicated and expanded in the time dimension to generate a temporal timbre feature sequence that matches the target Mel spectrum frame number. ; (2) Pitch feature encoding: Receive discretized frame-level pitch sequences First, it is mapped into dense vectors using a learnable embedding lookup table, and then contextual modeling is performed through a multi-layer convolutional stack to extract pitch representations with local smoothness and dynamic continuity. This ensures the smoothness and local correlation of the pitch trajectory; (3) Content feature upsampling: Receive frame-level content features output by the pre-trained Conformer content encoder. Since its original temporal resolution is usually lower than the target Mel spectrum, an upsampling network consisting of multiple transposed convolutional modules is introduced to accurately upsample the low-frequency content features to the target frame length. The aligned content feature sequence is obtained. ; (4) Effective speech region mask generation: Non-silence mask is introduced. As an auxiliary condition, the mask is generated based on pitch features. Frames with a valid (non-zero) fundamental frequency value are marked as "valid", and the rest are marked as "silent". After expansion, the binary mask forms a temporal signal with the same dimension as other features, which is used to guide the generation model to suppress artifacts in silent areas and focus on detail reconstruction in sound areas. S32. Multimodal feature splicing and fusion: (1) Feature splicing and pre-normalization: The aligned timbre features, content features, pitch features and effective region masks are spliced together in the feature channel dimension to form combined conditional features. Subsequently, a pre-layer normalization is applied to the spliced features to eliminate the dimensional differences and distribution shifts between different feature sources, thereby improving training stability and accelerating convergence. (2) Unified spatial projection and post-normalization: Using a linear projection, the high-dimensional combined condition features are mapped to the preset hidden layer dimension, realizing the deep fusion of heterogeneous features, enabling information from different modalities to interact with each other, and adapting to the input requirements of the subsequent generation model in terms of dimension. The projected features are normalized in a second layer to further stabilize the feature distribution. The final output fusion condition sequence is directly used as the condition input of the flow matching model.
5. The method for stream matching vocal enhancement based on feature decoupling and masking reconstruction according to claim 4, characterized in that: In step S4, the original Mel spectrogram is randomly masked to obtain masking features; The specific process is as follows: Design a dynamic masking training strategy: (1) Continuous segment masking strategy: For an input with a length of Complete Mel spectrogram A random start time is selected, and a continuous time frame is masked. The length of the masked area is determined proportionally to the total duration. The proportionality coefficient Sampling is performed within a preset interval (e.g., [0.3, 0.5]) to obtain the masked Mel spectrum. ; (2) Masking signal construction: Synchronously generate a binary masking marker vector with the same time dimension as the Mel spectrum. , among which, if the first If the frame is not masked, then If the first If the frame is masked, then This marker clearly indicates the area that the model needs to reconstruct, enhancing its ability to detect missing locations. Subsequently, and The masking features of the splicing structure model .
6. The method for stream matching vocal enhancement based on feature decoupling and masking reconstruction according to claim 5, characterized in that: In step S5, a flow matching generative model is constructed by inputting conditional features, masking features, and the original Mel-spectrum image into the model for training, thereby obtaining a trained flow matching model. The specific process is as follows: A flow matching framework based on continuous normalization is adopted. Compared with the traditional diffusion model, this framework does not require a multi-step denoising process. It can complete the mapping from noise to data with only a single forward propagation, which has the advantages of more stable training, faster sampling speed and more efficient gradient propagation. (1) Definition of probabilistic path: The definition is derived from the noise distribution Data distribution Probability path, for a uniform distribution Mid-sampling time step Random noise with the same shape as the target spectrum is sampled from the standard normal distribution, and the intermediate noisy sample is obtained by linear interpolation. : in For Gaussian noise, when hour, For real data; when hour, This is pure noise; (2) Velocity field prediction: The core of the model is to train a parameterized velocity field. This function describes the instantaneous velocity of a sample evolving along a probability path, guiding the sample from a straight line from the data distribution to the noise distribution. During training, the model needs to mask each sample. and conditional features Given the input, predict the state from the current state. Reaching complete data The velocity field satisfies the ordinary differential equation (ODE): To achieve this goal, a U-Net architecture is used as the velocity field predictor, comprising a downsampling encoder, an intermediate bottleneck layer, and an upsampling decoder, with temporal embedding... Broadcast as a global scalar condition to all network layers; (3) Loss function design: The goal of training is to minimize the mean square error (MSE) between the predicted velocity field and the target velocity field. The loss function is defined as: The loss is calculated element-wise across the entire Mel spectrogram, forcing the model to accurately infer the evolution direction required to regress from the current noise state to the complete spectrum at any time step, based on known context, masking features, and multidimensional control conditions. Through the above training strategy, the model not only learns to generate high-quality Mel spectrograms from pure noise, but also has the ability to perform context-aware intelligent completion under some observation conditions.
7. The method for stream matching vocal enhancement based on feature decoupling and masking reconstruction according to claim 6, characterized in that: In step S6, a hybrid inference condition is constructed, and the amateur singing voice is beautified based on the trained stream matching model. The beautified Mel spectrogram is then input into the vocoder to obtain the beautified singing voice. The specific process is as follows: S61. Construction of Mixed Conditions and Temporal Separation: (1) Masking Mel spectrum construction: Construct a spliced Mel spectrum of length. The first half is the Mel spectrum of real amateur singing, serving as a known acoustic context reference; the second half is the region of the enhanced singing to be generated. (2) Masking mark construction: Construct the corresponding masking mark sequence Setting the first part to 1 indicates that the original model should be retained, while setting the second part to 0 indicates that the model needs to be reconstructed in this area. and The masking features are obtained by splicing. ; (3) Construction of multi-dimensional hybrid control conditions: In order to preserve the personal singing characteristics of amateur singers and achieve the accurate pitch and rhythm of professional singers, the condition features of each dimension are spliced together: For the timbre condition, the voiceprint of the same singer is embedded in both the front and back sections to ensure that the timbre identity is strictly consistent before and after beautification; for the pitch condition, the original pitch of the amateur singer is used in the first half and the professional standard pitch is used in the second half to guide the model to output an accurate and stable melody outline in the generation area; for the content condition, the content features of the second half are obtained by mapping the content features of amateur audio to the time axis of professional audio through the time warping algorithm to ensure that the generated lyrics content is consistent with the original input, but the rhythm and rhyme are aligned with the professional singing; the validity mask marks the valid vocal areas of amateur and professional sections respectively to help the model suppress artifacts in the silent section and focus on details in the vocal section; The above four types of mixed condition features are then integrated according to the fusion process described in S32 to generate the final inference control conditions. , serving as the core guiding signal of the flow matching model; S62, Stream Matching Sampling Generation: The Euler method is used to solve the ordinary differential equations of flow matching, enabling deterministic generation of the Mel spectrum from random noise to the target: (1) Initialization: Sample initial noise from the standard normal distribution and discretize the continuous time interval [0,1]. The time intervals are usually set to 20-50; (2) Euler Iterative Update: From arrive Iteration, or evolution from noise to data, involves continuously updating the state using the Euler method; (3) Extraction: Completed After one iteration, the complete reconstructed Mel spectrum is obtained, and only the latter half is truncated as the final beautified Mel spectrum output. S63, High-quality enhanced audio synthesis: (1) Inverse normalization: The beautified Mel spectrum generated in step S52 is inverse normalized, that is, the inverse transformation of the normalization operation in S12 is performed to restore its numerical range to the input domain expected by the HiFiGAN vocoder. (2) Vocoder synthesis: The denormalized Mel spectrum is input into the pre-trained HiFiGAN generator to synthesize the final beautified singing audio with natural sound quality, rich details and high fidelity.