Robust song conversion method and system for real song scene
By constructing a robust vocal conversion model, the problems of harmony interference and pitch error in real song scenarios are solved, achieving a stable generation of the target singer's timbre and improving the stability and naturalness of vocal conversion, making it suitable for industrial applications.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- GIANT MOBILE TECH CO LTD
- Filing Date
- 2026-02-06
- Publication Date
- 2026-05-01
AI Technical Summary
Existing vocal conversion methods are easily affected by harmonic interference, pitch extraction errors, and timbre aliasing in real song scenarios, resulting in unstable timbre, pitch shift, and distortion, making it difficult to meet the needs of industrial applications.
By generating enhanced samples to simulate real separation errors, a feature encoding module for frozen parameters, a fundamental frequency-aware timbre modulation module, and a generation module are designed. A phased training strategy is adopted to construct a robust vocal conversion model that can stably generate the target singer's timbre even in the presence of harmonic interference and fundamental frequency perturbation.
In real-world song scenarios, the stability and naturalness of the generated vocal results are significantly improved, effectively suppressing vocal distortion and artifacts, maintaining consistency between melody and lyrics, and demonstrating good engineering practical value.
Smart Images

Figure CN121963758A_ABST
Abstract
Description
A robust vocal conversion method and system for real-world song scenarios Technical Field
[0001] This invention relates to the field of vocal conversion technology, and in particular to a robust vocal conversion method and system for real-world song scenarios. Background Technology
[0002] Most existing vocal conversion methods are based on idealized input assumptions. When faced with real song scenarios, they are easily affected by problems such as harmony interference, pitch extraction errors, and timbre aliasing, resulting in unstable timbre, pitch deviation, and distortion in the conversion results, which are difficult to meet the needs of industrial applications.
[0003] Therefore, it is necessary to provide a robust vocal conversion method and system for real-world song scenarios, which converts the timbre of the source vocals into the timbre of the target singer while keeping the original melody and lyrics unchanged. Summary of the Invention
[0004] The purpose of this invention is to provide a robust vocal conversion method and system for real-world song scenarios, which converts the timbre of the source vocals into the timbre of the target singer while keeping the original melody and lyrics unchanged.
[0005] To address the problems existing in the prior art, this invention provides a robust vocal conversion method for real-world song scenarios, comprising the following steps:
[0006] S1: During the training data construction phase, augmented samples are generated based on the vocal track and multi-track music data using the following methods:
[0007] S11: Linearly mix the vocal track with the harmony track to construct an input vocal that includes harmonic interference, used to simulate real separation error;
[0008] S12: Introduce random perturbations to the fundamental frequency characteristics, including jitter, slip, and abrupt changes, to simulate the fundamental frequency estimation error in real songs;
[0009] S13: Use a retrieval-based pre-trained timbre conversion model to randomly transform the input singing voice, thereby weakening the source timbre information while keeping the lyrics and melody unchanged;
[0010] S2: The singing conversion model consists of a feature encoding module based on frozen parameters, a fundamental frequency sensing tone modulation module, and a generation module.
[0011] S3: A phased training strategy is adopted, including continuous pre-training and supervised fine-tuning, to obtain a robust vocal conversion model for real song scenarios.
[0012] Optionally, in the robust vocal conversion method for real-world song scenarios,
[0013] The feature encoding module includes a content encoder, a timbre encoder, and a fundamental frequency encoder, which are used to extract content features related to lyrics and melody, timbre features of the target singer, and fundamental frequency features that change over time, respectively.
[0014] The tone modulation module adopts a fundamental frequency-aware tone adaptive structure, which integrates global tone embedding with fundamental frequency features to generate a fine-grained tone representation that varies with pitch.
[0015] The generation module is based on a diffusion generation framework. It uses the initial noise features, time-frequency features, content features related to lyrics and melody, timbre features of the target singer, fundamental frequency features that change over time, fine-grained timbre representation that changes with pitch, and timbre and fundamental frequency conditions as joint inputs. It gradually generates the acoustic features of the target singing voice through a multi-layer self-attention and feedforward network.
[0016] Optionally, in the robust vocal conversion method for real-world song scenarios,
[0017] During the continuous pre-training phase, the model is trained using large-scale speech and singing data, enabling the model to learn basic acoustic modeling capabilities and complete the collaborative adaptation of various modules.
[0018] During the supervised fine-tuning phase, the model is trained based on enhanced real song data, enabling the model to stably generate high-quality target vocals even in the presence of harmonic interference and fundamental frequency perturbations.
[0019] The present invention also provides a robust vocal conversion system for real-world song scenarios, which is established by using the robust vocal conversion method described above.
[0020] Compared with the prior art, the present invention has the following advantages:
[0021] (1) When the model generated by the robust vocal conversion method for real song scenarios described in this invention is used for inference, the model can directly use vocal signals that may contain harmonics or residual vocal parts as input, and stably generate vocal results with the target singer's timbre while maintaining consistency with the original melody and lyrics. This invention, through model design and training strategies targeting the characteristics of real song vocals, makes the generation process more robust to the above-mentioned interference, thereby significantly improving the stability and naturalness of the conversion results.
[0022] (2) Regarding the generation effect, the conversion results of the present invention are compared with existing standalone vocal conversion models. Under ideal input conditions, the timbre consistency and melody preservation effect are close to those of the aforementioned models. However, under input conditions with harmonic interference, the present invention can effectively suppress voice distortion and artifacts, resulting in a more stable overall conversion effect. Furthermore, the degree of matching between the generated vocals and the original singing content was comprehensively evaluated for the overall conversion effect. The results show that the present invention can achieve satisfactory application results in real song vocal scenarios and has good engineering practical value. Attached Figure Description
[0023] Figure 1 is a diagram of the overall structure of the model provided in an embodiment of the present invention. Detailed Implementation
[0024] The specific embodiments of the present invention will now be described in more detail with reference to the accompanying drawings. The advantages and features of the present invention will become clearer from the following description. It should be noted that the drawings are all in a very simplified form and use non-precise proportions, and are only used to facilitate and clarify the illustration of the embodiments of the present invention.
[0025] In the following, if the methods described herein include a series of steps, the order of these steps presented herein is not necessarily the only order in which these steps can be performed, and some of the steps described may be omitted and / or some other steps not described herein may be added to the method.
[0026] Most existing vocal conversion methods are based on idealized input assumptions. When faced with real song scenarios, they are easily affected by problems such as harmony interference, pitch extraction errors, and timbre aliasing, resulting in unstable timbre, pitch deviation, and distortion in the conversion results, which are difficult to meet the needs of industrial applications.
[0027] To address the problems existing in the prior art, this invention provides a robust singing voice conversion method for real-world song scenarios. Singing Voice Conversion (SVC) aims to convert the timbre of the source vocals to that of the target singer while preserving the original melody and lyrics. However, in real-world song scenarios, the input audio is usually not an ideal, pure solo, but rather a mixed audio containing harmonies or residual background sounds.
[0028] As shown in Figure 1, the robust vocal conversion method for real-world song scenarios includes the following steps:
[0029] S1: Regarding data construction and generation: In response to the common problems of accompaniment residue, harmony superposition, and fundamental frequency instability in real songs, this invention designs a robust data construction method for singing conversion tasks.
[0030] Specifically, in the training data construction phase, augmented samples are generated based on the lead vocal track and multi-track music data in the following manner:
[0031] S11: Linearly mixes the lead vocal track with the harmonic / backing vocal track to construct an input vocal that includes harmonic interference, in order to simulate real separation errors;
[0032] S12: Introduce random perturbations, including jitter, glide, and jump, into the fundamental frequency (F0) feature to simulate fundamental frequency estimation errors in real songs;
[0033] S13: A retrieval-based voice conversion (RVC) pre-trained model is used to randomly transform the timbre of the input vocals, weakening the source timbre information while preserving the lyrics and melody. Through this data construction method, the model learns to suppress harmonic interference and adapt to fundamental frequency noise during the training phase, thereby improving its robustness in real-world song scenarios.
[0034] S2: Regarding the model structure design: The singing conversion model proposed in this invention consists of a feature encoder module with frozen parameters, a fundamental frequency-aware timbre adaptation module (F0-aware Timbre Adaptor), and a generation module (Diffusion Transformer, DiT). The feature encoder module includes a content encoder, a timbre encoder, and a fundamental frequency encoder (F0 Encoder), used to extract content features related to lyrics and melody, timbre features of the target singer, and fundamental frequency features that change over time, respectively. The timbre adaptation module employs a fundamental frequency-aware timbre adaptation structure (F0-aware Timbre Adaptor), fusing global timbre embedding with fundamental frequency features to generate a fine-grained timbre representation that changes with pitch. The generation module is based on the Flow Matching based DiffusionTransformer (DiT) framework. It uses the initial noise features, time-frequency features (Mel-spectrogram), the above-mentioned features, timbre and fundamental frequency conditions as joint inputs, and gradually generates the acoustic features of the target singing voice through a multi-layer self-attention and feedforward network.
[0035] S3: Regarding model training: This invention adopts a phased training strategy, including Continuous Pre-Training (CPT) and Supervised Fine-Tuning (SFT).
[0036] During the continuous pre-training phase, the model is trained using large-scale speech and singing data, enabling the model to learn basic acoustic modeling capabilities and complete the collaborative adaptation of various modules.
[0037] In the supervised fine-tuning phase, the model is trained based on enhanced real song data, enabling it to stably generate high-quality target vocals even in the presence of harmonic interference and fundamental frequency perturbations. Through the above training strategy, a robust data construction model for vocal conversion tasks is finally obtained.
[0038] The present invention also provides a robust vocal conversion system for real-world song scenarios, which is established by using the robust vocal conversion method described above.
[0039] In summary, compared with the prior art, the present invention has the following advantages:
[0040] (1) When the model generated by the robust vocal conversion method for real song scenarios described in this invention is used for inference, the model can directly use vocal signals that may contain harmonics or residual vocal parts as input, and stably generate vocal results with the target singer's timbre while maintaining consistency with the original melody and lyrics. This invention, through model design and training strategies targeting the characteristics of real song vocals, makes the generation process more robust to the above-mentioned interference, thereby significantly improving the stability and naturalness of the conversion results.
[0041] (2) Regarding the generation effect, the conversion results of the present invention are compared with existing standalone vocal conversion models. Under ideal input conditions, the timbre consistency and melody preservation effect are close to those of the aforementioned models. However, under input conditions with harmonic interference, the present invention can effectively suppress voice distortion and artifacts, resulting in a more stable overall conversion effect. Furthermore, the degree of matching between the generated vocals and the original singing content was comprehensively evaluated for the overall conversion effect. The results show that the present invention can achieve satisfactory application results in real song vocal scenarios and has good engineering practical value.
[0042] The above are merely preferred embodiments of the present invention and do not constitute any limitation on the present invention. Any equivalent substitutions or modifications made by those skilled in the art to the technical solutions and content disclosed in the present invention without departing from the scope of the present invention shall be deemed to have remained within the protection scope of the present invention.
Claims
1. A robust vocal conversion method for real-world song scenarios, characterized in that, Includes the following steps: S1: In the training data construction phase, based on the vocal track and multi-track music data, enhanced samples are generated in the following ways: S11: The vocal track and harmony track are linearly mixed to construct an input vocal track containing harmonic interference to simulate real separation errors; S12: Random perturbations are introduced into the fundamental frequency features, including jitter, slip, and abrupt changes, to simulate fundamental frequency estimation errors in real songs; S13: A retrieval-based pre-trained timbre conversion model is used to randomly transform the input vocal track, weakening the source timbre information while keeping the lyrics and melody unchanged; S2: The vocal conversion model is composed of a feature encoding module based on frozen parameters, a fundamental frequency-aware timbre modulation module, and a generation module; S3: A phased training strategy is adopted, including continuous pre-training and supervised fine-tuning, to obtain a robust vocal conversion model for real song scenarios.
2. The robust vocal conversion method for real-world song scenarios as described in claim 1, characterized in that, The feature encoding module includes a content encoder, a timbre encoder, and a fundamental frequency encoder, which are used to extract content features related to lyrics and melody, timbre features of the target singer, and fundamental frequency features that change over time, respectively. The tone modulation module adopts a fundamental frequency-aware timbre adaptive structure, which integrates global timbre embedding with fundamental frequency features to generate a fine-grained timbre representation that varies with pitch. The generation module is based on a diffusion generation framework, which uses diffusion initial noise features, time-frequency features, content features related to lyrics and melody, target singer timbre features, time-varying fundamental frequency features, fine-grained timbre representation that varies with pitch, and timbre and fundamental frequency conditions as joint inputs. It gradually generates the acoustic features of the target singing voice through a multi-layer self-attention and feedforward network.
3. The robust vocal conversion method for real-world song scenarios as described in claim 1, characterized in that, During the continuous pre-training phase, the model is trained using large-scale speech and singing data, enabling the model to learn basic acoustic modeling capabilities and complete the collaborative adaptation of various modules. During the supervised fine-tuning phase, the model is trained based on enhanced real song data, enabling the model to stably generate high-quality target vocals even in the presence of harmonic interference and fundamental frequency perturbations.
4. A robust vocal conversion system for real-world song scenarios, characterized in that, A robust vocal conversion system for real-world song scenarios is established using the robust vocal conversion method described in any one of claims 1-3.