Music accompaniment separation method and system based on sound structure perception

By using a music accompaniment separation method based on voice structure perception, and optimizing the separation process with multi-track music data and a voice separation backbone network, the problem of voice mixing in existing technologies is solved, and clear separation of the lead vocal part and overall coordination are achieved.

CN121983079APending Publication Date: 2026-05-05GIANT MOBILE TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
GIANT MOBILE TECH CO LTD
Filing Date
2026-02-06
Publication Date
2026-05-05

AI Technical Summary

Technical Problem

Existing music accompaniment separation technology ignores the differences in vocal structure between lead vocals and harmony, resulting in voice overlap in the separation results. This affects the integrity and semantic consistency of the lead vocal sequence in subsequent applications and lacks a constraint mechanism for the consistency of voice structure.

Method used

The model is trained using multi-track music data. It uses a frequency band independent encoder and a voice part separation backbone network to perceive the voice part structure. The separation results are optimized by combining global speaker coding and auxiliary structure modules. Multi-cycle, multi-scale and multi-resolution discriminators are used to model and score the voice part prediction branches.

Benefits of technology

It achieves cleaner and more clearly structured vocal separation, improves the usability and stability of the separation results in downstream generative audio tasks, and maintains the consistency of vocal roles and overall coordination throughout the entire song.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121983079A_ABST
    Figure CN121983079A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of music source separation, in particular to a music accompaniment separation method and system based on sound structure perception. According to the music accompaniment separation method and system based on sound part structure perception provided by the invention, while the music accompaniment is separated, a sound part structure perception mechanism is introduced, and the structural relationship among different sound parts in the music is modeled, so that a clearer and clearer-structure main singing sound part is obtained, and the music accompaniment separation efficiency is improved. And the availability and the stability of the separation result in the downstream generation type audio task are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of music source separation technology, and in particular to a music accompaniment separation method and system based on acoustic structure perception. Background Technology

[0002] Most current mainstream methods for separating vocals and accompaniment are based on deep learning models, achieving separation by performing end-to-end modeling of the mixed audio. While these methods have achieved some success in signal reconstruction metrics, they still have the following shortcomings:

[0003] Current music accompaniment separation technologies typically divide the music signal simply into two sound sources: "voice" and "accompaniment," ignoring the structural differences between the lead vocals and harmonies within the vocal range. In actual music production and application scenarios, the lead vocals usually carry the main semantic and melodic information, while the harmonies are used to enhance the overall listening experience and musical depth. Existing technologies often mix the lead vocals and harmonies into a single vocal track during the separation process, resulting in significant aliasing issues in the separation results (still containing a large amount of harmonic remnants). This affects the requirements for the integrity and consistency of the lead vocal sequence in subsequent applications, such as causing interference in tasks like vocal transitions, pitch extraction, and semantic alignment.

[0004] The separation results are difficult to meet the requirements of downstream tasks: In applications such as vocal conversion, pitch estimation, and lip-syncing, there are high requirements for the temporal consistency and semantic purity of the vocal sequence. However, the vocal results generated by existing separation methods are prone to introducing interference, which affects the overall generation effect.

[0005] There is a lack of constraints on the consistency of voice structure. Most existing technologies rely solely on reconstruction errors at the waveform or spectral level for optimization, without constraining the relationships between different voices from a musical structure perspective, which easily leads to confusion of voice roles.

[0006] Therefore, it is necessary to provide a music accompaniment separation method and system based on voice structure perception, which can combine voice structure information to constrain and optimize the music accompaniment separation process. Summary of the Invention

[0007] The purpose of this invention is to provide a music accompaniment separation method and system based on voice structure perception, which can constrain and optimize the music accompaniment separation process by combining voice structure information.

[0008] To address the problems existing in the prior art, this invention provides a music accompaniment separation method based on acoustic structure perception, comprising the following steps:

[0009] S1: To construct the angle of data generation, multi-track music data is used as the source of model training data. The multi-track music includes the lead vocal track, the backing vocal track, and the music accompaniment track.

[0010] S2: From the perspective of model structure design:

[0011] S21: The music accompaniment separation model takes the mixed music audio signal as input. The input audio first undergoes time-frequency transformation to obtain time-frequency feature representation, and then is sent to the frequency band independent encoder. The frequency band independent encoder divides and encodes the time-frequency features in the frequency band dimension to obtain the feature representation corresponding to each frequency band.

[0012] S22: The feature representation encoded by the frequency band independent encoder is input into the acoustic discrete backbone network for modeling.

[0013] S23: The features output by the vocal part separation backbone network are respectively connected to multiple vocal part prediction branches. Each vocal part prediction branch is independent of each other and is used to generate the separation results corresponding to the lead vocal part, backing vocal part and musical accompaniment vocal part respectively.

[0014] S3: Angle for model training:

[0015] S31: During the model training phase, the constructed multitrack music training data is input into the model for training;

[0016] S32: During training, the mixed music audio is encoded by time-frequency transformation and frequency band independent encoder and then input into the vocal part separate backbone network for forward calculation; after the model forward calculation, the prediction results of the lead vocal part, backing vocal part and music accompaniment part are output respectively, and the corresponding multi-track vocal part audio track is used as the supervision signal to calculate the vocal part separate reconstruction loss of the prediction results for parameter update.

[0017] S34: Separate music accompaniment based on the trained model.

[0018] Optionally, in the music accompaniment separation method based on voice structure perception, the angle for generating the data is constructed in the following way:

[0019] S11: During the data construction phase, only multi-track music with a single lead singer is selected for training, and samples with multiple lead singers singing in rounds, duets, or with unstable lead singer identities are removed.

[0020] S12: When extracting speaker codes, since the lead vocal part exists as an independent audio track in the multi-track data, the lead vocal track is directly used as the speaker feature extraction object, and the context-aware mask speaker coding model is used to encode and extract the lead vocal track.

[0021] Optionally, in the music accompaniment separation method based on voice structure perception, the voice separation backbone network structurally integrates global speaker coding as conditional information to participate in feature calculation.

[0022] Optionally, the music accompaniment separation method based on voice structure perception further includes the following steps:

[0023] S24: Establish an auxiliary structure module that is connected only during the training phase to perform structural modeling of the output results of the voice prediction branch. The auxiliary structure module consists of multiple discriminant sub-networks, specifically including a multi-period discriminator, a multi-scale discriminator, and a multi-resolution spectral discriminator, which respectively model and score the output results of the voice prediction branch from different periodic characteristics, different time scales, and different spectral resolutions.

[0024] Optionally, in the music accompaniment separation method based on vocal structure perception, during the training process in S32, the speaker code extracted from the lead vocal track is injected as global conditional information into the vocal separation backbone network. Specifically, it is applied to the intermediate layer features of the vocal separation backbone network through conditional modulation, and the global conditional information is discarded with a preset probability.

[0025] Optionally, in the music accompaniment separation method based on voice structure perception, during the training process of S32, the auxiliary structure module constructs the voice prediction branch output results obtained by the model prediction and the corresponding real voice track into multiple voice combination forms according to preset rules, and models the voice combination audio composed of the voice prediction branch output results and the corresponding reference voice combination audio input to the discrimination subnetwork.

[0026] Optionally, in the music accompaniment separation method based on voice structure perception, the voice combination forms include: combining the lead vocal part in the voice prediction branch output with the actual music accompaniment part, and using the combination of the actual lead vocal part and the actual music accompaniment part as the corresponding reference combination; and combining the lead vocal part in the voice prediction branch output with the actual backing vocal part, or combining the actual lead vocal part with the backing vocal part in the voice prediction branch output, and using the combination of the actual lead vocal part and the actual backing vocal part as the corresponding reference combination.

[0027] The discriminant subnetwork outputs a relative score based on the results of the aforementioned voice combination form and updates its parameters using a ranking-based loss function. Simultaneously, the voice-separated backbone network updates its parameters based on the output of the discriminant subnetwork with an optimization objective different from that of the discriminant subnetwork, thereby achieving joint training of the voice-separated backbone network and the auxiliary structure module.

[0028] The present invention also provides a music accompaniment separation system based on acoustic structure perception, which establishes a music accompaniment separation system using the aforementioned music accompaniment separation method to achieve music accompaniment separation.

[0029] Compared with the prior art, the present invention has the following advantages:

[0030] (1) While separating the musical accompaniment, the present invention introduces a vocal structure perception mechanism to model the structural relationship between different vocal parts in the music, thereby obtaining a cleaner and clearer vocal part and improving the availability and stability of the separation results in downstream generative audio tasks.

[0031] (2) When using the model described in this invention for reasoning, the model takes the mixed audio of the whole song as input, performs multi-source separation of the lead vocal part, backing vocal part and background music part under the global perception, and uses the lead vocal speaker code to conditionally correct the local vocal part results during the reasoning process, thereby ensuring the consistency of the role and structural stability of the lead vocal part in the whole song.

[0032] (3) In terms of separation quality, the present invention compares the extracted vocal part results with the output effect of the existing independent vocal separation model, and the separation effect remains at a similar level; at the same time, the separated background music results are compared with the output effect of the existing independent background music separation model, and the effect also remains at a similar level.

[0033] (4) The present invention evaluates the structural consistency and synergy between the lead vocal part and the background music from the perspective of the overall results. The results show that while maintaining the quality of each vocal part, the present invention can better maintain the stability of the vocal part role and the overall coordination within the whole song, thus demonstrating the advantages in terms of comprehensive processing efficiency and system consistency. Attached Figure Description

[0034] Figure 1 A flowchart of a music accompaniment separation method provided in an embodiment of the present invention. Detailed Implementation

[0035] The specific embodiments of the present invention will now be described in more detail with reference to the accompanying drawings. The advantages and features of the present invention will become clearer from the following description. It should be noted that the drawings are all in a very simplified form and use non-precise proportions, and are only used to facilitate and clarify the illustration of the embodiments of the present invention.

[0036] In the following, if the methods described herein include a series of steps, the order of these steps presented herein is not necessarily the only order in which these steps can be performed, and some of the steps described may be omitted and / or some other steps not described herein may be added to the method.

[0037] Most mainstream methods for separating vocals and accompaniment are based on deep learning models, which separate vocals from accompaniment by performing end-to-end modeling of the mixed audio. Although these methods have achieved some success in signal reconstruction metrics, they still have many shortcomings.

[0038] To address the problems existing in the prior art, this invention provides a music accompaniment separation method based on acoustic structure perception, such as... Figure 1 As shown, it includes the following steps:

[0039] S1: To construct the angle of data generation, multi-track music data is used as the source of model training data. The multi-track music includes the lead vocal track, the backing vocal track, and the music accompaniment track.

[0040] The method for constructing angles from the data is as follows:

[0041] S11: During the data construction phase, only multi-track music with a single lead singer is selected for training, and samples with multiple lead singers singing in rounds, duets, or with unstable lead singer identities are removed.

[0042] S12: When extracting speaker codes, since the lead vocal part exists as an independent audio track in the multi-track data, the lead vocal track is directly used as the speaker feature extraction object, and the context-aware masking (CAM++) speaker coding model is used to encode and extract the lead vocal track.

[0043] S2: From the perspective of model structure design:

[0044] S21: The music accompaniment separation model takes the mixed music audio signal as input. The input audio first undergoes time-frequency transformation to obtain time-frequency feature representation, and then is sent to the band-split encoder. The band-split encoder divides and encodes the time-frequency features in the band dimension to obtain the feature representation corresponding to each band.

[0045] S22: The feature representation encoded by the band-independent encoder is input into the modeling based on the Band-Split Roformer, which structurally integrates the global speaker encoding as conditional information to participate in feature calculation.

[0046] S23: The features output by the vocal part separation backbone network are respectively connected to multiple vocal part prediction branches. Each vocal part prediction branch is independent of each other and is used to generate the separation results corresponding to the lead vocal part, backing vocal part and musical accompaniment vocal part respectively.

[0047] S24: Establish an auxiliary structural module that is connected only during the training phase to perform structural modeling of the sound part prediction branch output. This auxiliary structural module consists of multiple discriminant sub-networks, specifically including a multi-period discriminator, a multi-scale discriminator, and a multi-resolution spectrogram discriminator. The multi-period discriminator, multi-scale discriminator, and multi-resolution spectrogram discriminator model and score the sound part prediction branch output from different periodic characteristics, different time scales, and different spectral resolutions, respectively. During the inference phase, this auxiliary structural module does not participate in model computation.

[0048] S3: Angle for model training:

[0049] S31: During the model training phase, the constructed multitrack music training data is input into the model for training;

[0050] S32: During training, the mixed music audio, after time-frequency transformation and encoding with a frequency band-independent encoder, is input into the vocal partition backbone network for forward computation. Simultaneously, the speaker code extracted from the lead vocal track is injected as global conditional information into the vocal partition backbone network, specifically through feature-wise modulation applied to the intermediate layer features of the backbone network. During training, this global conditional information is discarded with a preset probability (Condition Dropout).

[0051] After forward computation, the model outputs the prediction results for the lead vocals, backing vocals, and musical accompaniment, respectively. The corresponding multi-track audio tracks are used as supervision signals to calculate the separation reconstruction loss of the vocal parts based on the prediction results for parameter updates.

[0052] During the training process of S32, the auxiliary structure module constructs various cross-composed mixtures by combining the output results of the vocal part prediction branches obtained from the model prediction with the corresponding real vocal part tracks according to preset rules. The cross-composed audio composed of the vocal part prediction branch outputs is then input into the discriminant subnetwork for modeling. The vocal part combination forms include: combining the lead vocal part in the vocal part prediction branch output with the real musical accompaniment part, and using the combination of the real lead vocal part and the real musical accompaniment part as the corresponding reference combination; and combining the lead vocal part in the vocal part prediction branch output with the real backing vocal part, or combining the real lead vocal part with the backing vocal part in the vocal part prediction branch output, and using the combination of the real lead vocal part and the real backing vocal part as the corresponding reference combination.

[0053] The discriminant subnetwork outputs a relative score based on the results of the combined voice parts, and updates its parameters using a ranking loss function. Meanwhile, the voice-separated backbone network updates its parameters with an optimization objective different from that of the discriminant subnetwork, based on the output of the discriminant subnetwork, thereby achieving joint optimization of the voice-separated backbone network and the auxiliary structure module.

[0054] S34: Separate music accompaniment based on the trained model.

[0055] The present invention also provides a music accompaniment separation system based on acoustic structure perception, which establishes a music accompaniment separation system using the aforementioned music accompaniment separation method to achieve music accompaniment separation.

[0056] In summary, compared with the prior art, the present invention has the following advantages:

[0057] (1) While separating the musical accompaniment, the present invention introduces a vocal structure perception mechanism to model the structural relationship between different vocal parts in the music, thereby obtaining a cleaner and clearer vocal part and improving the availability and stability of the separation results in downstream generative audio tasks.

[0058] (2) When using the model described in this invention for reasoning, the model takes the mixed audio of the whole song as input, performs multi-source separation of the lead vocal part, backing vocal part and background music part under the global perception, and uses the lead vocal speaker code to conditionally correct the local vocal part results during the reasoning process, thereby ensuring the consistency of the role and structural stability of the lead vocal part in the whole song.

[0059] (3) In terms of separation quality, the present invention compares the extracted vocal part results with the output effect of the existing independent vocal separation model, and the separation effect remains at a similar level; at the same time, the separated background music results are compared with the output effect of the existing independent background music separation model, and the effect also remains at a similar level.

[0060] (4) The present invention evaluates the structural consistency and synergy between the lead vocal part and the background music from the perspective of the overall results. The results show that while maintaining the quality of each vocal part, the present invention can better maintain the stability of the vocal part role and the overall coordination within the whole song, thus demonstrating the advantages in terms of comprehensive processing efficiency and system consistency.

[0061] The above are merely preferred embodiments of the present invention and do not constitute any limitation on the present invention. Any equivalent substitutions or modifications made by those skilled in the art to the technical solutions and content disclosed in the present invention without departing from the scope of the present invention shall be deemed to have remained within the protection scope of the present invention.

Claims

1. A method for separating musical accompaniment based on voice structure perception, characterized in that, Includes the following steps: S1: To construct the angle of data generation, multi-track music data is used as the source of model training data. The multi-track music includes the lead vocal track, the backing vocal track, and the music accompaniment track. S2: From the perspective of model structure design: S21: The music accompaniment separation model takes the mixed music audio signal as input. The input audio first undergoes time-frequency transformation to obtain time-frequency feature representation, and then is sent to the frequency band independent encoder. The frequency band independent encoder divides and encodes the time-frequency features in the frequency band dimension to obtain the feature representation corresponding to each frequency band. S22: The feature representation encoded by the frequency band independent encoder is input into the acoustic discrete backbone network for modeling. S23: The features output by the vocal part separation backbone network are respectively connected to multiple vocal part prediction branches. Each vocal part prediction branch is independent of each other and is used to generate the separation results corresponding to the lead vocal part, backing vocal part and musical accompaniment vocal part respectively. S3: Angle for model training: S31: During the model training phase, the constructed multitrack music training data is input into the model for training; S32: During training, the mixed music audio is encoded by time-frequency transformation and frequency band independent encoder and then input into the vocal part separate backbone network for forward calculation; after the model forward calculation, the prediction results of the lead vocal part, backing vocal part and music accompaniment part are output respectively, and the corresponding multi-track vocal part audio track is used as the supervision signal to calculate the vocal part separate reconstruction loss of the prediction results for parameter update. S34: Separate music accompaniment based on the trained model.

2. The music accompaniment separation method based on voice structure perception as described in claim 1, characterized in that, The method for constructing angles from the data is as follows: S11: During the data construction phase, only multi-track music with a single lead singer is selected for training, and samples with multiple lead singers singing in rounds, duets, or with unstable lead singer identities are removed. S12: When extracting speaker codes, since the lead vocal part exists as an independent audio track in the multi-track data, the lead vocal track is directly used as the speaker feature extraction object, and the context-aware mask speaker coding model is used to encode and extract the lead vocal track.

3. The music accompaniment separation method based on voice structure perception as described in claim 1, characterized in that, The separate voice backbone network structurally integrates global speaker coding as conditional information for feature calculation.

4. The music accompaniment separation method based on voice structure perception as described in claim 1, characterized in that, It also includes the following steps: S24: Establish an auxiliary structure module that is connected only during the training phase to perform structural modeling of the output results of the voice prediction branch. The auxiliary structure module consists of multiple discriminant sub-networks, specifically including a multi-period discriminator, a multi-scale discriminator, and a multi-resolution spectral discriminator, which respectively model and score the output results of the voice prediction branch from different periodic characteristics, different time scales, and different spectral resolutions.

5. The music accompaniment separation method based on voice structure perception as described in claim 4, characterized in that, During the training process of S32, the speaker code extracted from the lead vocal track is injected into the vocal segment backbone network as global conditional information. Specifically, it is applied to the intermediate layer features of the vocal segment backbone network through conditional modulation, and the global conditional information is discarded with a preset probability.

6. The music accompaniment separation method based on voice structure perception as described in claim 5, characterized in that, During the training process of S32, the auxiliary structure module constructs the output results of the voice prediction branch obtained by the model prediction and the corresponding real voice track into a variety of voice combination forms according to preset rules, and models the voice combination audio composed of the voice prediction branch output results and the corresponding reference voice combination audio input to the discriminant subnetwork.

7. The music accompaniment separation method based on voice structure perception as described in claim 6, characterized in that, The vocal part combination methods include: combining the lead vocal part from the vocal part prediction branch output with the actual musical accompaniment part, and using the combination of the actual lead vocal part and the actual musical accompaniment part as the corresponding reference combination; and combining the lead vocal part from the vocal part prediction branch output with the actual backing vocal part, or combining the actual lead vocal part with the backing vocal part from the vocal part prediction branch output, and using the combination of the actual lead vocal part and the actual backing vocal part as the corresponding reference combination. The discriminant subnetwork outputs a relative score based on the results of the aforementioned voice combination form and updates its parameters using a ranking-based loss function. Simultaneously, the voice-separated backbone network updates its parameters based on the output of the discriminant subnetwork with an optimization objective different from that of the discriminant subnetwork, thereby achieving joint training of the voice-separated backbone network and the auxiliary structure module.

8. A music accompaniment separation system based on voice structure perception, characterized in that, A music accompaniment separation system is established using the music accompaniment separation method as described in any one of claims 1-7, thereby achieving music accompaniment separation.