Method for Extracting Polyphonic Music Singing Melody Based on Modeling of Musical Tone Signal Spectrogram
Through the method based on spectral diagram modeling of musical tone signal, the graph convolution network and significance function are used to solve the accuracy and robustness of singing melody extraction in multitone music, and the efficient melody extraction effect is achieved.
Patent Information
- Application Number
- CN202211120049.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-09-14
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2042-09-14
AI Technical Summary
The prior art is difficult to accurately extract the singing melody in multi-tone music, especially when the time domain and the frequency domain are superimposed on each other, and there is a lack of accurate theoretical support.
Using a method based on spectral diagram modeling of musical tone signals, the logarithmic frequency amplitude spectrum is obtained through normal Q transformation, the graph structure is constructed, and complex input and output mapping functions are learned using graph convolutional networks, and post-processing is combined with the significance function to fine-tune melody pitch estimation.
It achieves high accuracy and robustness of melody extraction, and can effectively handle singing melody extraction tasks in multi-tone music.
Smart Images

Figure CN115579018B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of audio signal processing, and in particular, to a method for extracting the melody of multi-tone music singing based on the modeling of the spectrogram of the musical tone signal. Background Art
[0002] Multi-tone music is a mixed signal of human voice and accompaniment, and there may be two or more sound sources sounding simultaneously, making the human voice and accompaniment superimposed on each other in both the time domain and the frequency domain, thus making it difficult to accurately extract the singing melody. At present, the perceptual attributes of the singing melody perceived by the human ear still cannot be accurately described, so that the modeling of melody extraction still lacks accurate theoretical support. Energy saliency and temporal continuity are two basic bases for melody extraction, and existing methods model energy saliency and temporal continuity in different ways. Existing melody extraction methods include the saliency method, the source separation method, and the machine learning method. The saliency method includes steps such as spectral analysis, multi-pitch estimation, and melody trajectory tracking. Usually, the saliency function is artificially set to model the energy saliency of the melody, and the scientificity and rationality of these saliency functions are difficult to guarantee. The source separation method first separates or enhances the singing component from the mixed signal, and then uses the single-pitch estimation method to estimate the melody pitch. The source separation problem belongs to the category of underdetermined problems and still cannot obtain satisfactory results, thus limiting the performance of such methods. The machine learning method includes traditional machine learning methods and deep learning methods. Since the melody is sometimes drowned by noise, the traditional machine learning methods have poor robustness, while the deep learning methods have the disadvantages of large parameter scale and poor interpretability. Summary of the Invention
[0003] According to the problems existing in the prior art, the present invention discloses a method for extracting the melody of multi-tone music singing based on the modeling of the spectrogram of the musical tone signal, which specifically includes the following steps:
[0004] Perform a constant-Q transform on the audio signal to obtain a logarithmic frequency amplitude spectrum, intercept the amplitude spectrum within a certain frequency range, splice the amplitude spectra of the consecutive odd frames before and after the i-th frame to obtain a spliced amplitude spectrum, and use the spliced amplitude spectrum as the input feature of the i-th frame, denoted as X i ;
[0005] Construct an adjacency matrix corresponding to the spliced amplitude spectrum;
[0006] Use each frequency point of the spliced amplitude spectrum as a node of the graph structure, and determine the edges according to the adjacency matrix, that is, the connection relationship of each node, so as to represent each frequency component of the musical tone signal with a graph structure;
[0007] Discretize the melody pitch frequency corresponding to the i-th frame signal to obtain a one-hot vector of the output label, and use the one-hot vector as the output of the graph convolutional network to obtain the output label Y corresponding to the input feature X of the i-th frame i i ;
[0008] Train the graph convolutional network to obtain optimal parameters;
[0009] Adopt the trained network parameters to perform melody pitch prediction on the test set, and take the frequency corresponding to the maximum value in the output nodes of the graph convolutional network as the preliminary melody pitch estimate;
[0010] Perform median filtering on the preliminary melody pitch sequence obtained by the graph convolutional network to obtain a smooth melody pitch trajectory;
[0011] Frame the audio signal, then zero-pad each frame signal and perform short-time Fourier transform to obtain the short-time Fourier transform magnitude spectrum;
[0012] Use the phase vocoder to correct the instantaneous amplitude and instantaneous frequency of the short-time Fourier transform magnitude spectrum;
[0013] Calculate the saliency value frame by frame according to the saliency function;
[0014] Take the band region composed of a certain frequency range centered on the smooth melody pitch trajectory as the final singing melody output candidate range, search for the maximum saliency value within the candidate range, and take the frequency corresponding to the maximum saliency value as the final singing melody output result of the non-zero frequency segment; for the zero output of the graph convolutional network, no correction is performed.
[0015] The saliency function is:
[0016]
[0017] where a i is the amplitude of the i-th spectral peak, Tr(a i ) is the amplitude threshold function, and w(b, h, f i ) is the weight function.
[0018] Due to the adoption of the above technical solutions, the present invention provides a multi-tone music singing melody extraction method based on the modeling of the musical tone signal spectrogram. This method performs constant Q transform on the mixed audio signal to obtain the logarithmic frequency amplitude spectrum, constructs an adjacency matrix based on the frequency point position relationship between the fundamental frequency and each harmonic component of the same musical tone source, obtains a graph structure, uses a graph convolutional network to learn the complex input-output mapping function, and takes the frequency corresponding to the maximum value in the output nodes of each frame of the graph convolutional network as the preliminary melody pitch estimation result of this frame; adopts a post-processing step to construct a saliency spectrogram and fine-tune the melody pitch estimation. Therefore, this method has achieved high accuracy and robustness. Brief Description of the Drawings
[0019] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments recorded in the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0020] Figure 1 Flow chart of the method of the present invention
[0021] Figure 2 Time-domain waveform diagram of a piece of music signal in the present invention;
[0022] Figure 3 Constant Q transform amplitude spectrum diagram of this piece of music signal of the present invention
[0023] Figure 4 Schematic diagram of the adjacency matrix in the present invention
[0024] Figure 5 Initial melody sequence estimation of the present invention
[0025] Figure 6 Post-processing saliency map of the present invention
[0026] Figure 7 Final singing melody extraction result diagram of the present invention Detailed implementation manners
[0027] To make the technical solutions and advantages of the present invention clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention:
[0028] The singing melody extraction method proposed by the present invention is as Figure 1 shown. First, perform the constant Q transform on the mixed audio signal to obtain the logarithmic frequency amplitude spectrum. Secondly, obtain the graph structure based on the frequency point position relationship between the fundamental wave and each harmonic component of the same musical sound source. Then, use the constant Q transform amplitude spectrum as the input of the graph convolutional network, convert the melody pitch into a one-hot vector, and use it as the output of the graph convolutional network. Use the graph convolutional network to learn the complex input-output mapping function, and use the frequency corresponding to the maximum value in each frame output node of the graph convolutional network as the preliminary melody pitch estimation result of this frame. Finally, adopt post-processing steps to construct a saliency spectrum diagram and fine-tune the melody pitch estimation.
[0029] Embodiment:
[0030] Considering that the melody has typical harmonicity, in the present invention, each frequency point in the music signal spectrum is represented by a node, and the harmonicity of the musical tone is represented by an edge. In this way, the internal relationship of each harmonic of a certain musical tone sound source can be represented by a graph structure, and then the extraction of the singing melody based on the graph is realized. In order to improve the smoothness of the melody pitch trajectory and reduce the quantization error, the present invention uses a pitch salience function to fine-tune the preliminary melody pitch estimation. The specific scheme includes the following steps:
[0031] S1: Given an audio signal, the time-domain waveform diagram is as Figure 2 shown. Using a frequency resolution of 12 frequency points / octave, perform a constant-Q transform on the audio signal to obtain a logarithmic frequency magnitude spectrum, as Figure 3 shown. Intercept the magnitude spectrum in the frequency range of 47.65 - 8141.46 Hz. Therefore, the magnitude spectrum of each frame of the audio signal has a total of 90 frequency points. Concatenate the continuous three frames of magnitude spectra from the previous frame to the next frame of each frame to obtain a concatenated magnitude spectrum with a length of 270, which is used as the input feature representation X i .
[0032] S2: Construct an adjacency matrix corresponding to the concatenated three-frame magnitude spectrum, as Figure 4 shown. The specific calculation formula is:
[0033]
[0034] where N = 90, h = 1, …, 5, i(j) = 1, …, 270.
[0035] S3: Each frequency point of the concatenated magnitude spectrum is used as a node of a graph structure, and the adjacency matrix defined by formula (1) determines the edges, that is, defines the connection relationship of each node. In this way, each component of the musical tone signal is represented by a graph structure.
[0036] S4: Discretize the melody pitch frequency corresponding to the i-th frame signal according to the resolution of 12 frequency points / octave to obtain a one-hot vector of the output label, and use it as the output of the graph convolutional network. In this way, the output label Y i corresponding to the input feature X i of the i-th frame is obtained.
[0037] S5: On the training set, perform parameter training. The loss function is selected as the binary cross-entropy function, the optimizer is selected as Adam, the learning rate is set to 0.001, train for 1000 epochs, the batch size is set to 256, and continuously iterate to obtain the optimal graph convolutional network parameters.
[0038] S6: Use the trained network parameters to perform melody pitch prediction on the test set, and use the frequency corresponding to the maximum value in the output nodes of the graph convolutional network as the preliminary melody pitch estimation, as Figure 5 shown.
[0039] S7: Median filter the preliminary melodic pitch sequence obtained by the graph convolutional network (with a filter window width of 7) to obtain a smooth melodic pitch trajectory.
[0040] S8: Frame the audio signal, with each frame containing 2048 points. Zero-pad each frame signal and perform a 8192-point short-time Fourier transform to obtain the short-time Fourier transform magnitude spectrum.
[0041] S9: Use a phase vocoder to correct the instantaneous amplitude and instantaneous frequency of the magnitude spectrum after the short-time Fourier transform.
[0042] S10: Calculate the saliency value frame by frame according to the saliency function, as Figure 6 shown; the saliency function is:
[0043]
[0044] where a i is the amplitude of the i-th spectral peak, Tr(a i ) is the amplitude threshold function, and w(b, h, f i ) is the weight function.
[0045] S11: Use the band region formed by the range of plus and minus 1.5 semitones centered on the smooth melodic pitch trajectory as the final singing melody output candidate range. Search for the maximum saliency value within this candidate range, and correct the non-zero output of the graph convolutional network with the frequency corresponding to the maximum saliency value. For the zero output of the graph convolutional network, no correction is performed. The result is as Figure 7 shown.
[0046] Considering that the fundamental wave and its harmonics in the logarithmic frequency domain are shift-invariant for different fundamental frequencies, the present invention constructs a graph structure in the logarithmic frequency domain to solve the problem of singing melody extraction, and automatically learns the parameters of the graph convolutional network in a data-driven mode to achieve the purpose of singing melody extraction with lightweight parameters.
[0047] The above is only a preferred specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention, according to the technical solution and inventive concept of the present invention, makes equivalent substitutions or changes, and should be covered by the protection scope of the present invention.
Claims
1. A method for extracting the melody of polyphonic music singing based on the modeling of the spectrogram of musical sound signals, characterized in that it includes: Performing a constant Q transform on the audio signal to obtain a logarithmic frequency amplitude spectrum, intercepting the amplitude spectrum within a certain frequency range, and splicing the amplitude spectrum of the i-th frame with the continuous three-frame amplitude spectra from the previous frame to the next frame of this frame to obtain a spliced amplitude spectrum, Use this spliced magnitude spectrum as the input feature for the i-th frame, denoted as X i ; Constructing an adjacency matrix corresponding to the splicing of the three-frame amplitude spectra, and the specific calculation formula is: where N = 90, h = 1, …, 5, i(j) = 1, …, 270; Regarding each frequency point of the spliced amplitude spectrum as a node of the graph structure, and determining the edges according to the adjacency matrix, that is, the connection relationships of each node, so as to represent each frequency component of the musical sound signal with a graph structure; Discretize the melodic pitch frequency corresponding to the i-th frame signal to obtain the one-hot vector of the output label, and use the one-hot vector as the output of the graph convolutional network to obtain the i-th frame input feature X i The corresponding output label Y i ; Training the graph convolutional network to obtain optimal parameters; Using the trained network parameters to predict the melody pitch on the test set, and taking the frequency corresponding to the maximum value among the output nodes of the graph convolutional network as the preliminary melody pitch estimate; Performing median filtering on the preliminary melody pitch sequence obtained by the graph convolutional network to obtain a smooth melody pitch trajectory; Framing the audio signal, then padding zeros to each frame signal and performing a short-time Fourier transform to obtain a short-time Fourier transform amplitude spectrum; Using a phase vocoder to correct the instantaneous amplitude and instantaneous frequency of the short-time Fourier transform amplitude spectrum; Calculating the saliency value frame by frame according to the saliency function; The saliency function is: where a i is the amplitude of the i-th spectral peak, Tr(a i ) is the amplitude threshold function, and w(b, h, f i ) is the weight function; Taking the band region formed by the range of plus and minus 1.5 semitones centered on the smooth melody pitch trajectory as the candidate range for the final singing melody output, searching for the maximum saliency value within this candidate range, and correcting the non-zero output of the graph convolutional network with the frequency corresponding to the maximum saliency value, and not correcting the zero output of the graph convolutional network.
Citation Information
Patent Citations
Underwater target identification method based on small sample training graph convolutional network
CN113111786A
Indoor passive moving target detection method based on graph convolutional neural network
CN114158004A