A Melody Extraction Method, System and Related Products Based on Deep Feature Extraction

By combining the multimodal feature extraction method of residual neural network, bidirectional long and short-term memory network and convolutional recurrent neural network, the problem of low melody extraction accuracy caused by the semi-supervised extreme learning machine network model relying on a single network structure is solved, and higher melody extraction accuracy and robustness are achieved.

CN120032614BActive Publication Date: 2025-06-20XI AN JIAOTONG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510498636.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-21
Publication Date
2025-06-20
Estimated Expiration
2045-04-21

AI Technical Summary

Technical Problem

In the prior art, the semi-supervised extreme learning machine network model relies on the output of a single network structure feature, resulting in a low accuracy of melody extraction.

Method used

Through the synergy between residual neural networks, maximum pooling layer, bidirectional long and short-term memory networks and convolutional recurrent neural networks, multimodal features of audio signals are extracted, and feature vectors are fused in series, and input them to the semi-supervised extreme learning machine network model for melody extraction.

Benefits of technology

It improves the accuracy and robustness of melody extraction, enhances the ability to characterize melody features, and solves the problem of insufficient expression of features of traditional models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120032614B_ABST
    Figure CN120032614B_ABST
Patent Text Reader

Abstract

The present invention discloses a melody extraction method, system and related products based on deep feature extraction, belonging to the technical field of music information retrieval. The present invention uses a residual neural network layer, a max pooling layer and a bidirectional long short-term memory network layer to extract the features of the second spectrogram to obtain a feature vector; uses a convolutional recurrent neural network layer to extract the features of the second spectrogram to obtain a feature vector, and concatenates and fuses the two into a feature vector; finally, the melody of the audio signal to be processed is extracted through a semi-supervised extreme learning machine network model. Through the synergistic effect of the residual neural network, the bidirectional long short-term memory network and the convolutional recurrent neural network, this method respectively extracts the deep temporal features and local frequency domain features of the second spectrogram, realizes multi-modal feature complementarity through concatenation and fusion, enhances the representation ability of melody features, solves the problem of insufficient feature expression caused by the traditional SSELM model relying on a single network structure, and improves the accuracy of melody extraction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of music information retrieval, and in particular to a melody extraction method, system and related products based on deep feature extraction. Background Art

[0002] In the field of music information retrieval, melody extraction is a crucial task. As one of the core elements of music expression, melody plays an indispensable role in many music applications such as automatic music transcription, humming search technology, music recommendation and personalized services. Accurately extracted melodies can not only provide a solid foundation for music creation and interpretation, but also bring users a more personalized and immersive music experience, which has a far-reaching impact on industries such as music streaming platforms, audio equipment manufacturers, and copyright management companies. However, with the increasing diversification of music application scenarios, higher requirements are placed on the accuracy of melody extraction technology.

[0003] In recent years, deep learning models have made remarkable progress in the field of melody extraction. With its powerful automatic feature learning ability, deep learning models can mine deep feature representations from massive music data and improve the accuracy of melody extraction through training on large-scale data sets. However, deep learning models rely heavily on large-scale manually labeled data sets for training, while labeled data sets that can be used for melody extraction are relatively scarce; in addition, the training process of deep learning models is often time-consuming and labor-intensive, requiring high-performance computing resources and large-capacity storage space. In addition, the capacity and diversity of the training set have a decisive impact on the performance of the deep learning model. If the training set is biased or insufficient, the generalization ability of the model in real scenarios will be limited. In order to overcome the limitations of deep learning models, the Semi-Supervised Extreme Learning Machine Network (SSELM) model provides a new solution for music main melody extraction. As a semi-supervised learning method, the SSELM model cleverly combines the advantages of labeled data and unlabeled data, improves the performance of the model by utilizing unlabeled data, and maintains the efficiency and generalization performance of the Extreme Learning Machine (ELM). Although the SSELM model has shown some potential in music melody extraction, it still has some shortcomings. Specifically, the SSELM model relies on the feature output of a single network structure, which leads to its relatively limited feature expression ability and makes it difficult to fully explore the rich information in the music signal, thus affecting the accuracy of melody extraction.

[0004] Therefore, how to achieve multimodal feature fusion to improve the accuracy of music main melody extraction has become a technical problem that technical personnel in this field urgently need to overcome. Summary of the Invention

[0005] The object of the present invention is to provide a melody extraction method, system and related products based on deep feature extraction, so as to overcome the problem of low accuracy of melody extraction caused by the single network structure feature output of the semi-supervised extreme learning machine network model in the prior art.

[0006] The present invention solves the above technical problems through the following technical solutions:

[0007] A melody extraction method based on deep feature extraction includes the following steps:

[0008] Preprocess the audio signal to be processed to obtain a first spectrogram, and based on the phase spectrum method, correct the first spectrogram to obtain a second spectrogram;

[0009] Extract the features of the second spectrogram through a residual neural network layer, a max pooling layer and a bidirectional long short-term memory network layer to obtain a feature vector ; extract the features of the second spectrogram through a convolutional recurrent neural network layer to obtain a feature vector , the feature vector and the feature vector have the same dimension; fuse the feature vector and the feature vector in a concatenated manner along the dimension of the feature vector to obtain a fused feature vector ;

[0010] Input the fused feature vector into the trained semi-supervised extreme learning machine network model to obtain the melody of the audio signal to be processed.

[0011] A further improvement of the present invention is that the preprocessing of the audio signal to be processed to obtain a first spectrogram and correcting the first spectrogram based on the phase spectrum method to obtain a second spectrogram specifically includes:

[0012] Resample the audio signal to be processed to obtain an audio signal A;

[0013] Frame the audio signal A to obtain a framed signal B; perform time-frequency conversion on the framed signal B using the short-time Fourier transform to generate a complex matrix;

[0014] Visualize the complex matrix to obtain a first spectrogram;

[0015] Based on the complex matrix, calculate the instantaneous frequency and instantaneous amplitude of the framed signal B using the phase spectrum method;

[0016] Based on the instantaneous frequency and instantaneous amplitude, correct the first spectrogram to obtain a second spectrogram.

[0017] A further improvement of the present invention lies in that: the sampling rate of resampling is 16 kHz; the frame processing is performed through a Hann window, the window length of the Hann window is set to 1024, and the window shift of the Hann window is set to 10 ms.

[0018] A further improvement of the present invention lies in that: the complex matrix is expressed as:

[0019]

[0020] wherein, is a complex matrix; n is a frame index; m is a frequency index; is the amplitude; is the phase; is the imaginary unit.

[0021] A further improvement of the present invention lies in that: the instantaneous frequency is expressed as:

[0022]

[0023] wherein, is the instantaneous frequency; is the sampling rate; is the window length;

[0024] is a correction coefficient, specifically:

[0025]

[0026] wherein, is the window shift of the Hann window; () is a correction function; is the phase angle at frame index n and frequency index m; is the phase angle at frame index n - 1 and frequency index m;

[0027] The instantaneous amplitude is expressed as:

[0028]

[0029] wherein, is the instantaneous amplitude; w() is the Hann window function; Q is the total number of frequency points in the short-time Fourier transform.

[0030] A further improvement of the present invention lies in that:

[0031] extracting the features of the second spectrogram through a residual neural network layer, a max pooling layer, and a bidirectional long short-term memory network layer to obtain a feature vector , specifically including:

[0032] Input the second spectrogram into the residual neural network layer to perform normalization, non-linear activation, 1×1 convolution, and 3×3 convolution processing in sequence to obtain a primary feature map;

[0033] Input the primary feature map into the max pooling layer to perform spatial dimension downsampling to obtain an intermediate feature sequence;

[0034] Input the intermediate feature sequence into the bidirectional long short-term memory network layer to perform temporal modeling, extract bidirectional features, and obtain a feature vector 。

[0035] The present invention also provides a melody extraction system based on deep feature extraction, including:

[0036] A first module for preprocessing a to-be-processed audio signal to obtain a first spectrogram, and based on the phase spectrum method, correcting the first spectrogram to obtain a second spectrogram;

[0037] A second module for extracting features of the second spectrogram through a residual neural network layer, a max pooling layer, and a bidirectional long short-term memory network layer to obtain a feature vector ; extracting features of the second spectrogram through a convolutional recurrent neural network layer to obtain a feature vector wherein the feature vector and the feature vector have the same dimension; fusing the feature vector and the feature vector in a concatenated manner along the dimension of the feature vector to obtain a fused feature vector ;

[0038] A third module for inputting the fused feature vector into a trained semi-supervised extreme learning machine network model to generate the melody of the to-be-processed audio signal.

[0039] The present invention also provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, the steps of the above-mentioned melody extraction method based on deep feature extraction are implemented.

[0040] The present invention also provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the above-mentioned melody extraction method based on deep feature extraction are implemented.

[0041] The present invention also provides a computer program product, including a computer program, and when the computer program is executed by a processor, the steps of the above-mentioned melody extraction method based on deep feature extraction are implemented.

[0042] Compared with the prior art, the positive and progressive effects of the present invention are as follows:

[0043] The melody extraction method based on deep feature extraction provided by the present invention corrects the first spectrogram based on the phase spectrum method to obtain the second spectrogram; uses a residual neural network layer, a max pooling layer, and a bidirectional long short-term memory network layer to extract the features of the second spectrogram to obtain a feature vector ; uses a convolutional recurrent neural network layer to extract the features of the second spectrogram to obtain a feature vector and concatenates and fuses the two into a feature vector ; finally, through a semi-supervised extreme learning machine network model, the melody of the audio signal to be processed is obtained. Through the synergistic effect of the residual neural network, the bidirectional long short-term memory network and the convolutional recurrent neural network, the deep temporal features and local frequency domain features of the spectrogram are respectively extracted, and multi-modal feature complementarity is realized through concatenation and fusion, enhancing the representation ability of the melody features, and solving the problem that the traditional SSELM model relies on a single network structure resulting in insufficient feature expression, thereby improving the accuracy of melody extraction.

[0044] Furthermore, the first spectrogram is generated by short-time Fourier transform, and the first spectrogram is corrected by the phase spectrum method, effectively suppressing the spectral ambiguity phenomenon generated by the traditional short-time Fourier transform, realizing the enhancement of the fundamental frequency component energy and the suppression of interference noise, generating the optimized second spectrogram, and providing a clear data basis for subsequent feature extraction. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] The accompanying drawings in the specification are used to provide a further understanding of the present invention, and constitute a part of the present invention. The schematic embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation to the present invention.

[0046] Figure 1 is a schematic flow chart of a melody extraction method based on deep feature extraction of the present invention;

[0047] Figure 2 is a flow block diagram of obtaining the fused feature vector of the present invention ;

[0048] Figure 3 is the error rate on the unlabeled training set and the validation set under different melody extraction methods;

[0049] Figure 4 is the average accuracy of different melody extraction methods on three test sets;

[0050] Figure 5 is the melody pitch map of the audio file from the dataset MIR-1K extracted by the method of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0051] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0052] In the description of the present invention, it should be understood that the terms "include" and "comprise" indicate the presence of the described features, wholes, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components, and / or their combinations.

[0053] It should also be understood that the terms used in the specification of the present invention are merely for the purpose of describing specific embodiments and are not intended to limit the present invention. As used in the specification of the present invention and the appended claims, unless the context clearly indicates otherwise, the singular forms "a", "an", and "the" are intended to include the plural forms.

[0054] It should be understood that although the terms first, second, third, etc. may be used in the embodiments of the present invention to describe preset ranges, etc., these preset ranges should not be limited to these terms. These terms are only used to distinguish the preset ranges from each other. For example, without departing from the scope of the embodiments of the present invention, the first preset range may also be referred to as the second preset range, and similarly, the second preset range may also be referred to as the first preset range.

[0055] Depending on the context, the word "if" as used herein may be interpreted as "when" or "while" or "in response to determining" or "in response to detecting". Similarly, depending on the context, the phrase "if determined" or "if detected (stated condition or event)" may be interpreted as "when determined" or "in response to determining" or "when detected (stated condition or event)" or "in response to detecting (stated condition or event)".

[0056] Various structural schematic diagrams according to the disclosed embodiments of the present invention are shown in the drawings. These figures are not drawn to scale, where for the purpose of clear expression, some details are enlarged and some details may be omitted. The shapes of various regions and layers shown in the figures and their relative sizes and positional relationships are merely exemplary, and may actually deviate due to manufacturing tolerances or technical limitations, and those skilled in the art can design regions / layers with different shapes, sizes, and relative positions according to actual needs.

[0057] The following further elaborates on the present invention in conjunction with the accompanying drawings and specific embodiments, which is an explanation rather than a limitation of the present invention.

[0058] See Figure 1 and Figure 2 , a melody extraction method based on deep feature extraction, comprising the following steps:

[0059] Preprocess the audio signal to be processed to obtain a first spectrogram, and based on the phase spectrum method, correct the first spectrogram to obtain a second spectrogram;

[0060] Extract the features of the second spectrogram through a residual neural network layer, a max pooling layer, and a bidirectional long short-term memory network layer to obtain a feature vector ; extract the features of the second spectrogram through a convolutional recurrent neural network layer to obtain a feature vector , the feature vector and the feature vector have the same dimension; fuse the feature vector and the feature vector in a concatenated manner along the dimension of the feature vector to obtain a fused feature vector ;

[0061] Input the fused feature vector into the trained semi-supervised extreme learning machine network model to obtain the melody of the audio signal to be processed.

[0062] Feature vector , where are respectively the features of the nth frame and the first dimension of the feature vector , the features of the nth frame and the second dimension of the feature vector , the features of the nth frame and the third dimension of the feature vector , …, the features of the nth frame and the Mth dimension of the feature vector ;

[0063] Feature vector , where are respectively the features of the nth frame and the first dimension of the feature vector , the features of the nth frame and the second dimension of the feature vector , the features of the nth frame and the third dimension of the feature vector , …, the features of the nth frame and the Mth dimension of the feature vector ; fuse the feature vector and the feature vector in a concatenated manner along the dimension of the feature vector to obtain a fused feature vector = , where are respectively the feature vectors The feature of the first dimension of the nth frame, the feature vector The feature of the second dimension of the nth frame, the feature vector The feature of the third dimension of the nth frame, …, the feature vector The feature of the Mth dimension of the nth frame.

[0064] Through the collaborative action of the residual neural network, the bidirectional long short-term memory network and the convolutional recurrent neural network, this method extracts the deep temporal features and local frequency domain features of the spectrogram respectively, and realizes the multimodal feature complementarity through concatenated fusion, enhancing the representation ability of the melody features, solving the problem of insufficient feature expression caused by the traditional SSELM model relying on a single network structure, and thus improving the accuracy of melody extraction.

[0065] Using the method of the present invention to extract the audio file fdps_4_01.wav from the dataset MIR-1K, see Figure 5 It can be seen that the method of the present invention can better outline the melody contour of the audio file, and the silent frames are assigned to one class, and it can be observed that there are also some silent frames with short time.

[0066] Specifically, the preprocessing of the audio signal to be processed is performed to obtain a first spectrogram, and the first spectrogram is corrected based on the phase spectrum method to obtain a second spectrogram, which specifically includes:

[0067] Resample the audio signal to be processed to obtain an audio signal A;

[0068] Perform frame splitting on the audio signal A to obtain a frame-split signal B; perform time-frequency conversion on the frame-split signal B using the short-time Fourier transform to generate a complex matrix;

[0069] Visualize the complex matrix to obtain a first spectrogram;

[0070] Based on the complex matrix, calculate the instantaneous frequency and instantaneous amplitude of the frame-split signal B using the phase spectrum method;

[0071] Based on the instantaneous frequency and instantaneous amplitude, correct the first spectrogram to obtain a second spectrogram.

[0072] Generating the first spectrogram through the short-time Fourier transform and correcting the first spectrogram using the phase spectrum method effectively suppresses the spectral blur phenomenon generated by the traditional short-time Fourier transform, realizes strengthening the energy of the fundamental frequency component and suppressing the interference noise, generates an optimized second spectrogram, and provides a clear data basis for subsequent feature extraction.

[0073] Specifically, the resampling rate is 16 kHz; the frame processing is performed through a Hann window, the window length of the Hann window is set to 1024, and the window shift of the Hann window is set to 10 ms.

[0074] Specifically, the complex matrix is expressed as:

[0075]

[0076] Wherein, is a complex matrix; n is the frame index; m is the frequency index; is the amplitude; is the phase; is the imaginary unit.

[0077] Specifically, the instantaneous frequency is expressed as:

[0078]

[0079] Wherein, is the instantaneous frequency; is the sampling rate; is the window length;

[0080] is the correction coefficient, specifically:

[0081]

[0082] Wherein, is the window shift of the Hann window; () is the correction function; is the phase angle at frame index n and frequency index m; is the phase angle at frame index n - 1 and frequency index m;

[0083] The instantaneous amplitude is expressed as:

[0084]

[0085] Wherein, is the instantaneous amplitude; w() is the Hann window function; Q is the total number of frequency points in the short-time Fourier transform.

[0086] Specifically, the features of the second spectrogram are extracted through a residual neural network layer, a max pooling layer, and a bidirectional long short-term memory network layer to obtain a feature vector , specifically including:

[0087] The second spectrogram is input into the residual neural network layer for normalization, non-linear activation, 1*1 convolution, and 3*3 convolution processing in sequence to obtain a primary feature map;

[0088] Input the primary feature map into the max pooling layer for spatial dimensional downsampling to obtain an intermediate feature sequence;

[0089] Input the intermediate feature sequence into the bidirectional long short-term memory network layer for temporal modeling to extract bidirectional features and obtain a feature vector .

[0090] Based on the same inventive concept, the present invention also provides a melody extraction system based on deep feature extraction, including:

[0091] The first module is used to preprocess the audio signal to be processed to obtain a first spectrogram, and correct the first spectrogram based on the phase spectrum method to obtain a second spectrogram;

[0092] The second module is used to extract the features of the second spectrogram through a residual neural network layer, a max pooling layer and a bidirectional long short-term memory network layer to obtain a feature vector ; extract the features of the second spectrogram through a convolutional recurrent neural network layer to obtain a feature vector , the feature vector and the feature vector have the same dimension; fuse the feature vector and the feature vector in a concatenated manner along the dimension of the feature vector to obtain a fused feature vector ;

[0093] The third module is used to input the fused feature vector into the trained semi-supervised extreme learning machine network model to generate the melody of the audio signal to be processed.

[0094] To verify the effectiveness of the method of the present invention, based on three public datasets, a comparative experiment is carried out between this method and 2 methods for melody extraction using deep learning in recent years. Among them, the first method is an extraction method that directly sends the features obtained by performing a constant Q transform on the audio file into a semi-supervised extreme learning machine network for training and prediction, and hereinafter, cqt is used to refer to the first method; the second method is an extraction method that directly sends the features obtained by performing a constant Q transform on the audio file into an extreme learning machine network for training and prediction, and hereinafter, stft is used to refer to the second method. The audio files of the three public datasets are mixed together and randomly divided into a training set (including a labeled training set and an unlabeled training set), a validation set and a test set, and there is no overlap between these datasets. For detailed information, see Table 1.

[0095] Table 1 Division of the training set, validation set and test set

[0096]

[0097] Among them, the dataset ISMIR04 contains 20 excerpted songs of different genres, such as jazz, pop music, opera, etc., with a sampling rate of 44.1 kHz. The duration of these recordings is about 20 seconds, and the melodic pitch marking interval is 5.8 ms. In this test, 15 out of the 20 excerpted songs are selected, excluding jazz and MIDI (Musical Instrument Digital Interface) files; the dataset MIREX05 includes 13 excerpted songs with a duration between 24 seconds and 39 seconds, a sampling rate of 44.1 kHz, and a melodic pitch marking interval of 10 ms. To compare the results under the same conditions, 12 of them are selected as the test set; the dataset MIR-1K contains 1000 song segments cut from 110 karaoke songs, with a duration of 4s - 13s, a sampling rate of 16 kHz, and a melodic pitch marking interval of 10 ms.

[0098] The total amount of audio data for these three datasets is 1033. Each dataset is randomly divided into four subsets, with 30 audio files as the labeled training set, 30 audio files as the unlabeled training set, 25 audio files as the validation set, and the other 948 audio files as the test set. All the audio files in the test set are unlabeled. Generally speaking, the more training samples, the higher the accuracy. To prove the generalization of the method proposed in the present invention, only a small part of the audio files (about 3%) are labeled, and the proportion of unlabeled audio files for testing is much larger. It can be seen that there are 5 audio files from the dataset ISMIR04 in the training set, 15 audio files in the test set, and only one audio file from the dataset MIREX05 is covered in the unlabeled training set, and the other 12 audio files belong to the test set. Since the dataset MIR-1K contains most of the records, it has records in all sets.

[0099] The second spectrogram and the fused feature vectors of this method are implemented based on Keras in Python 3.6 on a computer configured with an NVIDIA GTX 1080 Ti GPU (NVIDIA GeForce GTX 1080 Ti Graphics Processing Unit), and then the saved feature vectors are input into SSELM for training and testing to extract the melody.

[0100] The accuracy of melody extraction is evaluated from three metrics: Raw Pitch Accuracy (RPA), Raw Chroma Accuracy (RCA), and Overall Accuracy (OA). The definitions of these metrics are as follows:

[0101]

[0102] In the formula, (True Positives for pitch) represents the number of voiced frames where the difference between the extracted pitch and the true pitch is within half a semitone. GM (Ground - truth Melody frames) represents the total number of voiced frames in the test set.

[0103]

[0104] In the formula, (True Positive for chroma) represents the number of voiced frames where the extracted pitch and the true pitch belong to the same note.

[0105]

[0106] In the formula, (True Negative) represents the number of unvoiced frames correctly extracted, (Total Observations) represents the total number of frames in the test set, including both voiced and unvoiced frames.

[0107] To improve the accuracy of pitch extraction, a pitch label quantized by 6.25 cents (1 / 16 semitone) is converted to the frequency scale (Hz) for comparison with the true pitch. The relationship between frequency and cents is non - linear, and the conversion method is shown in formula (4).

[0108]

[0109] In the formula, i is the MIDI note number estimated by the main network for melody extraction. 69 represents the MIDI note number corresponding to a frequency of 440 Hz. The mir - eval library designed for target evaluation in MIR tasks in Python is used for calculation.

[0110] By calculating the accuracies of cqt, stft, and the method of the present invention on three datasets, it is evaluated whether this method helps to improve the performance of melody extraction. See Figure 3, under different melody extraction methods, the error rates of the unlabeled training set and the validation set can be found. It can be seen that the error rates of the method of the present invention on the unlabeled training set and the validation set are the lowest, which are 24.756% and 14.065% respectively, and are reduced by about 10% compared with cqt respectively.

[0111] As shown in Table 2, it can be found that on the test set of the dataset ISMIR04, the method of the present invention has achieved the highest accuracy rates of 53.45% and 54.68% on RPA and RCA respectively, which are increased by about 11.09% and 7.76% compared with cqt respectively; while on OA, it has achieved the second highest accuracy rate of 56.55%, which is only about 2.26% different from cqt. On the one hand, it is because the training set only contains 5 audio files of the dataset ISMIR04, and two of them belong to the unlabeled training set; on the other hand, the audio features of the dataset ISMIR04 are quite different from those of the other two datasets.

[0112] Table 2 Accuracy rates of different melody extraction methods on the test set of the dataset ISMIR04

[0113]

[0114] As shown in Table 3, it can be found that on the test set of the dataset MIREX05, the method of the present invention has achieved the highest accuracy rates of 68.3%, 71.9% and 72.23% on RPA, RCA and OA respectively. Since there is only one audio file from the MIREX05 dataset in the unlabeled training set, the robustness and generalization of the method of the present invention can be seen.

[0115] Table 3 Accuracy rates of different melody extraction methods on the test set of the dataset MIREX05

[0116]

[0117] As shown in Table 4, it can be found that on the test set of the dataset MIR-1K, the method of the present invention has achieved the highest accuracy rates of 88.52% and 88.76% on RCA and OA respectively, which are 14.39% and 12.03% higher than the cqt method respectively; while on RPA, it has achieved the second highest accuracy rate of 84.74%, which is only 0.31% different from the highest accuracy rate. The reason why the three methods have relatively high accuracy rates on the test set of the dataset MIR-1K is that there are labeled training sets and unlabeled training sets from the dataset MIR-1K, and the accuracy of the method of the present invention can be seen.

[0118] Table 4 Accuracy rates of different melody extraction methods on the test set of the dataset MIR-1K

[0119]

[0120] See also Figure 4 , provides the mean values ​​of the indicators on three public datasets. Overall, the method of the present invention obtains the highest RPA and RCA and the second highest OA, which are 18.5%, 16.51% and 10.96% higher than cqt, respectively. It can be seen that the method of the present invention still has significant robustness when the dataset is small. At the same time, RPA and RCA represent the accuracy between voiced frames, which can prove that the method of the present invention has good performance in melody extraction.

[0121] In summary, the method of the present invention combines multi-feature fusion with a semi-supervised extreme learning machine network, effectively solving the problems of single feature expression and strong dependence on labeled data in traditional melody extraction methods, significantly improving the accuracy and robustness of melody extraction in complex music scenes, and can be widely used in automatic music transcription, humming search and intelligent recommendation systems.

[0122] Based on the same inventive concept, an embodiment of the present application provides a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the melody extraction method based on deep feature extraction are implemented, wherein the memory may include a memory, such as a high-speed random access memory, and may also include a non-volatile memory, such as at least one disk memory, etc.; the processor, the network interface, and the memory are interconnected through an internal bus, and the internal bus may be an industrial standard architecture bus, a peripheral component interconnection standard bus, an extended industrial standard structure bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. The memory is used to store programs. Specifically, the program may include a program code, and the program code includes computer operation instructions. The memory may include a memory and a non-volatile memory, and provide instructions and data to the processor.

[0123] Based on the same inventive concept, an embodiment of the present application provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the melody extraction method based on deep feature extraction are implemented. Specifically, the computer-readable storage medium includes, but is not limited to, for example, a volatile memory and / or a non-volatile memory. The volatile memory may include a RAM (Random Access Memory) and / or a cache memory, etc. The non-volatile memory may include a ROM (Read Only Memory), a hard disk, a flash memory, an optical disk, a magnetic disk, etc.

[0124] Based on the same inventive concept, an embodiment of the present application provides a computer program product, which includes a computer program stored on a computer-readable storage medium. The computer program includes program instructions that, when executed by a computer device, cause the computer device to execute the steps of the above-mentioned melody extraction method based on deep feature extraction.

[0125] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method or a computer program product. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM (Compact Disc Read-Only Memory), optical memory, etc.) containing computer-usable program code.

[0126] The present invention is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each flow and / or block in the flowchart and / or block diagram, as well as the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer device or other programmable data processing devices generate a device for implementing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0127] These computer program instructions can also be stored in a computer-readable memory that can direct a computer device or other programmable data processing devices to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device, and the instruction device implements the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0128] These computer program instructions can also be loaded onto a computer device or other programmable data processing devices, so that a series of operation steps are executed on the computer device or other programmable devices to generate a process implemented by the computer device. Thus, the instructions executed on the computer device or other programmable devices provide steps for implementing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0129] Although the preferred embodiments of the present invention have been described, additional changes and modifications can be made to these embodiments by those skilled in the art once they learn of the basic inventive concept. Therefore, the appended claims are intended to be construed to include the preferred embodiments as well as all changes and modifications that fall within the scope of the present invention.

[0130] Obviously, those skilled in the art can make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalent technologies, the present invention is also intended to include these modifications and variations.

Claims

1. A melody extraction method based on deep feature extraction, characterized in that: The following steps are involved: Preprocessing the audio signal to be processed to obtain a first spectrum diagram, and correcting the first spectrum diagram based on a phase spectrum method to obtain a second spectrum diagram; The features of the second spectrum graph are extracted through the residual neural network layer, the maximum pooling layer and the bidirectional long short-term memory network layer to obtain the feature vector ; Extract the features of the second spectrum graph through the convolutional recurrent neural network layer to obtain the feature vector , the feature vector and the eigenvector The dimensions are the same; Fuse feature vectors in series along their dimensions and the eigenvector , and get the fused feature vector ; The fused feature vector Input into the trained semi-supervised extreme learning machine network model to obtain the melody of the audio signal to be processed; The preprocessing of the audio signal to be processed to obtain a first spectrum diagram, and the correction of the first spectrum diagram based on the phase spectrum method to obtain a second spectrum diagram specifically include: Resample the audio signal to be processed to obtain an audio signal A; Performing frame processing on the audio signal A to obtain a frame signal B; performing time-frequency conversion on the frame signal B using short-time Fourier transform to generate a complex matrix; Visualize the complex matrix and obtain the first spectrum diagram; Based on the complex matrix, the instantaneous frequency and instantaneous amplitude of the frame signal B are calculated by using the phase spectrum method; Based on the instantaneous frequency and the instantaneous amplitude, the first spectrum diagram is corrected to obtain a second spectrum diagram; The feature vector is obtained by extracting the features of the second spectrum graph through the residual neural network layer, the maximum pooling layer and the bidirectional long short-term memory network layer. , specifically including: The second spectrum map is input into the residual neural network layer for normalization, nonlinear activation, 1*1 convolution and 3*3 convolution processing in sequence to obtain a primary feature map; The primary feature map is input into the maximum pooling layer for spatial dimension downsampling to obtain an intermediate feature sequence; The intermediate feature sequence is input into the bidirectional long short-term memory network layer for time series modeling, and the bidirectional features are extracted to obtain the feature vector .

2. The melody extraction method based on deep feature extraction according to claim 1, characterized in that: The resampling sampling rate is 16 kHz; the framing process is performed through the Hann window, the window length of the Hann window is set to 1024, and the window shift of the Hann window is set to 10 ms.

3. A melody extraction method based on deep feature extraction according to claim 2, characterized in that: The complex matrix is ​​represented as: in, is a complex matrix; n is a frame index; m is a frequency index; is the amplitude; is the phase; Is an imaginary unit.

4. A melody extraction method based on deep feature extraction according to claim 3, characterized in that: The instantaneous frequency is expressed as: in, is the instantaneous frequency; is the sampling rate; is the window length; is the correction factor, specifically: in, The window shift for the Hann window; () is the correction function; is the phase angle at frame index n and frequency index m; is the phase angle at frame index n-1 and frequency index m; The instantaneous amplitude is expressed as: in, is the instantaneous amplitude; w() is the Hann window function; Q is the total number of frequency points in the short-time Fourier transform.

5. A melody extraction system based on deep feature extraction, characterized in that: The steps for executing the melody extraction method based on deep feature extraction as claimed in any one of claims 1 to 4 include: The first module is used to preprocess the audio signal to be processed to obtain a first spectrum diagram, and to correct the first spectrum diagram based on the phase spectrum method to obtain a second spectrum diagram; the preprocessing of the audio signal to be processed to obtain the first spectrum diagram, and to correct the first spectrum diagram based on the phase spectrum method to obtain the second spectrum diagram specifically includes: Resample the audio signal to be processed to obtain an audio signal A; Performing frame processing on the audio signal A to obtain a frame signal B; performing time-frequency conversion on the frame signal B using short-time Fourier transform to generate a complex matrix; Visualize the complex matrix and obtain the first spectrum diagram; Based on the complex matrix, the instantaneous frequency and instantaneous amplitude of the frame signal B are calculated by using the phase spectrum method; Based on the instantaneous frequency and the instantaneous amplitude, the first spectrum diagram is corrected to obtain a second spectrum diagram; The second module is used to extract the features of the second spectrum graph through the residual neural network layer, the maximum pooling layer and the bidirectional long short-term memory network layer to obtain the feature vector ; Extract the features of the second spectrum graph through the convolutional recurrent neural network layer to obtain the feature vector , the feature vector and the eigenvector The dimensions of the feature vectors are the same; the feature vectors are fused in series along the dimension of the feature vectors and the eigenvector , and get the fused feature vector The features of the second spectrum graph are extracted through the residual neural network layer, the maximum pooling layer and the bidirectional long short-term memory network layer to obtain the feature vector , specifically including: The second spectrum map is input into the residual neural network layer for normalization, nonlinear activation, 1*1 convolution and 3*3 convolution processing in sequence to obtain a primary feature map; The primary feature map is input into the maximum pooling layer for spatial dimension downsampling to obtain an intermediate feature sequence; The intermediate feature sequence is input into the bidirectional long short-term memory network layer for time series modeling, and the bidirectional features are extracted to obtain the feature vector ; The third module is used to transform the fused feature vector Input into the trained semi-supervised extreme learning machine network model to generate the melody of the audio signal to be processed.

6. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that: When the processor executes the computer program, the steps of the melody extraction method based on deep feature extraction as described in any one of claims 1 to 4 are implemented.

7. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the melody extraction method based on deep feature extraction as described in any one of claims 1 to 4 are implemented.

8. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the melody extraction method based on deep feature extraction as described in any one of claims 1 to 4 are implemented.

Citation Information

Patent Citations

  • BLSTM speech emotion recognition method based on multiple output feature fusion

    CN110164476A

  • Audio detection

    US20240412728A1