Audio classification method and system based on maximum and minimum peak and valley tracks of spectrum graph

By constructing the extremely large and small peak and valley feature matrix of the audio spectrum map, the problem of insufficient aggregation of peak and valley trajectory information in the existing technology is solved, and higher audio classification accuracy is achieved.

CN114842872BActive Publication Date: 2025-05-06YANGZHOU UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210373422.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-04-11
Publication Date
2025-05-06
Estimated Expiration
2042-04-11

AI Technical Summary

Technical Problem

The existing spectrum map-based audio classification algorithm lacks effective aggregation of peak and valley trajectory information in the spectrum map, resulting in a decrease in classification accuracy.

Method used

By constructing the extremely large and small peak and valley feature matrix of the spectrogram, combining the extremely large peak and extremely small valley information, the input convolutional neural network for feature aggregation is carried out to improve the accuracy of audio classification.

Benefits of technology

This method significantly improves the accuracy of audio classification by better utilizing the trajectory information in the spectrum graph, which is better than the traditional spectrum graph-based classification algorithm.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114842872B_ABST
    Figure CN114842872B_ABST
Patent Text Reader

Abstract

The present invention discloses an audio classification method and system based on the maximum and minimum peak and valley tracks of the spectrum graph. First, the audio is sliced ​​and the spectrum graph of each audio slice is calculated; then, based on the amplitude axis of the spectrum graph, the position and amplitude of the maximum value point of the amplitude are calculated and sorted, and the maximum position matrix and the maximum amplitude matrix are constructed respectively, and the maximum peak feature matrix is ​​constructed after connection; and the position and amplitude of the minimum value point of the amplitude are calculated and sorted, and the minimum position matrix and the minimum amplitude matrix are constructed respectively, and the minimum valley feature matrix is ​​constructed after connection, and then the maximum and minimum peak and valley feature matrix is ​​obtained; finally, the maximum and minimum peak and valley feature matrix is ​​input into the convolutional neural network, and the classification result of the audio data is output. The present invention conducts a more comprehensive exploration of the mutual relationship between the peak track of the spectrum graph and the valley track of the spectrum graph; the track features of the spectrum graph are aggregated before inputting the model, which can improve the accuracy of classification.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of audio recognition, and relates to an audio classification method and system based on maximum and minimum peak and valley tracks of a frequency spectrum graph. Background Art

[0002] Today's audio classification methods can be divided into the following categories according to the types of features used: audio classification algorithms based on waveforms, audio classification algorithms based on spectrograms, etc. The waveform-based audio classification algorithm directly uses the waveform of the audio as the input feature, and uses a neural network as a feature extractor. The waveform is a high-dimensional feature, and the deeper neural network designed for the waveform will have disadvantages such as long training time and easy overfitting. The spectrogram-based audio classification algorithm uses the intermediate feature representation of audio data - the spectrogram as the input feature, which can effectively utilize the information in the time domain and frequency domain of the spectrogram to improve the accuracy of audio classification. The core difficulty of the spectrogram classification algorithm lies in how to process the information in the spectrogram, aggregate or discard it, and construct the spectrogram features.

[0003] There are two feature extraction strategies for audio classification algorithms based on spectrograms: one is to use the neural network model in deep learning as a feature extractor to automatically extract features; the other is to manually extract features from the spectrogram based on domain knowledge. When manually extracting features from the spectrogram, it is easy to cause partial loss of audio features, resulting in a decrease in classification accuracy; in addition, previous audio classification algorithms based on spectrograms directly use deep learning models to extract features, without considering aggregating the peak and valley trajectory features of the spectrogram before inputting the model, and therefore lack exploration of the interconnection of trajectory information in the spectrogram. Summary of the invention

[0004] Purpose of the invention: In view of the shortcomings of the prior art, the purpose of the present invention is to provide an audio classification method and system based on the maximum and minimum peak and valley trajectories of the spectrum graph, using the peak and valley trajectories to describe the characteristic relationship of the audio spectrum graph, and aggregating the peak and valley trajectory features of the spectrum graph before inputting into the neural network model to improve the accuracy of classification.

[0005] Technical solution: To achieve the above-mentioned invention object, the present invention adopts the following technical solution:

[0006] The audio classification method based on the maximum and minimum peak and valley tracks of the spectrum graph includes the following steps:

[0007] (1) constructing a spectrogram feature for each audio: slicing the audio data to obtain multiple audio data slices, and calculating the spectrogram of each audio data slice;

[0008] (2) Based on the audio spectrum, the maximum peak feature matrix and the minimum valley feature matrix are calculated respectively to construct the final maximum and minimum peak valley feature matrix; including:

[0009] Based on the amplitude axis of the spectrum graph, the position and amplitude of the maximum amplitude point are calculated and sorted, and the maximum position matrix and the maximum amplitude matrix are constructed respectively. After connecting, the maximum peak feature matrix is ​​constructed;

[0010] Based on the amplitude axis of the spectrum graph, the position and amplitude of the minimum amplitude point are calculated and sorted, and the minimum position matrix and the minimum amplitude matrix are constructed respectively. After connecting, the minimum valley feature matrix is ​​constructed.

[0011] Connect the maximum peak feature matrix and the minimum valley feature matrix to obtain the maximum and minimum peak valley feature matrix;

[0012] (3) The final maximum and minimum peak and valley feature matrix is ​​input into the convolutional neural network, and the classification results of the audio data are output.

[0013] Preferably, in step (1), for an audio data x, the lth slice x l The discrete Fourier transform DFT is expressed as:

[0014]

[0015] Among them, X l [k] is the DFT coefficient of the kth frequency, k = 0…2N f -1,x l [m] is x l The amplitude value at the mth time point, m = 0...2N f -1, 2N f is the slice size, j is the imaginary part of the complex number, l=0…L-1, and L is the number of slices for each audio data.

[0016] Preferably, in step (2), the maximum and minimum peak and valley feature matrix in is the maximum position matrix, is the minimum position matrix, is the maximum amplitude matrix, is the minimum amplitude matrix, and p is the number of preset extreme values.

[0017] As a preference, The lth column Among them, H l is the set of frequency positions of all peaks in the lth slice spectrum, fH l From H l A vector consisting of the first p frequency positions selected after sorting in descending order of peak amplitude; The rth row and lth column of in

[0018] As a preference, The lth column Among them, S l is the set of frequency positions of all spectral valleys in the lth slice spectrum, fS l For S l A vector consisting of the first p frequency positions selected after sorting in ascending order of valley amplitude; The rth row and lth column of in

[0019] Preferably, in step (3), the convolutional neural network model used uses three convolution blocks, each convolution block includes a convolution layer and a pooling layer, and the three convolution blocks are followed by a multi-layer perceptron to further aggregate the features; and a softmax function is used as a classifier.

[0020] Based on the same inventive concept, the present invention provides an audio classification system based on the maximum and minimum peak and valley tracks of the spectrum graph, comprising:

[0021] A spectrogram feature construction module is used to construct a spectrogram feature for each audio: slice the audio data to obtain a plurality of audio data slices, and calculate a spectrogram of each audio data slice;

[0022] The maximum and minimum peak and valley feature matrix construction module is used to calculate the maximum peak feature matrix and the minimum valley feature matrix respectively based on the audio frequency spectrum diagram, and construct the final maximum and minimum peak and valley feature matrix; including: based on the amplitude axis of the frequency spectrum diagram, calculating the position and amplitude size of the maximum value point of the amplitude and sorting them, constructing the maximum position matrix and the maximum amplitude matrix respectively, and connecting them to construct the maximum peak feature matrix; based on the amplitude axis of the frequency spectrum diagram, calculating the position and amplitude size of the minimum value point of the amplitude and sorting them, constructing the minimum position matrix and the minimum amplitude matrix respectively, and connecting them to construct the minimum valley feature matrix; connecting the maximum peak feature matrix and the minimum valley feature matrix to obtain the maximum and minimum peak and valley feature matrix;

[0023] And a classification module is used to input the final maximum and minimum peak and valley feature matrix into the convolutional neural network and output the classification results of the audio data.

[0024] Based on the same inventive concept, the present invention provides a computer system, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the computer program is loaded into the processor, the steps of the audio classification method based on the maximum and minimum peak and valley trajectories of the spectrum graph are implemented.

[0025] Based on the same inventive concept, the present invention provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the steps of the audio classification method based on the maximum and minimum peak and valley trajectories of the spectrum graph are implemented.

[0026] Beneficial effects: The present invention first slices the audio data and calculates the spectrogram features of each slice; based on the maximum peak trajectory of the audio spectrogram, the minimum valley trajectory of the spectrogram is introduced to construct a maximum and minimum peak and valley feature matrix, while paying attention to the maximum peak information and the minimum valley information, so as to better utilize the trajectory information in the spectrogram. The experiment on the data shows that the method has superior performance, which is specifically manifested as follows:

[0027] 1) The method of the present invention is performed on audio slices, and after extracting the features of each slice, it is fused, focusing on the feature details of the audio spectrum.

[0028] 2) The present invention is based on the spectrogram features of audio. The spectrogram has a displayed time domain and frequency domain and has more information on the frequency axis. The maximum and minimum peak and valley feature matrix extracted therefrom can better express the connection between the audio data features.

[0029] 3) Continue to aggregate features by inputting the maximum and minimum peak and valley feature matrix into the neural network. This process makes better use of the relationship between the spectrum information, thus helping to improve the accuracy of classification.

[0030] The advantage of the present invention is that it uses the peak-valley trajectory to describe the characteristic relationship of the audio spectrogram, calculates the position matrix and the amplitude matrix at the same time, constructs the maximum peak matrix and the minimum valley matrix, and finally obtains the final maximum and minimum peak-valley characteristic matrix by connection, so that the relationship between the peak trajectory of the spectrogram and the valley trajectory of the spectrogram is more fully explored; in addition, the present invention aggregates the features of the spectrogram before inputting the model, which can improve the accuracy of classification. Experiments have proved the effectiveness of the method of the present invention, and its classification performance is significantly better than many previous classification algorithms based on spectrograms. BRIEF DESCRIPTION OF THE DRAWINGS

[0031] Figure 1 The figure is a schematic diagram of the method principle of the embodiment of the present invention. The figure illustrates in detail the execution process of the embodiment of the present invention, which consists of constructing audio slices, maximum position matrix, maximum amplitude matrix, minimum position matrix, minimum amplitude matrix, maximum and minimum peak and valley matrix, classification results, etc. DETAILED DESCRIPTION

[0032] The technical solution of the present invention is further described below in conjunction with the accompanying drawings and specific embodiments.

[0033] An audio classification method based on the maximum and minimum peak and valley trajectories of a spectrogram disclosed in an embodiment of the present invention adopts the peak and valley trajectories to describe the characteristic relationship of the audio spectrogram, and calculates the position matrix and the amplitude matrix through peak tracking and valley tracking to generate the final maximum and minimum peak and valley feature matrix. This overcomes the lack of exploration of the connection between the peak trajectory of the spectrogram and the valley trajectory of the spectrogram in previous audio classification algorithms based on spectrograms. The maximum and minimum peak and valley feature matrix is ​​constructed to improve the accuracy of classification. The final classification result can be obtained by inputting the maximum and minimum peak and valley feature matrix into the neural network model.

[0034] like Figure 1 As shown, the specific steps of the embodiment of the present invention are as follows:

[0035] (1) Construct spectrogram features for each audio

[0036] Slice the audio data to obtain multiple audio data slices, and calculate the spectrogram of each audio data slice. For an audio data x, we slice it (slice it by 2 seconds in this example) to obtain L slices of size 2N each. f Slice x l (l=0…L-1). Then the lth slice x l The discrete Fourier transform (DFT) coefficients of the kth frequency are:

[0037]

[0038] Where k = 0…2N f -1, indicating the kth frequency, m = 0...2N f -1, indicates the time point, x l [m] represents x l The amplitude value at the mth time point, j represents the imaginary part of the complex number, is a complex function.

[0039] (2) Based on the audio spectrum, the maximum peak feature matrix and the minimum valley feature matrix are calculated respectively to construct the final maximum and minimum peak valley feature matrix.

[0040] Specifically: calculate the maximum peak trajectory and minimum valley trajectory of the spectrum graph respectively to find the maximum and minimum peak and valley feature matrix First, determine the frequency positions of all spectrum peaks and valleys in the lth frame, arrange the frequency positions of the p highest peaks in the lth frame spectrum in descending order, arrange the frequency positions of the p lowest valleys in the lth frame spectrum in ascending order, and construct the maximum position matrix and the minimum position matrix; similarly, arrange the frequency amplitudes of the p highest peaks in the lth frame spectrum in descending order, and arrange the frequency amplitudes of the p lowest valleys in the lth frame spectrum in ascending order, and construct the maximum amplitude matrix and the minimum amplitude matrix. Finally, connect to obtain the maximum and minimum peak and valley feature matrices.

[0041] The specific process is as follows:

[0042] Based on the x obtained in step (1), we l , determine x l The frequency positions of all spectral peaks in , to construct the following set H l :

[0043] H l ={k:(|X l [k-1]<X l [k]|)∧(|X l [k]>X l [k+1]|}#(2)

[0044] Where 0≤k<(2N f -1),|X l [k]| represents X l [k] size. Similarly, we also determine x l The frequency positions of all spectral valleys in the spectral graph are constructed as follows: l :

[0045] S l ={k:(|X l [k-1]>X l [k]|)∧(|X l [k]<X l [k+1]|}#(3)

[0046] The number of spectral peaks in each slice is different, so not all spectral peaks are important. Therefore, only a fixed number (at most p) of the highest amplitude peaks can be determined from the spectrum of each slice, and the highest spectral peaks in each slice are used to construct the maximum amplitude position set.

[0047]

[0048] where |X l [k (0) ]|≥|X l [k (1) ]|≥…|X l [k (q) ]| and 0<q≤(p-1). Similarly, for the set S l , the minimum spectral valley of each slice is used to construct the minimum amplitude position set

[0049]

[0050] If for any x l, q<(p-1), then H l The highest frequency position (i.e. max(H l )) and S l The lowest frequency position (i.e. min(S l )) will be repeated p-1-q times. When only a small number of the highest amplitude spectrum peaks are considered (p = 10), the effect of this repetition process on the peak (valley) amplitude and position distribution can be ignored. The elements (the frequency positions of the p highest peaks in the spectrum) are sorted in descending order to construct the vector fH l ;

[0051] f l [0]≥fH l [1]≥…≥fH l [p-1]#(6)

[0052] right The elements (the frequency positions of the p lowest valleys in the spectrum) are sorted in ascending order to construct a vector fS l :

[0053] f l [0]≤fS l [1]≤…≤fS l [p-1]#(7)

[0054] in And r=0,…p-1,fH l Included The maximum frequency position of the sorted spectrum peaks; And r=0,…p-1. fS l Included The sorted minimum frequency positions of the mid-spectrum valleys. Vector fH l and fS l (l=0,…L-1) is used to construct the maximum position matrix of the audio slice and the minimum position matrix and The lth column is defined as:

[0055]

[0056]

[0057] Similarly, the maximum amplitude matrix constructed is defined as:

[0058]

[0059] in And l = 0, ... (L-1). Similarly, the minimum amplitude matrix constructed is defined as:

[0060]

[0061] Finally, we get the maximum and minimum peak and valley feature matrix

[0062] (3) The final maximum and minimum peak and valley feature matrix Input the convolutional neural network, use the softmax function as the classifier, use the binary cross entropy as the loss function for training, and output the classification result of the audio data. After three layers of convolution blocks and a multi-layer perceptron, the predicted classification result will be obtained. The loss is calculated with the true label, and the loss is minimized to back-propagate the training model. The objective function of the model is as follows:

[0063]

[0064] Where n represents the number of labels. This embodiment is for a binary classification task (music and non-music), so 2 is taken here. Represents the output after the i-th vector is input into the neural network; y i represents the true label.

[0065] In this example, the convolutional neural network specifically uses three convolutional blocks and a multi-layer perceptron structure. Each convolutional block contains a convolutional layer and a maximum pooling layer, followed by batch normalization to suppress overfitting. The convolution kernel sizes of the convolutional layers of the three convolutional blocks are 5*5, 3*3, and 3*3, respectively, with a step size of 2. All activation functions use the relu function, and the results after convolution are padded. The multi-layer perceptron uses three fully connected layers, with 1024, 5121, and 2 neurons, respectively. The first two fully connected layers are followed by dropout, with the parameter set to 0.5. After the maximum and minimum peak feature matrix passes through the convolutional neural network model, the classification result is obtained, and the model is trained by backpropagation through the loss function.

[0066] As shown in the following table, the table shows the classification performance of the embodiment of the present invention (abbreviated as MMSPT in English) under three music and speech mixed data sets: GTZAN Music / Speech collection, Scheirer-Slaney Music-Speech Corpus, and MUSAN. Dataset Each audio in the GTZAN Music / Speech collection data set has a duration of 30s. In order to ensure the quality of the audio, they are all extracted from radio, CDs and mp3 files. These audios are stored in the form of sampling rate 22050Hz, 16 bits, and mono. The music samples of this data set include the following genres: classical, country, disco, hip-hop, jazz, blues, reggae, pop, metal, etc. Among them, the "classical" category includes: choir, piano, string quartet and other categories; the "jazz" category includes: bigband, cool, fusion, piano, quaret, swing and other categories. The non-music samples of this data set include 3 categories: male voice, female voice, and sports sound. The Scheirer-Slaney Music-SpeechCorpus dataset was obtained by Eric Scheirer and Malcoh Slaney in the San Francisco Bay Area by digitally sampling an FM tuner (22.05kHz sampling rate and 16-bit mono), which includes radio stations, content styles, and uses different degrees of noise formation. The music samples and non-music samples of this dataset are both 20 minutes long, and each sample is 15 seconds long. In the non-music samples, the dataset contains a variety of characters and recording scenes, such as male and female speakers; "recorded in the studio", "on the phone", "quiet environment"; in the music samples, the dataset includes jazz, pop, country, salsa, reggae, classical, non-Western style, new age music, with or without vocals (pure music) and other music types. The MUSAN dataset consists of about 109 hours of audio composed of music, non-music, and noise. The sampling rate of all audio is 16kHz, and these audios are from OpenSLR (a website dedicated to hosting speech resources). The non-music corpus contains about 60 hours of speeches, of which 20 hours and 21 minutes are from Librivox reading speeches. Each of these 20 hours and 21 minutes of audio files is a recording of an entire chapter in a book (by the same speaker); half of these Librivox reading speeches are in English, and the other half contain 11 other languages. The remaining 40 hours and 1 minute of speeches are recordings from hearings, committees, and debates of the US government, and the language of these recordings is English. The music corpus contains 42 hours and 31 minutes of music, which is divided into Western art music (Baroque, Romanticism, Classical, etc.) and popular music (jazz, hip-hop, etc.).This experiment uses a total of 102 hours and 53 minutes of audio data, including music and non-music parts, as the data set. In the experiment, the slice specification of the audio data is 2s; after normalization, the dimension of each slice is 188*86.

[0067] Table 1: Performance (accuracy) of the MMSPT algorithm under different slice specifications of three datasets

[0068]

[0069] Table 2: Performance (accuracy) of MMSPT and other algorithms on three datasets

[0070]

[0071] Table 3: Performance (accuracy) of MMSPT ablation experiments on three datasets

[0072]

[0073] Table 1 shows the performance of MMSPT on three datasets with three different slice sizes. Table 2 shows the results of comparative experiments comparing the performance of four different methods on three datasets. Table 3 shows the results of ablation experiments comparing the performance of using only the maximum peak feature of the spectrogram, the minimum valley feature of the spectrogram, and both the maximum peak feature and the minimum valley feature (MMSPT) on three datasets.

[0074] Based on the same inventive concept, an audio classification system based on the maximum and minimum peak and valley trajectories of a spectrum graph disclosed in an embodiment of the present invention includes: a spectrum graph feature construction module, which is used to construct a spectrum graph feature for each audio: slice the audio data to obtain multiple audio data slices, and calculate the spectrum graph of each audio data slice; a maximum and minimum peak and valley feature matrix construction module, which is used to calculate the maximum peak feature matrix and the minimum valley feature matrix respectively based on the spectrum graph of the audio, and construct the final maximum and minimum peak and valley feature matrix; including: based on the amplitude axis of the spectrum graph, calculating the position and amplitude size of the maximum value points of the amplitude and sorting them, respectively constructing the maximum position matrix and the maximum amplitude matrix, and connecting them to construct the maximum peak feature matrix; based on the amplitude axis of the spectrum graph, calculating the position and amplitude size of the minimum value points of the amplitude and sorting them, respectively constructing the minimum position matrix and the minimum amplitude matrix, and connecting them to construct the minimum valley feature matrix; connecting the maximum peak feature matrix and the minimum valley feature matrix to obtain the maximum and minimum peak and valley feature matrix; and a classification module, which is used to input the final maximum and minimum peak and valley feature matrix into a convolutional neural network and output the classification result of the audio data.

[0075] Those skilled in the art can clearly understand that, for the convenience and simplicity of description, the specific working process of each module described above can refer to the corresponding process in the aforementioned method embodiment, and will not be repeated here. The division of the modules is only a logical function division, and there may be other division methods in actual implementation, for example, multiple modules can be combined or integrated into another system.

[0076] Based on the same inventive concept, an embodiment of the present invention discloses a computer system, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the computer program is loaded into the processor, the steps of the audio classification method based on the maximum and minimum peak and valley trajectories of the spectrum graph are implemented.

[0077] Based on the same inventive concept, an embodiment of the present invention discloses a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the steps of the audio classification method based on the maximum and minimum peak and valley trajectories of the spectrum graph are implemented.

[0078] Those skilled in the art can understand that the technical solution of the present invention, in essence or in other words, the part that contributes to the prior art, can be embodied in the form of a software product, which is stored in a storage medium and includes several instructions for enabling a computer system (which may be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in the embodiment of the present invention. The storage medium includes: various media that can store computer programs, such as a USB flash drive, a mobile hard disk, a read-only memory ROM, a random access memory RAM, a magnetic disk or an optical disk.

Claims

1. An audio classification method based on the maximum and minimum peak and valley tracks of the spectrum graph, characterized in that: The steps include: (1) constructing a spectrogram feature for each audio: slicing the audio data to obtain multiple audio data slices, and calculating the spectrogram of each audio data slice; (2) Based on the audio spectrum, the maximum peak feature matrix and the minimum valley feature matrix are calculated respectively to construct the final maximum and minimum peak valley feature matrix; including: Based on the amplitude axis of the spectrum graph, the position and amplitude of the maximum amplitude point are calculated and sorted, and the maximum position matrix and the maximum amplitude matrix are constructed respectively. After connecting, the maximum peak feature matrix is ​​constructed; Based on the amplitude axis of the spectrum graph, the position and amplitude of the minimum amplitude point are calculated and sorted, and the minimum position matrix and the minimum amplitude matrix are constructed respectively. After connecting, the minimum valley feature matrix is ​​constructed. Connect the maximum peak feature matrix and the minimum valley feature matrix to obtain the maximum and minimum peak valley feature matrix in is the maximum position matrix, is the minimum position matrix, is the maximum amplitude matrix, is the minimum amplitude matrix, p is the number of preset extreme values, and L is the number of slices of each audio data; The lth column l=0…L-1; where H l is the set of frequency positions of all peaks in the lth slice spectrum, fH l From H l A vector consisting of the first p frequency positions selected after sorting in descending order of peak amplitude; The rth row and lth column of in r=0,…(p-1),X l [h] is the DFT coefficient of the hth frequency of the lth slice; The lth column l=0…L-1; where S l is the set of frequency positions of all spectral valleys in the lth slice spectrum, fS l For S l A vector consisting of the first p frequency positions selected after sorting in ascending order of valley amplitude; The rth row and lth column of in r=0,…(p-1); (3) The final maximum and minimum peak and valley feature matrix is ​​input into the convolutional neural network, and the classification results of the audio data are output.

2. The audio classification method based on the maximum and minimum peak and valley trajectories of the spectrum graph according to claim 1 is characterized in that: In step (1), for an audio data x, the lth slice x l The discrete Fourier transform DFT is expressed as: Among them, X l [k] is x l DFT coefficients of the kth frequency, k = 0…2N f -1,x l [m] is x l The amplitude value at the mth time point, m = 0...2N f -1, 2N f is the slice size, j is the imaginary part of the complex number, and l=0…L-1.

3. The audio classification method based on the maximum and minimum peak and valley tracks of the spectrum graph according to claim 1 is characterized in that: In the step (3), the convolutional neural network model used uses three convolution blocks, each of which includes a convolution layer and a pooling layer. The three convolution blocks are followed by a multi-layer perceptron to further aggregate the features; and a softmax function is used as a classifier.

4. An audio classification system based on the maximum and minimum peak and valley tracks of the spectrum graph, characterized in that: include: A spectrogram feature construction module is used to construct a spectrogram feature for each audio: slice the audio data to obtain a plurality of audio data slices, and calculate a spectrogram of each audio data slice; The maximum and minimum peak and valley feature matrix construction module is used to calculate the maximum peak feature matrix and the minimum valley feature matrix based on the audio spectrum diagram to construct the final maximum and minimum peak and valley feature matrix; including: based on the amplitude axis of the spectrum diagram, calculating the position and amplitude size of the maximum amplitude point and sorting them, constructing the maximum position matrix and the maximum amplitude matrix respectively, and connecting them to construct the maximum peak feature matrix; based on the amplitude axis of the spectrum diagram, calculating the position and amplitude size of the minimum amplitude point and sorting them, constructing the minimum position matrix and the minimum amplitude matrix respectively, and connecting them to construct the minimum valley feature matrix; connecting the maximum peak feature matrix and the minimum valley feature matrix to obtain the maximum and minimum peak and valley feature matrix in is the maximum position matrix, is the minimum position matrix, is the maximum amplitude matrix, is the minimum amplitude matrix, p is the number of preset extreme values, and L is the number of slices of each audio data; The lth column l=0…L-1; where H l is the set of frequency positions of all peaks in the lth slice spectrum, fH l From H l A vector consisting of the first p frequency positions selected after sorting in descending order of peak amplitude; The rth row and lth column of in r=0,…(p-1),X l [h] is the DFT coefficient of the hth frequency of the lth slice; The lth column l=0…L-1; where S l is the set of frequency positions of all spectral valleys in the lth slice spectrum, fS l For S l A vector consisting of the first p frequency positions selected after sorting in ascending order of valley amplitude; The rth row and lth column of in r=0,…(p-1); And a classification module is used to input the final maximum and minimum peak and valley feature matrix into the convolutional neural network and output the classification results of the audio data.

5. A computer system comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the computer program is loaded into a processor, the steps of the audio classification method based on the maximum and minimum peak and valley trajectories of the spectrogram are implemented according to any one of claims 1 to 3.

6. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the audio classification method based on the maximum and minimum peak and valley trajectories of the spectrogram are implemented according to any one of claims 1 to 3.

Citation Information

Patent Citations

  • Music audio classification method based on convolutional recurrent neural network

    CN112199548A

  • Audio data classification

    US10424321B1