Audio hash generation method based on deep learning

Through the deep learning audio hash generation method, combined with spectrum analysis, principal component analysis and spatial attention mechanism, the performance instability of existing audio hashing methods under noise interference and frequency changes is solved, and efficient and robust audio retrieval and management are achieved.

CN120296196APending Publication Date: 2025-07-11UNIV OF SHANGHAI FOR SCI & TECH
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510216291.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-26
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

The existing audio hashing methods have unstable performance under conditions such as noise interference, echo reverb, time shift and frequency changes, making it difficult to balance high efficiency and accuracy, especially for the storage and calculation of long audio inputs.

Method used

A deep learning-based audio hash generation method is adopted, and the optimized model is trained through triple comparison, combining spectrum analysis, principal component analysis, spatial attention mechanism and dynamic pooling layer to generate robust audio hash values.

Benefits of technology

While ensuring efficient storage, it significantly improves the stability and accuracy of audio hashing, enhances the ability to distinguish audio, and is suitable for large-scale audio database retrieval and copyright protection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120296196A_ABST
    Figure CN120296196A_ABST
Patent Text Reader

Abstract

The invention discloses an audio hash generation method based on deep learning. The method comprises the following steps: S1, performing spectral analysis on an input audio signal and extracting a frequency cepstrum coefficient; s2, performing principal component analysis dimensionality reduction on the audio feature vector obtained in the above step; s3, weighting each fragment on the spectrum energy dimension through a space attention mechanism to optimize feature expression; s4, aligning the audio features with different lengths with the full connection layer through the dynamic pooling layer; and S5, according to the feature vector after dimension reduction, generating a final audio hash value through a size relationship of adjacent elements. According to the method disclosed by the invention, the method has remarkable robustness on common interferences such as time offset, frequency change and noise increase of the audio, the stability and the accuracy of audio hash are ensured while efficient storage is ensured, and the method is suitable for retrieval and copyright protection of a large-scale audio database.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of audio processing and deep learning, and particularly relates to an audio hash generation method based on deep learning. Background Art

[0002] With the rapid growth of multimedia data, the fields of audio retrieval and copyright protection have received extensive attention in recent years. The existing audio hash methods usually extract frequency domain features and combine traditional machine learning algorithms. However, it is difficult for the existing methods to maintain stable performance under conditions of noise interference, echo reverberation, time shift, and frequency change. In addition, long audio inputs pose higher requirements for storage and computing, and it is difficult for the existing methods to balance efficiency and accuracy. Summary of the Invention

[0003] Aiming at the deficiencies in the existing technology, the purpose of the present invention is to provide an audio hash generation method based on deep learning, which further optimizes the sensitivity of the model to similar audio and the discrimination ability for different audio through triplet contrast training. The method has significant robustness to common interferences such as time shift, frequency change, and added noise of audio, ensures the stability and accuracy of the audio hash while guaranteeing efficient storage, and is applicable to the retrieval and copyright protection of large-scale audio databases. To achieve the above object and other advantages of the present invention, an audio hash generation method based on deep learning is provided, including:

[0004] S1. Perform spectral analysis on the input audio signal and extract mel-frequency cepstral coefficients;

[0005] S2. Perform principal component analysis dimensionality reduction on the audio feature vector obtained in the above step;

[0006] S3. Weight each segment in the spectral energy dimension through a spatial attention mechanism to optimize the feature representation;

[0007] S4. Align audio features of different lengths with a fully connected layer through a dynamic pooling layer;

[0008] S5. Generate a final audio hash value according to the dimensionality-reduced feature vector through the size relationship of adjacent elements.

[0009] Preferably, the spectral analysis of the audio signal mentioned in step S1 is realized by combining the short-time Fourier transform (STFT) with a custom window function to optimize the resolution of feature extraction and generate a spectrogram suitable for the extraction of mel-frequency cepstral coefficients.

[0010] Preferably, in step S2, the audio features are dimensionally reduced by principal component analysis (PCA) to significantly reduce the GPU computational amount of the spatial attention mechanism in step S3 and avoid video memory overflow.

[0011] Preferably, the spatial attention mechanism in step S3 is used to weight the spectral segments to enhance the sensitivity of the features to important audio segments.

[0012] Preferably, the dynamic pooling layer in step S4 adaptively adjusts the pooling scale according to the feature length of the input audio to align audio feature vectors of any length to the fully connected layer, thereby meeting the input requirements of audio samples of different lengths.

[0013] Compared with the prior art, the beneficial effects of the present invention are as follows: compared with the prior art, it overcomes the problems of poor algorithm generalization ability, insufficient robustness against interference, and large computational overhead, and has the following beneficial effects: adopting a dataset retention operation strategy and triplet contrast learning to optimize the model, it has strong robustness against time shift, frequency adjustment, and noise interference of audio; combining dynamic pooling and spatial attention mechanism, it effectively aligns audio features of different lengths and enhances the recognition ability of key segments; the multiple dimensionality reduction mechanism reduces the computational overhead, is friendly to small computing power devices, and further improves the universality of the retrieval system; by generating a compact binary hash representation based on the size relationship of adjacent elements, it significantly reduces the storage cost and improves the retrieval efficiency.

[0014] In summary, the present invention realizes the perfect combination of high efficiency, low storage requirements, strong robustness, and high precision in the field of audio hash generation, can significantly improve the overall performance of the audio retrieval and management system, and provides an effective solution for the intelligent processing of audio data. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] Figure 1 is a flowchart of a method for generating an audio hash based on deep learning according to the present invention;

[0016] Figure 2 is a spectral energy feature analysis diagram of an embodiment of a method for generating an audio hash based on deep learning according to the present invention;

[0017] Figure 3 is a block diagram of the AudioHashNet network structure of an embodiment of a method for generating an audio hash based on deep learning according to the present invention;

[0018] Figure 4 is a t-SNE dimensionality reduction and visual hash effect diagram of a method for generating an audio hash based on deep learning according to the present invention;

[0019] Figure 5 is an experimental result diagram of the influence of PCA and attention mechanism on the model performance of a method for generating an audio hash based on deep learning according to the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0020] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0021] Referring to Figure 1 , a method for generating audio hashes based on deep learning, comprising the following steps:

[0022] S1. Perform spectral analysis on the input audio signal and extract mel-frequency cepstral coefficients; in this step, first, all input audio signals are resampled to a unified sampling rate (16000 Hz) to ensure the consistency of the sampling rate of audio from different sources. Then, perform short-time Fourier transform on the audio signal to convert the signal from the time domain to the frequency domain.

[0023] This process frames the audio signal through a window function and calculates the spectral information of each frame to obtain the frequency-domain representation of the signal. The frequency-domain information is mapped to the mel-scale through a filter bank. The mel-scale is a frequency scale that simulates the way the human ear perceives different frequencies and has non-linear characteristics, which can better represent the low-frequency and high-frequency information of the audio signal. The number of filter banks is set to 64, and the window function size is set to 1024 to balance the spectral resolution and the computational complexity of AudioHashNet. The frequency spectrogram generated in this process is a two-dimensional matrix, where each column represents the spectrum of a time frame, and the number of rows represents the number of filters.

[0024] The short-time Fourier transform (STFT) is combined with a custom window function to optimize the resolution of feature extraction and generate a spectrogram suitable for the extraction of mel-frequency cepstral coefficients, as Figure 2 shown. Its formula is:

[0025]

[0026] In the formula: X(t,f) represents the spectral feature at time t and frequency f, n is the discrete time index, t is the index in time, representing the moment of signal analysis, x[n] is the time-domain signal, and w[n - t] is the window function. The spectrogram obtained through STFT is used to generate mel-frequency cepstral coefficient features.

[0027] S2. Perform principal component analysis on the audio feature vectors obtained in the above steps for dimensionality reduction; after the frequency spectrogram is extracted, high-dimensional features are generated, usually including multiple frequency channels (such as 64 channels) and time frames. However, there may be redundant information in these features. For example, some channels change little between different audios, and the features of some channels contribute less to the judgment of audio similarity. This step can effectively extract the most discriminative principal components in the spectrogram, compress the dimensionality of the feature channels, thereby reducing the computational cost, while retaining the key feature information for hash generation. In particular, the principal components in the low-frequency band correspond to the key content in the audio signal, such as speech or instrument characteristics, and are of higher importance in the hash generation process. Specifically, the present invention reduces the number of feature channels from 64 to 32, and the dimensionality-reduced features provide efficient input for the subsequent attention mechanism and hash generation, and also improve the computational efficiency of the fully connected layer, significantly shortening the training and inference time of the model. The specific steps are as follows:

[0028] (1) Construction of the feature matrix

[0029] The input frequency spectrogram is a four-dimensional tensor X with dimensions (B, C, H, W), where:

[0030] · B represents the batch size;

[0031] · C represents the number of channels (such as 64);

[0032] · H and W are the time frame and frequency dimensions respectively.

[0033] (2) Calculation of the covariance matrix

[0034] To identify the correlation between feature channels, calculate the covariance matrix ∑ of the feature matrix X:

[0035]

[0036] where:

[0037] · n = B × H × W is the total number of feature samples;

[0038] · x i is the i-th row vector in the feature matrix;

[0039] · μ is the mean vector of all feature samples.

[0040] The covariance matrix Σ describes the relationship between feature channels, and the magnitude of the eigenvalues reflects the variance of information in each direction.

[0041] (3) Extraction and dimensionality reduction of the principal components

[0042] Select the top k principal component directions with the largest eigenvalues and construct the dimensionality reduction matrix W k :

[0043] W k = [q1, q2, …, q k 。

[0044] Project the original features onto the dimensionality-reduced space:

[0045] X' = XW k

[0046] The channel dimension of the feature matrix X' after dimensionality reduction of matrix X is 32.

[0047] S3. Weight each segment in the spectral energy dimension through a spatial attention mechanism to optimize the feature representation; the characteristics of the audio signal energy distribution on the frequency spectrogram are concentrated and different. For example, in some frequency bands, such as the formant region of speech or the specific frequency range of musical instruments, they have higher value for distinguishing audio content. To highlight the characteristics of these key regions, the present invention introduces a spatial attention mechanism, which dynamically weights different parts of the frequency spectrogram by generating an attention weight map, thereby enhancing the model's ability to focus on important features.

[0048] To generate the attention weight map, the present invention combines the global information of average pooling and max pooling, calculates the global information of average pooling and max pooling respectively for the frequency spectrogram features after PCA dimensionality reduction, and concatenates them into a feature map A. Subsequently, the feature map is input into a convolutional layer with a convolutional kernel size of k×k to generate the final attention weight matrix A. In particular, the present invention sets k to 7 to balance the receptive field and the amount of calculation. The spatial attention mechanism in step S3 is used to weight the spectral segments and enhance the sensitivity of the features to important audio segments. The attention weight matrix A is calculated by the following formula:

[0049] A = spftmax(W f F + b)

[0050] where: W f is the weight matrix, F is the spectral feature matrix, b is the bias term, the weight distribution of the feature vector in the spectral energy dimension is adjusted through the attention mechanism, and the softmax function is used to convert the output into weight values to ensure that the sum of all weights is 1..

[0051] S4. Align audio features of different lengths with the fully connected layer through a dynamic pooling layer; the dynamic pooling layer adaptively adjusts the pooling scale according to the feature length of the input audio to align audio feature vectors of any length to the fully connected layer, so as to meet the input requirements of audio samples of different lengths. After the matrix A of the spatial attention mechanism, the time and frequency matrix sizes are (H, W), and the set target size after pooling is (H t , W t)。Through the adaptive average pooling operation, the feature map dynamically adjusts the pooling window size and stride in both the time and frequency dimensions, and is proportionally compressed to the target size.

[0052] S5. Generate the final audio hash value based on the size relationship between adjacent elements of the feature vector after dimensionality reduction. In this step, to convert the dimensionality-reduced audio feature vector into a binary hash representation with compactness and robustness, a hash generation method based on the size relationship between adjacent elements is adopted. This method can directly map continuous numerical features into binary codes while retaining the relative sorting information of the input features, enhancing the ability to distinguish audio similarities.

[0053] First, after the audio input passes through the convolutional layer, PCA dimensionality reduction layer, spatial attention mechanism, and dynamic pooling, it enters the fully connected layer to generate a feature vector H = [h1, h2, …, h n , where h is the vector element. In particular, n is set to 128 in this network. This feature vector has been optimized to have a strong discriminative ability semantic representation during the network training process. To convert this feature vector into a compact and robust binary hash representation, compare the size relationship between every two adjacent elements h i and h i+1 : If h i > h i+1 , then the corresponding hash bit is set to 1, otherwise it is set to 0; the hash generation rule is:

[0054]

[0055] Through this comparison operation, the generated hash value has the following characteristics:

[0056] (1) Compactness: The feature vector is directly mapped into a binary code through adjacent element comparison, significantly reducing the storage requirement.

[0057] (2) Robustness: The size relationship between adjacent elements can resist small perturbations of the input features to a certain extent, such as noise and time shift.

[0058] (3) Semantic consistency: The feature sorting information retains the relative distribution of the input features and can effectively distinguish different semantic categories of audio.

[0059] A neural network model for generating audio hash, characterized by including:

[0060] An analysis and extraction module for performing spectral analysis on the input audio signal and extracting mel-frequency cepstral coefficients;

[0061] A dimensionality reduction module for performing principal component analysis dimensionality reduction on the extracted audio feature vector;

[0062] The weight distribution adjustment module is used to weight each segment in the spectral energy dimension through a spatial attention mechanism to optimize the feature representation;

[0063] The hash generation module is used to generate a final audio hash value according to the dimensionality-reduced feature vector by the size relationship of adjacent elements; the hash generation module constructs the audio hash value in the following way:

[0064] For the dimensionality-reduced feature vector H = [h1, h2, …, h n , compare the size relationship of every two adjacent elements h i and h i+1 : if h i > h i+1 , the corresponding hash bit is set to 1; otherwise it is set to 0;

[0065] The hash generation rule is:

[0066]

[0067] The binary hash representation generated by this method has compactness and robustness to audio transformations.

[0068] The dataset generation module is used to generate a variety of audio perturbation samples based on the dataset retention operation strategy; including time shift, frequency adjustment and noise interference to enhance the robustness of the model to audio perturbations.

[0069] To improve the robustness and generalization ability of the model, the dataset generation module is designed. This module generates variant audio samples with diverse features by performing a variety of perturbation operations on the original audio samples to simulate different actual scenarios. The generated diverse dataset not only retains the semantic information of the original audio, but also enhances the model's adaptability to common interferences such as noise, time shift, and frequency change.

[0070] To ensure the diversity of the dataset, each operation can be applied alone or superimposed in pairs or in multiple combinations. For example, volume adjustment and frequency adjustment are applied simultaneously to simulate more complex scenario changes. The data augmentation strategies include but are not limited to:

[0071] (1) Time shift

[0072] Shift the entire audio signal forward or backward by a fixed duration (such as 2S) to simulate the time alignment deviation of the audio in different recording devices or playback delay scenarios.

[0073] (2) Frequency adjustment

[0074] Generate high or low pitch variants by changing the sampling rate, for example, adjusting to 90% or 110% of the original sampling rate to simulate the frequency distortion effect of recording or playback devices.

[0075] (3) Speed adjustment

[0076] By stretching or compressing the audio signal in the time domain, the playback speed change of the device is simulated. Specifically, the time-domain stretching (slowing down the speed) or compression (speeding up the speed) of the signal will change the rhythm and pitch of the audio, thereby enhancing the model's adaptability to audio speed changes. For example, the audio speed can vary between 0.8 times and 1.2 times, generating different speaking speeds or music speeds.

[0077] (4) Volume adjustment

[0078] By amplifying or reducing the loudness of the audio signal, the differences in the volume settings of the playback device or the sensitivity of the recording device are simulated. For example, increasing or decreasing the volume (such as ±20%).

[0079] (5) Noise injection

[0080] By adding Gaussian white noise to the audio signal, the signal-to-noise ratio (SNR) range is set to 10 - 30 dB, which is used to enhance the model's robustness to environmental noise interference.

[0081] (6) Filtering operation

[0082] By simulating the differences in the frequency response of the device through high-pass, low-pass, and band-pass filters, for example:

[0083] · High-pass filter: Removes low-frequency noise (such as background hum).

[0084] · Low-pass filter: Weakens high-frequency components (such as sharp noise).

[0085] · Band-pass filter: Highlights the signal within a specific frequency range and enhances specific semantic content.

[0086] (7) Echo effect

[0087] It is achieved by adding a delayed signal and a reverberation component to the original audio signal. For example, a delayed signal can be added to the later part of the audio, and the delay time (such as 100 ms to 500 ms) and the reverberation intensity can be adjusted to simulate the echo effects of different spaces. The echo effect usually causes changes in the time characteristics of the audio, making the audio feel more distant or blurred.

[0088] The triplet training module is used to optimize the model through the triplet contrast learning strategy, perform similarity optimization on the hash representation generated from the input audio samples, and make the hash distances of similar audios smaller and the hash distances of different audios larger. To improve the performance of the audio hashing model, the present invention adopts the triplet contrast learning strategy to optimize the hash representation generated by the model. It makes the hash distances of similar audios smaller and the hash distances of different audios larger by comparing the hash distances of similar samples and dissimilar samples. This method can significantly improve the discrimination ability of the model for audio samples and enhance the accuracy of the hash representation in retrieval and matching tasks.

[0089] (1) Sample preparation

[0090] After the content retention operation in step S6, given an anchor audio sample (original audio), m positive samples (similar to the anchor audio), and m negative samples (different from the anchor audio), they jointly form a set of 2m + 1 contrast training sets.

[0091] (2) Feature extraction and hash generation

[0092] Extract the feature vectors of the anchor sample, positive samples, and negative samples through the trained deep neural network (AudioHashNet). The output feature of each audio sample is a 128-dimensional vector, and a fixed-length hash representation (such as a 128-bit binary hash value) is generated through the fully connected layer. For each triplet sample pair, we calculate its hash representation and measure the similarity through the Hamming distance.

[0093] Furthermore, the neural network model AudioHashNet for audio hash generation has the following features:

[0094]

[0095]

[0096] (3) Loss function calculation

[0097] By optimizing the loss function, it forces the hash values of similar samples to be closer and the hash values of dissimilar samples to be farther apart.

[0098] Furthermore, the triplet contrast loss function L triplet Optimizes the model, specifically as follows:

[0099]

[0100] In the formula: h i is the hash representation of the target audio (Anchor), h i + and h i -They are the positive samples similar to the target audio and the negative samples not similar respectively. d represents the Hamming distance, and α represents the preset hyperparameter boundary distance. By minimizing the loss, it is ensured that the distance between similar audios is smaller and the distance between different audios is larger. Further, the hyperparameter α is set to 5.

[0101] (4) Training and optimization

[0102] By calculating the triplet loss, it is optimized using the Adam gradient descent algorithm, which is set to 0.001; the batch size is set to 32, and the number of training epochs is set to 1000 to update the network parameters in this way. The model selects triplets from the dataset for training in each cycle. During the training process, the network gradually adjusts its weights to make the hash value distances of similar samples smaller and the hash value distances of dissimilar samples larger.

[0103] Finally, the obtained hash values after training can be used in application scenarios such as audio retrieval, audio deduplication, and copyright protection.

[0104] To verify the technical effect of this method, the following experimental analysis is carried out:

[0105] The experimental environment is a PC with an Intel(R) Core(TM) i9-10900X CPU, 32GB of DDR4 3200 memory, and an NVIDIA GeForce RTX 3090. The experimental data comes from a self-built dataset, which has 1000 original audio resources and 16000 transformed audios generated through the S6 content retention operation.

[0106] The experimental content includes t-SNE visualization analysis, Hamming distance distribution histogram analysis, and average Hamming distance comparison and retrieval performance evaluation.

[0107] 1. t-SNE visualization analysis

[0108] t-SNE is a dimensionality reduction and visualization algorithm, which is often used to map high-dimensional data to a low-dimensional space for visual display. To visually show the distribution of audio samples in the hash space, the audio hash model of the present invention is used to extract the 128-bit hash of six randomly selected original audio samples and transformed samples, and t-SNE dimensionality reduction analysis is carried out, as Figure 4 shown.

[0109] The experimental results show that audio samples of different categories present an obvious clustering structure on the two-dimensional plane. Audio samples of the same category are closely clustered in the hash space, and there is a large spatial separation between audio samples of different categories. The results show that the audio hash model of the present invention can effectively distinguish different types of audios and their content retention operation samples, and achieve robust hash coding and retrieval.

[0110] 2. Hamming Distance Distribution Histogram Analysis

[0111] The Hamming distance distribution histogram can visually display the distribution of audio samples in the hash space and quantify the audio retrieval and matching performance of the model. To further verify the discrimination performance and matching accuracy of audio samples in the hash space, the Hamming distance distributions of 499,500 pairs among 1,000 different audio samples and the Hamming distance distribution of 120,000 pairs of similar samples were statistically analyzed.

[0112] The experimental results are as Figure 5 shown. The Hamming distances of similar audio sample pairs are concentrated between 0 and 20 bits, while the Hamming distances of different audio sample pairs are mainly concentrated between 20 and 75 bits. The histogram presents a bimodal distribution structure, indicating that the model can accurately distinguish different audio samples. In the database matching task, the hash representation has high reliability and accuracy.

[0113] 3. Ablation Analysis of PCA and Attention Mechanism

[0114] The average Hamming distance is a statistical metric for measuring the overall similarity between two sets of audio hash representations. By randomly selecting a sample from the original audio database and calculating the average Hamming distance between it and each sample in the database, the discrimination ability of the model for different samples is quantified. By calculating the average Hamming distance between it and each sample in the audio perturbation set, the robustness and applicability of the model for similar samples are quantified.

[0115] To verify the effectiveness of PCA dimensionality reduction and the attention mechanism in the audio hash model, a comparative experiment was conducted. The audio retrieval performance of two model configurations, namely without enabling PCA and the attention mechanism and with enabling PCA and the attention mechanism, was tested respectively, and the retrieval results were quantitatively analyzed through the average Hamming distance.

[0116] The experimental results are as Figure 5 shown. In the 100-song dataset, the impact of the enabled model on the hash performance is significantly improved. The average Hamming distance of negative sample pairs increases to 44.6, much higher than 35.7 when the mechanism is not enabled, indicating that the model's discrimination ability for different categories of audio samples is enhanced. The maximum Hamming distance of positive samples decreases from 9 to 5, and the minimum Hamming distance of negative samples increases from 13 to 19, which further verifies the improvement of the model's performance discrimination.

[0117] The experimental results show that the present invention has significant audio discrimination ability and high retrieval stability, performs well in the audio retrieval task, and meets the retrieval and copyright protection requirements in large-scale audio databases.

[0118] The number of devices and the processing scale described here are used to simplify the description of the present invention. Applications, modifications, and variations of the present invention will be apparent to those skilled in the art.

[0119] Although the embodiments of the present invention have been disclosed as above, they are not limited to the applications listed in the specification and the embodiments. It can be fully applied to various fields suitable for the present invention. For those familiar with the art, additional modifications can be easily achieved. Therefore, without departing from the general concept defined by the claims and the equivalent scope, the present invention is not limited to specific details and the illustrated examples described herein.

Claims

1. An audio hash generation method based on deep learning, characterized in that It includes the following steps: S1. Perform spectral analysis on the input audio signal and extract mel-frequency cepstral coefficients; S2. Perform principal component analysis dimensionality reduction on the audio feature vectors obtained in the above steps; S3. Weight each segment in the spectral energy dimension through a spatial attention mechanism to optimize the feature representation; S4. Align audio features of different lengths with the fully connected layer through a dynamic pooling layer; S5. Generate the final audio hash value based on the dimensionality-reduced feature vectors through the size relationship of adjacent elements.

2. The audio hash generation method based on deep learning according to claim 1, characterized in that The specific steps of step S1 include the following steps: All input audio signals are resampled to a unified sampling rate, the audio signal is subjected to a short-time Fourier transform to convert the signal from the time domain to the frequency domain; By combining the short-time Fourier transform with a custom window function to optimize the resolution of feature extraction, and generating a spectrogram suitable for the extraction of mel-frequency cepstral coefficients, and using the spectrogram to generate mel-frequency cepstral coefficient features.

3. The audio hash generation method based on deep learning according to claim 1, characterized in that, The specific steps of step S2 include the following steps: S21. Construct a feature matrix; S22. Calculate the covariance matrix of the feature matrix; S23. Perform the extraction and dimensionality reduction of the principal components, specifically: Select the top k principal component directions with the largest eigenvalues, construct a dimensionality reduction matrix; Project the original features into the dimensionality reduction space.

4. The audio hash generation method based on deep learning according to claim 1, characterized in that The specific steps of step S3 include the following steps: S31. For the frequency spectrogram features after PCA dimensionality reduction, combine the global information of average pooling and max pooling, and calculate the global information of average pooling and max pooling respectively; S32. Concatenate them into a feature map; S33. Input the feature map into a convolutional layer with a convolutional kernel size of k×k to generate the final attention weight matrix.

5. The method for generating audio hash based on deep learning according to claim 1, wherein, The specific steps of step S4 include the following steps: S41. The dynamic pooling layer adaptively adjusts the pooling scale according to the feature length of the input audio to align audio feature vectors of any length to the fully connected layer; S42. Set the time and frequency matrix sizes through the attention weight matrix, and then obtain the target size after pooling; S43. Through the adaptive average pooling operation, the feature map dynamically adjusts the pooling window size and stride in both the time and frequency dimensions, and is proportionally compressed to the target size.

6. A neural network model for generating audio hashes, characterized in that, It includes: An analysis and extraction module for performing spectral analysis on the input audio signal and extracting mel-frequency cepstral coefficients; A dimensionality reduction module for performing principal component analysis dimensionality reduction on the extracted audio feature vectors; An adjustment weight distribution module for weighting each segment in the spectral energy dimension through a spatial attention mechanism to optimize the feature representation; A hash generation module for generating the final audio hash value based on the dimensionality-reduced feature vectors through the size relationship of adjacent elements; A dataset generation module for generating various audio perturbation samples based on the dataset retention operation strategy; A triplet training module for optimizing the model through a triplet contrast learning strategy, performing similarity optimization on the hash representations generated from the input audio samples, making the hash distances of similar audios smaller and the hash distances of different audios larger.

Citation Information

Cited By

  • Sound effect quality evaluation method and device, equipment and storage medium

    CN121459848A