Audio feature extraction method and device based on neural network
Through multi-level acoustic feature extraction and dynamic time feature analysis based on neural network, the problem of poor accuracy and reliability of vehicle speech detection in complex noise environments is solved, and multi-scale characterization and noise suppression of vehicle speech signals are realized, which significantly improves detection performance.
Patent Information
- Application Number
- CN202510173284.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-17
- Publication Date
- 2025-05-13
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Vehicle voice detection has poor accuracy and reliability in complex noise environments, and it is difficult for the prior art to effectively deal with voice detection problems under different noise conditions.
The audio feature extraction method based on neural network is adopted, and deep time frequency features are constructed through multi-level acoustic feature extraction and dynamic time feature analysis, combined with principal component analysis, dynamic time feature extraction network and bilayer neural network classifier to realize multi-scale characterization and noise suppression of vehicle-mounted voice signals.
It significantly enhances the system's noise resistance, improves the accuracy and reliability of vehicle-mounted voice detection, and can effectively deal with voice detection problems under different noise conditions.
Smart Images

Figure CN119993193A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of audio processing, and in particular to a method and device for extracting audio features based on a neural network. Background Art
[0002] With the continuous development of intelligent driving technology, in-vehicle voice interaction systems have become an important way of human-computer interaction, among which in-vehicle voice detection is a key link in voice interaction systems. However, there are many complex noise interferences in the in-vehicle environment, such as engine noise, road noise, wind noise, etc., which seriously affect the accuracy and reliability of voice detection.
[0003] Traditional speech detection methods mainly rely on manually designed acoustic features, such as short-time energy and zero-crossing rate. These features have poor robustness in complex noise environments and limited expressiveness, making it difficult to accurately characterize the time-frequency characteristics of speech signals in vehicle environments. At the same time, existing feature extraction methods often use a single feature representation method, which cannot fully utilize the information of speech signals at different time scales, resulting in limited detection performance. In addition, the type and intensity of noise in the vehicle environment are dynamically changing. Existing speech detection methods lack the ability to adapt to the noise environment and are difficult to effectively handle speech detection problems under different noise conditions. Summary of the invention
[0004] The main purpose of the present invention is to provide an audio feature extraction method and device based on a neural network. The present invention effectively suppresses various noise interferences in the vehicle environment and enhances the anti-noise ability of the system.
[0005] To achieve the above object, the present invention provides an audio feature extraction method based on a neural network, comprising the following steps: The original audio signal collected by the vehicle microphone is processed by frame division and fast Fourier transform to obtain a frequency domain feature sequence; Extracting short-time energy, zero-crossing rate, spectrum centroid and Mel-frequency cepstrum coefficients from the frequency domain feature sequence, and combining the first-order difference and second-order difference features of each feature to obtain a multi-dimensional acoustic feature point sequence; Arranging the multidimensional acoustic feature point sequence in time sequence to construct an original feature matrix, and performing principal component analysis on the original feature matrix to obtain a feature description matrix; Inputting the feature description matrix into a dynamic time feature extraction network to perform deep time-frequency feature analysis to obtain deep time-frequency features of the vehicle-borne speech; Speech detection and classification are performed based on the deep time-frequency features of the in-vehicle speech, and the in-vehicle speech detection result is output.
[0006] The present invention also provides an audio feature extraction device based on a neural network, comprising: A transformation module is used to perform frame processing and fast Fourier transformation on the original audio signal collected by the vehicle microphone to obtain a frequency domain feature sequence; An extraction module is used to extract short-time energy, zero-crossing rate, spectrum centroid and Mel-frequency cepstrum coefficients from the frequency domain feature sequence, and combine the first-order difference and second-order difference features of each feature to obtain a multi-dimensional acoustic feature point sequence; A construction module, used to arrange the multi-dimensional acoustic feature point sequence in time sequence, construct an original feature matrix, and perform principal component analysis on the original feature matrix to obtain a feature description matrix; An analysis module, used for inputting the feature description matrix into a dynamic time feature extraction network to perform deep time-frequency feature analysis to obtain deep time-frequency features of the vehicle-mounted speech; The classification module is used to perform speech detection and classification based on the deep time-frequency features of the in-vehicle speech and output the in-vehicle speech detection result.
[0007] The present invention also provides a computer device, comprising a memory and a processor, wherein the memory stores a computer program, and the processor implements the steps of any one of the above methods when executing the computer program.
[0008] The present invention also provides a computer-readable storage medium on which a computer program is stored. When the computer program is executed by a processor, the steps of any of the above-mentioned methods are implemented.
[0009] In summary, the technical solution provided by the present invention realizes the multi-scale characterization of the time-frequency characteristics of the vehicle-mounted speech signal through multi-level acoustic feature extraction and dynamic time feature analysis, thereby enhancing the expression ability of the features; principal component analysis is used to reduce the dimension and optimize the features, thereby reducing the feature redundancy and improving the discrimination performance of the features; a dynamic time feature extraction network is designed, and the ability to extract features of different time scales is enhanced through multi-branch parallel convolution and attention mechanism; hierarchical clustering and adaptive filtering strategies are introduced to effectively suppress various noise interferences in the vehicle-mounted environment and enhance the anti-noise ability of the system; a two-layer neural network classifier based on feature decoupling and mutual information maximization is proposed, thereby realizing the effective utilization of features of different levels of speech signals; through variational autoencoders and Bayesian posterior inference, the reliability of the detection results is improved, so that the system has better generalization ability. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] Figure 1 1 is a schematic diagram of the steps of an audio feature extraction method based on a neural network in one embodiment of the present invention; Figure 2 It is a structural block diagram of an audio feature extraction device based on a neural network in one embodiment of the present invention.
[0011] The realization of the purpose, functional features and advantages of the present invention will be further explained in conjunction with embodiments and with reference to the accompanying drawings. DETAILED DESCRIPTION
[0012] In order to make the purpose, technical solution and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.
[0013] Reference Figure 1 , this embodiment provides an audio feature extraction method based on a neural network, comprising the following steps: S1, performs frame processing and fast Fourier transform on the original audio signal collected by the vehicle microphone to obtain a frequency domain feature sequence; Among them, the original audio signal collected by the vehicle microphone is framed and divided into multiple short-time frames. These frames have a certain length (such as 20 milliseconds or 25 milliseconds), and in order to avoid losing key information, there is a certain degree of overlap between the frames. Through framing, the audio signal is regarded as a signal segment with stable characteristics in a short time range. After the framing is completed, the first audio frame sequence is input into the first-order high-pass filter for pre-emphasis processing to enhance the high-frequency part of the audio signal, compensate for the loss of high-frequency components by the recording equipment, and suppress the background noise of the low-frequency part. The pre-emphasis operation is achieved through the following filters: , The specific value of is adjusted according to the actual application scenario. The audio signal after pre-emphasis processing forms a second audio frame sequence. The second audio frame sequence is processed by a window function to reduce the influence of the boundary effect caused by the frame operation on the spectrum estimation. The window function includes a Hamming window, a Hanning window or a Heiman window. The window function effectively smoothes the boundary by applying a weighted operation to each frame signal, thereby improving the accuracy of the spectrum calculation and generating a third audio frame sequence. The third audio frame sequence is subjected to a fast Fourier transform to convert the time domain signal into a frequency domain signal to obtain an initial spectrum sequence. The amplitude spectrum is calculated for the initial spectrum sequence to extract the energy distribution characteristics of the signal. The amplitude spectrum is normalized. By dividing each frequency component of the amplitude spectrum by its maximum value, the range of the eigenvalue is ensured to be between [0,1], the influence of the absolute value difference is eliminated, and the robustness of the algorithm is enhanced. The normalized spectrum sequence is input into a triangular filter bank for sub-band decomposition. The Mel filter is a typical triangular filter bank, which simulates the ability of the human ear to distinguish different frequencies and maps the spectrum energy on the frequency axis to the Mel scale. The design characteristics of the Mel filter are high low-frequency resolution and low high-frequency resolution, which conforms to the auditory characteristics of the human ear. Through the filter bank processing, the original spectrum is divided into several sub-bands to obtain a sub-band energy sequence. The sub-band energy sequence is logarithmically compressed. By taking the logarithmic value of each sub-band energy component, the difference between high-energy and low-energy components is effectively reduced, so that the sensitivity of subsequent processing to energy changes is balanced, and a logarithmic energy sequence is obtained. The logarithmic energy sequence is subjected to a P-order discrete cosine transform. The correlation of the frequency domain feature sequence is minimized through the transformation, and more representative low-order features are extracted to obtain a frequency domain feature sequence, where the typical value of P is 13.
[0014] S2, extracting short-time energy, zero-crossing rate, spectrum centroid and Mel-frequency cepstrum coefficients from the frequency domain feature sequence, and combining the first-order difference and second-order difference features of each feature to obtain a multi-dimensional acoustic feature point sequence; Specifically, the short-time energy of each frame signal is calculated for the frequency domain feature sequence to measure the energy intensity of the signal in each frame. The short-time energy is obtained by accumulating the square of the amplitude of each frame signal and normalizing it. It reflects the distribution of signal energy in time and can help distinguish the voiced and silent segments of speech, thereby enhancing the effectiveness of the signal. The number of symbol changes of adjacent sampling points is calculated for the frequency domain feature sequence, and the number of changes is divided by 2 times the frame length to obtain the zero-crossing rate sequence. The zero-crossing rate is an important feature that reflects the change of frequency components in the signal. By counting the number of symbol changes in each frame signal and normalizing it, the zero-crossing rate can reflect the density of high-frequency components, especially in distinguishing speech signals from background noise. In order to capture the frequency characteristics of the signal, the spectral centroid is calculated. This feature describes the central position of the distribution of spectral energy. The height of the spectral centroid can reflect the dominance of high-frequency or low-frequency components in the signal. The frequency domain feature sequence is input into the Mel filter bank for filtering to simulate the perception characteristics of the human ear to different frequencies. The filtering result is processed by logarithmic operation and discrete cosine transform to generate a Mel frequency cepstrum coefficient sequence. Mel-frequency cepstral coefficients are a feature used for speech signal analysis. They can effectively compress spectral information and retain important features of speech. They have extremely high recognition capabilities in tasks such as speech recognition. The short-time energy sequence, zero-crossing rate sequence, spectral centroid sequence, and Mel-frequency cepstral coefficient sequence are combined to obtain a basic feature sequence. In order to capture the temporal trend of the signal, the first-order difference is calculated for the basic feature sequence to obtain the rate of change of the signal feature. The first-order difference reflects the speed of the change of the signal feature over time, thereby providing dynamic characteristic information. On this basis, the first-order difference sequence is calculated again to obtain the second-order difference feature, which describes the acceleration information of the signal change and reflects the complex pattern of the dynamic change of the signal. The basic feature sequence, the first-order difference feature sequence, and the second-order difference feature sequence are feature spliced to generate the final multi-dimensional acoustic feature point sequence. This sequence integrates static and dynamic features to describe the time-varying characteristics and spectral distribution of the signal.
[0015] S3, arranging the multi-dimensional acoustic feature point sequence in time sequence, constructing an original feature matrix, and performing principal component analysis on the original feature matrix to obtain a feature description matrix; It should be noted that the feature vector is constructed for the multidimensional acoustic feature point sequence according to the time series relationship. Each feature point represents the acoustic feature of the audio signal at a certain moment, and a feature group is formed by combining M adjacent feature points together. These feature groups are arranged in time order to obtain a preliminary original feature sequence, which effectively captures the local correlation of the signal in the time dimension. The original feature sequence is rearranged into a matrix form, namely the original feature matrix. By setting the number of rows of the matrix to N, representing the total number of feature points, and the number of columns to D, representing the dimension of each feature point, an N×D-dimensional first feature matrix is obtained to describe the time series characteristics of the audio signal and the distribution of each acoustic feature. The first feature matrix is standardized. The mean of each column of the original feature matrix is subtracted from the column to eliminate the mean offset of different features; the result is divided by the standard deviation of the column to eliminate the scale difference between the feature values. After this step, the obtained standardized feature matrix has zero mean and unit standard deviation, ensuring that features of different dimensions are compared and analyzed at the same scale. After the standardization is completed, in order to extract the main change pattern of the feature, the covariance matrix is calculated for the standardized feature matrix. The covariance matrix is a D×D dimensional square matrix that describes the correlation between each feature dimension. Through this calculation, the linear relationship between each feature and other features is quantified. Solve the characteristic equation of the covariance matrix to obtain D eigenvalues and their corresponding D eigenvectors. The eigenvalue represents the contribution of each principal component to the overall variance, while the eigenvector describes the direction of each principal component. Calculate the cumulative contribution rate of the eigenvalue. The cumulative contribution rate is the cumulative proportion of the first few values of the eigenvalues to the total eigenvalues after the eigenvalues are arranged in descending order. According to the set target cumulative contribution rate (such as 95%), select the first K eigenvalues and their corresponding eigenvectors required when the cumulative contribution rate reaches the target value. These eigenvectors form a feature transformation matrix, which is used to map the original high-dimensional feature space to a lower-dimensional subspace. Multiply the standardized feature matrix with the above feature transformation matrix to obtain a K-dimensional reduced-dimensional feature matrix. Normalize the reduced-dimensional feature matrix. By mapping each element in the matrix to the range of [0,1], the amplitude differences between different features are eliminated, ensuring that all features are in the same numerical range, and obtaining a feature description matrix.
[0016] S4, inputting the feature description matrix into the dynamic time feature extraction network to perform deep time-frequency feature analysis to obtain the deep time-frequency features of the vehicle speech; Specifically, the feature description matrix is input into the multi-branch parallel convolution layer of the dynamic temporal feature extraction network. In this layer, three parallel branches are designed, each of which uses convolution kernels of different sizes to capture the multi-scale information in the feature matrix. The first branch uses 64 1×1 convolution kernels. Small convolution kernels can extract local linear combination features and capture the direct correlation between features in a fine-grained space; the second branch uses 96 3×3 convolution kernels to capture local context information with a larger receptive field; the third branch uses 128 5×5 convolution kernels to further expand the receptive field to extract feature relationships in a larger range. Through the parallel processing of these three branches, feature information of different scales is captured simultaneously in the same layer of the network to generate multi-scale feature maps. The multi-scale feature map is input into the dynamic temporal attention module of the dynamic temporal feature extraction network, which consists of a temporal self-attention submodule and a channel attention submodule. The temporal self-attention submodule focuses on analyzing the dependency of feature maps in the temporal dimension. By calculating the attention weight of features at a specific time point to features at other time points, the parts with key temporal relationships in the audio signal are highlighted. The channel attention submodule weights the features of the channel dimension, aiming to highlight the feature channels with high discriminative power while suppressing irrelevant or redundant features. The combination of these two submodules generates a double-enhanced feature map. The double-enhanced feature map is subjected to a dynamic time convolution operation to generate an adaptive feature map. The dynamic time convolution operation dynamically adjusts the weight distribution during the convolution process, enabling the network to adaptively capture complex temporal patterns according to the feature changes of different audio signals. The adaptive feature map is input into the multi-scale feature pyramid network of the dynamic time feature extraction network for processing. The multi-scale feature pyramid network aims to hierarchically extract the different time scale features of the audio signal. Each layer of the network focuses on a different time scale, from short-term patterns to long-term dependencies, to obtain multi-scale time-frequency features. The multi-scale time-frequency features are cross-scale feature aggregation, and an aggregated feature vector is generated by fusing features at different scales. The aggregated feature vector is subjected to temporal dependency analysis to model the complex temporal patterns in the audio signal. The contextual relationship of the feature vector in the time dimension is captured by a deep neural network to generate temporal modeling features. In order to improve the feature expression ability, the temporal modeling features are processed by residual dense connections. Residual dense connection directly accumulates or connects features between different network layers, retains the original information in the previous layers of the network, reduces the gradient vanishing problem, enhances the diversity and expression ability of features, and obtains deep fusion features. Time-frequency attention enhancement is performed on deep fusion features. By adjusting the attention distribution in the time and frequency dimensions, the network can focus on the most useful information areas for the classification and analysis of vehicle speech features, and obtain deep time-frequency features of vehicle speech.
[0017] S5, performs speech detection and classification based on the deep time-frequency features of the in-vehicle speech, and outputs the in-vehicle speech detection results.
[0018] Among them, the deep time-frequency features of vehicle speech are divided into multiple subsequences, each of which corresponds to the time-frequency information of the audio signal within a specific time range. Each subsequence is windowed, and the edge effect is smoothed by the window function and the local characteristics of the time-frequency feature segmented sequence are enhanced to obtain the time-frequency feature segmented sequence. The noise feature analysis is performed on the time-frequency feature segmented sequence, focusing on the characteristics of three common vehicle environmental noises: engine noise, road noise and wind noise. By calculating the power spectrum density and harmonic ratio of each noise, the spectral characteristics and energy distribution laws of these noises are extracted, and then a noise feature library is established. The power spectrum density is used to characterize the frequency domain distribution of noise, while the harmonic ratio reflects the ratio of the periodic component of noise to background noise. Based on the noise feature library, the Mahalanobis distance between features is calculated to construct a feature distance matrix. The Mahalanobis distance is a distance measurement method that considers the covariance relationship and can effectively evaluate the similarity between different noise features. By constructing a feature distance matrix, the degree of difference between different noise categories is quantified. The feature distance matrix is input into the BIRCH clustering algorithm for iterative calculation, and a hierarchical clustering tree is gradually generated. The BIRCH clustering algorithm can efficiently process a large amount of data and gradually cluster features in a hierarchical manner to provide hierarchical results for the classification of different noise categories. The hierarchical clustering tree is split and merged to optimize the noise category classification results. By analyzing the feature differences and similarities between nodes, the overly detailed nodes are merged, and the fuzzy categories are split to ensure that the final noise category classification results are highly accurate and clear. The noise category classification results are matched with the time-frequency feature segmentation sequence, and the correlation between each time-frequency feature and each noise category is calculated to obtain the noise feature weight, which reflects the proportion of noise components in each time-frequency feature segment. In the dynamic noise reduction stage, the time-frequency feature segmentation sequence is input into the Wiener filter. Based on the previously calculated noise feature weight, the filter coefficient is dynamically adjusted to adaptively suppress the noise signal and retain the speech component. The dynamic adjustment mechanism can optimize the performance of the filter according to the real-time changes of the environmental noise and significantly improve the noise reduction effect. By denoising each subsequence, a denoised subsequence is generated, and these denoised subsequences are spliced and reconstructed using the overlap-addition method to obtain the denoised speech features. The overlap-add method can effectively smooth the boundaries between different subsequences and preserve the integrity and coherence of speech to the maximum extent. The noise reduction speech features are input into a two-layer neural network classifier for speech detection and classification. The classifier extracts and combines features through a two-layer structure. The first layer performs preliminary classification of noise reduction speech features to capture simple patterns and relationships, while the second layer further optimizes and comprehensively analyzes the preliminary classification results to ensure the accuracy and robustness of the classification results. Through this process, the classifier can distinguish between speech signals and non-speech signals in a vehicle environment and output vehicle speech detection results.
[0019] The noise-reduced speech features are input into the first sub-detection network of the two-layer neural network classifier. The first sub-detection network captures multi-dimensional fine-grained features and macro-level speech patterns from the input features through its designed local and global feature extraction modules. Local features can reflect the detailed information in the speech signal in a short period of time, such as instantaneous phoneme changes, while global features capture speech rhythm and prosodic patterns over a longer time range. Through this step, a multi-dimensional feature representation is obtained. The multi-dimensional feature representation is feature decoupled to decompose the complex multi-dimensional features into independent phoneme-level, syllable-level and prosodic-level features, which correspond to different levels of information in the speech signal. Among them, the phoneme-level feature mainly focuses on the detailed changes of a single phoneme in the speech signal, the syllable-level feature reflects the speech characteristics of a larger unit composed of phonemes, and the prosodic-level feature describes the global rhythm, stress and intonation pattern of the speech. Through feature decoupling, these levels of information are extracted separately, so as to more carefully characterize the diversity and complexity of the speech signal. Through subspace mapping, the phoneme-level, syllable-level and rhythm-level features are projected into orthogonal feature spaces respectively to ensure their independence and complementarity, and generate decoupled feature vectors. The decoupled feature vectors are input into the mutual information maximization module of the two-layer neural network classifier, and the mutual information scores between features at different levels are calculated by contrastive learning. The mutual information score can quantify the degree of correlation between features at each level, generate a feature correlation matrix, and describe the common and unique information of features at each level in the speech signal. Based on the feature correlation matrix, adaptive fusion weights are constructed. The first layer of fusion features are generated by weighted combination of the decoupled feature vectors. The first layer of fusion features are input into the second sub-detection network of the two-layer neural network classifier. The second sub-detection network adopts a variational autoencoder structure, which has the unique advantage of being able to learn potential probability distribution parameters from complex features. Through the encoder module, the first layer of fusion features are mapped to a latent space, and the corresponding probability distribution parameters are generated to describe the distribution form of the features in the latent space. In the latent space, the probability distribution parameters are sampled by reparameterization technology to generate latent variables. The latent variables are reconstructed into feature vectors through the decoder module. These reconstructed feature vectors contain the core information of the input features while effectively suppressing noise and redundant features. The reconstruction error of the reconstructed feature vectors and the initial denoised speech features are calculated, and the KL divergence loss is combined to optimize the representation ability and distribution consistency of the features. Through this process, the speech detection probability is generated, indicating the probability distribution of speech signals in different categories. The speech detection probability is inferred by Bayesian posteriori. Combined with the prior knowledge of the speech detection task, the optimal decision threshold is calculated through Bayesian inference to achieve accurate speech detection classification. The output vehicle speech detection result combines the advantages of deep feature extraction, latent space modeling and probabilistic inference, and is highly accurate and robust.
[0020] In one example, the original audio signal collected by the vehicle microphone is framed and fast Fourier transformed to obtain a frequency domain feature sequence, including: The original audio signal collected by the vehicle microphone is divided into frames to obtain a first audio frame sequence; Inputting the first audio frame sequence into a first-order high-pass filter for signal pre-emphasis processing to obtain a second audio frame sequence; Performing window function processing on the second audio frame sequence to obtain a third audio frame sequence, and performing fast Fourier transform on the third audio frame sequence to obtain an initial spectrum sequence; Calculate the amplitude spectrum of the initial spectrum sequence, and normalize the amplitude spectrum to obtain a normalized spectrum sequence; The normalized spectrum sequence is input into the triangular filter bank for sub-band decomposition to obtain a sub-band energy sequence, and the sub-band energy sequence is logarithmically compressed to obtain a logarithmic energy sequence; Perform P-order discrete cosine transform on the logarithmic energy sequence to obtain the frequency domain feature sequence, where P is 13.
[0021] In this example, the original audio signal collected by the car microphone is framed. The input audio signal is set to ,in is the signal value, is the index of the sampling point (in the range of ,in is the total length of the sampled signal). The continuous signal is divided into short-time frames to ensure the short-time stability of the signal. Assume that the length of each frame is sampling points, the frame shift is ,in Indicates the overlap ratio of frames. By dividing the frames, frames, and the obtained signal of each frame is expressed as ,in is the frame index, is the sampling point index within the frame. The input is sent to a first-order high-pass filter for pre-emphasis. The filter suppresses low-frequency components and enhances high-frequency components. The formula is: ; in is the output signal after filtering, is the pre-emphasis coefficient, and Represent the signal values of the current sampling point and the previous sampling point respectively. After filtering, the second audio frame sequence obtained is The spectrum leakage problem caused by the boundary effect is reduced by window function processing. The commonly used Hamming window formula is: ; in is the window function value, is the sampling point index within the window, is the length of the window. The frame signal after the window function is ,in is the third audio frame sequence after windowing. Apply fast Fourier transform to convert the time domain signal into frequency domain signal. The mathematical formula of Fourier transform is: ; in is the frequency domain signal The complex value of the frequency components, is the frequency index (in the range of ), is a complex exponential function, representing the basis function in the frequency domain. After Fourier transform, the amplitude spectrum of the signal is calculated , that is, take the modulus of the complex spectrum to represent the energy of each frequency component. In order to eliminate the energy difference between different frames, the amplitude spectrum needs to be normalized. The formula of the normalized amplitude spectrum is: ; in is the normalized amplitude spectrum value, is the maximum value of the amplitude spectrum, which is used to normalize the amplitude to the interval [0,1]. The input is sent to a triangular filter bank (Mel filter bank) for subband decomposition. The filter bank consists of The center frequency of the filter is distributed according to the Mel scale. The Mel scale calculation formula is: ; in is the frequency. The output of the filter bank is the energy value in each frequency band, which is called the subband energy sequence ,in is the filter index. Logarithmic compression is performed to reduce the difference in energy range. The formula is: ; in is the logarithmic energy, is a small constant used to avoid numerical problems with logarithms (usually taken as ). Logarithmic energy series Discrete cosine transform is performed to remove the correlation between frequency bands and extract compact features. The formula for discrete cosine transform is: ; in It is The cepstral coefficients, Is the dimension index of the feature (usually , is the number of Mel filter banks. Through the above steps, the frequency domain feature sequence obtained It is a compact representation of each frame and is used for speech detection or classification tasks.
[0022] In one example, short-time energy, zero-crossing rate, spectrum centroid, and Mel-frequency cepstrum coefficients are extracted from the frequency domain feature sequence, and the first-order difference and second-order difference features of each feature are combined to obtain a multi-dimensional acoustic feature point sequence, including: Calculate the short-time energy of each frame signal for the frequency domain feature sequence, divide the square sum of each frame signal by the frame length, and obtain the short-time energy sequence; The number of symbol changes of adjacent sampling points is calculated for the frequency domain feature sequence, and the number of changes is divided by 2 times the frame length to obtain a zero-crossing rate sequence. The ratio of the frequency-weighted amplitude spectrum to the sum of the amplitude spectrum is calculated for the frequency domain feature sequence to obtain a spectrum centroid sequence. The frequency domain feature sequence is input into the Mel filter bank for filtering to obtain the filtering result, and the filtering result is subjected to logarithmic operation and discrete cosine transform to obtain the Mel frequency cepstrum coefficient sequence; The short-time energy sequence, zero-crossing rate sequence, spectrum centroid sequence and Mel-frequency cepstrum coefficient sequence are combined to obtain the basic feature sequence; The difference between adjacent frames of the basic feature sequence is calculated to obtain a first-order difference feature sequence, and the difference between adjacent frames of the first-order difference feature sequence is calculated to obtain a second-order difference feature sequence; The basic feature sequence, the first-order difference feature sequence and the second-order difference feature sequence are concatenated to obtain a multi-dimensional acoustic feature point sequence.
[0023] In this example, first, the short-time energy of each frame signal is calculated to characterize the energy distribution of the signal in each frame. The calculation formula of short-time energy is: ; in, Indicates The short-term energy of the frame, is the total number of frequency components per frame, The frequency domain signal By calculating the sum of squares of each frame and normalizing the frame length, we get a time series , reflecting the time-varying energy characteristics of the signal. Calculate the zero-crossing rate of the signal, which is used to capture the changes in frequency components. The zero-crossing rate is defined as the number of times the symbols of adjacent sampling points change in each frame of the signal, and the formula is: ; in, Indicates The zero-crossing rate of the frame, It is a condition indicator function, which takes the value 1 when the condition is met, otherwise it takes the value 0. Indicates the sign change of two adjacent points. By dividing the number of sign changes by , normalize the zero crossing rate so that its value is within a fixed range. Calculate the spectrum centroid to describe the center of gravity of the spectrum energy. The formula for the spectrum centroid is: ; in, Indicates The spectral centroid of the frame, is the frequency index, is the value of the amplitude spectrum. The spectrum centroid reflects the distribution center of the signal energy in frequency. For example, for a signal whose spectrum amplitude is concentrated in the low-frequency area, the spectrum centroid will be biased towards low frequencies. Input Mel filter bank, which consists of several triangular filters whose center frequencies are evenly distributed on the Mel scale. Through weighted processing of Mel filter, we get subband energy sequence ,in is the filter index. The subband energy is calculated as: ; in It is The weight function of the Mel filter. Then the subband energy is logarithmically operated to obtain the logarithmic energy sequence: ; in is the logarithmic energy value, is a small positive value used to avoid numerical problems with logarithmic operations (such as ). Perform discrete cosine transform on the logarithmic energy sequence and extract a compact Mel-frequency cepstral coefficient (MFCC) sequence: ; in Indicates The cepstral coefficients, is the number of filter banks. , zero-crossing rate sequence , spectral centroid sequence and Mel-frequency cepstral coefficient sequence Combine to form a basic feature sequence. Calculate the difference between adjacent frames of the basic feature sequence, the formula is: ; in Indicates Differential features of frames, is the feature vector of the basic feature sequence. Similarly, the difference between adjacent frames is calculated for the differential feature to obtain the second-order differential feature: ; The basic feature sequence, the first-order difference feature sequence and the second-order difference feature sequence are concatenated to form the final multi-dimensional acoustic feature point sequence.
[0024] In one example, a multi-dimensional acoustic feature point sequence is arranged in time sequence to construct an original feature matrix, and a principal component analysis is performed on the original feature matrix to obtain a feature description matrix, including: Construct feature vectors for the multi-dimensional acoustic feature point sequence according to the time sequence relationship, and group the adjacent M feature points into feature groups to obtain the original feature sequence; The original feature sequence is rearranged into an N×D dimensional original feature matrix, where N is the total number of feature points and D is the dimension of each feature point, to obtain the first feature matrix; Subtract the mean of each column of the first feature matrix and divide the result by the standard deviation of the column to obtain a standardized feature matrix; Calculate the D×D dimensional covariance matrix based on the standardized feature matrix to obtain the feature covariance matrix; Solve the characteristic equation for the characteristic covariance matrix to obtain D eigenvalues and corresponding D eigenvectors; Calculate the cumulative contribution rate based on the eigenvalue, select the first K eigenvalues and eigenvectors corresponding to when the cumulative contribution rate exceeds the target value, and construct the characteristic transformation matrix; The standardized feature matrix is multiplied by the feature transformation matrix to obtain a K-dimensional reduced dimension feature matrix, and the K-dimensional reduced dimension feature matrix is normalized to map each element to the [0,1] interval to obtain a feature description matrix.
[0025] In this example, according to the timing relationship, the adjacent Feature points Combined into a feature group, the dimension of each feature group is . Define a new feature group as , and its calculation formula is: ; Where concat represents the concatenation operation of vectors. By using the sliding window method, the entire multi-dimensional acoustic feature point sequence is traversed to generate feature groups, and obtain the original feature sequence ,in is the total number of feature groups. The original feature sequence is rearranged into a The characteristic matrix of The constructed matrix is: ; matrix Each row of represents a feature group, and each column represents a feature dimension. Each column of is standardized to eliminate the dimensional differences of different feature dimensions. The formula for the standardization operation is: ; in It is The mean of the column, is the standard deviation of the jth column, Indicates Line The value of the column, Represents the standardized Column. After getting the standardized feature matrix After that, the covariance matrix is calculated to capture the linear relationship between the features. The formula for calculating the covariance matrix is: ; in is the covariance matrix, express The transpose of the covariance matrix Solve for the eigenvalues and eigenvectors. The characteristic equation is: ; in It is eigenvalues, is the corresponding eigenvector. The eigenvalue represents the variance contribution of each principal component, and the eigenvector defines the direction of the principal component. Based on the eigenvalue, the cumulative contribution rate is calculated: ; in It is before The cumulative contribution rate of the eigenvalues. When the target value (such as 95%) is reached, select The eigenvalues and their corresponding eigenvectors constitute the characteristic transformation matrix: ; Multiply the standardized feature matrix by the feature transformation matrix to complete the dimensionality reduction operation: ; in is the feature matrix after dimensionality reduction. Normalize and map each element to the interval [0,1]. The normalization formula is: ; in and They are The minimum and maximum values of the column, is the normalized feature description matrix.
[0026] In one example, the feature description matrix is input into a dynamic time feature extraction network for deep time-frequency feature analysis to obtain deep time-frequency features of vehicle-borne speech, including: The feature description matrix is input into the multi-branch parallel convolution layer of the dynamic temporal feature extraction network. The multi-branch parallel convolution layer contains three parallel branches. The first branch uses 64 1×1 convolution kernels, the second branch uses 96 3×3 convolution kernels, and the third branch uses 128 5×5 convolution kernels to obtain a multi-scale feature map. The multi-scale feature map is input into the dynamic temporal attention module of the dynamic temporal feature extraction network. The dynamic temporal attention module includes a temporal self-attention submodule and a channel attention submodule. The attention weights of the temporal dimension and the channel dimension are calculated to obtain a double enhanced feature map. Perform dynamic temporal convolution operation on the double enhanced feature map to obtain an adaptive feature map, and input the adaptive feature map into the multi-scale feature pyramid network of the dynamic temporal feature extraction network to extract features of different time scales and obtain multi-scale time-frequency features; Perform cross-scale feature aggregation on multi-scale time-frequency features to obtain aggregated feature vectors, and perform temporal dependency analysis on the aggregated feature vectors to obtain temporal modeling features; The time series modeling features are densely connected with residuals to obtain deep fusion features, and the deep fusion features are enhanced with time-frequency attention to obtain deep time-frequency features of in-vehicle speech.
[0027] In this example, the feature description matrix Input the multi-branch parallel convolution layer of the dynamic temporal feature extraction network, which extracts multi-scale features through three parallel branches. Each branch uses a convolution kernel of different sizes, respectively. and The first branch adopts Convolution, the convolution operation formula is: ; in is the output of the first branch, Is the convolution kernel, containing 64 Convolution kernel, is the bias term, * represents the convolution operation, is an activation function (such as ReLU). This branch is used to capture the linear combination of local features. The second branch uses Convolution, the convolution operation is: ; in , , containing 96 The third branch uses Convolution, the formula is: ; in These two branches are used to capture larger contextual information and complex patterns. The output feature maps of these three branches are concatenated along the channel dimension to obtain a multi-scale feature map: ; in The multi-scale feature map is input into the dynamic temporal attention module, which contains the temporal self-attention submodule and the channel attention submodule. The temporal self-attention submodule calculates the temporal attention weight by capturing the global correlation of features in the temporal dimension. The formula is: ; ; in They are query, key and value matrices, respectively, and are composed of multi-scale feature maps Through linear transformation, we get is the temporal attention weight matrix, is the feature after temporal enhancement. The channel attention submodule highlights important channel information by calculating the weight of the feature in the channel dimension. The formula is: ; ; in is the channel weight, pool represents the global average pooling operation, and © is the element-by-element multiplication. The dynamic temporal attention module outputs a dual enhanced feature map: ; The dual enhanced feature maps are fed into a dynamic temporal convolution operation to capture complex temporal patterns, as follows: ; in is a dynamically adjusted convolution kernel, is an adaptive feature map. The adaptive feature map is input into the multi-scale feature pyramid network, which extracts features of different time scales through layered convolution. The formula is: ; in is the layer index, It is the pyramid network The convolution kernel of the layer generates multi-scale time-frequency features. The multi-scale time-frequency features are aggregated across scales to obtain the aggregated feature vector: ; The aggregated feature vectors are analyzed for temporal dependencies to capture the characteristic patterns over a long period of time and generate temporal modeling features. Finally, the temporal modeling features are enhanced through residual dense connections to express the features: ; And enhanced by time-frequency attention, the formula is: ; Finally, the deep time-frequency features of vehicle speech are obtained.
[0028] In one example, speech detection and classification are performed based on the deep time-frequency features of vehicle-borne speech, and the vehicle-borne speech detection results are output, including: The deep time-frequency features of the vehicle speech are divided into multiple subsequences, and each subsequence is subjected to windowing processing to obtain a time-frequency feature segmented sequence; Perform characteristic analysis of engine noise, road noise and wind noise on the time-frequency feature segmented sequence, and calculate the power spectrum density and harmonic ratio of each noise to obtain a noise feature library; The Mahalanobis distance between features is calculated based on the noise feature library, and the feature distance matrix is constructed. The feature distance matrix is input into the BIRCH clustering algorithm for iterative calculation to obtain a hierarchical clustering tree. Perform node splitting and merging operations on the hierarchical clustering tree to obtain the noise category division result, and perform feature matching between the noise category division result and the time-frequency feature segmentation sequence to obtain the noise feature weight; The time-frequency feature segmented sequence is input into the Wiener filter, the filter coefficient is dynamically adjusted according to the noise feature weight to obtain the denoised subsequence, and the denoised subsequence is reconstructed into the denoised speech feature through the overlap-addition method; The noise reduction speech features are input into a two-layer neural network classifier for speech detection and classification, and the vehicle-mounted speech detection results are output.
[0029] In this example, the deep time-frequency feature matrix Divide into multiple subsequences according to the time dimension, each subsequence length is , the subsequence obtained by segmentation is ,in Indicates subsequences, a total of Subsequences. Apply window function weighted processing to each subsequence to reduce the impact of boundary effects. The window function uses Hamming window, which is defined as: ; in is the sampling point index within the window, is the length of the window. After weighting by the window function, the subsequence is updated to: ; in is the windowed time-frequency feature segmented sequence. The time-frequency feature segmented sequence is subjected to characteristic analysis of engine noise, road noise and wind noise. The power spectral density (PSD) of each noise is calculated, and the formula is: ; in is the frequency index, is the length of the time series, It is Frame in The power spectral density reflects the distribution of noise energy with frequency. The harmonic ratio is calculated to evaluate the periodic component of the noise, which is defined as: ; Through the above calculations, the spectral characteristics of engine noise, road noise and wind noise are extracted to build a noise feature library. Based on the noise feature library, the Mahalanobis distance between features is calculated, which is defined as: ; in and are two sets of eigenvectors, is the covariance matrix. The Mahalanobis distance quantifies the differences between features and generates a feature distance matrix , which represents the distance relationship between all subsequences. The feature distance matrix is input into the BIRCH clustering algorithm for iterative calculation. The BIRCH algorithm constructs a hierarchical clustering tree by gradually merging and splitting clusters. The goal of clustering is to divide noise subsequences into different categories based on feature similarity. After obtaining the hierarchical clustering tree, the noise category division results are optimized by splitting and merging the nodes. Each subsequence is marked as belonging to a specific noise category. The noise category division results are matched with the time-frequency feature segmentation sequence, and the correlation between each segment feature and the noise category is calculated to obtain the noise feature weight. The time-frequency feature segmented sequence is input into the Wiener filter, and the filter coefficient is dynamically adjusted by the noise feature weight to perform noise reduction. The output of the Wiener filter is: ; Where SNR is the signal-to-noise ratio, dynamically adjusting the gain of the filter. All denoising subsequences are reconstructed into a complete denoising speech feature matrix through overlap-addition. The denoising speech features are input into a two-layer neural network classifier for speech detection and classification. The first layer of the network is used to extract nonlinear features, and the formula is: ; in is the hidden layer output, and are weights and biases. The second layer of the network completes the classification, and the formula is: ; in is a classification probability vector, which indicates the probability that the speech signal belongs to different categories. The output classification result is the vehicle-mounted speech detection result.
[0030] In one example, the noise reduction speech features are input into a two-layer neural network classifier for speech detection and classification, and the vehicle-mounted speech detection results are output, including: The noise reduction speech features are input into the first sub-detection network of the double-layer neural network classifier to extract local and global features, and obtain a multi-dimensional feature representation; Decouple the multidimensional feature representation, extract phoneme-level features, syllable-level features and rhythm-level features respectively, and project the phoneme-level features, syllable-level features and rhythm-level features into the orthogonal feature space through subspace mapping to obtain the decoupled feature vector; The decoupled feature vector is input into the mutual information maximization module of the two-layer neural network classifier, and the mutual information scores between features at different levels are calculated by contrastive learning to obtain the feature correlation matrix; Adaptive fusion weights are constructed based on the feature correlation matrix, and the decoupled feature vectors are weighted combined to obtain the first layer of fusion features; The first layer fusion features are input into the second sub-detection network of the two-layer neural network classifier. The second sub-detection network adopts a variational autoencoder structure. The encoder maps the features to the latent space to obtain probability distribution parameters. The probability distribution parameters are reparameterized and sampled to generate latent variables, and the features are reconstructed through the decoder to obtain a reconstructed feature vector. The reconstruction error is calculated by combining the reconstructed feature vector with the denoised speech features, and combined with the KL divergence loss to obtain the speech detection probability; Perform Bayesian posterior inference on the speech detection probability, calculate the optimal decision threshold based on prior knowledge, and output the in-vehicle speech detection result.
[0031] In this example, the noise reduction speech feature matrix Input the first sub-detection network of the two-layer neural network classifier, which is used to extract local and global features. Local feature extraction is achieved through a one-dimensional convolution operation, capturing feature patterns within a short time range, and its formula is: ; in is a local feature representation, is the local convolution kernel, is the length of the convolution kernel, is the number of output channels, is the bias term, is an activation function (such as ReLU). Global feature extraction is achieved through global average pooling and multi-head self-attention mechanism, and its formula is: ; in are query, key, and value matrices, respectively, and are represented by The linear transformation is obtained, and softmax is used to normalize the attention weight. The local and global features are concatenated to obtain a multi-dimensional feature representation: ; in It is a complete multi-dimensional feature representation. The multi-dimensional feature representation is decoupled to extract phoneme-level, syllable-level and rhythm-level features respectively. Phoneme-level features are extracted by convolution in a short time range, and the formula is: ; in , Represents the dimension of phoneme-level features. Syllable-level features are extracted through convolution with a larger receptive field: ; in , Represents the dimension of syllable-level features. Rhythm-level features are extracted through global pooling: ; in , Represents the dimension of the rhythm-level feature. Through subspace mapping, the phoneme-level, syllable-level, and rhythm-level features are projected into the orthogonal feature space. The formula is: ; in phoneme,syllable,prosody is a mapping matrix to ensure orthogonality. The decoupled feature vector is input into the mutual information maximization module, and the mutual information scores between features at different levels are calculated through contrastive learning. The mutual information score is calculated by the following formula: ; in represents entropy, They are feature vectors at different levels. Construct feature correlation matrix based on mutual information score . Based on the feature correlation matrix, calculate the adaptive fusion weight , and perform weighted combination of the decoupled features to generate the first layer of fusion features: ; The first layer of fusion features is input into the second sub-detection network, which adopts a variational autoencoder structure. The encoder maps the features to the latent space to obtain the probability distribution parameters: ; in are the mean and standard deviation of the latent space, is the dimension of the latent variable. Sampling is performed using the reparameterization technique, and the formula is: ; in is the noise of standard normal distribution. The decoder reconstructs the feature vector through the latent variables: ; And calculate the reconstruction error: ; Combined with KL divergence loss: ; The total loss is: ; The speech detection probability is calculated based on the reconstruction error, and the optimal decision threshold is determined by combining Bayesian posterior inference with prior knowledge. The formula is: ; in, Speech Represents a given fusion feature The posterior probability that the signal belongs to the speech category ("Speech") when . This value reflects the system's confidence that the current signal is speech and is the target result output by the classifier; Speech) means that when the signal is speech ("Speech"), the fusion feature is observed The probability of is called the likelihood function, which is estimated by the variational autoencoder model. The distribution of , for example through a Gaussian distribution: ; in and are the parameters generated by the encoder, representing the mean and standard deviation respectively; Speech) represents the prior probability of speech category, that is, the probability that the signal belongs to the speech category when there is no feature information. Combined with the prior knowledge, the optimal decision threshold is calculated and the vehicle speech detection result is output.
[0032] The present invention also includes the following data enhancement steps: dividing the original audio signal into a foreground speech segment and a background noise segment, performing speech activity detection through an endpoint detection algorithm to obtain a speech segmentation sequence; performing time domain transformation enhancement on the speech segmentation sequence, including time stretching, speed perturbation and pitch conversion, to obtain a time domain enhancement sequence; performing frequency domain mixing enhancement on the time domain enhancement sequence, generating a variety of spectrum representations through spectrum masking and frequency jittering, to obtain a frequency domain enhancement sequence; performing adversarial noise injection on the frequency domain enhancement sequence, synthesizing real vehicle-mounted environmental noise through a generative adversarial network, and mixing the noise signal with the speech signal at different signal-to-noise ratios to obtain an adversarial enhancement sequence. Strong sequence; input the adversarial enhancement sequence into the multi-view consistency learning module, generate sample representations of multiple viewpoints through different data transformations, and obtain a multi-view feature sequence; construct a contrastive learning task based on the multi-view feature sequence, and obtain the contrast loss function by maximizing the mutual information between different viewpoints of the same sample and minimizing the mutual information between different samples; combine the contrast loss function with the classification loss function of the two-layer neural network classifier to construct a joint optimization objective, update the model parameters through the back propagation algorithm, and obtain the optimized feature extraction model; apply the optimized feature extraction model to new test samples for vehicle-mounted speech feature extraction and detection.
[0033] Reference Figure 2 , this embodiment provides an audio feature extraction device based on a neural network, comprising: Transformation module 1, used for performing frame processing and fast Fourier transform on the original audio signal collected by the vehicle microphone to obtain a frequency domain feature sequence; Extraction module 2 is used to extract short-time energy, zero-crossing rate, spectrum centroid and Mel-frequency cepstrum coefficients from the frequency domain feature sequence, and combine the first-order difference and second-order difference features of each feature to obtain a multi-dimensional acoustic feature point sequence; Construction module 3 is used to arrange the multi-dimensional acoustic feature point sequence in time sequence, construct an original feature matrix, and perform principal component analysis on the original feature matrix to obtain a feature description matrix; Analysis module 4, used for inputting the feature description matrix into the dynamic time feature extraction network to perform deep time-frequency feature analysis to obtain the deep time-frequency features of the vehicle-borne speech; The classification module 5 is used to perform speech detection and classification based on the deep time-frequency features of the in-vehicle speech and output the in-vehicle speech detection results.
[0034] In this embodiment, for the specific implementation of each unit in the above device embodiment, please refer to the above method embodiment, which will not be described in detail here.
[0035] It should be noted that, in this article, the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, device, article or method including a series of elements includes not only those elements, but also includes other elements not explicitly listed, or also includes elements inherent to such process, device, article or method. In the absence of further restrictions, an element defined by the sentence "includes a ..." does not exclude the presence of other identical elements in the process, device, article or method including the element.
[0036] The above description is only a preferred embodiment of the present invention, and does not limit the patent scope of the present invention. Any equivalent structure or equivalent process transformation made by using the contents of the present invention specification and drawings, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present invention.
Claims
1. A method for extracting audio features based on a neural network, characterized in that: The following steps are involved: The original audio signal collected by the vehicle microphone is processed by frame division and fast Fourier transform to obtain a frequency domain feature sequence; Extracting short-time energy, zero-crossing rate, spectrum centroid and Mel-frequency cepstrum coefficients from the frequency domain feature sequence, and combining the first-order difference and second-order difference features of each feature to obtain a multi-dimensional acoustic feature point sequence; Arranging the multidimensional acoustic feature point sequence in time sequence to construct an original feature matrix, and performing principal component analysis on the original feature matrix to obtain a feature description matrix; Inputting the feature description matrix into a dynamic time feature extraction network to perform deep time-frequency feature analysis to obtain deep time-frequency features of the vehicle-borne speech; Speech detection and classification are performed based on the deep time-frequency features of the in-vehicle speech, and the in-vehicle speech detection result is output.
2. The method for extracting audio features based on a neural network according to claim 1, characterized in that: The original audio signal collected by the vehicle microphone is subjected to frame processing and fast Fourier transform to obtain a frequency domain feature sequence, including: The original audio signal collected by the vehicle microphone is divided into frames to obtain a first audio frame sequence; Inputting the first audio frame sequence into a first-order high-pass filter for signal pre-emphasis processing to obtain a second audio frame sequence; Performing window function processing on the second audio frame sequence to obtain a third audio frame sequence, and performing fast Fourier transform on the third audio frame sequence to obtain an initial spectrum sequence; Calculating an amplitude spectrum for the initial frequency spectrum sequence, and normalizing the amplitude spectrum to obtain a normalized frequency spectrum sequence; Inputting the normalized spectrum sequence into a triangular filter bank for sub-band decomposition to obtain a sub-band energy sequence, and performing logarithmic compression processing on the sub-band energy sequence to obtain a logarithmic energy sequence; Perform a P-order discrete cosine transform on the logarithmic energy sequence to obtain a frequency domain feature sequence, where P is 13.
3. The method for extracting audio features based on a neural network according to claim 2, characterized in that: The short-time energy, zero-crossing rate, spectrum centroid and Mel-frequency cepstrum coefficients are extracted from the frequency domain feature sequence, and the first-order difference and second-order difference features of each feature are combined to obtain a multi-dimensional acoustic feature point sequence, including: Calculating the short-time energy of each frame signal for the frequency domain feature sequence, dividing the sum of squares of each frame signal by the frame length to obtain a short-time energy sequence; The frequency domain characteristic sequence is calculated for the number of symbol changes of adjacent sampling points, and the number of changes is divided by 2 times the frame length to obtain a zero-crossing rate sequence, and the frequency domain characteristic sequence is calculated for the ratio of the frequency-weighted amplitude spectrum to the sum of the amplitude spectrum to obtain a spectrum centroid sequence; Inputting the frequency domain feature sequence into a Mel filter bank for filtering to obtain a filtering result, and performing logarithmic operation and discrete cosine transform on the filtering result to obtain a Mel frequency cepstrum coefficient sequence; Performing feature combination on the short-time energy sequence, the zero-crossing rate sequence, the spectrum centroid sequence and the Mel-frequency cepstrum coefficient sequence to obtain a basic feature sequence; Calculating the difference between adjacent frames of the basic feature sequence to obtain a first-order difference feature sequence, and calculating the difference between adjacent frames of the first-order difference feature sequence to obtain a second-order difference feature sequence; The basic feature sequence, the first-order difference feature sequence and the second-order difference feature sequence are feature concatenated to obtain a multi-dimensional acoustic feature point sequence.
4. The method for extracting audio features based on a neural network according to claim 3, characterized in that: The step of arranging the multidimensional acoustic feature point sequence in time sequence to construct an original feature matrix, and performing principal component analysis on the original feature matrix to obtain a feature description matrix includes: Constructing a feature vector for the multidimensional acoustic feature point sequence according to a time sequence relationship, forming a feature group with M adjacent feature points, and obtaining an original feature sequence; Rearranging the original feature sequence into an N×D dimensional original feature matrix, where N is the total number of feature points and D is the dimension of each feature point, to obtain a first feature matrix; Subtract the mean of each column of the first feature matrix from the mean of the column, and divide the result by the standard deviation of the column to obtain a standardized feature matrix; Calculate a D×D dimensional covariance matrix based on the standardized feature matrix to obtain a feature covariance matrix; Solving the characteristic equation for the characteristic covariance matrix to obtain D eigenvalues and corresponding D eigenvectors; Calculate the cumulative contribution rate based on the eigenvalues, select the first K eigenvalues and eigenvectors corresponding to when the cumulative contribution rate exceeds the target value, and construct a feature transformation matrix; The standardized feature matrix is multiplied by the feature transformation matrix to obtain a K-dimensional reduced-dimensional feature matrix, and the K-dimensional reduced-dimensional feature matrix is normalized to map each element to the interval [0, 1] to obtain a feature description matrix.
5. The method for extracting audio features based on a neural network according to claim 4, characterized in that: The step of inputting the feature description matrix into a dynamic time feature extraction network to perform deep time-frequency feature analysis to obtain deep time-frequency features of vehicle-borne speech includes: Inputting the feature description matrix into a multi-branch parallel convolutional layer of a dynamic temporal feature extraction network, wherein the multi-branch parallel convolutional layer comprises three parallel branches, wherein the first branch uses 64 1×1 convolution kernels, the second branch uses 96 3×3 convolution kernels, and the third branch uses 128 5×5 convolution kernels, to obtain a multi-scale feature map; Inputting the multi-scale feature map into the dynamic time attention module of the dynamic time feature extraction network, wherein the dynamic time attention module includes a temporal self-attention submodule and a channel attention submodule, and calculating the attention weights of the temporal dimension and the channel dimension to obtain a dual enhanced feature map; Performing a dynamic time convolution operation on the dual enhanced feature map to obtain an adaptive feature map, and inputting the adaptive feature map into a multi-scale feature pyramid network of the dynamic time feature extraction network to extract features of different time scales to obtain multi-scale time-frequency features; Performing cross-scale feature aggregation on the multi-scale time-frequency features to obtain an aggregated feature vector, and performing temporal dependency analysis on the aggregated feature vector to obtain a temporal modeling feature; The time series modeling features are subjected to residual dense connection to obtain deep fusion features, and the deep fusion features are subjected to time-frequency attention enhancement to obtain deep time-frequency features of the in-vehicle speech.
6. The method for extracting audio features based on a neural network according to claim 5, characterized in that: The performing speech detection and classification based on the deep time-frequency features of the vehicle-borne speech and outputting the vehicle-borne speech detection result includes: Dividing the in-vehicle speech deep time-frequency features into multiple subsequences, and performing windowing processing on each subsequence to obtain a time-frequency feature segmented sequence; Performing characteristic analysis of engine noise, road noise and wind noise on the time-frequency characteristic segmented sequence, and calculating the power spectrum density and harmonic ratio of each noise to obtain a noise characteristic library; Calculating the Mahalanobis distance between features based on the noise feature library, constructing a feature distance matrix, and inputting the feature distance matrix into the BIRCH clustering algorithm for iterative calculation to obtain a hierarchical clustering tree; Perform node splitting and merging operations on the hierarchical clustering tree to obtain noise category division results, and perform feature matching on the noise category division results and the time-frequency feature segmentation sequence to obtain noise feature weights; Inputting the time-frequency feature segmented sequence into a Wiener filter, dynamically adjusting the filter coefficient according to the noise feature weight to obtain a denoised subsequence, and reconstructing the denoised subsequence into a denoised speech feature by an overlap-addition method; The noise reduction speech features are input into a double-layer neural network classifier for speech detection and classification, and the vehicle-mounted speech detection result is output.
7. The method for extracting audio features based on a neural network according to claim 6, characterized in that: The step of inputting the noise reduction speech features into a double-layer neural network classifier for speech detection and classification, and outputting the vehicle-mounted speech detection result, comprises: Inputting the noise reduction speech features into a first sub-detection network of a double-layer neural network classifier to extract local and global features to obtain a multi-dimensional feature representation; Performing feature decoupling on the multidimensional feature representation, extracting phoneme-level features, syllable-level features, and rhythm-level features respectively, and projecting the phoneme-level features, syllable-level features, and rhythm-level features into an orthogonal feature space through subspace mapping to obtain a decoupled feature vector; The decoupled feature vector is input into the mutual information maximization module of the two-layer neural network classifier, and the mutual information scores between features at different levels are calculated by contrastive learning to obtain a feature correlation matrix; Constructing adaptive fusion weights based on the feature correlation matrix, performing weighted combination on the decoupled feature vectors, and obtaining a first layer of fusion features; Inputting the first layer of fusion features into the second sub-detection network of the two-layer neural network classifier, the second sub-detection network adopts a variational autoencoder structure, and maps the features to a latent space through an encoder to obtain probability distribution parameters; Reparameterization sampling is performed on the probability distribution parameters to generate latent variables, and features are reconstructed through a decoder to obtain a reconstructed feature vector, a reconstruction error is calculated by combining the reconstructed feature vector with the noise reduction speech feature, and a KL divergence loss is combined to obtain a speech detection probability; The speech detection probability is subjected to Bayesian posterior inference, an optimal decision threshold is calculated in combination with prior knowledge, and an in-vehicle speech detection result is output.
8. An audio feature extraction device based on a neural network, characterized in that: For implementing the steps of the method according to any one of claims 1 to 7, the device comprises: A transformation module is used to perform frame processing and fast Fourier transformation on the original audio signal collected by the vehicle microphone to obtain a frequency domain feature sequence; An extraction module is used to extract short-time energy, zero-crossing rate, spectrum centroid and Mel-frequency cepstrum coefficients from the frequency domain feature sequence, and combine the first-order difference and second-order difference features of each feature to obtain a multi-dimensional acoustic feature point sequence; A construction module, used to arrange the multi-dimensional acoustic feature point sequence in time sequence, construct an original feature matrix, and perform principal component analysis on the original feature matrix to obtain a feature description matrix; An analysis module, used for inputting the feature description matrix into a dynamic time feature extraction network to perform deep time-frequency feature analysis to obtain deep time-frequency features of the vehicle-mounted speech; The classification module is used to perform speech detection and classification based on the deep time-frequency features of the in-vehicle speech and output the in-vehicle speech detection result.
Citation Information
Cited By
Voice packet loss processing method and device based on deep learning and variational mode decomposition, equipment and storage medium
CN120472939A
Voice packet loss processing method, device, equipment and storage medium based on deep learning and variational mode decomposition
CN120472939B
GIS disconnecting switch opening and closing state monitoring method based on multi-dimensional voiceprint features
CN120954449A
Audio and video optimization processing method and system for handheld terminal
CN121037602A
Audio and video optimization processing method and system for handheld terminal
CN121037602B