A new weighted mel frequency cepstral coefficient feature fusion method
By performing differential operations and linear weighted fusion on MFCC features, and combining principal component analysis and histogram statistics to optimize the weights, the problem of poor feature separability in existing technologies is solved, thereby improving the recognition accuracy.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- HARBIN ENG UNIV
- Filing Date
- 2023-09-25
- Publication Date
- 2026-04-28
AI Technical Summary
Existing weight selection methods result in poor separability of fused features during feature fusion, limiting the improvement in recognition accuracy and failing to effectively consider the distribution of sample features.
By performing difference operations on the MFCC features, first-order and second-order difference features are obtained, and then linear weighted fusion is performed. The weights are calculated by combining principal component analysis and histogram statistics, and the feature distribution distance is optimized to determine the weights α1 and α2, thus constructing the weighted Mel frequency cepstral feature WMFCC.
It improves the robustness of features and recognition accuracy. By fusing features to include more time-dimensional information, it enhances the separability of features and improves the recognition accuracy of the recognition algorithm.
Smart Images

Figure CN117251822B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of underwater acoustic target recognition, specifically involving a novel weighted Mel frequency cepstral feature fusion method. Background Technology
[0002] In recent decades, the classification of ship radiated noise has attracted much attention. The complex underwater noise sources, the reverberation effects of the ocean surface and seabed, and the lack of prior information about the target make the classification problem difficult to solve. A popular solution is to improve the robustness of the features.
[0003] Feature extraction based on the principles of human hearing is one of the main research areas in passive sonar target recognition. The human ear can distinguish the voices of different speakers because the cochlea has a non-linear perception of frequency. Typical auditory features are extracted by designing a set of Mel filter banks (MFBs) that conform to the perceptual characteristics of the human ear, thereby extracting the envelope features of the signal spectrum. Numerous studies have shown that various vocalizations of marine organisms and the radiated noise of different types of ships have unique sound levels and spectral shapes; therefore, auditory features have been widely applied in the classification and recognition of underwater acoustic targets. For example, grouper, as a type of marine fish, emits low-frequency sounds between 50Hz and 350Hz during spawning; merchant ships, under normal navigation conditions, produce peaks in the continuous spectrum between 50Hz and 150Hz due to propeller cavitation. Simultaneously, the first-order difference coefficients (ΔMFCC) and second-order difference coefficients (ΔFCC) of the Mel frequency cepstral coefficients (MFCC) are also involved. 2 Multi-domain fusion (MFCC) can better represent the characteristics of a target in the time dimension. Real-world underwater acoustic target signals are usually nonlinear, non-Gaussian, and non-stationary, so it is difficult for any single feature to fully describe them. A common solution to this problem is to improve the robustness of the features through multi-domain fusion.
[0004] Currently, commonly used multi-domain fusion methods include feature concatenation methods and feature fusion methods. MFCC, ΔMFCC, Δ 2 MFCC can effectively describe the auditory characteristics of underwater targets in both the frequency and time domains. The feature concatenation method directly concatenates the three features. While this increases the learning information of the recognition algorithm, excessively high MFCC feature dimensionality can lead to a significant drop in recognition performance. In contrast, the feature fusion method uses a linear weighted summation to combine MFCC, ΔMFCC, and Δ... 2 MFCC and other features are merged into a new feature that incorporates the time-varying information of MFCC, reducing the computational complexity of the recognition algorithm. However, in existing methods, ΔMFCC and Δ... 2The weights in MFCC are primarily selected based on the recognition accuracy of the algorithm. However, this accuracy-based weight selection method does not consider the feature distribution of the samples, resulting in poor separability of the fused features. Furthermore, it is difficult to establish a direct correlation between the fused features and the recognition error rate. Therefore, the existing weight selection methods have limited effectiveness in improving recognition accuracy. Proposing a new weight selection method and feature fusion method is therefore essential. Summary of the Invention
[0005] The purpose of this invention is to address the problem that the fused features obtained based on existing weight selection methods have poor separability and limited effect on improving recognition accuracy, and to propose a new weighted Mel-frequency cepstral feature fusion method.
[0006] The technical solution adopted by the present invention to solve the above-mentioned technical problems is as follows:
[0007] A novel weighted Mel frequency cepstral feature fusion method, the method specifically includes the following steps:
[0008] Step 1: After performing frame segmentation on the radiated noise signal, window each frame signal obtained from the frame segmentation process to obtain the windowed signal for each frame.
[0009] Step 2: Perform Fourier transform on each windowed signal to obtain the windowed signals of each frame after Fourier transform;
[0010] Step 3: Set up the Mel filter bank, which includes L Mel filters. The L Mel filters divide the spectrum into L frequency bands.
[0011] Then, the windowed signals of each frame after Fourier transform are passed through the Mel filter bank to obtain the output energy of each windowed signal after Fourier transform through each Mel filter.
[0012] Step 4: For the windowed signal X of the m-th frame after Fourier transform... m (ω), for X m (ω) Take the logarithm of the energy output from each Mel filter, and use the logarithmic results to construct the logarithmic energy matrix E of the m-th frame signal. m ;
[0013] Similarly, the logarithmic energy matrix of each frame signal is obtained;
[0014] Step 5: Perform Discrete Cosine Transform on the logarithmic energy matrix of each frame of signal, and obtain the MFCC features of each frame of signal based on the Discrete Cosine Transform results.
[0015] Step 6: Perform differential operations on the MFCC features of each frame of signal to obtain the first-order differential features of each frame of signal; then perform differential operations on the first-order differential features of each frame of signal to obtain the second-order differential features of each frame of signal.
[0016] Step 7: The MFCC features of each frame signal are concatenated to obtain the MFCC features of the radiated noise signal; the first-order difference features of each frame signal are concatenated to obtain the first-order difference features of the radiated noise signal; and the second-order difference features of each frame signal are concatenated to obtain the second-order difference features of the radiated noise signal.
[0017] Step 8: Perform weighted fusion of the MFCC features, first-order difference features, and second-order difference features of the radiated noise signal to obtain the weighted Mel frequency cepstral features of the radiated noise signal.
[0018] Furthermore, the specific process of step one is as follows:
[0019] Step 11: Represent the radiated noise signal as x(t). After performing frame segmentation on x(t), the m-th frame signal is x_m. m (t), m=1,2,…,M, M is the number of frames, and the duration of each frame signal is T;
[0020] Step 22: Window each frame of the signal separately. The windowed signal for the m-th frame is then:
[0021]
[0022] in, Let h(t) be the windowed signal of the m-th frame, and h(t) be the Hamming window function, where m = 1, 2, ..., M.
[0023] Furthermore, the Fourier transform process of the windowed signal in the m-th frame is as follows:
[0024]
[0025] Among them, X m (ω) represents the windowed signal of the m-th frame after Fourier transform, e is the base of the natural logarithm, i is the imaginary unit, and ω is the frequency.
[0026] Furthermore, the Mel filter is a triangular filter.
[0027] Furthermore, the output energy of the windowed signal of the m-th frame after Fourier transform, after passing through the l-th Mel filter, is:
[0028]
[0029] Among them, e m,l Let g be the output energy of the windowed signal of the m-th frame after Fourier transform, passed through the l-th Mel filter. l(ω) is the l-th Mel filter, l = 0, 1, ..., L-1.
[0030] Furthermore, the statement regarding X m (ω) Take the logarithm of the energy output from each Mel filter, and use the logarithmic results to construct the logarithmic energy matrix E of the m-th frame signal. m Specifically:
[0031]
[0032] in, Indicates X m (ω) is the logarithm of the energy output after the l-th Mel filter, and lg is the logarithm to the base 10;
[0033] Then the logarithmic energy matrix E of the m-th frame signal m for:
[0034]
[0035] Furthermore, the specific process of step five is as follows:
[0036]
[0037]
[0038] Where u is the decorrelation coefficient after converting the output energy of the Mel filter into logarithmic energy, u = 0, 1, ..., L-1, E m (l) is E m The l-th element in F, c(u) is the orthogonalization factor; m (u) is the result of performing a discrete cosine transform on the logarithmic energy matrix of the m-th frame signal;
[0039] Take F m The first 13 bits of (u) are used as the MFCC feature of the m-th frame signal, that is, the MFCC feature of the m-th frame signal is
[0040] Furthermore, the step of performing a differential operation on the MFCC features of each frame signal to obtain the first-order differential features of each frame signal is as follows:
[0041]
[0042] in, Let N be the first-order differential feature of the m-th frame signal, where N is a constant. The MFCC features of the signal in frame m+n are... Let mn be the MFCC characteristics of the signal in frame mn;
[0043] The first-order difference features of each frame of signal are subjected to a difference operation to obtain the second-order difference features of each frame of signal; specifically:
[0044]
[0045] in, Let be the second-order difference feature of the m-th frame signal. The first-order differential feature of the signal in the (m+n)th frame is... This represents the first-order differential feature of the mn-th frame signal.
[0046] Furthermore, the constant N takes the value of 1 or 2.
[0047] Furthermore, the specific process of step eight is as follows:
[0048]
[0049] Among them, F WMFCC For the weighted Mel-frequency cepstral characteristics of the radiated noise signal, F MFCC For the MFCC characteristics of the radiated noise signal, F ΔMFCC The first-order difference feature of the radiated noise signal. The second-order difference feature of the radiated noise signal is represented by α1 and α2, which are the weights.
[0050] Furthermore, the weights α1 and α2 are calculated as follows:
[0051] Step 1: Extract the weighted Mel-frequency cepstral features of the radiated noise signal of target O1, then reduce the dimensionality of the weighted Mel-frequency cepstral features to one dimension through principal component analysis, and process the one-dimensional weighted Mel-frequency cepstral features using histogram statistics to obtain the j-th frame signal x in target O1. j The probability density is obtained by deriving the probability density function p(x) based on the probability density of each frame of the signal. j |O1,α1,α2);
[0052] Step 2: Extract the weighted Mel-frequency cepstral features of the radiated noise signal of target O2, then reduce the dimensionality of the weighted Mel-frequency cepstral features to one dimension through principal component analysis, and process the one-dimensional weighted Mel-frequency cepstral features using histogram statistics to obtain the j-th frame signal x′ in target O2. j The probability density is obtained by deriving the probability density function p(x′) based on the probability density of each frame of the signal. j |O2,α1,α2);
[0053] Step 3: The Bartoff distance J between the weighted Mel frequency cepstral features of the radiated noise signal of target O1 and the weighted Mel frequency cepstral features of the radiated noise signal of target O2. B for:
[0054]
[0055] definition
[0056] Then the weights α1 and α2 satisfy:
[0057]
[0058] The beneficial effects of this invention are:
[0059] This invention obtains first-order difference features by performing difference operations on MFCC features, and then obtains second-order difference features by performing difference operations on the first-order difference features. The MFCC features, first-order difference features, and second-order difference features are then linearly weighted to construct a fused feature with the same dimension as the MFCC features. The fused feature of this invention contains a large amount of information in the time dimension, thus exhibiting stronger robustness compared to MFCC features. Furthermore, by calculating the distance between the feature probability density distribution functions of two different target classes and determining the weights α1 and α2 by finding the maximum value of the feature distribution distance, the fused feature ensures maximized feature separability. Utilizing the fused feature can further improve the recognition accuracy. Attached Figure Description
[0060] Figure 1 This is a flowchart of the feature extraction process for weighted Mel frequency cepstral coefficients.
[0061] Figure 2a These are visualization results of the MFCC characteristics of two actual ship noise signals;
[0062] Figure 2b These are visualization results of the WMFCC characteristics of two actual ship noise signals;
[0063] Figure 3a These are the MFCC characteristic distributions of two actual ship noise signals;
[0064] Figure 3b These are the WMFCC characteristic distributions of two actual ship noise signals. Detailed Implementation
[0065] The present application will now be described in further detail with reference to specific embodiments and accompanying drawings. Obviously, the described embodiments are merely a part of the embodiments of the present invention, and not all of them. Other embodiments obtained by those skilled in the art based on the embodiments of the present invention without inventive effort are all within the scope of protection of the present invention.
[0066] Specific Implementation Method 1: Combination Figure 1This embodiment describes a novel weighted Mel frequency cepstral feature fusion method, which specifically includes the following steps:
[0067] Step 1: After performing frame segmentation on the radiated noise signal, window each frame signal obtained from the frame segmentation process to obtain the windowed signal for each frame.
[0068] The specific process of step one is as follows:
[0069] Step 11: Represent the radiated noise signal as x(t). After performing frame segmentation on x(t), the m-th frame signal is x_m. m (t), m=1,2,…,M, M is the number of frames, and the duration of each frame signal is T;
[0070] Step 22: Window each frame of the signal separately. The windowed signal for the m-th frame is then:
[0071]
[0072] in, Let h(t) be the windowed signal of the m-th frame, and h(t) be the Hamming window function, where m = 1, 2, ..., M.
[0073] Step 2: Perform Fourier Transform (FFT) on each windowed signal frame to obtain the windowed signals of each frame after Fourier Transform;
[0074] The Fourier transform process of the windowed signal in the m-th frame is as follows:
[0075]
[0076] Among them, X m (ω) represents the windowed signal of the m-th frame after Fourier transform, e is the base of the natural logarithm, i is the imaginary unit, and ω is the frequency.
[0077] The windowed signals of each frame are converted to the frequency domain using Fourier transform.
[0078] Step 3: Set up a Mel filter bank, which includes L Mel filters (triangular filters) that divide the spectrum into L frequency bands.
[0079] Then, the windowed signals of each frame after Fourier transform are passed through the Mel filter bank to obtain the output energy of each windowed signal after Fourier transform through each Mel filter.
[0080] The output energy of the windowed signal of the m-th frame after Fourier transform, after passing through the l-th Mel filter, is:
[0081]
[0082] Among them, e m,l Let g be the output energy of the windowed signal of the m-th frame after Fourier transform, passed through the l-th Mel filter. l (ω) is the l-th Mel filter, l = 0, 1, ..., L-1.
[0083] Step 4: For the windowed signal X of the m-th frame after Fourier transform... m (ω), for X m (ω) Take the logarithm of the energy output from each Mel filter, and use the logarithmic results to construct the logarithmic energy matrix E of the m-th frame signal. m ;
[0084] Similarly, the logarithmic energy matrix of each frame signal is obtained;
[0085] Specifically:
[0086]
[0087] in, Indicates X m (ω) is the logarithm of the energy output after the l-th Mel filter, and lg is the logarithm to the base 10;
[0088] Then the logarithmic energy matrix E of the m-th frame signal m for:
[0089]
[0090] This implementation method, by taking the logarithm, can better simulate the human ear's perception of sound.
[0091] Step 5: Perform Discrete Cosine Transform (DCT) on the logarithmic energy matrix of each frame of signal, and obtain the MFCC features of each frame of signal based on the DCT results.
[0092]
[0093]
[0094] Where u is the decorrelation coefficient after converting the output energy of the Mel filter into logarithmic energy, u = 0, 1, ..., L-1, E m (l) is E m The l-th element in F, c(u) is the orthogonalization factor; m (u) is the result of performing a discrete cosine transform on the logarithmic energy matrix of the m-th frame signal;
[0095] Take F m The first 13 bits of (u) are used as the MFCC feature of the m-th frame signal, that is, the MFCC feature of the m-th frame signal is
[0096] Step 6: Perform differential operations on the MFCC features of each frame of signal to obtain the first-order differential features of each frame of signal; then perform differential operations on the first-order differential features of each frame of signal to obtain the second-order differential features of each frame of signal.
[0097]
[0098] in, Let N be the first-order differential feature of the m-th frame signal, where N is a constant, taking the value 1 or 2. The MFCC features of the signal in frame m+n are... Let mn be the MFCC characteristics of the signal in frame mn;
[0099]
[0100] in, Let be the second-order difference feature of the m-th frame signal. The first-order differential feature of the signal in the (m+n)th frame is... This represents the first-order differential feature of the mn-th frame signal.
[0101] Step 7: The MFCC features of each frame signal are concatenated to obtain the MFCC features of the radiated noise signal; the first-order difference features of each frame signal are concatenated to obtain the first-order difference features of the radiated noise signal; and the second-order difference features of each frame signal are concatenated to obtain the second-order difference features of the radiated noise signal.
[0102] Step 8: Perform weighted fusion of the MFCC features, first-order difference features, and second-order difference features of the radiated noise signal to obtain the weighted Mel frequency cepstral features (WMFCC) of the radiated noise signal.
[0103]
[0104] Among them, F WMFCC For the weighted Mel-frequency cepstral characteristics of the radiated noise signal, F MFCC For the MFCC characteristics of the radiated noise signal, F ΔMFCC The first-order difference feature of the radiated noise signal. The second-order difference feature of the radiated noise signal is represented by α1 and α2, which are the weights.
[0105] This invention can construct a WMFCC-based recognition model using machine learning algorithms, and the recognition performance of this model is better than that of a recognition model that learns MFCC.
[0106] The values of α1 and α2 affect the feature quality of WMFCC. An important criterion for feature quality evaluation is feature distance, a measure of feature separability. Bach distance, as a measure of similarity between samples of different classes for the same feature, fully reflects the statistical properties of features. A smaller Bach distance indicates poor feature separability; conversely, a larger distance indicates better feature separability. The weight calculation method of this invention is as follows:
[0107] Step 1: Extract the weighted Mel-frequency cepstral features of the radiated noise signal of target O1, then reduce the dimensionality of the weighted Mel-frequency cepstral features to one dimension using principal component analysis (PCA), and process the one-dimensional weighted Mel-frequency cepstral features using histogram statistics (HIS) to obtain the j-th frame signal x in target O1. j The probability density (PDF) is obtained from the probability density of each frame signal, and the probability density function p(x) is derived from the probability density of each frame signal. j |O1,α1,α2);
[0108] Step 2: Extract the weighted Mel-frequency cepstral features of the radiated noise signal of target O2, then reduce the dimensionality of the weighted Mel-frequency cepstral features to one dimension using principal component analysis (PCA), and process the one-dimensional weighted Mel-frequency cepstral features using histogram statistics (HIS) to obtain the j-th frame signal x′ in target O2. j The probability density (PDF) is obtained from the probability density of each frame signal, and the probability density function p(x′) is derived from it. j |O2,α1,α2);
[0109] Step 3: The Bartoff distance J between the weighted Mel frequency cepstral features of the radiated noise signal of target O1 and the weighted Mel frequency cepstral features of the radiated noise signal of target O2. B for:
[0110]
[0111] definition When p(x) j |O1,α1,α2) and p(x j When '|O2,α1,α2) completely overlap, f(α1,α2)=1, that is, J B =0; as the overlapping region gradually decreases, f(α1,α2)→0, that is, J B →∞. An important criterion for feature quality assessment is feature distance, a measure of feature separability. Bach distance, as a measure of similarity between samples of different classes for the same feature, fully reflects the statistical properties of features. When the Bach distance is small, the feature's separability is poor; conversely, the feature's separability is good. Therefore, to ensure feature separability, a high Bach distance is needed. B →∞, which is equivalent to It has a minimum value;
[0112] Then the weights α1 and α2 satisfy:
[0113]
[0114] The above equation is an unconstrained convex quadratic optimization problem, which can be solved using a linear random search algorithm to obtain the weight vector corresponding to the global minimum.
[0115] Experimental Section
[0116] Using the publicly available "ShipsEar" dataset, radiated noise audio from "Adventure of the Seas" and "Mar de Cangas" was selected. "Adventure of the Seas" refers to ocean liners, while "Mar de Cangas" refers to ferries. Recordings for "Adventure of the Seas" and "Mar de Cangas" were taken for 244 seconds and 405 seconds respectively. One second of the effective signal time was extracted as a single sample (one frame), yielding 202 and 313 samples respectively, representing 39.2% and 60.8% of the total samples. During feature extraction, the Hamming window length was set to 1024 points with 50% overlap, and 40 triangular filters were used.
[0117] Feature visualization and feature distribution are two typical feature analysis methods. Figure 2a and Figure 2b Feature visualization results obtained using the t-SNE algorithm for two types of targets are presented. From Figure 2a and Figure 2b The results show that the spatial overlap of WMFCCs in the feature space is significantly smaller than that of MFCCs. Figure 3a and Figure 3b The feature distributions of the two target classes in the one-dimensional case are also given, and the overlap area of WMFCC is smaller than that of MFCC. Table 1 shows the two feature distances of WMFCC and MFCC.
[0118] Table 1. Two characteristic distances for MFCC and WMFCC
[0119]
[0120] As can be seen, WMFCC outperforms MFCC in both feature distances. Therefore, both feature analysis results demonstrate that the feature separability of the WMFCC of this invention is superior to that of the traditional feature MFCC.
[0121] Table 2 shows the identification results for the two types of targets.
[0122] Table 2 Identification results of two actual ship noise signals
[0123]
[0124] Table 2 shows that, under the same recognition algorithm, the recognition rate of WMFCC is 0.6% to 12.5% higher than that of MFVV. On the other hand, KNN has the best recognition rate under the same features. Therefore, the recognition results show that WMFCC performs better than traditional feature MFCC.
[0125] The above examples of the present invention are merely illustrative of the computational model and process of the present invention, and are not intended to limit the implementation of the present invention. Those skilled in the art will recognize that other variations or modifications can be made based on the above description. It is impossible to exhaustively list all possible implementations here. Any obvious variations or modifications derived from the technical solutions of the present invention are still within the scope of protection of the present invention.
Claims
1. A novel weighted Mel-frequency cepstral feature fusion method, characterized in that, The method specifically includes the following steps: Step 1: After performing frame segmentation on the radiated noise signal, window each frame signal obtained from the frame segmentation process to obtain the windowed signal for each frame. Step 2: Perform Fourier transform on each windowed signal to obtain the windowed signals of each frame after Fourier transform; Step 3: Configure the Mel filter bank, which includes... A Mel filter, A Mel filter divides the spectrum into One frequency band; Then, the windowed signals of each frame after Fourier transform are passed through the Mel filter bank to obtain the output energy of each windowed signal after Fourier transform through each Mel filter. Step 4: For the fourth step after the Fourier transform... Frame windowing signal ,right Taking the logarithm of the energy output from each Mel filter, and using the logarithm results to construct the first... Logarithmic energy matrix of frame signal ; Similarly, the logarithmic energy matrix of each frame signal is obtained; Step 5: Perform Discrete Cosine Transform on the logarithmic energy matrix of each frame of signal, and obtain the MFCC features of each frame of signal based on the Discrete Cosine Transform results. Step 6: Perform differential operations on the MFCC features of each frame of signal to obtain the first-order differential features of each frame of signal; then perform differential operations on the first-order differential features of each frame of signal to obtain the second-order differential features of each frame of signal. Step 7: The MFCC features of each frame signal are concatenated to obtain the MFCC features of the radiated noise signal; the first-order difference features of each frame signal are concatenated to obtain the first-order difference features of the radiated noise signal; and the second-order difference features of each frame signal are concatenated to obtain the second-order difference features of the radiated noise signal. Step 8: Weighted fusion of the MFCC features, first-order difference features, and second-order difference features of the radiated noise signal to obtain the weighted Mel frequency cepstral features of the radiated noise signal. The specific process of step eight is as follows: in, The weighted Mel-frequency cepstral characteristics of the radiated noise signal. The MFCC characteristics of the radiated noise signal, The first-order difference feature of the radiated noise signal. This represents the second-order difference characteristic of the radiated noise signal. and For weights; The weight and The calculation method is as follows: Step 1: Extract the target The weighted Mel-frequency cepstral features of the radiated noise signal are obtained, and then principal component analysis is used to reduce the dimensionality of the weighted Mel-frequency cepstral features to one dimension. Histogram statistics are then used to process the one-dimensional weighted Mel-frequency cepstral features to obtain the target... The first in Frame signal The probability density is obtained by deriving the probability density function from the probability density of each frame of the signal. ; Step 2: Extract the target The weighted Mel-frequency cepstral features of the radiated noise signal are obtained, and then principal component analysis is used to reduce the dimensionality of the weighted Mel-frequency cepstral features to one dimension. Histogram statistics are then used to process the one-dimensional weighted Mel-frequency cepstral features to obtain the target... The first in Frame signal The probability density is obtained by deriving the probability density function from the probability density of each frame of the signal. ; Step 3, Objective Weighted Mel-frequency cepstral characteristics of the radiated noise signal and the target Bach distance of the weighted Mel frequency cepstral feature of the radiated noise signal for: definition ; Then weight and satisfy: 。 2. The novel weighted Mel-frequency cepstral feature fusion method according to claim 1, characterized in that, The specific process of step one is as follows: Step 11: Represent the radiated noise signal as ,right The first frame obtained after frame segmentation Frame signal is , , The number of frames is [number], and the duration of each frame is [duration]. ; Step 22: Window each frame of signal separately, then the... The frame windowing signal is: in, For the first Frame windowing signal, For Hamming window functions, .
3. A novel weighted Mel-frequency cepstral feature fusion method according to claim 2, characterized in that, The first The Fourier transform process of the windowed frame signal is as follows: in, The th Fourier transform Frame windowing signal, is the base of the natural logarithm. The imaginary unit, For frequency.
4. A novel weighted Mel-frequency cepstral feature fusion method according to claim 3, characterized in that, The Fourier transform of the first Frame windowing signal after the first The output energy of the Mel filter is: in, The th Fourier transform Frame windowing signal after the first The output energy of a Mel filter For the first A Mel filter, .
5. A novel weighted Mel-frequency cepstral feature fusion method according to claim 4, characterized in that, The pair Taking the logarithm of the energy output from each Mel filter, and using the logarithm results to construct the first... Logarithmic energy matrix of frame signal Specifically: in, Indicates to After the first The energy of the Mel filter output is the logarithm of the result. It is a logarithm with base 10; Then the first Logarithmic energy matrix of frame signal for: 。 6. A novel weighted Mel-frequency cepstral feature fusion method according to claim 5, characterized in that, The specific process of step five is as follows: in, It is the decorrelation coefficient obtained by converting the output energy of the Mel filter into logarithmic energy. , yes The first in One element, It is the orthogonalization factor; For the first The result of performing a discrete cosine transform on the logarithmic energy matrix of the frame signal; Pick The first 13 as the number The MFCC characteristics of the frame signal, i.e., the first The MFCC characteristics of the frame signal are .
7. A novel weighted Mel-frequency cepstral feature fusion method according to claim 6, characterized in that, The MFCC features of each frame of signal are differentially processed to obtain the first-order differential features of each frame of signal. Specifically: in, For the first First-order difference characteristics of frame signals, It is a constant. For the first MFCC characteristics of frame signals For the first MFCC characteristics of frame signals; The first-order difference features of each frame of signal are subjected to a difference operation to obtain the second-order difference features of each frame of signal; specifically: in, For the first Second-order difference characteristics of frame signals, For the first First-order difference characteristics of frame signals, For the first First-order difference characteristics of frame signals.
8. A novel weighted Mel-frequency cepstral feature fusion method according to claim 7, characterized in that, The constant The value can be 1 or 2.
Citation Information
Patent Citations
Construction site special vehicle identification method based on improved MFCC
CN112927716A