Xylem cavitation ultrasonic signal purification and MFCC feature improvement extraction method
Through the improved Mel filter bank formula, inverse discrete cosine transformation, F-ratio formula and PCA linear dimensionality reduction, the xylem cavitation signal characteristics with strong characterization capabilities were extracted, solving the problems of noise removal and feature extraction, and achieving efficient waveform recognition and denoising effects.
Patent Information
- Application Number
- CN202510303632.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-14
- Publication Date
- 2025-07-01
AI Technical Summary
The prior art is difficult to effectively remove noise in xylem cavitation signals, and the existing MFCC feature extraction methods are not suitable for most waveform signals.
The xylem cavitation ultrasonic signal purification and MFCC feature improvement extraction methods are proposed, including improved Mel filter bank formula, inverse discrete cosine transformation, F ratio formula and PCA linear dimensionality reduction to extract feature vectors with strong characterization capabilities.
It realizes efficient feature extraction and denoising effect of xylem cavitation signal, and improves the recognition performance and training efficiency of neural network models.
Smart Images

Figure CN120234602A_ABST
Abstract
Description
Technical Field
[0001] The present invention discloses a method for purifying xylem cavitation ultrasonic signals and improving the extraction of MFCC features, belonging to the fields of deep learning and feature extraction. Background Art
[0002] Most woody plants transport internal water through xylem vessels. When the water potential in the xylem vessels drops to a certain threshold, nanoscale bubbles in the xylem vessels expand and fill the entire vessel, which is the phenomenon of xylem cavitation. During the cavitation process of plants, ultrasonic signals are radiated, which has a direct connection with the water use situation of plants. Therefore, studying and analyzing the waveform signal characteristics generated by xylem cavitation is one of the main research approaches to explore plant cavitation, and it has important guiding significance for analyzing the water use mechanism, water demand characteristics, and drought tolerance of plants, and is an important research field for coping with the ecological system damage caused by climate change.
[0003] The processes of xylem cavitation and embolism repair are accompanied by ultrasonic signal radiation. Current non-destructive observation methods, such as CT, nuclear magnetic resonance, and even synchrotron radiation imaging, are still unable to directly observe the cavitation phenomenon. Therefore, studying the ultrasonic signals generated by xylem cavitation is one of the main research approaches to explore plant cavitation. The xylem vessels are very thin, and the ultrasonic emission during the cavitation process of the xylem sap inside is a microscopic process, which is numerous but very weak. And the detection system for collecting signals inevitably has ADDIN such as thermal noise and electromagnetic noise, which results in a large amount of noise pollution in the collected cavitation signals even after using various noise reduction measures and signal amplitude discrimination systems. Therefore, it is very important to use advanced neural network technology for waveform recognition to remove the noise signals in the acoustic radiation of plant cavitation.
[0004] During the deep learning training process, feature extraction is one of the key steps determining the performance of the model. Through effective feature extraction, the model can more accurately capture the important information in the data, thereby improving the representation ability of the input data during the subsequent training process. The quality and representation of features directly affect the training speed, generalization ability, and final classification or regression accuracy of the network. Inappropriate or insufficient feature extraction may lead to the model being unable to fully utilize the potential patterns in the data, increasing the risk of overfitting. Therefore, a suitable feature extraction method is crucial for constructing a robust and efficient neural network model.
[0005] The main methods for feature extraction of audio samples include MFCC (Mel Frequency Cepstral Coefficient), Fbank, Spectrogram, etc. Among them, MFCC is one of the most commonly used audio preprocessing methods in voiceprint recognition. However, its applicable frequency spectrum range is narrow and it is not suitable for feature extraction of most waveform signals. Therefore, for different recognition targets and training tasks, researchers usually need to adjust and improve the audio preprocessing method to match different recognition tasks and achieve an ideal training effect as much as possible. Summary of the Invention
[0006] Aiming at the problems existing in the prior art, the present invention proposes a method for purifying xylem cavitation ultrasonic signals and improving MFCC features extraction, which realizes the efficient feature extraction of xylem cavitation ultrasonic signals and is further applied to the training of neural network models, thereby realizing efficient and intelligent waveform recognition and denoising.
[0007] The present invention adopts the following technical solutions:
[0008] A method for purifying xylem cavitation ultrasonic signals and improving MFCC features extraction includes proposing three groups of improved Mel filter bank formulas applicable to the frequency range of the cavitation ultrasonic signals after frequency reduction; introducing the inverse discrete cosine transform processing after extracting MFCC features; using the F-ratio formula to rank the contribution degrees of feature dimensions and selecting the top 100-dimensional feature vectors; using PCA for linear dimensionality reduction to 40 dimensions.
[0009] The specific implementation process of the improved Mel filter bank formula is as follows:
[0010] Since a large part of the frequency domain distribution range of the Robinia pseudoacacia cavitation signal still exceeds the upper limit of the frequency band processing of the IMFCC filter bank, and its distribution in the frequency domain range is complex and the current filter bank is sparse in most of the cavitation signal frequency bands, therefore, combining the ideas of MFCC, IMFCC, and Mid-MFCC, three groups of filter bank formulas matching the frequency domain and distribution of the cavitation signal are proposed as follows:
[0011]
[0012] In the formula, F mel represents the Mel frequency domain scale, and f is the frequency in Hz.
[0013] The specific implementation process of the inverse discrete cosine transform is as follows:
[0014] After MFCC feature extraction, inverse discrete cosine transform (IDCT) processing is performed. Compared with directly removing DCT, IDCT can restore some lost high-frequency information while maintaining the simplicity of feature extraction and the noise reduction effect, increase spectral details, making the model more sensitive to signal features. In addition, the features processed by IDCT have a higher correlation with the original audio spectrum.
[0015] The specific implementation process of the F-ratio formula is as follows:
[0016]
[0017] In the formula: F Fisher is the Fisher ratio of each dimension of the feature parameter. The larger its value, the better the separability of a certain dimension. σ between is the between-class variance of the feature component, indicating its degree of dispersion. Its expression is as shown in (10); σ within is the within-class variance of the feature component, indicating the degree of aggregation of the features. Its expression is as shown in (11).
[0018]
[0019]
[0020] In the formula, k represents the dimension of the feature parameter, k = 1, 2,..., 40; M represents the number of categories of the speech feature sequence; m k represents the mean of the k-th component of the speech feature over all classes; m k (i) represents the mean of the k-th component of the i-th class of the speech feature; c k (i) represents the k-th component of the i-th class of the speech feature sequence; ω i represents the speech feature sequence of the i-th class; n i represents the number of samples of each class.
[0021] Fisher criterion ratio calculations are performed on the 120-dimensional feature vectors of single-frame xylem ultrasonic signals respectively, and they are sorted according to their contribution degrees. The first 40, 60, 80, 90, 100, and 110 dimensions are selected for experiments respectively. According to the evaluation results, the first 100 dimensions with the best performance are selected for further processing.
[0022] The specific implementation process of principal component analysis linear dimensionality reduction is as follows:
[0023] PCA achieves dimensionality reduction by finding several projection directions such that the variance of the high-dimensional feature vectors projected onto these directions is the largest and the projection directions are uncorrelated with each other.
[0024] Suppose there are m ultrasonic signal samples and their respective n voiceprint recognition eigenvalues. Denote the j-th eigenvalue of the i-th signal as x ij , and then the feature quantity matrix X = (x ij ) m×n can be constructed, where the rows represent the number of samples and the columns represent the dimensions.
[0025] Then the steps of using PCA to reduce the dimension of the feature vectors are as follows:
[0026] (1) Perform the de-centralization processing on each column of the matrix X, that is, subtract the mean p of the column from each element x ij of the column. j
[0027]
[0028] (2) Use the new matrix Y to calculate the covariance matrix representing the relationship between each feature dimension
[0029]
[0030] (3) Calculate the eigenvalues of the covariance matrix Z and their corresponding unit eigenvectors, and denote the eigenvalues as λ1, λ2 ……, λ n , and the unit eigenvectors as α1, α2 ……, α n .
[0031] (4) Sort the eigenvalues in descending order of magnitude, and at the same time arrange their corresponding eigenvectors in the same order. Determine the number of principal components by calculating the cumulative contribution rate of the principal components. The cumulative contribution rate formula is:
[0032]
[0033] In the formula, λ q is the q-th eigenvalue in the new sorting, and ε k is the cumulative contribution rate of the first k principal components. When ε k reaches 85%, the number of principal components is determined to be k, and at the same time, this also determines that the dimension of the new feature quantity after dimension reduction is k.
[0034] (5) Take the first k rows of the eigenvectors in the new sorting to form the matrix R = [r1, r2 ……, r k , and calculate the new feature quantity matrix W:
[0035] W = XR#(16)
[0036] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0037] (1) By improving the MFCC feature extraction, this method extracts the feature vectors of the locust cavitation signal with strong characterization ability.
[0038] (2) The feature vectors extracted by this method are applied to an excellent voice-based voiceprint recognition model for training, and good recognition performance and denoising effects are achieved. Brief Description of the Drawings
[0039] Figure 1 is the overall flowchart of the present invention
[0040] Figure 2 is F in the filter bank mel1 corresponding relationship curve and filter bank distribution
[0041] Figure 3 is F in the filter bank mel2 corresponding relationship curve and filter bank distribution
[0042] Figure 4 is F in the filter bank mel3 corresponding relationship curve and filter bank distribution
[0043] Figure 5 is the signal diagram collected during the cavitation process of Robinia pseudoacacia
[0044] Figure 6 is the noise diagram collected during the cavitation process of Robinia pseudoacacia Detailed Implementation Manner
[0045] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention; obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0046] STEP1. The specific implementation process of improving the Mel filter bank formula is as follows: Since the frequency domain distribution range of the plant cavitation signal still has a large part exceeding the frequency band processing upper limit of the IMFCC filter bank, and its distribution in the frequency domain is complex and the current filter bank is sparse in most of the cavitation signal frequency bands, therefore, combining the ideas of MFCC, IMFCC, and Mid-MFCC, three groups of filter bank formulas matching the frequency domain and distribution of the cavitation signal are proposed as follows:
[0047]
[0048] In the formula, F mel represents the Mel frequency domain scale, and f is the frequency in Hz.
[0049] STEP 2. Inverse Discrete Cosine Transform: IDCT processing is adopted to improve the feature representation ability while maintaining the efficient use of computing resources. For xylem ultrasonic signals, due to the high sampling frequency and the large proportion of high-frequency components in the effective signals, their spectral distributions are complex and the correlations between different frequency components are strong. Only using DCT compression will result in the loss of important spectral details, thereby reducing the quality of the extracted features. Completely removing DCT processing will retain too much redundant information, leading to an increase in feature dimensions, further increasing the computational cost and model complexity, and reducing the overall efficiency. Therefore, performing inverse discrete cosine transform (IDCT) processing after MFCC feature extraction has become an improved method. Compared with directly removing DCT, IDCT can restore some lost high-frequency information, increase spectral details, make the model more sensitive to signal features, and maintain the simplicity of feature extraction and the noise reduction effect. In addition, the features processed by IDCT have a higher correlation with the original audio spectrum.
[0050] STEP 3. F-ratio formula: Sort the total 120-dimensional features generated by the three groups of filter banks according to the contribution degree, and then select a certain number of dimensions according to the sorting results for PCA data fusion to reduce the dimension to 40. The idea of improving the MFCC to enrich the filter bank distribution has greatly increased the feature dimensions. After the improved MFCC, a total of 120-dimensional feature vectors are obtained. If they are directly superimposed, there will be a large amount of duplicate information, increasing the computational complexity of training and the recognition time. In addition, the contribution of each-dimensional feature parameter to recognition is different. Some parameters may contain less information, and some may contain redundant information. If they are directly superimposed and treated equally, it will ultimately have a negative impact on the recognition performance. Therefore, it is necessary to evaluate the influence degree of each-dimensional parameter on the recognition result, select the parameters that contribute the most to the recognition effect, and then combine the three groups of features as new feature parameters. Based on the Fisher criterion, the F-ratio formula for measuring the effectiveness of speaker feature parameters is proposed:
[0051]
[0052] The F-ratio formula can sort the total 120-dimensional features generated by the three groups of filter banks according to the contribution degree, and then select a certain number of dimensions according to the sorting results for PCA data fusion to reduce the dimension to 40. However, due to the different numbers of selected dimensions, the feature richness and information content of the finally input data are also different, resulting in different model training results. To determine the optimal number of dimensions for xylem ultrasonic signal feature extraction and model training, on the basis of improving the filter bank and introducing the IDCT module, model training comparisons were carried out for different selected numbers of dimensions.
[0053] STEP 4. Linear dimensionality reduction by principal component analysis: The dimension range of the conventional input of the speech recognition neural network model that selects MFCC for feature extraction is 12 - 40 dimensions. If the input dimension for model training is too high, it may lead to unstable training, overfitting of the model, or excessive consumption of training resources (such as computing, memory, etc.). Therefore, it is necessary to perform fusion dimensionality reduction processing on the selected multi-dimensional data to 40 dimensions. Since there is overlap between the filter banks of the improved MFCC, the extracted feature dimensions can be regarded as linearly correlated. Therefore, it is decided to use a linear dimensionality reduction algorithm. Commonly used dimensionality reduction algorithms include Principal Component Analysis (PCA) and Linear Discriminant Analysis (LDA). Although Linear Discriminant Analysis is applicable to supervised learning and classification problems, for data containing k categories, LDA can at most reduce the data to k - 1 dimensions. When the classification label has only 2 types, LDA can only reduce the data to 1 dimension. Therefore, the principal component analysis method is used as the dimensionality reduction algorithm.
[0054] As described above, it is only the preferred specific implementation manner of the present invention; however, the protection scope of the present invention is not limited thereto.
[0055] Any person skilled in the art within the technical scope disclosed by the present invention should cover any equivalent replacement or change made according to the technical solution of the present invention and its improved concept within the protection scope of the present invention.
Claims
1. Xylem cavitation ultrasonic signal purification and MFCC feature improvement extraction method, characterized in that: The following steps are involved: S1. Combining the ideas of MFCC, IMFCC and Mid-MFCC, three sets of improved Mel filter bank formulas suitable for the frequency range of cavitation ultrasonic signals after frequency reduction are proposed. S2. Considering the balance between algorithm effect and storage overhead, inverse discrete cosine transform processing is introduced after extracting MFCC features to restore some lost high-frequency information and increase spectrum details while maintaining efficient utilization of computing resources. S3. Use the F ratio formula to sort the contribution of feature dimensions, and select the top 100 feature vectors based on the sorting results. S4. In order to reduce system complexity and time consumption, the principal component analysis (PCA) method is used to reduce the first 100 dimensions to 40 dimensions.
2. The method for purifying xylem cavitation ultrasonic signals and improving MFCC feature extraction according to claim 1, characterized in that The width of the improved filter bank changes more smoothly, which is more consistent with the frequency domain range of the xylem cavitation signal and more suitable for the feature extraction of the xylem ultrasonic cavitation frequency band. Since the distribution range of the Robinia pseudoacacia cavitation signal in the frequency domain still exceeds the upper limit of the frequency band processing of the IMFCC filter bank, and the distribution in its frequency domain is complex, and the current filter bank is sparsely distributed in most of the cavitation signal frequency bands, the ideas of MFCC, IMFCC and Mid-MFCC are combined to propose three sets of filter bank formulas suitable for the frequency domain and distribution of cavitation signals as follows: Where F mel represents the Mel frequency domain scale, and f is the frequency in Hz.
3. The method for purifying xylem cavitation ultrasonic signals and improving MFCC feature extraction according to claim 2, characterized in that: After extracting MFCC features, IDCT processing is used to improve the feature representation capability while maintaining efficient use of computing resources. For xylem ultrasonic signals, due to the high sampling frequency and the large proportion of high-frequency components of effective signals, the spectrum distribution is complex and the correlation between different frequency components is strong. Using only discrete cosine transform (DCT) compression will lead to the loss of important spectrum details, thereby reducing the quality of the extracted features; while completely removing DCT processing will retain too much redundant information, resulting in an increase in feature dimensions, thereby increasing computational overhead and model complexity, and reducing overall efficiency. For this reason, inverse discrete cosine transform (IDCT) processing after MFCC feature extraction has become an improved method. Compared with directly removing DCT, IDCT can restore some of the lost high-frequency information while maintaining the simplicity of feature extraction and noise reduction effect, increase spectrum details, and make the model more sensitive to signal features. In addition, the features processed by IDCT have a higher correlation with the original audio spectrum.
4. The method for purifying xylem cavitation ultrasonic signals and improving MFCC feature extraction according to claim 3, characterized in that: The F ratio formula can sort the 120-dimensional features generated by the three filter groups by contribution, and then select a certain number of dimensions based on the sorting results to perform PCA data fusion to reduce the dimensions to 40. However, due to the different number of dimensions selected, the feature richness and information content of the final input data are also different, resulting in different model training results. Through model training comparison, we determined that the fusion of the first 100-dimensional feature vectors can achieve the best model training performance.
5. The method for purifying xylem cavitation ultrasonic signals and improving MFCC feature extraction according to claim 4, characterized in that: The conventional input dimension range of the speech recognition neural network model using MFCC as feature extraction is 12-40 dimensions. Too high an input dimension for model training may lead to unstable training, overfitting of the model, or excessive consumption of training resources (computation, memory, etc.). Therefore, it is necessary to fuse and reduce the selected multi-dimensional data to 40 dimensions. Since there is overlap between the filter groups of the improved MFCC, the extracted feature dimensions can be regarded as linearly correlated, so it was decided to use the principal component analysis (PCA) dimensionality reduction algorithm.