A bone conduction speech conversion method based on spectral envelope mapping

By combining a spectral envelope mapping method and a GRU neural network with timbre weighted filters and codebook search, the problem of bone-guided speech conversion in high-frequency missing and noisy environments is solved, achieving high-quality air-conducted speech generation, which is suitable for steel, coal mining and power industry scenarios.

CN116403596BActive Publication Date: 2026-04-24DALIAN UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
DALIAN UNIV OF TECH
Filing Date
2023-03-03
Publication Date
2026-04-24

AI Technical Summary

Technical Problem

Existing bone conduction speech conversion technology suffers from problems such as high frequency missing, poor speech intelligibility, low clarity, and inability to convert in real time under strong noise, especially in steel, coal mining, and power industries where it cannot meet the demand for high-quality speech.

Method used

A spectral envelope mapping-based method is adopted. By preprocessing and linear prediction analysis of bone conduction speech signals, combined with GRU neural network mapping of LSF coefficients of air conduction speech signals, and using timbre weighted filter and codebook search method to optimize excitation signal, a high-quality air conduction speech signal is finally generated through post-filter.

Benefits of technology

It improves the accuracy and robustness of bone conduction speech conversion, enabling real-time conversion to high-quality air conduction speech in noisy environments, reducing invalid noise interference, and restoring mid-to-high frequency information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116403596B_ABST
    Figure CN116403596B_ABST
Patent Text Reader

Abstract

The application discloses a bone conduction speech conversion method based on spectral envelope mapping, comprising the following steps: pre-processing and linear prediction analysis of the bone conduction speech signal, and calculating the LP filter coefficient; mapping the LSF coefficient of the air conduction speech signal corresponding to the bone conduction speech signal by using the trained neural network; controlling the minimum error of the original bone conduction speech signal and the synthesized air conduction speech signal according to the timbre weighting characteristic of the bone conduction speech characteristic; estimating the integer pitch of the bone conduction speech signal through the timbre weighting filter, obtaining the reference signal by passing the linear prediction residual signal through the timbre weighting filter, estimating the fractional pitch to obtain the adaptive codebook vector; obtaining the new reference signal by subtracting the adaptive codebook vector from the reference signal, searching for the optimal excitation in the fixed codebook; synthesizing the air conduction speech signal by combining the optimal excitation and the LP filter of the air conduction speech signal, and correcting the air conduction speech signal.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to bone conduction speech conversion methods, and more particularly to a bone conduction speech conversion method based on spectral envelope mapping. Background Technology

[0002] Bone-guided speech, formed by the bones of the human body, can reduce the intake of ambient noise. However, bone-guided speech sounds muffled, suffers from severe high-frequency loss, has poor intelligibility, and low clarity, falling far short of air-guided speech. Bone-guided speech conversion is a technique that converts bone-guided speech into air-guided speech. This technology can compensate for the high-frequency loss in bone-guided speech and has promising prospects and applications. Existing technologies propose reconstruction filters to achieve bone-guided speech conversion, constructing the filter by extracting the amplitude spectra of bone-guided and air-guided speech separately. However, its effectiveness is poor. Existing technologies use LP models to convert the spectral envelope of bone-guided speech into the spectral envelope of air-guided speech, combining it with the original bone-guided excitation signal to synthesize speech. However, this conversion method is too simplistic and lacks accuracy.

[0003] In recent years, deep neural networks have yielded many achievements in speech signal processing, including bone-guided speech conversion. For example, existing technologies use fully connected networks to map cepstral coefficients between bone-guided and air-guided speech; existing technologies use multilayer perceptrons to map LSF coefficients. However, these mapping networks are relatively simple, do not consider the temporal characteristics of speech, and do not take into account the influence of excitation, but directly use bone-guided excitation.

[0004] Current technologies employ WaveNet vocoders for bone-guided speech conversion. These technologies decompose speech into three features: fundamental frequency, spectral envelope, and aperiodic ratio, and then use a Gaussian mixture model to map these features. While deep learning-based methods achieve good results, they suffer from drawbacks such as excessive computation, inability to perform real-time conversion, and limited commercial viability. In noisy environments such as steel mills, coal mines, and power plants, high-quality speech is required, and current speech conversion technologies largely fail to meet these needs. Summary of the Invention

[0005] To address the problems existing in the prior art, this invention discloses a bone-guided speech conversion method based on spectral envelope mapping, specifically including the following steps:

[0006] Preprocessing and linear prediction analysis are performed on bone conduction speech signals. LP filter coefficients are calculated, converted into line spectrum pair LSP coefficients, and then converted into line spectrum frequency LSF coefficients.

[0007] The trained neural network is used to map the LSF coefficients of the air conduction speech signal corresponding to the bone conduction speech signal, and the LP filter coefficients of the air conduction speech signal are obtained based on the LSF coefficients of the air conduction speech signal.

[0008] The optimal excitation signal is searched using a comprehensive analysis method, and the error between the original bone conduction speech signal and the synthesized air conduction speech signal is minimized based on the timbre weighting characteristics of bone conduction speech.

[0009] The bone conduction speech signal is passed through a timbre-weighted filter to estimate the integer pitch, the linear prediction residual signal is passed through the timbre-weighted filter to obtain the reference signal, and the fractional pitch is estimated to obtain the adaptive codebook vector.

[0010] Subtract the adaptive codebook vector from the reference signal to obtain a new reference signal, and search for the optimal excitation in the fixed codebook;

[0011] The optimal excitation is combined with an LP filter of the air-conducted speech signal to synthesize an air-conducted speech signal, and the air-conducted speech signal is then corrected.

[0012] Furthermore, when using a neural network to map the LSF coefficients of the air conduction speech signal corresponding to the bone conduction speech signal: during the training phase, the LSF coefficients of the bone conduction speech signal and the air conduction speech signal at the same time are calculated, fed into the neural network for training, and the weights after training are saved.

[0013] During the mapping phase, the LSF coefficients of the input 10-dimensional bone conduction speech signal are... The forward operation is performed layer by layer in the following manner:

[0014]

[0015]

[0016]

[0017]

[0018]

[0019] Output To obtain the LSF coefficients of the 10-dimensional airborne speech signal obtained through mapping, the LSF coefficients of the airborne speech signal are corrected to ensure that the 10th-order coefficients satisfy... ;

[0020] The LSP coefficients of the air conduction speech signal are obtained from the LSF coefficients of the mapped air conduction speech signal, and then weighted with the LSP coefficients of the bone conduction speech signal to obtain the LSP coefficients of the final synthesized air conduction speech signal.

[0021]

[0022] in

[0023] Then, the LP filter for the air-conducted speech signal is calculated based on these LSP coefficients. and linear prediction coefficient

[0024] To minimize the error between the original bone conduction speech signal and the synthesized air conduction speech signal based on the timbre weighting characteristics of bone conduction speech, the timbre weighting filter is defined as follows:

[0025]

[0026] The bone conduction signal is weighted by timbre, with a higher weight for the low-frequency component and a lower weight for the high-frequency component. The linear prediction coefficients of the original bone conduction speech signal were combined. The linear prediction coefficients of the air-conducted speech signal obtained by mapping Perform timbre weighting.

[0027] When synthesizing the airborne speech signal by combining the optimal excitation with an LP filter of the airborne speech signal, and then correcting the airborne speech signal, a codebook search method is used to search for the optimal excitation signal. ;

[0028]

[0029] in, For adaptive codebook vectors, For adaptive codebook gain, For fixed codebook vectors, To maintain a fixed codebook gain, the optimal excitation signal will be... LP filters that transmit speech signals through the air , obtain air-conducted speech signal ;

[0030] The air-conducted speech signal is corrected by passing it through a bandpass filter with a passband of 1.5–2.5 kHz and a low-pass filter with a cutoff frequency of 3.5 kHz.

[0031] By employing the above technical solutions, this invention provides a bone-guided speech conversion method based on spectral envelope mapping. This method, based on a source-filter model, maps the LSF coefficients of the bone-guided speech using a GRU neural network. The GRU network learns the sequential features of the speech, improving mapping accuracy. Furthermore, it applies timbre weighting based on the characteristics of bone-guided speech, making the low-frequency portion of the compensated speech closely resemble the bone-guided speech, while recovering the effective air-guided information in the high-frequency portion. A codebook search method is used to optimally select the excitation signal, improving robustness and reducing invalid noise in the high-frequency portion. Post-filtering is applied to the compensated speech generated from the optimal excitation and the mapped spectral envelope, making the missing mid-to-high-frequency information in the bone-guided speech more apparent and reducing interference from invalid noise. The overall algorithm has low computational complexity and can be implemented in real time. Attached Figure Description

[0032] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0033] Figure 1 This is a flowchart of the method of the present invention;

[0034] Figure 2 This is a schematic diagram of the spectral envelope mapping network structure in this invention;

[0035] Figure 3 This is the frequency response diagram of the timbre-weighted filter in this invention;

[0036] Figure 4 This is the frequency response diagram of the bandpass filter in this invention;

[0037] Figure 5 This is the frequency response diagram of the low-pass filter in this invention;

[0038] Figure 6 This is the spectrogram of the original bone conduction speech signal in this invention;

[0039] Figure 7 This is the spectrogram of the converted air-conducted speech signal in this invention. Detailed Implementation

[0040] To make the technical solutions and advantages of the present invention clearer, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention:

[0041] like Figure 1The bone-guided speech conversion method based on spectral envelope mapping shown includes the following steps:

[0042] S1: Input bone conduction speech signal with a sampling frequency of 8kHz A high-pass filter with a cutoff frequency of 150Hz is used to avoid overflow and filter out low-frequency components. The filtered signal is then used... express.

[0043] In linear predictive analysis, the expression for the LP filter is:

[0044]

[0045] in, These are the coefficients of the LP filter.

[0046] For each frame of bone conduction speech signal, a frame-by-frame windowed autocorrelation method is used for linear prediction analysis. The Hamming window is used as the window function, and its expression is as follows:

[0047]

[0048] Calculate the autocorrelation function of windowed speech. use Solving for LP filter coefficients That is, to solve the following system of equations:

[0049]

[0050] This uses the classic Levinson-Durbin algorithm to obtain the LP filter coefficients.

[0051] S2: To facilitate neural network mapping and interpolation, the LP filter coefficients need to be converted into line spectrum pair LSP coefficients. The line spectrum pair LSP coefficients are defined as the roots of the following two polynomials.

[0052]

[0053]

[0054] polynomial It is symmetrical. It is antisymmetric, defining a new polynomial

[0055]

[0056]

[0057] In the above equation, the roots of these polynomials lie on the unit circle, and each polynomial has 5 conjugate complex roots. They can be expressed as follows:

[0058]

[0059]

[0060] in, .parameter These are called line spectral frequency (LSF) coefficients, which satisfy the following form of order. .

[0061] Using Chebyshev polynomials , The line spectrum pair LSP coefficients can be calculated by evaluating them. To facilitate neural network mapping, the line spectrum frequency LSF coefficients can then be obtained from the line spectrum pair LSP coefficients.

[0062] The network consists of four fully connected neural network layers and one GRU network layer. Its structure is as follows: Figure 2 As shown:

[0063] During the training phase, parallel corpora of bone-conducted and air-conducted speech samples from 15 men and 4 women were collected as a corpus. Data augmentation was performed on the corpus by adding common noises such as those from shopping malls, cafes, streets, and offices to the air-conducted speech signals, as well as specific environmental noises such as those from power plants, mines, and steel mills. Noise reduction was then applied. This makes the LSF coefficients of the mapped air-conducted speech signals more robust. The LSF coefficients of the bone-conducted and air-conducted speech signals at the same time in the corpus were calculated and learned using a GRU network to obtain temporal information before and after the speech. Unlike the classic BP neural network, this method uses the ReLU activation function commonly used in deep networks and employs the Adam optimization method to train the model, using the mean absolute error (MAE) loss function. The final network size is as follows: 10 nodes in the input layer; 32 nodes in the first hidden layer; 64 nodes in the second hidden layer; 64 dimensions of input features for the GRU network; 64 dimensions of output features for the GRU network; 32 nodes in the third hidden layer; and 10 nodes in the output layer.

[0064] During the mapping phase, the LSF coefficients of the input 10-dimensional bone conduction speech signal are... Perform forward operations layer by layer according to the following formula.

[0065]

[0066]

[0067]

[0068]

[0069]

[0070] Finally, based on the values ​​of the output layer These are the LSF coefficients of a 10-dimensional air-conduction speech signal. After mapping the LSF coefficients of the bone conduction speech signal to the LSF coefficients of the air-conduction speech signal, it is also necessary to ensure that the 10th-order coefficients satisfy... This ensures that the synthesized speech does not produce distortion or popping sounds.

[0071] The LSP coefficients of the air conduction speech signal are obtained from the LSF coefficients of the air conduction speech signal, and then the LSP coefficients of the final synthesized air conduction speech signal are obtained by weighting them with the LSP coefficients of the bone conduction speech signal.

[0072]

[0073] in

[0074] Then, the LP filter is calculated from these LSP coefficients. and linear prediction coefficient

[0075] S3: The timbre-weighted filter is defined as...

[0076]

[0077] The frequency response of the timbre-weighted filter is as follows Figure 3 Bone conduction speech signals are characterized by severe loss of high-frequency frequencies. To improve the accuracy of subsequent search stimuli, the bone conduction speech signal is weighted by timbre, with higher weights for low-frequency components and lower weights for high-frequency components. Furthermore, it incorporates the linear prediction coefficients of bone conduction speech signals. The linear prediction coefficients of the air-conducted speech signal obtained by mapping Perform timbre weighting.

[0078] Weighted bone conduction speech signals are

[0079]

[0080] The autocorrelation function of the weighted bone conduction speech signal is:

[0081]

[0082] S4: respectively in Calculate the autocorrelation function within these three ranges The maximum value, keep the largest. Then normalize to complete integer gene estimation. .

[0083] More precise fractional gene estimation is performed based on integer gene estimation. The fractional pitch estimation search aims to minimize the weighted mean square error between bone conduction and air conduction speech signals, transforming it into... To obtain the maximum value.

[0084]

[0085] in For reference signal, The unit impulse response of the timbre-weighted filter based on the characteristics of bone conduction speech. Indicates incentive delay and The convolution yields the integer gene delay. and fractional gene delay

[0086] Calculate the adaptive codebook vector Upsampling filter The coefficients are shown in Table 1, and .

[0087] Table 1 Upsampling Filter coefficient

[0088]

[0089]

[0090] Calculate adaptive codebook gain

[0091]

[0092] in

[0093]

[0094] S5: Fixed Codebook Structure and Search: In a fixed codebook, each codebook vector contains 4 non-zero pulses, each pulse having an amplitude of 1 or -1. The fixed codebook vector is...

[0095]

[0096] The reference signal for fixed codebook search is

[0097]

[0098] The basic principle for fixed codebook search is to input bone conduction speech signals. To minimize the mean square error of the air-conducted speech signal, it can be transformed into the following formula to obtain its maximum value.

[0099]

[0100] in This is a lower triangular Toepliz convolution matrix, where all diagonal elements are... The lower diagonal elements are as follows:

[0101]

[0102]

[0103] Search for fixed codebook gain To minimize the weighted mean square error between bone conduction speech signals and air conduction speech signals.

[0104]

[0105] in

[0106]

[0107] S6: The excitation signal for the current subframe is

[0108]

[0109] Signal Through LP filter The synthesized air-conducted speech signal was obtained. This involves synthesizing an airborne speech signal using the spectral envelope of the airborne speech signal obtained from the optimal excitation and network mapping. Finally, the synthesized airborne speech signal is processed through a... Figure 4 The passband is The purpose of the bandpass filter is to enhance the timbre of the frequency bands where bone conduction speech signals are lost. Then, through... Figure 5 The cutoff frequency is A low-pass filter to remove high-frequency noise.

[0110] To verify the effectiveness of this method, a large number of bone conduction voice samples were tested. Figure 6 This is the original bone conduction speech signal spectrogram. Figure 7 This is the spectrogram of the converted air-conducted speech signal. The results show that the converted... The mid-to-high frequency components recovered valid information.

[0111] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in the present invention, based on the technical solution and inventive concept of the present invention, should be covered within the scope of protection of the present invention.

Claims

1. A bone-guided speech conversion method based on spectral envelope mapping, characterized in that... include: Preprocessing and linear prediction analysis are performed on bone conduction speech signals. LP filter coefficients are calculated, converted into line spectrum pair LSP coefficients, and then converted into line spectrum frequency LSF coefficients. The trained neural network is used to map the LSF coefficients of the air conduction speech signal corresponding to the bone conduction speech signal, and the LP filter coefficients of the air conduction speech signal are obtained based on the LSF coefficients of the air conduction speech signal. The optimal excitation signal is searched using a comprehensive analysis method, and the error between the original bone conduction speech signal and the synthesized air conduction speech signal is minimized based on the timbre weighting characteristics of bone conduction speech. The bone conduction speech signal is passed through a timbre-weighted filter to estimate the integer pitch, the linear prediction residual signal is passed through the timbre-weighted filter to obtain the reference signal, and the fractional pitch is estimated to obtain the adaptive codebook vector. Subtract the adaptive codebook vector from the reference signal to obtain a new reference signal, and search for the optimal excitation in the fixed codebook; The optimal excitation is combined with an LP filter of the air-conducted speech signal to synthesize an air-conducted speech signal, which is then corrected. To minimize the error between the original bone conduction speech signal and the synthesized air-conducted speech signal, the timbre-weighted filter is defined as follows: The bone conduction signal is weighted by timbre, with a higher weight for the low-frequency component and a lower weight for the high-frequency component. The linear prediction coefficients of the original bone conduction speech signal were combined. The linear prediction coefficients of the air-conducted speech signal obtained by mapping Perform timbre weighting; When synthesizing the airborne speech signal by combining the optimal excitation with an LP filter of the airborne speech signal, and then correcting the airborne speech signal, a codebook search method is used to search for the optimal excitation signal. ; in, For adaptive codebook vectors, For adaptive codebook gain, For fixed codebook vectors, To maintain a fixed codebook gain, the optimal excitation signal will be... LP filters that transmit speech signals through the air , obtain air-conducted speech signal ; The air-conducted speech signal is corrected by passing it through a bandpass filter with a passband of 1.5–2.5 kHz and a low-pass filter with a cutoff frequency of 3.5 kHz.

2. The bone-guided speech conversion method based on spectral envelope mapping according to claim 1, characterized in that: When using a neural network to map the LSF coefficients of the air conduction speech signal corresponding to the bone conduction speech signal: during the training phase, the LSF coefficients of the bone conduction speech signal and the air conduction speech signal at the same time are calculated, fed into the neural network for training, and the weights after training are saved. During the mapping phase, the LSF coefficients of the input 10-dimensional bone conduction speech signal are... The forward operation is performed layer by layer in the following manner: Output To obtain the LSF coefficients of the 10-dimensional airborne speech signal obtained through mapping, the LSF coefficients of the airborne speech signal are corrected to ensure that the 10th-order coefficients satisfy... ; The LSP coefficients of the air conduction speech signal are obtained from the LSF coefficients of the mapped air conduction speech signal, and then weighted with the LSP coefficients of the bone conduction speech signal to obtain the LSP coefficients of the final synthesized air conduction speech signal. in Then, the LP filter for the air-conducted speech signal is calculated based on these LSP coefficients. and linear prediction coefficient .

Citation Information

Patent Citations

  • Linear prediction speech coding method and speech synthesis method

    CN103050121A

  • Bone conduction voice blind enhancement method based on codec architecture and recurrent neural network

    CN108986834A