A method for voice bandwidth expansion and related devices
By predicting high-frequency power spectra and simulating human vocal tract characteristics, the method enhances VOIP audio quality by producing wideband signals that accurately represent speaker traits, addressing the issue of degraded audio quality in VOIP systems.
Patent Information
- Application Number
- CN202111335506.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-11-11
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2041-11-11
AI Technical Summary
When the existing voice bandwidth expansion method synthesizes broadband voice signals, it is difficult to reflect the language characteristics of the speaker, resulting in poor voice quality.
By predicting the narrowband voice signal frames at high frequency power spectrum, the broadband power spectrum is obtained, and the filter parameters are used to simulate the human voice channel, filter out the tone information and channel shape information, and generate a broadband voice signal frame similar to human voice.
The fidelity and voice quality of broadband voice signal frames are improved, making the synthesized voice signal sound more natural and reflects the language characteristics of the speaker.
Smart Images

Figure CN116110424B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of voice communication technologies, and particularly to a voice bandwidth expansion method and related devices. Background Art
[0002] In a voice communication system based on Voice over Internet Protocol (VOIP), for example, when conducting multi-person audio and video calls through instant messaging software, etc., the receiving end may receive narrowband voice signals due to certain circumstances, resulting in an obvious decline in sound quality.
[0003] In order to improve the subjective experience of listening to voice at the receiving end, the most natural method is to perform voice bandwidth expansion on the narrowband voice signal, and then artificially synthesize a broadband voice signal.
[0004] However, the current voice bandwidth expansion methods only try to recover the lost high-frequency information as much as possible. The synthesized broadband voice signal gives a rather rigid and mechanical listening feeling to users, making it difficult to reflect the language characteristics of the speaker, and its performance in terms of fidelity is poor, thus resulting in poor voice quality. Summary of the Invention
[0005] To solve the above technical problems, this application provides a voice bandwidth expansion method and related devices, which can obtain broadband voice signal frames with a listening feeling similar to human vocalization, can reflect the language characteristics of the speaker, thereby improving the fidelity of the broadband voice signal frames and improving the voice quality.
[0006] The embodiments of this application disclose the following technical solutions:
[0007] In a first aspect, the embodiments of this application provide a voice bandwidth expansion method, and the method includes:
[0008] Perform high-frequency power spectrum prediction based on a narrowband voice signal frame to be processed to obtain the high-frequency power spectrum corresponding to the narrowband voice signal frame to be processed;
[0009] Stitch the high-frequency power spectrum and the low-frequency power spectrum corresponding to the narrowband voice signal frame to be processed to obtain a broadband power spectrum;
[0010] Perform spectral envelope calculation on the broadband power spectrum, and determine filter parameters according to the calculation result;
[0011] Perform analysis filtering based on the narrowband voice signal frame to be processed and the filter parameters to generate a low-frequency excitation signal corresponding to the narrowband voice signal frame to be processed, where the analysis filtering is used to filter out the timbre information and vocal tract shape information in the narrowband voice signal frame to be processed;
[0012] Expand the low-frequency excitation signal to obtain a high-frequency excitation signal;
[0013] Synthesize according to the low-frequency excitation signal and the high-frequency excitation signal to obtain a broadband excitation signal;
[0014] Perform synthesis filtering based on the broadband excitation signal and the filter parameters to obtain a broadband speech signal frame, and the synthesis filtering is used to add timbre information and vocal tract shape information to the broadband excitation signal.
[0015] In a second aspect, an embodiment of the present application provides a voice bandwidth expansion device, and the device includes a prediction unit, a splicing unit, a determination unit, a generation unit, and a filtering unit:
[0016] The prediction unit is configured to perform high-frequency power spectrum prediction according to the narrowband speech signal frame to be processed, and obtain the high-frequency power spectrum corresponding to the narrowband speech signal frame to be processed;
[0017] The splicing unit is configured to splice the high-frequency power spectrum and the low-frequency power spectrum corresponding to the narrowband speech signal frame to be processed to obtain a broadband power spectrum;
[0018] The determination unit is configured to calculate a spectral envelope of the broadband power spectrum and determine filter parameters according to the calculation result;
[0019] The generation unit is configured to perform analysis filtering based on the narrowband speech signal frame to be processed and the filter parameters to generate a low-frequency excitation signal corresponding to the narrowband speech signal frame to be processed, and the analysis filtering is used to filter out timbre information and vocal tract shape information in the narrowband speech signal frame to be processed;
[0020] The determination unit is further configured to expand the low-frequency excitation signal to obtain a high-frequency excitation signal; synthesize according to the low-frequency excitation signal and the high-frequency excitation signal to obtain a broadband excitation signal;
[0021] The filtering unit is configured to perform synthesis filtering based on the broadband excitation signal and the filter parameters to obtain a broadband speech signal frame, and the synthesis filtering is used to add timbre information and vocal tract shape information to the broadband excitation signal.
[0022] In a third aspect, an embodiment of the present application provides a device for voice bandwidth expansion, and the device includes a processor and a memory:
[0023] The memory is used to store program code and transmit the program code to the processor;
[0024] The processor is configured to execute the method described in the first aspect according to the instructions in the program code.
[0025] Fourthly, an embodiment of the present application provides a computer-readable storage medium for storing program codes for executing the method described in the first aspect.
[0026] Fifthly, an embodiment of the present application provides a computer program product including a computer program which, when executed by a processor, implements the method described in the first aspect.
[0027] As can be seen from the above technical solutions, after obtaining a narrowband speech signal frame to be processed, high-frequency power spectrum prediction is performed based on the narrowband speech signal frame to be processed to obtain the high-frequency power spectrum corresponding to the narrowband speech signal frame to be processed. The high-frequency power spectrum and the low-frequency power spectrum corresponding to the narrowband speech signal frame to be processed are spliced to obtain a broadband power spectrum, thereby expanding the bandwidth of the narrowband speech signal frame to be processed to a certain extent. Based on the analysis of human speech generation, the vocal tract (mouth, throat) of humans is equivalent to a filter, and the vibrations generated by the vocal tract produce speech signals with different timbres through different-shaped vocal tracts. Therefore, the present application calculates the spectral envelope of the broadband power spectrum and determines the filter parameters according to the calculation results to simulate the human vocal tract through the filter parameters. Then, analysis filtering is performed based on the narrowband speech signal frame to be processed and the filter parameters to filter out the timbre information and vocal tract shape information in the narrowband speech signal frame to be processed to generate a low-frequency excitation signal corresponding to the narrowband speech signal frame to be processed. The low-frequency excitation signal is expanded to obtain a high-frequency excitation signal, and then a broadband excitation signal is determined by synthesizing the low-frequency excitation signal and the high-frequency excitation signal. This broadband excitation signal is equivalent to the vibrations generated by the vocal tract during the human speech generation process. Since the filter parameters simulate the human vocal tract, synthesis filtering is performed based on the broadband excitation signal and the filter parameters to add timbre information and vocal tract shape information to the broadband excitation signal to obtain a broadband speech signal frame, which is equivalent to processing the vibrations generated by the vocal tract through the corresponding vocal tract, and then obtaining a broadband speech signal frame similar to human vocalization. It can be seen that the speech bandwidth expansion method provided by this solution can obtain a broadband speech signal frame with a listening sensation similar to human vocalization, which can reflect the language characteristics of the speaker, thereby improving the fidelity of the broadband speech signal frame and the speech quality. Description of the Drawings
[0028] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0029] Figure 1 It is an example diagram of different speech signals provided by an embodiment of the present application;
[0030] Figure 2 Schematic diagram of the system architecture of a voice bandwidth expansion method provided by an embodiment of the present application;
[0031] Figure 3 Flowchart of a voice bandwidth expansion method provided by an embodiment of the present application;
[0032] Figure 4 Flow structure block diagram of a voice bandwidth expansion method provided by an embodiment of the present application;
[0033] Figure 5 Flow structure block diagram of obtaining a wideband excitation signal based on a low-frequency excitation signal provided by an embodiment of the present application;
[0034] Figure 6 Logic block diagram of a receiving end receiving the code streams of narrowband voice signals from multiple sending ends and sending them for playback provided by an embodiment of the present application;
[0035] Figure 7 System block diagram of a receiving end simultaneously receiving narrowband voice signals and wideband voice signals provided by an embodiment of the present application;
[0036] Figure 8 Structure diagram of a voice bandwidth expansion device provided by an embodiment of the present application;
[0037] Figure 9 Structure diagram of a smart phone provided by an embodiment of the present application;
[0038] Figure 10 Structure diagram of a server provided by an embodiment of the present application. Detailed implementation manners
[0039] The embodiments of the present application will be described below with reference to the accompanying drawings.
[0040] In a VOIP voice communication system, generally there are two situations that may cause the receiving end to receive narrowband voice signals, resulting in the loss of high-frequency information and thus an obvious decline in sound quality. The first situation is that when making a multi-party VOIP voice call, at least one party accesses through a traditional narrowband telephone, such as the Public Switched Telephone Network (PSTN), and thus other receiving ends can only receive narrowband voice signals. The second situation is that when making a VOIP voice call between two or more parties, due to the settings of the vocoder, such as when the network bandwidth is insufficient or the packet loss is serious, the vocoder operates in the narrowband coding mode and thus only sends the code stream of the narrowband voice signal. In order to improve the subjective experience of listening to voice at the receiving end, the most natural method is to perform voice bandwidth expansion on the narrowband voice and thus artificially synthesize a wideband voice signal.
[0041] For the sake of helping understanding, different voice signals will be introduced separately below. Refer to Figure 1 as shown in Figure 1 The figure marked in (a) in it is a broadband voice signal with a sampling rate of 16 kHz. According to the Nyquist sampling theorem, the bandwidth of the voice signal is half of the sampling rate, that is, the bandwidth of this broadband voice signal is 8 kHz. Figure 1 The figure marked in (b) in it is a narrowband voice signal with a sampling rate of 8 kHz. Therefore, the bandwidth of this narrowband voice signal is only 4 kHz (for example Figure 1 in the figure marked in (b) in it, the voice signal is basically concentrated below 4 kHz). Figure 1 The figure marked in (c) in it is a voice signal obtained by upsampling the narrowband voice signal sampled at 8 kHz to a sampling rate of 16 kHz. At this time, although the sampling rate has been increased, since the sampling was performed at a sampling rate of 8 kHz before, only the voice signal below 4 kHz can be collected. Therefore, even though the current sampling rate has become 16 kHz, the voice signal above 4 kHz is still lost (for example Figure 1 in the figure marked in (c) in it, it is completely black above 4 kHz and there is basically no voice signal). In this sense, the voice signal after upsampling is still a narrowband voice signal. It can be seen that the loss of high-frequency information (such as the voice signal above 4 kHz) in the narrowband voice signal leads to a decrease in the quality of the voice signal, and further reduces the subjective experience of listening to the voice.
[0042] Therefore, it is necessary to expand the bandwidth of the narrowband voice signal, for example Figure 1 as shown in the figure marked in (d) in it, which is a broadband voice signal after predicting and complementing the high-frequency information of the narrowband voice signal. It can be seen that according to the low-frequency information below 4 kHz, the high-frequency information above 4 kHz can be roughly predicted, and then a new broadband voice signal is obtained, improving the subjective auditory feeling.
[0043] The current voice bandwidth expansion method only tries to restore the lost high-frequency information as much as possible. The synthesized broadband voice signal gives the user a rather rigid and mechanical listening feeling, and it is difficult to reflect the language characteristics of the speaker. It performs poorly in terms of fidelity, and thus leads to poor voice quality.
[0044] To solve the above technical problems, an embodiment of the present application provides a method for voice bandwidth expansion. The method predicts the high-frequency power spectrum corresponding to the narrowband voice to be processed, and then obtains the broadband power spectrum. Then, based on the broadband power spectrum, spectral envelope calculation is performed to obtain filter parameters, and the human vocal tract is simulated through the filter parameters, so as to obtain a broadband voice signal frame with a similar auditory sense to human voice based on the filter parameters. The broadband voice signal frame obtained in this way can reflect the language characteristics of the speaker, thereby improving the fidelity of the broadband voice signal frame and the voice quality.
[0045] To facilitate the understanding of the technical solution of the present application, the method for voice bandwidth expansion provided in the embodiment of the present application will be introduced below in combination with an actual application scenario.
[0046] See Figure 2 , Figure 2 which is a schematic diagram of the system architecture for a method for voice bandwidth expansion provided in an embodiment of the present application. The system architecture includes a receiving end 201 and a transmitting end 202. Among them, the receiving end 201 and the transmitting end 202 conduct a voice call through a network. The receiving end 201 and the transmitting end 202 can be terminal devices with voice call functions, such as smartphones, tablets, laptops, desktop computers, smart watches, vehicle terminals, smart TVs, landline phones, etc. In the scenario of a multi-person audio and video call through instant messaging software, etc., instant messaging software can be installed on the receiving end 201 and the transmitting end 202.
[0047] When the transmitting end 202 and the receiving end 201 conduct a voice call, due to the above two situations, the receiving end may play according to the narrowband voice signal, thereby reducing the subjective experience of the user listening to the voice on the receiving end 201. Therefore, after the receiving end 201 obtains the narrowband voice signal frame to be processed, it can perform high-frequency power spectrum prediction according to the narrowband voice signal frame to be processed to obtain the high-frequency power spectrum corresponding to the narrowband voice signal frame to be processed, and splice the high-frequency power spectrum and the low-frequency power spectrum corresponding to the narrowband voice signal frame to be processed to obtain a broadband power spectrum, thereby expanding the bandwidth of the narrowband voice signal frame to be processed to a certain extent and complementing the high-frequency information.
[0048] Among them, the narrowband voice signal frame to be processed is the basic unit for voice bandwidth expansion. The narrowband voice signal obtained by the receiving end 201 includes multiple narrowband voice signal frames, and each narrowband voice signal frame can be used as the narrowband voice signal frame to be processed to execute the method for voice bandwidth expansion provided in the embodiment of the present application.
[0049] Based on the analysis of human speech generation, the human vocal tract (oral cavity, throat) is equivalent to a filter. The vibrations generated by the vocal tract produce speech signals with different timbres through vocal tracts of different shapes. Therefore, the receiving end 201 can calculate the spectral envelope of the broadband power spectrum and determine the filter parameters according to the calculation results, and simulate the human vocal tract through the filter parameters. Then, based on the narrowband speech signal frame to be processed and the filter parameters, analysis and filtering are performed to filter out the timbre information and vocal tract shape information in the narrowband speech signal frame to be processed to generate a low-frequency excitation signal corresponding to the narrowband speech signal frame to be processed. The low-frequency excitation signal is expanded to obtain a high-frequency excitation signal, and then a broadband excitation signal is determined through synthesis based on the low-frequency excitation signal and the high-frequency excitation signal. This broadband excitation signal is equivalent to the vibrations generated by the vocal tract during the human speech generation process. Since the filter parameters simulate the human vocal tract, a broadband speech signal frame is obtained through synthesis and filtering based on the broadband excitation signal and the filter parameters, which is equivalent to processing the vibrations generated by the vocal tract through the corresponding vocal tract, and then a broadband speech signal frame similar to human vocalization is obtained.
[0050] It should be noted that the method provided in the embodiments of the present application can also be executed by a server. For example, if the narrowband speech signal obtained by the receiving end 201 is caused by the above second case, then what the sending end 202 sends to the receiving end 201 through the network is the narrowband speech signal. In this case, the server can first execute the speech bandwidth expansion method provided in the embodiments of the present application on the narrowband speech signal frames included in the narrowband speech signal, and then send the obtained broadband speech signal frames to the receiving end 201 for playback. Of course, it can also be executed in cooperation by the server and the receiving end 201, and the embodiments of the present application do not limit this. Among them, the server can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server providing cloud computing services. The receiving end 201, the sending end 202, and the server can be directly or indirectly connected through wired or wireless communication methods, and the present application does not limit this here.
[0051] Next, taking the receiving end as the execution entity as an example, the speech bandwidth expansion method provided in the embodiments of the present application will be introduced in detail with reference to the accompanying drawings.
[0052] See Figure 3 , Figure 3 shows a flowchart of a speech bandwidth expansion method, and the method includes:
[0053] S301. Predict the high-frequency power spectrum according to the narrowband speech signal frame to be processed to obtain the high-frequency power spectrum corresponding to the narrowband speech signal frame to be processed.
[0054] In a VOIP multi-party voice call, the receiving end of a certain party may receive narrowband voice signals from other sending ends. The other sending ends may be landline phones or mobile phones, which send narrowband voice signals through wired or wireless telecommunication networks, such as G.711 or G.729a, etc. It is also possible that other computers or mobile devices accessing the Internet encode only the low-frequency narrowband energy of the voice in a specific mode through a VOIP application program. For example, the low bitrate mode of opus, and so on.
[0055] As a result, it may cause the receiving end to obtain narrowband voice signals, thereby affecting the subjective auditory experience of the user at the receiving end. To avoid this situation, the receiving end can perform voice bandwidth expansion on the narrowband voice signals. Since the narrowband voice signals can include multiple narrowband voice signal frames, each narrowband voice signal frame can be used as a narrowband voice signal frame to be processed, and the narrowband voice signal frame to be processed is used as the smallest unit for performing the voice bandwidth expansion method.
[0056] When performing voice bandwidth expansion, first, the receiving end can perform high-frequency power spectrum prediction based on the narrowband voice signal frame to be processed to obtain the high-frequency power spectrum corresponding to the narrowband voice signal frame to be processed. In a possible implementation manner, the receiving end can use a pre-trained prediction model, that is, an artificial intelligence (AI) model for high-frequency power spectrum prediction. Specifically, reference can be made to Figure 4 as shown Figure 4 which shows a flowchart structural block diagram of a voice bandwidth expansion method. The receiving end extracts features from the narrowband voice signal frame to be processed to obtain a corresponding feature vector, and then performs high-frequency power spectrum prediction based on the feature vector and the prediction model (see Figure 4 shown in 401 therein) to obtain the high-frequency power spectrum corresponding to the narrowband voice signal frame to be processed.
[0057] It can be seen that the embodiments of the present application may relate to the field of artificial intelligence. Artificial intelligence (AI) is a comprehensive technology in computer science. It attempts to understand the essence of intelligence and produce a new intelligent machine that can react in a way similar to human intelligence. The basic technologies of artificial intelligence generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, and mechatronics. The software technologies of artificial intelligence mainly include several major directions such as computer vision technology, speech technology, natural language processing technology, and machine learning / deep learning, autonomous driving, and intelligent transportation. The embodiments of the present application, for example, relate to speech technology, and particularly to speech signal feature extraction technology, so as to extract features from the narrowband voice signal frame to be processed to obtain a corresponding feature vector.
[0058] For another example, when it comes to Machine Learning (ML) technology, machine learning is the core of artificial intelligence and the fundamental way to endow computers with intelligence, and its applications cover all fields of artificial intelligence. Machine learning and deep learning generally include technologies such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and rote learning. In the embodiments of the present application, a prediction model can be trained using machine learning.
[0059] Among them, feature extraction can be implemented through a feature extraction module (see Figure 4 shown in 402). In a possible implementation manner, the implementation manner of performing feature extraction on the narrowband speech signal frame to be processed to obtain a corresponding feature vector can be to divide the spectrum of the narrowband speech signal frame to be processed into multiple sub-bands, calculate the power spectrum corresponding to each sub-band respectively, so as to determine the feature vector according to the power spectrum.
[0060] The multiple sub-bands can be K sub-bands, where K is an integer greater than 1. The power spectrum corresponding to each sub-band can be a logarithmic power spectrum or other forms of power spectra, and the present embodiment does not limit this. If the power spectrum corresponding to each sub-band is a logarithmic power spectrum, the calculated logarithmic power spectrum of the corresponding sub-band can be represented by Pl t (k), k = 1, 2,..., K, where t represents that the narrowband speech signal frame to be processed is the t-th narrowband speech signal frame, and k represents the k-th sub-band. The sub-band division method can be uniform division according to frequency or division according to the bark frequency band that conforms to the human ear auditory response.
[0061] In a possible implementation manner, after obtaining the power spectrum corresponding to each sub-band, the power spectrum corresponding to each sub-band can be used as the feature vector. In another possible implementation manner, in order to improve the prediction accuracy, the logarithmic power spectra for a period of time before and after the current moment can also be combined together to form a larger feature vector. For example, the logarithmic power spectra corresponding to the previous p frames (i.e., the p frames before the narrowband speech signal frame to be processed) and the subsequent q frames (i.e., the q frames after the narrowband speech signal frame to be processed) are combined together to obtain the final feature vector F = [Pl t-p (k)Pl t (k)Pl t+q (k)], where t represents that the narrowband speech signal frame to be processed is the t-th narrowband speech signal frame, t - p represents the (t - p)-th narrowband speech signal frame, t + q represents the (t + q)-th narrowband speech signal frame, Pl t-p (k) represents the power spectrum corresponding to each sub-band in the (t - p)-th narrowband speech signal frame, Pl t (k) represents the power spectrum corresponding to each sub-band in the t-th narrowband speech signal frame, Plt+q (k) represents the power spectrum corresponding to each sub - band in the narrow - band speech signal frame of the (t + q)-th frame. In this way, more information can be included in the feature vector, thereby improving the prediction accuracy.
[0062] It should be noted that the above - mentioned prediction model can include a forward multi - layer perceptron network (Deep Neural Networks, DNN), a recurrent neural network (Recurrent Neural Network, RNN), a convolutional neural network (Convolutional Neural Networks, CNN), etc., as well as a combined network model of one or more of the above - mentioned neural network layers. The prediction model trains the model parameters through pre - prepared training data, and then interacts with the input feature vector F and the model parameters, so as to be able to predict the high - frequency power spectrum of the narrow - band speech signal frame to be processed. Among them, the high - frequency power spectrum prediction can be realized by a high - frequency power spectrum prediction module (see Figure 4 shown in 403 in
[0063] The above - mentioned high - frequency power spectrum prediction module predicts the high - frequency power spectrum according to the input feature vector F and the model parameters, and according to the model structure specified by the model parameters, such as the DNN model. In a possible implementation manner, the method of obtaining the high - frequency power spectrum corresponding to the narrow - band speech signal frame to be processed based on the feature vector and the prediction model can be to input the feature vector into the prediction model, and the prediction model outputs the power value of each frequency point in the high - frequency band of the narrow - band speech signal frame to be processed, and then constructs the high - frequency power spectrum based on the power value of each frequency point in the high - frequency band of the narrow - band speech signal frame to be processed. For example, the high - frequency power spectrum of the narrow - band speech signal frame to be processed is expressed as PSDh t (z), z = 1, 2, …, N / 2, where N is the number of frequency points of the broadband spectrum, N / 2 represents that the number of frequency points of the high - frequency band is half of the number of frequency points of the entire broadband spectrum, and z represents the z - th frequency point in the high - frequency band.
[0064] In another possible implementation, the method for obtaining the high-frequency power spectrum corresponding to the narrowband speech signal frame to be processed based on the eigenvector and the prediction model can also be to input the eigenvector into the prediction model, and the prediction model outputs the average power value of each sub-band in the high-frequency band of the narrowband speech signal frame to be processed. Then, based on the average power value of each sub-band in the high-frequency band of the narrowband speech signal frame to be processed, a high-frequency power spectrum is constructed. At this time, the output of the prediction model is not the power values of all N / 2 frequency points of the high-frequency power spectrum, but the high-frequency band is divided into multiple sub-bands (for example, S sub-bands) in a preset manner, and the prediction model predicts the average power value of each sub-band as the power value of the frequency points included in the corresponding sub-band. Then, based on the average power value of each sub-band, a high-frequency power spectrum is constructed. In this way, it is not necessary to calculate the power value of each frequency point, thus greatly reducing the calculation amount and improving the prediction efficiency.
[0065] S302. Concatenate the high-frequency power spectrum and the low-frequency power spectrum corresponding to the narrowband speech signal frame to be processed to obtain a broadband power spectrum.
[0066] In the embodiment of the present application, the low-frequency power spectrum corresponding to the narrowband speech signal frame to be processed can be represented by PSDl t (z), where t represents that the narrowband speech signal frame to be processed is the t-th narrowband speech signal frame, and z represents the z-th frequency point in the low-frequency band. Then, through the broadband power spectrum synthesis module (see Figure 4 shown in 404), the above low-frequency power spectrum is concatenated with the predicted high-frequency power spectrum to obtain a broadband power spectrum PSD t = [PSDl t PSDh t , where PSDl t represents the low-frequency power spectrum, and PSDh t represents the high-frequency power spectrum. If the low-frequency power spectrum and the high-frequency power spectrum are respectively a 256-point real vector, the concatenated broadband power spectrum is a 512-point real vector.
[0067] In a possible implementation, before concatenating the high-frequency power spectrum and the low-frequency power spectrum corresponding to the narrowband speech signal frame to be processed to obtain a broadband power spectrum, the narrowband speech signal frame to be processed can be upsampled, and then the low-frequency power spectrum is calculated based on the upsampled narrowband speech signal frame to be processed.
[0068] Among them, the upsampling process can be implemented through an upsampling module (see Figure 4 shown in 405), Figure 4The upsampling module in [[ ]] performs upsampling on the narrowband speech signal frame to be processed. For example, the sampling rate of a general narrowband speech signal is 8 kHz. The upsampling module uses a standard upsampling method, such as the interpolation filtering method, to increase the sampling rate to 16 kHz, thereby obtaining the narrowband speech signal frame to be processed after upsampling.
[0069] In this way, the frequency points of the narrowband speech signal frame to be processed after upsampling can be aligned with the high-frequency power spectrum, facilitating subsequent splicing of the low-frequency power spectrum and the high-frequency power spectrum of the narrowband speech signal frame to be processed.
[0070] Then, through Figure 4 the low-frequency power spectrum calculation module in [[ ]] (see Figure 4 shown in 406 in [[ ]]) performs time-frequency domain conversion on the narrowband speech signal frame to be processed after upsampling after windowing. The window function used for windowing generally uses a Hanning window or a Hamming window, etc. The time-frequency domain conversion generally uses the discrete Fourier transform. The time-frequency domain conversion obtains a spectrum, and the low-frequency band spectrum coefficients of the spectrum, that is, the spectrum coefficients below 4 kHz, are used to calculate the low-frequency power spectrum PSDl t (z) = POW(ABS(s(z))), z = 1, 2,..., N / 2, where N / 2 corresponds to the highest frequency point of the low-frequency band, that is, 4 kHz. s(z) is the spectrum coefficient corresponding to the z-th frequency point obtained by Fourier transform, ABS represents taking the absolute value, taking the absolute value of the spectrum coefficient can obtain the amplitude spectrum, and POW represents squaring each absolute value, thereby obtaining the low-frequency power spectrum.
[0071] S303. Calculate the spectral envelope of the broadband power spectrum, and determine the filter parameters according to the calculation result.
[0072] In a possible implementation manner, calculating the spectral envelope of the broadband power spectrum obtains the linear prediction error and the linear prediction coefficients. At this time, the linear prediction error and the linear prediction coefficients are the calculation results, and then the filter parameters are determined according to the linear prediction error and the linear prediction coefficients. For example, the linear prediction error and the linear prediction coefficients can be directly used as the filter parameters, or the linear prediction error and the linear prediction coefficients can be transformed to obtain the filter parameters.
[0073] In the embodiments of the present application, the spectral envelope calculation can be performed through Figure 4 the broadband spectral envelope calculation module in [[ ]] (see Figure 4 shown in 407 in [[ ]]). For example, the linear prediction error and the linear prediction coefficients representing the spectral envelope can be calculated through the classical Levinson-Durbin recursive algorithm. The filter formed by the linear prediction error and the linear prediction coefficients as the filter parameters can be expressed as where G is the linear prediction error, α iare the P-order linear prediction coefficients, where P is a preset integer, generally between 10 and 30. The H(Z) filter can reflect the amplitude and spectral envelope of the speech signal frame and is usually called the all-pole vocal tract response filter. Filtering through H(Z) is called synthesis filtering, and filtering through its inverse filter is called analysis filtering.
[0074] S304. Perform analysis filtering based on the to-be-processed narrowband speech signal frame and the filter parameters to generate a low-frequency excitation signal corresponding to the to-be-processed narrowband speech signal frame.
[0075] After obtaining the filter parameters, the receiving end can perform analysis filtering based on the to-be-processed narrowband speech signal frame and the filter parameters. Through analysis filtering, the timbre information and vocal tract shape information in the to-be-processed narrowband speech signal frame can be filtered out to generate a low-frequency excitation signal corresponding to the to-be-processed narrowband speech signal frame. This low-frequency excitation signal is obtained by filtering out the timbre information and vocal tract shape information from the to-be-processed narrowband speech signal frame and is equivalent to the basic low-frequency vibration generated by the vocal tract.
[0076] If the to-be-processed narrowband speech signal frame has been upsampled, then in S304, analysis filtering can be performed based on the upsampled to-be-processed narrowband speech signal frame and the filter parameters to generate a low-frequency excitation signal corresponding to the narrowband speech signal frame. Among them, the low-frequency excitation signal is the linear prediction residual obtained by analysis filtering, and analysis filtering can be performed through the Figure 4 analysis filtering module in (see Figure 4 shown in 408 in
[0077] S305. Expand the low-frequency excitation signal to obtain a high-frequency excitation signal.
[0078] S306. Synthesize according to the low-frequency excitation signal and the high-frequency excitation signal to obtain a broadband excitation signal.
[0079] The receiving end can determine the broadband excitation signal according to the low-frequency excitation signal. In one possible implementation, the method of determining the broadband excitation signal according to the low-frequency excitation signal can be to expand the low-frequency excitation signal to obtain a high-frequency excitation signal, and synthesize according to the low-frequency excitation signal and the high-frequency excitation signal to obtain a broadband excitation signal.
[0080] Specifically, Figure 4 the broadband excitation signal generation module in (see Figure 4 shown in 409 inFigure 5 as shown
[0081] Figure 5 The high-frequency excitation expansion module in (see Figure 5 shown as 501 in) performs a modulation frequency shift operation on the low-frequency excitation signal: u h (z) = u l (z) * 2cos(Ωz), where u l (z) is the low-frequency excitation signal, and u h (z) is the high-frequency excitation signal obtained after the frequency shift operation. When Ω = π, it means that the Nyquist frequency is used for modulation. At this time, the high-frequency excitation signal is the mirror image of the low-frequency excitation signal in the frequency spectrum.
[0082] In a possible implementation manner, the method for synthesizing a wideband excitation signal according to the low-frequency excitation signal and the high-frequency excitation signal may be to perform high-pass filtering on the high-frequency excitation signal to obtain a high-frequency excitation signal with low-frequency noise filtered out, perform delay compensation on the low-frequency excitation signal, and then add the high-frequency excitation signal with low-frequency noise filtered out and the low-frequency excitation signal after delay compensation to obtain the wideband excitation signal.
[0083] Specifically, the high-frequency excitation signal obtained after high-frequency excitation expansion passes through a high-pass filter with a cut-off frequency of 4 kHz (see Figure 5 shown as 502 in) to filter out low-frequency noise. The delay compensation module (see Figure 5 shown as 503 in) is used to compensate for the delay caused by operations such as high-pass filtering. Finally, the wideband excitation signal generation module (see Figure 5 shown as 504 in, Figure 5 504 in is equivalent to Figure 4 409 in) adds the high-frequency excitation signal and the low-frequency excitation signal to obtain the wideband excitation signal u(z).
[0084] S307. Synthesize and filter based on the wideband excitation signal and the filter parameters to obtain a wideband speech signal frame.
[0085] In the embodiments of the present application, the wideband excitation signal can be subjected to synthetic filtering CONV(u, h) through the Figure 4 synthetic filtering module in (see Figure 4 shown as 410 in). The synthetic filtering can add timbre information and vocal tract shape information on the basis of the wideband excitation signal to obtain a synthesized wideband speech signal frame, that is, it is equivalent to the vocal tract affecting the wideband excitation signal, so that the wideband excitation signal passes through the action of the vocal tract to obtain a wideband speech signal frame with timbre information and vocal tract shape information. Where CONV represents the convolution filtering operation, u is the wideband excitation signal, and h is the filter parameter of the synthesis filter H.
[0086] In a possible implementation, after obtaining the broadband excitation signal, the broadband excitation signal can be shaped. In this way, when executing S307, the broadband speech signal frame can be obtained through synthetic filtering based on the shaped broadband excitation signal and the filter parameters.
[0087] Specifically, refer to Figure 4 as shown in Figure 4 the excitation signal shaping module in Figure 4 411 in Use a preset filter to perform filtering operations on the broadband excitation signal to avoid speech distortion caused by excessive energy in some frequency bands. At the same time, at some frequency points in the high-frequency band, according to the energy magnitude of the excitation signal in this frequency band, randomly generated noise is superimposed proportionally, making the excitation signal closer to the characteristics of human pronunciation in the high-frequency range. Thus, the shaped broadband excitation signal obtained is
[0088] By shaping the broadband excitation signal, speech distortion caused by excessive energy in some frequency bands can be avoided. At the same time, the shaped broadband excitation signal can be closer to the characteristics of human pronunciation in the high-frequency range, improving the auditory experience.
[0089] As can be seen from the above technical solution, after obtaining the narrowband speech signal frame to be processed, high-frequency power spectrum prediction is performed according to the narrowband speech signal frame to be processed to obtain the high-frequency power spectrum corresponding to the narrowband speech signal frame to be processed. The high-frequency power spectrum and the low-frequency power spectrum corresponding to the narrowband speech signal frame to be processed are spliced to obtain a broadband power spectrum, thereby expanding the bandwidth of the narrowband speech signal frame to be processed to a certain extent. Based on the analysis of human speech generation, the vocal tract (oral cavity, throat) of humans is equivalent to a filter, and the vibrations generated by the vocal tract produce speech signals with different timbres through vocal tracts of different shapes. Therefore, the present application calculates the spectral envelope of the broadband power spectrum, determines the filter parameters according to the calculation results, and simulates the human vocal tract through the filter parameters. Then, analysis filtering is performed based on the narrowband speech signal frame to be processed and the filter parameters to filter out the timbre information and vocal tract shape information in the narrowband speech signal frame to be processed to generate a low-frequency excitation signal corresponding to the narrowband speech signal frame to be processed. The low-frequency excitation signal is expanded to obtain a high-frequency excitation signal, and then a broadband excitation signal is determined by synthesizing the low-frequency excitation signal and the high-frequency excitation signal. This broadband excitation signal is equivalent to the vibration generated by the vocal tract during the human speech generation process. Since the filter parameters simulate the human vocal tract, synthesis filtering is performed based on the broadband excitation signal and the filter parameters to add timbre information and vocal tract shape information to the broadband excitation signal to obtain a broadband speech signal frame, which is equivalent to processing the vibration generated by the vocal tract through the corresponding vocal tract, and then obtaining a broadband speech signal frame similar to human vocalization. It can be seen that the speech bandwidth expansion method provided by this solution can obtain a broadband speech signal frame with a listening feeling similar to human vocalization, which can reflect the language characteristics of the speaker, thereby improving the fidelity of the broadband speech signal frame and the speech quality.
[0090] In a possible implementation manner, the receiving end may obtain multiple narrowband speech signals, and these multiple narrowband speech signals respectively originate from corresponding sending ends. In this case, in order to avoid executing the speech bandwidth expansion method provided by the embodiments of the present application for each narrowband speech signal, when obtaining the narrowband speech signal frame to be processed, narrowband mixing may be performed on the narrowband speech signal frames from different narrowband speech channels to obtain a narrowband speech signal frame after narrowband mixing, and then the narrowband speech signal frame after narrowband mixing is used as the narrowband speech signal frame to be processed.
[0091] For the sake of understanding, Figure 6 is a logic block diagram for the receiving end to receive the code streams of narrowband speech signals from multiple sending ends and send them for playback. Assume that M sending ends all send code streams of narrowband speech coding. Therefore, each downlink narrowband speech channel of the receiving end will receive and cache the code stream of the narrowband speech signal of the corresponding sending end. For example, Figure 6It includes narrowband voice channels 1, ……, narrowband voice channel M. Narrowband voice channel 1 is used to receive and cache the bitstream of the narrowband voice signal sent by transmitter 1, ……, narrowband voice channel M is used to receive and cache the bitstream of the narrowband voice signal sent by transmitter M.
[0092] Then, each voice coding frame in the bitstream is decoded by a narrowband decoder to obtain a narrowband voice signal frame (the length of each frame is approximately 10 ms to 30 ms). Then, the narrowband voice signal frames of the M narrowband voice channels are mixed in narrowband to form a narrowband voice signal frame (i.e., the narrowband voice signal frame after narrowband mixing). This narrowband voice signal frame after narrowband mixing is used as the narrowband voice signal frame to be processed. After performing voice bandwidth expansion on the narrowband voice signal frame to be processed, it is sent to the speaker for playback. The narrowband decoder can be a decoder that only supports narrowband voice decoding, such as a decoder conforming to the G.729a standard, etc., or it can be a wideband voice decoder operating in narrowband mode, such as opus decoding to obtain a narrowband voice signal in low bitrate mode.
[0093] By the above method, only one narrowband voice signal frame after narrowband mixing is subjected to voice bandwidth expansion, without separately performing voice bandwidth expansion on the narrowband voice signal frames from each narrowband voice channel, thereby reducing the computational load and improving the efficiency of voice bandwidth expansion.
[0094] In a possible implementation, the receiver may receive narrowband voice signals and wideband voice signals simultaneously. Refer to Figure 7 as shown. Figure 7 shows a system block diagram of the receiver receiving narrowband voice signals and wideband voice signals simultaneously. At this time, the receiver not only receives the bitstreams of the narrowband voice signals sent by M transmitters, but also the bitstreams of the wideband voice signals sent by at most N other transmitters. The bitstreams of the narrowband voice signals sent by the M transmitters are respectively received and cached by M narrowband voice channels, and the bitstreams of the wideband voice signals sent by the N transmitters are respectively received and cached by N wideband voice channels (wideband voice channels 1, ……, wideband voice channel N).
[0095] In this case, since the N transmitters are already sending the bitstreams of broadband voice signals, the method for obtaining the narrowband voice signal frames to be processed can be to obtain the target voice signal frames from different voice channels, and determine which target voice signal frames of which voice channels are narrowband voice signal frames and which are broadband voice signal frames based on the energy values of the high-frequency signals. If the energy value of the high-frequency signal is less than the preset energy threshold, it can be determined that the target voice signal frame is a narrowband voice signal frame. Therefore, the target voice signal frames with energy values of the high-frequency signals less than the preset energy threshold can be selected from the target voice signal frames from different voice channels as the narrowband voice signal frames to be processed. At this time, the corresponding voice channel is a narrowband voice channel, such as Figure 7 the narrowband voice channel 1, …… narrowband voice channel M in
[0096] If the energy value of the high-frequency signal reaches the preset energy threshold, it can be determined that the target voice signal frame is a broadband voice signal frame. Therefore, the target voice signal frames with energy values of the high-frequency signals reaching the preset energy threshold are selected from the target voice signal frames from different voice channels as the target broadband voice signal frames (such as Figure 7 the target voice signal frames obtained after decoding by the broadband decoder 1, …… broadband decoder N in Figure 7 the broadband voice channel 1, …… broadband voice channel N in). Then, the broadband voice signal frames obtained by performing synthesis filtering are mixed with the target broadband voice signal frames in broadband, and played according to the broadband voice signal frames obtained after broadband mixing.
[0097] Through the above method, voice bandwidth expansion can be performed only on the narrowband voice signal frames, avoiding voice bandwidth expansion on the broadband voice signal frames, thereby avoiding wasting computing resources and solving the technical problem of reducing voice quality due to voice bandwidth expansion on the broadband voice signal frames.
[0098] It should be noted that one implementation method of narrowband mixing and broadband mixing can be to add the corresponding sample points in this frame of data by weighting. For example, if the sampling rate of the narrowband voice signal frame is 8 kHz and each frame is 20 ms, then each frame has 8 * 20 = 160 sample points, that is, the narrowband voice signal frame X1 of the first narrowband voice channel = [x1(1), x1(2), …, x1(160)], the narrowband voice signal frame X2 of the second narrowband voice channel = [x2(1), x2(2), …, x2(160)], and so on. Suppose there are narrowband voice signal frames from M narrowband voice channels that need to be mixed in narrowband, then the narrowband voice signal frame after narrowband mixing is: where w m(j), where j = 1, 2, …, 160 are preset or dynamically set weights, and x m (j) represents the amplitude of the narrowband speech signal corresponding to the m-th narrowband speech channel at the j-th sample point.
[0099] The way of wideband mixing is similar to that of narrowband mixing, only replacing the relevant representation of the narrowband speech signal frame with the relevant representation of the wideband speech signal frame, which will not be elaborated here in detail.
[0100] Next, the speech bandwidth expansion method provided by the embodiments of the present application will be introduced in combination with an actual application scenario. In the scenario of multi-person audio and video calls based on instant messaging software, due to some reasons such as insufficient network bandwidth or serious packet loss, the narrowband speech signal may be sent from the sending end to the receiving end. In this case, in order to prevent the receiving end from being affected by the lack of high-frequency information in the narrowband speech signal and affecting the subjective experience of the user at the receiving end when listening to the speech, the embodiments of the present application provide a speech bandwidth expansion method, which specifically includes:
[0101] After the narrowband speech signal frame to be processed undergoes feature extraction, together with a pre-trained prediction model, high-frequency power spectrum prediction is performed to obtain the predicted high-frequency power spectrum. At the same time, the narrowband speech signal frame to be processed after undergoing upsampling processing is obtained for calculating the low-frequency power spectrum. The wideband power spectrum synthesis module splices the previously calculated low-frequency power spectrum and the predicted high-frequency power spectrum to obtain the wideband power spectrum. The wideband spectral envelope calculation module calculates the spectral gain and envelope according to the wideband power spectrum, specifically by calculating the linear prediction error and the linear prediction coefficient. According to the obtained linear prediction error and linear prediction coefficient, the filter parameters are predicted, and then the obtained filter parameters can be used for analysis filtering and synthesis filtering. The analysis filtering module performs analysis filtering on the narrowband speech signal frame after upsampling processing to obtain the narrowband excitation signal. The wideband excitation signal generation module generates a high-frequency excitation signal according to the narrowband excitation signal and merges it with the original narrowband excitation signal to obtain the wideband excitation signal. The excitation signal shaping module shapes the wideband excitation signal and then re-synthesizes the wideband speech signal frame through the synthesis filtering module. Finally, the synthesized wideband speech signal frame is sent to the speaker at the receiving end for playback. In this way, the user at the receiving end hears the wideband speech signal frame with the high-frequency information complemented, thereby improving the subjective experience of the user when listening to the speech.
[0102] In addition, since the human vocal tract is simulated using filter parameters, a broadband speech signal frame is obtained by performing synthesis filtering based on a broadband excitation signal and filter parameters. This is equivalent to processing the vibrations generated by the throat through the corresponding vocal tract, and then obtaining a broadband speech signal frame similar to human vocalization. It can be seen that the speech bandwidth expansion method provided by this solution can obtain a broadband speech signal frame with a listening experience similar to human vocalization, which can reflect the speaker's language characteristics, thereby improving the fidelity of the broadband speech signal frame and the speech quality.
[0103] Based on the speech bandwidth expansion method provided in the foregoing embodiments, an embodiment of the present application further provides a speech bandwidth expansion device. Refer to Figure 8 , the speech bandwidth expansion device 800 includes a prediction unit 801, a splicing unit 802, a determination unit 803, a generation unit 804, and a filtering unit 805:
[0104] The prediction unit 801 is configured to perform high-frequency power spectrum prediction on the narrowband speech signal frame to be processed, and obtain the high-frequency power spectrum corresponding to the narrowband speech signal frame to be processed;
[0105] The splicing unit 802 is configured to splice the high-frequency power spectrum and the low-frequency power spectrum corresponding to the narrowband speech signal frame to be processed to obtain a broadband power spectrum;
[0106] The determination unit 803 is configured to calculate the spectral envelope of the broadband power spectrum, and determine filter parameters according to the calculation result;
[0107] The generation unit 804 is configured to perform analysis filtering based on the narrowband speech signal frame to be processed and the filter parameters to generate a low-frequency excitation signal corresponding to the narrowband speech signal frame to be processed, and the analysis filtering is used to filter out the timbre information and vocal tract shape information in the narrowband speech signal frame to be processed;
[0108] The determination unit 803 is further configured to expand the low-frequency excitation signal to obtain a high-frequency excitation signal; synthesize the low-frequency excitation signal and the high-frequency excitation signal to obtain a broadband excitation signal;
[0109] The filtering unit 805 is configured to perform synthesis filtering based on the broadband excitation signal and the filter parameters to obtain a broadband speech signal frame, and the synthesis filtering is used to add timbre information and vocal tract shape information to the broadband excitation signal.
[0110] In a possible implementation manner, the determination unit 803 is configured to:
[0111] Calculate the spectral envelope of the broadband power spectrum to obtain a linear prediction error and a linear prediction coefficient, and the linear prediction error and the linear prediction coefficient are the calculation results;
[0112] Determine the filter parameters according to the linear prediction error and the linear prediction coefficients.
[0113] In a possible implementation, the determining unit 803 is specifically configured to:
[0114] Perform high-pass filtering on the high-frequency excitation signal to obtain a high-frequency excitation signal with low-frequency noise filtered out;
[0115] Perform delay compensation on the low-frequency excitation signal;
[0116] Add the high-frequency excitation signal with low-frequency noise filtered out and the delay-compensated low-frequency excitation signal to obtain the wideband excitation signal.
[0117] In a possible implementation, the apparatus further includes:
[0118] An upsampling unit, configured to perform upsampling processing on the to-be-processed narrowband speech signal frame before splicing the high-frequency power spectrum and the low-frequency power spectrum corresponding to the to-be-processed narrowband speech signal frame to obtain a wideband power spectrum;
[0119] A calculation unit, configured to calculate the low-frequency power spectrum based on the upsampled to-be-processed narrowband speech signal frame;
[0120] The generating unit 804 is specifically configured to:
[0121] Perform analysis filtering based on the upsampled to-be-processed narrowband speech signal frame and the filter parameters to generate a low-frequency excitation signal corresponding to the narrowband speech signal frame.
[0122] In a possible implementation, the apparatus further includes an obtaining unit, and the obtaining unit obtaining the to-be-processed narrowband speech signal frame includes:
[0123] Obtain target speech signal frames from different speech channels;
[0124] Select, from the target speech signal frames from different speech channels, a target speech signal frame whose energy value of the high-frequency signal is less than a preset energy threshold as the to-be-processed narrowband speech signal frame.
[0125] In a possible implementation, the apparatus further includes:
[0126] A selecting unit, configured to select, from the target speech signal frames from different speech channels, a target speech signal frame whose energy value of the high-frequency signal reaches the preset energy threshold as a target wideband speech signal frame;
[0127] A broadband mixing unit for broadband mixing the broadband speech signal frames obtained by synthesis filtering with the target broadband speech signal frames;
[0128] A playback unit for playing according to the broadband speech signal frames obtained after broadband mixing.
[0129] In a possible implementation manner, the apparatus further includes an acquisition unit, and the acquisition unit acquires the to-be-processed narrowband speech signal frames, including:
[0130] Performing narrowband mixing on the narrowband speech signal frames from different narrowband speech channels to obtain narrowband speech signal frames after narrowband mixing;
[0131] Taking the narrowband speech signal frames after narrowband mixing as the to-be-processed narrowband speech signal frames.
[0132] In a possible implementation manner, the prediction unit 801 is specifically configured to:
[0133] Performing feature extraction on the to-be-processed narrowband speech signal frames to obtain corresponding feature vectors;
[0134] Based on the feature vectors and a prediction model, obtaining the high-frequency power spectrum corresponding to the to-be-processed narrowband speech signal frames.
[0135] In a possible implementation manner, the prediction unit 801 is specifically configured to:
[0136] Dividing the spectrum of the to-be-processed narrowband speech signal frames into multiple sub-bands;
[0137] Calculating the power spectrum corresponding to each sub-band respectively;
[0138] Determining the feature vectors according to the power spectrum.
[0139] In a possible implementation manner, the prediction unit 801 is specifically configured to:
[0140] Inputting the feature vectors into the prediction model, and outputting the power values of each frequency point of the to-be-processed narrowband speech signal frames in the high-frequency band through the prediction model;
[0141] Based on the power values of each frequency point of the to-be-processed narrowband speech signal frames in the high-frequency band, constructing the high-frequency power spectrum;
[0142] Alternatively, inputting the feature vectors into the prediction model, and outputting the average power values of each sub-band of the to-be-processed narrowband speech signal frames in the high-frequency band through the prediction model;
[0143] Construct the high-frequency power spectrum based on the average power value of each sub-band of the narrowband voice signal frame to be processed in the high-frequency band.
[0144] As can be seen from the above technical solution, after obtaining the narrowband voice signal frame to be processed, the high-frequency power spectrum is predicted according to the narrowband voice signal frame to be processed, and the high-frequency power spectrum corresponding to the narrowband voice signal frame to be processed is obtained. The high-frequency power spectrum and the low-frequency power spectrum corresponding to the narrowband voice signal frame to be processed are spliced to obtain the broadband power spectrum, so as to expand the bandwidth of the narrowband voice signal frame to be processed to a certain extent. Based on the analysis of human voice generation, the human vocal tract (oral cavity, throat) is equivalent to a filter, and the vibrations generated by the vocal tract generate voice signals with different timbres through different-shaped vocal tracts. Therefore, the present application calculates the spectral envelope of the broadband power spectrum, determines the filter parameters according to the calculation results, and simulates the human vocal tract through the filter parameters. Then, based on the narrowband voice signal frame to be processed and the filter parameters, analysis filtering is performed to filter out the timbre information and vocal tract shape information in the narrowband voice signal frame to be processed to generate a low-frequency excitation signal corresponding to the narrowband voice signal frame to be processed. The low-frequency excitation signal is expanded to obtain a high-frequency excitation signal, and then the broadband excitation signal is determined according to the low-frequency excitation signal and the high-frequency excitation signal. This broadband excitation signal is equivalent to the vibration generated by the vocal tract during the human voice generation process. Since the filter parameters simulate the human vocal tract, synthetic filtering is performed based on the broadband excitation signal and the filter parameters to add timbre information and vocal tract shape information to the broadband excitation signal to obtain a broadband voice signal frame, which is equivalent to processing the vibration generated by the vocal tract through the corresponding vocal tract, and then obtaining a broadband voice signal frame similar to human vocalization. It can be seen that the voice bandwidth expansion method provided by this solution can obtain a broadband voice signal frame with a listening feeling similar to human vocalization, which can reflect the language characteristics of the speaker, thereby improving the fidelity of the broadband voice signal frame and the voice quality.
[0145] The embodiment of the present application also provides a device for voice bandwidth expansion. This device can be a receiving end. Taking the receiving end as a smart phone as an example:
[0146] Figure 9 What is shown is a block diagram of a part of the structure of the smart phone provided by the embodiment of the present application. Refer to Figure 9, the smart phone includes components such as a Radio Frequency (RF) circuit 910, a memory 920, an input unit 930, a display unit 940, a sensor 950, an audio circuit 960, a wireless fidelity (WiFi) module 970, a processor 980, and a power supply 990. The input unit 930 may include a touch panel 931 and other input devices 932, the display unit 940 may include a display panel 941, and the audio circuit 960 may include a speaker 961 and a microphone 962. Those skilled in the art can understand that Figure 9 the structure of the smart phone shown in
[0147] does not limit the smart phone and may include more or fewer components than shown in the figure, or combine certain components, or have different component arrangements. The memory 920 can be used to store software programs and modules. The processor 980 executes various functional applications and data processing of the smart phone by running the software programs and modules stored in the memory 920. The memory 920 may mainly include a program storage area and a data storage area. Among them, the program storage area can store an operating system, application programs required for at least one function (such as a sound playback function, an image playback function, etc.); the data storage area can store data created according to the use of the smart phone (such as audio data, a phone book, etc.). In addition, the memory 920 may include a high-speed random access memory and may also include a non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other non-volatile solid-state storage devices.
[0148] The processor 980 is the control center of the smart phone, connects various parts of the entire smart phone using various interfaces and lines, executes various functions of the smart phone and processes data by running or executing the software programs and / or modules stored in the memory 920, and calling the data stored in the memory 920, thereby monitoring the smart phone as a whole. Optionally, the processor 980 may include one or more processing units; preferably, the processor 980 may integrate an application processor and a modem processor. Among them, the application processor mainly processes the operating system, user interface, and application programs, etc., and the modem processor mainly processes wireless communication. It can be understood that the above-mentioned modem processor may not be integrated into the processor 980.
[0149] In this embodiment, the processor 980 in the smart phone may execute the following steps:
[0150] Perform high-frequency power spectrum prediction according to the narrowband voice signal frame to be processed, and obtain the high-frequency power spectrum corresponding to the narrowband voice signal frame to be processed;
[0151] Concatenate the high-frequency power spectrum and the low-frequency power spectrum corresponding to the narrowband speech signal frame to be processed to obtain a broadband power spectrum;
[0152] Calculate the spectral envelope of the broadband power spectrum and determine the filter parameters according to the calculation results;
[0153] Perform analysis filtering based on the narrowband speech signal frame to be processed and the filter parameters to generate a low-frequency excitation signal corresponding to the narrowband speech signal frame to be processed, where the analysis filtering is used to filter out the timbre information and vocal tract shape information in the narrowband speech signal frame to be processed;
[0154] Expand the low-frequency excitation signal to obtain a high-frequency excitation signal;
[0155] Synthesize according to the low-frequency excitation signal and the high-frequency excitation signal to obtain a broadband excitation signal;
[0156] Perform synthesis filtering based on the broadband excitation signal and the filter parameters to obtain a broadband speech signal frame, where the synthesis filtering is used to add timbre information and vocal tract shape information to the broadband excitation signal.
[0157] The embodiments of the present application further provide a server. Please refer to Figure 10 as shown Figure 10 is a structural diagram of the server 1000 provided by the embodiments of the present application. The server 1000 may vary greatly due to configuration or performance differences, and may include one or more central processing units (CPUs for short) 1022 (for example, one or more processors) and a memory 1032, and one or more storage media 1030 for storing application programs 1042 or data 1044 (for example, one or more mass storage devices). Among them, the memory 1032 and the storage media 1030 can be transient storage or persistent storage. The program stored in the storage media 1030 may include one or more modules (not marked in the figure), and each module may include a series of instruction operations on the server. Further, the central processor 1022 may be set to communicate with the storage media 1030 and execute a series of instruction operations in the storage media 1030 on the server 1000.
[0158] The server 1000 may further include one or more power supplies 1026, one or more wired or wireless network interfaces 1050, one or more input / output interfaces 1058, and / or, one or more operating systems 1041, such as Windows Server TM , Mac OS X TM , Unix TM , LinuxTM , FreeBSD TM and so on.
[0159] The steps performed by the server in the foregoing embodiments may be implemented based on Figure 10 the structure shown.
[0160] According to one aspect of the present application, there is provided a computer-readable storage medium for storing program code for executing the voice bandwidth expansion method described in each of the foregoing embodiments.
[0161] According to one aspect of the present application, there is provided a computer program product or computer program, which includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the methods provided in various alternative implementations of the foregoing embodiments.
[0162] The terms "first", "second", "third", "fourth", etc. (if any) in the specification of the present application and the above drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way may be interchanged under appropriate circumstances, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0163] In several embodiments provided by the present application, it should be understood that the disclosed systems, devices and methods may be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division, and there may be other division methods in actual implementation. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other may be through some interfaces, and the indirect coupling or communication connection of the device or unit may be in an electrical, mechanical or other form.
[0164] The unit described as a separation component may or may not be physically separated. The component shown as a unit may or may not be a physical unit, that is, it may be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0165] In addition, each functional unit in various embodiments of the present application may be integrated in a processing unit, may exist separately as individual physical units, or two or more units may be integrated in one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of a software functional unit.
[0166] If the above-mentioned integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present application. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs that can store program codes.
[0167] As described above, the above embodiments are only used to illustrate the technical solutions of the present application and are not intended to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of various embodiments of the present application.
Claims
1. A method for voice bandwidth expansion, characterized in that, The method includes: Performing high-frequency power spectrum prediction on the narrowband speech signal frame to be processed to obtain the high-frequency power spectrum corresponding to the narrowband speech signal frame to be processed; Stitching the high-frequency power spectrum and the low-frequency power spectrum corresponding to the narrowband speech signal frame to be processed to obtain a broadband power spectrum; Calculating the spectral envelope of the broadband power spectrum and determining filter parameters according to the calculation result; Performing analysis filtering based on the narrowband speech signal frame to be processed and the filter parameters to generate a low-frequency excitation signal corresponding to the narrowband speech signal frame to be processed, where the analysis filtering is used to filter out timbre information and vocal tract shape information in the narrowband speech signal frame to be processed; Expanding the low-frequency excitation signal to obtain a high-frequency excitation signal; Performing synthesis based on the low-frequency excitation signal and the high-frequency excitation signal to obtain a broadband excitation signal; Performing synthesis filtering based on the broadband excitation signal and the filter parameters to obtain a broadband speech signal frame, where the synthesis filtering is used to add timbre information and vocal tract shape information to the broadband excitation signal.
2. The method according to claim 1, characterized in that Calculating the spectral envelope of the broadband power spectrum and determining filter parameters according to the calculation result, including: Calculating the spectral envelope of the broadband power spectrum to obtain a linear prediction error and linear prediction coefficients, where the linear prediction error and the linear prediction coefficients are the calculation results; Determining the filter parameters according to the linear prediction error and the linear prediction coefficients.
3. The method according to claim 1, wherein The performing synthesis based on the low-frequency excitation signal and the high-frequency excitation signal to obtain a broadband excitation signal includes: Performing high-pass filtering on the high-frequency excitation signal to obtain a high-frequency excitation signal with low-frequency noise filtered out; Performing delay compensation on the low-frequency excitation signal; Adding the high-frequency excitation signal with low-frequency noise filtered out and the delay-compensated low-frequency excitation signal to obtain the broadband excitation signal.
4. The method according to claim 1, characterized in that, Before stitching the high-frequency power spectrum and the low-frequency power spectrum corresponding to the narrowband speech signal frame to be processed to obtain a broadband power spectrum, the method further includes: Performing upsampling processing on the narrowband speech signal frame to be processed; Calculating the low-frequency power spectrum based on the narrowband speech signal frame after upsampling processing; The performing analysis filtering based on the narrowband speech signal frame to be processed and the filter parameters to generate a low-frequency excitation signal corresponding to the narrowband speech signal frame to be processed includes: Performing analysis filtering based on the narrowband speech signal frame after upsampling processing and the filter parameters to generate a low-frequency excitation signal corresponding to the narrowband speech signal frame.
5. The method according to any one of claims 1 to 4, characterized in that Obtaining the narrowband speech signal frame to be processed includes: Obtaining target speech signal frames from different speech channels; Selecting, from the target speech signal frames from different speech channels, a target speech signal frame with an energy value of the high-frequency signal less than a preset energy threshold as the narrowband speech signal frame to be processed.
6. The method according to claim 5, wherein The method further includes: Selecting, from the target speech signal frames from different speech channels, a target speech signal frame with an energy value of the high-frequency signal reaching the preset energy threshold as a target broadband speech signal frame; Performing broadband mixing on the broadband speech signal frame obtained by performing synthesis filtering and the target broadband speech signal frame. Play according to the broadband voice signal frames obtained after broadband mixing.
7. The method according to any one of claims 1-4, characterized in that, Obtain the narrowband voice signal frames to be processed, including: Perform narrowband mixing on the narrowband voice signal frames from different narrowband voice channels to obtain narrowband voice signal frames after narrowband mixing; Use the narrowband voice signal frames after narrowband mixing as the narrowband voice signal frames to be processed.
8. The method according to any one of claims 1 to 4, characterized in that, The high-frequency power spectrum prediction based on the narrowband voice signal frames to be processed to obtain the high-frequency power spectrum corresponding to the narrowband voice signal frames to be processed includes: Extract features from the narrowband voice signal frames to be processed to obtain corresponding feature vectors; Based on the feature vectors and the prediction model, obtain the high-frequency power spectrum corresponding to the narrowband voice signal frames to be processed.
9. The method according to claim 8, characterized in that, The extracting features from the narrowband voice signal frames to be processed to obtain corresponding feature vectors includes: Divide the spectrum of the narrowband voice signal frames to be processed into multiple subbands; Calculate the power spectrum corresponding to each subband; Determine the feature vectors according to the power spectrum.
10. The method according to claim 8, characterized in that, The obtaining the high-frequency power spectrum corresponding to the narrowband voice signal frames to be processed based on the feature vectors and the prediction model includes: Input the feature vectors into the prediction model, and output the power values of each frequency point of the narrowband voice signal frames to be processed in the high-frequency band through the prediction model; Construct the high-frequency power spectrum based on the power values of each frequency point of the narrowband voice signal frames to be processed in the high-frequency band; Alternatively, input the feature vectors into the prediction model, and output the average power values of each subband of the narrowband voice signal frames to be processed in the high-frequency band through the prediction model; Construct the high-frequency power spectrum based on the average power values of each subband of the narrowband voice signal frames to be processed in the high-frequency band.
11. A voice bandwidth expansion device, characterized in that, The device includes a prediction unit, a splicing unit, a determination unit, a generation unit, and a filtering unit: The prediction unit is used to perform high-frequency power spectrum prediction according to the narrowband voice signal frames to be processed to obtain the high-frequency power spectrum corresponding to the narrowband voice signal frames to be processed; The splicing unit is used to splice the high-frequency power spectrum and the low-frequency power spectrum corresponding to the narrowband voice signal frames to be processed to obtain a broadband power spectrum; The determination unit is used to calculate the spectral envelope of the broadband power spectrum and determine the filter parameters according to the calculation results; The generation unit is used to perform analysis filtering based on the narrowband voice signal frames to be processed and the filter parameters to generate a low-frequency excitation signal corresponding to the narrowband voice signal frames to be processed, and the analysis filtering is used to filter out the timbre information and vocal tract shape information in the narrowband voice signal frames to be processed; The determination unit is further used to expand the low-frequency excitation signal to obtain a high-frequency excitation signal; synthesize the broadband excitation signal according to the low-frequency excitation signal and the high-frequency excitation signal; The filtering unit is used to perform synthesis filtering based on the broadband excitation signal and the filter parameters to obtain broadband voice signal frames, and the synthesis filtering is used to add timbre information and vocal tract shape information to the broadband excitation signal.
12. A device for voice bandwidth expansion, characterized in that, The device includes a processor and a memory: The memory is used to store program code and transmit the program code to the processor; The processor is used to execute the method according to any one of claims 1-10 based on the instructions in the program code.
13. A computer-readable storage medium, characterized in that, The computer-readable storage medium is used to store program code, and the program code is used to execute the method according to any one of claims 1-10.
14. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the method according to any one of claims 1-10.
Citation Information
Patent Citations
Sound synthesizer
CA1170370A
Excitation signal generation in bandwidth spreading and signal reconstruction method and apparatus
CN101458930A