Voice bandwidth expansion method oriented to actual communication scene

By combining traditional source filter models and lightweight deep neural networks, a cascading processing framework is used to expand voice bandwidth, which solves the problems of computing complexity and model size in the existing technology, and achieves efficient and real-time voice bandwidth expansion effect.

CN120108415APending Publication Date: 2025-06-06GUANGZHOU UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510264624.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-06
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

Existing voice bandwidth scaling technologies have limitations in computing complexity and model size, making it difficult to achieve real-time deployment on mobile devices.

Method used

A cascaded processing framework is adopted, and the traditional voice bandwidth expansion technology based on the source filter model is initially expanded, and then a lightweight deep neural network model is built, and the model performance is optimized through multi-scale discriminator adversarial training, and finally refined spectrum optimization and feature enhancement are carried out through deep neural networks.

Benefits of technology

It significantly reduces the computational complexity, simplifies the model structure, implements high-quality voice bandwidth expansion, and supports real-time processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120108415A_ABST
    Figure CN120108415A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of voice signal processing, and particularly discloses an actual communication scene-oriented voice bandwidth expansion method, which adopts a coarse-to-fine double-layer architecture, and comprises the following steps of: firstly, carrying out preliminary expansion on a narrow-band voice signal by utilizing a traditional voice bandwidth expansion technology based on a source filter model to obtain a roughly expanded voice; then constructing a lightweight deep neural network model, wherein the model comprises an encoder module, a decoder module and a lightweight packet LSTM (GLSTM) module; in the model training process, a multi-scale discriminator is used for adversarial training, and the model performance is optimized; and finally, performing refined spectrum optimization and feature enhancement on the roughly expanded speech through the optimized lightweight deep neural network model, and outputting high-quality broadband speech. According to the invention, the traditional signal processing and lightweight deep learning technologies are fused, the calculation complexity is significantly reduced, the high quality of voice expansion is ensured, and the voice signals can be processed in real time.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of voice signal processing technology, and more specifically, to a voice bandwidth extension method for actual communication scenarios. Background Art

[0002] With the development of communication technology, users' requirements for voice communication quality are constantly increasing. Traditional voice communication systems are limited to a narrow bandwidth of 4kHz, which results in the loss of high-frequency components of voice, dull sound and insufficient clarity. In order to improve voice quality, the International Telecommunication Union has proposed a broadband voice transmission standard (8kHz bandwidth), but broadband transmission requires high costs and complex infrastructure support. Therefore, voice bandwidth extension technology has emerged. By restoring or synthesizing the missing high-frequency components at the receiving end, it can be deployed directly on the terminal device without modifying the transmission network.

[0003] The core of speech bandwidth extension is to reconstruct the high-frequency components of 4-8kHz from narrowband signals. Early methods were based on the source filter model, decomposing and expanding speech signals through linear predictive coding, but they suffered from insufficient spectrum continuity and perceived quality defects. In recent years, deep learning methods have made breakthroughs due to their powerful nonlinear mapping capabilities, but their high computational complexity and large models limit their real-time deployment on mobile devices.

[0004] Therefore, a voice bandwidth extension method for actual communication scenarios is provided. Summary of the invention

[0005] In order to solve the above technical problems, this application is proposed.

[0006] Specifically, according to one aspect of the present application, a voice bandwidth extension method for an actual communication scenario is provided, which includes:

[0007] S1. Using the traditional speech bandwidth extension technology based on the source filter model to preliminarily expand the input narrowband speech signal to obtain a roughly expanded speech;

[0008] S2. Construct a lightweight deep neural network model, wherein the model includes an encoder module, a decoder module, and a lightweight grouped LSTM (GLSTM) module;

[0009] S3. During the training process of the lightweight deep neural network model, using a multi-scale discriminator to perform adversarial training to optimize model performance;

[0010] S4. Use a lightweight deep neural network model optimized through adversarial training to perform refined spectrum optimization and feature enhancement on the coarsely extended speech to obtain broadband speech.

[0011] Preferably, S1 comprises: S11, performing a preprocessing operation of framing and windowing on the narrowband speech signal; S12, performing zero interpolation upsampling and low-pass filtering on the narrowband speech signal after framing and windowing in sequence to obtain the low-frequency part of the broadband speech signal; S13, passing the narrowband speech signal after framing and windowing through a high-frequency extension module and high-pass filtering in sequence to obtain the high-frequency part of the broadband speech signal; S14, time-aligning and superimposing the low-frequency part of the broadband speech signal and the high-frequency part of the broadband speech signal to obtain a roughly extended speech.

[0012] Preferably, in S2, the encoder module and the lightweight grouped LSTM (GLSTM) module share parameters, the encoder module includes a ConvGLU module, and the decoder module includes a DeconvGLU module; wherein, the ConvGLU module and the DeconvGLU module both include a convolutional layer and a gated linear unit (GLU), the ConvGLU module divides the input into two branches for processing, one of which obtains the output of the branch through a convolution operation and a gating mechanism, and the other branch only uses a convolution operation; the DeconvGLU module also divides the input into two branches for processing, one of which obtains the output of the branch through a deconvolution operation and a gating mechanism, and the other branch only uses a deconvolution operation.

[0013] Preferably, in S3, the multi-scale discriminator comprises a plurality of parallel sub-discriminators, which share the same network topology and process speech features of different time resolutions respectively.

[0014] Preferably, the S4 comprises: S41, using short-time Fourier transform (STFT) to extract the amplitude spectrum representation of the roughly extended speech and the interpolated speech respectively; S42, stacking the amplitude spectrum representation of the roughly extended speech and the interpolated speech together as the input of the lightweight deep neural network model; S43, using the ConvGLU module through the encoder module to downsample the input and extract features; S44, further extracting and processing features through the lightweight grouped LSTM module; S45, using the DeconvGLU module through the decoder module to upsample and predict the masks of the roughly extended speech and the interpolated speech respectively; S46, multiplying the two predicted masks and the two amplitude spectrum representations element by element to obtain two optimized amplitude spectrum components; S47, adding the two optimized amplitude spectrum components as the amplitude spectrum of the final output of the model; S48, combining the amplitude spectrum of the final output of the model with the phase spectrum of the interpolated speech; S49, using the inverse short-time Fourier transform (I The combined spectral representation is converted back to the time domain signal using STFT to obtain wideband speech.

[0015] Compared with the prior art, the present application provides a voice bandwidth extension method for actual communication scenarios. The method adopts an efficient cascade processing framework. In the first stage, the input narrowband voice signal is preliminarily bandwidth expanded using a traditional signal processing method based on a source filter model. The model extracts the spectral envelope and excitation characteristics of the voice signal through linear prediction analysis, and uses efficient band replication technology to generate preliminary high-frequency components. This step can achieve a rough reconstruction of the spectrum at an extremely low computational cost, providing a good initial estimate for subsequent processing. In the second stage, a lightweight deep neural network is designed to perform refined spectrum optimization and feature enhancement on the preliminarily expanded voice signal. Through this staged processing strategy, the deep neural network only needs to focus on compensating for the residual information, thereby greatly reducing the computational complexity of the model. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] By describing the embodiments of the present application in more detail in conjunction with the accompanying drawings, the above and other purposes, features and advantages of the present application will become more apparent. The accompanying drawings are used to provide a further understanding of the embodiments of the present application and constitute a part of the specification. Together with the embodiments of the present application, they are used to explain the present application and do not constitute a limitation of the present application. In the accompanying drawings, the same reference numerals generally represent the same components or steps.

[0017] Figure 1 The figure illustrates a technical route flow chart according to an embodiment of the present application.

[0018] Figure 2 The figure shows a schematic diagram of a traditional voice bandwidth extension method in the first stage according to an embodiment of the present application.

[0019] Figure 3 The figure illustrates a model structure diagram of the second stage according to an embodiment of the present application.

[0020] Figure 4 The diagram illustrates the structure diagrams of the ConvGLU module and the DeconvGLU module according to an embodiment of the present application.

[0021] Figure 5 The figure shows a structure diagram of a multi-scale discriminator according to an embodiment of the present application.

[0022] Figure 6 The figure illustrates a result diagram of voice bandwidth extension according to an embodiment of the present application. DETAILED DESCRIPTION

[0023] Below, the embodiments according to the present application will be described in detail with reference to the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments of the present application, and it should be understood that the present application is not limited to the exemplary embodiments described here.

[0024] Example:

[0025] Figure 1 The figure shows a technical route flow chart according to an embodiment of the present application, such as Figure 1 As shown in the figure, this method adopts an efficient cascade processing framework. In the first stage, the input narrowband speech signal is initially bandwidth expanded using the traditional signal processing method based on the source filter model to obtain roughly expanded speech. In the second stage, the roughly expanded speech obtained from the first stage is input into a lightweight deep neural network, and the powerful nonlinear mapping ability of the deep neural network is used to perform refined spectrum optimization and feature enhancement, and finally the expanded speech is output. At the same time, in order to further improve the quality of the expanded speech, a multi-scale discriminator is introduced for adversarial training.

[0026] like Figure 2 and Figure 3 As shown, according to the embodiment of the present application, the voice bandwidth extension method for actual communication scenarios includes a first stage (step S1) and a second stage (steps S2, S3 and S4), and the specific implementation process is: S1, using the traditional voice bandwidth extension technology based on the source filter model to preliminarily expand the input narrowband voice signal to obtain a roughly extended voice; S2, constructing a lightweight deep neural network model, the model includes an encoder module, a decoder module and a lightweight grouped LSTM (GLSTM) module; S3, in the training process of the lightweight deep neural network model, using a multi-scale discriminator to perform adversarial training to optimize the model performance; S4, using the lightweight deep neural network model optimized by adversarial training to perform refined spectrum optimization and feature enhancement on the roughly extended voice to obtain broadband voice.

[0027] In the embodiment of the present application, the S1 uses the traditional voice bandwidth extension technology based on the source filter model to perform preliminary expansion on the input narrowband voice signal to obtain roughly expanded voice. Specifically, the S1 includes: S11, pre-processing operation of frame windowing on the narrowband voice signal; S12, zero interpolation up-sampling and low-pass filtering on the narrowband voice signal after frame windowing in sequence to obtain the low-frequency part of the broadband voice signal; S13, passing the narrowband voice signal after frame windowing through a high-frequency expansion module and high-pass filtering in sequence to obtain the high-frequency part of the broadband voice signal; S14, time-aligning and superimposing the low-frequency part of the broadband voice signal and the high-frequency part of the broadband voice signal to obtain roughly expanded voice.

[0028] Among them, the high-frequency extension module includes: performing linear prediction analysis on the narrow-band speech signal after frame division and windowing to obtain a narrow-band speech spectrum envelope and a narrow-band excitation signal; performing spectrum mirror expansion on the narrow-band excitation signal to generate a wide-band excitation signal; and using linear prediction synthesis technology to synthesize the wide-band excitation signal and the narrow-band speech spectrum envelope into a high-frequency extension component.

[0029] That is, in the first stage, an efficient speech bandwidth extension method based on the source filter model is used to extend the high-frequency part. The core idea is to combine the extended excitation signal with the original spectrum envelope to generate an extended speech signal. This method only extends the excitation signal by means of spectral mirroring, without the need for additional statistical methods to process the spectrum envelope, so it introduces almost no delay and has little impact on the overall algorithm. In addition, this method does not contain any learnable parameters and only performs predefined operations on the input data, which effectively avoids the error accumulation problem common in the two-stage method. This design provides an alternative for subsequent processing, significantly reduces the difficulty of neural networks in estimating high-frequency components, simplifies the structure of the second-stage deep neural network, and reduces the complexity of the learning process without losing performance. More importantly, the preliminary extension method based on the source filter model can reasonably adjust the physical characteristics of speech, such as formant structure and harmonic characteristics, to ensure that the extended speech is more reasonable in acoustic characteristics, and provides a high-quality initial estimate for subsequent neural network processing.

[0030] In the embodiment of the present application, the S2, constructs a lightweight deep neural network model, the model includes an encoder module, a decoder module and a lightweight grouped LSTM (GLSTM) module. It should be understood that the encoder is responsible for extracting features, and the decoder is responsible for recovering signals. The two work together to effectively improve the quality of voice bandwidth extension. Therefore, constructing an encoder-decoder architecture can efficiently process sequence data and is widely used in speech processing, image processing and other fields. In addition, LSTM (long short-term memory network) can capture the time series information in the speech signal, and the grouped LSTM further reduces the amount of calculation by processing the features in groups, while retaining the long-term and short-term dependencies of the speech signal. Moreover, the lightweight design enables the GLSTM module to significantly reduce the computational complexity while maintaining performance, which is suitable for real-time speech processing. Therefore, when constructing a lightweight deep neural network model, a lightweight GLSTM module is further added to optimize the computational efficiency of the model, while retaining the time dependency of the speech signal, so that the model can better handle the dynamic changes of the speech signal.

[0031] Specifically, in S2, the lightweight deep neural network model consists of an encoder and two decoders, and two stacked lightweight grouped LSTMs are inserted between the encoder and the decoder. The encoder module and the GLSTM module share parameters, while the two decoders predict the masks of the two inputs respectively. The encoder module uses 5 ConvGLU modules to downsample the input, and the decoder module has the opposite structure to the encoder module and uses 5 DeconvGLU modules for upsampling. The linear layer is stacked after the last DeconvGLU block to convert the features into masks of the amplitude spectrogram. In addition, it is worth noting that all convolution operations use causal convolution, which means that no future information is used, ensuring the strict causality of the model when processing speech signals, thereby meeting the key requirements of real-time processing.

[0032] In particular, if Figure 4 As shown, the ConvGLU module and the DeconvGLU module are an improved structure in the convolutional neural network (CNN), which combines the convolution layer and the gated linear unit (GLU). It divides the input into two branches for processing, one of which obtains the output of the branch through convolution operation and gating mechanism, and the other branch only uses convolution operation, and then multiplies the outputs of the two branches element by element to obtain the final output. The DeconvGLU module also divides the input into two branches for processing, one of which obtains the output of the branch through deconvolution operation and gating mechanism, and the other branch only uses deconvolution operation. GLU introduces stronger nonlinearity and improves the model's expressiveness. In addition, the gating mechanism can effectively regulate the flow of information and improve model performance.

[0033] In the embodiment of the present application, in the training process of the lightweight deep neural network model, the multi-scale discriminator is used for adversarial training to optimize the model performance. It should be understood that in order to further improve the quality of the extended speech, the multi-scale discriminator is introduced for adversarial training. The structure of the discriminator module is as follows Figure 5 As shown in the figure, the discriminator module adopts a hierarchical multi-scale architecture, which consists of three parallel sub-discriminators that share the same network topology but process speech features of different time resolutions respectively. Specifically, hierarchical downsampling is achieved by implementing an average pooling (Avg Pool) operation with a kernel size of 4 on the original speech signal. This multi-scale processing mechanism enables each sub-discriminator to focus on speech features in different frequency bands: the first discriminator processes the original high-resolution signal to capture the fine time-frequency features of speech; the second discriminator processes the 2x downsampled signal and focuses on the medium-scale spectral features; the third discriminator processes the 4x downsampled signal and mainly learns the overall spectral profile of speech.

[0034] In particular, in S3, a weighted sum of reconstruction loss and adversarial loss is used as the training loss. Reconstruction loss is applied in the time domain and frequency domain. The time domain reconstruction loss uses the mean square error (MSE) between the expanded speech waveform and the true wideband speech waveform to match the overall shape and phase of the speech waveform, as shown in the following formula:

[0035]

[0036] Where y represents broadband speech, Indicates extended speech.

[0037] For the frequency domain reconstruction loss, a multi-resolution STFT loss is introduced, including the spectrum convergence loss between extended speech and wideband speech, and the L1 loss of the logarithmic magnitude spectrogram with different FFT analysis parameters. The formula is expressed as:

[0038]

[0039] Where S(x;θ) represents the magnitude spectrum of x, m is the number of resolutions, and θ i is the STFT parameter at each resolution, ||·|| F and ||·|| 1 They are the Frobenius norm and the L1 norm respectively.

[0040] In adversarial training, the generator G and the discriminator D are trained alternately. The discriminator is trained to distinguish between real speech and extended speech more effectively, while the generator strives to make the extended speech closer to the real speech. The adversarial loss is shown in the following formula:

[0041]

[0042] Where y represents the real wideband speech, and G(x) represents the extended speech generated by the proposed method.

[0043] In order to further improve the similarity between the samples generated by the generator and the true values, a feature matching loss is also introduced, which is defined as the L1 loss between the discriminator feature maps of the wideband speech waveform and the extended speech waveform, as shown in the following formula:

[0044]

[0045] Where T represents the number of layers in the discriminator, represents the feature map output of the i-th layer of the k-th discriminator block, N i Indicates the number of units in each layer.

[0046] To summarize, the losses of the generator and discriminator are as follows:

[0047]

[0048] Among them, λ wav ,λ STFT ,λ FM Set them to 100, 0.5, and 10 respectively.

[0049] In the embodiment of the present application, the S4 uses a lightweight deep neural network model optimized by adversarial training to perform refined spectrum optimization and feature enhancement on the roughly extended speech to obtain broadband speech. Specifically, the S4 includes: S41, using short-time Fourier transform (STFT) to extract the amplitude spectrum representation of the roughly extended speech and the interpolated speech respectively; S42, stacking the amplitude spectrum representation of the roughly extended speech and the interpolated speech together as the input of the lightweight deep neural network model; S43, using the ConvGLU module through the encoder module to downsample the input and extract features; S44, further extracting and processing features through a lightweight grouped LSTM module;

[0050] S45. Use the DeconvGLU module to perform upsampling through the decoder module to predict the masks of the coarsely extended speech and the interpolated speech respectively; S46. Multiply the two predicted masks and the two amplitude spectrum representations element by element to obtain two optimized amplitude spectrum components; S47. Add the two optimized amplitude spectrum components as the amplitude spectrum of the final output of the model; S48. Combine the amplitude spectrum of the final output of the model with the phase spectrum of the interpolated speech; S49. Use the inverse short-time Fourier transform (iSTFT) to convert the combined spectrum representation back to the time domain signal to obtain broadband speech.

[0051] In the second stage, the deep neural network only needs to focus on compensating the residual information of the preliminary expanded speech without learning the complete spectral mapping from scratch.

[0052] Therefore, the experiment focuses on the bandwidth expansion of communication speech, that is, expanding the narrowband speech with a bandwidth of 4kHz to the wideband speech with a bandwidth of 8kHz (increasing the sampling rate from 8kHz to 16kHz). The results of speech bandwidth expansion are shown in Figure 6 As shown in the figure, (a) is the time domain waveform of the narrowband speech signal, (b) is the time domain waveform of the broadband speech signal, and (c) is the time domain waveform of the extended speech signal generated by the proposed method. It can be seen that the extended speech and broadband speech have more details than the narrowband speech. (d) is the amplitude spectrum of the narrowband speech signal, (e) is the amplitude spectrum of the broadband speech signal, and (f) is the amplitude spectrum of the extended speech signal generated by the proposed method. It can be seen that the high-frequency part of the extended speech in Figure (f) is well restored.

[0053] In summary, the voice bandwidth extension method for actual communication scenarios according to the embodiment of the present application is explained, which adopts a "coarse to fine" two-layer architecture. First, the narrowband voice signal is preliminarily expanded using the traditional voice bandwidth extension technology based on the source filter model to obtain a roughly extended voice; then a lightweight deep neural network model is constructed, which includes an encoder module, a decoder module and a lightweight grouped LSTM (GLSTM) module; during the model training process, a multi-scale discriminator is used for adversarial training to optimize the model performance; finally, the roughly extended voice is refined by the optimized lightweight deep neural network model. Spectral optimization and feature enhancement are performed to output high-quality broadband voice. This application integrates traditional signal processing with lightweight deep learning technology, significantly reduces computational complexity, ensures the high quality of voice expansion, and can process voice signals in real time.

[0054] It is obvious to those skilled in the art that the present invention is not limited to the details of the above exemplary embodiments, and that the present invention can be implemented in other specific forms without departing from the spirit or essential characteristics of the present invention.

[0055] Finally, it should be noted that the above embodiments are only used to illustrate the technical solution of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solution of the present invention can be modified or replaced by equivalents without departing from the spirit of the technical solution of the present invention.

Claims

1. A voice bandwidth extension method for actual communication scenarios, characterized in that: include: S1. Using the traditional speech bandwidth extension technology based on the source filter model to preliminarily expand the input narrowband speech signal to obtain a roughly expanded speech; S2. Construct a lightweight deep neural network model, wherein the model includes an encoder module, a decoder module, and a lightweight grouped LSTM (GLSTM) module; S3. During the training process of the lightweight deep neural network model, using a multi-scale discriminator to perform adversarial training to optimize model performance; S4. Use a lightweight deep neural network model optimized through adversarial training to perform refined spectrum optimization and feature enhancement on the coarsely extended speech to obtain broadband speech.

2. The voice bandwidth extension method for actual communication scenarios according to claim 1, characterized in that: Said S1 comprises: S11, performing a preprocessing operation of framing and windowing on the narrowband speech signal; S12, performing zero interpolation up-sampling and low-pass filtering on the narrowband speech signal after frame division and windowing in sequence to obtain a low-frequency part of the broadband speech signal; S13, passing the narrowband speech signal after frame division and windowing through a high-frequency extension module and a high-pass filter in sequence to obtain a high-frequency part of the broadband speech signal; S14, time-aligning and superimposing the low-frequency part of the broadband speech signal and the high-frequency part of the broadband speech signal to obtain a roughly extended speech.

3. The voice bandwidth extension method for actual communication scenarios according to claim 2 is characterized in that: The high-frequency extension module comprises: Performing linear prediction analysis on the narrowband speech signal after frame division and windowing to obtain the narrowband speech spectrum envelope and narrowband excitation signal; Performing spectrum mirroring expansion on the narrowband excitation signal to generate a broadband excitation signal; The broadband excitation signal and the narrowband speech spectrum envelope are synthesized into a high-frequency extension component by using a linear prediction synthesis technique.

4. The voice bandwidth extension method for actual communication scenarios according to claim 3 is characterized in that: In S2, the encoder module and the lightweight grouped LSTM (GLSTM) module share parameters, the encoder module includes a ConvGLU module, and the decoder module includes a DeconvGLU module; Among them, the ConvGLU module and the DeconvGLU module both include a convolution layer and a gated linear unit (GLU). The ConvGLU module divides the input into two branches for processing, one of which obtains the output of the branch through a convolution operation and a gating mechanism, and the other branch only uses a convolution operation; the DeconvGLU module also divides the input into two branches for processing, one of which obtains the output of the branch through a deconvolution operation and a gating mechanism, and the other branch only uses a deconvolution operation.

5. The voice bandwidth extension method for actual communication scenarios according to claim 4, characterized in that: In S3, the multi-scale discriminator includes a plurality of parallel sub-discriminators, which share the same network topology and process speech features with different time resolutions respectively.

6. The voice bandwidth extension method for actual communication scenarios according to claim 5, characterized in that: In S3, a weighted sum of the reconstruction loss and the adversarial loss is used as the training loss, and the reconstruction loss is applied in the time domain and the frequency domain; The time domain reconstruction loss uses the mean square error (MSE) between the expanded speech waveform and the true wideband speech waveform to match the overall shape and phase of the speech waveform, and is expressed as: Where y represents broadband speech, Indicates extended speech; For the frequency domain reconstruction loss, a multi-resolution STFT loss is introduced, including the spectrum convergence loss between extended speech and wideband speech, and the L1 loss of the logarithmic magnitude spectrogram with different FFT analysis parameters. The formula is expressed as: Where S(x;θ) represents the magnitude spectrum of x, m is the number of resolutions, and θ i is the STFT parameter at each resolution, ||·|| F and ||·||1 are the Frobenius norm and L1 norm respectively; The adversarial loss is given by the following formula: Where y represents the real wideband speech, and G(x) represents the extended speech generated by the proposed method.

7. The voice bandwidth extension method for actual communication scenarios according to claim 6, characterized in that: In said S3, a feature matching loss is also introduced, which is defined as the L1 loss between the discriminator feature maps of the wideband speech waveform and the extended speech waveform, as shown in the following formula: Where T represents the number of layers in the discriminator, represents the feature map output of the i-th layer of the k-th discriminator block, N i Indicates the number of units in each layer.

8. The voice bandwidth extension method for actual communication scenarios according to claim 7, characterized in that: The losses of the generator and discriminator are shown in the following formulas: Among them, λ wav ,λ STFT ,λ FM Set them to 100, 0.5, and 10 respectively.

9. The voice bandwidth extension method for actual communication scenarios according to claim 8, characterized in that: The S4 comprises: S41, using short-time Fourier transform (STFT) to extract the amplitude spectrum representation of the roughly extended speech and the interpolated speech respectively; S42, stacking the amplitude spectrum representations of the roughly expanded speech and the interpolated speech together as input of the lightweight deep neural network model; S43, down-sample the input using the ConvGLU module through the encoder module to extract features; S44, further extracting and processing features through a lightweight grouping LSTM module; S45, upsampling is performed by using a DeconvGLU module through a decoder module, and masks of the roughly extended speech and the interpolated speech are predicted respectively; S46, multiplying the two predicted masks and the two amplitude spectrum representations element by element respectively to obtain two optimized amplitude spectrum components; S47, adding the two optimized amplitude spectrum components as the amplitude spectrum finally output by the model; S48, combining the amplitude spectrum finally output by the model with the phase spectrum of the interpolated speech; S49. Use inverse short-time Fourier transform (iSTFT) to convert the combined spectral representation back to a time domain signal to obtain wideband speech.

10. The voice bandwidth extension method for actual communication scenarios according to claim 9, characterized in that: The convolution operation of the lightweight deep neural network model is a causal convolution operation.