A speech spectrum reconstruction method, system, terminal, and medium combining time-domain half-wave rectification and weighted Gaussian mixture model decoder.
By combining time-domain half-wave rectification and weighted Gaussian mixture model decoder, the speech spectrum reconstruction method solves the problem of low reconstruction accuracy of high-frequency components on edge devices, and achieves high-efficiency, low-latency, and high-quality speech reconstruction results.
Patent Information
- Application Number
- CN202511588692.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-03
- Publication Date
- 2026-01-06
- Estimated Expiration
- 2045-11-03
AI Technical Summary
Existing voice bandwidth expansion technologies are difficult to deploy efficiently on edge devices, and the high-frequency component reconstruction accuracy is low in complex acoustic environments, failing to meet the demand for high-quality voice.
By combining time-domain half-wave rectification and weighted Gaussian mixture model decoder, high-frequency details of the positive half-cycle are extracted by performing time-domain half-wave rectification on low-resolution audio, and Gaussian component parameters are generated using weighted Gaussian mixture model decoder. The audio time-domain waveform is reconstructed by combining frame-level frequency distribution weights and phase.
It achieves high-fidelity, high-frequency reconstruction while reducing model complexity, making it suitable for edge devices, improving speech perception quality and reconstruction accuracy, and supporting causal inference and low-latency deployment.
Smart Images

Figure CN121054016B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of speech signal processing technology, and in particular to a speech spectrum reconstruction method, system, terminal, and medium that combines time-domain half-wave rectification and weighted Gaussian mixture model decoder. Background Technology
[0002] Bandwidth Extension (BWE) technology can reconstruct the missing high-frequency components in narrowband speech signals, significantly improving speech perception quality. It is widely used in telecommunications, voice assistants, and other scenarios, and is a key technology for improving the voice interaction experience. However, existing BWE technology has the following core problems:
[0003] 1. Performance limitations of traditional signal processing methods: Early BWE methods relied on spectral envelope modeling, statistical mapping, or nonlinear signal processing. Although they achieved high-frequency reconstruction by mining low-frequency correlations, they relied too much on statistical assumptions and simplified signal models. In complex acoustic environments, the high-frequency component reconstruction accuracy was low, making it difficult to meet the requirements of high-quality speech.
[0004] 2. Deployment difficulties of deep learning methods: In recent years, the mainstream deep learning BWE models are divided into frequency domain and time domain. Although the reconstruction effect is better than traditional methods, the models are large in scale and have high computational complexity, making it difficult to deploy on resource-constrained edge platforms such as mobile phones and IoT devices.
[0005] 3. Lightweight models still have room for optimization: There are already models that achieve adaptive bandwidth expansion through amplitude repair networks and phase optimization networks. The lightweight version of these models focuses on amplitude repair to reduce complexity, but there are still two shortcomings: First, the input signal is obtained directly from low-sampling-rate speech, resulting in low high-frequency information density, which limits the reconstruction accuracy; second, the model complexity still cannot meet the extreme lightweight requirements of edge devices.
[0006] In summary, existing BWE technologies struggle to balance high performance and high efficiency, necessitating a new BWE approach that balances high-fidelity reconstruction, low computational cost, and adaptability to edge deployment. Summary of the Invention
[0007] The technical problem to be solved by this invention is to address the above-mentioned deficiencies of the prior art by providing a speech spectrum reconstruction method, system, terminal, and medium that combines time-domain half-wave rectification and weighted Gaussian mixture model decoder. The technical solution adopted by this invention is as follows:
[0008] In a first aspect, the present invention provides a speech spectrum reconstruction method combining a time-domain half-wave rectifier and a weighted Gaussian mixture model decoder, the method comprising:
[0009] A low-resolution audio signal is acquired, a time-domain half-wave rectification operation is performed on the low-resolution audio signal to extract high-frequency details in the positive half-cycle, and a mixed amplitude spectrum is obtained based on the amplitude spectrum of the low-resolution audio signal and the amplitude spectrum of the rectified signal.
[0010] Based on the encoder, the output is the GRU feature corresponding to the mixed amplitude spectrum. The GRU feature is input to the weighted Gaussian mixture model decoder. Each Gaussian component is constrained by three parallel linear layers in the weighted Gaussian mixture model decoder and three constraint designs are applied to generate several sets of Gaussian component parameters.
[0011] The frame-level frequency distribution weights are calculated based on the Gaussian component parameters and combined with the mixed amplitude spectrum to obtain the model output amplitude spectrum. The model output amplitude spectrum and the amplitude spectrum of the low-resolution audio are input into the frequency band guiding masking module in the weighted Gaussian mixture model decoder to obtain the final amplitude spectrum. The audio time-domain waveform is reconstructed based on the phase of the final amplitude spectrum and the low-resolution audio through short-time Fourier inverse transform.
[0012] In one implementation, low-resolution audio is acquired, and a time-domain half-wave rectification operation is performed on the low-resolution audio to extract high-frequency details of the positive half-cycle, including:
[0013] Upsample the low-resolution audio to increase the sampling rate to the target value to obtain the time-domain waveform corresponding to the low-resolution audio.
[0014] Perform time-domain half-wave rectification on the time-domain waveform corresponding to low-resolution audio, retain the positive half-cycle of the signal, and extract high-frequency details of the positive half-cycle.
[0015] In one implementation, based on the encoder, the output is the GRU feature corresponding to the mixed amplitude spectrum, including:
[0016] Input the mixed amplitude spectrum into Linear2ER and output the ERB band spectrum;
[0017] The output ERB band spectrum is input into the encoder to capture the inter-frame temporal dependencies of the speech spectrum and obtain GRU features. The last layer of the encoder is a grouped gated cyclic unit.
[0018] In one implementation, the Gaussian component parameters include the weight, mean, and standard deviation of each Gaussian component; the three constraint designs include: weighted constraint, mean constraint, and standard deviation constraint, wherein the weighted constraint is used to constrain the magnitude of each Gaussian component, the mean constraint is used to constrain the position of each Gaussian component in the frequency band, and the standard deviation constraint is used to constrain the radiation range of each Gaussian component.
[0019] In one implementation, the frame-level frequency distribution weights are represented as:
[0020]
[0021] in, The density function representing each Gaussian component, Representing the One audio frame, For frequency points in the frequency domain, Representing the Gaussian components The total number of Gaussian components. For the first The first audio frame The weights of each Gaussian component, For the first The first audio frame The mean of the Gaussian components, For the first The first audio frame The standard deviation of each Gaussian component.
[0022] In one implementation, the audio time-domain waveform reconstructed via inverse short-time Fourier transform is represented as follows:
[0023]
[0024] in, For high-resolution audio, For the final amplitude spectrum, For low-resolution audio, the phase... For imaginary number ranges, satisfying .
[0025] In one implementation, during the model training phase for reconstructing the audio time-domain waveform, the discriminator's loss function is:
[0026]
[0027] in, , , ;
[0028] The waveform loss is expressed as:
[0029]
[0030] in, This represents the total number of audio frames. For a specific frame of audio;
[0031] The multi-resolution short-time Fourier transform loss is expressed as:
[0032]
[0033] in, For a certain short-time Fourier transform parameter, For the first Spectral convergence loss of the short-time Fourier transform parameters No. The logarithmic magnitude spectrum L1 norm loss of the short-time Fourier transform parameters. =1;
[0034] To combat losses, there are generator-adversarial losses and discriminator-adversarial losses. To address adversarial losses, there are generator adversarial losses and discriminator adversarial losses. The generator adversarial loss is expressed as:
[0035]
[0036] The discriminator adversarial loss is expressed as:
[0037]
[0038] in, For realistic low-resolution input speech. For the generated speech, This represents the output of the discriminator for real low-resolution input speech. This represents the discriminator's response to the generated output. Represents L2 distance, represent L2 distance to all -1 vectors represent L2 distance to a vector consisting entirely of 1s represent L2 distance to all -1 vectors represent L2 distance to the all-1 vector;
[0039] The feature matching loss is expressed as:
[0040]
[0041] in, To determine the number of discriminator layers, Representative discriminator number Layer features.
[0042] Secondly, embodiments of the present invention also provide a speech spectrum reconstruction system combining time-domain half-wave rectification and a weighted Gaussian mixture model decoder. The system is used to implement the steps of the speech spectrum reconstruction method combining time-domain half-wave rectification and a weighted Gaussian mixture model decoder described above. The system includes:
[0043] The data preprocessing module is used to acquire low-resolution audio, perform time-domain half-wave rectification on the low-resolution audio, extract high-frequency details of the positive half-cycle, and obtain a mixed amplitude spectrum based on the amplitude spectrum of the low-resolution audio and the amplitude spectrum of the rectified signal.
[0044] The high-frequency reconstruction module is used to output GRU features corresponding to the mixed amplitude spectrum based on the encoder, input the GRU features to the weighted Gaussian mixture model decoder, and constrain each Gaussian component by applying three constraints through three parallel linear layers in the weighted Gaussian mixture model decoder to generate several sets of Gaussian component parameters.
[0045] The waveform reconstruction and output module is used to calculate the frame-level frequency distribution weights based on the Gaussian component parameters, combine them with the mixed amplitude spectrum to obtain the model output amplitude spectrum, input the model output amplitude spectrum and the amplitude spectrum of the low-resolution audio to the frequency band guiding masking module in the weighted Gaussian mixture model decoder to obtain the final amplitude spectrum, and reconstruct the audio time-domain waveform based on the final amplitude spectrum and the phase of the low-resolution audio through short-time Fourier inverse transform.
[0046] Thirdly, embodiments of the present invention also provide a terminal, wherein the terminal includes a memory, a processor, and a speech spectrum reconstruction program combining a time-domain half-wave rectification and a weighted Gaussian mixture model decoder stored in the memory and executable on the processor. When the processor executes the speech spectrum reconstruction program combining a time-domain half-wave rectification and a weighted Gaussian mixture model decoder, it implements the steps of the speech spectrum reconstruction method combining a time-domain half-wave rectification and a weighted Gaussian mixture model decoder in any of the above-mentioned schemes.
[0047] Fourthly, embodiments of the present invention also provide a computer-readable storage medium, wherein the computer-readable storage medium stores a speech spectrum reconstruction program combining a time-domain half-wave rectifier and a weighted Gaussian mixture model decoder, the speech spectrum reconstruction program combining a time-domain half-wave rectifier and a weighted Gaussian mixture model decoder implementing the steps of the speech spectrum reconstruction method combining a time-domain half-wave rectifier and a weighted Gaussian mixture model decoder as described in any of the above-described schemes on the computer-readable storage medium.
[0048] Beneficial Effects: Compared with existing technologies, this invention provides a speech spectrum reconstruction method combining time-domain half-wave rectification and a weighted Gaussian mixture model decoder. First, low-resolution audio is acquired. Time-domain half-wave rectification is performed on the low-resolution audio to extract high-frequency details in the positive half-cycle. Based on the amplitude spectrum of the low-resolution audio and the amplitude spectrum of the rectified signal, a mixed amplitude spectrum is obtained. Then, based on the encoder, GRU features corresponding to the mixed amplitude spectrum are output. These GRU features are input to the weighted Gaussian mixture model decoder. Each Gaussian component is constrained through three parallel linear layers and three constraint designs, generating several sets of Gaussian component parameters. Finally, frame-level frequency distribution weights are calculated based on the Gaussian component parameters and combined with the mixed amplitude spectrum to obtain the model output amplitude spectrum. The model output amplitude spectrum and the amplitude spectrum of the low-resolution audio are input to the frequency band guiding masking module in the weighted Gaussian mixture model decoder to obtain the final amplitude spectrum. Based on the phase of the final amplitude spectrum and the low-resolution audio, the audio time-domain waveform is reconstructed using an inverse short-time Fourier transform.
[0049] This invention innovatively proposes HWB-Net (High-Performance and Efficient Hybrid Waveform Bandwidth Extension Method-network), a network that achieves high-fidelity high-frequency reconstruction while significantly reducing model complexity. It is suitable for various scenarios such as speech bandwidth extension, speech enhancement, and speech restoration. Based on this, this invention combines half-wave rectification nonlinear operations with a lightweight neural network architecture to increase the input information density of the neural network. Furthermore, by replacing the decoder in the common encoder-decoder architecture with a weighted Gaussian mixture model decoder, it significantly reduces model parameters and computational complexity while improving the accuracy of high-frequency speech component reconstruction, achieving a balance between high performance and high efficiency. Attached Figure Description
[0050] Figure 1 This is a flowchart of a preferred embodiment of the speech spectrum reconstruction method combining time-domain half-wave rectification and weighted Gaussian mixture model decoder provided in this invention.
[0051] Figure 2 This is a flowchart illustrating the speech spectrum reconstruction method combining time-domain half-wave rectification and weighted Gaussian mixture model decoder provided in an embodiment of the present invention.
[0052] Figure 3 This is a schematic diagram illustrating the practical application of the speech spectrum reconstruction method combining time-domain half-wave rectification and weighted Gaussian mixture model decoder provided in this embodiment of the invention.
[0053] Figure 4 The flowchart of the weighted Gaussian mixture model decoder in the speech spectrum reconstruction method combining time-domain half-wave rectification and weighted Gaussian mixture model decoder provided in the embodiments of the present invention is as follows.
[0054] Figure 5 The schematic diagram of the speech spectrum reconstruction system combining time-domain half-wave rectification and weighted Gaussian mixture model decoder provided in the embodiment of the present invention.
[0055] Figure 6 A schematic diagram of a terminal provided in an embodiment of the present invention. Detailed Implementation
[0056] To make the objectives, technical solutions, and effects of this invention clearer and more explicit, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative of the invention and are not intended to limit the invention.
[0057] The flowchart shown in the attached diagram is for illustrative purposes only and does not necessarily include all content, operations, or steps, nor does it require execution in the described order. For example, some operations or steps can be broken down, combined, or partially merged, so the actual execution order may change depending on the actual situation.
[0058] It should be understood that the terminology used in this specification is for the purpose of describing particular embodiments only and is not intended to limit the invention. As used in this specification and the appended claims, the singular forms “a,” “an,” and “the” are intended to include the plural forms unless the context clearly indicates otherwise.
[0059] It should be understood that, in order to clearly describe the technical solutions of the embodiments of the present invention, the terms "first" and "second" are used in the embodiments of the present invention to distinguish identical or similar items with essentially the same function and effect. For example, "first control information" and "second control information" are only used to distinguish different control information and do not limit their order.
[0060] Those skilled in the art will understand that the words "first" and "second" do not limit the quantity or the order of execution, and that the words "first" and "second" do not necessarily imply that they are different.
[0061] It should also be understood that the term “and / or” as used in this specification and the appended claims refers to any combination of one or more of the associated listed items and all possible combinations, and includes such combinations.
[0062] The main innovative points of this invention include:
[0063] 1. Expand input information by using half-wave rectification (HWR) to enhance high-frequency information density. Specifically, half-wave rectification is performed on the low-resolution speech signal in the time domain to extract high-frequency details in the positive half-cycle. These details are then combined with the original signal as model input, increasing the high-frequency information content without increasing computational load.
[0064] 2. Lightweight Decoder: In this embodiment, a Weighted Gaussian Mixture Model (WGMM) decoder is used instead of the original decoder. By predicting Gaussian component parameters and adjusting high-frequency information using multiple Gaussian models, the sensing quality is maintained or improved while reducing the number of parameters.
[0065] 3. Collaborative optimization: The combination of HWR and WGMM significantly improves DNSMOS (Deep Noise Suppression Mean Opinion Score) and PESQ (Perceptual Evaluation of Speech Quality) metrics, making the perceived speech quality close to that of wideband speech, while supporting causal inference and low-latency deployment.
[0066] In practical applications, the speech spectrum reconstruction method combining time-domain half-wave rectification and weighted Gaussian mixture model decoder in this embodiment can be applied to terminals, such as mobile phones, IoT devices, and other intelligent products. Figure 1 As shown in the figure, the method in this embodiment specifically includes the following steps:
[0067] Step S100: Obtain low-resolution audio, perform time-domain half-wave rectification on the low-resolution audio, extract high-frequency details of the positive half-cycle, and obtain a mixed amplitude spectrum based on the amplitude spectrum of the low-resolution audio and the amplitude spectrum of the rectified signal.
[0068] In practical applications, combined with Figure 2 As shown in the diagram, this embodiment first acquires low-resolution audio, then upsamples the low-resolution audio, increasing the sampling rate to a target value (e.g., 16000), to obtain the time-domain waveform corresponding to the low-resolution audio. Next, a time-domain half-wave rectification operation is performed on the time-domain waveform corresponding to the low-resolution audio, preserving the positive half-cycle of the signal, extracting high-frequency details of the positive half-cycle, and actively supplementing high-frequency features to address the problem of insufficient high-frequency information in the input of existing models. The half-wave rectification operation in this embodiment is defined as follows:
[0069]
[0070] The time-domain waveform is obtained after half-wave rectification of low-resolution audio. This is the time-domain waveform after upsampling the low-resolution audio.
[0071] Furthermore, this embodiment describes the amplitude spectrum of the rectified signal. amplitude spectrum of low-resolution audio A Short-Time Fourier Transform (STFT) was performed using a Hanning window with a frame length of 512 sampling points and a frame shift of 256 sampling points, yielding 257-dimensional amplitude and phase spectrum features, thus obtaining the mixed amplitude spectrum. , combined Figure 3 As shown, the mixed amplitude spectrum Amplitude spectrum of low-resolution audio Amplitude spectrum of rectified signal The element-wise summation. This embodiment will mix the amplitude spectrum. As the actual input to the model, it can enrich the model's understanding of the spectral features of speech signals, especially enhance the detailed representation of high-frequency components, laying the foundation for high-fidelity reconstruction, and this operation has no additional computational cost.
[0072] Step S200: Based on the encoder, output the GRU features corresponding to the mixed amplitude spectrum, input the GRU features to the weighted Gaussian mixture model decoder, and constrain each Gaussian component by applying three constraints through three parallel linear layers in the weighted Gaussian mixture model decoder to generate several sets of Gaussian component parameters.
[0073] like Figure 3 As shown, this embodiment innovatively proposes HWB-Net (High-Performance and Efficient Hybrid Waveform Bandwidth Extension Method-network). HWB-Net can achieve high-fidelity high-frequency reconstruction while significantly reducing model complexity, and is suitable for various scenarios such as speech bandwidth extension, speech enhancement, and speech restoration. Figure 3 The gray area represents the structure of HWB-Net. In this embodiment, a Weighted Gaussian Mixture Model (WGMM) decoder replaces the decoder in the traditional encoder-decoder architecture. Therefore, combined with... Figure 3 As shown, HWB-Net includes Linear2ERB (linear frequency-equivalent rectangular bandwidth mapping module), an encoder, and a Weighted Gaussian Mixture Model (WGMM) decoder. The WGMM decoder comprises a WGMM module and a frequency band guiding masking module. Specifically, the functions of each module are described below:
[0074] Linear2ERB: Due to differences in human hearing perception of resolution at different frequencies—specifically, high perception of low-frequency resolution and low perception of high-frequency resolution—the input mixed amplitude spectrum... There is high-frequency redundancy, and it does not match auditory perception. To address this, the core function of Linear2ERB is to map the STFT (Short-Time Fourier Transform) spectrum of the linear frequency axis to the perceptual frequency axis based on ERB (Equivalent Rectangular Bandwidth), providing subsequent modules with feature inputs that are more consistent with human auditory perception.
[0075] encoder (specifically) Figure 3 (Light blue section): The 1D convolutional layer (Conv1D) of the encoder is the core feature encoding module of HWB-Net. Its function is to extract local features and compress dimensions of ERB features, providing compact and semantically rich spectral features for subsequent decoding modules, while supporting real-time streaming inference. The last layer of the encoder is the Grouped Gated Recurrent Unit (Grouped GRU), which is used to capture the inter-frame temporal dependencies of the speech spectrum (such as the continuity of the speech fundamental frequency and the dynamic changes of formants), while taking into account the requirements of model lightweighting and streaming inference, providing more temporally consistent feature support for subsequent amplitude completion and phase optimization.
[0076] Frequency band guided masking module: The main goal of this module is to solve the problem of excessive modification of low-frequency components in adaptive bandwidth scenarios. Due to fluctuations in the effective bandwidth of the input speech (e.g., dynamic changes between 8kHz and 48kHz), traditional amplitude completion models tend to make unnecessary adjustments to existing reliable low-frequency components (e.g., below 8kHz), leading to speech distortion. The frequency band guided masking module achieves accurate differentiation between "high-frequency components that need to be completed" and "low-frequency components that need to be retained" by dynamically generating frequency band-specific gain masks, thus ensuring high-quality amplitude spectrum output.
[0077] In practical applications, the input to Linear2ERB in this embodiment is the spectrum after the STFT step described above: a 257-dimensional linear frequency spectrum (i.e., a mixed amplitude spectrum). The output is a 128-dimensional ERB band spectrum, which is mapped using a triangular ERB filter bank. This ERB band spectrum is then input into the encoder to capture the inter-frame temporal dependencies of the speech spectrum, yielding GRU features. Specifically, the encoder contains four 1D convolutional layers and one group-gated recurrent unit (GRU). In the 1D convolutional layers, layers 1 and 2 have 128 input channels and 128 output channels, while layers 3 and 4 have 128 input channels and 64 output channels. All convolutional kernels are 3 in size, and the ReLU activation function is used. The GRU has an input dimension of 64 and a hidden layer dimension of 64. The final output GRU features are represented as follows: .
[0078] Next, this embodiment will use GRU features. The input is fed to the Weighted Gaussian Mixture Model decoder, also known as the WGMM decoder. The WGMM decoder comprises the WGMM module and a band-guided masking module. The WGMM module contains three parallel linear layers, each used to predict the mean, standard deviation, and weight of each Gaussian component, thus obtaining the Gaussian component parameters. These parameters are then used to generate the corresponding density function, which is used to adjust the high-frequency signal. Specifically, this is combined with... Figure 4 As shown, GRU features After inputting into three parallel linear layers, three constraint designs are applied to constrain each Gaussian component, generating several sets of Gaussian component parameters. , Representing the One audio frame, Representing the Gaussian components For the first The first audio frame The weights of each Gaussian component, For the first The first audio frame The mean of the Gaussian components, For the first The first audio frame The standard deviation of each Gaussian component. The three constraint designs include: weighted constraint, mean constraint, and standard deviation constraint, as detailed below:
[0079] Weighted Constraints: Weights This is used to constrain the magnitude of each Gaussian component. Unlike traditional Gaussian mixture models, the weighted Gaussian mixture model does not require the weights to sum to 1; it only requires that the weight of each component is greater than 1. It can flexibly scale the contribution of each Gaussian component to high-frequency reconstruction;
[0080] Mean constraint: mean This is used to constrain the position of each Gaussian component in the frequency band. For example... Figure 4 As shown, the preset base mean The target high-frequency band is evenly distributed at equal intervals, and the final mean value is... , This is the mean offset for model learning, ensuring that the mean focuses on the high-frequency target band;
[0081] Standard deviation constraint: Standard deviation Used to constrain the radiation range of each Gaussian component. For example... Figure 4 As shown, the preset baseline standard deviation ( To ensure numerical stability, the final standard deviation... , The standard deviation bias of the model learning is empirically shown to be fixed. It can improve training stability while ensuring performance.
[0082] Step S300: Calculate the frame-level frequency distribution weights based on the Gaussian component parameters, combine them with the mixed amplitude spectrum to obtain the model output amplitude spectrum, input the model output amplitude spectrum and the amplitude spectrum of the low-resolution audio to the frequency band guiding masking module in the weighted Gaussian mixture model decoder to obtain the final amplitude spectrum, and reconstruct the audio time-domain waveform based on the phase of the final amplitude spectrum and the low-resolution audio through short-time Fourier inverse transform.
[0083] After obtaining the Gaussian component parameters, this embodiment can calculate the frame-level frequency distribution weight based on the Gaussian component parameters.
[0084] The frame-level frequency distribution weights are represented as follows:
[0085]
[0086] in, The density function representing each Gaussian component, Representing the One audio frame, For frequency points in the frequency domain, Representing the Gaussian components The total number of Gaussian components. For the first The first audio frame The weights of each Gaussian component, For the first The first audio frame The mean of the Gaussian components, For the first The first audio frame The standard deviation of the Gaussian components. Specifically, for the first Gaussian component... Each audio frame, this embodiment through... We perform a weighted summation of Gaussian components to model the frequency distribution characteristics of the speech frame, quantify the amplitude distribution weight of the speech frame at each frequency point, and thus provide an accurate frequency domain modeling basis for high-frequency reconstruction of speech.
[0087] Furthermore, combined Figure 4 As shown, the distribution density of each Gaussian component can be obtained based on the frame-level frequency distribution weights. After extending to the entire time dimension, it is obtained by combining the Hadamard product with the mixed amplitude spectrum. By combining these, the amplitude spectrum of the model output is obtained. The Hadamard product is a matrix multiplication operation that multiplies two matrices of equal size at the same positions. Next, the model outputs the amplitude spectrum. amplitude spectrum of the low-resolution audio The frequency band guiding masking module in the weighted Gaussian mixture model decoder is input to obtain the final amplitude spectrum. The frequency band guidance masking module in this embodiment includes three one-dimensional convolutions and is used to receive two feature inputs (i.e., the above-mentioned...). and The algorithm generates two different masks through a dual-path approach, then fuses them using one-dimensional convolution and a sigmoid function to produce the final mask. Finally, this mask is combined with the input to generate the predicted amplitude spectrum. .
[0088] Furthermore, this embodiment is based on the final amplitude spectrum. The phase of the low-resolution audio is used to reconstruct the audio time-domain waveform through short-time inverse Fourier transform, thus obtaining high-resolution audio.
[0089] Combination Figure 3 As shown, during the model training phase for reconstructing the audio time-domain waveform, the discriminator employs a multi-scale STFT discriminator. During the training phase, the phase of the real high-resolution input speech signal is used. The final amplitude spectrum obtained from the prediction High-resolution audio is obtained through inverse STFT transformation. The formula is , This represents the inverse STFT transform. During the inference phase, because the phase of the true high-resolution input speech signal cannot be obtained, the phase of the low-resolution audio signal is used. The final amplitude spectrum obtained from the prediction Perform inverse STFT, the formula is as follows , For imaginary number ranges, satisfying During the inference phase, the phase of the low-resolution audio is obtained by phase flipping. Assuming the signal frequency of the low-resolution audio is 0-2kHz, it can be obtained from 2-4kHz by phase flipping. The relevant formula is Phase[2000:4000]=Phase[0:2000]+π, where Phase represents the phase. Repeating the above operation, the phase of the final reconstruction can be obtained.
[0090] In designing the loss function of the discriminator, this embodiment employs a four-loss weighted summation to ensure accurate recovery of the amplitude spectrum and phase. Specifically, this embodiment sets the waveform loss... Multi-resolution STFT loss Combating losses Feature matching loss The details are as follows:
[0091] Waveform loss Computational reconstruction of audio temporal speech With real high-resolution input speech signal The L2 distance is calculated using the following formula: , This represents the total number of audio frames. For a given audio frame, it is used to match the overall shape and phase of the waveform;
[0092] Multi-resolution STFT loss Combining spectral convergence loss (SC) and logarithmic amplitude spectrum L1 norm loss, three sets of STFT parameters are used (FFT bins (frequency components in the frequency domain data obtained after Fast Fourier Transform) = {512, 1024, 2048}, hop length (distance between adjacent windows) = {50, 120, 240}, window size (window size) = {240, 600, 1200}), the formula is as follows: ,in, For a certain short-time Fourier transform parameter, For the first Spectral convergence loss of the short-time Fourier transform parameters No. The logarithmic magnitude spectrum L1 norm loss of the short-time Fourier transform parameters. =1.
[0093] Combat losses The relativistically averaged least squares GAN (Generative Adversarial Network) is employed. The adversarial loss includes generator adversarial loss and discriminator adversarial loss. The discriminator D is used to distinguish between the real high-resolution input speech and the reconstructed speech. The generator adversarial loss is expressed as:
[0094]
[0095] The discriminator adversarial loss is expressed as:
[0096]
[0097] in, For realistic low-resolution input speech. For the generated speech, This represents the output of the discriminator for real low-resolution input speech. This represents the discriminator's response to the generated output. L2 distance (i.e., the straight-line distance between two points in Euclidean space, which is a commonly used metric for measuring the similarity between vectors (or sample points)). represent L2 distance to all -1 vectors represent L2 distance to a vector consisting entirely of 1s represent L2 distance to all -1 vectors represent L2 distance to the all-1 vector.
[0098] Feature matching loss Calculate the L1 distance (also known as Manhattan distance, which is a measure of the difference between vectors (or sample points) between the true features output by each layer of the discriminator and the generated features. The formula is:
[0099]
[0100] in, To determine the number of discriminator layers, Representative discriminator number Layer features.
[0101] Furthermore, this embodiment also provides system evaluation using more comprehensive objective metrics, including:
[0102] Spectral difference index: Log Spectral Distance (LSD), which quantifies the difference between the reconstructed spectrum and the true spectrum; the smaller the value, the better.
[0103] Perceptual quality: Deep Noise Suppression Mean Opinion Score (DNSMOS, conforming to P.808 standard), higher values are better, range: [0, 5]; Perceptual Evaluation of Speech Quality (PESQ), higher values are better, range: [-0.5, 4.5]; Non-Intrusive Speech Quality Assessment (NISQA), higher values are better, range: [0, 5];
[0104] Deployment efficiency: number of parameters (Para.) and number of multiply-accumulate operations per second (MACs) are used to measure edge adaptability.
[0105] The specific details are shown in Table 1 below. Table 1 compares the performance of HWB in this embodiment with other bandwidth extension methods, and the optimal performance index is highlighted in bold. In Table 1, the upward arrows indicate that the larger the value of the corresponding index, the better, and the downward arrows indicate that the smaller the value of the corresponding index, the better. In Table 1, Sinc, AERO, BAE, BAE-Lite, and BAE-Lite-Small are other different bandwidth extension methods. Sinc is a traditional interpolation benchmark method, AERO is a method based on a high-performance but bulky frequency domain model, BAE is a lightweight and spectrally accurate method, and BAE-Lite and BAE-Lite-Small both belong to the lightweight bandwidth extension model of the BAE (Blind Audio Bandwidth Extension) series. BAE-Lite is the first lightweight version of the BAE model, and BAE-Lite-Small is a second compressed version of BAE-Lite.
[0106] Table 1
[0107]
[0108] As shown in Table 1, HWB achieved the best performance in speech perception quality. To better compare the results, this embodiment compressed the previously best-performing BAE-Lite model to the same size as HWB (BAE-Lite-Small). Compared to BAE-Lite, DNSMOS improved by 0.13 (4.03%), PESQ by 0.22 (8.36%), and NISQA by 0.27 (7.52%). Although LSD is slightly higher than BAE-Lite, the perception quality is more in line with human auditory needs and meets the actual application scenarios of edge BWE.
[0109] Furthermore, this embodiment also provides ablation verification, as shown in Table 2. In Table 2, +HWR indicates combining half-wave rectification (HWR) with the BAE-Lite method, and +WGMM indicates combining a weighted Gaussian mixture model (WGMM) with the BAE-Lite method. The effect of HWR: Without increasing parameters or computational load, DNSMOS is improved by 3.06%, proving that HWR can effectively supplement high-frequency information and improve reconstruction accuracy. The effect of WGMM: Reducing parameters from 0.57M to 0.20M (a reduction of 64.91%) and MACs from 0.035G / s to 0.013G / s (a reduction of 62.86%), while DNSMOS only decreases by 0.59% and NISQA improves by 2.66%, proving that WGMM can maintain or even improve sensing quality while significantly reducing complexity.
[0110] Table 2
[0111]
[0112] Based on the ablation results in Table 2, HWR and WGMM are the core innovations for achieving a balance between "high performance and high efficiency", and their synergistic effect can maximize model performance.
[0113] The method in this embodiment has the characteristics of high fidelity, low latency, and low parameters, and can be widely used in:
[0114] Real-time communication scenarios: such as voice calls and video conferencing, improving voice clarity under narrowband networks;
[0115] Edge voice devices: such as smart speakers, in-vehicle voice assistants, and wearable devices, enable wideband voice output with limited hardware resources;
[0116] Voice restoration scenarios: such as high-frequency restoration of old recordings and low-quality voice files, restoring voice details;
[0117] Telecom operator networks: Integrate into base stations or terminal equipment to optimize the bandwidth expansion effect of 2G / 3G narrowband voice.
[0118] Based on the above embodiments, the present invention also provides a speech spectrum reconstruction system combining time-domain half-wave rectification and a weighted Gaussian mixture model decoder, wherein the system is used to implement the method steps in the method embodiments. Specifically, as follows... Figure 5 As shown, the system includes a data preprocessing module 10, a high-frequency reconstruction module 20, and a waveform reconstruction and output module 30. Specifically, the data preprocessing module 10 is used to acquire low-resolution audio, perform time-domain half-wave rectification on the low-resolution audio, extract high-frequency details of the positive half-cycle, and obtain a mixed amplitude spectrum based on the amplitude spectrum of the low-resolution audio and the amplitude spectrum of the rectified signal. The high-frequency reconstruction module 20 is used to output GRU features corresponding to the mixed amplitude spectrum based on the encoder, input the GRU features to the weighted Gaussian mixture model decoder, and constrain each Gaussian component by applying three constraints through three parallel linear layers in the weighted Gaussian mixture model decoder, generating several sets of Gaussian component parameters. The waveform reconstruction and output module 30 is used to calculate the frame-level frequency distribution weight based on the Gaussian component parameters, combine it with the mixed amplitude spectrum to obtain the model output amplitude spectrum, input the model output amplitude spectrum and the amplitude spectrum of the low-resolution audio to the frequency band guiding masking module in the weighted Gaussian mixture model decoder to obtain the final amplitude spectrum, and reconstruct the audio time-domain waveform based on the final amplitude spectrum and the phase of the low-resolution audio through short-time Fourier inverse transform.
[0119] The speech spectrum reconstruction system combining time-domain half-wave rectification and weighted Gaussian mixture model decoder in this embodiment is based on the same principle as the steps in the above method embodiments, and will not be repeated here.
[0120] Based on the above embodiments, the present invention also provides a terminal, the principle block diagram of which can be as follows: Figure 6 As shown. The terminal may include one or more processors 100 ( Figure 6 (Only one is shown in the diagram), memory 101, and computer program 102 stored in memory 101 and executable on one or more processors 100. For example, a speech spectrum reconstruction program combining a time-domain half-wave rectifier and a weighted Gaussian mixture model decoder. When one or more processors 100 execute computer program 102, they can implement the various steps in the speech spectrum reconstruction method embodiment combining a time-domain half-wave rectifier and a weighted Gaussian mixture model decoder. Alternatively, when one or more processors 100 execute computer program 102, they can implement the functions of various modules / units in the speech spectrum reconstruction system embodiment combining a time-domain half-wave rectifier and a weighted Gaussian mixture model decoder, which is not limited here.
[0121] In one embodiment, the processor 100 may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor.
[0122] In one embodiment, memory 101 may be an internal storage unit of an electronic device, such as a hard drive or RAM. Memory 101 may also be an external storage device of the electronic device, such as a plug-in hard drive, smart media card (SMC), secure digital card (SD), flash card, etc. Furthermore, memory 101 may include both internal and external storage units. Memory 101 is used to store computer programs and other programs and data required by the terminal. Memory 101 can also be used to temporarily store data that has been output or will be output.
[0123] Those skilled in the art will understand that Figure 6 The block diagram shown is merely a partial structural diagram related to the present invention and does not constitute a limitation on the terminal to which the present invention is applied. A specific terminal may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0124] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments of the above methods. Any references to memory, storage, operational databases, or other media used in the embodiments provided by this invention can include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in a variety of forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual operating data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), RAMbus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.
[0125] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A speech spectrum reconstruction method combining time-domain half-wave rectification and weighted Gaussian mixture model decoder, characterized in that, The method comprises: obtaining low-resolution audio, performing a time-domain half-wave rectification operation on the low-resolution audio, extracting high-frequency details in the positive half cycle, and obtaining a hybrid amplitude spectrum based on the amplitude spectrum of the low-resolution audio and the amplitude spectrum of the rectified signal; based on the encoder, outputting GRU features corresponding to the hybrid amplitude spectrum, inputting the GRU features into the weighted Gaussian mixture model decoder, and designing three constraints to constrain each Gaussian component through three parallel linear layers in the weighted Gaussian mixture model decoder, to generate a plurality of groups of Gaussian component parameters; based on the Gaussian component parameters, calculating frame-level frequency distribution weights, combining the hybrid amplitude spectrum to obtain a model output amplitude spectrum, inputting the model output amplitude spectrum and the amplitude spectrum of the low-resolution audio into the frequency band guided masking module in the weighted Gaussian mixture model decoder to obtain a final amplitude spectrum, and reconstructing the audio time-domain waveform based on the final amplitude spectrum and the phase of the low-resolution audio through inverse short-time Fourier transform.
2. The method for speech spectrum reconstruction combining time-domain half-wave rectification with weighted Gaussian mixture model decoder according to claim 1, characterized in that, obtaining low-resolution audio, performing a time-domain half-wave rectification operation on the low-resolution audio, extracting high-frequency details in the positive half cycle, comprising: upsampling the low-resolution audio to a target sampling rate to obtain a time-domain waveform corresponding to the low-resolution audio; performing a time-domain half-wave rectification operation on the time-domain waveform corresponding to the low-resolution audio to retain the positive half cycle and extract high-frequency details in the positive half cycle.
3. The method for speech spectrum reconstruction combined with time-domain half-wave rectification and weighted Gaussian mixture model decoder according to claim 1, characterized in that, based on the encoder, outputting GRU features corresponding to the hybrid amplitude spectrum, comprising: inputting the hybrid amplitude spectrum into Linear2ER to output an ERB band spectrum; inputting the output ERB band spectrum into the encoder to capture the inter-frame time sequence dependence of the speech spectrum to obtain GRU features, wherein the last layer of the encoder is a grouped gated recurrent unit.
4. The method for speech spectrum reconstruction incorporating time-domain half-wave rectification and weighted Gaussian mixture model decoder of claim 1, wherein, The Gaussian component parameters include the weight, mean and standard deviation of each Gaussian component; the three constraint designs include: weight constraint, mean constraint and standard deviation constraint, wherein the weight constraint is used to constrain the size of each Gaussian component, the mean constraint is used to constrain the position of each Gaussian component in the frequency band, and the standard deviation constraint is used to constrain the radiation range of each Gaussian component.
5. The method for speech spectrum reconstruction incorporating time-domain half-wave rectification and weighted Gaussian mixture model decoder of claim 1, wherein, The frame-level frequency distribution weight is represented as: wherein denotes a density function of each Gaussian component, denotes the th speech frame, is a frequency bin in the frequency domain, denotes the th Gaussian component, is the total number of Gaussian components, is a weight of the th Gaussian component in the th speech frame, is a mean of the th Gaussian component in the th speech frame, is a standard deviation of the th Gaussian component in the th speech frame.
6. The method for speech spectrum reconstruction incorporating time-domain half-wave rectification and weighted Gaussian mixture model decoder of claim 1, wherein, The reconstructed audio time-domain waveform is represented as: wherein is a high resolution audio, is a final amplitude spectrum, is a phase of a low resolution audio, is an imaginary bin, satisfying .
7. The method for speech spectrum reconstruction combining time-domain half-wave rectification with weighted Gaussian mixture model decoder according to claim 6, characterized in that, In the model training stage of the reconstructed audio time-domain waveform, the loss function of the discriminator is: wherein , , ; For the waveform loss, denoted as: wherein, is the total number of audio frames, is a certain frame of audio; The multi-resolution short-time Fourier transform loss is denoted as: wherein is a spectral convergence loss for a certain short-time Fourier transform parameter, is a log-magnitude spectrum LI norm loss for the th short-time Fourier transform parameter, is a log-magnitude spectrum LI norm loss for the th short-time Fourier transform parameter, = 1. To combat losses, including generator adversarial loss and discriminator adversarial loss, the generator adversarial loss is represented as: The discriminator adversarial loss is represented as: wherein, is the real low-resolution input speech, is the generated speech, represents the output of the discriminator for the real low-resolution input speech, represents the output of the discriminator for the generated output, represents the L2 distance, represents the L2 distance to the all-1 vector, represents the L2 distance to the all-1 vector, represents the L2 distance to the all-1 vector, represents the L2 distance to the all-1 vector. For the feature matching loss, denoted as Lmatch, is given by wherein, is the number of layers of the discriminator, represents the first layer of the discriminator, represents the layer feature.
8. A speech spectrum reconstruction system incorporating time-domain half-wave rectification with a weighted Gaussian mixture model decoder, characterized by, The system is used to implement the steps of the speech spectrum reconstruction method combining time-domain half-wave rectification and weighted Gaussian mixture model decoder according to any one of claims 1-7, and the system comprises: a data preprocessing module for obtaining low-resolution audio, performing a time-domain half-wave rectification operation on the low-resolution audio, extracting high-frequency details in the positive half cycle, and obtaining a hybrid amplitude spectrum based on the amplitude spectrum of the low-resolution audio and the amplitude spectrum of the rectified signal; a high-frequency reconstruction module configured to output GRU features corresponding to the mixed-amplitude spectrum based on the encoder, input the GRU features into a weighted Gaussian mixture model decoder, constrain each Gaussian component by three parallel linear layers in the weighted Gaussian mixture model decoder and applying three constraint designs, and generate a plurality of sets of Gaussian component parameters; a waveform reconstruction and output module configured to calculate frame-level frequency distribution weights based on the Gaussian component parameters, combine the mixed-amplitude spectrum to obtain a model output amplitude spectrum, input the model output amplitude spectrum and the amplitude spectrum of the low-resolution audio into a band guided masking module in the weighted Gaussian mixture model decoder to obtain a final amplitude spectrum, and reconstruct an audio time-domain waveform by inverse short-time Fourier transform based on the final amplitude spectrum and the phase of the low-resolution audio.
9. A terminal, characterized by comprising: The terminal comprises a memory, a processor, and a speech spectrum reconstruction program combined with the time-domain half-wave rectification and the weighted Gaussian mixture model decoder stored in the memory and executable on the processor. When the processor executes the speech spectrum reconstruction program combined with the time-domain half-wave rectification and the weighted Gaussian mixture model decoder, the steps of the speech spectrum reconstruction method combined with the time-domain half-wave rectification and the weighted Gaussian mixture model decoder according to any one of claims 1-7 are implemented.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a speech spectrum reconstruction program combined with the time-domain half-wave rectification and the weighted Gaussian mixture model decoder. The speech spectrum reconstruction program combined with the time-domain half-wave rectification and the weighted Gaussian mixture model decoder implements the steps of the speech spectrum reconstruction method combined with the time-domain half-wave rectification and the weighted Gaussian mixture model decoder according to any one of claims 1-7 on the computer-readable storage medium.
Citation Information
Patent Citations
Band-width spreading method and system for voice or audio signal
CN101140759A
Audio compression and reconstruction method, device, equipment and medium
CN120526782A