Speech reconstruction method and system based on entropy coding residual quantization and spectrum repair

By introducing entropy-coded residual quantization and spectrum restoration techniques into the neural vocoder, the problems of insufficient multi-scale feature capture and robust code stream organization in speech signal processing at low bit rates are solved, achieving efficient speech reconstruction and stable communication results.

CN121096348AActive Publication Date: 2025-12-09QILU UNIVERSITY OF TECHNOLOGY (SHANDONG ACADEMY OF SCIENCES) +1

Patent Information

Application Number
CN202511631176.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-10
Publication Date
2025-12-09
Estimated Expiration
2045-11-10

AI Technical Summary

Technical Problem

Existing neural vocoders suffer from problems such as insufficient multi-scale feature capture and fusion, information loss, uneven codebook utilization, insufficient spectral phase information modeling, and poor robustness of code stream organization under low bit rate conditions, resulting in decreased synthesized speech quality and poor communication stability.

Method used

An entropy-based residual quantization and spectrum restoration method is adopted. By introducing a dynamically gated residual vector quantization mechanism at the encoding end, bits are adaptively allocated. At the decoding end, spectrum restoration technology is used to predict residuals in the logarithmic amplitude spectral domain and perform confidence-gated fusion. An anchor point mechanism and a layer mask context entropy coding scheme are set at the bitstream organization level to improve the robustness of the system.

Benefits of technology

It improves the compression efficiency and quality of speech reconstruction at extremely low bit rates, enhances speech clarity and intelligibility, strengthens the robustness and stability of the system under complex network conditions, reduces quantization distortion and stuttering, and improves the reliability and subjective naturalness of communication.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121096348A_ABST
    Figure CN121096348A_ABST
Patent Text Reader

Abstract

The invention provides a voice reconstruction method and system based on entropy coding residual quantization and frequency spectrum restoration, and relates to the technical field of artificial intelligence voice signal processing, and the method comprises the steps: obtaining an original voice waveform, inputting the voice waveform into a neural voice coding and decoding model, firstly entering a coder to map the input voice waveform into acoustic potential representation, and then entering a frequency spectrum restoration model; performing residual quantization on the acoustic potential characterization layer by layer through a residual vector quantization module, introducing a gating-based dynamic layer number selection mechanism and entropy regularization constraint, enabling bits to be adaptively distributed among different voice segments, reconstructing reconstructed acoustic features of the potential characterization, inputting the reconstructed acoustic features into a decoder, restoring the reconstructed acoustic features into a time domain waveform, and outputting the time domain waveform. And mapping to a logarithmic magnitude spectrum domain through a spectrum repairing module, predicting a residual error in the logarithmic magnitude spectrum domain and performing confidence gating fusion to obtain a complex spectrum, and outputting after time domain synthesis to obtain reconstructed speech. According to the invention, high fidelity, intelligibility and transmission reliability of the voice can be considered at an extremely low bit rate.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present disclosure relates to the technical field of artificial intelligence speech signal processing, in particular to a speech reconstruction method and system based on entropy coding residual quantization and spectrum repair. BACKGROUND

[0002] The statements in this section merely provide background information related to the present disclosure and do not necessarily constitute prior art.

[0003] Low-bitrate speech coding technology has important application value in modern communication systems, especially in special environments such as satellite communication, short-wave communication, underwater acoustic communication, and secure communication. Due to the problems of limited bandwidth, poor channel quality, and complex environmental noise in these scenarios, how to achieve high-quality speech transmission at extremely low code rate has become a core research problem in the field of speech signal processing. Taking satellite communication as an example, speech signals need to cross the atmosphere and space environment, and the transmission bandwidth is extremely limited; in military or emergency communication, devices often need to work under low power consumption and high interference conditions, and speech signals not only need to be clear and understandable, but also need to occupy as few resources as possible. Therefore, low-bitrate speech coding not only improves communication efficiency, but also plays a key role in secure transmission and energy-saving applications.

[0004] With the development of deep learning technology, neural vocoder has gradually become a research hotspot in the field of speech coding and decoding. A typical neural vocoder uses a convolutional neural network at the encoding end to downsample the input speech, designs a multi-level vector quantizer at the quantization end to map continuous features to discrete code word indices, uses a decoder at the decoding end to upsample and reconstruct, generates time-domain speech, and introduces a discriminator at the training stage to enhance the realism of the generator through adversarial training, thereby improving the naturalness and clarity of the synthesized speech.

[0005] However, the existing neural vocoder still has the following problems: (1) Speech signals themselves have multi-scale features, and different time scales and frequency ranges contain rich speech details, but the existing model is insufficient in capturing and fusing multi-scale features, which can easily cause information loss or redundancy, thereby affecting the quality of the synthesized speech; (2) In the up-sampling and down-sampling process, high-frequency details often cannot be effectively restored, resulting in blurred and distorted decoded speech, especially under low-bitrate conditions, where too much detail is compressed, making the naturalness of the speech significantly decrease; (3) In the quantization stage, the problem of uneven codebook utilization is prominent, with some code words being idle for a long time and others being used excessively frequently, causing a decline in expression ability and quantization distortion; (4) There is a lack of modeling of phase information in the frequency spectrum, which often leads to perceptual defects such as metallic sound or transient distortion in the synthesized speech; (5) From the perspective of bit stream organization, most of the current methods lack efficient side information compression and robustness design. Once there is packet loss or asynchronization, the decoding end is often difficult to recover quickly, affecting the stability and reliability of actual communication. SUMMARY

[0006] In order to solve the above problems, the present disclosure proposes a speech reconstruction method and system based on entropy coding residual quantization and spectral repair. By introducing a residual vector quantization mechanism based on dynamic gating at the encoding end, adaptive bit allocation is realized. At the decoding end, the spectral repair technology is used to predict the residual in the log amplitude spectrum domain and perform confidence gating fusion, combined with phase gradient repair to make up for the loss of high frequencies. Anchor mechanism and layer mask context entropy coding scheme are set at the bit stream organization level to realize random access and packet loss fallback, and improve the robustness of the system under complex network conditions.

[0007] According to some embodiments, the present disclosure adopts the following technical solutions: The speech reconstruction method based on entropy coding residual quantization and spectral repair comprises: Obtaining an original speech waveform and performing preprocessing; Inputting the preprocessed speech waveform into a neural speech coding and decoding model to output a reconstructed speech; The processing process of the preprocessed speech waveform in the neural speech coding and decoding model comprises: The input speech waveform first enters an encoder. The encoder adopts a multi-layer convolutional residual structure and hierarchical downsampling to map the input speech waveform to an acoustic latent representation. The acoustic latent representation is input into a residual vector quantization module to perform residual quantization on the acoustic latent representation layer by layer. A dynamic layer number selection mechanism based on gating and an entropy regularization constraint are introduced to enable adaptive bit allocation between different speech segments. The quantized features are obtained, and the reconstructed acoustic features of the latent representation are obtained by reconstruction. The reconstructed acoustic features are input into a decoder, which is restored to a time domain waveform through hierarchical upsampling and residual blocks. The time domain waveform is mapped to the log amplitude spectrum domain through a spectral repair module, the residual is predicted in the log amplitude spectrum domain, and the confidence gating fusion is performed to obtain a complex spectrum. After time domain synthesis, the reconstructed speech is output.

[0008] According to some embodiments, the present disclosure adopts the following technical solutions: The speech reconstruction system based on entropy coding residual quantization and spectral repair comprises: A data acquisition module for obtaining an original speech waveform and performing preprocessing; A speech reconstruction module for inputting the preprocessed speech waveform into a neural speech coding and decoding model to output a reconstructed speech; The processing process of the preprocessed speech waveform in the neural speech coding and decoding model comprises: The input speech waveform first enters an encoder, which uses a multi-layer convolutional residual structure and hierarchical down-sampling to map the input speech waveform to an acoustic latent representation, inputs the acoustic latent representation to a residual vector quantization module, performs residual quantization on the acoustic latent representation layer by layer, introduces a dynamic number of layers selection mechanism based on gating and an entropy regularization constraint to adaptively allocate bits between different speech segments, obtains quantized features, reconstructs the acoustic features of the latent representation to obtain reconstructed acoustic features, inputs the reconstructed acoustic features to a decoder, restores the time-domain waveform through hierarchical up-sampling and residual blocks, maps the time-domain waveform to the log amplitude spectrum domain through a spectral repair module, predicts the residual in the log amplitude spectrum domain and performs confidence gated fusion to obtain a complex spectrum, and then outputs the reconstructed speech after time-domain synthesis.

[0009] According to some embodiments, the present disclosure adopts the technical scheme as follows: A non-transitory computer-readable storage medium for storing computer instructions, which, when executed by a processor, implements the speech reconstruction method based on entropy coding residual quantization and spectral repair.

[0010] According to some embodiments, the present disclosure adopts the technical scheme as follows: An electronic device comprising a processor, a memory, and a computer program; wherein the processor is connected with the memory, and the computer program is stored in the memory; when the electronic device is running, the processor executes the computer program stored in the memory to enable the electronic device to implement the speech reconstruction method based on entropy coding residual quantization and spectral repair.

[0011] Compared with the prior art, the present disclosure has the beneficial effects that: The speech reconstruction method based on entropy coding residual quantization and spectral repair of the present disclosure introduces a residual vector quantization mechanism based on dynamic gating at the encoding end to achieve adaptive bit allocation, thereby improving compression efficiency and codebook utilization under limited code rate; at the decoding end, the spectral repair mechanism is used to predict the residual in the log amplitude spectrum domain and perform confidence gated fusion, combined with phase gradient repair to make up for high frequency loss; at the same time, an anchor mechanism and a layer mask context entropy coding scheme are set at the code stream organization level to realize random access and packet loss fallback, and improve the robustness of the system under complex network conditions.

[0012] The speech reconstruction method based on entropy coding residual quantization and spectral repair of the present disclosure introduces a gating-based dynamic layer number selection and entropy regularization constraint in the residual vector quantization link, so that the bits can be adaptively allocated according to the complexity of the speech content: for difficult segments such as explosive sound, fricative and fast transient, the system automatically increases the effective layer depth to ensure the details; for the silent and smooth segments, the layer number is reduced to save bits. This mechanism significantly alleviates the "over-provisioning / under-provisioning" contradiction caused by fixed layer number, improves the balance and effective capacity of code word usage, reduces the accumulation and diffusion of quantization distortion on the time axis, and makes the reconstructed speech more stable in clarity and intelligibility. It can maintain high speech quality under very low code rate conditions, and is suitable for low-bandwidth speech communication, satellite and short-wave transmission, Internet of Things edge nodes, and speech archive compression storage scenarios.

[0013] The speech reconstruction method based on entropy coding residual quantization and spectral repair of the present disclosure, the spectral repair mechanism at the decoding end takes the residual refinement in the log magnitude spectrum domain as the core, and is supplemented by confidence gating and an optional phase correction path, which can compensate for the loss of high-frequency details and the accumulation of artifacts caused by quantization. Compared with schemes that only rely on time domain or single amplitude spectrum constraints, this module presents a more natural energy distribution at the harmonic structure, fricative boundary and transient zero-crossing, significantly reducing perceptual defects such as "metallic sound" and "grit feeling", while performing cautious gain control on distorted areas to avoid over-repair-induced pumping and ringing, thereby achieving higher subjective naturalness and spatiality.

[0014] The speech reconstruction method based on entropy coding residual quantization and spectral repair of the present disclosure, at the code stream organization level, the context probability of the layer mask is modeled and compressed using arithmetic coding, which significantly reduces the side information overhead; meanwhile, anchor packets are inserted at fixed intervals to record the quantization state and phase reference, so that the decoding end has the ability of random starting and fast packet loss rollback. This design significantly improves the continuous playback and interactive experience under complex network conditions without significantly increasing the overall code rate, reduces the duration of stuttering, sound interruption and significant distortion, and improves the stability and availability of end-to-end services.

[0015] The speech reconstruction method based on entropy coding residual quantization and spectral repair of the present disclosure, the multi-domain multi-scale discriminator and mutual exclusion regularization in the training stage enable the system to constrain the generated quality from three dimensions of periodic structure, time scale and spectral texture, avoiding the loss of training efficiency caused by overlapping of discriminator focus. Combined with multi-resolution STFT and feature matching loss, the model can still maintain high harmonic consistency, detail resolution and coherent speech flow under low code rate, and both subjective and objective indicators are systematically improved.

[0016] The speech reconstruction method based on entropy coding residual quantization and spectral repair of the present disclosure has advantages on the engineering and deployment side. Dynamic layer selection and context compression reduce the average code rate and computational burden, the anchor system facilitates synchronization maintenance and version evolution in long flow, the spectral repair mechanism adopts a lightweight U-Net structure, which is convenient for joint use with quantization, distillation, pruning and other inference optimization, so as to realize low delay and low power consumption on mobile terminals and edge devices. In summary, the present disclosure can balance speech quality, robustness and transmission efficiency in the scene of 0.6-0.8kbps and other ultra-low code rates, and has significant application and industrialization value. BRIEF DESCRIPTION OF DRAWINGS

[0017] The accompanying drawings, which form a part of the present disclosure, are used to provide further understanding of the present disclosure, and the schematic embodiments of the present disclosure and the description thereof are used to explain the present disclosure, and do not constitute improper limitations on the present disclosure.

[0018] Figure 1 A structural schematic diagram of a neural speech coding and decoding model of an embodiment of the present disclosure; Figure 2 A flowchart of entropy coding dynamic residual quantization and spectral repair of an embodiment of the present disclosure; Figure 3 A flowchart of a spectral repair mechanism of an embodiment of the present disclosure; Figure 4 A schematic diagram of an entropy coding dynamic residual quantization module of an embodiment of the present disclosure. DETAILED DESCRIPTION

[0019] The present disclosure will be further described below in combination with the drawings and embodiments.

[0020] It should be noted that the following detailed description is exemplary and is intended to provide further explanation of the present disclosure. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which the present disclosure belongs.

[0021] It should be noted that the terms used herein are only for the purpose of describing specific embodiments, and are not intended to limit the exemplary embodiments according to the present disclosure. As used herein, the singular form is intended to include the plural form unless the context clearly indicates otherwise, and furthermore, it should be understood that when the terms "comprise" and / or "include" are used in the specification, there is a reference to the presence of a feature, step, operation, device, component and / or combinations thereof.

[0022] Embodiment 1 In an embodiment of the present disclosure, a speech reconstruction method based on entropy coding residual quantization and spectral repair is provided, and the method steps include: Step 1: Obtain the original speech waveform and perform preprocessing; Step two: input the preprocessed speech waveform into the neural speech coding model to output reconstructed speech; The processing procedure of the preprocessed speech waveform in the neural speech coding model comprises: The input speech waveform first enters the encoder, which adopts a multi-layer convolution residual structure and hierarchical down-sampling to map the input speech waveform into acoustic latent representation. The acoustic latent representation is input into the residual vector quantization module, which performs residual quantization on the acoustic latent representation layer by layer. A dynamic layer number selection mechanism based on gating and an entropy regularization constraint are introduced to adaptively allocate bits among different speech segments. The quantized features are reconstructed to obtain reconstructed acoustic features of the latent representation. The reconstructed acoustic features are input into the decoder, which is restored to a time-domain waveform through hierarchical up-sampling and residual blocks. The time-domain waveform is mapped to the log amplitude spectrum domain through the spectral repair module. The residual is predicted in the log amplitude spectrum domain and fused through confidence gating to obtain a complex spectrum. After time-domain synthesis, the reconstructed speech is output.

[0023] As an embodiment, the present disclosure is based on a speech reconstruction method based on entropy coding residual quantization and spectral repair. The method introduces a residual vector quantization mechanism based on dynamic gating at the encoding end to achieve adaptive bit allocation, thereby improving compression efficiency and codebook utilization under limited code rate. At the decoding end, the spectral repair mechanism is used to predict the residual in the log amplitude spectrum domain and fuse it through confidence gating, combined with phase gradient repair to compensate for high-frequency loss. Meanwhile, an anchor mechanism and a layer mask context entropy coding scheme are set at the code stream organization level to realize random access and packet loss fallback, improving the robustness of the system under complex network conditions.

[0024] The present disclosure designs a neural speech coding model, which comprises an encoder, a residual vector quantization module, a code stream organization module, a decoder, and a spectral repair module. The encoder and the decoder adopt a mirror SEANet structure, with a channel base width of , and a hierarchical down-sampling ratio of [10, 8, 4, 1]. The residual vector quantization module mainly includes a residual vector quantizer (RVQ), which contains layer codebooks, the first layer codebook size is , and the quantization dimension is aligned with the latent dimension. The multi-scale STFT constraint uses several groups of , representing the number of points of Fourier transform (i.e., window length), Representing the frame shift length, the model can learn the spectral characteristics of speech at different scales by using multiple sets of parameters, typically (512, 80), (1024, 160), and (256, 40). The Spectrum Restoration Module (SRM) adopts a lightweight U-Net architecture with 4 layers each for encoding and decoding, and a channel width of 16-128. Its output includes the logarithmic magnitude spectrum residual. With confidence plot [0,1], simultaneously outputs phase gradient correction amount in extended configuration. The discriminator employs a multi-period (MPD) and multi-scale (MSD) structure, combined with a multi-scale STFT discriminator (MS-STFTD) to form a multi-domain constraint. A layer mask context model and an arithmetic encoder are set on the bitstream side, with fixed intervals... Frame insertion anchor packet (GOP anchor). 50-200 frames can be used to balance latency and redundancy. The specific implementation process of the speech reconstruction method based on entropy coding residual quantization and spectrum repair disclosed in this paper is as follows: Step 1: Obtain the raw speech waveform and perform preprocessing; The original speech sampling rate was set to 8kHz, and the frame shift and segment length were set for mini-batch training and streaming inference, with reading and writing typically performed in 20ms granularity.

[0025] Furthermore, the preprocessing steps include: normalizing or peak limiting the input raw speech waveform, removing the DC component, and performing light pre-noise reduction as needed.

[0026] During the training phase, long speech segments are divided into fixed-length segments (e.g., 1.0s-1.5s) and randomly pruned and perturbed with random gain. To improve robustness, room reverberation, band-limited filtering, and background noise can be added with a certain probability. During the inference phase, audio blocks are collected in streaming mode using a circular buffer. Each time a basic frame shift is collected, an encoding-quantization-packing process is triggered.

[0027] Step 2: Input the preprocessed speech waveform into the neural speech encoding / decoding model, and output the reconstructed speech. The neural speech encoding / decoding model includes an encoder, a residual vector quantization module, a code stream organization module, a decoder, and a spectrum restoration module. The specific processing steps include: Step 21: First, the encoder employs a multi-layer convolutional residual structure and hierarchical downsampling to map the input speech waveform into an acoustic latent representation, including: The speech waveform input into the neural speech coding model first enters an encoder, which adopts an SEANet structure and is composed of a plurality of residual convolution units and hierarchical down-sampling modules. Specifically, the residual convolution units are residual blocks containing point-wise convolution, dilated convolution and normalization / activation sequence, and skip connections are reserved between layers for gradient and information flow. In addition, modules such as long short-term memory network (LSTM) can be integrated into the model to enhance the modeling capability of time series. After sampling by the hierarchical down-sampling module, the speech signal obtains a latent tensor with a set time sequence length of To facilitate the operation of the residual vector quantizer RVQ, the tensor dimensions are uniformly arranged to obtain the acoustic latent representation , where B is the batch size, D is the number of channels of the latent representation, T' is the length of the time frame after down-sampling.

[0028] Step 22: input the acoustic latent representation into the residual vector quantization module, perform residual quantization on the acoustic latent representation layer by layer, introduce a dynamic layer number selection mechanism based on gating and an entropy regularization constraint, and adaptively allocate bits between different speech segments, including: The residual vector quantization module is a residual vector quantizer (RVQ) containing a layer codebook, the size of the layer codebook is , and the quantization dimension is aligned with the latent dimension.

[0029] The residual vector quantizer (RVQ) performs residual quantization on the acoustic latent representation layer by layer, the layer quantizes the current residual with the codebook to obtain the index and the reconstructed vector , and updates the residual .

[0030] The present disclosure is to adaptively select the number of layers under a given code rate and introduce a dynamic gating mechanism based on Gumbel-Softmax (dynamic layer number selection mechanism based on gating): a lightweight scoring network is used to generate an importance score for each layer, and the Gumbel-Softmax normalization obtains the gating probability . The time dimension is averaged to obtain the segment-level score , and the target code rate ​The mapping relationship between them can be calculated using layer depth. The implementation method is based on the average bit budget per layer. The bit rate consumption of the approximation layer ( (for frames per second), searching for the satisfying The smallest ,in Amortize bits for side information (layer mask, anchor points, etc.).

[0031] During training, a straight-through estimator (STE) is used to calculate the gradient of discrete choices; during inference, scores are calculated at the fragment level. Select from high to low Layer as an activation layer set Unselected layers do not generate codewords and only participate in regularization and distillation loss during training. To balance codeword usage, an empirical distribution is applied to each codebook. Adding normalized entropy regularization It is also assigned a small weight to avoid disturbing the main objective.

[0032] Step 23: The bitstream organization module encapsulates the quantization index and layer mask, uses the context model for probabilistic modeling, and employs arithmetic coding compression to reduce information overhead; Specifically, the stream organization module will activate the layer set. Encoded as a length of binary layer mask ,in Indicates the first The layer is activated. The layer mask exhibits Markov property over time, using a first-order context model. The conditional distribution is estimated and written into the bitstream using arithmetic coding, thereby constraining the side information bits to a minimal overhead range. Subsequently, the codeword index sequences of each layer are written in layer order. Quantitative features are obtained; To improve error recovery capabilities, every The frame is written to the anchor packet, which contains an RVQ residual state summary, necessary normalized statistics, and phase reference hash. The bitstream header records metadata such as sampling rate, encoder version, STFT group, and discriminator signature, facilitating compatibility between different model versions at the decoding end.

[0033] Step 24: The decoder reads the quantization features and first parses the header information and anchor points; if it is a random start-up or packet loss recovery scenario, it recovers the RVQ state starting from the most recent anchor point, and then decodes the layer mask in chronological order. By combining the codeword indexes of each layer, the quantized features are reconstructed and then stacked layer by layer to obtain the reconstructed acoustic features of the latent representation. and organize it into The layout prioritizes channels that align with the decoder input.

[0034] Step 25: Reconstructing acoustic features The decoder of the SEANet structure is restored to the time-domain waveform by hierarchical upsampling and residual blocks To reduce the boundary effect, the overlapping-addition (OLA) strategy is used to splice the continuous frames into a complete stream; in streaming mode, the sliding window alignment ensures phase continuity, and a very light limiter is used to suppress occasional spikes.

[0035] Step 26: Map the time-domain waveform to the log-magnitude spectrum domain through the spectral repair module, predict the residual and perform confidence gating fusion in the log-magnitude spectrum domain to obtain the complex spectrum; Specifically, the time-domain waveform is projected into the frequency domain by multiple sets of short-time Fourier transform (STFT) to obtain the complex spectrum . The amplitude part is taken as the logarithm and a numerical stabilizing term is added to obtain As the input of the spectral repair module SRM, the input is processed through a light U-Net structure, which contains a down-sampling encoding path composed of multiple convolution, normalization and ReLU activation function units, and an up-sampling decoding path containing skip connection, so as to effectively fuse multi-scale spectral features. After processing through the U-Net structure, the residual and the confidence map are obtained.

[0036] The repaired log-magnitude is obtained in a gated manner , and is exponentiated to . In the configuration with the phase branch enabled, the phase gradient correction amount is predicted simultaneously to form the phase . The final complex spectrum is , which is synthesized in the time domain through inverse short-time Fourier transform (ISTFT) with the same window type and overlap rate as the encoding end, and the reconstructed speech waveform is output. The output of the phase branch is usually limited and scaled to prevent the introduction of unstable phase jumps; the confidence map is used to suppress the over-repair of SRM in uncertain areas and reduce "pumping" and "metallic taste".

[0037] As an embodiment, the speech reconstruction method based on entropy coding residual quantization and spectral repair of the present disclosure concatenates the entire process as described in steps 21-26 into an end-to-end graph in the training stage. Waveform L1, multi-resolution STFT, feature matching and adversarial loss are used jointly, and multi-period discriminator, multi-scale discriminator and multi-scale STFT discriminator are used for multi-domain optimization. Specifically as follows: The reconstruction loss adopts waveform domain Joint with multi-resolution STFT distance, the spectral distance selects the log-magnitude and phase consistency measure, the original speech is x , the reconstructed speech is G ( z ), the reconstruction loss is: L_recon = || x - G ( z )||_1 + α * L_stft( x , G ( z )) The adversarial loss adopts the Logistic form, the discriminator includes three types of MPD, MSD and MS-STFTD, which respectively constrain the periodic structure, time scale and spectral structure, and the Logistic form of the adversarial loss is as follows, where G is the generator, D is the discriminator: L_adv( G , D ) = E_x[log D( x )] + E_z[log(1 - D( G ( z )))] The feature matching loss is introduced to stabilize the training and improve the perceptual quality. The loss is realized by minimizing the L1 distance of the feature maps in the intermediate layers of the discriminator network, and the formula is as follows: L_fm( G , D ) = E _{ x , z}[Σ_{ i =1 to L} (1 / N _ i ) * || D _ i ( x ) - D _ i ( G ( z ))||_1] The mutual exclusion regular or mutual information penalty is added between the discriminators to encourage them to learn complementary features rather than simple repetition, which is realized by minimizing the cosine similarity between the features extracted by different discriminators (such as MPD and MSD): L_mutex = |cos(f_mpd( x ), f_msd(x ))| where f_mpd and f_msd represent the feature vectors extracted from the MPD and MSD, respectively.

[0038] RVQ side uses codeword entropy regularization to suppress the "hot codeword" phenomenon, and dynamically gates the selection with sparsification or temperature annealing. In terms of training strategy, the high-resolution discriminator and phase branch are initially closed, and then gradually opened after convergence of the amplitude and low-resolution spectral items. The segment length and batch size are linearly pulled up according to the GPU memory; the optimizer uses Adam / AdamW, and the initial learning rate magnitudes and cooperates with cosine annealing or multi-stage decay. To approach the landing scene, noise / reverb / band-limited distortion can be injected at a probability of 10-20%; for end-side deployment, distillation, INT8 quantization, and structured pruning can be used as auxiliary means.

[0039] Further, for scenarios that further squeeze bits, combine voice activity detection (VAD) and transient detector (such as burst / scratch sound detection) to do time-varying modulation on the gating score, appropriately increase the upper limit for voiced segments or high-complexity transient segments, and reduce the time axis layer mask sequence as a Markov chain for context entropy coding. This mechanism significantly reduces the time average of side information and total code rate, while maintaining quality in key segments.

[0040] The decoding end continuously monitors the continuity of the layer mask and codeword stream, and when a loss or CRC check failure is detected, immediately falls back to the nearest anchor point, reconstructs the RVQ state and continues decoding from there. If the loss occurs within the anchor point interval and the context probability is high, a soft padding strategy can be used to perform maximum a posteriori (MAP) estimation on the missing mask and attempt a smooth transition; if it fails, it is forced to fall back to ensure overall consistency of the sound quality.

[0041] To ensure compatibility in long-term evolution, the model version, STFT group, RVQ architecture summary, and SRM configuration signature are recorded in the code stream header; the decoding end can establish a backward compatibility table, which can still be decoded stably under the condition that the SRM is closed or the phase branch is disabled for old version code streams. For end-side devices, the target code rate, layer upper limit, and whether to enable the phase branch can be agreed upon during session establishment through capability negotiation.

[0042] Simulation experiment The disclosed LibriSpeech corpus is used, which has a sampling rate of 16 kHz, and is first down-sampled to 8 kHz. The training data comes from the 800 hours of speech in the LibriSpeech corpus training set, and the test data comes from the 10 hours of speech in the LibriSpeech corpus test set. In order to further evaluate the generalization ability of the model, additional speech tests are performed on the LJSpeech and THCHS-30 data sets.

[0043] Objective index selection uses PESQ (perceptual evaluation of speech quality), STOI (short-time objective intelligibility), and LSD (log spectral distance) as objective indexes for experiments.

[0044] Subjective index selection MUSHRA is a commonly used subjective scoring method in speech enhancement and speech quality evaluation. The subjects hear multiple audio samples simultaneously, including the sample to be tested, the hidden reference, and the anchor point, and directly compare and score, and then calculate the average of all subject scores to obtain the corresponding MUSHRA value. The hidden reference refers to the original reference signal (highest quality) that is not processed in the test, but the identity of which is not explicitly indicated to the subjects, and is used to test whether the score is objective. In an ideal case, the reference should get the highest score. The anchor point refers to intentionally adding low-quality audio (such as low-bit rate encoding or severely distorted samples) as a "baseline" for scoring, to help calibrate the subjects' scoring scale. The MUSHRA value scoring interval is 0-100.

[0045] To evaluate the subjective quality, the present disclosure adopts the MUSHRA index: 10 listeners (5 females and 5 males, aged 23-32) score 20 randomly selected utterances from the LibriSpeech test set.

[0046] The experimental results show that the present disclosure significantly outperforms existing methods in both objective indicators and subjective evaluation at a super-low code rate of 0.6-0.8 kbps, and has important application prospects.

[0047] Embodiment 2 In an embodiment of the present disclosure, a speech reconstruction system based on entropy coding residual quantization and spectral repair is provided, comprising: A data acquisition module for acquiring an original speech waveform and performing preprocessing; A speech reconstruction module for inputting the preprocessed speech waveform into a neural speech coding and decoding model to output a reconstructed speech; The processing process of the preprocessed speech waveform in the neural speech coding and decoding model includes: The input speech waveform first enters the encoder, the encoder adopts a multi-layer convolution residual structure and hierarchical down-sampling to map the input speech waveform into an acoustic latent representation, the acoustic latent representation is input into a residual vector quantization module, residual quantization is performed on the acoustic latent representation layer by layer, a dynamic number of layers selection mechanism based on gating and an entropy regularization constraint are introduced to adaptively allocate bits among different speech segments, and quantized features are obtained, the quantized features are reconstructed to obtain reconstructed acoustic features of the latent representation, the reconstructed acoustic features are input into the decoder, and are restored into a time domain waveform through hierarchical up-sampling and a residual block, the time domain waveform is mapped to a log amplitude spectrum domain through a spectrum repair module, a prediction residual is predicted in the log amplitude spectrum domain, and confidence gating fusion is performed to obtain a complex spectrum, and the reconstructed speech is output after time domain synthesis.

[0048] Embodiment 3 In an embodiment of the present disclosure, a non-transitory computer-readable storage medium is provided for storing computer instructions, which, when executed by a processor, implement the speech reconstruction method based on entropy coding residual quantization and spectrum repair.

[0049] Embodiment 4 In an embodiment of the present disclosure, an electronic device is provided, comprising a processor, a memory, and a computer program; wherein the processor is connected with the memory, and the computer program is stored in the memory; when the electronic device is running, the processor executes the computer program stored in the memory, so that the electronic device implements the speech reconstruction method based on entropy coding residual quantization and spectrum repair.

[0050] The present disclosure is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present disclosure. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and the combination of flows and / or blocks in the flowcharts and / or block diagrams can be implemented by computer program instructions. These computer program instructions can be provided to a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate a means for implementing the functions specified in the flowcharts and / or block diagrams. Figure 1 The means for implementing one or more functions specified in one or more flows and / or blocks. Figure 1 The means for implementing one or more functions specified in one or more flows and / or blocks.

[0051] These computer program instructions can also be loaded onto a computer or other programmable data processing device to cause a series of operation steps to be performed on the computer or other programmable data processing device to generate a computer-implemented process, so that the instructions executed on the computer or other programmable data processing device provide a means for implementing the functions specified in the flowcharts and / or block diagrams. Figure 1 The means for implementing one or more functions specified in one or more flows and / or blocks.Figure 1 the steps of the functions specified in the one or more blocks.

[0052] Although the specific embodiments of the present disclosure have been described in detail above with reference to the accompanying drawings, the above description is not intended to limit the scope of protection of the present disclosure, and those skilled in the art should understand that various modifications or changes made on the basis of the technical solutions of the present disclosure without creative labor are still within the scope of protection of the present disclosure.

Claims

1. Speech reconstruction method based on entropy coding residual quantization and spectral repair, characterized in that, include: Acquire the raw speech waveform and perform preprocessing; The preprocessed speech waveform is input into the neural speech encoding and decoding model, and the reconstructed speech is output. The processing of the preprocessed speech waveform in the neural speech encoding / decoding model includes: The input speech waveform first enters the encoder, which employs a multi-layer convolutional residual structure and hierarchical downsampling to map the input speech waveform into an acoustic latent representation. The acoustic latent representation is then input to the residual vector quantization module, which performs residual quantization layer by layer on the acoustic latent representation. A gated dynamic layer selection mechanism and entropy regularization constraint are introduced to adaptively allocate bits among different speech segments, resulting in quantized features. The quantized features are then reconstructed to obtain the reconstructed acoustic features of the latent representation. The reconstructed acoustic features are input to the decoder, which is then restored to a time-domain waveform through hierarchical upsampling and residual blocks. The time-domain waveform is then mapped to the logarithmic amplitude spectrum domain through a spectrum restoration module. The residual is predicted in the logarithmic amplitude spectrum domain and confidence-gated fusion is performed to obtain a complex spectrum. Finally, after time-domain synthesis, the reconstructed speech is output.

2. The speech reconstruction method based on entropy coding residual quantization and spectral repair as claimed in claim 1, wherein, The process of acquiring the original speech waveform and performing preprocessing includes: The original speech waveform is acquired at a set sampling rate. The acquired original speech waveform is then normalized or peak-limited, DC components are removed, and light pre-noise reduction is performed as a preprocessing operation.

3. The speech reconstruction method based on entropy coding residual quantization and spectral repair of claim 1, wherein, The input speech waveform first enters the encoder, which employs a multi-layer convolutional residual structure and hierarchical downsampling to map the input speech waveform into an acoustic latent representation, including: The encoder adopts the SEANet structure, which consists of several residual convolutional units cascaded with a hierarchical downsampling structure. The residual convolutional unit contains pointwise convolution, dilated convolution, and normalization / activation order. After multiple levels of downsampling, a latent tensor with a set time length is obtained. The latent tensor is then unified in dimension to obtain an acoustic latent representation.

4. The speech reconstruction method based on entropy coding residual quantization and spectral repair of claim 1, wherein, The process involves inputting the acoustic latent representation into the residual vector quantization module, performing residual quantization layer by layer on the acoustic latent representation, and introducing a gate-based dynamic layer selection mechanism and entropy regularization constraints to adaptively allocate bits among different speech segments, including: Residual quantization is performed layer by layer on the acoustic latent representation. At each layer, the current residual is quantized with the nearest neighbor of the set codebook to obtain the index and the reconstruction vector, and the residual is updated. To adaptively select the number of layers at a given code rate, a dynamic layer selection mechanism based on gating is introduced. A lightweight scoring network generates an importance score for each layer, and a dynamic gating normalization obtains a gating probability. The time dimension is averaged to obtain a segment-level score, and the mapping relationship between the target code rate and the available layer depth is calculated to select the top n layers as the activation layer set. To balance codeword usage, normalized entropy regularization is added to the empirical distribution on each codebook, with small weights.

5. The speech reconstruction method based on entropy coding residual quantization and spectral repair of claim 4, wherein, The activation layer set is encoded into a binary layer mask of length . The layer mask exhibits Markov property on the time axis. The conditional distribution is estimated using a first-order context model and written into the bit stream using arithmetic coding, thereby constraining the side information bits to a very small overhead range. Subsequently, the codeword index sequence of each layer is written in layer order to obtain the quantized features.

6. The speech reconstruction method based on entropy coding residual quantization and spectral repair as claimed in claim 1, wherein, The reconstructed acoustic features obtained by reconstructing the quantized features to obtain the potential representation include: Read the quantization features and first parse the header information and anchor points; if it is a random start-up or packet loss recovery scenario, start the recovery of RVQ state from the nearest anchor point, then solve the layer mask and codeword index of each layer in time order, reconstruct the quantization features and stack them layer by layer to obtain the reconstructed acoustic features of the potential representation, and arrange them into a channel priority layout that conforms to the decoder input.

7. The speech reconstruction method based on entropy coding residual quantization and spectral repair as claimed in claim 1, wherein, The time-domain waveform is mapped to a log-magnitude spectrum domain by a spectrum repair module, residual errors are predicted in the log-magnitude spectrum domain, and confidence gated fusion is performed to obtain a complex spectrum, including: The reconstructed acoustic feature of the latent representation is input into a decoder of the SEANet structure, restored to a time-domain waveform through hierarchical upsampling and a residual block, projected to a frequency domain through multiple groups of STFT, and the amplitude part is taken as a logarithm and added with a numerical stabilizing item, residual errors and a confidence map are obtained through a light U-Net, a repaired log-magnitude is obtained in a gated manner, a phase gradient correction amount is predicted, a phase is formed, and a final complex spectrum is obtained based on the phase.

8. Speech reconstruction system based on entropy coding residual quantization and spectral repair characterized in that, Comprise: The data acquisition module is used for acquiring an original speech waveform and performing preprocessing; The speech reconstruction module is used for inputting the preprocessed speech waveform into a neural speech coding and decoding model to output a reconstructed speech. The processing process of the preprocessed speech waveform in the neural speech coding and decoding model comprises: The input speech waveform first enters an encoder, the encoder adopts a multi-layer convolution residual structure and hierarchical downsampling to map the input speech waveform to an acoustic latent representation, the acoustic latent representation is input into a residual vector quantization module, residual quantization is performed on the acoustic latent representation layer by layer, a dynamic layer number selection mechanism based on gating and an entropy regularization constraint are introduced to adaptively allocate bits between different speech segments, a quantized feature is obtained, the quantized feature is reconstructed to obtain a reconstructed acoustic feature of the latent representation, the reconstructed acoustic feature is input into a decoder, restored to a time-domain waveform through hierarchical upsampling and a residual block, mapped to a log-magnitude spectrum domain through a spectrum repair module, residual errors are predicted in the log-magnitude spectrum domain, and confidence gated fusion is performed to obtain a complex spectrum, and the reconstructed speech is output after time-domain synthesis.

9. A non-transitory computer-readable storage medium, comprising: The non-transitory computer readable storage medium is used for storing computer instructions, and the computer instructions are executed by the processor to implement the speech reconstruction method based on entropy coding residual quantization and spectrum repair according to any one of claims 1-7.

10. An electronic device, comprising: Comprise: A processor, a memory and a computer program; wherein the processor is connected with the memory, the computer program is stored in the memory, and when the electronic device is running, the processor executes the computer program stored in the memory to enable the electronic device to implement the speech reconstruction method based on entropy coding residual quantization and spectrum repair according to any one of claims 1-7.

Citation Information

Patent Citations

  • Voice compression method and system based on multi-scale residual attention

    CN118335092A

  • Audio coding and decoding method and device based on neural network, equipment and storage medium

    CN119152863A

  • Audio encoding and decoding method and system

    CN120089147A

  • Voice signal compression method and device, equipment and medium

    CN120375835A

  • 3D Gaussian sputtering compression method and system based on residual quantization and dynamic pruning

    CN120411264A

Cited By

  • Streaming spatial audio separation method, streaming spatial audio separation equipment and vehicle-mounted audio system

    CN121687094A

  • Voice reconstruction method and system based on gating re-calibration and route weighting

    CN121983072A