The invention provides a voice
reconstruction method and
system based on entropy coding residual quantization and
frequency spectrum restoration, and relates to the technical field of
artificial intelligence voice
signal processing, and the method comprises the steps: obtaining an original voice waveform, inputting the voice waveform into a neural voice coding and decoding model, firstly entering a coder to map the input voice waveform into acoustic potential representation, and then entering a
frequency spectrum restoration model; performing residual quantization on the acoustic potential characterization layer by layer through a residual
vector quantization module, introducing a gating-based dynamic layer number selection mechanism and entropy regularization constraint, enabling bits to be adaptively distributed among different voice segments, reconstructing reconstructed acoustic features of the potential characterization, inputting the reconstructed acoustic features into a decoder, restoring the reconstructed acoustic features into a
time domain waveform, and outputting the
time domain waveform. And mapping to a logarithmic magnitude spectrum domain through a spectrum repairing module, predicting a residual error in the logarithmic magnitude spectrum domain and performing confidence gating fusion to obtain a complex spectrum, and outputting after
time domain synthesis to obtain reconstructed speech. According to the invention,
high fidelity, intelligibility and transmission reliability of the voice can be considered at an extremely low
bit rate.