Auto-regressive Neural Network Speech Reconstruction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional speech coding systems using parametric coders struggle to generate high-quality wide-band speech from narrow-band speech, as they lack sufficient information to accurately reconstruct the signal, even with wide-band extensions.

Innovation Solution

Employing a decoder auto-regressive generative neural network that uses parametric coding parameters to conditionally generate speech, reducing the need for transmitting waveform information and allowing for high-quality speech reconstruction with reduced data transmission.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Loss of energy

If conventional parametric coding is used to compress speech, then data transmission rate is reduced, but speech quality deteriorates because narrow-band parameters cannot reconstruct wide-band speech

Engineering Contradiction:
Improvedata transmission rateVSAvoidspeech reconstruction quality
Core Design Contradiction:
Loss of energyVSManufacturing precision

Solution Approach 1:

The patent introduces an auto-regressive generative neural network as an intermediary between the parametric decoder and the final speech output. This neural network takes the limited parametric coding parameters as input and generates high-quality wide-band speech by learning the mapping from narrow-band parameters to wide-band waveforms, effectively bridging the information gap without requiring additional transmission data

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent transforms the speech representation from traditional parametric coding parameters to a neural network-generated waveform representation. By changing the output parameters from narrow-band spectral parameters to wide-band time-domain waveforms through the neural network, the system achieves high-quality speech reconstruction while maintaining low-bitrate parametric encoding

Inventive Principle:
Principle #35Parameter changes

2Manufacturing precision

If wide-band extension is applied after parametric decoding, then speech quality improves, but the amount of data transmitted remains high because additional waveform information must be sent

Engineering Contradiction:
Improvespeech reconstruction qualityVSAvoiddata transmission volume
Core Design Contradiction:
Manufacturing precisionVSQuantity of substance

Solution Approach 1:

The patent extracts only the essential parametric coding parameters (such as pitch, gain, and spectral parameters) for transmission, removing the need to transmit additional waveform data. The auto-regressive generative neural network then synthesizes the complete wide-band speech from these extracted parameters, achieving wide-band quality without the data overhead of traditional wide-band extension methods

Inventive Principle:
Principle #2Taking out (Extraction)

3Measurement precision

If more parametric parameters are transmitted to improve reconstruction accuracy, then speech quality improves, but transmission complexity and data rate increase

Engineering Contradiction:
Improvereconstruction accuracyVSAvoidcoding parameter complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent replaces the traditional mechanical approach of transmitting multiple detailed parametric parameters with a neural network-based generative model. Instead of increasing parameter transmission to improve accuracy, the system uses the neural network's learned representations to infer missing information, substituting parameter transmission with intelligent synthesis

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Data Source

PatentUS12062380B2Speech coding using auto-regressive generative neural networks
Publication Date: 2024.08.13 GOOGLE LLC
  • US12062380B2 patent drawing
  • US12062380B2 patent drawing
  • US12062380B2 patent drawing

AI summary

Methods, systems, and apparatus, including computer programs encoded on computer storage media, for coding speech using neural networks. One of the methods includes obtaining a bitstream of parametric coder parameters characterizing spoken speech; generating, from the parametric coder parameters, a conditioning sequence; generating a reconstruction of the spoken speech that includes a respective speech sample at each of a plurality of decoder time steps, comprising, at each decoder time step: processing a current reconstruction sequence using an auto-regressive generative neural network, wherein the auto-regressive generative neural network is configured to process the current reconstruction to compute a score distribution over possible speech sample values, and wherein the processing comprises conditioning the auto-regressive generative neural network on at least a portion of the conditioning sequence; and sampling a speech sample from the possible speech sample values.