Unified Speech Audio Codec Module Selection

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing speech and audio codecs fail to provide optimal performance when integrated, as they are optimized for specific signal characteristics, leading to suboptimal processing of unified speech/audio signals in communication and broadcasting services.

Innovation Solution

A unified codec apparatus and method that selects between speech and audio encoding/decoding modules based on input signal characteristics, using a combination of Code Excitation Linear Prediction (CELP) and Modified Discrete Cosine Transform (MDCT) operations, with initialization and multiplexing to generate and decode bitstreams without distortion.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Manufacturing precision

If a unified codec uses a single encoding/decoding module optimized for speech signals, then speech processing performance is improved, but audio signal processing performance deteriorates

Engineering Contradiction:
Improvespeech processing performanceVSAvoidaudio signal processing capability
Core Design Contradiction:
Manufacturing precisionVSAdaptability or versatility

Solution Approach 1:

The codec dynamically switches between speech-optimized module and audio-optimized module based on the characteristics of the input signal. The system determines whether each input signal is a speech signal or an audio signal and selects the appropriate module accordingly, enabling adaptive optimization for different signal types rather than using a static single-module design

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The unified codec achieves multi-functionality by integrating both speech-optimized and audio-optimized encoding/decoding modules within a single system. This allows the codec to handle both speech signals and audio signals effectively, providing universal processing capability across different signal types while maintaining optimal performance for each

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Manufacturing precision

If a unified codec uses a single encoding/decoding module optimized for audio signals, then audio processing performance is improved, but speech signal processing performance deteriorates

Engineering Contradiction:
Improveaudio processing performanceVSAvoidspeech signal processing capability
Core Design Contradiction:
Manufacturing precisionVSAdaptability or versatility

Solution Approach 1:

The codec dynamically switches between audio-optimized module and speech-optimized module based on the characteristics of the input signal. The system determines whether each input signal is an audio signal or a speech signal and selects the appropriate module accordingly, enabling adaptive optimization for different signal types rather than using a static single-module design

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The unified codec achieves multi-functionality by integrating both audio-optimized and speech-optimized encoding/decoding modules within a single system. This allows the codec to handle both audio signals and speech signals effectively, providing universal processing capability across different signal types while maintaining optimal performance for each

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Adaptability or versatility

If the codec switches between different encoding/decoding modules for each frame, then signal processing adaptability is improved, but signal distortion increases due to module structure changes

Engineering Contradiction:
Improvesignal processing adaptabilityVSAvoidsignal quality
Core Design Contradiction:
Adaptability or versatilityVSManufacturing precision

Solution Approach 1:

The codec performs preliminary determination of whether the current frame is a speech frame or an audio frame before selecting the encoding/decoding module. This advance identification allows the system to pre-select the appropriate module structure, avoiding mid-processing switches that would cause distortion and ensuring continuous processing with a single consistent module structure for each frame

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentEP2302623B1Apparatus for encoding and decoding of integrated speech and audio
Publication Date: 2020.04.01 ELECTRONICS & TELECOMM RES INST
  • EP2302623B1 patent drawingFigure 1
  • EP2302623B1 patent drawingFigure 2
  • EP2302623B1 patent drawingFigure 3

AI summary

Provided is an apparatus for integrally encoding and decoding a speech signal and an audio signal. An encoding apparatus for integrally encoding a speech signal and an audio signal, may include: a module selection unit to analyze a characteristic of an input signal and to select a first encoding module for encoding a first frame of the input signal; a speech encoding unit to encode the input signal according to a selection of the module selection unit and to generate a speech bitstream; an audio encoding unit to encode the input signal according to the selection of the module selection unit and to generate an audio bitstream; and a bitstream generation unit to generate an output bitstream from the speech encoding unit or the audio encoding unit according to the selection of the module selection unit.