Personalized Speech Decoding via Speaker Group Segmentation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current speech signal encoding and decoding technologies fail to effectively adapt to individual speaker characteristics, leading to suboptimal speech quality restoration across varying environmental conditions.

Innovation Solution

A method and apparatus that encode and decode speech signals by determining the speaker group of the input speech signal, generating a bit stream with encrypted speaker information, and using a personalized neural vocoder trained for each speaker group to improve speech quality restoration.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Device complexity

If a single non-personalized decoder is used for all speakers, then device complexity is reduced, but speech quality restoration deteriorates

Engineering Contradiction:
Improvedecoder structureVSAvoidspeech quality restoration
Core Design Contradiction:
Device complexityVSManufacturing precision

Solution Approach 1:

The decoder is segmented into multiple specialized decoders, each trained for a specific speaker group. The system divides the general decoding task into specialized sub-tasks based on speaker characteristics, allowing each decoder to optimize for its target group while maintaining overall system efficiency.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system dynamically selects the appropriate decoder based on real-time speaker identification. The decoder configuration changes adaptively according to the detected speaker group, transitioning from a static single-decoder approach to a dynamic multi-decoder selection mechanism that optimizes performance for each speaker.

Inventive Principle:
Principle #15Dynamics

2Manufacturing precision

If speaker group information is extracted and encoded, then speech quality restoration is improved, but data processing complexity increases

Engineering Contradiction:
Improvespeech quality restorationVSAvoiddata processing
Core Design Contradiction:
Manufacturing precisionVSDevice complexity

Solution Approach 1:

Speaker group information is extracted and encoded in advance during the encoding phase. This preliminary classification of speaker characteristics allows the decoder to quickly select the appropriate model without performing complex analysis during real-time speech restoration, shifting processing complexity to the encoding stage.

Inventive Principle:
Principle #10Preliminary action

3Manufacturing precision

If multiple personalized decoders are maintained for different speaker groups, then speech quality restoration is improved, but memory requirements increase

Engineering Contradiction:
Improvespeech quality restorationVSAvoidmemory storage
Core Design Contradiction:
Manufacturing precisionVSQuantity of substance

Solution Approach 1:

The multiple decoders are designed to handle universal speech restoration tasks across different speaker groups. Each decoder is trained on diverse data within its speaker group, allowing it to generalize effectively. This multi-functionality approach enables a single decoder to serve multiple purposes within its designated speaker group, reducing the need for excessively specialized models.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS20250104724A1Method and apparatus for encoding/decoding neural network-based personalized speech
Publication Date: 2025.03.27 ELECTRONICS & TELECOMM RES INST
  • US20250104724A1 patent drawing
  • US20250104724A1 patent drawing
  • US20250104724A1 patent drawing

AI summary

A method and apparatus for encoding/decoding a neural network-based personalized speech are provided. The method includes outputting a first bit stream in which an input speech signal is encrypted, based on the input speech signal, and outputting a second bit stream in which speaker information of the input speech signal is encrypted, based on the input speech signal.