Speech Signal Pre-Augmented Filtering for Cross-Network Intelligibility

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

The existing speech encoding/decoding processes across different networks result in significant reduction of audio signal quality due to multiple lossy encoding/decoding processes, leading to poor speech intelligibility during cross-network and cross-platform voice communications.

Innovation Solution

A method that involves feature recognition on speech signals to determine user groups with distinct voice characteristics, followed by pre-augmented filtering using specific filter coefficients to minimize signal loss, thereby enhancing speech signal clarity before cascade encoding/decoding.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If multiple cascade encoding/decoding processes are performed to support cross-network and cross-platform voice communications, then compatibility between different network terminals is improved, but speech signal quality and intelligibility deteriorate

Engineering Contradiction:
Improvecompatibility between different network terminalsVSAvoidspeech signal quality
Core Design Contradiction:
Adaptability or versatilityVSLoss of information

Solution Approach 1:

The patent applies pre-augmented filtering to the speech signal before it undergoes cascade encoding/decoding processes. This preliminary action compensates for the expected quality loss by enhancing specific frequency components in advance, ensuring that the speech signal maintains better intelligibility after passing through multiple encoding/decoding stages across different networks

Inventive Principle:
Principle #10Preliminary action

2Loss of information

If pre-augmented filtering is applied to compensate for signal loss, then speech signal quality is improved, but processing complexity increases

Engineering Contradiction:
Improvespeech signal qualityVSAvoidprocessing complexity
Core Design Contradiction:
Loss of informationVSDevice complexity

Solution Approach 1:

The patent applies pre-augmented filtering with different filter coefficients tailored to specific user groups (e.g., male voice, female voice, child voice). Instead of applying a general complex processing to all signals, the system selectively applies targeted filtering based on voice characteristics, reducing unnecessary processing complexity while maintaining quality improvement for relevant frequency components

Inventive Principle:
Principle #3Local quality

Data Source

PatentUS11605394B2Speech signal cascade processing method, terminal, and computer-readable storage medium
Publication Date: 2023.03.14 TENCENT TECHNOLOGY (SHENZHEN) CO LTD
  • US11605394B2 patent drawing
  • US11605394B2 patent drawing
  • US11605394B2 patent drawing

AI summary

A method for improving speech signal intelligibility is performed at a device. A speech signal is obtained. A correspondence between the speech signal and a respective user group among different user groups having distinct voice characteristics is identified. Pre-encoding signal augmentation is performed on the speech signal with a respective pre-augmentation filtering coefficient that corresponds to the respective user group to obtain a group-specific pre-augmented speech signal. The device encodes the pre-augmented speech signal for subsequent transmission through the voice communication channel. An encoded version of the pre-augmented speech signal has reduced loss of signal quality as compared to an encoded version of the speech signal that is obtained without the pre-encoding signal augmentation.