Speech Signal Pre-Augmented Filtering for Cross-Network Intelligibility
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The existing speech encoding/decoding processes across different networks result in significant reduction of audio signal quality due to multiple lossy encoding/decoding processes, leading to poor speech intelligibility during cross-network and cross-platform voice communications.
Innovation Solution
A method that involves feature recognition on speech signals to determine user groups with distinct voice characteristics, followed by pre-augmented filtering using specific filter coefficients to minimize signal loss, thereby enhancing speech signal clarity before cascade encoding/decoding.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If multiple cascade encoding/decoding processes are performed to support cross-network and cross-platform voice communications, then compatibility between different network terminals is improved, but speech signal quality and intelligibility deteriorate
Solution Approach 1:
The patent applies pre-augmented filtering to the speech signal before it undergoes cascade encoding/decoding processes. This preliminary action compensates for the expected quality loss by enhancing specific frequency components in advance, ensuring that the speech signal maintains better intelligibility after passing through multiple encoding/decoding stages across different networks
2Loss of information
If pre-augmented filtering is applied to compensate for signal loss, then speech signal quality is improved, but processing complexity increases
Solution Approach 1:
The patent applies pre-augmented filtering with different filter coefficients tailored to specific user groups (e.g., male voice, female voice, child voice). Instead of applying a general complex processing to all signals, the system selectively applies targeted filtering based on voice characteristics, reducing unnecessary processing complexity while maintaining quality improvement for relevant frequency components
Data Source
AI summary
A method for improving speech signal intelligibility is performed at a device. A speech signal is obtained. A correspondence between the speech signal and a respective user group among different user groups having distinct voice characteristics is identified. Pre-encoding signal augmentation is performed on the speech signal with a respective pre-augmentation filtering coefficient that corresponds to the respective user group to obtain a group-specific pre-augmented speech signal. The device encodes the pre-augmented speech signal for subsequent transmission through the voice communication channel. An encoded version of the pre-augmented speech signal has reduced loss of signal quality as compared to an encoded version of the speech signal that is obtained without the pre-encoding signal augmentation.


