Digital Vocoder Trill Sound Modulation Enhancement

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Digital land mobile radios, particularly those using narrowband vocoders, distort speech sounds with high modulation rates, such as the alveolar trill, leading to intelligibility issues in languages like Spanish and Italian, due to low frame energy analysis rates and aliasing artifacts.

Innovation Solution

The implementation of pre- and post-processing techniques, including frame shifting, energy parameter modification, time expansion/compression, and modulation enhancement filtering, to enhance the modulation index of trill sounds without modifying the vocoder, by detecting and aligning with modulation nulls and adjusting the vocoder parameters.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Quantity of substance

If narrowband vocoders are used for digital radio transmission, then bandwidth efficiency is improved, but speech fidelity for high modulation rate sounds deteriorates

Engineering Contradiction:
Improvebandwidth efficiencyVSAvoidspeech fidelity
Core Design Contradiction:
Quantity of substanceVSManufacturing precision

Solution Approach 1:

The system performs preliminary detection of trill sounds before vocoding and applies pre-processing techniques (frame shifting, time expansion) to prepare the signal for better vocoder handling, preventing distortion before it occurs

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system dynamically changes vocoder parameters (frame shift amount, time expansion factor) based on detected speech characteristics, particularly for trill sounds, to optimize the balance between bandwidth efficiency and speech fidelity

Inventive Principle:
Principle #35Parameter changes

2Manufacturing precision

If vocoder frame rate is increased to improve trill sound encoding, then speech fidelity is improved, but data rate increases

Engineering Contradiction:
Improvetrill sound encoding accuracyVSAvoiddata rate
Core Design Contradiction:
Manufacturing precisionVSQuantity of substance

Solution Approach 1:

The system performs preliminary time expansion on trill sounds before vocoding, which effectively lowers the modulation rate and allows accurate encoding at the standard vocoder frame rate without increasing data rate

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system changes the time expansion parameter dynamically based on detected trill characteristics, adjusting the effective sampling rate for trill sounds while maintaining the standard vocoder operating parameters for other speech

Inventive Principle:
Principle #35Parameter changes

3Manufacturing precision

If frame shifting is applied to align with modulation nulls, then trill sound intelligibility is improved, but processing complexity increases

Engineering Contradiction:
Improvetrill sound intelligibilityVSAvoidsignal processing complexity
Core Design Contradiction:
Manufacturing precisionVSDevice complexity

Solution Approach 1:

The system applies frame shifting selectively only to detected trill sounds rather than all speech, and uses localized processing around modulation nulls, reducing overall processing complexity while maintaining intelligibility improvement where needed

Inventive Principle:
Principle #3Local quality

Data Source

PatentEP3080805B1Method and apparatus for enhancing the modulation index of speech sounds passed through a digital vocoder
Publication Date: 2019.11.13 MOTOROLA SOLUTIONS INC
  • EP3080805B1 patent drawingFigure 1
  • EP3080805B1 patent drawingFigure 2
  • EP3080805B1 patent drawingFigure 3

AI summary

A method and apparatus for enhancing modulation of certain speech sounds, such as trill sounds, are provided for radios which utilize digital vocoders. A digitized speech stream is sampled and the sampling is adjusted to determine, detect and enhance trill nulls in the digitized voice stream by one or more of: frame shifting the digitized speech input stream prior to vocoding, time expanding a digitized speech steam prior to vocoding, time compressing a digitized speech output stream after vocoding, and/or modulation enhancement and filtering of the a digitized speech output stream after vocoding.