Bias-Corrected Speech Level Determination via Gaussian Parametric Modeling

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional methods for estimating speech level in audio signals are heavily biased by signal-to-noise ratio and amplitude compression, leading to inaccurate measurements.

Innovation Solution

A method that uses a parametric spectral model, specifically a Gaussian model, to determine speech level distributions across frequency bands, correcting for biases introduced by noise and compression by employing predetermined correction values based on a reference speech model.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If conventional speech level estimation methods are used, then the measurement process is simple, but the measurement precision deteriorates due to bias from noise and compression

Engineering Contradiction:
Improvespeech level measurement accuracyVSAvoidmeasurement system complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent applies preliminary action by pre-calculating correction values based on a reference speech model before actual speech level measurement. The system stores correction values corresponding to different signal-to-noise ratios and compression ratios, which are then applied during measurement to eliminate bias without requiring complex real-time analysis. This resolves the contradiction by preparing correction data in advance, maintaining measurement simplicity while improving accuracy.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent changes parameters by introducing correction values that compensate for signal-to-noise ratio and compression ratio variations. Instead of directly measuring speech level from the distorted signal, the system transforms the measurement process by applying parameter-based corrections derived from a reference model. This allows accurate speech level determination despite variations in noise and compression conditions, resolving the precision-accuracy contradiction.

Inventive Principle:
Principle #35Parameter changes

2Adaptability or versatility

If signal-to-noise ratio varies, then adaptability to different environments is improved, but measurement precision deteriorates due to bias introduction

Engineering Contradiction:
Improveenvironmental adaptabilityVSAvoidspeech level measurement accuracy
Core Design Contradiction:
Adaptability or versatilityVSMeasurement precision

Solution Approach 1:

The patent addresses this contradiction by changing the measurement parameters through correction values that are specifically tailored to different signal-to-noise ratios and compression ratios. The system maintains adaptability to various environmental conditions while preserving measurement precision by applying condition-specific corrections rather than using a fixed measurement approach. This resolves the contradiction by making the measurement system adaptable without sacrificing accuracy.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent employs feedback by comparing the measured speech signal characteristics against a reference speech model to determine appropriate correction values. The system uses the observed signal conditions (noise level, compression ratio) to select or calculate correction factors that compensate for bias. This feedback mechanism enables the system to adapt to different environments while maintaining precise measurements, resolving the contradiction between adaptability and precision.

Inventive Principle:
Principle #23Feedback

3Reliability

If amplitude compression is applied, then signal robustness is improved, but measurement precision deteriorates due to bias in level estimation

Engineering Contradiction:
Improvesignal robustnessVSAvoidspeech level measurement accuracy
Core Design Contradiction:
ReliabilityVSMeasurement precision

Solution Approach 1:

The patent resolves this contradiction by introducing correction values that specifically address the bias introduced by amplitude compression. The system measures signal characteristics including compression ratio and applies corresponding parameter changes through correction factors derived from a reference model. This allows the system to maintain signal robustness through compression while correcting the measurement bias, thereby preserving measurement precision despite the presence of compression.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent uses an intermediary approach by introducing correction values as a mediating element between the compressed speech signal and the final level measurement. Rather than directly measuring the compressed signal or attempting to reverse compression, the system uses correction values as an intermediary to compensate for compression-induced bias. This intermediary correction mechanism maintains both signal robustness and measurement precision, resolving the contradiction.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentEP2828853B1Method and system for bias corrected speech level determination
Publication Date: 2018.09.12 DOLBY LABORATORIES LICENSING CORP
  • EP2828853B1 patent drawingFigure 1
  • EP2828853B1 patent drawingFigure 2
  • EP2828853B1 patent drawingFigure 3~4

AI summary

Method for measuring level of speech determined by an audio signal in a manner which corrects for and reduces the effect of modification of the signal by the addition of noise thereto and/or amplitude compression thereof, and a system configured to perform any embodiment of the method. In some embodiments, the method includes steps of generating frequency banded, frequency-domain data indicative of an input speech signal, determining from the data a Gaussian parametric spectral model of the speech signal, and determining from the parametric spectral model an estimated mean speech level and a standard deviation value for each frequency band of the data; and generating speech level data indicative of a bias corrected mean speech level for each frequency band, including using at least one correction value to correct the estimated mean speech level for the frequency band, where each correction value has been predetermined using a reference speech model.