Audio Enhancement via Pupil Dilation Cognitive Load

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

State-of-the-art audio codecs face challenges in accurately applying dialog enhancement, as boosting dialog can harm audio quality, and existing dialog detectors are not fully accurate, leading to unboosted dialog and boosted non-dialog, with no universally applicable solution for evaluating quality.

Innovation Solution

A method that selectively applies speech/dialog enhancement based on the cognitive load of the listener, measured by pupil dilation, allowing enhancement only when needed, thus avoiding unnecessary enhancement and improving understanding without degrading overall audio quality.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If dialog enhancement is applied to boost dialog in audio signals, then the intelligibility of dialog is improved, but the overall perceived audio quality deteriorates

Engineering Contradiction:
Improveintelligibility of dialogVSAvoidperceived audio quality
Core Design Contradiction:
Measurement precisionVSObject-affected harmful factors

Solution Approach 1:

The patent applies dialog enhancement selectively only to time segments where dialog is detected, rather than uniformly to the entire audio signal. The processing unit identifies dialog portions and applies enhancement parameters specifically to those segments, leaving non-dialog segments unchanged. This localised approach preserves overall audio quality while improving dialog intelligibility where needed.

Inventive Principle:
Principle #3Local quality

2Speed

If dialog enhancement parameters are provided in the encoded bitstream, then the processing speed is improved, but the accuracy of dialog detection deteriorates

Engineering Contradiction:
Improveprocessing speedVSAvoidaccuracy of dialog detection
Core Design Contradiction:
SpeedVSMeasurement precision

Solution Approach 1:

The patent performs dialog detection and determines enhancement parameters during the encoding phase, storing this information in the bitstream. The decoding unit then retrieves pre-determined parameters without performing real-time detection, significantly reducing processing complexity and improving decoding speed while maintaining detection accuracy through the encoder's more powerful processing capabilities.

Inventive Principle:
Principle #10Preliminary action

3Device complexity

If simple dialog boosting is applied, then the device complexity is reduced, but the adaptability to individual listeners and listening conditions deteriorates

Engineering Contradiction:
Improvecomplexity of enhancement algorithmVSAvoidadaptability to listener needs
Core Design Contradiction:
Device complexityVSAdaptability or versatility

Solution Approach 1:

The patent employs dynamic enhancement parameters that can be adjusted based on different listening conditions and individual listener characteristics. The system allows for adaptive tuning where parameters such as gain, bandwidth, and time constants can be modified according to the specific audio scene, listener preferences, and environmental conditions, moving from static to dynamic adaptation.

Inventive Principle:
Principle #15Dynamics

4Ease of operation

If dialog enhancement is applied without accurate dialog detection, then the ease of operation is improved, but the loss of information increases

Engineering Contradiction:
Improvesimplicity of enhancement applicationVSAvoidunboosted dialog and boosted non-dialog
Core Design Contradiction:
Ease of operationVSLoss of information

Solution Approach 1:

The patent implements a feedback mechanism where the encoded bitstream contains dialog detection information generated during encoding. The decoding unit uses this feedback information to identify which time segments contain dialog and applies enhancement only to those segments. This feedback loop ensures accurate dialog identification without requiring complex real-time detection at the decoder, maintaining both simplicity and accuracy.

Inventive Principle:
Principle #23Feedback

Data Source

PatentEP3444820B1Speech/dialog enhancement controlled by pupillometry
Publication Date: 2024.02.07 DOLBY INTERNATIONAL AB
  • EP3444820B1 patent drawingFigure 1
  • EP3444820B1 patent drawingFigure 2
  • EP3444820B1 patent drawingFigure 3

AI summary

The present disclosure relates to methods for processing a decoded audio signal and for selectively applying speech/dialog enhancement to the decoded audio signal. The present disclosure also relates to a method of operating a headset for computer-mediated reality. A method of processing a decoded audio signal comprises obtaining a measure of a cognitive load of a listener that listens to a rendering of the audio signal, determining whether speech/dialog enhancement shall be applied based on the obtained measure of the cognitive load, and performing speech/dialog enhancement based on the determination. A method of operating a headset for computer-mediated reality comprises obtaining eye-tracking data of a wearer of the headset, determining a measure of a cognitive load of the wearer of the headset based on the eye-tracking data, and outputting an indication of the cognitive load of the wearer of the headset. The present disclosure further relates to corresponding apparatus and systems, and to methods of operating such apparatus and systems.