Dynamic Audio Equalization Target Profiles via Semantic Clustering

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing dynamic equalization (DEQ) techniques require manual selection of target profiles, which can lead to perceptual degradation of audio content due to inappropriate choices, especially when multiple reference audio content items are available.

Innovation Solution

Automatically generate and select target profiles by clustering reference audio content items based on audio characteristics and semantic labels, using methods like k-means clustering and Euclidean distance metrics to ensure appropriate equalization.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If manual selection of target profiles is used, then user control and flexibility are improved, but perceptual degradation occurs due to inappropriate choices

Engineering Contradiction:
Improveuser controlVSAvoidaudio quality
Core Design Contradiction:
Ease of operationVSReliability

Solution Approach 1:

The system automatically generates multiple target profiles from reference audio content and autonomously selects the most appropriate profile based on similarity metrics, eliminating the need for manual user selection while preventing perceptual degradation through algorithmic accuracy

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system computes similarity metrics between input audio content and available target profiles to automatically select the best match, using quantitative feedback to ensure appropriate profile selection and maintain audio quality

Inventive Principle:
Principle #23Feedback

2Adaptability or versatility

If automatic generation of multiple target profiles is implemented, then adaptability to different audio content is improved, but device complexity increases

Engineering Contradiction:
Improveprofile matchingVSAvoidprocessing complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The system segments reference audio content into frequency bands and extracts features for each band, creating multiple target profiles through systematic division of the audio spectrum and organized storage of profile data

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system transforms audio content into parameter space by extracting spectral contour parameters, dynamic range parameters, and other quantitative features, enabling automatic profile generation and selection through parameter comparison rather than complex audio analysis

Inventive Principle:
Principle #35Parameter changes

3Manufacturing precision

If clustering algorithms are used to generate target profiles, then manufacturing precision of profile selection is improved, but loss of time in processing increases

Engineering Contradiction:
Improveprofile generation accuracyVSAvoidprocessing time
Core Design Contradiction:
Manufacturing precisionVSLoss of time

Solution Approach 1:

The system pre-generates multiple target profiles from reference audio content and stores them in a database before actual DEQ processing, so that during runtime only fast similarity comparisons are needed rather than full profile generation

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system extracts only the essential features (spectral contour, dynamic range, quantiles) needed for profile generation and selection, separating the critical audio characteristics from the full audio content to reduce processing complexity and time

Inventive Principle:
Principle #2Taking out (Extraction)

Data Source

PatentUS12375053B2Automatic generation and selection of target profiles for dynamic equalization of audio content
Publication Date: 2025.07.29 DOLBY LABORATORIES LICENSING CORP
  • US12375053B2 patent drawing
  • US12375053B2 patent drawing
  • US12375053B2 patent drawing

AI summary

In an embodiment, a method comprises: filtering reference audio content items to separate the reference audio content items into different frequency bands; for each frequency band, extracting a first feature vector from at least a portion of each of the reference audio content items, wherein the first feature vector includes at least one audio characteristic of the reference audio content items; obtaining at least one semantic label from at least a portion of each of the reference audio content items; obtaining a second feature vector consisting of the first feature vectors per frequency band and the at least one semantic label; generating, based on the second feature vector, cluster feature vectors representing centroids of clusters; separating the reference audio content items according to the cluster feature vectors; and computing an average target profile for each cluster based on the reference audio content items in the cluster.