Sound Source Localization via Room Shape Estimation and Template Matching

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional sound source localization systems suffer from low accuracy and slow response in various environments, necessitating a novel approach to improve accuracy and speed.

Innovation Solution

A sound source localization system integrating a microphone array, a room shape estimator, a lookup table, and a localizer, which determines the location of a sound source by utilizing a room shape map, template voice features, and similarity measures.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If conventional sound source localization systems are used, then the system structure is simple, but the localization accuracy and response speed are low

Engineering Contradiction:
Improvelocalization accuracyVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The system performs preliminary actions by pre-calculating and storing room shape templates and voice feature templates in lookup tables before actual localization occurs. The room shape estimator pre-processes acoustic environment data to create location maps, and the voice feature extractor pre-processes speech signals to create reference templates. This preliminary preparation enables faster and more accurate real-time localization without increasing computational complexity during the actual localization process.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The localization system is segmented into distinct functional modules: room shape estimator, voice feature extractor, lookup tables, and localizer. Each module performs a specific function - the room shape estimator determines acoustic environment characteristics, the voice feature extractor processes speech signals, the lookup tables store pre-computed references, and the localizer performs final localization. This segmentation allows each component to be optimized independently while maintaining overall system accuracy and efficiency.

Inventive Principle:
Principle #1Segmentation

2Speed

If conventional sound source localization systems are used, then the system complexity is low, but the response speed is slow

Engineering Contradiction:
Improveresponse speedVSAvoidsystem complexity
Core Design Contradiction:
SpeedVSDevice complexity

Solution Approach 1:

The system performs preliminary actions by pre-calculating and storing room shape templates and voice feature templates in lookup tables before actual localization occurs. The room shape estimator pre-processes acoustic environment data to create location maps, and the voice feature extractor pre-processes speech signals to create reference templates. This preliminary preparation enables faster and more accurate real-time localization without increasing computational complexity during the actual localization process.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system creates copies of reference data by generating template voice features for virtual sound sources at different locations and storing them in lookup tables. These templates serve as reference copies that can be quickly compared against actual incoming signals. The room shape templates act as copied representations of acoustic environment characteristics, enabling rapid matching and localization without performing complex calculations for each new sound source detection.

Inventive Principle:
Principle #26Copying

3Measurement precision

If conventional sound source localization systems are used, then the processing mechanism is simple, but the localization accuracy in various environments is low

Engineering Contradiction:
Improvelocalization accuracyVSAvoidprocessing mechanism complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The system changes parameters by adapting to different acoustic environments through the room shape estimator, which determines room shape characteristics and updates the location maps in the lookup tables. The system modifies its reference templates based on the specific acoustic properties of each environment, allowing it to maintain high localization accuracy across diverse settings. This parameter adaptation enables the system to optimize its localization performance for each unique acoustic environment.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The system performs preliminary actions by pre-calculating and storing room shape templates and voice feature templates in lookup tables before actual localization occurs. The room shape estimator pre-processes acoustic environment data to create location maps, and the voice feature extractor pre-processes speech signals to create reference templates. This preliminary preparation enables faster and more accurate real-time localization without increasing computational complexity during the actual localization process.

Inventive Principle:
Principle #10Preliminary action

Applied Scientific Principles

This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.

Function Achieved in This Case

The system achieves improved accuracy and faster response in localizing sound sources by leveraging multiple mechanisms, including room shape estimation and template voice feature comparison.

Implementation Method 1

a microphone array 11 composed of a plurality of microphones 111 each configured to convert sound wave into a corresponding voice signal

Methodology Applied
Scientific EffectAcoustic transduction:

Data Source

PatentUS12323779B2Sound source localization system
Publication Date: 2025.06.03 HIMAX TECH LTD
  • US12323779B2 patent drawing
  • US12323779B2 patent drawing
  • US12323779B2 patent drawing

AI summary

A sound source localization system includes a microphone array composed of a plurality of microphones each converting sound wave into a corresponding voice signal; a room shape estimator that determines a room shape including a location map and a corresponding template voice feature map composed of template voice features associated with a virtual sound source disposed at different locations respectively, and outputs a room reliability indicating confidence about the determined room shape; a lookup table (LUT) that pre-stores the location map and the corresponding template voice feature map; and a localizer that determines a location of a sound source according to the room reliability and similarity between a voice feature associated with the sound source and the template voice features of the template voice feature map.