Mono-to-Stereo Audio Processor Using Delayed Signal Summation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing audio signal processing methods fail to effectively convert mono signals to stereo for headphone listening without causing in-head localization, which is unpleasant and unnatural, and often require excessive processing power.
Innovation Solution
An audio processor that uses a combination of delays and summations to generate stereo output signals from a mono input, with specific delay ranges and optional filtering to simulate acoustic reflections, allowing for out-of-head localization with minimal processing power.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of manufacture
If traditional stereo reverberation processing is applied to convert mono signal to stereo, then a more pleasant sound is achieved, but the processing power requirement becomes excessively large
Solution Approach 1:
The patent segments the reverberation processing into multiple independent delay lines with specific delay times (e.g., 10ms, 20ms, 30ms, 40ms, 50ms, 60ms, 70ms, 80ms, 90ms, 100ms). Each delay line processes a portion of the reverberation effect independently, allowing the system to achieve complex stereo reverberation with manageable processing requirements rather than using a single complex processing block.
Solution Approach 2:
The patent employs periodic delay patterns where delay lines are arranged in sequences with regular time intervals. The delay times follow periodic patterns (e.g., increasing by fixed intervals), creating a structured reverberation effect that is computationally efficient while maintaining natural sound characteristics.
2Ease of operation
If binaural synthesis with HRTFs is used to achieve out-of-head localization, then sound can be placed outside the head, but the sound color (timbre) is adversely affected and location precision remains poor
Solution Approach 1:
The patent applies different delay characteristics to different frequency components of the audio signal. Low-frequency components receive different delay treatments compared to high-frequency components, allowing each frequency range to achieve optimal localization效果 while preserving natural timbre characteristics specific to that frequency range.
Solution Approach 2:
The patent uses dynamic delay time adjustments where the delay periods are not fixed but vary based on the audio content and desired spatial effect. This allows the system to adaptively optimize localization precision and timbre preservation for different types of audio signals and listening conditions.
3Adaptability or versatility
If frequency band selective allocation is used to convert mono to stereo, then stereo content is created, but in-head localization occurs and excessive filters are required
Solution Approach 1:
The patent extracts the essential localization-critical frequency components from the audio signal and processes them through specific delay lines, while leaving other frequency components to be handled by simpler processing paths. This extraction approach creates effective stereo content without requiring extensive filtering of the entire frequency spectrum.
Solution Approach 2:
The patent designs delay lines that serve multiple functions simultaneously: they provide time-based stereo separation, frequency-dependent localization control, and implicit filtering through their delay characteristics. This multi-functionality reduces the need for separate dedicated filter components while achieving comprehensive stereo processing.
Data Source
Figure 1~2
Figure 3~4
Figure 5~6
AI summary
An audio processor (or mono-to-stereo converter) arranged to receive a single channel audio input signal X and generate a set of stereo audio output signals L, R in response. The outputs L, R are based on four delayed versions S1, S2, S3, S4 of the input signal X. S1 is delayed by delay d1 in relation to X, and S2 is delayed by delay d2 in relation to S1. S3 is delayed by a delay d3 in relation to X, and S4 is delayed by delay d4 in relation to S3. The output L is then generated as a sum of X, S1, and S4, while the output R is generated as a sum of X, S2, and S3. Delays d1 and d3 are selected to be different and within a range of 20 ms to 100 ms, e.g. d1=50 ms and d3 =60 ms. Delays d2 and d4 are selected to be within 50 µs to 1 ms, e.g. 450-650 µs. Such processor produces a stereo signal suited for headphone listening without the feeling of in-head localization and still with a natural timbre. Additionally, low-pass or band-pass filters and appropriate gains can be applied for further refinement. The audio processor can be implemented with a low signal processing requirement and is thus suited as mono-to-stereo converter in portable equipment such as mobile phones, hearing aids etc.