Polyphonic Note Detection Using Narrow-Band Signal Segmentation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing methods for polyphonic note detection in music signals face challenges in accurately distinguishing between notes due to overlapping harmonics and noise, often resulting in reduced information processing and increased latency, especially when dealing with simultaneous notes.
Innovation Solution
The method preserves all available information by using a bank of narrow-band pass filters to generate time-domain signals, allowing for the extraction of various features such as signal envelopes and statistics, rather than relying solely on energy averages, enabling more robust decision-making without significant computational complexity.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of manufacture
If a straightforward Fourier transform with equidistant frequency bands is used for note detection, then the method is simple and computationally efficient, but it fails to accurately detect notes when multiple notes occupy the same frequency band due to harmonic overlap and beat phenomena
Solution Approach 1:
The frequency spectrum is segmented into multiple narrow bands within each constant-Q band, allowing the system to resolve individual spectral components that would otherwise overlap in broader frequency bins. This segmentation enables accurate detection of multiple notes even when their harmonics fall within the same constant-Q band.
Solution Approach 2:
The patent transitions from analyzing only frequency domain energy distribution to incorporating time-domain analysis by examining the temporal evolution of spectral components. By analyzing how spectral peaks evolve over time across multiple narrow bands, the system can distinguish between genuine note components and beat phenomena, resolving ambiguities that frequency-domain-only methods cannot solve.
2Device complexity
If energy averaging is used to characterize frequency bands, then computational complexity is reduced, but information about individual spectral components and their temporal evolution is lost
Solution Approach 1:
Instead of computing a single averaged energy value for each frequency band, the patent extracts individual spectral peaks and their corresponding energies from each narrow band. This extraction preserves information about distinct spectral components while maintaining computational efficiency by focusing only on significant peaks rather than processing all frequency bins uniformly.
Solution Approach 2:
The patent changes the parameter representation from averaged energy values to a set of discrete spectral peak frequencies and their corresponding energies. This parameter transformation allows the system to maintain detailed information about individual spectral components while reducing the data volume by only retaining significant peaks above a threshold, thus balancing information preservation with computational efficiency.
3Measurement precision
If the number of frequency bands is significantly increased to improve detection resolution, then note detection accuracy improves, but processing time and latency increase
Solution Approach 1:
The patent applies different analysis resolutions to different frequency regions by using constant-Q bands with non-uniform width (narrower at low frequencies, wider at high frequencies). This local adaptation of resolution matches the perceptual characteristics of human hearing and the physical characteristics of musical instruments, providing high resolution where needed while reducing processing load in regions where lower resolution suffices.
Solution Approach 2:
Instead of uniformly increasing the number of frequency bands across the entire spectrum, the patent applies finer segmentation only within constant-Q bands where spectral overlap problems occur most frequently. This partial application of high-resolution analysis focuses computational resources on the most problematic frequency regions while maintaining coarser analysis elsewhere, thus improving detection accuracy without proportionally increasing processing latency.
Data Source
Figure 1
Figure 2
Figure 3
AI summary
This is a method and installation in which a time-domain digital audio signal is split into a plurality of narrow-band time-domain digital audio signals confined to specific frequency bands, short-term segments of which are temporarily stored in memory. The method comprises the use of signal processing algorithms for extracting multiple signal features from said short-term segments in a fixed sequence or upon request from a decision-making algorithm. Said decision-making algorithm makes tentative or final decisions about the type of occupancy of frequency bands resulting from the extracted features. Said decision- making algorithm may request from said signal processing algorithms further specific feature extractions from specific short-term segments and make further tentative or final decisions about the type of occupancy of frequency bands resulting from the requested features. Next, said decision-making algorithm stores its tentative decisions and makes final decisions about band occupancy for processing together with results from later short-term segments. Eventually, said decision-making algorithm outputs final decisions derived from current and past short-segments in the form of a set of notes having been played over some recent time interval, together with information as to the timing of each note from the set.