Audio Scene Mapping via Frequency Segmentation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current systems for processing audio signals from multiple recording devices in an event space lead to inefficiencies due to redundant processing and bandwidth load, as each device captures and uploads the same audio signal, resulting in increased processing capacity and power requirements without significant quality improvement for the end user.
Innovation Solution
An apparatus and method that receive audio signals from multiple devices, scale and combine them based on energy estimation and frequency range, to generate a combined audio signal representation, thereby reducing redundant processing and optimizing bandwidth usage by assigning each device a specific frequency range and processing only a portion of the audio spectrum.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If each recording device captures and uploads the complete audio signal, then the audio quality is preserved, but the processing load and bandwidth consumption increase significantly
Solution Approach 1:
The audio frequency spectrum is divided into multiple frequency bands, with each recording device capturing only a specific frequency range rather than the complete spectrum. This segmentation allows multiple devices to capture different portions of the audio signal simultaneously, reducing redundant processing while maintaining overall audio quality when the frequency bands are combined at the server.
2Reliability
If multiple devices record the same audio scene, then the audio coverage is improved, but the bandwidth load increases due to redundant data transmission
Solution Approach 1:
Each recording device is assigned a specific frequency band to capture, creating a segmented division of the audio spectrum. This ensures comprehensive audio coverage across all frequency ranges while minimizing bandwidth consumption, as each device transmits only its designated frequency portion rather than duplicating the complete audio signal.
Solution Approach 2:
Different frequency bands are captured by different devices based on their local optimization for specific frequency ranges. This local quality approach allows the system to achieve global audio coverage efficiency by having each device specialize in capturing its assigned frequency band with optimal quality.
3Measurement precision
If all recording devices process the complete audio signal, then the processing accuracy is maintained, but the power consumption increases
Solution Approach 1:
The processing task is segmented across multiple devices, with each device responsible for processing only its assigned frequency band. This segmentation maintains processing accuracy within each frequency range while significantly reducing the computational load and power consumption for each individual device compared to processing the complete audio spectrum.
4Loss of information
If multiple devices upload full audio signals, then the data completeness is ensured, but the bitrate required increases
Solution Approach 1:
The audio data is segmented into frequency bands, with each device uploading only its assigned portion. This segmentation ensures that the combined data from all devices is complete across the full frequency spectrum while minimizing the total bitrate required, as redundant frequency information is eliminated.
Solution Approach 2:
The segmented frequency band data from multiple recording devices is merged at the server to reconstruct the complete audio signal. This combining process ensures data completeness across all frequency ranges while achieving efficient bandwidth utilization, as the merged result contains each frequency band only once rather than multiple redundant copies.
Data Source
AI summary
An apparatus comprising: at least one processor and at least one memory including computer code for one or more programs, the at least one memory and the computer code configured to with the at least one processor cause the apparatus to at least perform: receiving at least two signals comprising at least two audio signals from at least two recording apparatus recording within an audio scene an audio source, wherein the first of the at least two audio signals is configured to represent a first frequency range and the second of the at least two audio signals is configured to represent a second frequency range; scaling the at least two audio signals; and combining the at least two audio signals to generate a combined audio signal representation of the audio source.


