Audio Mixer Using Unencrypted Power Levels for Speaker Selection

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conferencing systems face computational overload and noise issues due to the increasing number of participants, requiring efficient processing of multiple Real-Time Protocol (RTP) audio packet streams and decryption of Secure Real-Time Protocol (SRTP) packets to determine loudest speakers.

Innovation Solution

Incorporating unencrypted multi-frequency power level and timebase information in audio packets allows the audio mixer to identify the loudest speakers without decryption, discarding other streams and mixing only the loudest, thereby reducing processing load and noise.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If the audio mixer decrypts and decodes all SRTP packets to determine power levels, then accurate speaker identification is achieved, but processing capacity and bandwidth are overwhelmed

Engineering Contradiction:
Improvespeaker identification accuracyVSAvoidprocessing capacity
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The system performs preliminary action by including unencrypted power level information in the packet headers before decryption and decoding. This allows the audio mixer to identify the N loudest speakers based on pre-computed power levels, avoiding the need to decrypt and decode all packets to determine speaker volume. The power level calculation is performed in advance by the endpoint devices and embedded in the packet metadata.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The packet structure is segmented into encrypted payload portions and unencrypted header portions containing power level information. This segmentation allows the audio mixer to access power level data without processing the encrypted audio content, separating the identification function from the decryption function and reducing overall processing requirements.

Inventive Principle:
Principle #1Segmentation

2Reliability

If all audio streams from multiple participants are processed, then complete audio mixing is achieved, but computational intensity increases significantly

Engineering Contradiction:
Improveaudio mixing completenessVSAvoidcomputational intensity
Core Design Contradiction:
ReliabilityVSPower

Solution Approach 1:

The system applies partial action by processing only the N loudest speaker streams for full decryption and mixing, rather than processing all participant streams. The unencrypted power level information enables the audio mixer to select which streams require full processing, performing computations on a subset of data that is sufficient to achieve the mixing objective while reducing computational intensity.

Inventive Principle:
Principle #16Partial or excessive action

Solution Approach 2:

The power level information is extracted from the audio data and placed in unencrypted packet headers, allowing the system to separate the identification of important streams from the processing of all streams. This extraction enables selective processing where only the most relevant audio streams (N loudest speakers) undergo full decryption and mixing operations.

Inventive Principle:
Principle #2Taking out (Extraction)

3Productivity

If unencrypted power level information is included in packets, then processing load is reduced, but security is compromised

Engineering Contradiction:
Improveprocessing efficiencyVSAvoidsecurity
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The packet structure implements local quality by encrypting only the critical audio payload while leaving the power level information in the header unencrypted. This allows the system to optimize processing efficiency for metadata while maintaining security for the actual audio content. The unencrypted power level data is insufficient to reconstruct or infer the encrypted audio payload, providing localized security where it is most needed.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The unencrypted power level information acts as an intermediary that enables the audio mixer to make processing decisions without accessing the encrypted audio content. This intermediary data structure allows the system to bridge the gap between security requirements (encrypted payloads) and processing efficiency requirements (unencrypted metadata for quick analysis).

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentEP2100397B1Audio conferencing utilizing packets with unencrypted power level information
Publication Date: 2013.07.03 CISCO TECHNOLOGY INC
  • EP2100397B1 patent drawingFigure 1
  • EP2100397B1 patent drawingFigure 2
  • EP2100397B1 patent drawingFigure 3

AI summary

In one embodiment, a method that includes receiving a plurality of packet streams input from different endpoints, packets of each stream including encrypted and unencrypted portions, the unencrypted portion containing audio power level information. The audio power level information contained in the packets of each of the packet streams is then compared to select N packet streams with loudest audio. The N packet streams are then decrypted to obtain audio content, and the audio content of the N packet streams mixed to produce one or more output packet streams. It is emphasized that this abstract is provided to comply with the rules requiring an abstract that will allow a searcher or other reader to quickly ascertain the subject matter of the technical disclosure.