Local Media Rendering for Multi-Party Calls

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional central rendering in multi-party calls, such as voice conferences, faces challenges with high processing capacity and latency due to the need for unique 3D positional audio rendering for each participant, especially in large and dynamic conferences, and relies on costly central resources.

Innovation Solution

A method where a Media Server determines and delivers a maximum number of media streams to Client User Equipment for local rendering, with clients requesting and prioritizing media streams based on geographical properties and signal strength, allowing for separate transmission of rendering information and media streams, thereby reducing processing load on the central device.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If central rendering is used to provide unique 3D positional audio streams for each participant, then true stereo or 3D positional audio can be delivered to each client, but the processing capacity requirement for the central voice mixing device becomes very large

Engineering Contradiction:
Improve3D positional audio rendering capabilityVSAvoidprocessing capacity
Core Design Contradiction:
Adaptability or versatilityVSPower

Solution Approach 1:

The patent extracts the 3D positional audio rendering function from the central server and transfers it to the client devices. The central server only performs basic functions like receiving audio streams and transmitting positional information, while clients perform the computationally intensive rendering locally using the extracted audio data and positional information to generate personalized 3D audio streams.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

Each client device performs self-service by rendering its own personalized 3D positional audio stream locally. Instead of relying on the central server to provide rendered audio for each participant, each client independently processes the audio streams and applies the appropriate spatial audio effects based on its own position and orientation, thereby distributing the processing load.

Inventive Principle:
Principle #25Self-service

2Adaptability or versatility

If central rendering creates unique 3D positional audio signals for each participant, then true stereo or 3D positional audio can be provided, but the system architecture becomes complicated and requires very large processing capacity

Engineering Contradiction:
Improveunique media stream per clientVSAvoidsystem architecture
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent extracts the complex rendering operations from the central server architecture and relocates them to client devices. The central server architecture is simplified to primarily handle media stream reception and transmission, while clients handle the complex spatial audio processing locally, thereby reducing overall system complexity.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent segments the audio processing function between central server and client devices. The central server handles basic media streaming functions, while clients handle 3D positional audio rendering. This segmentation distributes complexity across multiple devices rather than concentrating it at the central server.

Inventive Principle:
Principle #1Segmentation

3Productivity

If central rendering is used for highly interactive applications, then media streams can be processed centrally, but latency in positional information makes faithful voice rendering impossible

Engineering Contradiction:
Improvemedia stream processingVSAvoidlatency
Core Design Contradiction:
ProductivityVSLoss of time

Solution Approach 1:

The patent extracts the time-sensitive audio rendering operations from the central server and performs them locally at client devices. This extraction eliminates the transmission delay and processing latency associated with central rendering, enabling real-time audio processing with minimal latency that is crucial for faithful voice rendering in highly interactive applications.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

Instead of having the central server process and render audio streams for each client, the patent inverts the approach by having clients request and process audio streams locally. This inversion of the processing location fundamentally reduces latency by eliminating the round-trip time between server and client for rendering operations.

Inventive Principle:
Principle #13The other way round (Inversion)

Data Source

PatentEP2661857B1Local media rendering
Publication Date: 2016.06.01 TELEFONAKTIEBOLAGET LM ERICSSON (PUBL)
  • EP2661857B1 patent drawingFigure 1
  • EP2661857B1 patent drawingFigure 2
  • EP2661857B1 patent drawingFigure 3

AI summary

The invention involves local media rendering of a multi -party call, performed by a Client User Equipment (1). The media is encoded by each party in the call, and sent as a media stream to a Media server (2), and the media server receives a request for media streams from each Client User Equipment, each media stream in the request associated with a client priority. The Media server selects the media streams to send to each Client User Equipment, based on the request, and further such that the number of streams does not exceed a determined maximum number, which is based e.g. on the available bandwidth.