Spatial Audio Codec for Simultaneous Translation Playback

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current audio codec systems face challenges in providing real-time language translation with low latency and natural conversation flow, especially in immersive audio environments like virtual reality, due to sequential translation methods that double call time and disrupt conversational flow.

Innovation Solution

The implementation of secondary audio tracks encoded and decoded using spatial audio techniques, allowing for simultaneous playback of original and translated audio signals, and user-controlled rendering of these tracks to enhance conversation flow and reduce latency.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If sequential translation methods are used, then language translation can be achieved, but call time is doubled and conversational flow is disrupted

Engineering Contradiction:
Improvelanguage translation accuracyVSAvoidcall time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The audio signal is segmented into multiple independent tracks (primary track with original audio, secondary track with translated audio). Each track is processed and transmitted separately, allowing the receiver to play them simultaneously rather than sequentially, thus reducing overall processing time while maintaining translation accuracy.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The solution transitions from sequential time-based processing to parallel spatial processing by introducing multiple audio tracks that can be rendered simultaneously in different spatial positions. This dimensional change from time-sequential to space-parallel processing eliminates the doubling of call time while preserving translation quality.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Ease of operation

If multiple audio tracks are played simultaneously, then conversational flow is improved, but audio rendering complexity increases

Engineering Contradiction:
Improveconversational flowVSAvoidaudio rendering complexity
Core Design Contradiction:
Ease of operationVSDevice complexity

Solution Approach 1:

Spatial metadata acts as an intermediary that simplifies the rendering process. Instead of complex multi-track mixing, the spatial metadata provides pre-calculated position information that guides the renderer to place audio tracks in appropriate spatial positions, reducing rendering complexity while maintaining natural conversational flow.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The system changes the parameter representation from complex time-synchronized multi-track control to simpler spatial position parameters. By encoding track identification and spatial position as metadata parameters, the system reduces rendering complexity while enabling simultaneous playback that improves conversational flow.

Inventive Principle:
Principle #35Parameter changes

3Adaptability or versatility

If spatial audio decoding is used, then immersive audio experience is enhanced, but processing requirements increase

Engineering Contradiction:
Improveimmersive audio capabilityVSAvoidprocessing requirements
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

Spatial analysis and metadata generation are performed in advance during the encoding phase. This preliminary action prepares the spatial information before transmission, reducing the processing burden on the receiving device while maintaining immersive audio capabilities through efficient spatial audio decoding.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS12067992B2Audio codec extension
Publication Date: 2024.08.20 NOKIA TECHNOLOGIES OY
  • US12067992B2 patent drawing
  • US12067992B2 patent drawing
  • US12067992B2 patent drawing

AI summary

An apparatus comprising means configured to: receive a primary track comprising at least one audio signal; receive at least one secondary track, each of the at least one secondary track comprising at least one audio signal, wherein the at least one secondary track is based on the primary track; and decode and render the primary track and the at least one secondary track using spatial audio decoding.