Spatial Audio Rendering for Simultaneous Speech Translation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional mediated reality systems face challenges in optimally rendering audio in virtual sound spaces, particularly in mono voice calls with speech-to-speech translation, leading to a sequential voice experience that doubles the time required for communication and disrupts the natural flow of discussion.

Innovation Solution

Simultaneously render first and second audio content as virtual sound objects with differing spatial positions, leveraging the 'cocktail party effect' to allow users to distinguish and focus on one audio source, enhancing the listening experience by controlling spatial rendering to avoid overlapping voices.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If speech-to-speech translation is provided in a mono voice call, then users can understand each other's language, but the communication time is doubled and the natural flow of discussion is interrupted

Engineering Contradiction:
Improvelanguage understandingVSAvoidcommunication time
Core Design Contradiction:
Adaptability or versatilityVSLoss of time

Solution Approach 1:

The audio signal is segmented into original speech and translated speech as separate virtual sound objects, allowing them to be rendered simultaneously in different spatial positions rather than sequentially, thus reducing communication time while maintaining language understanding

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent transitions from a mono (single channel) voice call to a spatial audio system with multiple virtual sound objects positioned in different spatial locations, adding a spatial dimension that allows simultaneous playback of original and translated speech without interference

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Adaptability or versatility

If original speech and translated speech are mixed in a mono channel, then both languages are available, but the speeches become impossible to understand due to the interfering talker problem

Engineering Contradiction:
Improvelanguage availabilityVSAvoidspeech intelligibility
Core Design Contradiction:
Adaptability or versatilityVSReliability

Solution Approach 1:

The speech signals are segmented into separate virtual sound objects that can be spatially separated, preventing the interfering talker problem by allowing users to focus on one speech at a time through spatial cues

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

Spatial rendering acts as an intermediary mechanism that separates the original and translated speech in the virtual sound space, allowing both to coexist without interference by positioning them at different virtual locations

Inventive Principle:
Principle #24Intermediary (Mediator)

3Adaptability or versatility

If sequential voice experience is used with translation, then language translation is provided, but the user experience is frustrating and discussion flow is constantly interrupted

Engineering Contradiction:
Improvetranslation supportVSAvoiduser experience
Core Design Contradiction:
Adaptability or versatilityVSEase of operation

Solution Approach 1:

The patent enables continuous simultaneous playback of original and translated speech through spatial rendering, eliminating the sequential interruptions that frustrate users and maintaining the natural flow of discussion

Inventive Principle:
Principle #20Continuity of useful action

Data Source

PatentEP3720149B1An apparatus, method, computer program or system for rendering audio data
Publication Date: 2025.11.19 NOKIA TECHNOLOGIES OY
  • EP3720149B1 patent drawingFigure 1A~4
  • EP3720149B1 patent drawingFigure 5~6
  • EP3720149B1 patent drawingFigure 7

AI summary

Certain examples of the present invention relate to rendering of audio data. Certain examples provide an apparatus 500 comprising means 501 for causing: receiving 401 first audio data 701 representative of first audio content 101; receiving 402 second audio data 702 representative of second audio content 102, wherein the second audio content 102 is derived from the first audio content 101; rendering 403 the first audio data 701 as a first virtual sound object 601 in a virtual sound scene 600 such that it is spatially rendered with a first virtual position 601o,601l within the virtual sound scene 600; rendering the second audio data 702 as a second virtual sound object 602 in the virtual sound scene 600 such that it is spatially rendered with a second virtual position 602o,602l within the virtual sound scene 600; and controlling the spatial rendering of the first and second virtual sound objects 601, 602 such that the first and second virtual positions 601o,601l;602o,602l differ.