Audio Video Translation for Multiple Listeners

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current audio-video content systems do not allow multiple viewers to listen to audio in their preferred languages simultaneously, limiting personal experiences in multi-cultural settings like airports or movie theaters.

Innovation Solution

A system where a display device sends audio in different languages to connected devices using machine learning to recognize listeners and translate audio on the fly, using neural networks and near-field communication to correlate listener characteristics with preferred languages.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Device complexity

If a single audio source is used for multiple listeners, then device complexity is reduced, but adaptability to different language preferences deteriorates

Engineering Contradiction:
Improveaudio system complexityVSAvoidlanguage preference adaptability
Core Design Contradiction:
Device complexityVSAdaptability or versatility

Solution Approach 1:

The audio output is segmented and routed to different recipient devices (headphones, smartglasses, speakers) associated with different listeners. Each recipient device receives audio in the listener's preferred language, achieved by separating the audio stream distribution according to device identity and language preference stored in memory.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The display device is designed to serve multiple functions: it can present video to a group while simultaneously providing audio in different languages to different individual listeners through their respective recipient devices. The system universally handles multiple language outputs from a single audio source.

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Adaptability or versatility

If audio is translated in real-time for each listener, then adaptability to language preferences is improved, but processing time increases

Engineering Contradiction:
Improvelanguage preference adaptabilityVSAvoidaudio processing time
Core Design Contradiction:
Adaptability or versatilityVSLoss of time

Solution Approach 1:

Language preferences for different recipient devices are determined and stored in advance in memory before audio playback begins. The system pre-associates each device identifier with its preferred language, eliminating the need for real-time language determination during audio playback.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

Instead of translating audio in real-time, the system creates separate audio copies in different languages and routes them to appropriate recipient devices. Pre-translated audio tracks are stored and distributed according to device preference, avoiding real-time translation processing delays.

Inventive Principle:
Principle #26Copying

3Adaptability or versatility

If multiple audio sources are used for different languages, then adaptability to language preferences is improved, but device complexity increases

Engineering Contradiction:
Improvelanguage support capabilityVSAvoidaudio source configuration
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

Multiple audio sources for different languages are merged into a single display device that can output to multiple recipient devices. The display device consolidates the functionality of multiple audio sources while maintaining the ability to provide different language audio to different listeners through their respective devices.

Inventive Principle:
Principle #5Merging (Combining)

Data Source

PatentUS11443737B2Audio video translation into multiple languages for respective listeners
Publication Date: 2022.09.13 SONY GROUP CORP
  • US11443737B2 patent drawing
  • US11443737B2 patent drawing
  • US11443737B2 patent drawing

AI summary

An audio source such as a display device configured to present AV content can present the video and send the audio in different languages to the respective devices of different listeners. For example, a device/TV/source can send audio in different languages to connected headphones/smartglasses with speakers/devices/sink. Furthermore, machine learning may be employed both to recognize listeners and correlate them to likely languages and to mimic voices in the played-back audio. Or, the source AV display device may send language in only the selected language of the display device to each listener device, with each receiving listener device converting the audio to the preferred language of the respective listener on the fly.