Spatial Audio Mixing for Teleoperator Event Localization

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Teleoperators of autonomous vehicle fleets face challenges in accurately localizing events in real-time due to the inability to combine and process audio data from multiple sensors effectively, leading to delayed or inaccurate decision-making during remote operations.

Innovation Solution

An audio data processing system that combines multiple audio channels from various sensors into a reduced number of channels for output via speakers, providing a spatialized, immersive sound experience, allowing teleoperators to perceive events as if they were present in the vehicle's environment, even when other sensors are inoperable or occluded.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If multiple audio channels from various sensors are combined and processed into a reduced number of channels, then the teleoperator's ability to localize events and perceive the environment immersively is improved, but the device complexity and processing requirements increase

Engineering Contradiction:
Improveevent localization accuracyVSAvoidaudio processing system complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent combines multiple audio channels from various sensors (microphones, audio processors) into a reduced number of output channels. This merging process integrates spatial audio data to create an immersive sound field that enables accurate event localization, directly resolving the contradiction by achieving high measurement precision through systematic combination of multiple input sources.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent transforms two-dimensional audio channel data into a three-dimensional immersive sound field experienced by the teleoperator. By creating spatial audio representations that simulate real-world acoustic environments, the system enhances event localization accuracy by adding spatial dimensionality to the audio processing, allowing operators to perceive direction and distance of events.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Productivity

If audio data is processed in real-time to provide immersive feedback, then the teleoperator's decision-making speed is improved, but the processing time and computational load increase

Engineering Contradiction:
Improvedecision-making speedVSAvoidaudio processing time
Core Design Contradiction:
ProductivityVSLoss of time

Solution Approach 1:

The patent performs preliminary processing of audio data by combining multiple channels into reduced output channels before transmission to the teleoperator. This pre-processing approach prepares the immersive audio feedback in advance, reducing the computational burden during critical decision-making moments and enabling faster response times without sacrificing processing quality.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent extracts essential spatial audio information from multiple sensor channels and consolidates it into a reduced set of output channels. By taking out only the critical spatial characteristics needed for event localization and immersive perception, the system achieves real-time processing capability while maintaining high decision-making speed, avoiding unnecessary computational overhead.

Inventive Principle:
Principle #2Taking out (Extraction)

3Device complexity

If individual audio channels are listened to separately, then the processing complexity is reduced, but the ability to localize events accurately is degraded

Engineering Contradiction:
Improveaudio processing simplicityVSAvoidevent localization accuracy
Core Design Contradiction:
Device complexityVSMeasurement precision

Solution Approach 1:

The patent merges individual audio channels into integrated spatial audio outputs that preserve directional and positional information. By combining channels rather than processing them separately, the system maintains event localization accuracy while presenting the information in a unified immersive format that the teleoperator can interpret intuitively without managing multiple separate audio streams.

Inventive Principle:
Principle #5Merging (Combining)

Data Source

PatentUS11480961B1Immersive sound for teleoperators
Publication Date: 2022.10.25 ZOOX INC
  • US11480961B1 patent drawing
  • US11480961B1 patent drawing
  • US11480961B1 patent drawing

AI summary

Immersive experiences for users are described herein. In an example, audio data from a plurality of audio sensors associated with a vehicle can be received by an audio data processing system. The audio data processing system can combine individual captured audio channels (e.g., from the plurality of audio sensors) into two or more audio channels for output via two or more speakers proximate a user. A first audio channel of the two or more audio channels can be output via a first speaker and second audio channel of the two or more audio channels to be output via a second speaker, wherein output of the first audio channel and the second audio channel causes a resulting sound corresponding to at least a portion of a sound scene associated with the vehicle. In an example, a user computing device operable by the user can receive an input from the user.