Vehicle Sound Separation for Privacy-Respecting Audio Localization

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Autonomous vehicles face challenges in accurately detecting and identifying important sounds in complex driving environments while preserving privacy, as existing sound recognition technologies struggle with multiple sound sources and legal requirements to protect personal conversations.

Innovation Solution

Implementing a sound separation model that separates audio data into elemental sounds, redacts private speech, and processes relevant sounds for vehicle navigation, using a combination of on-board microphones and machine learning models to ensure privacy and accuracy.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If sound separation model is implemented to separate audio data into elemental sounds, then detection precision of important sounds is improved, but device complexity increases

Engineering Contradiction:
Improvedetection precisionVSAvoiddevice complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The audio data is segmented into multiple elemental sounds using a sound separation model. The model decomposes mixed audio signals into distinct sound sources (e.g., sirens, speech, traffic noise), allowing precise identification and detection of important sounds while filtering out irrelevant background noise.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

A sound separation model acts as an intermediary processing layer between the audio sensors and the detection system. This intermediary component separates and categorizes different sound elements before they reach the final detection stage, improving precision while managing complexity through modular architecture.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Loss of information

If all audio data is processed for sound detection, then detection completeness is improved, but loss of time increases due to processing large volumes of data

Engineering Contradiction:
Improvedetection completenessVSAvoidprocessing time
Core Design Contradiction:
Loss of informationVSLoss of time

Solution Approach 1:

The system extracts only the relevant elemental sounds from the complete audio data stream. By using sound separation to identify and extract specific sound categories (such as emergency vehicle sirens or important environmental sounds), the system maintains detection completeness for critical events while reducing processing time by ignoring irrelevant audio content.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

Instead of processing all audio data with equal depth, the system applies partial processing - performing detailed analysis only on extracted relevant sound elements while using lighter processing for other audio components. This selective approach maintains detection completeness for important sounds while minimizing overall processing time.

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS12479466B2Privacy-respecting detection and localization of sounds in autonomous driving applications
Publication Date: 2025.11.25 WAYMO LLC
  • US12479466B2 patent drawing
  • US12479466B2 patent drawing
  • US12479466B2 patent drawing

AI summary

The described aspects and implementations enable privacy-respecting detection, separation, and localization of sounds in vehicle environments. The techniques include obtaining, using audio detector(s) of a vehicle, a sound recording that includes a plurality of elemental sounds (ESs) in a driving environment of the vehicle, and processing, using a sound separation model, the sound recording to separate individual ESs of the plurality of ESs. The techniques further include identifying a content of individual ESs and causing a driving path of the vehicle to be modified in view of the identified content of the individual ESs. Further techniques include rendering speech imperceptibly by redacting temporal portions of the speech, using sound recognition models to identify and discard recordings of speech, and driving at speeds that exceed threshold speeds at which speech becomes imperceptible from noise masking.