Electric Vehicle Audio Classification Using Video Labels

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

The challenge of accurately identifying electric vehicles using deep-learning systems is hindered by the scarcity and complexity of labeled audio data, which is crucial for training and evaluating such systems, particularly in real-world applications where noise interference complicates the task.

Innovation Solution

A synchronized camera and microphone array setup is used to collect and label vehicle audio and video data, enabling the training of a neural network to differentiate between electric and non-electric vehicles based on audio patterns, with techniques like beam forming and noise isolation to enhance accuracy.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If deep-learning systems are trained to identify electric vehicles using audio data, then vehicle classification accuracy is improved, but the complexity of data labeling and noise interference handling worsens

Engineering Contradiction:
Improvevehicle classification accuracyVSAvoiddata labeling complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent uses video data as an intermediary to automatically generate audio labels. A video classification model identifies vehicle types (electric vs. non-electric) from visual features, and these classifications are transferred to corresponding audio segments as labels. This mediator approach eliminates the need for manual audio labeling while improving classification accuracy.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent replaces manual audio data labeling (mechanical process) with automated video-based classification systems. The video classification model automatically generates labels for audio data, substituting the labor-intensive manual annotation process with an automated computational approach.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Adaptability or versatility

If audio data is collected from real-world roadways to train neural networks, then the variety and realism of training data is improved, but noise interference and data quality worsen

Engineering Contradiction:
Improvetraining data varietyVSAvoidnoise interference
Core Design Contradiction:
Adaptability or versatilityVSObject-affected harmful factors

Solution Approach 1:

The patent converts noise interference (harmful factor) into a feature for distinction. By training the model with both noisy real-world audio and the corresponding video-based ground truth labels, the system learns to identify and utilize noise patterns that differentiate electric vehicles from non-electric vehicles in real-world conditions.

Inventive Principle:
Principle #22Blessing in disguise (Convert harm into benefit)

Solution Approach 2:

Video data serves as an intermediary that provides clean ground truth labels despite noisy audio conditions. The video classification model accurately identifies vehicle types regardless of audio noise, and these reliable labels are used to train the audio classification model to be robust against noise interference.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Measurement precision

If manual labeling of audio data is performed to train deep-learning models, then label accuracy is improved, but the time and resources required worsen

Engineering Contradiction:
Improvelabel accuracyVSAvoiddata preparation time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent performs preliminary video classification to generate audio labels before training the audio model. By pre-classifying vehicles using video data and transferring these labels to audio segments, the system prepares accurate labels automatically without requiring subsequent manual annotation, saving significant time and resources.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

Video classification acts as an intermediary process that automatically generates accurate audio labels. This mediator approach replaces manual labeling while maintaining high label accuracy, as the video-based classification provides reliable ground truth that can be directly applied to audio data.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS20250218187A1Methods and systems for classifying vehicles as electric or nonelectric based on audio
Publication Date: 2025.07.03 ROBERT BOSCH GMBH
  • US20250218187A1 patent drawing
  • US20250218187A1 patent drawing
  • US20250218187A1 patent drawing

AI summary

Methods and systems for training a neural network to identify an electric vehicle based on audio. Video data is generated from a camera with a field of view including a roadway. Audio data is generated from a microphone, the audio data associated with vehicles traveling across the roadway. The video data is segmented into segments, each having a start time and a finish time that corresponds to a respective vehicle traveling across the roadway in and out of the field of view. Each video segment is labeled with a label indicating the respective vehicle in that segment as either an electric vehicle or a non-electric vehicle. The audio data is segmented into segments, each having a start time and end time associated with a respective one of the video segments. A neural network is trained based on the audio segments and the labels of the associated video segments.