Multi-Modal Road Section Recognition Under Irregular Noise
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing methods for identifying road sections during vehicle testing are inaccurate due to irregular noise and interference from abnormal sounds, particularly when using neural network architectures and sound frequency analysis.
Innovation Solution
A multi-modal cognitive mechanism that combines acoustic spectrum density distribution maps, spectrogram image processing, and machine learning models to identify road section switching points, utilizing preprocessing filters, image morphology, and long short-term memory (LSTM) models for enhanced accuracy.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of manufacture
If pre-applied rules based on sound loudness and frequency are used to identify road sections, then the method is simple to implement, but the accuracy deteriorates when there are many road shapes or abnormal sounds cause interference
Solution Approach 1:
The patent divides the road section identification task into three distinct modes: acoustic spectrum density distribution map analysis, spectrogram image processing, and machine learning prediction. Each mode processes the audio signal independently through different methodologies, and their results are combined to achieve high-accuracy road section identification that overcomes the limitations of any single method
Solution Approach 2:
The patent creates a composite recognition system that integrates three different recognition approaches (acoustic spectrum analysis, spectrogram image processing, and machine learning prediction). This composite mechanism combines the strengths of each individual method while compensating for their respective weaknesses, achieving both high accuracy and robustness against abnormal sounds
2Adaptability or versatility
If a neural network architecture is used to identify road sections, then the system can handle complex patterns, but the accuracy deteriorates due to irregular noise in the drive test process
Solution Approach 1:
The patent segments the neural network-based recognition into a dedicated machine learning mode that specifically predicts expected sounds at each frame and calculates similarity with actual sounds. This segmented approach allows the neural network to focus on learning temporal patterns while being combined with other modes that handle different aspects of sound analysis, improving overall accuracy despite irregular noise
Solution Approach 2:
The patent introduces an intermediary similarity calculation mechanism that compares predicted sounds from the neural network with actual sounds. This intermediary step acts as a filter, only accepting neural network predictions that closely match actual audio, thereby improving accuracy by rejecting predictions corrupted by irregular noise
3Device complexity
If single-mode recognition methods are used, then the system complexity is low, but the reliability deteriorates due to interference from abnormal sounds and irregular noise
Solution Approach 1:
The patent merges three separate recognition modes into a unified multi-modal cognitive mechanism. Each mode processes audio signals through different methods (acoustic spectrum density, spectrogram image processing, and machine learning prediction), and their results are combined to produce the final road section identification. This merging approach significantly improves reliability by ensuring that abnormal sounds and irregular noise must simultaneously interfere with all three independent methods to cause misidentification
Solution Approach 2:
The patent implements a feedback mechanism where the results from three different recognition modes are combined and cross-validated. The system uses the consensus or majority voting from multiple independent analyses to determine the final road section identification, providing feedback that enhances reliability and reduces the impact of interference from abnormal sounds
Data Source
AI summary
In an approach for road section recognition using multi-modal cognitive mechanism, a processor receives an audio signal from a road test. A processor processes the audio signal to generate an acoustic spectrum density distribution map to identify a respective at least one road section switching point in a first mode. A processor processes a spectrogram of the audio signal to identify the respective at least one road section switching point in a second mode. A processor uses a machine learning model to predict an expected sound at each frame of the audio signal, to calculate a similarity between the expected sound and an actual sound, and to identify the respective at least one road switching point when the similarity is lower than a pre-set similarity threshold in a third mode. A processor combines results of the three modes to obtain a final set of road section switching points.


