Signal Data Augmentation Using Phase-Randomized Surrogates
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current data augmentation techniques for signal classification are either too domain-specific, requiring extensive knowledge and time, or too generic, failing to generate new realizations of stochastic processes, leading to inefficient and resource-intensive dataset expansion.
Innovation Solution
A data-driven, agnostic data augmentation method that applies random noise to the phase spectrum of non-relevant frequency bands identified through discrete Fourier transform, preserving the most relevant frequency components of the signal, thus generating new signals with statistical similarity to the original class.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If domain-specific data augmentation techniques are used, then the generation of realistic signal variations is improved, but the complexity and time required for implementation increases
Solution Approach 1:
The patent uses surrogate signals as copies or replicas of the original signal, created by replacing the phase spectrum with random values while preserving the amplitude spectrum. This copying approach generates realistic signal variations without requiring complex domain-specific transformations, as the surrogate signals inherently capture the statistical properties of the original signal class.
Solution Approach 2:
The patent modifies the phase spectrum parameter of the signal while keeping the amplitude spectrum unchanged. This parameter change approach allows generation of diverse signal variations by randomizing only the phase component, which is sufficient to create realistic variations without needing to model complex domain-specific transformations.
2Productivity
If generic data augmentation techniques are used, then the implementation time is reduced, but the effectiveness in generating meaningful variability decreases
Solution Approach 1:
The surrogate signal copying method is computationally efficient as it only requires Fourier transform operations and random phase generation, avoiding complex domain-specific transformations. This maintains high productivity while the statistical properties of the amplitude spectrum ensure the generated variability is meaningful and representative of the original signal class.
Solution Approach 2:
By changing only the phase spectrum parameter randomly, the method achieves fast computation without complex transformations. The amplitude spectrum is preserved to maintain statistical similarity, ensuring the generated variability is effective for improving model generalization.
3Device complexity
If simple noise addition is used, then the computational complexity is minimized, but the ability to mimic real-world signal variations is insufficient
Solution Approach 1:
Instead of simple noise addition, the patent creates surrogate signals by copying the amplitude spectrum and replacing only the phase spectrum with random values. This approach is computationally simple yet effectively mimics real-world signal variations by preserving the statistical structure while introducing realistic variability through phase randomization.
4Reliability
If extensive data augmentation processes are applied, then the dataset variability is increased, but the time and resources required increase significantly
Solution Approach 1:
The surrogate signal method generates multiple varied copies from a single original signal through simple phase randomization, achieving high dataset variability without extensive processing. Each surrogate is created quickly using Fast Fourier Transform operations, making the process scalable to generate large numbers of augmented samples efficiently.
Solution Approach 2:
The surrogate data augmentation method is universally applicable to any signal type without requiring domain-specific parameters or transformations. This multi-functionality allows the same simple process to be applied across different signal domains, reducing the need for extensive domain-specific augmentation pipelines and reducing overall processing time.
Data Source
AI summary
A computer implemented data augmentation method comprising receiving a dataset to be processed and, upon the received dataset being unclassified into classes, performing a clustering algorithm to partition the dataset whereby clusters formed are interpreted as the signal classes. The method further includes forming a sample dataset by gathering, for each class of a plurality of classes, at least two sample signals then applying a discrete Fourier transform (DFT) to each sample signal of the sample dataset. The method includes computing frequency parameters of each sample signal to determine, based on a spectral coherence threshold, frequency bands: relevant bands that characterizes a class. The method further includes injecting random noise in a phase spectrum of the non-relevant frequency bands of each sample signal of the sample dataset, to generate a set of augmented sample signals, and applying an inverse DFT, in each of the generated augmented sample signals.


