A data enhancement method and system for audio signals
By systematically simulating various situations of audio signals in the actual environment, a comprehensive data enhancement method is adopted to solve the problems of overfitting, data finiteness, complex environmental impact and insufficient generalization capabilities of deep learning models when processing audio signals, and improve the generalization capabilities and accuracy of the model.
Patent Information
- Application Number
- CN202410960729.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-07-17
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2044-07-17
AI Technical Summary
Deep learning models face problems such as overfitting, data finiteness, complex environmental impact and insufficient generalization capabilities when processing audio signals, resulting in poor performance in practical applications.
A comprehensive audio signal data enhancement method is adopted to systematically simulate various situations encountered by audio signals in the actual environment through six steps, including adding random noise, simulated energy attenuation, transformed frequency response, simulated multipath interference, performing short-time Fourier transform and simulated Doppler frequency bias.
Through these steps, the generated data is closer to the complex environment of the real world, improving the diversity and robustness of the input dataset of the model, thereby improving the generalization ability and accuracy of the model.
Smart Images

Figure CN118782069B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to but is not limited to the field of data enhancement technology, and in particular relates to a data enhancement method and system for audio signals. Background Art
[0002] Deep learning algorithms have demonstrated excellent performance in a variety of computer vision tasks, and powerful machine learning systems work best when trained on large amounts of data. However, limited labeled data can lead to overfitting problems, which hinder the network's performance on unseen data. Overfitting means that the model performs well on the training data but poorly on the test data. To address this problem, various generalization techniques have been proposed, including element culling, normalization, and data augmentation. However, overfitting remains a huge challenge, and the standard training process can only learn important features but cannot learn less important features required for generalization. In this case, small, invisible perturbations added to the input image can deceive the network, causing it to fail to recognize the correct features in the image.
[0003] Deep learning algorithms have demonstrated excellent performance in various computer vision tasks, especially in the presence of large amounts of labeled data, and are able to achieve highly accurate predictions. However, when labeled data is limited, the model is prone to overfitting, which means that it performs well on the training data but poorly on the test data. Overfitting occurs when the model remembers the details and noise of the training data during training and fails to learn the features required for generalization. This makes it difficult for the model to make accurate predictions on unseen data, thus hindering its practical application.
[0004] In order to solve the overfitting problem, various generalization techniques have been proposed, including element elimination, normalization, and data augmentation. Element elimination reduces the complexity of the model by randomly ignoring some input units, normalization reduces the training time and complexity of the model by standardizing the data distribution, and data augmentation increases the size and diversity of the data set by generating diverse training samples. However, these methods still face huge challenges in practical applications, especially in complex and changing environments.
[0005] In the complex and ever-changing indoor environment, the collection of sound signals is affected by many factors. Physical hardware such as speakers and microphones will cause differences in the frequency response of sound signals, energy attenuation will occur when sound propagates in the medium, inevitable multipath interference and superposition of environmental noise, and Doppler frequency deviation caused by dynamic collection, all of which make the characteristics of sound signals more complex and difficult to capture. In order to enable deep learning models to have good generalization capabilities in different environments, different distances, different movement speeds, and different terminals, data enhancement for audio signals is particularly important.
[0006] However, the actual collected data cannot cover all situations, which further exacerbates the difficulty of generalization of deep learning models. Although existing technical means can alleviate the overfitting problem to a certain extent, they still cannot completely solve the generalization challenges encountered by deep learning models in practical applications. Small, invisible perturbations added to the input image can deceive the network, causing it to be unable to recognize the correct features in the image, indicating that the existing technology still needs to be improved in terms of the robustness and generalization ability of processing input data.
[0007] In summary, the technical problems of existing technologies in deep learning algorithms mainly include:
[0008] 1) Overfitting problem: The model performs well on the training data, but performs poorly on the test data and is difficult to generalize to unseen data.
[0009] 2) Data limitation: The labeled data is limited, and the data actually collected cannot cover all situations, making it difficult to provide enough training samples.
[0010] 3) Influence of complex environment: In a complex and changing environment, sound signals are affected by many factors, which increases the difficulty of signal processing.
[0011] 4) Insufficient generalization ability: Existing technical means are difficult to comprehensively improve the generalization ability of the model and cannot cope with changes in different environments and different terminals.
[0012] In order to improve the generalization ability of deep learning models, it is urgent to solve the above technical problems and achieve more stable and reliable performance through innovative data augmentation methods and more robust model training techniques. Summary of the invention
[0013] In view of the problems existing in the prior art, the present invention provides a data enhancement method and system for an audio signal.
[0014] The present invention is implemented as follows: a data enhancement method for an audio signal, characterized in that the data enhancement method for an audio signal specifically comprises:
[0015] S1: normalize the input time domain signal data, generate a set of normalized random noise, and generate new signal data with a specified signal-to-noise ratio according to the corresponding attenuation parameters and noise coefficient;
[0016] S2: Generate a random suppression curve in advance, whose length is consistent with the length of the data to be enhanced, and transform different frequency responses through the random suppression curve;
[0017] S3: Delay and superimpose the signals to simulate multipath interference;
[0018] S4: The collected time domain signal is passed through a bandpass filter to reduce the environmental noise in the non-target frequency band, and the filtered time domain signal is converted into a time-frequency image through short-time Fourier transform (STFT) to obtain a time-frequency diagram of the signal;
[0019] S5: offset the time-frequency diagram according to the random frequency offset, complete the Doppler frequency offset simulation, and then cut out the target frequency band image;
[0020] S6: Normalize the image to prevent the value from affecting the model training.
[0021] Furthermore, the S1 simulates the time domain signal s after the environmental noise is simulated. N (t) can be expressed as:
[0022] s N (t)=s(t)+Noise(t)
[0023] Where Noise(t) is a random value that changes with time, and s(t) represents the original time domain signal. The present invention proposes to set a random attenuation coefficient α in the collected time domain signal to simulate the time domain signal s after energy attenuation. E (t) can be expressed as:
[0024] s E (t) = α·s(t)
[0025] While simulating the different degrees of energy attenuation of the signal propagating in the medium, random additive noise is combined to simulate signals with different signal-to-noise ratios.
[0026] Furthermore, S2 sets a random suppression curve in the collected time domain signal, and after point multiplication with the collected time domain signal, readjusts the amplitude difference between each frequency point, and generates a random curve by sliding a window and averaging the random value in the random vector to simulate the time domain signal s after the hardware frequency response. R (t) can be expressed as:
[0027] s R (t) = curve(t)·s(t)
[0028] Where curve(t) is the random inhibition curve that changes with time.
[0029] Furthermore, in S3, original signals with different delays and different intensities are superimposed on the collected time domain signals to simulate the echo phenomenon in the actual environment and simulate the time domain signal s after multipath interference. M (t) can be expressed as:
[0030]
[0031] Where is the number of paths of random P machines, λ p is the random intensity coefficient of different paths, τ p is the random delay of different paths.
[0032] Further, the S4, short-time Fourier transform is specifically as follows:
[0033] Assuming that the length of the sliding window is WL and the moving step of the window is SL, the time resolution is (SL / Fs)s, the frequency resolution is (Fs / WL)Hz, and the detailed steps of STFT are as follows:
[0034] First, the sliding window starts from the starting point of the received sound signal data. At this time, the window function is t = τ 0 As the center, the signal is processed by windowing function:
[0035] y(t)=x(t)·w(t-τ 0 )
[0036] Then, Fourier transform is performed on the data to obtain the PSD matrix of the sound data of the first window. Indicates that the received signal is delayed by (0, τ 0 ]:
[0037]
[0038] Where x(t) represents the received signal The sound data segment, w is the Hamming window function, f m Depends on the sampling rate (Frequency of Sampling, Fs) of the smartphone, ranging from 0Hz to Fs / 2Hz, f m and τ 0 The definitions are as follows:
[0039]
[0040] τ 0 =(WL / 2) / Fs
[0041] Finally, the PSD matrix representing the sound data segment of the nth window is calculated as follows:
[0042]
[0043] In the formula, The received signal is delayed by (τ n-1 , τ n ], where x(t n ) and τ n The definition is as follows:
[0044] x(t n )=R[(n-1)×SL:WL+(n-1)×SL]
[0045] τ n =[WL / 2+(n-1)×SL] / Fs
[0046] When performing short-time Fourier transform, the step size of the window sliding determines the time resolution of the input image, and the size of the window determines the frequency resolution of the input image. A window sliding step size of 1 ms is equivalent to a resolution of 0.34 m. In order to facilitate fast Fourier transform, the window size is selected to be 512.
[0047] Another object of the present invention is to provide a data enhancement system for an audio signal, the system specifically comprising:
[0048] A signal preprocessing module, used for normalizing the input time domain signal data;
[0049] A noise generation module, used to generate random noise;
[0050] Signal simulation module, used to simulate different frequency responses and multipaths of signals;
[0051] The filtering module is used to pass the collected time domain signal through a bandpass filter to reduce the environmental noise in non-target frequency bands;
[0052] Fourier transform module, used to convert the filtered time domain signal into a time-frequency image;
[0053] The frequency deviation module is used to shift the time-frequency diagram according to the random frequency deviation to complete the simulation of Doppler frequency deviation;
[0054] Output module, used to output data enhancement results.
[0055] The present invention also provides an audio signal data enhancement system, the system comprising:
[0056] A normalization processing module is used to normalize the input time domain signal data and generate a set of normalized random noises, and generate new signal data with a specified signal-to-noise ratio by combining the attenuation parameter and the noise coefficient;
[0057] A random suppression curve generation module is used to generate a random suppression curve consistent with the length of data to be enhanced, and transform different frequency responses through the curve;
[0058] Multipath interference simulation module, used to delay and superimpose signals to simulate multipath interference;
[0059] The bandpass filter and short-time Fourier transform module is used to pass the collected time domain signal through a bandpass filter to reduce the environmental noise in the non-target frequency band, and convert the filtered time domain signal into a time-frequency image through short-time Fourier transform to obtain the time-frequency diagram of the signal;
[0060] The Doppler frequency shift simulation and image normalization module is used to shift the time-frequency diagram according to the random frequency shift, complete the Doppler frequency shift simulation, crop the target frequency band image, and then normalize the image.
[0061] The normalization processing module further comprises:
[0062] An environmental noise simulation unit, used for simulating environmental noise and generating a time domain signal after simulating environmental noise;
[0063] The energy attenuation simulation unit is used to set a random attenuation coefficient, simulate the energy attenuation of the signal when it propagates in the medium, and generate a time domain signal after the simulated energy attenuation.
[0064] The random inhibition curve generation module further comprises:
[0065] The frequency response simulation unit is used to generate a random suppression curve, perform point multiplication with the collected time domain signal, readjust the amplitude difference between each frequency point, and simulate the time domain signal after the hardware frequency response.
[0066] The bandpass filtering and short-time Fourier transform module further comprises:
[0067] A bandpass filter unit is used to pass the collected time domain signal through a bandpass filter to reduce environmental noise in non-target frequency bands;
[0068] The short-time Fourier transform unit is used to convert the filtered time domain signal into a time-frequency image. Specifically, the signal is processed by a sliding window with a window function, and the data is Fourier transformed to obtain the PSD matrix of the sound data.
[0069] In combination with the above technical solutions and the technical problems solved, the advantages and positive effects of the technical solutions to be protected by the present invention are as follows:
[0070] First, the present invention adds a certain amount of random noise to the collected time domain signal, which can improve the diversity of environmental noise in the data set, enrich the infinitely changing noise in the environment, and enhance the input data set of the model. The present invention proposes to set a random attenuation coefficient α in the collected time domain signal, which can improve the diversity of attenuation in different situations in the data set and enhance the input data set of the model.
[0071] The present invention sets a random suppression curve in the collected time domain signal, and after point multiplication with the collected time domain signal, readjusts the amplitude difference between each frequency point, which can improve the limited frequency response in the data set and enhance the input data set of the model.
[0072] The present invention superimposes original signals with different time delays and different intensities on the collected time domain signals to simulate the echo phenomenon in the actual environment, thereby improving the diversity of multipath interference in the data set and enhancing the input data set of the model.
[0073] The present invention introduces random frequency deviation to simulate the influence of Doppler effect, which can improve the signal diversity of the data set under different relative motion conditions and enhance the input data set of the model.
[0074] Second, in the field of audio signal processing, data enhancement is an important means to improve the generalization ability and robustness of the model. Existing audio data enhancement methods often focus on a single transformation, such as adding noise or changing signal strength, but rarely comprehensively consider the various complex situations encountered by audio signals in actual environments, such as multipath interference, hardware frequency response differences, and Doppler effect. When processing audio signals in actual scenarios, these methods cannot fully simulate the complexity of the real world, thus affecting the performance of the model.
[0075] The present invention proposes a comprehensive audio signal data enhancement method, which systematically simulates various situations encountered by audio signals in actual environments through six steps. First, the adaptability of the signal to environmental noise is enhanced by adding random noise with a specified signal-to-noise ratio and simulating energy attenuation. Secondly, the frequency response differences of different hardware are simulated using random suppression curves. Next, multipath interference is simulated by superimposing signals with different delays and intensities. Then, the signal is converted into a time-frequency diagram using a short-time Fourier transform (STFT), and a frequency deviation simulation is performed to reproduce the Doppler effect. Finally, the image is normalized to ensure the consistency of the data and the effectiveness of the model training.
[0076] The significant technical innovation of the present invention lies in its comprehensiveness and systematicness. It not only takes into account single environmental factors such as noise and signal attenuation, but also deeply simulates the impact of hardware characteristics (such as frequency response differences) and physical phenomena (such as the Doppler effect) on audio signals. Through the application of short-time Fourier transform, the present invention also realizes the detailed analysis and processing of the time-frequency characteristics of audio signals, further improving the effect of data enhancement and the generalization ability of the model.
[0077] The audio signal data enhancement method of the present invention has significant value in practical applications. It can generate training data that is closer to the complex environment of the real world, help machine learning models better adapt to various practical scenarios, and improve the accuracy and robustness of the model in tasks such as audio recognition, speech enhancement, and sound event detection. In addition, the method is also highly flexible and scalable, and can adjust parameters and transformation strategies according to the needs of specific application scenarios, providing strong support for research and application in the field of audio signal processing.
[0078] Third, the present invention proposes a data enhancement method for audio signals, which realizes effective enhancement of audio signals through a series of steps to solve the technical problems in the prior art and achieve significant technical progress. First, the method processes the input time domain signal data by normalization, generates normalized random noise at the same time, generates new signal data with a specified signal-to-noise ratio according to the corresponding attenuation parameters and noise coefficient, simulates different environmental noise and energy attenuation conditions, and thus improves the diversity and robustness of the data.
[0079] Secondly, by generating a random suppression curve in advance, this method can transform different frequency responses and simulate the time domain signal after the hardware frequency response. The random suppression curve is consistent with the length of the data to be enhanced, and the random value jumps are smoothed by a sliding window, so that the amplitude difference between each frequency point is readjusted. This processing method effectively simulates the frequency response characteristics of the actual hardware, making the generated data more realistic and representative.
[0080] Furthermore, by delaying and superimposing the signal, the simulation of multipath interference is realized. Specifically, the original signal with different delays and different intensities is superimposed on the collected time domain signal to simulate the echo phenomenon in the actual environment. This process can generate more complex multipath interference signals, making the data-enhanced signal more consistent with the propagation characteristics in the actual environment, thereby improving the generalization ability and accuracy of the model in complex environments.
[0081] Finally, a bandpass filter is used to reduce the environmental noise in the non-target frequency band, and the short-time Fourier transform (STFT) is used to convert the filtered time domain signal into a time-frequency image. The time-frequency image is further offset according to the random frequency deviation to complete the simulation of the Doppler frequency deviation, and then the target frequency band image is cropped, and finally the image is normalized. This method not only improves the diversity and complexity of signal data, but also ensures the quality and consistency of the generated data through a series of sophisticated processing steps, providing a reliable data foundation for subsequent model training. Through the above series of innovative steps, the present invention significantly improves the effect of audio signal data enhancement and provides a new solution for the fields of signal processing and machine learning.
[0082] Fourth, the technical solution of the present invention solves multiple technical problems and achieves significant technical progress through carefully designed parameters, algorithms and mathematical models. The following is a detailed description of these aspects:
[0083] Attenuation coefficient and noise coefficient: By adjusting these parameters, the present invention can simulate signals with different signal-to-noise ratios, thereby solving the problem of signals being interfered by noise to varying degrees in actual environments.
[0084] The length and shape of the random suppression curve: These parameters are used to simulate the frequency response differences of the hardware, so that the enhanced data can more realistically reflect the impact of different hardware devices on the audio signal.
[0085] Delay and strength coefficient: When simulating multipath interference, by adjusting these parameters, you can simulate the echo and signal attenuation in different environments.
[0086] STFT window size and step size: The choice of these parameters directly affects the resolution and computational efficiency of the time-frequency diagram. By setting them properly, the computational speed can be improved while maintaining high resolution.
[0087] Normalization algorithm: ensures that the amplitude of the input signal is within a uniform range to facilitate subsequent processing and analysis.
[0088] Random noise generation algorithm: Generates random noise that matches the original signal to simulate noise interference in the actual environment.
[0089] Random suppression curve generation algorithm: adjusts the frequency response of the signal by generating a random curve to simulate the characteristics of the hardware device.
[0090] Delay superposition algorithm: Simulates multipath interference and echo phenomena by superimposing signals of different delays and strengths.
[0091] STFT algorithm: Converts time domain signals into time-frequency diagrams to more effectively extract the features of audio signals.
[0092] Signal plus noise model: The original signal is combined with random noise through a mathematical model to generate a new signal-to-noise ratio signal.
[0093] Hardware frequency response model: Use random suppression curves to simulate the frequency response characteristics of hardware devices, making the enhanced data closer to the actual situation.
[0094] Multipath interference model: A mathematical model is used to simulate the multipath interference that signals encounter in actual environments.
[0095] Time-frequency conversion model: Use mathematical tools such as STFT to convert time domain signals into time-frequency diagrams to facilitate subsequent feature extraction and analysis.
[0096] Insufficient data diversity: By introducing random parameters and algorithms, the diversity of the data set is increased, enabling the model to better adapt to various practical scenarios.
[0097] Environmental interference simulation is not realistic: Through carefully designed mathematical models and algorithms, factors such as noise, hardware response and multipath interference in the actual environment are more realistically simulated.
[0098] Difficulty in feature extraction: Use mathematical tools such as STFT to convert time domain signals into time-frequency diagrams to more effectively extract the features of audio signals.
[0099] Improved data enhancement effect: Through the technical solution of the present invention, enhanced data that is closer to the actual environment can be generated, thereby improving the generalization ability of the model.
[0100] Improved model training efficiency: Since the augmented data is more representative, the model training speed can be accelerated and the training effect can be improved.
[0101] Improved model performance: The model trained based on the enhanced dataset of the present invention shows better performance and stability in practical applications.
[0102] In summary, the technical solution of the present invention solves multiple technical problems through carefully designed parameters, algorithms and mathematical models, and has achieved significant technical progress, providing strong support for audio signal processing and machine learning model training. BRIEF DESCRIPTION OF THE DRAWINGS
[0103] Figure 1 is a flow chart of a data enhancement method for an audio signal provided by an embodiment of the present invention;
[0104] Figure 2 It is data enhancement based on random additive noise provided by an embodiment of the present invention;
[0105] Figure 3 The data enhancement based on the random reduction coefficient provided by the embodiment of the present invention;
[0106] Figure 4 The data enhancement based on the random suppression curve provided by the embodiment of the present invention;
[0107] Figure 5 The data enhancement based on random delay superposition provided by the embodiment of the present invention;
[0108] Figure 6 The random frequency shift-based data enhancement provided by the embodiment of the present invention;
[0109] Figure 7 is a module diagram of a data enhancement system for audio signals provided by an embodiment of the present invention; DETAILED DESCRIPTION
[0110] In order to make the purpose, technical solution and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0111] like Figure 1 As shown, an embodiment of the present invention provides a data enhancement method for an audio signal, the method specifically comprising:
[0112] S1: normalize the input time domain signal data, generate a set of normalized random noise, and generate new signal data with a specified signal-to-noise ratio according to the corresponding attenuation parameters and noise coefficient;
[0113] S2: Generate a random suppression curve in advance, whose length is consistent with the length of the data to be enhanced, and transform different frequency responses through the random suppression curve;
[0114] S3: Delay and superimpose the signals to simulate multipath interference;
[0115] S4: The collected time domain signal is passed through a bandpass filter to reduce the environmental noise in the non-target frequency band, and the filtered time domain signal is converted into a time-frequency image through short-time Fourier transform (STFT) to obtain a time-frequency diagram of the signal;
[0116] S5: offset the time-frequency diagram according to the random frequency offset, complete the Doppler frequency offset simulation, and then cut out the target frequency band image;
[0117] S6: Normalize the image to prevent the value from affecting the model training.
[0118] The S1, noise figure is set as follows:
[0119] Adding a certain amount of random noise to the collected time domain signal can increase the diversity of environmental noise in the data set, enrich the infinitely changing noise in the environment, and enhance the input data set of the model.
[0120] Time domain signal s after simulating environmental noise N (t) can be expressed as:
[0121] s N (t)=s(t)+Noise(t)
[0122] Where Noise(t) is a random value that changes with time, and s(t) represents the original time domain signal. From the perspective of frequency, random noise is a full-frequency noise, which is ultimately reflected in the model input as Figure 2 shown.
[0123] The attenuation parameters are set as follows:
[0124] For the signal of a single audio speaker, the signal energy collected from different angles, different distances, and whether there are objects blocking it is different. Assuming that the noise in the environment is constant, different signal energies result in different signal-to-noise ratios of the collected data.
[0125] The present invention proposes to set a random attenuation coefficient α in the collected time domain signal, which can improve the diversity of attenuation in different situations in the data set and enhance the input data set of the model. E (t) can be expressed as:
[0126] s E (t) = α·s(t)
[0127] Since the input of the model will be normalized, the effect of simple coefficient attenuation will be reduced. Therefore, when simulating the energy attenuation of the signal in the medium, it is necessary to combine random additive noise to simulate different signal-to-noise ratio signals. Figure 3 shown.
[0128] The S2, random inhibition curve is set as follows:
[0129] Frequency response is an important feature of the signal, which describes the relative difference in the amplitude of the captured sound signal at different frequencies. However, this feature is irrelevant to audio ranging and may even interfere with it. The data source only contains data collected by a small number of devices, and the detection accuracy of the trained model for data collected by other devices will be reduced.
[0130] The present invention proposes to set a random suppression curve in the collected time domain signal, and after point multiplication with the collected time domain signal, the amplitude difference between each frequency point is readjusted, which can improve the limited frequency response in the data set and enhance the input data set of the model.
[0131] The random curve is generated by sliding the window in the random vector and averaging the random value to smooth the random value jump. The time domain signal s after simulating the hardware frequency response R (t) can be expressed as:
[0132] s R (t) = curve(t)·s(t)
[0133] Where curve(t) is the random inhibition curve that changes with time. The final effect is reflected in the model input, such as Figure 4 shown.
[0134] The simulation settings of S3, multipath interference are as follows:
[0135] Multipath interference is also a propagation feature that is strongly related to the environment. Data collected in different places in different environments are subject to different echo reverberation interference, which mainly depends on the direction of the speaker and the position and direction of the surrounding objects. Echoes of different delays and intensities will be superimposed on the collected signal, which is inevitable.
[0136] The present invention proposes to superimpose original signals with different delays and different intensities on the collected time domain signals to simulate the echo phenomenon in the actual environment, which can improve the diversity of multipath interference in the data set and enhance the input data set of the model.
[0137] The time domain signal s after simulating multipath interference M (t) can be expressed as:
[0138]
[0139] Where is the number of paths of random P machines, λ p is the random intensity coefficient of different paths, τ p is the random delay of different paths. It is finally reflected in the model input as Figure 5 shown.
[0140] The S4, short-time Fourier transform is as follows:
[0141] Assuming the length of the sliding window is WL and the moving step of the window is SL, the time resolution is (SL / Fs)s and the frequency resolution is (Fs / WL)Hz. The detailed steps of STFT are as follows:
[0142] First, the sliding window starts from the starting point of the received sound signal data. At this time, the window function is t = τ 0 As the center, the signal is processed by windowing function:
[0143] y(t)=x(t)·w(t-τ 0 )
[0144] Then, Fourier transform is performed on the data to obtain the PSD matrix of the sound data of the first window. Indicates that the received signal is delayed by (0, τ 0 ] is the vector at ].
[0145]
[0146] Where x(t) represents the sound data segment of the received signal R(0:WL]; w is the Hamming window function; f m Depends on the sampling rate (Frequency ofSampling, Fs) of the smartphone, ranging from 0Hz to Fs / 2Hz.m and τ 0 The definitions are as follows:
[0147]
[0148] τ 0 =(WL / 2) / Fs
[0149] Finally, the PSD matrix representing the sound data segment of the nth window is calculated as follows:
[0150]
[0151] In the formula, The received signal is delayed by (τ n-1 , τ n ], where x(t n ) and τ n The definition is as follows:
[0152] x(t n )=R[(n-1)×SL:WL+(n-1)×SL]
[0153] τ n =[WL / 2+(n-1)×SL] / Fs
[0154] When performing short-time Fourier transform, the step size of the window sliding determines the time resolution of the input image, while the size of the window determines the frequency resolution of the input image. Taking 1ms as the step size of the window sliding is equivalent to a resolution of 0.34m, and in order to facilitate fast Fourier transform, the window size is selected as 512.
[0155] S5, introducing random frequency deviation to simulate the influence of Doppler effect, can improve the signal diversity of the data set for different relative motion conditions, and can enhance the input data set of the model, which is finally reflected in the model input such as Figure 6 shown.
[0156] like Figure 7 As shown, an embodiment of the present invention provides a data enhancement system for an audio signal, specifically comprising:
[0157] A signal preprocessing module, used for normalizing the input time domain signal data;
[0158] A noise generation module, used to generate random noise;
[0159] Signal simulation module, used to simulate different frequency responses and multipaths of signals;
[0160] The filtering module is used to pass the collected time domain signal through a bandpass filter to reduce the environmental noise in non-target frequency bands;
[0161] Fourier transform module, used to convert the filtered time domain signal into a time-frequency image;
[0162] The frequency deviation module is used to shift the time-frequency diagram according to the random frequency deviation to complete the simulation of Doppler frequency deviation;
[0163] Output module, used to output data enhancement results.
[0164] In order to fully verify the performance of the data enhancement method proposed in the present invention, the present invention designed two groups of experiments for comparison from multiple angles.
[0165] (1) Experimental description and explanation
[0166] The present invention takes into account different situations such as different experimental environments and the use of different equipment terminals to fully verify the data enhancement method for audio signals proposed in the present invention.
[0167] First of all, the present invention selected three different experimental scenes for experiments. The first scene is the corridor environment of an office building, which is about 35 meters long and 2 meters wide. This is a typical narrow and long indoor environment, which is relatively closed and has less interference. It is an ideal experimental scene. The second scene is an underground parking lot, where a large number of vehicles are parked and the floor height is low. There are interferences such as vehicle starting, moving, and honking in this environment, and there are more ventilation ducts making noise. The space is relatively transparent and the layout is complex, but it is a typical indoor application scene. The third scene is a shopping mall environment, which is an open environment with a large atrium in the middle, and there are many cylindrical columns on each floor, and the floor height is relatively high.
[0168] Secondly, the present invention selected four different terminal devices for experiments, three of which are mobile phones issued on the market, and another one uses a work badge as a terminal. The three mobile phones are Xiaomi 5s Plus, Huawei Nova7, and Vivo X80 Pro. These three mobile phones come from different manufacturers and have different ideas in microphone selection. And the release time of the three mobile phones is also different. Xiaomi 5s Plus was released in 2016, Huawei Nova7 was released in 2020, and Vivo X80 Pro was released in 2022, so the device usage time is also different. Different mobile phone manufacturers and different usage times have greatly enriched the diversity of device terminals.
[0169] The specific process of the experiment is as follows. Taking the corridor scene as an example, four speaker devices are used in the experiment of the present invention, three of which are placed at one end of the scene as the base stations to be measured. Because the time synchronization between the mobile phone device and the speaker terminal is not stable, it is difficult to directly obtain the distance measurement. The present invention places the remaining base station together with the terminal, and at such a suitable close distance, it can basically ensure that the arrival time of the base station is accurate.
[0170] (2) Comparison and analysis of static ranging effects in different scenarios
[0171] Taking the test in the corridor as an example, the audio base station equipment is placed at one end of the corridor, and marks are made at distances of 5 meters, 10 meters, 20 meters, and 30 meters, respectively, and the above four different equipment terminals are used for data collection. The statistical results of static ranging in different scenarios are shown in Table 1. Overall, the experimental results in the underground parking lot and the shopping mall are much worse than those in the corridor experiment. This is because in the real environment, the layout of the positioning space becomes complicated, the signal transmission path becomes complex and diverse, and the movement and interaction of people, vehicles and objects in the environment will also generate a lot of noise and interference. The maximum error comes from Xiaomi 5s Plus. In the other three terminal devices, the maximum error is about 1 meter. From the perspective of root mean square error and standard deviation, the results of the shopping mall are worse. This is because the selected shopping mall venue is connected to the outside world, which is equivalent to a semi-open environment, and the sound attenuation will be more serious. Although in this experiment, there are some differences in the maximum jitter in each scene, which is caused by the different scene environments, the overall ranging performance is stable, and the root mean square difference of different devices is maintained at around 0.200 meters. For example, the root mean square difference of Xiaomi 5s Plus in a shopping mall is 0.380 meters.
[0172] Table 1 Statistics of static ranging errors in different scenarios (unit: meter)
[0173]
[0174] (3) Comparison and analysis of dynamic ranging effects in different scenarios
[0175] After comparing the static ranging performance, the dynamic ranging effect cannot be ignored. After all, in actual use, the tracked and located device is in motion most of the time. The test scenarios, equipment and methods are basically the same as those of the static test. In the dynamic ranging comparison experiment, each experiment starts from about 30 meters away from the base station, approaches the base station equipment at different speeds, and when the distance is only about 5 meters, it moves away from the base station and returns to the starting point, and then immediately approaches and returns, a total of two round trips. The movement process is uniform linear motion, and each time it passes a ground mark point, the arrival timestamp is recorded, so that the true value of the distance measurement at the mark point can be confirmed according to the time. Each set of experiments gives a quantitative statistical table of the error results, and the statistical indicators are one standard deviation and two standard deviations, abbreviated as CDF-68% and CDF-95%.
[0176] The specific data statistics are shown in Table 2. In each scenario, the algorithm proposed by the present invention fluctuates by less than 0.5 meters. In terms of ranging accuracy, the algorithm proposed by the present invention is also very advantageous. The distance jump detected by Xiaomi 5s Plus in the underground parking lot is the largest, while other terminals are relatively stable.
[0177] Table 2 Statistics of dynamic ranging errors in different scenarios (unit: meter)
[0178]
[0179] It should be noted that the embodiments of the present invention can be implemented by hardware, software, or a combination of software and hardware. The hardware part can be implemented using dedicated logic; the software part can be stored in a memory and executed by an appropriate instruction execution system, such as a microprocessor or dedicated design hardware. It can be understood by a person of ordinary skill in the art that the above-mentioned devices and methods can be implemented using computer executable instructions and / or contained in a processor control code, such as a carrier medium such as a disk, CD or DVD-ROM, a programmable memory such as a read-only memory (firmware), or a data carrier such as an optical or electronic signal carrier. Such code is provided on the carrier medium. The device and its modules of the present invention can be implemented by hardware circuits such as very large-scale integrated circuits or gate arrays, semiconductors such as logic chips, transistors, etc., or programmable hardware devices such as field programmable gate arrays, programmable logic devices, etc., can also be implemented by software executed by various types of processors, and can also be implemented by a combination of the above-mentioned hardware circuits and software, such as firmware.
[0180] The above description is only a specific implementation mode of the present invention, but the protection scope of the present invention is not limited thereto. Any modifications, equivalent substitutions and improvements made by any technician familiar with the technical field within the technical scope disclosed by the present invention and within the spirit and principle of the present invention should be covered by the protection scope of the present invention.
Claims
1. A data enhancement method for an audio signal, characterized in that: The method specifically includes: S1: normalize the input time domain signal data, generate a set of normalized random noise, and generate new signal data with a specified signal-to-noise ratio according to the corresponding attenuation parameters and noise coefficient; S2: Generate a random suppression curve in advance, whose length is consistent with the length of the time domain signal data obtained by processing in step S1, and transform different frequency responses through the random suppression curve; S3: delaying and superimposing the time domain signal obtained by step S2 to simulate multipath interference; S4: The time domain signal obtained by S3 is passed through a bandpass filter to reduce the environmental noise in the non-target frequency band, and the filtered time domain signal is converted into a time-frequency image through short-time Fourier transform (STFT) to obtain a time-frequency diagram of the signal; S5: offset the time-frequency diagram according to the random frequency offset, complete the Doppler frequency offset simulation, and then cut out the target frequency band image; S6: Normalize the image to prevent the value from affecting the model training; S2, the time domain signal s obtained by processing in S1 E (t) sets a random suppression curve, which is consistent with the time domain signal s E After point multiplication (t), the amplitude difference between each frequency point is readjusted. The random curve is generated by sliding the window and averaging the random vector to smooth the random value jumps. The time domain signal s after simulating the hardware frequency response is R (t) can be expressed as: s R (t)=curve(t)·s E (t) Where curve(t) is the random inhibition curve that changes with time; The S3, in the time domain signal s R (t) superimposes signals s with different delays and strengths R (t-τ p ), simulating the echo phenomenon in the actual environment, simulating the time domain signal s after multipath interference M (t) can be expressed as: Where P is the number of random paths, γ P is the random intensity coefficient of different paths, τ p is the random delay of different paths.
2. The data enhancement method for an audio signal according to claim 1, characterized in that: S1, the time domain signal s after simulating environmental noise N (t) can be expressed as: s N (t)=s(t)+Noise(t) Where Noise(t) is a random value that changes with time, s(t) represents the original time domain signal; N (t) sets a random attenuation coefficient α to simulate the time domain signal s after energy attenuation E (t) can be expressed as: s E (t)=α·s N (t) While simulating the different degrees of energy attenuation of the signal propagating in the medium, random additive noise is combined to simulate signals with different signal-to-noise ratios.
3. The data enhancement method for an audio signal according to claim 1, characterized in that: The S4, short-time Fourier transform is as follows: Assuming that the length of the sliding window is WL and the moving step of the window is SL, the time resolution is (SL / Fs)s, the frequency resolution is (Fs / WL)Hz, and the detailed steps of STFT are as follows: First, start sliding the window from the starting point of the received sound signal data. At this time, the window function is centered at t = τ0 and performs windowing function processing on the signal: y(t)=x(t)·w(t-τ0) Then, Fourier transform is performed on the data to obtain the PSD matrix of the sound data of the first window. Represents the vector of the received signal at the time delay (0,τ0]: Where x(t) represents the sound data segment of the received signal R(0:WL], w is the Hamming window function, and f m Depends on the sampling rate (Frequency of Sampling, Fs) of the smartphone, ranging from 0Hz to Fs / 2Hz, f m and τ0 are defined as follows: τ0=(WL / 2) / Fs Finally, the PSD matrix representing the sound data segment of the nth window is calculated as follows: In the formula, The received signal is delayed by (τ n-1 , τ n ], where x(t n ) and τ n The definition is as follows: x(t n )=R[(n-1)×SL:WL+(n-1)×SL] t n =[WL / 2+(n-1)×SL] / Fs When performing short-time Fourier transform, the step size of the window sliding determines the time resolution of the input image, and the size of the window determines the frequency resolution of the input image. A window sliding step size of 1 ms is equivalent to a resolution of 0.34 m. In order to facilitate fast Fourier transform, the window size is selected to be 512.
4. A data enhancement system for an audio signal according to claims 1-3, characterized in that: The system specifically includes: A normalization processing module is used to normalize the input time domain signal data and generate a set of normalized random noises, and generate new signal data with a specified signal-to-noise ratio by combining the attenuation parameter and the noise coefficient; A random suppression curve generation module is used to generate a random suppression curve consistent with the length of data to be enhanced, and transform different frequency responses through the curve; Multipath interference simulation module, used to delay and superimpose signals to simulate multipath interference; The bandpass filter and short-time Fourier transform module is used to pass the collected time domain signal through a bandpass filter to reduce the environmental noise in the non-target frequency band, and convert the filtered time domain signal into a time-frequency image through short-time Fourier transform to obtain the time-frequency diagram of the signal; The Doppler frequency shift simulation and image normalization module is used to shift the time-frequency diagram according to the random frequency shift, complete the Doppler frequency shift simulation, crop the target frequency band image, and then normalize the image.
5. The audio signal data enhancement system according to claim 4, characterized in that: The normalization processing module comprises: An environmental noise simulation unit, used for simulating environmental noise and generating a time domain signal after simulating environmental noise; The energy attenuation simulation unit is used to set a random attenuation coefficient, simulate the energy attenuation of the signal when it propagates in the medium, and generate a time domain signal after the simulated energy attenuation.
6. The audio signal data enhancement system according to claim 4, characterized in that: The random inhibition curve generation module further comprises: The frequency response simulation unit is used to generate a random suppression curve, perform point multiplication with the time domain signal, readjust the amplitude difference between each frequency point, and simulate the time domain signal after the hardware frequency response.
7. The audio signal data enhancement system according to claim 4, characterized in that: The bandpass filtering and short-time Fourier transform module includes: A bandpass filter unit is used to pass the collected time domain signal through a bandpass filter to reduce environmental noise in non-target frequency bands; The short-time Fourier transform unit is used to convert the filtered time domain signal into a time-frequency image. Specifically, the signal is processed by a sliding window with a window function, and the data is Fourier transformed to obtain the PSD matrix of the sound data.
Citation Information
Patent Citations
Radar anti-interference detection optimization system and method based on simulation technology
CN114755644A
Audio indoor positioning method and system based on CDMA (Code Division Multiple Access)
CN117793889A