Multi-source periodic audio signal processing method and device and storage medium
Through noise reduction, DC offset removal, resampling and main frequency extraction, the peak and short-term average amplitude alignment of multi-source periodic audio signals is performed, and amplitude standardization is performed, which solves the problems of amplitude difference, time offset and periodic feature loss, and improves the accuracy and efficiency of signal alignment.
Patent Information
- Application Number
- CN202510136415.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-07
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2045-02-07
AI Technical Summary
The prior art has problems of amplitude difference, time offset and periodic feature loss when dealing with multi-source periodic audio signals, which affects the alignment and standardization effects of signals.
By obtaining one of the multiple audio signals as a reference signal, noise reduction, DC offset removal and resampling are performed, the main frequency is extracted to determine the signal period, and the signal is aligned using peak alignment and short-time average amplitude alignment methods, and the amplitude standardization is finally performed to ensure that the amplitude range of the signal is consistent.
The timing alignment accuracy and efficiency of multi-source periodic audio signals is improved, the natural dynamic range of the signal is ensured, it is suitable for complex or noise interference scenarios, and the periodic structure of long-period audio signals is maintained.
Smart Images

Figure CN119964588A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of signal technology, and in particular to a multi-source periodic audio signal processing method, device and storage medium. Background Art
[0002] Audio signal processing is the basis of time domain analysis. The existing technology has the following problems in the timing processing of periodic audio signals:
[0003] Amplitude differences of multi-source signals: When processing multi-source audio signals, amplitude differences often occur, for example, due to different pickup amplification factors or different microphone placement positions. Traditional methods often result in large amplitude differences when processing multi-source signals, which in turn affects the alignment and standardization effects.
[0004] Multi-source signal alignment problem: When multiple different sound collectors are recording at the same time, the signals may be time-shifted due to device transmission delay or other factors. Traditional processing methods are inefficient and inaccurate in solving the time shift problem.
[0005] Integrity issues of periodic signals: For audio signals with longer periodicity (such as industrial machine sounds or natural sounds lasting more than a few seconds), traditional standardization processing methods may ignore the periodic characteristics of the signal, resulting in the loss of important periodic information during data processing. This not only affects the usability of the signal, but may also lead to inaccurate analysis results. Summary of the invention
[0006] In view of the problems existing in the above-mentioned prior art, the present invention provides a multi-source periodic audio signal processing method, device and storage medium. The technical solution is as follows:
[0007] In a first aspect, a method for processing a multi-source periodic audio signal is provided, comprising the following steps:
[0008] Acquire multiple collected audio signals, take one of the multiple audio signals as a reference signal and record it as a first audio signal, and take the remaining signals as signals to be processed and record them as second audio signals;
[0009] Performing noise reduction and DC offset removal processing on the first audio signal and the second audio signal;
[0010] Resampling the first audio signal and the second audio signal using an interpolation algorithm, wherein the sampling frequency after resampling is the least common multiple of the sampling rates of the first audio signal and the second audio signal;
[0011] Extracting audio features from the first audio signal and the second audio signal, including extracting the main frequency of the audio signal;
[0012] Determine the period of the audio signal based on the main frequency of the audio signal, and perform signal alignment on the audio signal with a single period as the minimum processing unit, wherein the signal alignment includes alignment based on the peak signal and alignment based on the short-time average amplitude value of the period;
[0013] The two aligned periodic audio signals are amplitude-normalized to ensure that the amplitude ranges of the two signals are consistent.
[0014] In some implementations, extracting the main frequency of the audio signal includes:
[0015] The fast Fourier transform (FFT) is used to calculate the spectrum of the signal, and the frequency with the largest amplitude in the spectrum is extracted as the main frequency f 0 .
[0016] In some implementations, the step of aligning the audio signal with a single cycle as the minimum processing unit includes:
[0017] For the first audio signal and the second audio signal, respectively, according to the main frequency f of the audio signal 0 Determine the period T 0 , according to the period T 0 , finding the start and end points of the cycles of the first audio signal and the second audio signal in the time domain, and extracting complete cycle segments of the first audio signal and complete cycle segments of the second audio signal respectively;
[0018] For the first audio signal and the second audio signal, respectively obtain peak points in each cycle, form a peak sequence of the first audio signal as a first peak sequence, and form a peak sequence of the second audio signal as a second peak sequence, wherein the peak points in each cycle are determined based on an extreme value method;
[0019] calculating a preliminary time shift of the first audio signal and the second audio signal based on the first peak sequence and the second peak sequence, Where N is the number of peak signals, t ref,i is the time of the i-th peak of the reference signal, i.e., the first audio signal, ref represents the reference signal, t target,i represents the time of the i-th peak of the signal to be processed, i.e., the second audio signal, and target represents the target signal;
[0020] According to the calculated preliminary time offset Δt peak , a time offset, i.e., a translation of the time axis, is applied to the target signal to perform preliminary peak alignment between the target signal and the reference signal.
[0021] In some implementations, the signal alignment of the audio signal with a single cycle as the minimum processing unit further includes:
[0022] After the target signal and the reference signal are initially peak aligned,
[0023] The first audio signal and the second audio signal are divided into a plurality of first windows according to their respective periods, each of which includes a complete period segment;
[0024] For the first window, the signal is divided twice using a second window having a length smaller than that of the first window to obtain a plurality of second window segments in the first window;
[0025] For the second window, calculating the short-time average amplitude of the signal in the second window based on the signal amplitudes of all samples in the second window;
[0026] The short-time average amplitude of the second window is weighted, and after weighted fusion, a weighted average amplitude value of the first window is obtained; a window weighted average amplitude sequence of the first audio signal and a window weighted average amplitude sequence of the second audio signal are obtained;
[0027] Based on the window weighted average amplitude sequence of the first audio signal and the window weighted average amplitude sequence of the second audio signal, periodic window alignment is performed. The alignment process is: with the first audio signal fixed, the time axis of the second audio signal is gradually adjusted by time shifting, so that the error of the window weighted average amplitude sequence of the two is minimized. The error of the window weighted average amplitude sequence of the two is Wherein, M is the number of periodic windows of the aligned segment after the first audio signal and the second audio signal are aligned, and k is the kth periodic window in the aligned segment.
[0028] In some implementations, the short-time average amplitude of the second window is weighted, and a large weight is given to the second window containing the peak value, that is, the closer the second window is to the peak value, the larger the weight is.
[0029] In some implementations, the performing noise reduction and DC offset removal processing on the first audio signal and the second audio signal includes:
[0030] Perform wavelet transform on the signal to obtain wavelet transform coefficients;
[0031] Based on a preset wavelet coefficient threshold, removing the portion of the wavelet coefficient whose amplitude is lower than the preset wavelet coefficient threshold;
[0032] Perform inverse wavelet transform to restore the denoised signal;
[0033] Calculate the mean of the signal and subtract it from the signal;
[0034] Center-clips the signal, removing the portion that exceeds a preset signal amplitude threshold.
[0035] In some implementations, resampling the first audio signal and the second audio signal using an interpolation algorithm, wherein the sampling frequency after resampling is the least common multiple of the sampling rates of the first audio signal and the second audio signal, includes:
[0036] Get the least common multiple f of the sampling rates of the first audio signal and the second audio signal LCM ;
[0037] The first audio signal and the second audio signal are resampled based on a B-spline interpolation algorithm.
[0038] In a second aspect, a multi-source periodic audio signal processing device is provided, comprising:
[0039] An audio signal acquisition unit, used to acquire multiple collected audio signals, take one of the multiple audio signals as a reference signal and record it as a first audio signal, and record the remaining signals as signals to be processed and record them as second audio signals;
[0040] A pre-processing unit, configured to perform noise reduction and DC offset removal processing on the first audio signal and the second audio signal;
[0041] A resampling unit, configured to resample the first audio signal and the second audio signal by using an interpolation algorithm, wherein the sampling frequency after resampling is the least common multiple of the sampling rates of the first audio signal and the second audio signal;
[0042] A signal feature acquisition unit, configured to extract audio features from the first audio signal and the second audio signal, including extracting the main frequency of the audio signal;
[0043] A signal alignment unit, configured to determine the period of the audio signal based on the main frequency of the audio signal, and perform signal alignment on the audio signal with a single period as the minimum processing unit, wherein the signal alignment includes alignment based on a peak signal and alignment based on a short-term average amplitude value of a period;
[0044] The signal standardization unit is used to standardize the amplitudes of the two aligned periodic audio signals to ensure that the amplitude ranges of the two signals are consistent.
[0045] In some embodiments, the signal alignment unit includes: a preliminary peak alignment unit and an optimization alignment unit.
[0046] The preliminary peak alignment unit is used to perform the following steps:
[0047] For the first audio signal and the second audio signal, respectively, according to the main frequency f of the audio signal 0 Determine the period T 0 , according to the period T 0, finding the start and end points of the cycles of the first audio signal and the second audio signal in the time domain, and extracting complete cycle segments of the first audio signal and complete cycle segments of the second audio signal respectively;
[0048] For the first audio signal and the second audio signal, respectively obtain peak points in each cycle, form a peak sequence of the first audio signal as a first peak sequence, and form a peak sequence of the second audio signal as a second peak sequence, wherein the peak points in each cycle are determined based on an extreme value method;
[0049] calculating a preliminary time shift of the first audio signal and the second audio signal based on the first peak sequence and the second peak sequence, Where N is the number of peak signals, t ref,i is the time of the i-th peak of the reference signal, i.e., the first audio signal, ref represents the reference signal, t target,i represents the time of the i-th peak of the signal to be processed, i.e., the second audio signal, and target represents the target signal;
[0050] According to the calculated preliminary time offset Δt peak , apply a time offset, i.e., a time axis translation, to the target signal to perform preliminary peak alignment between the target signal and the reference signal
[0051] The optimization alignment unit is used to perform the following steps:
[0052] After the target signal and the reference signal are preliminarily peak aligned, the first audio signal and the second audio signal are divided into a plurality of first windows according to their respective periods, each of the first windows including a complete period segment;
[0053] For the first window, the signal is divided twice using a second window having a length smaller than that of the first window to obtain a plurality of second window segments in the first window;
[0054] For the second window, calculating the short-time average amplitude of the signal in the second window based on the signal amplitudes of all samples in the second window;
[0055] The short-time average amplitude of the second window is weighted, and after weighted fusion, a weighted average amplitude value of the first window is obtained; a window weighted average amplitude sequence of the first audio signal and a window weighted average amplitude sequence of the second audio signal are obtained;
[0056] Based on the window weighted average amplitude sequence of the first audio signal and the window weighted average amplitude sequence of the second audio signal, periodic window alignment is performed. The alignment process is: with the first audio signal fixed, the time axis of the second audio signal is gradually adjusted by time shifting, so that the error of the window weighted average amplitude sequence of the two is minimized. The error of the window weighted average amplitude sequence of the two is Wherein, M is the number of periodic windows of the aligned segment after the first audio signal and the second audio signal are aligned, and k is the kth periodic window in the aligned segment.
[0057] According to a third aspect, a computer-readable storage medium is provided, on which computer instructions are stored. When the instructions are executed by a processor, the steps of the multi-source periodic audio signal processing method according to the first aspect are implemented.
[0058] The present invention provides a multi-source periodic audio signal processing method, device and storage medium, which have the following beneficial effects:
[0059] 1. The present invention improves the accuracy and efficiency of timing alignment through innovative methods such as peak alignment, short-time average amplitude difference (STAAD) sequence calculation, and periodic feature detection, especially in complex or noisy scenes.
[0060] 2. The present invention combines DC offset removal, center clipping and sophisticated amplitude normalization methods to effectively balance the amplitude differences of multi-source audio signals and ensure that the signals maintain a natural dynamic range during the alignment process.
[0061] 3. The present invention first performs a rapid preliminary alignment through peak alignment, and further refines the alignment accuracy using the STAAD method, which can effectively improve the robustness and accuracy of the alignment, especially in the presence of noise, small errors or irregular periodic changes.
[0062] 4. The present invention ensures that the periodic structure of the long-period audio signal is not lost through periodic window division and short-time average amplitude calculation, and is particularly suitable for long-period audio signal processing in a multi-source complex environment. BRIEF DESCRIPTION OF THE DRAWINGS
[0063] Figure 1 It is a flowchart of a method for processing multi-source periodic audio signals in an embodiment of the present application;
[0064] Figure 2 is a flow chart of a signal alignment method in an embodiment of the present application;
[0065] Figure 3 It is a structural schematic diagram of a multi-source periodic audio signal processing device in an embodiment of the present application. DETAILED DESCRIPTION
[0066] It should be understood that the specific embodiments described herein are only used to explain the present invention, and are not used to limit the present invention.
[0067] The present invention provides a method for processing a multi-source periodic audio signal, comprising the following steps:
[0068] Step 1, obtaining multiple collected audio signals, taking one of the multiple audio signals as a reference signal and recording it as a first audio signal, and recording the remaining signals as signals to be processed and recording them as second audio signals;
[0069] Step 2, performing noise reduction and DC offset removal processing on the first audio signal and the second audio signal;
[0070] Step 3, resampling the first audio signal and the second audio signal using an interpolation algorithm, wherein the sampling frequency after resampling is the least common multiple of the sampling rates of the first audio signal and the second audio signal;
[0071] Step 4, extracting audio features from the first audio signal and the second audio signal, including extracting the main frequency of the audio signal;
[0072] Step 5, determining the period of the audio signal based on the main frequency of the audio signal, and performing signal alignment on the audio signal with a single period as the minimum processing unit, wherein the signal alignment includes alignment based on the peak signal and alignment based on the short-time average amplitude value of the period;
[0073] Step 6: normalize the amplitudes of the two aligned periodic audio signals to ensure that the amplitude ranges of the two signals are consistent.
[0074] In the embodiment of the present application, a more sophisticated amplitude normalization method is provided by combining DC offset removal and center clipping processing to effectively balance the amplitude differences of multi-source signals. This method can adapt to the gain differences of different recording devices and pickups, and maintain the natural dynamic range of the signal during the processing process, avoiding signal distortion or excessive compression caused by amplitude differences.
[0075] In the embodiment of the present application, by introducing the peak alignment method and the short-time average amplitude sequence (STAAD) alignment calculation, combined with the periodic characteristics of the signal, the signal is accurately aligned in time. In this process, the time domain and frequency domain information based on the periodic characteristics are used, and the calculation efficiency of the timing alignment is improved at the same time, avoiding the dependence on large amount of calculation in the traditional method. In addition, the correlation calculation method of the STAAD sequence is used to improve the alignment accuracy and reduce the influence of noise, solving the problem of low accuracy of the prior art under noise interference.
[0076] In one implementation, in the above step 2, the performing of noise reduction and DC offset removal processing on the first audio signal and the second audio signal includes:
[0077] Step 21, performing wavelet transform on the signal to obtain wavelet transform coefficients;
[0078] Step 22, based on the preset wavelet coefficient threshold, remove the part of the wavelet coefficient whose amplitude is lower than the preset wavelet coefficient threshold; let x(t) be the original signal, W ψ (x(t)) is the wavelet transform coefficient, and the wavelet coefficient threshold T is preset. The part of the wavelet coefficient whose amplitude is lower than the preset wavelet coefficient threshold T is removed as follows:
[0079] Step 23, perform inverse wavelet transform to restore the denoised signal:
[0080] Step 24, calculate the mean of the signal and subtract it from the signal: x 2 (t) = x 1 (t)-μ(x), where μ(x) is the mean of the original signal;
[0081] Step 25, perform center clipping on the signal to remove the portion exceeding the preset signal amplitude threshold: obtain the signal x after noise reduction and DC offset removal processing 3 (t):
[0082] In the audio signal preprocessing, the present invention adds DC offset removal after denoising, which effectively minimizes the signal difference, thereby improving the overall effect of signal alignment and standardization.
[0083] In one implementation, in the above step 3, resampling the first audio signal and the second audio signal using an interpolation algorithm, wherein the sampling frequency after resampling is the least common multiple of the sampling rates of the first audio signal and the second audio signal, includes:
[0084] Get the least common multiple f of the sampling rates of the first audio signal and the second audio signal LCM ;
[0085] The first audio signal and the second audio signal are resampled based on a B-spline interpolation algorithm.
[0086] The present invention adopts B-spline interpolation technology to provide higher accuracy and smoother results in the audio resampling process, thereby effectively solving the problem of audio quality degradation caused by inconsistent sampling rates. The use of B-spline interpolation technology instead of traditional linear interpolation methods can provide higher interpolation accuracy, especially in the processing of high-frequency signals, which can better maintain the smoothness of the signal and avoid signal distortion. Through precise sampling rate adjustment, the seamless connection of audio signals between different devices and systems is ensured, and the overall audio quality is improved.
[0087] In one embodiment, in the above step 4, extracting the main frequency of the audio signal includes: using fast Fourier transform (FFT) to calculate the spectrum of the signal, and extracting the frequency with the largest amplitude in the spectrum as the main frequency f 0 .
[0088] In one implementation, in the above step 5, the signal alignment of the audio signal is performed with a single cycle as the minimum processing unit, wherein the peak-based alignment includes:
[0089] Step 501: for the first audio signal and the second audio signal, respectively, according to the main frequency f of the audio signal 0 Determine the period T 0 , according to the period T 0 , finding the start and end points of the cycles of the first audio signal and the second audio signal in the time domain, and extracting complete cycle segments of the first audio signal and complete cycle segments of the second audio signal respectively;
[0090] Step 502: for the first audio signal and the second audio signal, respectively obtain peak points in each cycle, form a peak sequence of the first audio signal as a first peak sequence, and form a peak sequence of the second audio signal as a second peak sequence, wherein the peak points in each cycle are determined based on an extreme value method;
[0091] Step 503: Calculate a preliminary time offset between the first audio signal and the second audio signal based on the first peak sequence and the second peak sequence. Where N is the number of peak signals, t ref,i is the time of the i-th peak of the reference signal, i.e., the first audio signal, ref represents the reference signal, t target,i represents the time of the i-th peak of the signal to be processed, i.e., the second audio signal, and target represents the target signal;
[0092] Step 504: based on the calculated preliminary time offset Δt peak , a time offset, i.e., a translation of the time axis, is applied to the target signal to perform preliminary peak alignment between the target signal and the reference signal.
[0093] In this embodiment, peak alignment is used. The signal peak point represents the maximum or minimum point in the signal cycle. By detecting the local peak of the signal, the important time points in the signal can be determined. The peak value detection algorithm is used to find the peak value in each cycle, and the main peak time in each cycle is obtained. For the reference signal and the target signal, their peak time points are detected respectively to obtain the peak sequence P ref and P target , that is, the time position of each peak. Compare the peak point sequence P ref and P target, the time offset between the two signals is preliminarily calculated. By calculating the difference between the peak points, a preliminary alignment offset Δt is obtained peak , N is the number of peak points, the calculated Δt peak is an overall timing offset.
[0094] In one implementation, in the above step 5, the signal alignment of the audio signal with a single cycle as the minimum processing unit also includes alignment based on the short-term average amplitude value of the cycle, including the following steps:
[0095] Step 505, after the target signal and the reference signal are preliminarily peak aligned, the first audio signal and the second audio signal are divided into a plurality of first windows according to their respective periods, each of the first windows including a complete period segment;
[0096] Step 506, with respect to the first window, divide the signal twice using a second window having a length smaller than that of the first window to obtain a plurality of second window segments in the first window;
[0097] Step 507, for the second window, calculating the short-time average amplitude of the signal in the second window based on the signal amplitudes of all samples in the second window;
[0098] Step 508, weighting the short-time average amplitude of the second window, and obtaining the weighted average amplitude value of the first window through weighted fusion; obtaining the window weighted average amplitude sequence of the first audio signal and the window weighted average amplitude sequence of the second audio signal;
[0099] Step 509, periodic window alignment is performed based on the window weighted average amplitude sequence of the first audio signal and the window weighted average amplitude sequence of the second audio signal. The alignment process is: with the first audio signal fixed, the time axis of the second audio signal is gradually adjusted by time shifting, so that the error of the window weighted average amplitude sequences of the two is minimized. The error of the window weighted average amplitude sequences of the two is Wherein, M is the number of periodic windows of the aligned segment after the first audio signal and the second audio signal are aligned, and k is the kth periodic window in the aligned segment.
[0100] In the embodiment of the present application, after peak alignment, the overall periodic characteristics of the signal have been roughly aligned, but there may still be slight deviations or noise effects. At this stage, the STAAD (short-time average amplitude) sequence is used to further optimize the alignment of the signal. During the alignment process, the time axis of the target signal is gradually adjusted by time shifting until the difference between the STAAD sequences of the two is minimized, the target signal is time-shifted in a small range, and then the optimized MSE is calculated. This process is iterated to gradually minimize the MSE until the difference between the STAAD sequences of the two converges to a minimum value.
[0101] Traditional signal standardization methods usually ignore the periodic characteristics of the signal, which is easy to cause the loss of periodic information, especially in the processing of long-period signals. However, the embodiment of the present application can maintain the periodic structure of the signal by introducing methods such as periodic feature detection and intra-period window division, thereby ensuring that key information is not lost during the alignment and standardization process, which is particularly suitable for the processing of long-period audio signals such as industrial machine sounds and natural environment sounds.
[0102] In this embodiment, a rapid preliminary alignment is first performed through peak alignment, and the alignment accuracy is further refined using the STAAD method, which can effectively improve the robustness and accuracy of the alignment, especially in the presence of noise, small errors or irregular periodic changes.
[0103] Specifically, in the above step 508, the short-time average amplitude of the second window is weighted, a larger weight is given to the second window containing the peak, and a smaller weight is given to the second window without the peak, that is, the closer the second window is to the peak, the larger the weight is.
[0104] In the embodiment of the present application, the short-time amplitude of the weighted peak point is used to enhance the local features. The STAAD method measures the degree of signal alignment by calculating the average amplitude difference of the signal in the short-time window. In order to eliminate noise and instability, STAAD in the embodiment of the present application weights important windows (such as windows containing peaks) to optimize signal alignment. Unlike the peak alignment method, STAAD does not rely solely on a single peak point, but considers the amplitude distribution of the entire signal, thereby more comprehensively optimizing the signal alignment effect. The STAAD method can effectively eliminate small alignment errors caused by factors such as noise, inaccurate peak position, and periodic drift.
[0105] In one implementation, in step 6, the amplitude of the two aligned periodic audio signals is normalized to ensure that the amplitude ranges of the two signals are consistent, including: Among them, x aligned (t) represents the aligned signal, μ and σ are the mean and standard deviation of the signal, respectively.
[0106] Based on the above multi-source periodic audio signal processing method embodiment, the present application embodiment provides a multi-source periodic audio signal processing device 100, including:
[0107] The audio signal acquisition unit 101 is used to acquire multiple collected audio signals, take one of the multiple audio signals as a reference signal and record it as a first audio signal, and record the remaining signals as signals to be processed and record them as second audio signals;
[0108] A pre-processing unit 102, configured to perform noise reduction and DC offset removal processing on the first audio signal and the second audio signal;
[0109] The resampling unit 103 is used to resample the first audio signal and the second audio signal by using an interpolation algorithm, wherein the sampling frequency after resampling is the least common multiple of the sampling rates of the first audio signal and the second audio signal;
[0110] A signal feature acquisition unit 104, configured to extract audio features from the first audio signal and the second audio signal, including extracting the main frequency of the audio signal;
[0111] A signal alignment unit 105 is used to determine the period of the audio signal based on the main frequency of the audio signal, and perform signal alignment on the audio signal with a single period as the minimum processing unit, wherein the signal alignment includes alignment based on the peak signal and alignment based on the short-term average amplitude value of the period;
[0112] The signal standardization unit 106 is used to perform amplitude standardization on the two aligned periodic audio signals to ensure that the amplitude ranges of the two signals are consistent.
[0113] Specifically, the signal alignment unit 105 includes: a preliminary peak alignment unit 1051 and an optimization alignment unit 1052.
[0114] The preliminary peak alignment unit 1051 is used to perform the following steps:
[0115] For the first audio signal and the second audio signal, respectively, according to the main frequency f of the audio signal 0 Determine the period T 0 , according to the period T 0 , finding the start and end points of the cycles of the first audio signal and the second audio signal in the time domain, and extracting complete cycle segments of the first audio signal and complete cycle segments of the second audio signal respectively;
[0116] For the first audio signal and the second audio signal, respectively obtain peak points in each cycle, form a peak sequence of the first audio signal as a first peak sequence, and form a peak sequence of the second audio signal as a second peak sequence, wherein the peak points in each cycle are determined based on an extreme value method;
[0117] calculating a preliminary time shift of the first audio signal and the second audio signal based on the first peak sequence and the second peak sequence, Where N is the number of peak signals, t ref,i is the time of the i-th peak of the reference signal, i.e., the first audio signal, ref represents the reference signal, t target,i represents the time of the i-th peak of the signal to be processed, i.e., the second audio signal, and target represents the target signal;
[0118] According to the calculated preliminary time offset Δt peak , apply a time offset, i.e., a time axis translation, to the target signal to perform preliminary peak alignment between the target signal and the reference signal
[0119] The optimization alignment unit 1052 is used to perform the following steps:
[0120] After the target signal and the reference signal are preliminarily peak aligned, the first audio signal and the second audio signal are divided into a plurality of first windows according to their respective periods, each of the first windows including a complete period segment;
[0121] For the first window, the signal is divided twice using a second window having a length smaller than that of the first window to obtain a plurality of second window segments in the first window;
[0122] For the second window, calculating the short-time average amplitude of the signal in the second window based on the signal amplitudes of all samples in the second window;
[0123] The short-time average amplitude of the second window is weighted, and after weighted fusion, a weighted average amplitude value of the first window is obtained; a window weighted average amplitude sequence of the first audio signal and a window weighted average amplitude sequence of the second audio signal are obtained;
[0124] Based on the window weighted average amplitude sequence of the first audio signal and the window weighted average amplitude sequence of the second audio signal, periodic window alignment is performed. The alignment process is: with the first audio signal fixed, the time axis of the second audio signal is gradually adjusted by time shifting, so that the error of the window weighted average amplitude sequence of the two is minimized. The error of the window weighted average amplitude sequence of the two is Wherein, M is the number of periodic windows of the aligned segment after the first audio signal and the second audio signal are aligned, and k is the kth periodic window in the aligned segment.
[0125] For the specific definition of a multi-source periodic audio signal processing device, reference may be made to the definition of a multi-source periodic audio signal processing method mentioned above, which will not be repeated here.
[0126] In an embodiment of the present application, a computer-readable storage medium is provided, on which computer instructions are stored, characterized in that when the instructions are executed by a processor, the steps of the multi-source periodic audio signal processing method are performed. Computer-readable storage media include: permanent and non-permanent, removable and non-removable media, which are tangible devices that can retain and store instructions for use by instruction execution devices. Computer-readable storage media include: electronic storage devices, magnetic storage devices, optical storage devices, electromagnetic storage devices, semiconductor storage devices, and any suitable combination of the above. Computer-readable storage media include: phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), non-volatile random access memory (NVRAM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassette storage, magnetic tape disk storage or other magnetic storage devices, memory sticks, mechanical encoding devices (such as punched cards or raised structures in grooves with instructions recorded thereon), or any other non-transmission medium that can be used to store information that can be accessed by a computing device.
[0127] The following describes the experimental process and experimental results of the multi-source periodic audio signal processing method provided in the embodiment of the present application.
[0128] 1. Dataset Introduction
[0129] The experiment uses the DCASE2024 Task2 dataset. The DCASE (Detection and Classification of Acoustic Scenes and Events) task is mainly used to evaluate the performance of audio signal processing algorithms. Task 2 is unsupervised abnormal sound detection for machine condition monitoring. The dataset contains normal and abnormal sounds of 7 real machines (fan, gearbox, bearing, slider, toy car, toy train, valve). Each recording is a 10-second mono audio clip that includes the sound of the target machine and the environmental sound.
[0130] 2. Effect evaluation indicators
[0131] In order to evaluate the model performance from multiple perspectives, the AUC and PAUC indicators are calculated using the abnormal score of the test set samples, and the normal / abnormal judgment results of the test set samples are used as evaluation indicators of the model detection effect.
[0132] AUC(source): The AUC (Area Under the Curve) indicator of the source domain audio sample is used to measure the ability to distinguish normal and abnormal sounds.
[0133] AUC(target): The AUC indicator of the target domain audio sample, which is used to measure the model's adaptability to unknown machine states.
[0134] pAUC(source,target): partial AUC indicator, which comprehensively evaluates the performance of source domain and target domain samples.
[0135] TOTAL score: A comprehensive performance indicator that serves as an overall measure of the final model performance.
[0136] 3. Comparative experiment
[0137] In order to verify the effectiveness of the periodic audio signal processing method of the present invention, this experiment conducted a performance comparison test based on the DCASE2024 Task 2 dataset to evaluate the improvement of downstream task performance (anomaly detection) after alignment and standardization.
[0138] The experiment is divided into the following two groups for comparison:
[0139] a. Baseline method: Directly use the official baseline code (baseline_MAHALA and baseline_MSE) provided by DCASE to process audio signals and extract features.
[0140] b. The method of the present invention: On the basis of the baseline method, the audio signal alignment and standardization method of the present invention (our_MAHALA and our_MSE) is added, including audio signal preprocessing (denoising, DC offset removal, center clipping processing), sampling rate adjustment (resampling based on B-spline interpolation method to ensure signal consistency), timing alignment (periodic signal alignment through peak detection and time offset optimization), and standardization (amplitude normalization to ensure signal consistency).
[0141] 4. Experimental results
[0142] The detection performance of the proposed method in the source domain and target domain audio samples is significantly improved compared with the baseline method, and is better than the baseline method overall. The results show that the proposed method can effectively improve the alignment and standardization effect of periodic audio signals. The anomaly detection performance on the DCASE2024 Task 2 dataset is significantly better than the baseline method. The proposed method has broad application prospects and can be used in a variety of audio processing tasks to improve the robustness and accuracy of signal processing.
[0143] Table 1 Comparison results between the method of the present invention and the baseline method
[0144]
[0145] The present invention is not limited to the above-mentioned specific implementation modes. Various changes made by ordinary technicians in this field based on the above-mentioned concepts without creative work are all within the protection scope of the present invention.
Claims
1. A method for processing multi-source periodic audio signals, characterized in that: The steps include: Acquire multiple collected audio signals, take one of the multiple audio signals as a reference signal and record it as a first audio signal, and take the remaining signals as signals to be processed and record them as second audio signals; Performing noise reduction and DC offset removal processing on the first audio signal and the second audio signal; Resampling the first audio signal and the second audio signal using an interpolation algorithm, wherein the sampling frequency after resampling is the least common multiple of the sampling rates of the first audio signal and the second audio signal; Extracting audio features from the first audio signal and the second audio signal, including extracting the main frequency of the audio signal; Determine the period of the audio signal based on the main frequency of the audio signal, and perform signal alignment on the audio signal with a single period as the minimum processing unit, wherein the signal alignment includes alignment based on the peak signal and alignment based on the short-time average amplitude value of the period; The two aligned periodic audio signals are amplitude-normalized to ensure that the amplitude ranges of the two signals are consistent.
2. The multi-source periodic audio signal processing method according to claim 1, characterized in that: The extracting the main frequency of the audio signal comprises: The fast Fourier transform (FFT) is used to calculate the spectrum of the signal, and the frequency with the largest amplitude in the spectrum is extracted as the main frequency f0.
3. The multi-source periodic audio signal processing method according to claim 2, characterized in that: The signal alignment of the audio signal with a single cycle as the minimum processing unit includes: For the first audio signal and the second audio signal, respectively, determine a period T0 according to the main frequency f0 of the audio signal, find the period start and end points of the first audio signal and the second audio signal in the time domain according to the period T0, and respectively extract a complete period segment of the first audio signal and a complete period segment of the second audio signal; For the first audio signal and the second audio signal, respectively obtain peak points in each cycle, form a peak sequence of the first audio signal as a first peak sequence, and form a peak sequence of the second audio signal as a second peak sequence, wherein the peak points in each cycle are determined based on an extreme value method; calculating a preliminary time shift of the first audio signal and the second audio signal based on the first peak sequence and the second peak sequence, Where N is the number of peak signals, t ref,i is the time of the i-th peak of the reference signal, i.e., the first audio signal, ref represents the reference signal, t target,i represents the time of the i-th peak of the signal to be processed, i.e., the second audio signal, and target represents the target signal; According to the calculated preliminary time offset Δt peak , a time offset, i.e., a translation of the time axis, is applied to the target signal to perform preliminary peak alignment between the target signal and the reference signal.
4. The multi-source periodic audio signal processing method according to claim 3, characterized in that: The signal alignment of the audio signal with a single cycle as the minimum processing unit also includes: After the target signal and the reference signal are initially peak aligned, The first audio signal and the second audio signal are divided into a plurality of first windows according to their respective periods, each of which includes a complete period segment; For the first window, the signal is divided twice using a second window having a length smaller than that of the first window to obtain a plurality of second window segments in the first window; For the second window, calculating the short-time average amplitude of the signal in the second window based on the signal amplitudes of all samples in the second window; The short-time average amplitude of the second window is weighted, and after weighted fusion, a weighted average amplitude value of the first window is obtained; a window weighted average amplitude sequence of the first audio signal and a window weighted average amplitude sequence of the second audio signal are obtained; Based on the window weighted average amplitude sequence of the first audio signal and the window weighted average amplitude sequence of the second audio signal, periodic window alignment is performed. The alignment process is: with the first audio signal fixed, the time axis of the second audio signal is gradually adjusted by time shifting, so that the error of the window weighted average amplitude sequence of the two is minimized. The error of the window weighted average amplitude sequence of the two is Wherein, M is the number of periodic windows of the aligned segment after the first audio signal and the second audio signal are aligned, and k is the kth periodic window in the aligned segment.
5. The multi-source periodic audio signal processing method according to claim 4, characterized in that: The short-time average amplitude of the second window is weighted, and a large weight is given to the second window containing the peak value, that is, the closer the second window is to the peak value, the larger the weight is.
6. The multi-source periodic audio signal processing method according to claim 1, characterized in that: The performing noise reduction and DC offset removal processing on the first audio signal and the second audio signal includes: Perform wavelet transform on the signal to obtain wavelet transform coefficients; Based on a preset wavelet coefficient threshold, removing the portion of the wavelet coefficient whose amplitude is lower than the preset wavelet coefficient threshold; Perform inverse wavelet transform to restore the denoised signal; Calculate the mean of the signal and subtract it from the signal; Center-clips the signal, removing the portion that exceeds a preset signal amplitude threshold.
7. The multi-source periodic audio signal processing method according to claim 1, characterized in that: The method of resampling the first audio signal and the second audio signal by using an interpolation algorithm, wherein the sampling frequency after resampling is the least common multiple of the sampling rates of the first audio signal and the second audio signal, comprises: Get the least common multiple f of the sampling rates of the first audio signal and the second audio signal LCM ; The first audio signal and the second audio signal are resampled based on a B-spline interpolation algorithm.
8. A multi-source periodic audio signal processing device, characterized in that: include: An audio signal acquisition unit, used to acquire multiple collected audio signals, take one of the multiple audio signals as a reference signal and record it as a first audio signal, and record the remaining signals as signals to be processed and record them as second audio signals; A pre-processing unit, configured to perform noise reduction and DC offset removal processing on the first audio signal and the second audio signal; A resampling unit, configured to resample the first audio signal and the second audio signal by using an interpolation algorithm, wherein the sampling frequency after resampling is the least common multiple of the sampling rates of the first audio signal and the second audio signal; A signal feature acquisition unit, configured to extract audio features from the first audio signal and the second audio signal, including extracting the main frequency of the audio signal; A signal alignment unit, configured to determine the period of the audio signal based on the main frequency of the audio signal, and perform signal alignment on the audio signal with a single period as the minimum processing unit, wherein the signal alignment includes alignment based on a peak signal and alignment based on a short-term average amplitude value of a period; The signal standardization unit is used to standardize the amplitudes of the two aligned periodic audio signals to ensure that the amplitude ranges of the two signals are consistent.
9. The multi-source periodic audio signal processing device according to claim 1, characterized in that: The signal alignment unit includes: a preliminary peak alignment unit and an optimization alignment unit. The preliminary peak alignment unit is used to perform the following steps: For the first audio signal and the second audio signal, respectively, determine a period T0 according to the main frequency f0 of the audio signal, find the period start and end points of the first audio signal and the second audio signal in the time domain according to the period T0, and respectively extract a complete period segment of the first audio signal and a complete period segment of the second audio signal; For the first audio signal and the second audio signal, respectively obtain peak points in each cycle, form a peak sequence of the first audio signal as a first peak sequence, and form a peak sequence of the second audio signal as a second peak sequence, wherein the peak points in each cycle are determined based on an extreme value method; calculating a preliminary time shift of the first audio signal and the second audio signal based on the first peak sequence and the second peak sequence, Where N is the number of peak signals, t ref,i is the time of the i-th peak of the reference signal, i.e., the first audio signal, ref represents the reference signal, t target,i represents the time of the i-th peak of the signal to be processed, i.e., the second audio signal, and target represents the target signal; According to the calculated preliminary time offset Δt peak , apply a time offset, i.e., a time axis translation, to the target signal to perform preliminary peak alignment between the target signal and the reference signal The optimization alignment unit is used to perform the following steps: After the target signal and the reference signal are preliminarily peak aligned, the first audio signal and the second audio signal are divided into a plurality of first windows according to their respective periods, each of the first windows including a complete period segment; For the first window, the signal is divided twice using a second window having a length smaller than that of the first window to obtain a plurality of second window segments in the first window; For the second window, calculating the short-time average amplitude of the signal in the second window based on the signal amplitudes of all samples in the second window; The short-time average amplitude of the second window is weighted, and after weighted fusion, a weighted average amplitude value of the first window is obtained; a window weighted average amplitude sequence of the first audio signal and a window weighted average amplitude sequence of the second audio signal are obtained; Based on the window weighted average amplitude sequence of the first audio signal and the window weighted average amplitude sequence of the second audio signal, periodic window alignment is performed. The alignment process is: with the first audio signal fixed, the time axis of the second audio signal is gradually adjusted by time shifting, so that the error of the window weighted average amplitude sequence of the two is minimized. The error of the window weighted average amplitude sequence of the two is Wherein, M is the number of periodic windows of the aligned segment after the first audio signal and the second audio signal are aligned, and k is the kth periodic window in the aligned segment.
10. A computer-readable storage medium having computer instructions stored thereon, characterized in that: When the instructions are executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Data processing method, device and equipment and storage medium
CN112086095A
Multi-channel audio processing method, reading method, audio device and readable storage medium
CN118509770A
Synchronization of audio signals from distributed devices
US10743107B1
360-degree multi-source location detection, tracking and enhancement
US20190355373A1
Method For Time Aligning In-Band On-Channel Digital Radio Audio With FM Radio Audio
US20250015911A1