A plug-in modular digital audio processing device and its use method
Through plug-in modular design and frequency shift compensation algorithm, the flexibility and multi-microphone management problems of audio equipment are solved, high-quality audio processing and stability are achieved, eliminating howling phenomena and improving user experience.
Patent Information
- Application Number
- CN202510771637.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-11
- Publication Date
- 2025-08-22
- Estimated Expiration
- 2045-06-11
AI Technical Summary
Existing audio data processing equipment cannot be flexibly adjusted and expanded, and it is difficult to effectively manage multiple microphone signals, resulting in sound interference and howling problems, affecting audio quality and user experience.
It adopts a modular design of plug-in, including a plug-in expansion unit, an audio processing algorithm storage unit, an audio digital signal collection and processing unit and a microphone intelligent management unit. It uses frequency shift compensation algorithm and stability factors to analyze sound interference to achieve full opening of microphone without whistling.
It realizes modular flexibility and scalability, provides high-quality and stable audio processing effects, eliminates sound interference, and improves user experience.
Smart Images

Figure CN120299473B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of digital audio processing, and in particular to a plug-in modular digital audio processing device and a method for using the same. Background Art
[0002] A digital audio processor is a digital audio signal processing device. It converts multi-channel analog input signals into digital signals and then applies a series of tunable algorithms to these digital signals to meet application requirements such as improving sound quality, matrix mixing, noise cancellation, echo cancellation, and feedback cancellation. With the prevalence of modern communication methods such as video conferencing and teleconferencing, audio quality has become a critical factor affecting meeting effectiveness and user experience. In multi-microphone scenarios, effectively managing multiple microphone signals and processing complex audio data presents technical challenges.
[0003] Most current audio data processing devices use fixed configurations, making them difficult to flexibly adjust or expand based on actual needs. Furthermore, their functional scalability and system upgradeability are poor. Traditional audio data processing devices typically use fixed configurations, making it difficult to flexibly adjust or expand their functionality based on actual needs. Adding new audio processing functions or improving performance often requires replacing the entire device, increasing costs and maintenance difficulties. Furthermore, in multi-microphone applications, devices in related technologies often struggle to effectively manage multiple microphones, leading to acoustic interference between microphones and, in turn, howling. This not only impacts audio quality but also reduces the overall user experience.
[0004] Currently, no effective solutions have been proposed for the problems in related technologies. Summary of the Invention
[0005] In view of this, the present invention provides a plug-in card modular digital audio processing device and a method of use to solve the above-mentioned problems.
[0006] In order to solve the above problems, the specific technical solutions adopted by the present invention are as follows:
[0007] According to one aspect of the present invention, there is provided a plug-in modular digital audio processing device, comprising:
[0008] A plug-in expansion unit is used to provide a standardized card slot for inserting a preset hardware expansion card;
[0009] An audio processing algorithm storage unit, used to store multiple audio processing algorithms and feature labels corresponding to the audio processing algorithms;
[0010] An audio digital signal collection and processing unit, configured to collect audio signal data based on the plug-in expansion unit and process the audio signal data using an audio processing algorithm to obtain a sound signal from the microphone;
[0011] The microphone intelligent management unit is used to manage multiple microphones simultaneously and analyze sound interference using a frequency shift compensation algorithm with a stability factor to ensure that the microphones are fully open without howling.
[0012] Preferably, the audio digital signal collection and processing unit includes:
[0013] An audio collection and preprocessing module, configured to collect audio signal data based on the plug-in expansion unit and preprocess the collected audio signal data;
[0014] A processing algorithm selection module is used to match and select an audio processing algorithm in an audio processing algorithm storage unit based on the preprocessed audio signal data according to a combined feature analysis method;
[0015] The audio processing module is used to perform audio processing on the pre-processed audio signal data based on the selected audio processing algorithm to obtain the sound signal of the microphone.
[0016] Preferably, the matching and selecting of the audio processing algorithm in the audio processing algorithm storage unit based on the pre-processed audio signal data according to the combined feature analysis method includes:
[0017] Performing frame processing on the preprocessed audio signal data, and multiplying each frame of the audio signal after framing by a preset window function to perform windowing processing;
[0018] Perform fast Fourier transform on each frame of the windowed signal to obtain the spectrum of each frame of the audio signal, and extract the Mel frequency cepstral coefficient features and gammatone frequency cepstral coefficient features;
[0019] Based on the Mel-frequency cepstral coefficient features and the gammatone-frequency cepstral coefficient features, the support vector machine recursive feature elimination method is used to match and select the optimal audio processing algorithm.
[0020] Preferably, the matching and selection of the optimal audio processing algorithm based on the Mel-frequency cepstral coefficient feature and the gammatone frequency cepstral coefficient feature using the support vector machine recursive feature elimination method includes:
[0021] Mel-frequency cepstral coefficient features and gammatone frequency cepstral coefficient features are standardized respectively, and the support vector machine recursive feature elimination method is used to remove redundant and irrelevant features to obtain feature subsets;
[0022] Based on the preset classification model, the feature subset is classified to obtain the classification label, and the similarity between the classification label and the feature label is calculated. The audio processing algorithm is selected according to the similarity calculation result.
[0023] Preferably, the microphone intelligent management unit includes:
[0024] A feature analysis module is used to collect the sound signals of each microphone and analyze the interference characteristics between the sound signals of different microphones using short-time Fourier transform of the dichotomy method;
[0025] The signal compensation module is used to compensate for the interference signals between microphones based on the interference characteristics between the sound signals of different microphones, using a frequency shift compensation algorithm and combining a stability factor to obtain compensated multi-channel microphone sound signals;
[0026] The signal output module is used to output the compensated multi-channel microphone signals in real time to achieve full microphone opening without howling.
[0027] Preferably, the collecting of sound signals from each microphone and analyzing interference characteristics between sound signals from different microphones using a short-time Fourier transform with a dichotomy method includes:
[0028] Use short-time Fourier transform to perform time-frequency analysis on each microphone signal and generate a time-frequency graph.
[0029] Based on the generated time-frequency graph, the frequency range in the time-frequency graph is divided using the dichotomy method to obtain two groups of sub-frequency bands;
[0030] For each sub-band, select the signals from two different microphones and calculate their cross-correlation function within the sub-band. Based on the calculation results, determine whether there is an interference feature in each sub-band;
[0031] For sub-frequency bands with interference features, interference feature extraction is performed to obtain interference features. For sub-frequency bands without interference features, the dichotomy method is used again to divide the sub-frequency bands without interference features, and interference feature extraction is performed again until the preset conditions are met.
[0032] Preferably, the method of compensating the interference signals between microphones by using a frequency shift compensation algorithm based on the interference characteristics between the sound signals of different microphones and combining the stability factor to obtain the compensated multi-channel microphone sound signals includes:
[0033] Calculate the horizontal sound path difference, frequency shift and stability factor between each pair of microphone combinations based on the interference characteristics between the microphone sound signals;
[0034] Determine the waveguide invariant according to the characteristics of the propagation environment, and calculate the time delay compensation between each pair of microphone combinations according to the waveguide invariant and the horizontal sound path difference;
[0035] Perform frequency shifting on the signals received by each pair of microphones according to the frequency shift amount, and adjust the phase of the frequency-shifted signals according to the calculated delay compensation amount to obtain a corrected signal.
[0036] Based on the stability factor, weighted compensation processing is performed on the correction signal to obtain a compensated microphone sound signal.
[0037] Preferably, performing a frequency shift operation on the signal received by each pair of microphone combinations according to the frequency shift amount, and performing phase adjustment on the frequency-shifted signal according to the calculated delay compensation amount to obtain a corrected signal includes:
[0038] The signal received by each pair of microphones is transformed using the wavelet transform method to obtain a frequency domain signal;
[0039] Based on the frequency shift amount, the frequency of the frequency domain signal is adjusted, and the frequency shift operation is performed on the frequency domain signal using interpolation technology;
[0040] The phase of the frequency-shifted signal is corrected using the delay compensation amount to obtain a corrected signal.
[0041] Preferably, the performing weighted compensation processing on the correction signal based on the stability factor to obtain the compensated microphone sound signal includes:
[0042] The stability factor is used as the weighting coefficient of each microphone and the weighting coefficient is normalized;
[0043] The normalized weighted coefficient is used to perform weighted processing on the corrected signal of the microphone to obtain a compensated microphone sound signal.
[0044] According to another aspect of the present invention, a method for using a plug-in modular digital audio processing device is provided, comprising the following steps:
[0045] S1. Insert the preset hardware expansion card into the standardized card slot provided by the plug-in expansion unit;
[0046] S2. collecting audio signal data based on the plug-in expansion unit and processing the audio signal data using an audio processing algorithm to obtain a sound signal from the microphone;
[0047] S3. Manage multiple microphones simultaneously and use the frequency shift compensation algorithm of the stability factor to analyze the acoustic interference to achieve no howling when the microphones are fully opened.
[0048] The beneficial effects of the present invention are:
[0049] 1. The present invention has modular flexibility and scalability, rich audio processing algorithm support, efficient audio digital signal collection and processing, and intelligent microphone management to achieve full-open without howling, etc., which enables the device to adapt to the audio processing needs in different scenarios, provide high-quality, stable and reliable audio processing effects, and enhance the user's overall experience.
[0050] 2. The present invention uses framing, windowing, fast Fourier transform and feature extraction, combined with the support vector machine recursive feature elimination method, to accurately remove redundant features and obtain feature subsets. It then calculates the similarity between the classification model and the preset feature labels, intelligently matches and selects the optimal audio processing algorithm, thereby ensuring that the audio signal is processed in a targeted and optimized manner, significantly improving the accuracy and efficiency of audio processing.
[0051] 3. The present invention uses the short-time Fourier transform of the dichotomy method to deeply analyze the interference characteristics between the sound signals of different microphones and accurately locate the interference frequency band; then combines the frequency shift compensation algorithm with the stability factor to perform frequency shift, time delay compensation and phase adjustment on the interference signal, effectively eliminating the sound interference phenomenon; finally, through the weighted compensation processing of the stability factor, the signal quality is optimized to ensure that there is no howling when multiple microphones are fully turned on, which significantly improves the stability and audio quality of the audio system. BRIEF DESCRIPTION OF THE DRAWINGS
[0052] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative work. In the drawings:
[0053] Figure 1 This is a principle block diagram of a plug-in modular digital audio processing device according to an embodiment of the present invention;
[0054] Figure 2 The present invention is a flowchart of a method for using a plug-in modular digital audio processing device according to an embodiment of the present invention.
[0055] In the picture:
[0056] 1. Plug-in expansion unit; 2. Audio processing algorithm storage unit; 3. Audio digital signal collection and processing unit; 4. Microphone intelligent management unit. DETAILED DESCRIPTION
[0057] In order to enable those skilled in the art to better understand the technical solutions in this application, the following will clearly and completely describe the technical solutions in the embodiments of this application in conjunction with the drawings in the embodiments of this application. Obviously, the described embodiments are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments of this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of this application.
[0058] According to an embodiment of the present invention, a plug-in card modular digital audio processing device and a method of using the same are provided.
[0059] Specifically, the plug-in modular digital audio processing device can use WM-CK300, which is a distributed network audio system processing service center. It is an enterprise-level core AI audio algorithm processing platform that integrates powerful hardware processing capabilities and rich and powerful audio algorithms. It provides 128x128 channels of ultra-low latency, uncompressed network audio input, output and processing capabilities, supports 16x16 RTSP network audio streams, and the audio matrix routing can reach 144*144. It is extensible and compatible with DANTE, AES67 and other protocols, and supports a maximum of 128x128 channels.
[0060] The WM-CK300 adopts a configurable plug-in card design with a total of 9 card slots, which can be flexibly configured with 8 analog input and output cards, and an optional DANTE expansion card. At the same time, the system functions are software-defined, and the audio processing functions and system architecture can be customized. It supports environmental control management and can centrally manage network microphones, network speakers, audio nodes and other devices.
[0061] It adopts a 1+1 system hot standby design and supports two AC power supplies at the same time. It can demonstrate the reliability of the system for applications of any scale and is a solution with excellent performance and reliability.
[0062] With its large network audio channel access capability and powerful processing functions, it is suitable for a variety of large-scale application venues: large command centers, conference centers, administrative centers, corporate headquarters buildings, theme parks, transportation hubs, etc.
[0063] The present invention will now be further described with reference to the accompanying drawings and specific embodiments. Figure 1 As shown, according to one embodiment of the present invention, a plug-in modular digital audio processing device is provided, comprising:
[0064] A plug-in expansion unit 1, used to provide a standardized card slot for inserting a preset hardware expansion card;
[0065] Specifically, the standardized card slots adhere to uniform physical dimensions and electrical interface standards, ensuring that expansion cards with different functions can be easily and quickly inserted and removed. Users can insert expansion cards with different functions as needed, including the 128-channel Dante expansion card Amu 128D, the 4-channel mic / line input card Amu 4I, the 4-channel line output card Amu 4O, and more.
[0066] The above expansion card contains multi-channel audio signal acquisition capabilities and can collect signals from multiple microphones or audio sources at the same time.
[0067] Audio processing algorithm storage unit 2, used to store multiple audio processing algorithms and feature labels corresponding to the audio processing algorithms;
[0068] It should be noted that the audio processing algorithm storage unit 2 stores hundreds of rich audio processing algorithms, supports 5A audio algorithm, AI denoising, AI automatic gain and other commonly used audio algorithms, threshold automatic mixing, gain sharing automatic mixing, dynamic compressor, etc.
[0069] The audio digital signal collection and processing unit 3 is used to collect audio signal data based on the plug-in expansion unit and process the audio signal data using an audio processing algorithm to obtain a sound signal from the microphone;
[0070] As a preferred embodiment, the audio digital signal collection and processing unit 3 includes:
[0071] An audio collection and preprocessing module, configured to collect audio signal data based on the plug-in expansion unit and preprocess the collected audio signal data;
[0072] It should be noted that the preprocessing operations include: denoising, gain control, sampling rate conversion, etc.
[0073] Denoising uses filtering algorithms (such as high-pass and low-pass filtering) to remove ambient noise or device noise. Gain control dynamically adjusts signal amplitude to avoid clipping or distortion while improving the signal-to-noise ratio of low-level signals. Sampling rate conversion unifies the signal sampling rate (such as 48kHz) to ensure consistency in algorithm input.
[0074] A processing algorithm selection module is used to match and select an audio processing algorithm in an audio processing algorithm storage unit based on the preprocessed audio signal data according to a combined feature analysis method;
[0075] As a preferred embodiment, the matching and selecting of the audio processing algorithm in the audio processing algorithm storage unit based on the preprocessed audio signal data according to the combined feature analysis method includes:
[0076] Performing frame processing on the preprocessed audio signal data, and multiplying each frame of the audio signal after framing by a preset window function to perform windowing processing;
[0077] Specifically, audio signals change continuously over time. In order to perform short-time spectrum analysis, they need to be divided into multiple short-time frames. Frame processing can ensure that the audio signal is approximately stable within each frame, thereby meeting the prerequisites for Fourier transform.
[0078] Specifically, the frame length is typically set to 20-40ms (e.g., 32ms) to ensure the signal is sufficiently stable within the frame. The frame shift (the amount of overlap between adjacent frames) is typically set to half the frame length (e.g., 16ms) to increase signal continuity and reduce spectral leakage. The sliding window technique is then used to segment the continuous audio signal into multiple frames of equal length.
[0079] Windowing can reduce spectrum leakage and improve the accuracy of spectrum analysis. The window function smoothes the edges of the signal and reduces the spectrum sidelobes caused by signal truncation.
[0080] Perform fast Fourier transform on each frame of the windowed signal to obtain the spectrum of each frame of the audio signal, and extract the Mel frequency cepstral coefficient features and gammatone frequency cepstral coefficient features;
[0081] Among them, the Fast Fourier Transform (FFT) can convert the time domain signal into the frequency domain signal, thereby obtaining the spectral characteristics of the audio signal. Spectral analysis is the basis for extracting audio features. Specifically, it includes: performing an FFT transform on each frame of the windowed signal to obtain a complex spectrum, and then taking the modulus value of the FFT result to obtain the amplitude spectrum;
[0082] Based on the Mel-frequency cepstral coefficient features and the gammatone-frequency cepstral coefficient features, the support vector machine recursive feature elimination method is used to match and select the optimal audio processing algorithm.
[0083] As a preferred embodiment, the method of matching and selecting the optimal audio processing algorithm based on the Mel-frequency cepstral coefficient features and the gammatone frequency cepstral coefficient features using the support vector machine recursive feature elimination method includes:
[0084] Mel-frequency cepstral coefficient features and gammatone frequency cepstral coefficient features are standardized respectively, and the support vector machine recursive feature elimination method is used to remove redundant and irrelevant features to obtain feature subsets;
[0085] It should be noted that Mel-frequency cepstral coefficient features and Gammatone-frequency cepstral coefficient features have different dimensions and value ranges. Directly using them can cause certain features to dominate model training, affecting the accuracy of feature selection and algorithm matching. Therefore, normalization can eliminate the dimensionality effect and bring all features to the same scale. Normalization includes Z-Score normalization and Min-Max normalization.
[0086] Among them, the Mel-frequency cepstral coefficient features and the gammatone frequency cepstral coefficient features also contain redundant or irrelevant information, which increases the computational complexity and may reduce the accuracy of the algorithm matching. The support vector machine recursive feature elimination method removes redundant and irrelevant features by iteratively training the SVM model and eliminating features with smaller weights, which can effectively remove redundant and irrelevant features. Specifically, it includes the following steps:
[0087] The standardized Mel-frequency cepstral coefficient features and gammatone frequency cepstral coefficient features are merged into a high-dimensional feature vector as the initial feature set. The current feature set is then used to train the SVM classifier. Based on the decision function coefficient (or gradient) of the SVM classifier, the weight of each feature is calculated, the feature with the smallest weight (or several features with the smallest weight) is removed, and the feature set is updated. The iteration is stopped when the preset number of features is reached, the model performance no longer improves, or all features are eliminated.
[0088] Based on the preset classification model, the feature subset is classified to obtain the classification label, and the similarity between the classification label and the feature label is calculated. The audio processing algorithm is selected according to the similarity calculation result.
[0089] Specifically, based on a preset classification model, the feature subset is classified, and similarity is calculated between the sub-labels and the feature labels. The audio processing algorithm is selected according to the similarity calculation result, including the following steps:
[0090] First, a random forest classification model is trained using a labeled audio dataset (containing Mel-frequency cepstral coefficient features, Gammatone frequency cepstral coefficient features, and corresponding classification labels);
[0091] Then, the feature subset obtained by SVM-RFE is input into the trained random forest classification model to obtain the classification label.
[0092] Then use the cosine similarity, Euclidean distance or Jaccard similarity algorithm to calculate the similarity between the classification label and the feature label. Finally, based on the similarity calculation results, the audio processing algorithm corresponding to the feature label with the highest similarity is selected as the optimal algorithm.
[0093] The audio processing module is used to perform audio processing on the pre-processed audio signal data based on the selected audio processing algorithm to obtain the sound signal of the microphone.
[0094] Specifically, the audio signal data is input into the audio processing algorithm and processed according to the logical flow of the audio processing algorithm. The processing process includes filtering, noise reduction, gain adjustment, feature extraction and other operations.
[0095] The microphone intelligent management unit 4 is used to manage multiple microphones simultaneously and analyze the sound interference using a frequency shift compensation algorithm of a stability factor to achieve the goal of no howling when the microphones are fully opened.
[0096] As a preferred embodiment, the microphone intelligent management unit 4 includes:
[0097] A feature analysis module is used to collect the sound signals of each microphone and analyze the interference characteristics between the sound signals of different microphones using short-time Fourier transform of the dichotomy method;
[0098] As a preferred embodiment, the collecting of sound signals from each microphone and analyzing interference characteristics between sound signals from different microphones using a short-time Fourier transform with a dichotomy method includes:
[0099] Use short-time Fourier transform to perform time-frequency analysis on each microphone signal and generate a time-frequency graph.
[0100] It should be noted that the time-frequency characteristics of each microphone signal are analyzed by short-time Fourier transform (STFT) to generate a time-frequency diagram, which provides a basis for subsequent interference feature analysis.
[0101] Based on the generated time-frequency graph, the frequency range in the time-frequency graph is divided using the dichotomy method to obtain two groups of sub-frequency bands;
[0102] It should be noted that before dividing the frequency range in the time-frequency graph into multiple sub-bands using the bisection method, it is necessary to determine the frequency interval to be analyzed based on the frequency range of the time-frequency graph, then bisection the determined frequency interval to obtain two sub-bands. The frequency range of each sub-band is recorded.
[0103] For each sub-band, select the signals from two different microphones and calculate their cross-correlation function within the sub-band. Based on the calculation results, determine whether there is interference feature in each sub-band;
[0104] It should be noted that the time-frequency representation of the two selected microphone signals in the sub-band is cross-correlated and the similarity of the two signals in time delay can be determined. x ( t )and y ( t ), whose cross-correlation function R xy ( τ ) in the delay τ The calculation formula at is:
[0105] ;
[0106] For sub-frequency bands with interference features, interference feature extraction is performed to obtain interference features. For sub-frequency bands without interference features, the dichotomy method is used again to divide the sub-frequency bands without interference features, and interference feature extraction is performed again until the preset conditions are met.
[0107] Specifically, feature extraction is performed on sub-bands with interference features, and sub-bands without interference features are further divided until all sub-bands are fully analyzed. Specific steps:
[0108] For sub-frequency bands with interference features, the interference features, such as interference delay, interference amplitude, and interference phase, are extracted. For sub-frequency bands without interference features, the binary division method is used again to obtain smaller sub-frequency bands. The above steps are repeated to calculate the cross-correlation function and determine the interference features for the new sub-frequency bands.
[0109] The preset conditions include the minimum frequency width of the sub-band, the maximum number of divisions, the confidence level of the interference feature extraction, etc. When the preset conditions are met, the division and interference feature extraction process is stopped.
[0110] The signal compensation module is used to compensate for the interference signals between microphones based on the interference characteristics between the sound signals of different microphones, using a frequency shift compensation algorithm and combining a stability factor to obtain compensated multi-channel microphone sound signals;
[0111] As a preferred embodiment, the interference signals between microphones are compensated by using a frequency shift compensation algorithm based on the interference characteristics between the sound signals of different microphones and combining with a stability factor to obtain the compensated multi-channel microphone sound signals, including:
[0112] Calculate the horizontal sound path difference, frequency shift and stability factor between each pair of microphone combinations based on the interference characteristics between the microphone sound signals;
[0113] It should be noted that the horizontal sound path difference refers to the difference in distance between two microphones in the horizontal direction, which results in different arrival times of the sound signals. The horizontal sound path difference between each pair of microphones is calculated based on the geometric layout of the microphone array and the known microphone position information.
[0114] Frequency shift refers to the difference in signal frequency between different microphones due to relative motion between the sound source and the microphones or multipath propagation. The frequency shift between each pair of microphone signals is calculated using methods such as cross-correlation analysis and Doppler shift estimation.
[0115] The stability factor measures the stability of the microphone signal's interference signature, specifically whether the interference pattern changes over time. The stability factor is calculated by analyzing signal characteristics such as correlation and power spectral density to assess the stability of the interference signature.
[0116] Determine the waveguide invariant according to the characteristics of the propagation environment, and calculate the time delay compensation between each pair of microphone combinations according to the waveguide invariant and the horizontal sound path difference;
[0117] It should be noted that the waveguide invariant is a parameter related to the propagation characteristics of sound waves in a waveguide (such as air, water, or other media). The waveguide invariant of the propagation environment can then be determined through experimental measurement or theoretical calculation. The formula for calculating the time delay compensation between each microphone pair based on the waveguide invariant and the horizontal acoustic path difference is:
[0118] ;
[0119] Where, τ (Δ r , r j ) represents the delay compensation amount, β s and β p represents the parameters related to phase velocity and group velocity; Δ r represents the horizontal path difference between the two microphones, r j Indicates the j The distance of the microphone;
[0120] Perform frequency shifting on the signals received by each pair of microphones according to the frequency shift amount, and adjust the phase of the frequency-shifted signals according to the calculated delay compensation amount to obtain a corrected signal.
[0121] As a preferred embodiment, the method of performing a frequency shift operation on the signal received by each pair of microphone combinations according to the frequency shift amount, and performing phase adjustment on the frequency-shifted signal according to the calculated delay compensation amount to obtain a corrected signal includes:
[0122] The signal received by each pair of microphones is transformed using the wavelet transform method to obtain a frequency domain signal;
[0123] Specifically, the time domain signal received by each pair of microphone combinations is transformed continuously by using wavelet basis functions (such as Daubechies wavelet, Morlet wavelet, etc.) to obtain the representation of the signal at different scales and frequencies, and then the frequency domain signal is obtained by reconstructing the wavelet coefficients of a specific scale or frequency interval.
[0124] Based on the frequency shift amount, the frequency of the frequency domain signal is adjusted, and the frequency shift operation is performed on the frequency domain signal using interpolation technology;
[0125] Specifically, in the frequency domain, frequency shift is achieved by shifting the spectrum of the frequency domain signal. Specifically, each frequency component of the frequency domain signal is multiplied by a complex exponential factor ( f is the frequency shift, t is the time variable, j is an imaginary unit); and then linear interpolation is used to perform frequency shift operation on the frequency domain signal.
[0126] The phase of the frequency-shifted signal is corrected using the delay compensation amount to obtain a corrected signal.
[0127] It should be noted that the delay compensation amount represents the time difference of the sound signal propagating between different microphones. In the frequency domain, the delay compensation amount can be converted into a phase adjustment amount. For each frequency component in the frequency domain f , phase adjustment amount It can be calculated by the following formula:
[0128] ;
[0129] For each frequency component in the frequency domain f , the corrected frequency component is multiplied by a phase factor (in f is the frequency shift, τ is the time delay, j is an imaginary unit), the corrected signal can be obtained.
[0130] Based on the stability factor, weighted compensation processing is performed on the correction signal to obtain a compensated microphone sound signal.
[0131] As a preferred embodiment, the weighted compensation processing is performed on the correction signal based on the stability factor to obtain the compensated microphone sound signal, which includes:
[0132] The stability factor is used as the weighting coefficient of each microphone, and the weighting coefficient is normalized; the normalization process is to ensure the rationality and consistency of the weighting coefficient.
[0133] The normalized weighted coefficient is used to perform weighted processing on the corrected signal of the microphone to obtain a compensated microphone sound signal.
[0134] It should be noted that at each time point (or frequency point), the corrected signal from each microphone is multiplied by its corresponding normalized weighting coefficient. All weighted signals are then summed to produce the compensated sound signal. After weighting, the resulting multi-microphone sound signals are combined into a single, high-quality signal. This signal fully utilizes the information from multiple microphones, improving the signal-to-noise ratio and reducing the effects of interference and noise.
[0135] The signal output module is used to output the compensated multi-channel microphone signals in real time to achieve full microphone opening without howling.
[0136] like Figure 2 According to another embodiment of the present invention, a method for using a plug-in modular digital audio processing device is provided, comprising the following steps:
[0137] S1. Insert the preset hardware expansion card into the standardized card slot provided by the plug-in expansion unit;
[0138] S2. collecting audio signal data based on the plug-in expansion unit and processing the audio signal data using an audio processing algorithm to obtain a sound signal from the microphone;
[0139] S3. Manage multiple microphones simultaneously and use the frequency shift compensation algorithm of the stability factor to analyze the acoustic interference to achieve no howling when the microphones are fully opened.
[0140] In summary, with the help of the above technical solutions of the present invention, the present invention has modular flexibility and scalability, rich audio processing algorithm support, efficient audio digital signal collection and processing, and microphone intelligent management to achieve full-open without howling and other functions, which can enable the device to adapt to the audio processing needs in different scenarios, provide high-quality, stable and reliable audio processing effects, and enhance the overall user experience. The present invention can accurately remove redundant features and obtain feature subsets through framing, windowing, fast Fourier transform and feature extraction, combined with the support vector machine recursive feature elimination method, and then intelligently match and select the optimal audio processing algorithm through the similarity calculation between the classification model and the preset feature label, thereby ensuring that the audio signal is processed in a targeted and optimized manner, significantly improving the accuracy and efficiency of audio processing. The present invention uses the short-time Fourier transform of the dichotomy method to deeply analyze the interference characteristics between the sound signals of different microphones and accurately locate the interference frequency band; then combines the frequency shift compensation algorithm with the stability factor to perform frequency shift, time delay compensation and phase adjustment on the interference signal, effectively eliminating the acoustic interference phenomenon; finally, through the weighted compensation processing of the stability factor, the signal quality is optimized to ensure that there is no howling when multiple microphones are fully turned on, which significantly improves the stability and audio quality of the audio system.
[0141] Those skilled in the art will appreciate that embodiments of the present invention may be provided as methods, systems, or computer program products. Thus, the present invention may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, optical storage, etc.) containing computer-usable program code.
[0142] The specific embodiments described above further illustrate the objectives, technical solutions and beneficial effects of the present invention in detail. It should be understood that the above description is only a specific embodiment of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.
Claims
1. A plug-in modular digital audio processing device, characterized in that: include: A plug-in expansion unit is used to provide a standardized card slot for inserting a preset hardware expansion card; An audio processing algorithm storage unit, used to store multiple audio processing algorithms and feature labels corresponding to the audio processing algorithms; An audio digital signal collection and processing unit, configured to collect audio signal data based on the plug-in expansion unit and process the audio signal data using an audio processing algorithm to obtain a sound signal from the microphone; The microphone intelligent management unit is used to manage multiple microphones simultaneously and analyze acoustic interference using a frequency shift compensation algorithm with a stability factor to achieve full microphone operation without howling. It includes: The feature analysis module is used to collect the sound signals of each microphone, perform time-frequency analysis on each microphone signal separately using short-time Fourier transform, and generate a time-frequency graph; based on the generated time-frequency graph, the frequency range in the time-frequency graph is divided into two groups of sub-frequency bands using the dichotomy method; for each sub-frequency band, the signals of two different microphones are selected, and their cross-correlation functions within the sub-frequency band are calculated. Based on the calculation results, it is determined whether there is an interference feature in each sub-frequency band; for sub-frequency bands with interference features, interference features are extracted to obtain interference features. For sub-frequency bands without interference features, the dichotomy method is used again to divide the sub-frequency bands without interference features, and interference features are extracted again until the preset conditions are met; The signal compensation module is used to compensate for the interference signals between microphones based on the interference characteristics between the sound signals of different microphones, using a frequency shift compensation algorithm and combining a stability factor to obtain compensated multi-channel microphone sound signals; The signal output module is used to output the compensated multi-channel microphone signals in real time to achieve full microphone opening without howling.
2. The plug-in modular digital audio processing device according to claim 1, characterized in that: The audio digital signal collection and processing unit includes: An audio collection and preprocessing module, configured to collect audio signal data based on the plug-in expansion unit and preprocess the collected audio signal data; A processing algorithm selection module is used to match and select an audio processing algorithm in an audio processing algorithm storage unit based on the preprocessed audio signal data according to a combined feature analysis method; The audio processing module is used to perform audio processing on the pre-processed audio signal data based on the selected audio processing algorithm to obtain the sound signal of the microphone.
3. The plug-in modular digital audio processing device according to claim 2, characterized in that: The method of matching and selecting an audio processing algorithm in an audio processing algorithm storage unit based on the pre-processed audio signal data according to a combined feature analysis method includes: Performing frame processing on the preprocessed audio signal data, and multiplying each frame of the audio signal after framing by a preset window function to perform windowing processing; Perform fast Fourier transform on each frame of the windowed signal to obtain the spectrum of each frame of the audio signal, and extract the Mel frequency cepstral coefficient features and gammatone frequency cepstral coefficient features; Based on the Mel-frequency cepstral coefficient features and the gammatone-frequency cepstral coefficient features, the support vector machine recursive feature elimination method is used to match and select the optimal audio processing algorithm.
4. The plug-in modular digital audio processing device according to claim 3, characterized in that: The method of matching and selecting the optimal audio processing algorithm based on the Mel-frequency cepstral coefficient features and the gammatone frequency cepstral coefficient features by using the support vector machine recursive feature elimination method includes: Mel-frequency cepstral coefficient features and gammatone frequency cepstral coefficient features are standardized respectively, and the support vector machine recursive feature elimination method is used to remove redundant and irrelevant features to obtain feature subsets; Based on the preset classification model, the feature subset is classified to obtain the classification label, and the similarity between the classification label and the feature label is calculated. The audio processing algorithm is selected according to the similarity calculation result.
5. The plug-in modular digital audio processing device according to claim 4, characterized in that: The method of compensating the interference signals between the microphones by using a frequency shift compensation algorithm based on the interference characteristics between the sound signals of different microphones and combining the stability factor to obtain the compensated multi-channel microphone sound signals includes: Calculate the horizontal sound path difference, frequency shift and stability factor between each pair of microphone combinations based on the interference characteristics between the microphone sound signals; Determine the waveguide invariant according to the characteristics of the propagation environment, and calculate the time delay compensation between each pair of microphone combinations according to the waveguide invariant and the horizontal sound path difference; Perform frequency shifting on the signals received by each pair of microphones according to the frequency shift amount, and adjust the phase of the frequency-shifted signals according to the calculated delay compensation amount to obtain a corrected signal. Based on the stability factor, weighted compensation processing is performed on the correction signal to obtain a compensated microphone sound signal.
6. The plug-in modular digital audio processing device according to claim 5, characterized in that: The step of performing a frequency shift operation on the signal received by each pair of microphone combinations according to the frequency shift amount, and performing phase adjustment on the frequency-shifted signal according to the calculated delay compensation amount to obtain a corrected signal includes: The signal received by each pair of microphones is transformed using the wavelet transform method to obtain a frequency domain signal; Based on the frequency shift amount, the frequency of the frequency domain signal is adjusted, and the frequency shift operation is performed on the frequency domain signal using interpolation technology; The phase of the frequency-shifted signal is corrected using the delay compensation amount to obtain a corrected signal.
7. The plug-in modular digital audio processing device according to claim 6, characterized in that: The weighted compensation processing is performed on the correction signal based on the stability factor to obtain the compensated microphone sound signal, which includes: The stability factor is used as the weighting coefficient of each microphone and the weighting coefficient is normalized; The normalized weighted coefficient is used to perform weighted processing on the corrected signal of the microphone to obtain a compensated microphone sound signal.
8. A method for using a plug-in card modular digital audio processing device, for implementing the use of the plug-in card modular digital audio processing device according to any one of claims 1 to 7, characterized in that: The following steps are involved: S1. Insert the preset hardware expansion card into the standardized card slot provided by the plug-in expansion unit; S2. collecting audio signal data based on the plug-in expansion unit and processing the audio signal data using an audio processing algorithm to obtain a sound signal from the microphone; S3. Manage multiple microphones simultaneously and use the frequency shift compensation algorithm of the stability factor to analyze the acoustic interference to achieve no howling when the microphones are fully opened.
Citation Information
Patent Citations
Method and device for improving audio processing performance
CN104980337A