Plug-in card type modular digital audio processing device and use method
Through plug-in modular design and intelligent microphone management, the flexibility of audio equipment and multi-microphone signal management problems are solved, high-quality audio processing and stability are achieved, eliminating howling phenomena and improving user experience.
Patent Information
- Application Number
- CN202510771637.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-11
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2045-06-11
AI Technical Summary
Existing audio data processing equipment cannot flexibly adjust and expand functions, and it is difficult to effectively manage multiple microphone signals, resulting in sound interference and howling problems, affecting audio quality and user experience.
The plug-in-card modular design adopts a plug-in-card expansion unit, an audio processing algorithm storage unit, an audio digital signal collection and processing unit and a microphone intelligent management unit. It uses frequency shift compensation algorithm and stability factors to analyze acoustic interference, and combines the support vector machine recursive feature elimination method and dichotomy short-time Fourier transform to achieve full opening of the microphone without screaming.
It realizes modular flexibility and scalability, provides high-quality audio processing effects, eliminates sound interference, and improves the stability and user experience of the audio system.
Smart Images

Figure CN120299473A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of digital audio processing. Specifically, it particularly relates to a plug-in modular digital audio processing device and its usage method. Background Art
[0002] A digital audio processor is a digital audio signal processing device. It first converts the analog signals input from multiple channels into digital signals, and then performs a series of tunable algorithmic processes on the digital signals to meet application requirements such as improving sound quality, matrix mixing, noise cancellation, echo cancellation, and feedback cancellation. With the popularization of modern communication methods such as video conferencing and telephone conferencing, audio quality has become a key factor affecting the meeting effect and user experience. Especially in the multi-microphone application scenario, effectively managing multiple microphone signals and processing complex audio data has become a technical problem.
[0003] Most current audio data processing devices adopt a fixed configuration, which cannot be flexibly adjusted according to actual needs, and have poor functional expandability and system upgradability. Traditional audio data processing devices usually adopt a fixed configuration, making it difficult to be flexibly adjusted or expand functions according to actual needs. When new audio processing functions need to be added or performance needs to be improved, the entire device often needs to be replaced, increasing costs and maintenance difficulties. Moreover, in the multi-microphone application scenario, the devices in related technologies often have difficulty effectively managing multiple microphones, resulting in acoustic interference between microphones and thus causing a howling problem. This not only affects the audio quality but also reduces the overall user experience.
[0004] For the problems in related technologies, no effective solutions have been proposed yet. Summary of the Invention
[0005] In view of this, the present invention provides a plug-in modular digital audio processing device and its usage method to solve the problems mentioned above.
[0006] To solve the above problems, the specific technical solutions adopted by the present invention are as follows: According to one aspect of the present invention, there is provided a plug-in modular digital audio processing device, including: A plug-in expansion unit for providing a standardized card slot for inserting a preset hardware expansion card; An audio processing algorithm storage unit for storing a variety of audio processing algorithms and the characteristic tags corresponding to the audio processing algorithms; An audio digital signal collection and processing unit for collecting audio signal data based on the plug-in expansion unit and processing the audio signal data using the audio processing algorithms to obtain the sound signals of the microphones; A microphone intelligent management unit is used to manage several microphones simultaneously and analyze acoustic interference using a frequency shift compensation algorithm with a stability factor to achieve non-squealing when all microphones are fully open.
[0007] Preferably, the audio digital signal collection and processing unit includes: An audio collection and preprocessing module is used to collect audio signal data based on a plug-in expansion unit and preprocess the collected audio signal data; A processing algorithm selection module is used to match and select an audio processing algorithm in the audio processing algorithm storage unit based on the preprocessed audio signal data according to the combined feature analysis method; An audio processing module is used to perform audio processing on the preprocessed audio signal data based on the selected audio processing algorithm to obtain the sound signal of the microphone.
[0008] Preferably, the matching and selecting of the audio processing algorithm in the audio processing algorithm storage unit based on the preprocessed audio signal data according to the combined feature analysis method includes: Perform frame division processing on the preprocessed audio signal data, and perform windowing processing on each frame of the divided audio signal by multiplying it by a preset window function; Perform a fast Fourier transform on each windowed frame signal to obtain the spectrum of each frame of audio signal, and extract the Mel frequency cepstral coefficient feature and the gamma-tone frequency cepstral coefficient feature; Based on the Mel frequency cepstral coefficient feature and the gamma-tone frequency cepstral coefficient feature, use the support vector machine recursive feature elimination method to match and select the optimal audio processing algorithm.
[0009] Preferably, the matching and selecting of the optimal audio processing algorithm based on the Mel frequency cepstral coefficient feature and the gamma-tone frequency cepstral coefficient feature using the support vector machine recursive feature elimination method includes: Perform standardization processing on the Mel frequency cepstral coefficient feature and the gamma-tone frequency cepstral coefficient feature respectively, and use the support vector machine recursive feature elimination method to remove redundant and irrelevant features; obtain a feature subset; Based on a preset classification model, classify the feature subset to obtain a classification label, calculate the similarity between the sub-label and the feature label, and select an audio processing algorithm according to the similarity calculation result.
[0010] Preferably, the microphone intelligent management unit includes: A feature analysis module is used to collect the sound signals of each microphone and analyze the interference characteristics between the sound signals of different microphones using the short-time Fourier transform of the dichotomy method; A signal compensation module, which is used to compensate the interference signal between microphones according to the interference characteristics between the sound signals of different microphones, using a frequency shift compensation algorithm and combining with a stability factor, so as to obtain the compensated multi-channel microphone sound signals; A signal output module, which is used to output the compensated multi-channel microphone signals in real time to achieve non-whistling when all microphones are fully open.
[0011] Preferably, collecting the sound signals of each microphone and using the short-time Fourier transform of the dichotomy method to analyze the interference characteristics between the sound signals of different microphones includes: Performing time-frequency analysis on each channel of microphone signal respectively by using the short-time Fourier transform to generate a time-frequency diagram, Based on the generated time-frequency diagram, dividing the frequency range in the time-frequency diagram by using the dichotomy method to obtain two groups of sub-frequency bands; For each sub-frequency band, selecting the signals of two different microphones and calculating their cross-correlation function within the sub-frequency band, and judging whether there are interference characteristics in each sub-frequency band according to the calculation results; For the sub-frequency bands with interference characteristics, extracting the interference characteristics to obtain the interference characteristics. For the sub-frequency bands without interference characteristics, using the dichotomy method to divide the sub-frequency bands without interference characteristics again and re-performing the interference characteristic extraction until the preset conditions are met.
[0012] Preferably, compensating the interference signal between microphones according to the interference characteristics between the sound signals of different microphones, using a frequency shift compensation algorithm and combining with a stability factor to obtain the compensated multi-channel microphone sound signals includes: Calculating the horizontal sound path difference, frequency shift amount and stability factor between each pair of microphone combinations according to the interference characteristics between the microphone sound signals; Determining the waveguide invariant according to the characteristics of the propagation environment, and calculating the time delay compensation amount between each pair of microphone combinations according to the waveguide invariant and the horizontal sound path difference; Performing a frequency shift operation on the signals received by each pair of microphone combinations according to the frequency shift amount, and performing a phase adjustment on the frequency-shifted signals according to the calculated time delay compensation amount to obtain a corrected signal; Performing a weighted compensation process on the corrected signal based on the stability factor to obtain the compensated microphone sound signals.
[0013] Preferably, performing a frequency shift operation on the signals received by each pair of microphone combinations according to the frequency shift amount, and performing a phase adjustment on the frequency-shifted signals according to the calculated time delay compensation amount to obtain a corrected signal includes: Performing a transformation process on the signals received by each pair of microphone combinations by using the wavelet transform method to obtain a frequency domain signal; Adjust the frequency of the frequency-domain signal based on the frequency shift amount, and perform a frequency shift operation on the frequency-domain signal using interpolation technology; Perform phase correction on the frequency-shifted signal using the time delay compensation amount to obtain a corrected signal.
[0014] Preferably, the weighted compensation process for the corrected signal based on the stability factor to obtain the compensated microphone sound signal includes: Use the stability factor as the weighting coefficient for each microphone, and perform normalization processing on the weighting coefficient; Use the normalized weighting coefficient to perform a weighting process on the corrected signal of the microphone to obtain the compensated microphone sound signal.
[0015] According to another aspect of the present invention, there is provided a method for using a plug-in modular digital audio processing device, including the following steps: S1. Insert a preset hardware expansion card into the standardized card slot provided by the plug-in expansion unit; S2. Collect audio signal data based on the plug-in expansion unit, and use an audio processing algorithm to process the audio signal data to obtain the sound signal of the microphone; S3. Manage several microphones simultaneously, and use the frequency shift compensation algorithm with the stability factor to analyze acoustic interference to achieve no howling when all microphones are fully open.
[0016] The beneficial effects of the present invention are: 1. The present invention has multiple functions such as modular flexibility and scalability, rich audio processing algorithm support, efficient audio digital signal collection and processing, and intelligent microphone management to achieve no howling when fully open. It enables the device to adapt to audio processing requirements in different scenarios, provide high-quality, stable and reliable audio processing effects, and improve the overall user experience.
[0017] 2. Through frame division, windowing, fast Fourier transform and feature extraction, combined with the support vector machine recursive feature elimination method, the present invention can accurately remove redundant features and obtain a feature subset. Then, through the similarity calculation between the classification model and the preset feature labels, it can intelligently match and select the optimal audio processing algorithm, thereby ensuring that the audio signal is processed with strong pertinence and optimized effects, and significantly improving the accuracy and efficiency of audio processing.
[0018] 3. The present invention uses the short-time Fourier transform of the dichotomy method to deeply analyze the interference characteristics between the sound signals of different microphones, accurately locate the interference frequency band; then, combined with the frequency shift compensation algorithm and the stability factor, it performs frequency shift, time delay compensation and phase adjustment on the interference signal, effectively eliminating the acoustic interference phenomenon; finally, through the weighted compensation process of the stability factor, it optimizes the signal quality, ensures no howling when multiple microphones are fully open, and significantly improves the stability and audio quality of the audio system. Brief Description of the Drawings
[0019] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings. In the drawings: Figure 1 is a schematic block diagram of a plug-in modular digital audio processing device according to an embodiment of the present invention; Figure 2 is a flowchart of the usage method of a plug-in modular digital audio processing device according to an embodiment of the present invention.
[0020] In the figure: 1. Plug-in expansion unit; 2. Audio processing algorithm storage unit; 3. Audio digital signal collection and processing unit; 4. Microphone intelligent management unit. Detailed Embodiments
[0021] In order to enable those skilled in the art of this technology to better understand the technical solutions in this application, the following will clearly and completely describe the technical solutions in the embodiments of this application in conjunction with the drawings in the embodiments of this application. Obviously, the described embodiments are only some embodiments of this application, rather than all embodiments. Based on the embodiments in this application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of this application.
[0022] According to an embodiment of the present invention, a plug-in modular digital audio processing device and a usage method are provided.
[0023] Specifically, the plug-in modular digital audio processing device can use WM-CK300 as a distributed network audio system processing service center, which is a core AI audio algorithm processing platform for enterprises. It integrates powerful hardware processing capabilities, rich and powerful various audio algorithms, provides 128x128 channel ultra-low latency, uncompressed network audio input, output and processing capabilities, supports 16x16 RTSP network audio streams, the audio matrix routing can reach 144*144, can be extended to be compatible with protocols such as DANTE and AES67, and supports a maximum of 128x128 channels.
[0024] The WM-CK300 adopts a configurable plug-in card design with a total of 9 card slots. It can flexibly configure 8 analog input and output cards, and can be optionally equipped with DANTE expansion cards. At the same time, the system functions are software-defined, allowing users to customize audio processing functions and system architectures. It supports environmental control management and can centrally manage devices such as network microphones, network speakers, and audio nodes.
[0025] It adopts a 1+1 system hot standby design and supports dual AC power supplies simultaneously. It can demonstrate the reliability of the system in any scale of application scenarios and is an excellent solution with excellent performance and reliability.
[0026] It has a large network audio channel access capacity and powerful processing functions, and is suitable for diverse large-scale application places: large command centers, conference centers, administrative centers, corporate headquarters buildings, theme parks, transportation hubs, etc.
[0027] The present invention will be further described in conjunction with the accompanying drawings and specific embodiments. As Figure 1 shown, according to an embodiment of the present invention, a plug-in modular digital audio processing device is provided, including: A plug-in expansion unit 1 for providing standardized card slots for inserting preset hardware expansion cards; Specifically, the standardized card slots follow unified physical dimensions and electrical interface standards to ensure that expansion cards with different functions can be inserted and removed conveniently and quickly. Users can insert expansion cards with different functions according to their needs, including: 128-channel Dante expansion card Amu 128D, 4-channel microphone / line input card Amu 4I, 4-channel line output card Amu 4O, etc.
[0028] Among the above-mentioned expansion cards, there is a multi-channel audio signal acquisition ability, which can collect signals from multiple microphones or audio sources simultaneously.
[0029] An audio processing algorithm storage unit 2 for storing various audio processing algorithms and feature tags corresponding to the audio processing algorithms; It should be noted that the audio processing algorithm storage unit 2 stores hundreds of rich audio processing algorithms, supports 5A audio algorithms, AI noise reduction, AI automatic gain, and other common audio algorithms, such as threshold automatic mixing, gain-sharing automatic mixing, and dynamic compressors.
[0030] An audio digital signal collection and processing unit 3 for collecting audio signal data based on the plug-in expansion unit and processing the audio signal data using the audio processing algorithms to obtain the sound signals of the microphones; As a preferred embodiment, the audio digital signal collection and processing unit 3 includes: An audio collection preprocessing module, which is used to collect audio signal data based on a plug-in expansion unit and preprocess the collected audio signal data; It should be noted that the preprocessing operations include: denoising, gain control, sampling rate conversion, etc.
[0031] Among them, denoising uses a filtering algorithm (such as high-pass / low-pass filtering) to remove environmental noise or device background noise. Gain control dynamically adjusts the signal amplitude to avoid clipping or distortion and at the same time improves the signal-to-noise ratio of low-level signals. Sampling rate conversion unifies the signal sampling rate (such as 48 kHz) to ensure the consistency of the algorithm input.
[0032] A processing algorithm selection module, which is used to match and select an audio processing algorithm in the audio processing algorithm storage unit based on the preprocessed audio signal data according to the combined feature analysis method; As a preferred implementation manner, the matching and selecting of an audio processing algorithm in the audio processing algorithm storage unit based on the preprocessed audio signal data according to the combined feature analysis method includes: Performing frame segmentation on the preprocessed audio signal data, and multiplying each frame of the segmented audio signal by a preset window function for windowing processing; Specifically, the audio signal changes continuously over time. In order to perform short-time spectrum analysis, it needs to be divided into multiple short-time frames. Frame segmentation can ensure that the audio signal is approximately stationary within each frame, thus meeting the prerequisite conditions for Fourier transform.
[0033] Specifically, the frame length is usually set to 20 - 40 ms (such as 32 ms) to ensure that the signal is stationary enough within the frame. The frame shift (the overlap amount between adjacent frames) is usually set to half of the frame length (such as 16 ms) to increase the signal continuity and reduce spectral leakage. Then, using the sliding window technique, the continuous audio signal is divided into multiple equal-length frames.
[0034] Among them, windowing processing can reduce spectral leakage and improve the accuracy of spectrum analysis. The window function smooths the edges of the signal and reduces the spectral side lobes generated due to signal truncation.
[0035] Performing a fast Fourier transform on each windowed frame signal to obtain the spectrum of each frame of audio signal, and extracting Mel frequency cepstral coefficient features and gamma-tone frequency cepstral coefficient features; Among them, the fast Fourier transform (FFT) can convert the time-domain signal into a frequency-domain signal, thereby obtaining the spectral characteristics of the audio signal. Spectrum analysis is the basis for extracting audio features. Specifically, it includes: performing FFT transformation on each windowed frame signal to obtain the spectrum in complex form, and then taking the modulus value of the FFT result to obtain the amplitude spectrum; Based on the Mel-frequency cepstral coefficient (MFCC) features and gammatone frequency cepstral coefficient (GFCC) features, the support vector machine recursive feature elimination method is used to match and select the optimal audio processing algorithm.
[0036] As a preferred embodiment, the method of matching and selecting the optimal audio processing algorithm based on the Mel-frequency cepstral coefficient features and gammatone frequency cepstral coefficient features using the support vector machine recursive feature elimination method includes: Normalize the Mel-frequency cepstral coefficient features and gammatone frequency cepstral coefficient features respectively, and use the support vector machine recursive feature elimination method to remove redundant and irrelevant features; obtain a feature subset; It should be noted that the Mel-frequency cepstral coefficient features and gammatone frequency cepstral coefficient features have different dimensions and value ranges. Direct use will cause some features to dominate in model training, affecting the accuracy of feature selection and algorithm matching. Therefore, through normalization, the influence of dimensions can be eliminated, so that all features have the same scale. Among them, the normalization process includes: Z-Score normalization and Min-Max normalization.
[0037] Among them, the Mel-frequency cepstral coefficient features and gammatone frequency cepstral coefficient features may also contain redundant or irrelevant information, which will increase the computational complexity and may reduce the accuracy of algorithm matching. The support vector machine recursive feature elimination method removes redundant and irrelevant features by iteratively training the SVM model and eliminating features with smaller weights, which can effectively remove redundant and irrelevant features. Specifically, it includes the following steps: Merge the normalized Mel-frequency cepstral coefficient features and gammatone frequency cepstral coefficient features into a high-dimensional feature vector as the initial feature set, then use the current feature set to train the SVM classifier, calculate the weight of each feature based on the decision function coefficients (or gradients) of the SVM classifier, remove the feature with the smallest weight (or several features with the smallest weights), update the feature set, and stop the iteration when the preset number of features is reached, the model performance no longer improves, or all features are eliminated.
[0038] Based on a preset classification model, classify the feature subset to obtain classification labels, calculate the similarity between the classification labels and the feature labels, and select the audio processing algorithm according to the similarity calculation results.
[0039] Specifically, the steps of classifying the feature subset based on a preset classification model, calculating the similarity between the classification labels and the feature labels, and selecting the audio processing algorithm according to the similarity calculation results include: First, use the labeled audio dataset (including Mel-frequency cepstral coefficient features, gammatone frequency cepstral coefficient features, and corresponding classification labels) to train a random forest classification model; Then, the feature subset obtained by SVM-RFE is input into the trained random forest classification model to obtain classification labels.
[0040] Then, using the cosine similarity, Euclidean distance, or Jaccard similarity algorithm, calculate the similarity between the classification labels and the feature labels. Finally, based on the calculation results of the similarity, select the audio processing algorithm corresponding to the feature label with the highest similarity as the optimal algorithm.
[0041] The audio processing module is used to perform audio processing on the preprocessed audio signal data based on the selected audio processing algorithm to obtain the sound signal of the microphone.
[0042] Specifically, input the audio signal data into the audio processing algorithm and process it according to the logical flow of the audio processing algorithm. The processing process includes operations such as filtering, noise reduction, gain adjustment, and feature extraction.
[0043] The microphone intelligent management unit 4 is used to manage several microphones simultaneously and analyze sound interference using the frequency shift compensation algorithm of the stability factor to achieve non-whistling when all microphones are on.
[0044] As a preferred embodiment, the microphone intelligent management unit 4 includes: The feature analysis module is used to collect the sound signals of each microphone and analyze the interference characteristics between the sound signals of different microphones using the short-time Fourier transform of the bisection method; As a preferred embodiment, the collecting the sound signals of each microphone and analyzing the interference characteristics between the sound signals of different microphones using the short-time Fourier transform of the bisection method includes: Perform time-frequency analysis on each microphone signal using the short-time Fourier transform to generate a time-frequency diagram. It should be noted that by analyzing the time-frequency characteristics of each microphone signal through the short-time Fourier transform (STFT) to generate a time-frequency diagram, it provides a basis for subsequent interference feature analysis.
[0045] Based on the generated time-frequency diagram, divide the frequency range in the time-frequency diagram using the bisection method to obtain two groups of sub-frequency bands; It should be noted that before dividing the frequency range in the time-frequency diagram into multiple sub-frequency bands using the bisection method, it is necessary to determine the frequency interval to be analyzed according to the frequency range of the time-frequency diagram, and then bisect the determined frequency interval to obtain two sub-frequency bands. And record the frequency range of each sub-frequency band.
[0046] For each sub-frequency band, select the signals of two different microphones and calculate their cross-correlation function within this sub-frequency band. According to the calculation results, judge whether there are interference characteristics in each sub-frequency band; It should be noted that by performing cross - correlation calculation on the time - frequency representations of the selected two - channel microphone signals within the sub - band, the similarity of the two signals in terms of time delay can be obtained. Among them, for two signals x ( t ) and y ( t ), their cross - correlation function R xy ( τ ) at time delay τ has the following calculation formula: ; For the sub - bands with interference characteristics, interference feature extraction is performed to obtain the interference features. For the sub - bands without interference characteristics, the dichotomy method is used again to divide the sub - bands without interference characteristics and re - perform interference feature extraction until the preset conditions are met.
[0047] Specifically, for the sub - bands with interference characteristics, feature extraction is performed, and for the sub - bands without interference characteristics, further division is carried out until all sub - bands are fully analyzed. Specific steps: For the sub - bands with interference characteristics, extract interference features such as interference time delay, interference amplitude, interference phase, etc.; for the sub - bands without interference characteristics, use the dichotomy method again to divide them into smaller sub - bands; repeat the above steps to calculate the cross - correlation function and judge the interference characteristics of the new sub - bands.
[0048] Among them, the preset conditions include the minimum frequency width of the sub - band, the maximum number of divisions, the confidence level of interference feature extraction, etc. When the preset conditions are met, the division and interference feature extraction processes are stopped.
[0049] The signal compensation module is used to compensate the interference signals between microphones according to the interference characteristics between different microphone sound signals, using the frequency shift compensation algorithm and combining with the stability factor to obtain the compensated multi - channel microphone sound signals; As a preferred embodiment, the compensating the interference signals between microphones according to the interference characteristics between different microphone sound signals, using the frequency shift compensation algorithm and combining with the stability factor to obtain the compensated multi - channel microphone sound signals includes: Calculating the horizontal sound path difference, frequency shift amount and stability factor between each pair of microphone combinations according to the interference characteristics between microphone sound signals; It should be noted that the horizontal sound path difference refers to the distance difference between two microphones in the horizontal direction, which causes the difference in the arrival time of sound signals. By using the geometric layout of the microphone array and the known microphone position information, the horizontal sound path difference between each pair of microphones is calculated.
[0050] The frequency shift amount refers to the difference in the signal frequencies received by different microphones due to the relative motion between the sound source and the microphone or the multipath propagation effect. Methods such as cross-correlation analysis and Doppler frequency shift estimation are used to calculate the frequency shift amount between each pair of microphone signals.
[0051] The stability factor is used to measure the stability of the interference characteristics of the microphone signals, that is, whether the interference pattern changes over time. By analyzing features such as the correlation and power spectral density of the signals, the stability factor is calculated to evaluate the stability of the interference characteristics.
[0052] According to the characteristics of the propagation environment, the waveguide invariant is determined, and the time delay compensation amount between each pair of microphone combinations is calculated based on the waveguide invariant and the horizontal acoustic path difference; It should be noted that the waveguide invariant is a parameter related to the propagation characteristics of sound waves in a waveguide (such as media like air, water, etc.). Then, through experimental measurement or theoretical calculation, the waveguide invariant of the propagation environment can be determined. Among them, the calculation formula for the time delay compensation amount between each pair of microphone combinations based on the waveguide invariant and the horizontal acoustic path difference is: ; In the formula, τ (Δ r , r j ) represents the time delay compensation amount, β s and β p represent parameters related to the phase velocity and group velocity; Δ r represents the horizontal acoustic path difference between two microphones, r j represents the j distance of the microphone; According to the frequency shift amount, frequency shift operations are performed on the signals received by each pair of microphone combinations, and phase adjustments are made to the frequency-shifted signals according to the calculated time delay compensation amount to obtain corrected signals; As a preferred implementation manner, the step of performing frequency shift operations on the signals received by each pair of microphone combinations according to the frequency shift amount and making phase adjustments to the frequency-shifted signals according to the calculated time delay compensation amount to obtain corrected signals includes: Using the wavelet transform method to perform transformation processing on the signals received by each pair of microphone combinations to obtain frequency-domain signals; Specifically, using wavelet basis functions (such as Daubechies wavelets, Morlet wavelets, etc.), continuous wavelet transform is performed on the time-domain signals received by each pair of microphone combinations to obtain the representation of the signals at different scales and frequencies, and then the frequency-domain signals are obtained by reconstructing the wavelet coefficients in a specific scale or frequency interval.
[0053] Adjust the frequency of the frequency-domain signal based on the frequency shift amount, and perform a frequency shift operation on the frequency-domain signal using interpolation techniques; Specifically, in the frequency domain, the frequency shift is achieved by translating the spectrum of the frequency-domain signal. Specifically, each frequency component of the frequency-domain signal is multiplied by a complex exponential factor ( f is the frequency shift amount, t is the time variable, j is the imaginary unit); then, a linear interpolation is used to perform a frequency shift operation on the frequency-domain signal.
[0054] Use the time delay compensation amount to correct the phase of the frequency-shifted signal to obtain a corrected signal.
[0055] It should be noted that the time delay compensation amount represents the time difference in the propagation of the sound signal between different microphones. In the frequency domain, the time delay compensation amount can be converted into a phase adjustment amount. For each frequency component f in the frequency domain, the phase adjustment amount can be calculated by the following formula: ; For each frequency component f in the frequency domain, the corrected frequency component is multiplied by a phase factor (where f is the frequency shift amount, τ is the time delay, j is the imaginary unit), and the corrected signal can be obtained.
[0056] Based on the stability factor, perform a weighted compensation process on the corrected signal to obtain a compensated microphone sound signal.
[0057] As a preferred embodiment, the performing a weighted compensation process on the corrected signal based on the stability factor to obtain a compensated microphone sound signal includes: Use the stability factor as the weighting coefficient for each microphone, and perform a normalization process on the weighting coefficient; the normalization process is to ensure the rationality and consistency of the weighting coefficient.
[0058] Use the normalized weighting coefficient to perform a weighting process on the corrected signal of the microphone to obtain a compensated microphone sound signal.
[0059] It should be noted that for each time point (or frequency point), the corrected signal of each microphone is multiplied by its corresponding normalized weighting coefficient, and then all the weighted signals are added together to obtain a compensated sound signal. After the weighting process, the multiple-channel microphone sound signals are synthesized into a high-quality sound signal. This signal makes full use of the information of multiple microphones, improves the signal-to-noise ratio of the signal, and reduces the influence of interference and noise.
[0060] A signal output module for real - time output of the compensated multi - channel microphone signals to achieve non - whistling when all microphones are fully open.
[0061] As Figure 2 shown, according to another embodiment of the present invention, a method for using a plug - in modular digital audio processing device is provided, including the following steps: S1. Insert a preset hardware expansion card into the standardized card slot provided by the plug - in expansion unit; S2. Collect audio signal data based on the plug - in expansion unit and process the audio signal data using an audio processing algorithm to obtain the sound signals of the microphones; S3. Simultaneously manage several microphones and analyze acoustic interference using a frequency shift compensation algorithm with a stability factor to achieve non - whistling when all microphones are fully open.
[0062] In summary, by means of the above - mentioned technical solutions of the present invention, the present invention has functions such as modular flexibility and scalability, rich support for audio processing algorithms, efficient collection and processing of audio digital signals, and intelligent management of microphones to achieve non - whistling when all are fully open. It enables the device to adapt to the audio processing requirements in different scenarios, provides high - quality, stable and reliable audio processing effects, and enhances the overall user experience. Through frame division, windowing, fast Fourier transform and feature extraction, combined with the support vector machine recursive feature elimination method, the present invention can accurately remove redundant features and obtain a feature subset. Then, through the similarity calculation between the classification model and the preset feature labels, it can intelligently match and select the optimal audio processing algorithm, thereby ensuring that the audio signals are processed with strong pertinence and optimized effects, significantly improving the accuracy and efficiency of audio processing. The present invention uses the short - time Fourier transform of the dichotomy method to deeply analyze the interference characteristics between the sound signals of different microphones, accurately locate the interference frequency band; and then combines the frequency shift compensation algorithm with the stability factor to perform frequency shift, time delay compensation and phase adjustment on the interference signals, effectively eliminating the acoustic interference phenomenon; finally, through the stability factor weighted compensation processing, the signal quality is optimized to ensure no whistling when all multi - microphones are fully open, significantly improving the stability and audio quality of the audio system.
[0063] Those skilled in the art should understand that the embodiments of the present invention can be provided as methods, systems or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer - usable storage media (including but not limited to disk storage, optical storage, etc.) containing computer - usable program code.
[0064] The specific embodiments described above further elaborate on the object, technical solution, and beneficial effects of the present invention. It should be understood that the above are only specific embodiments of the present invention and are not used to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A plug-in modular digital audio processing device, characterized in that including: A plug-in expansion unit for providing a standardized card slot for inserting a preset hardware expansion card; An audio processing algorithm storage unit for storing multiple audio processing algorithms and corresponding feature tags for the audio processing algorithms; An audio digital signal collection and processing unit for collecting audio signal data based on the plug-in expansion unit and processing the audio signal data using an audio processing algorithm to obtain a sound signal of a microphone; A microphone intelligent management unit for simultaneously managing several microphones and analyzing acoustic interference using a frequency shift compensation algorithm with a stability factor to achieve non-whistling when all microphones are turned on.
2. The plug-in modular digital audio processing device according to claim 1, wherein The audio digital signal collection and processing unit includes: An audio collection and preprocessing module for collecting audio signal data based on the plug-in expansion unit and preprocessing the collected audio signal data; A processing algorithm selection module for matching and selecting an audio processing algorithm in the audio processing algorithm storage unit based on the preprocessed audio signal data according to the combined feature analysis method; An audio processing module for performing audio processing on the preprocessed audio signal data based on the selected audio processing algorithm to obtain a sound signal of a microphone.
3. The plug-in modular digital audio processing device according to claim 2, characterized in that, The matching and selecting of an audio processing algorithm in the audio processing algorithm storage unit based on the preprocessed audio signal data according to the combined feature analysis method includes: Performing frame division processing on the preprocessed audio signal data, and multiplying each frame of the divided audio signal by a preset window function for windowing processing; Performing a fast Fourier transform on each windowed frame signal to obtain the spectrum of each frame of audio signal, and extracting Mel frequency cepstral coefficient features and gamma-tone frequency cepstral coefficient features; Based on the Mel frequency cepstral coefficient features and gamma-tone frequency cepstral coefficient features, using the support vector machine recursive feature elimination method to match and select the optimal audio processing algorithm.
4. The plug-in modular digital audio processing device according to claim 3, characterized in that, The matching and selecting of the optimal audio processing algorithm based on the Mel frequency cepstral coefficient features and gamma-tone frequency cepstral coefficient features using the support vector machine recursive feature elimination method includes: Performing standardization processing on the Mel frequency cepstral coefficient features and gamma-tone frequency cepstral coefficient features respectively, and using the support vector machine recursive feature elimination method to remove redundant and irrelevant features; obtaining a feature subset; Based on a preset classification model, classifying the feature subset to obtain a classification label, and calculating the similarity between the sub-label and the feature label, and selecting an audio processing algorithm according to the similarity calculation result.
5. A plug-in modular digital audio processing device according to claim 1, characterized in that, The microphone intelligent management unit includes: A feature analysis module for collecting the sound signals of each microphone and analyzing the interference characteristics between the sound signals of different microphones using the short-time Fourier transform of the dichotomy method; A signal compensation module for compensating the interference signals between microphones according to the interference characteristics between the sound signals of different microphones, using the frequency shift compensation algorithm and combining with a stability factor to obtain compensated multi-channel microphone sound signals; A signal output module for real-time outputting the compensated multi-channel microphone signals to achieve non-whistling when all microphones are turned on.
6. The plug-in modular digital audio processing device according to claim 5, characterized in that, Collecting the sound signals of each microphone and analyzing the interference characteristics between the sound signals of different microphones by using the short-time Fourier transform of the dichotomy method includes: Performing time-frequency analysis on each microphone signal respectively by using the short-time Fourier transform to generate a time-frequency diagram; Based on the generated time-frequency diagram, dividing the frequency range in the time-frequency diagram by using the dichotomy method to obtain two groups of sub-frequency bands; For each sub-frequency band, selecting the signals of two different microphones and calculating the cross-correlation function between them within this sub-frequency band, and judging whether there are interference characteristics in each sub-frequency band according to the calculation results; For the sub-frequency bands with interference characteristics, extracting the interference characteristics to obtain the interference characteristics. For the sub-frequency bands without interference characteristics, using the dichotomy method again to divide the sub-frequency bands without interference characteristics and re-performing the interference characteristic extraction until the preset conditions are met.
7. An insert card type modular digital audio processing device according to claim 5, characterized in that According to the interference characteristics between the sound signals of different microphones, using the frequency shift compensation algorithm and combining with the stability factor to compensate the interference signals between the microphones to obtain the compensated multi-channel microphone sound signals includes: Calculating the horizontal sound path difference, frequency shift amount and stability factor between each pair of microphone combinations according to the interference characteristics between the microphone sound signals; Determining the waveguide invariant according to the characteristics of the propagation environment, and calculating the time delay compensation amount between each pair of microphone combinations according to the waveguide invariant and the horizontal sound path difference; Performing a frequency shift operation on the signals received by each pair of microphone combinations according to the frequency shift amount, and performing a phase adjustment on the frequency-shifted signals according to the calculated time delay compensation amount to obtain a corrected signal; Performing a weighted compensation process on the corrected signal based on the stability factor to obtain the compensated microphone sound signals.
8. An insertion card type modular digital audio processing device according to claim 7, characterized in that According to the frequency shift amount, performing a frequency shift operation on the signals received by each pair of microphone combinations, and performing a phase adjustment on the frequency-shifted signals according to the calculated time delay compensation amount to obtain a corrected signal includes: Performing a transformation process on the signals received by each pair of microphone combinations by using the wavelet transform method to obtain a frequency-domain signal; Based on the frequency shift amount, adjusting the frequency of the frequency-domain signal and performing a frequency shift operation on the frequency-domain signal by using the interpolation technique; Performing a phase correction on the frequency-shifted signal by using the time delay compensation amount to obtain a corrected signal.
9. The plug-in modular digital audio processing device according to claim 7, characterized in that Based on the stability factor, performing a weighted compensation process on the corrected signal to obtain the compensated microphone sound signals includes: Taking the stability factor as the weighting coefficient of each microphone and performing a normalization process on the weighting coefficient; Using the normalized weighting coefficient to perform a weighting process on the corrected signal of the microphone to obtain the compensated microphone sound signals.
10. A method for using a plug-in modular digital audio processing device, which is used to implement the use of the plug-in modular digital audio processing device described in any one of claims 1-9, characterized in that, Including the following steps: S1. Inserting a preset hardware expansion card into the standardized card slot provided by the plug-in expansion unit; S2. Collecting audio signal data based on the plug-in expansion unit and processing the audio signal data by using an audio processing algorithm to obtain the sound signals of the microphones; S3. Simultaneously managing several microphones and analyzing the sound interference by using the frequency shift compensation algorithm of the stability factor to achieve no howling when all microphones are turned on.
Citation Information
Patent Citations
Method and device for improving audio processing performance
CN104980337A
Speech recognition method based on power spectrum Gabor feature sequence recursive model
CN107103913A
Continuous mobile microphone array sound source positioning method and device
CN118191732A
Speech processing of audio signal
EP4386749A1
Speech data reproduction system, speech data reproduction program and utterance practice device
JP2008134455A