Self-adaptive dynamic audio processing method, storage medium and terminal

By employing an adaptive dynamic audio processing method, utilizing frame-by-frame processing and performance evaluation of microphone arrays, acoustic scenes are identified and microphone selection is optimized. This solves the problem of echo and noise interference in complex environments that traditional audio processing methods encounter, achieving stability and real-time adaptability in audio output.

CN121884840APending Publication Date: 2026-04-17SHENZHEN MINRRAY IND CORP LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SHENZHEN MINRRAY IND CORP LTD
Filing Date
2026-01-20
Publication Date
2026-04-17

AI Technical Summary

Technical Problem

Traditional audio processing methods cannot adapt to dynamic changes in the acoustic environment, making it difficult to balance processing effectiveness, adaptability, and real-time performance. In particular, in complex real-time audio interaction systems, echoes and noise interference are severe.

Method used

An adaptive dynamic audio processing method is adopted. By performing frame-by-frame processing on the reference signal and microphone signal, energy and correlation coefficients are calculated, acoustic scenes are identified, and the optimal microphone is selected based on performance indicators. Combined with filters and machine learning algorithms, real-time optimization is performed to achieve echo and noise suppression.

Benefits of technology

It enables real-time adaptation to sound source movement and noise abrupt changes in complex acoustic environments, ensuring the stability and adaptability of audio output quality, and improving the real-time performance and adaptability of processing effects.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121884840A_ABST
    Figure CN121884840A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of audio signal processing, in particular to a self-adaptive dynamic audio processing method, a storage medium and a terminal. The adaptive dynamic audio processing method comprises the following steps: acquiring a reference signal and each microphone signal of a microphone array, and framing the reference signal and each microphone signal to obtain a reference signal frame and a microphone signal frame; calculating the energy of each frame of signal, comparing the energy of the reference signal and the energy of the microphone signal with a far-end energy threshold and a near-end energy threshold respectively, comparing the correlation coefficients of the energy of the reference signal and the energy of the microphone signal with a correlation threshold, and judging four application scenes of far-end dominant, near-end dominant, double-talk and mute according to three groups of results; generating an estimated echo signal frame based on the reference signal frame, subtracting the frame from the microphone signal frame to obtain a residual signal frame, and calculating the energy of the residual signal frame; and calculating the performance index and score of each microphone in combination with the application scene and the residual signal energy, screening the optimal microphone, and obtaining and outputting a transmission signal according to the residual signal frame of the optimal microphone. According to the invention, stable audio output quality can be maintained.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of audio signal processing technology, and in particular to an adaptive dynamic audio processing method, storage medium, and terminal. Background Technology

[0002] In real-time audio interaction systems, microphone arrays are widely used to collect voice signals due to their spatial sound pickup advantages. However, the complex and varied acoustic environments in real-world applications lead to severe problems such as echo and noise interference.

[0003] Traditional audio processing methods mostly adopt static processing strategies, with fixed algorithm parameters and processing flow. They cannot adaptively adjust to dynamic changes in the acoustic environment (such as sound source movement, noise abrupt changes, and changes in acoustic path), making it difficult to balance processing effect, adaptability, and real-time performance. Summary of the Invention

[0004] This invention provides an adaptive dynamic audio processing method to solve the problem of balancing processing effect, adaptability, and real-time performance.

[0005] This invention discloses an adaptive dynamic audio processing method applied to a microphone array, wherein the microphone array includes multiple microphones; The adaptive dynamic audio processing method includes: Acquire a reference signal and the microphone signal of each microphone in the microphone array, and perform frame segmentation on the reference signal and each microphone signal to obtain at least one reference signal frame and at least one microphone signal frame; For each frame of the microphone signal, calculate its microphone signal energy; for each frame of the reference signal, obtain its reference signal energy; compare the reference signal energy at the current moment with the far-end energy threshold to obtain a first comparison result; compare the microphone signal energy with the near-end energy threshold to obtain a second comparison result; compare the correlation coefficient between the reference signal energy and the microphone signal energy with the correlation threshold to obtain a third comparison result. Based on the first comparison result, the second comparison result and the third comparison result, the application scenario at the current moment is obtained, and the application scenario includes any one of remote dominance, near-end dominance, dual-talk and mute; Based on the reference signal frame, obtain the estimated echo signal frame, subtract the estimated echo signal frame from the microphone signal frame corresponding to the reference signal frame to obtain the residual signal frame, and obtain the residual signal energy of the residual signal frame. The performance index of the microphone is calculated based on the application scenario and the residual signal energy, and the corresponding index score of the microphone is obtained based on the application scenario and the performance index. The optimal microphone in the microphone array is selected based on the index score of each microphone, and the transmission signal is obtained from the residual signal frame of the optimal microphone and then output.

[0006] Optionally, the step of selecting the optimal microphone in the microphone array based on the metric score of each microphone includes: The other microphones in the microphone array are used as candidate microphones, and the index score of the best microphone is used as the target score. The index score of each microphone in the microphone array is acquired in real time, and it is determined whether there is at least one candidate microphone whose index score exceeds the preset score difference threshold of the target score. If so, at least one candidate microphone whose index score exceeds the target score will be selected as the target microphone. Determine whether the duration for which the target microphone's score exceeds the target score exceeds a preset first duration threshold. If so, determine whether the duration for which the optimal microphone is used as the optimal microphone exceeds the second duration threshold. If so, the microphone with the highest index score among the target microphones will be selected as the new optimal microphone.

[0007] Optionally, the step of obtaining the estimated echo signal frame based on the reference signal frame includes: The filter coefficient matrix of the foreground filter is corrected based on the corresponding reference signal frame and the microphone signal frame using the minimum mean square error algorithm or the normalized minimum mean square error algorithm. The update rate and convergence speed of the foreground filter are modified based on the application scenario. Collect static noise, establish a noise baseline model based on the spectral characteristics of the static noise, and feed the noise baseline model back to the foreground filter; The mean square error of the foreground filter is monitored in real time. When the mean square error exceeds a preset error threshold, the filter coefficient matrix of the foreground filter is replaced with a candidate filter coefficient matrix. The foreground filter is used to obtain the estimated echo frame corresponding to the reference signal frame.

[0008] Optionally, the step of obtaining the estimated echo signal frame corresponding to the reference signal frame through the foreground filter includes: The full frequency domain is divided into multiple sub-bands, and the microphone signal frame is divided into multiple sub-microphone signal frames according to the multiple sub-bands; The plurality of sub-signal frames are respectively input into the corresponding foreground filters, and the filter coefficient matrix of the foreground filters is adjusted based on the frequency band corresponding to the currently input sub-microphone signal frame; Obtain the sub-estimated echo signal frame corresponding to each sub-microphone signal frame, and obtain the sub-residual signal frame corresponding to each sub-microphone signal frame based on the sub-estimated echo signal frame; The sub-residual signal frames are spliced ​​together to form the residual signal frame.

[0009] Optionally, after the step of obtaining at least one reference signal frame and at least one microphone signal frame, the process includes: Each microphone signal frame and / or reference signal frame is pre-emphasized and subjected to a short-time Fourier transform. The step of obtaining the transmitted signal based on the residual signal frame of the optimal microphone includes: Each residual signal frame of the optimal microphone is subjected to de-emphasis processing and inverse short-time Fourier transform, and then the processed residual signal frames are spliced ​​together to form a continuous transmission signal.

[0010] Optionally, after the step of selecting the optimal microphone in the microphone array based on the metric score of each microphone, the method further includes: Based on the microphone signal frame, the residual signal frame, and the reference signal frame of each microphone in the microphone array, and the preset performance indicators set for the microphone, the optimal scene classification is determined based on the preset feature threshold and the machine learning classifier. The optimal scene classification includes any one of the following: near-end speech dominance, far-end speech dominance, two-way talk, steady-state noise, and burst noise. Based on the optimal scene classification, a corresponding optimization algorithm is selected, and the residual signal frame is optimized using the optimization algorithm.

[0011] Optionally, the step of splicing the processed residual signal frames into a continuous transmission signal includes: The processed residual signal frame is subjected to residual echo suppression to obtain an echo-suppressed frame; The noise energy is estimated using spectral subtraction or a deep learning-based method, and noise suppression is performed on the echo-suppressed frame based on the noise energy to obtain a noise-suppressed frame. According to the application scenario, gain compensation is performed on the noise suppression frame to obtain the gain compensation frame; The gain-compensated frames are spliced ​​together to form the transmission signal.

[0012] Optionally, after the step of selecting the optimal microphone in the microphone array based on the index score of each microphone, the method further includes: The gain of the other microphones in the microphone array is adjusted based on the residual signal energy of the optimal microphone; Based on the pre-stored microphone frequency response compensation coefficient table, the frequency response deviation of each microphone is compensated by an FIR filter; The delay difference between the other microphones and the optimal microphone is calculated by cross-correlation analysis, and a delay compensation frame is inserted into the residual signal frame.

[0013] The present invention also discloses a storage medium storing a computer program, which, when executed by a processor, causes the processor to perform the steps of the method described above.

[0014] The present invention also discloses an audio transmission and playback terminal, including a memory and a processor. The memory stores a computer program, and when the computer program is executed by the processor, the processor performs the steps of the method described above.

[0015] The beneficial effects of the adaptive dynamic audio processing method, storage medium, and terminal provided in this invention are as follows: By processing the reference signal and microphone signal in frames, and by comparing the energy of the reference signal, the energy of the microphone signal, and the correlation coefficient and corresponding threshold of the two in real time, the system can accurately identify dynamic acoustic scenarios such as far-end dominance, near-end dominance, dual-talk, and silence. This breaks the single and fixed judgment logic of traditional methods. At the same time, relying on the real-time performance evaluation and optimal channel selection mechanism of the multi-channel microphone array, combined with channel equalization processing of energy, delay, and frequency response, the system can adapt to environmental changes such as sound source movement, noise abrupt changes, and acoustic path changes in real time. This ensures that no matter how the acoustic environment changes dynamically, the system can automatically switch to the optimal processing strategy and microphone channel, and always maintain stable audio output quality. Attached Figure Description

[0016] The technical solution of the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. In the accompanying drawings: Figure 1 This is a flowchart illustrating an embodiment of the adaptive dynamic audio processing method provided by the present invention; Figure 2 This is a schematic diagram of the internal structure of an audio transmission and playback terminal in one embodiment of the present invention. Detailed Implementation

[0017] It should be noted that, unless otherwise specified, the embodiments and features described in this application can be combined with each other. The preferred embodiments of the present invention will now be described in detail with reference to the accompanying drawings.

[0018] Please refer to the following: Figure 1 , Figure 1This is a flowchart illustrating an embodiment of the adaptive dynamic audio processing method provided by the present invention. The adaptive dynamic audio processing method provided by the present invention is applied to a microphone array, which includes multiple microphones, each of which can collect sound and convert the sound into a corresponding audio signal for transmission.

[0019] The adaptive dynamic audio processing method includes the following steps: S101: Acquire a reference signal and the microphone signal of each microphone in the microphone array, and perform frame segmentation on the reference signal and each microphone signal to obtain at least one reference signal frame and at least one microphone signal frame.

[0020] In a specific implementation scenario, suppose the microphone array includes The system uses three normally functioning microphones to simultaneously acquire microphone signals from each microphone. The reference signal is the clean, distant speech signal that will be played through the local speaker; it can be directly extracted from the data stream transmitted from the remote communication terminal. The reference signal is the original echo template, from which the echo signal caused by the echo can be calculated later. The microphone signal is the time-domain signal independently acquired by each microphone in the multi-channel microphone array, and can be denoted as... This refers to the microphone's number within the microphone array. The microphone signal includes not only the near-end speech signal generated by the user's speech, but also the echo signal formed by far-end speech reflected from the room and background noise signals from the environment.

[0021] Strict time synchronization calibration is performed on the reference signal and each microphone signal. Then, a unified framing strategy is used to perform framing processing on the synchronized reference signal and all microphone signals, dividing the continuous time-domain signal into discrete, overlapping short-time signal segments, obtaining at least one reference signal frame and at least one microphone signal frame time-domain aligned with the reference signal frame, forming a synchronized reference signal frame sequence. and microphone signal frame sequence , This is the frame number.

[0022] In this implementation scenario, each signal frame (including the reference signal frame and the microphone signal frame) has a frame length of 20ms and a frame shift of 10ms. The frame shift is 50% of the frame length to ensure the overlap between adjacent signal frames, avoid temporal breakpoints after signal framing, and improve processing continuity.

[0023] S102: For each microphone signal frame, calculate its microphone signal energy. For each reference signal frame, obtain its reference signal energy. Compare the reference signal energy at the current moment with the far-end energy threshold to obtain a first comparison result. Compare the microphone signal energy with the near-end energy threshold to obtain a second comparison result. Compare the reference signal energy with the microphone signal energy to obtain a third comparison result.

[0024] In a specific implementation scenario, for each microphone signal frame and reference signal frame obtained after framing, pre-emphasis processing and STFT (short-time Fourier transform) are sequentially performed according to a unified signal processing flow, converting the short-time discrete signal frame in the time domain into a frequency domain signal containing two-dimensional time-frequency features. It is important to note that the pre-emphasis coefficients, STFT window function, number of transform points, and other parameters of all signal frames must be strictly consistent to eliminate frequency domain feature misalignment caused by parameter differences.

[0025] Pre-emphasis processing can compensate for the high-frequency attenuation of the speech signal during pronunciation and transmission, increasing the proportion of high-frequency components in the signal. Since both the reference signal and the microphone signal include speech signals, a pre-emphasis filter is applied to each frame of the signal to enhance the high-frequency components of the speech and compensate for the high-frequency attenuation during speech transmission.

[0026] Perform STFT on each pre-emphasized signal frame, setting the window function (such as a Hamming window), the window length to be consistent with the frame length, and the overlap rate to 50% (determined by a frame shift of 10ms), converting the time-domain signal into a two-dimensional frequency-domain signal: frequency-domain microphone signal frame. and frequency domain reference signal frame .

[0027] In one implementation scenario, frequency domain vocal cord energy calculation can be performed to extract the far-end speech signal from the reference signal. Frequency points in the 150Hz~3.4kHz band are extracted from the reference signal frame sequence, the energy of this band is accumulated and converted to dB values, and 50Hz power hum and quantization noise above 4kHz are filtered out. The presence of a far-end speech signal is determined based on the frequency domain vocal cord energy. If no far-end speech signal is present, it means that no echo signal will be generated, and echo removal processing of the microphone is not required. Specifically, the calculation can be performed using the following formula:

[0028] in, For reference signal frame, This represents the energy of the vocal cords in the frequency domain.

[0029] A fixed threshold can be set. ,like If the signal is positive, then the presence of a distant speech signal is considered to exist; otherwise, the absence of a distant speech signal is considered to be absent.

[0030] Alternatively, exponential smoothing tracking can be used. In recent m (For example, m Minimum energy within 2 seconds The smoothing coefficient is set between 0.9 and 0.98; if If the signal is positive, then the presence of a distant speech signal is considered to exist; otherwise, the absence of a distant speech signal is considered to be absent. The value can be set according to actual needs, for example, 12dB.

[0031] The energy of the reference signal frame is calculated and converted to a dB value. The reference signal energy can be calculated using the following formula:

[0032] in, This refers to the number of sampling points per frame. For example, at a sampling rate of 16kHz, a 20ms frame length corresponds to N=320. For reference signal energy, This is the reference signal frame.

[0033] The reference signal energy is compared with the far-end energy threshold to obtain the first comparison result.

[0034] The microphone signal energy of the microphone signal frame is calculated using a similar algorithm to that used for the reference signal.

[0035] in, The number of sampling points is the frame length. For microphone signal energy, This is a microphone signal frame.

[0036] The microphone signal energy is compared with the near-end energy threshold to obtain a second comparison result.

[0037] Calculate the correlation coefficient between the microphone signal frame and the reference signal frame corresponding to the same frame. When only the far end is emitting sound, the reference signal will form an echo through acoustic paths such as room reflection and air propagation. This echo will be mixed into the microphone signal as a core component. If the near end is not emitting sound at this time, that is, there is no near-end speech signal, the waveform of the microphone signal and the waveform of the reference signal will be highly similar, and their normalized cross-correlation coefficient will also be at a high level. However, when the near end is emitting sound at the same time, the microphone signal will be superimposed with a near-end speech signal component that is completely unrelated to the reference signal. This component will destroy the similarity of the original waveform, thus causing a significant decrease in the cross-correlation coefficient between the two.

[0038] Specifically, the correlation coefficient between the reference signal energy and the microphone signal energy is calculated using the following formula:

[0039] in, The value range of is [0,1]. The closer the correlation coefficient is to 1, the higher the proportion of echo signal in the microphone signal. A third comparison result is obtained by comparing the correlation coefficient with a correlation threshold (e.g., 0.75).

[0040] S103: Based on the first comparison result, the second comparison result, and the third comparison result, obtain the application scenario at the current moment. The application scenario includes any one of remote dominance, near-end dominance, dual-talk, and mute.

[0041] In a specific implementation scenario, if the first comparison result shows that the reference signal energy is lower than the far-end energy threshold, the current application scenario is determined to be near-end dominant. In this case, the far end is silent, and the microphone signal is dominated by near-end speech and / or noise. If the first comparison result shows that the reference signal energy is higher than or equal to the far-end energy threshold, and the third comparison result shows that the correlation coefficient exceeds the high correlation threshold (e.g., 0.75), the current application scenario is determined to be far-end dominant, and the microphone signal is dominated by echo. If the first comparison result shows that the reference signal energy is higher than or equal to the far-end energy threshold, and the third comparison result shows that the correlation coefficient is lower than the low correlation threshold (e.g., 0.4), and the second comparison result shows that the microphone signal energy is higher than or equal to the near-end energy threshold, the current application scenario is determined to be dual-talk, where both the far and near ends speak simultaneously. In other cases, the current application scenario can be determined to be silent.

[0042] S104: Obtain the estimated echo signal frame based on the reference signal frame, subtract the estimated echo signal frame from the microphone signal frame corresponding to the reference signal frame to obtain the residual signal frame, and obtain the residual signal energy of the residual signal frame.

[0043] In a specific implementation scenario, an acoustic path model is performed on the reference signal frame using a foreground filter to generate a predicted echo signal frame with the same amplitude but opposite phase as the actual echo in the microphone signal. The foreground filter calls the filter coefficient matrix adapted to the current acoustic environment and processes the reference signal frame through convolution operations to obtain the predicted echo signal frame.

[0044] The estimated echo signal frame completely cancels out the actual echo in the microphone signal. In this case, the residual signal frame contains only near-end speech and background noise. If the acoustic environment changes abruptly (e.g., the echo path changes), a small amount of uncancelled echo components will remain in the residual signal frame. The average energy is obtained by summing the squares of all samples within the residual signal frame and dividing by the frame length. Specifically, the residual signal energy can be calculated using the following methods:

[0045] in, For residual signal energy, For residual signal frames, This represents the number of sampling points per frame.

[0046] In one embodiment, the foreground filter is optimized. Based on a synchronized reference signal frame, the minimum mean square error (LMS) or normalized minimum mean square error (NLMS) algorithm is used to iteratively update the filter coefficient matrix with the goal of minimizing the mean square error of the residual signal after cancellation. Specifically, the reference signal frame can be input into the foreground filter to generate a predicted echo signal with amplitude matching and phase opposite to the echo component in the microphone signal; secondly, the predicted echo signal is superimposed and canceled with the microphone signal frame to obtain the residual signal after echo cancellation; finally, the mean square error between the residual signal and the ideal output signal (echo-free near-end speech + environmental noise) is continuously calculated, and the weight distribution of the filter coefficient matrix is ​​dynamically adjusted based on the gradient descent direction of this error until the mean square error converges to a preset range, thereby achieving accurate cancellation of the echo of the current acoustic path.

[0047] The update rate and convergence speed of the foreground filter are modified based on the application scenario to balance echo cancellation effect and near-end speech fidelity. For example, when the application scenario is far-end single-talk, the update rate and convergence speed of the filter coefficients are increased to prioritize fast and deep echo cancellation; when the application scenario is two-talk, the update rate of the filter coefficients is reduced to limit the adjustment range of the coefficients and prevent the foreground filter from misjudging near-end speech components that are unrelated to the reference signal as echoes and suppressing them, thus ensuring the naturalness of near-end speech; when the application scenario is near-end single-talk / silence, the update of the filter coefficients is slowed down or even paused to prevent background noise from being mismodeled as echoes and causing speech distortion.

[0048] The background filter collects static noise (such as air conditioner operation noise and power supply noise). It continuously analyzes changes in static noise and echo paths (such as changes in room reflection coefficient and microphone position fine-tuning) within the current acoustic environment. Without requiring additional independent noise signal acquisition, it directly extracts the spectral features of static noise from the microphone signal frame and constructs a high-precision noise baseline model based on this feature data. This model is then fed back to the foreground filter in real time, providing it with noise feature references. This allows the foreground filter to accurately distinguish between echo signals and static noise signals during iterative coefficient correction, preventing the misidentification of static noise as echo and its subsequent cancellation. This, in turn, prevents distortion and muffled sound in the output speech.

[0049] The background filter receives and monitors the mean square error (MSE) data output by the foreground filter in real time. Based on the comparison between the error and a preset threshold, it executes differentiated processing strategies. If the MSE remains stable within the preset threshold, it is determined that the current acoustic environment has not changed significantly, and the background filter only updates the noise baseline model and the backup filter coefficient library at a very low rate to ensure system stability. If the MSE continuously exceeds the preset threshold (indicating a sudden change in the acoustic environment, such as a change in echo path or a sudden increase in noise intensity, meaning the current filter coefficient matrix is ​​no longer suitable for the new environment), the background filter immediately accelerates its learning of the acoustic characteristics of the new environment, generates a candidate filter coefficient matrix adapted to the new environment, and seamlessly switches the current filter coefficient matrix of the foreground filter to this candidate matrix, achieving rapid recovery of the echo cancellation effect and ensuring robust system operation in complex environments.

[0050] In other implementation scenarios, based on the characteristics of human hearing and the laws of echo propagation, the entire frequency domain (e.g., 20Hz~8kHz) is divided into multiple sub-bands. The sub-band type can be overlapping or non-overlapping (typically 16 sub-bands, with frequency intervals such as 125Hz, 250Hz…8kHz). Different frequency bands correspond to different echo characteristics. For example, the low-frequency sub-band (20Hz~500Hz) has slow echo reflection attenuation and is prone to forming continuous echoes, so it is important to ensure the integrity of the cancellation; the mid-to-high frequency sub-band (1kHz~8kHz) has fast echo attenuation, but is easily interfered with by sudden noises such as keyboard sounds, so it is necessary to balance convergence speed and anti-interference.

[0051] Frequency point data corresponding to each sub-band is extracted from the full-frequency domain signal, and the microphone signal frame is divided into multiple sub-microphone signal frames based on the multiple sub-bands. A corresponding foreground filter is set for the frequency band of each sub-band, and the filter coefficients of each foreground filter are matched with the corresponding sub-band frequency. For example, high-order filter coefficients (such as 256th order) are used for the low-frequency sub-band to ensure the capture of complete phase information of the echo and avoid low-frequency echo residue; low-order filter coefficients (such as 64th order) are used for the mid-high frequency sub-band to reduce computational complexity and speed up coefficient convergence.

[0052] Similarly, the foreground filter corresponding to each sub-band independently iteratively optimizes its own filter coefficient matrix with the goal of achieving either minimum mean square error (LMS) or normalized minimum mean square error (NLMS). The specific optimization process is basically the same as described above and will not be repeated here.

[0053] Each sub-band's corresponding foreground filter takes the sub-microphone signal frame of the corresponding frequency band as input and generates a sub-estimated echo signal frame with the same amplitude but opposite phase as the actual echo of that sub-band through filtering operations. The sub-microphone signal frames of the same sub-band are superimposed and canceled out with the sub-estimated echo signal frames to obtain the sub-residual signal frame of that sub-band. Ideally, this sub-residual signal frame has eliminated the echo components of the corresponding frequency band, retaining only near-end speech and background noise. All sub-residual signal frames are superimposed and stitched according to frequency band characteristics to generate a complete residual signal frame covering the entire frequency domain. For non-overlapping sub-bands, the frequency point intervals have no overlap, and the sub-residual signal frames can be directly filled into the corresponding intervals of the residual signal frame in sequence to obtain the complete residual signal. For overlapping sub-bands, the frequency point intervals of adjacent sub-bands overlap; direct stitching will cause energy superposition distortion in the overlapping area, requiring a weighted fusion strategy. For frequency points in the overlapping region, calculate the weighted average of two adjacent sub-residual signal frames. The weights can be allocated according to the sub-band overlap ratio (e.g., if there is 50% overlap, the weights are 0.5 for each).

[0054] S105: Calculate the corresponding microphone performance indicators based on the application scenario and residual signal energy, and obtain the corresponding microphone indicator score based on the application scenario and performance indicators.

[0055] In a specific implementation scenario, the corresponding microphone performance metrics are calculated based on the application scenario and residual signal energy. These performance metrics include signal-to-noise ratio (SNR) and ERLE (Echo Return Loss Enhancement). ERLE measures the additional enhancement effect of echo cancellation on the original echo path attenuation, representing the dB energy reduction of the echo signal after acoustic echo cancellation processing.

[0056] First, it is necessary to determine the specific content of the current performance indicators based on the application scenario, and calculate ERLE only in scenarios where the speaker is speaking from a distance (remote-led, two-way).

[0057] In an implementation scenario where the current application is remote-controlled or dual-talk, it is necessary to calculate the ERLE value. A minimal ε is introduced to avoid division by zero, and the ERLE value (in dB) is calculated; a higher value indicates better echo cancellation. A first-order IIR filter is used to smooth the ERLE value and reduce fluctuations. The ERLE can be calculated using the following algorithm:

[0058]

[0059] in, For microphone signal energy, For residual signal energy, The original ERLE value, The smoothed ERLE value. This is the smoothed ERLE value of the previous frame.

[0060] In other implementation scenarios, ERLE can be calculated using the following formula:

[0061] in, For the first The microphone The microphone signal energy of a frame. For the first The microphone The residual signal of a frame. The energy of the residual signal of a frame. It represents the ratio of the total energy before echo cancellation to the residual energy after cancellation.

[0062] In another implementation scenario, SNR is calculated. First, the noise power is dynamically estimated. It can be calculated using the following formula:

[0063] Wherein, VAD represents the Voice Activity Detection result. VAD=0 indicates that there is no near-end speech in the current microphone signal frame, indicating a silent state, and only background noise exists in the microphone signal frame. First-order IIR smoothing is used to update the noise power. A weight of 98% is used ( Inherit the noise power estimate from the previous frame. To ensure the stability of the noise reference; a weighting of 2% is used ( Absorb the microphone signal power of the current frame It tracks slow changes in noise (such as small fluctuations in air conditioning noise or electrical noise); thus enabling the noise baseline to be updated only in silent microphone signal frames, ensuring that the noise power does not include speech energy.

[0064] VAD=1 indicates the presence of near-end speech in the current microphone signal frame. The noise power estimate from the previous frame is directly used. No updates are made. This is to avoid misclassifying speech energy as noise, which would lead to an overestimation of noise power and ultimately distort the SNR calculation.

[0065] Then calculate the signal-to-noise ratio (SNR) using the following formula.

[0066]

[0067] in, For the first The microphone The microphone signal energy of a frame (including the total energy of speech, noise, and echo). For the estimated noise power Converting the power ratio to decibels (dB) aligns with the measurement conventions in the acoustics field (the human ear perceives sound intensity on a logarithmic scale, and dB is closer to the actual auditory experience).

[0068] In other implementation scenarios, the core purpose of performing linear normalization and truncation constraints on ERLE and SNR is to uniformly map the original index values ​​with different dimensions and effective ranges to the [0,1] interval, eliminate the range differences between the indicators, and facilitate the subsequent calculation of channel comprehensive scores.

[0069]

[0070] The original ERLE values ​​are linearly mapped from the [6,20] dB interval to the [0,1] interval.

[0071]

[0072] The original SNR value is linearly mapped from the [10,30] dB interval to the [0,1] interval.

[0073] In one implementation scenario, a quiet threshold can be set first. The residual signal energy of a microphone in the current frame is then checked against this threshold. If it is below the threshold, the microphone is considered to have failed to receive a valid sound signal and is therefore excluded from the selection of the best microphone. The best microphone is then chosen from among those microphones capable of receiving valid sound signals.

[0074] Assign corresponding weights to ERLE and SNR respectively. For example, the first weight for ERLE is... The second weight corresponding to SNR Normalized and The scores are obtained by multiplying each score by its corresponding weight value.

[0075] Furthermore, a metric score is calculated for each microphone based on different application scenarios. For example, it can be calculated using the following formula:

[0076] In far-end dominant scenarios, ERLE's primary weight is higher than SNR's secondary weight, with deep echo cancellation being the top priority and speech clarity being a secondary requirement. In near-end dominant scenarios, ERLE's primary weight is lower than SNR's secondary weight, ensuring near-end speech clarity is the top priority, and echo cancellation is a secondary requirement. In two-way conversation scenarios, ERLE's primary weight equals SNR's secondary weight; echo cancellation and speech clarity are equally important, requiring a balance between their performance to avoid excessive near-end speech suppression. In silent scenarios, noise suppression effectiveness is the sole evaluation criterion, and echo cancellation is meaningless.

[0077] S106: Select the optimal microphone in the microphone array based on the index score of each microphone, obtain the transmission signal based on the residual signal frame of the optimal microphone, and output the transmission signal.

[0078] In a specific implementation scenario, the index score of each microphone is calculated based on the above steps. The microphone with the highest index score is selected as the optimal microphone. The transmission signal is obtained from the residual signal frame of the optimal microphone and then output.

[0079] Specifically, an inverse pre-emphasis filter can be applied to the residual signal frame to cancel the pre-emphasis effect in the previous steps and restore the natural spectral characteristics of the speech. An ISTFT is then performed on the deemphasized frequency domain signal to restore it to the time domain. The final time-domain speech signal is output, ensuring that the processing latency meets the requirements of real-time voice communication / interaction.

[0080] In other implementation scenarios, since residual signal frames are extracted for each microphone in order to select the optimal microphone, a feature vector is then constructed based on the microphone signal frames, residual signal frames, and reference signal frames of each microphone in the microphone array, as well as preset performance indicators (such as ERLE and SNR) for the microphone. A combination of feature threshold screening and machine learning classifier is used to accurately determine the current voice interaction scenario classification (near-end voice dominance / far-end voice dominance / dual-talk / steady-state noise / burst noise).

[0081] First, extract basic features from the microphone signal frame / reference signal frame: Energy characteristics: short-time energy mean of each microphone signal (mean of the sum of squares of sampling points within a 20ms frame), energy fluctuation amplitude (standard deviation of energy over 10 consecutive frames), and near-end / far-end energy ratio (ratio of microphone signal frame energy to reference signal frame energy).

[0082] Spectral characteristics: spectral flatness of the signal (the ratio of the geometric mean to the arithmetic mean of the full-band spectral amplitude, with noise approaching 1 and speech approaching 0), high-frequency energy ratio (the proportion of energy above 2kHz to the total energy of the full-band), and sub-band energy distribution (dividing the signal into 16 sub-bands at 125Hz intervals and calculating the energy ratio of each sub-band, such as the proportion of low-frequency noise below 200Hz).

[0083] Temporal characteristics: signal duration (cumulative duration of continuous active frames), energy surge magnitude (the ratio of the energy of the current frame to the average energy of the last 50 frames; this value is ≥3 in sudden noise scenarios), and activity interval (the alternation frequency of speech frames and silent frames; there is a clear interval in speech scenarios, and no interval in steady-state noise).

[0084] Channel difference characteristics: energy difference coefficient between multiple microphone channels (the deviation ratio of each channel's energy from the average energy), correlation coefficient between the reference signal and each microphone signal (≥0.8 in the far-end dominant scenario, ∈[0.3,0.7] in the dual-talk scenario).

[0085] After standardizing the real-time ERLE value (reflecting echo cancellation effect) and SNR value (reflecting speech clarity and noise suppression effect) of each microphone, they are added to the feature vector as a supplementary dimension to form a composite feature representation of basic features + performance feedback, thereby improving the robustness of scene judgment.

[0086] First, the feature vectors are filtered using preset empirical thresholds (e.g., spectral flatness > 0.6 indicates a noisy scene, and near-end / far-end energy ratio > 5 indicates near-end dominance). Invalid data that clearly does not fit a certain scene type is removed, reducing the computational load of the classifier and lowering the probability of misclassification. The filtered feature vectors are then input into a trained machine learning classifier (lightweight models such as Support Vector Machine (SVM) and Random Forest are commonly used in engineering, balancing accuracy and real-time performance). Based on the "scene-feature" mapping relationship learned during the training phase, the classifier outputs the judgment result of the current scene (near-end speech dominance / far-end speech dominance / dual-talk / steady-state noise / burst noise).

[0087] Based on the optimal scenario classification, matching optimization algorithms are selected from the algorithm pool. The algorithm pool is functionally divided into four sub-libraries: echo cancellation, noise suppression, speech enhancement, and channel processing. Each sub-library contains differentiated algorithms adapted to different scenarios, and the matching of scenarios and algorithms follows the principle of prioritizing core needs. The algorithms in the algorithm pool mainly include: Far-end speech-dominated: Fast and deep echo cancellation, with the corresponding optimization algorithm being a fast convergence echo cancellation algorithm + high-frequency enhancement algorithm; Near-end speech-dominated: Improve speech clarity and suppress noise, with the corresponding optimization algorithm being a steady-state noise suppression algorithm + low signal-to-noise ratio speech enhancement algorithm; Dual-talk: Balance echo cancellation and near-end speech protection, with the corresponding optimization algorithm being a dual-talk protection echo cancellation algorithm + speech restoration algorithm; Steady-state noise: Suppress continuous background noise, with the corresponding optimization algorithm being a steady-state noise suppression algorithm + energy equalization algorithm; Burst noise: Eliminate transient interference and complete speech, with the corresponding optimization algorithm being a burst noise suppression algorithm + speech restoration algorithm.

[0088] Call the historical performance data corresponding to the candidate algorithm (such as whether ERLE is ≥15dB and SNR is ≥10dB in the last 100ms). If the performance meets the standard, the algorithm is directly enabled. If the candidate algorithm fails to meet the performance standard due to sudden environmental changes (such as a sudden increase in noise), immediately call the downgraded robust algorithm from the algorithm pool (such as switching to the robust echo cancellation algorithm when the dual-talk protection algorithm fails), and at the same time fine-tune the algorithm parameters (such as increasing the noise suppression threshold) until the performance recovers.

[0089] When switching algorithms, an overlapping frame transition strategy is adopted, which weights and superimposes the last 10ms frame of the previous algorithm with the first 10ms frame of the new algorithm to control the switching delay within 5ms and avoid dropouts and pops in the output signal. The matching optimization algorithm is enabled to perform targeted processing (such as echo suppression, noise filtering, high frequency enhancement, etc.) on the residual signal frame of the optimal microphone, and finally outputs a high-fidelity voice signal.

[0090] In other implementation scenarios, after obtaining the matching optimization algorithm, the optimization algorithm is used to suppress residual echoes in the residual signal frame, further eliminating residual echoes that are not completely canceled in the residual signal frame (such as acoustic path abrupt changes, echo components missed in two-way scenarios), and avoiding echo interference from far-end speech to near-end listening or transmission quality. After secondary suppression, the residual echo energy in the residual signal frame is reduced to below the noise level, resulting in an echo-suppressed frame.

[0091] Based on the noise type corresponding to the scene, the noise energy is accurately estimated, and environmental noise (steady-state noise such as air conditioner noise, sudden noise such as keyboard typing) in the echo-suppressed frame is filtered out, improving the signal-to-noise ratio of the speech signal. Noise energy can be estimated using spectral subtraction (adapted to steady-state noise scenes) or deep learning-based methods (adapted to complex / sudden noise scenes). The frequency domain signal amplitude of the echo-suppressed frame is compared with the estimated noise energy. If the signal amplitude is less than or equal to the noise energy, the signal amplitude of that frequency band is attenuated proportionally (attenuation coefficient 0.3~0.5); otherwise, the amplitude remains unchanged or is slightly enhanced (enhancement coefficient 1.0~1.2). With noise energy accurately filtered out and the spectral characteristics of the speech signal fully preserved, the noise-suppressed frame is obtained.

[0092] To address the issue of speech energy attenuation in different scenarios, adaptive gain compensation is performed to ensure uniform loudness and clearness of the speech signal, and to avoid sudden volume changes caused by scene switching. This results in gain compensation frames with uniform speech loudness, balanced energy across the entire frequency band, and no sudden volume changes.

[0093] Continuous gain-compensated frames (each 20ms, including a 50% overlap area) are spliced ​​sequentially in the time domain to avoid audio dropouts and pops caused by frame breaks, ensuring the continuity and smoothness of the transmitted signal. The spliced ​​continuous time-domain signal is then standardized according to transmission protocol requirements (e.g., 16kHz sampling rate, 16-bit bit depth, mono / stereo), ultimately yielding a directly transmittable signal.

[0094] In other implementation scenarios, the performance score of each microphone in the microphone array is calculated in real time and in parallel according to the period consistent with the signal framing (e.g., one frame every 20ms). When the performance score of the currently optimal microphone is found to be consistently low, the optimal microphone needs to be switched in a timely manner to ensure stable voice transmission quality. Under the premise of ensuring system output stability, the system dynamically switches to the microphone channel with the best performance in the current acoustic environment to avoid voice distortion and stuttering caused by instantaneous signal fluctuations or frequent switching.

[0095] Specifically, from all microphones in the microphone array, the currently used optimal microphone is removed, and all remaining microphone channels are marked as candidate microphones as potential switching options. The real-time metric score of the currently optimal microphone is extracted and defined as the target score, which is used as the benchmark for performance comparisons of all subsequent candidate microphones.

[0096] A pre-defined score difference threshold is set, which is the minimum difference between the candidate microphone score and the target score. This threshold is used to filter candidate microphones with significantly better performance. A short-term score cache queue is maintained for each microphone (e.g., caching the metric scores within the last 500ms) for subsequent duration verification, avoiding misjudgments caused by sudden changes in scores within a single frame.

[0097] In each frame, a score comparison is performed on all candidate microphones one by one, calculating the difference between the candidate microphone's score and the target score. If the score difference exceeds a preset score difference threshold, the candidate microphone is marked as a microphone to be verified. If the score difference of all candidate microphones is less than the preset score difference threshold, it is determined that there are no candidate microphones that meet the conditions in this round, and the verification process for the current frame ends directly. All microphones that meet the score difference threshold condition are integrated into the target microphone set.

[0098] For each microphone in the target microphone set, an independent duration timer is started from the moment it first meets the score difference threshold condition to obtain the duration of the score advantage. If the score difference value of the microphone falls below the score difference threshold in a certain frame, the timer is immediately reset and restarted. When the duration of the score advantage reaches the first duration threshold, the duration of the score advantage of each candidate microphone in the target microphone set is checked to see if it exceeds the first duration threshold. Candidate microphones with a duration of score advantage below the first duration threshold are removed. Brief increases in score caused by transient acoustic interference (such as sudden door slamming or keyboard typing) are filtered out to ensure that the performance advantage of the target microphone is stable and genuine.

[0099] Maintain an optimal microphone in-service timer, starting from the moment the last optimal microphone switch was successful. The timer continuously accumulates the in-service time of the current optimal microphone until the next switch occurs. Determine if the in-service duration of the current optimal microphone is greater than or equal to a preset second duration threshold. If the in-service duration of the current optimal microphone does not reach the second duration threshold, terminate the current switch process and maintain the current optimal microphone. If the in-service duration of the current optimal microphone reaches the second duration threshold, select the microphone with the highest metric score from all remaining candidate microphones in the current target microphone set and define it as the new optimal microphone. Set a stability protection period for the current optimal microphone to prevent switching again shortly after a switch because the target microphone meets the conditions, thus preventing the system from frequently changing the optimal channel in a short period and causing issues such as stuttering and popping in voice output.

[0100] Furthermore, reset the in-situ timer for the best microphone, update the timing start to the moment the switch was successful, clear the duration timers for all target microphones, and clear the short-term score cache queue for candidate microphones.

[0101] In other implementation scenarios, although the optimal channel has been selected as the final signal output carrier in a multi-channel audio system, comprehensive equalization of the energy, delay, and frequency response of the remaining candidate channels is a fundamental requirement to ensure stable system operation. Equalization ensures that candidate channels maintain consistent energy levels, temporal synchronization, and frequency response characteristics with the optimal channel, laying the foundation for seamless channel switching in the event of changes in the acoustic environment and preventing issues such as sudden volume changes, timbre distortion, or core algorithm failure after switching. It also eliminates performance deviations in the microphone hardware itself, allowing the feature comparison data of the multi-channels to reflect the differences in the real acoustic environment, supporting the accuracy of upper-level scene recognition. Furthermore, it ensures the robustness of algorithms such as multi-channel noise suppression and beamforming, improving the processing effect of the optimal channel signal through equalized multi-channel data. In addition, the equalized candidate channels can serve as redundant backup resources with uniform performance, quickly taking over in case of optimal channel hardware failure, and facilitating resource optimization and scheduling within the system, ultimately achieving intelligent and robust operation of the multi-channel audio system.

[0102] First, the residual signal energy of the residual signal frame of the currently selected optimal microphone is used as the energy benchmark value. Simultaneously, the remaining residual signal energies of all other candidate microphone residual signal frames in the microphone array are extracted. Based on this energy benchmark value and the remaining residual signal energies, the signal gain of each candidate microphone is dynamically adjusted. The adjustment range is determined using the gain calculation formula (candidate microphone target gain = energy benchmark value ÷ current residual signal energy of candidate microphone) to ensure that the residual signal energy of each candidate microphone reaches 90%-110% of the optimal microphone energy benchmark value, achieving multi-channel energy consistency. At the same time, strict gain constraints are set, controlling the maximum gain upper limit to within 3 times to prevent low-energy candidate microphones from amplifying background noise due to excessive gain, ensuring that the signal-to-noise ratio of the adjusted signal is not compromised.

[0103] Secondly, relying on the system's pre-stored microphone frequency response compensation coefficient table (which is pre-calibrated using a standard sweep signal: a sweep signal of 20Hz~8kHz is played for each microphone, the frequency response deviation at each frequency point is recorded, and a one-to-one mapping relationship between frequency and compensation coefficient is established), frequency response deviation compensation is performed on the residual signal frames of each candidate microphone and the optimal microphone. The frequency components of the residual signal are analyzed in real time, and the accurate compensation coefficients for the corresponding frequencies are called from the compensation coefficient table according to the current frequency band distribution of the signal. The residual signals of each microphone are corrected in the frequency domain by using a linear phase FIR filter (to avoid introducing signal phase distortion). The core objective of the compensation is to ensure that the frequency response curves of all microphones are highly consistent with the frequency response characteristics of the optimal microphone, control the frequency response deviation at each frequency point within ±1dB, eliminate spectral distortion caused by differences in microphone hardware, and ensure the tonal uniformity of multi-channel signals.

[0104] Finally, using the optimal microphone as the time-domain synchronization reference channel, the time-domain delay difference between each candidate microphone and the optimal microphone residual signal frame is accurately calculated through a cross-correlation analysis algorithm. A sliding cross-correlation operation is performed on the two types of residual signal frames to find the time offset corresponding to the maximum value of the cross-correlation coefficient. This offset is the actual delay difference between the two (in microseconds). Based on the calculated delay difference, a corresponding delay compensation frame is inserted into the residual signal frame of the candidate microphone. If the delay difference of the candidate microphone is Δt, an additional compensation frame of duration Δt is buffered before its residual signal frame to ensure that the residual signal of the candidate microphone is completely aligned with the residual signal of the optimal microphone in the time domain. Finally, the delay difference of the residual signal frames of all microphones (optimal + candidate) is controlled within 1μs to ensure the time-domain synchronization of multi-channel signals during subsequent scene recognition, algorithm scheduling, and channel switching, and to avoid algorithm performance degradation or signal discontinuity during switching caused by phase deviation.

[0105] As described above, this embodiment processes the reference signal and microphone signal in frames. By comparing the energy of the reference signal, the energy of the microphone signal, and the correlation coefficient between the two with the corresponding threshold in real time, it accurately identifies dynamic acoustic scenarios such as far-end dominance, near-end dominance, dual-talk, and silence. This breaks away from the single and fixed judgment logic of traditional methods. At the same time, relying on the real-time performance evaluation and optimal channel selection mechanism of the multi-channel microphone array, combined with channel equalization processing of energy, delay, and frequency response, it can adapt to environmental changes such as sound source movement, noise abrupt changes, and changes in acoustic path in real time. This ensures that no matter how the acoustic environment changes dynamically, the system can automatically switch to the optimal processing strategy and microphone channel to maintain stable audio output quality.

[0106] Figure 2 A schematic diagram of the internal structure of an audio transmission and playback terminal in one embodiment is shown. Figure 2 As shown, the audio transmission and playback terminal includes a processor, a memory, and a network interface connected via a system bus. The memory includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores an operating system and may also store a computer program. When executed by the processor, this computer program enables the processor to implement an adaptive dynamic audio processing method. The internal memory may also store a computer program, which, when executed by the processor, enables the processor to execute an application of the adaptive dynamic audio processing method. Those skilled in the art will understand that... Figure 2 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the electronic device to which the present application is applied. The specific electronic device may include more or fewer components than shown in the figure, or combine certain components, or have different component arrangements.

[0107] In one embodiment, an audio transmission and playback terminal is provided, including a memory and a processor. The memory stores a computer program, which, when executed by the processor, causes the processor to perform the steps described above.

[0108] In one embodiment, a computer-readable storage medium is provided storing a computer program that, when executed by a processor, causes the processor to perform the steps described above.

[0109] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments described above. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), RAMbus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and RAMbus dynamic RAM (RDRAM), etc.

[0110] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0111] It should be understood that the above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit them. Those skilled in the art can modify the technical solutions described in the above embodiments, or make equivalent substitutions for some of the technical features; and all such modifications and substitutions should fall within the protection scope of the appended claims of the present invention.

Claims

1. An adaptive dynamic audio processing method, characterized in that, Applied to a microphone array, the microphone array comprising multiple microphones; The adaptive dynamic audio processing method includes: Acquire a reference signal and the microphone signal of each microphone in the microphone array, and perform frame segmentation on the reference signal and each microphone signal to obtain at least one reference signal frame and at least one microphone signal frame; For each frame of the microphone signal, calculate its microphone signal energy; for each frame of the reference signal, obtain its reference signal energy; compare the reference signal energy at the current moment with the far-end energy threshold to obtain a first comparison result; compare the microphone signal energy with the near-end energy threshold to obtain a second comparison result; compare the correlation coefficient between the reference signal energy and the microphone signal energy with the correlation threshold to obtain a third comparison result. Based on the first comparison result, the second comparison result and the third comparison result, the application scenario at the current moment is obtained, and the application scenario includes any one of remote dominance, near-end dominance, dual-talk and mute; Based on the reference signal frame, obtain the estimated echo signal frame, subtract the estimated echo signal frame from the microphone signal frame corresponding to the reference signal frame to obtain the residual signal frame, and obtain the residual signal energy of the residual signal frame. The performance index of the microphone is calculated based on the application scenario and the residual signal energy, and the corresponding index score of the microphone is obtained based on the application scenario and the performance index. The optimal microphone in the microphone array is selected based on the index score of each microphone, and the transmission signal is obtained from the residual signal frame of the optimal microphone and then output.

2. The adaptive dynamic audio processing method according to claim 1, characterized in that, The step of selecting the optimal microphone in the microphone array based on the index score of each microphone includes: The other microphones in the microphone array are used as candidate microphones, and the index score of the best microphone is used as the target score. The index score of each microphone in the microphone array is acquired in real time, and it is determined whether there is at least one candidate microphone whose index score exceeds the preset score difference threshold of the target score. If so, at least one candidate microphone whose index score exceeds the target score will be selected as the target microphone. Determine whether the duration for which the target microphone's score exceeds the target score exceeds a preset first duration threshold. If so, determine whether the duration for which the optimal microphone is used as the optimal microphone exceeds the second duration threshold. If so, the microphone with the highest index score among the target microphones will be selected as the new optimal microphone.

3. The adaptive dynamic audio processing method according to claim 1, characterized in that, The step of obtaining the estimated echo signal frame based on the reference signal frame includes: The filter coefficient matrix of the foreground filter is corrected based on the corresponding reference signal frame and the microphone signal frame using the minimum mean square error algorithm or the normalized minimum mean square error algorithm. The update rate and convergence speed of the foreground filter are modified based on the application scenario. Collect static noise, establish a noise baseline model based on the spectral characteristics of the static noise, and feed the noise baseline model back to the foreground filter; The mean square error of the foreground filter is monitored in real time. When the mean square error exceeds a preset error threshold, the filter coefficient matrix of the foreground filter is replaced with a candidate filter coefficient matrix. The foreground filter is used to obtain the estimated echo frame corresponding to the reference signal frame.

4. The adaptive dynamic audio processing method according to claim 3, characterized in that, The step of obtaining the estimated echo signal frame corresponding to the reference signal frame through the foreground filter includes: The full frequency domain is divided into multiple sub-bands, and the microphone signal frame is divided into multiple sub-microphone signal frames according to the multiple sub-bands; The plurality of sub-signal frames are respectively input into the corresponding foreground filters, and the filter coefficient matrix of the foreground filters is adjusted based on the frequency band corresponding to the currently input sub-microphone signal frame; Obtain the sub-estimated echo signal frame corresponding to each sub-microphone signal frame, and obtain the sub-residual signal frame corresponding to each sub-microphone signal frame based on the sub-estimated echo signal frame; The sub-residual signal frames are spliced ​​together to form the residual signal frame.

5. The adaptive dynamic audio processing method according to claim 1, characterized in that, After obtaining at least one reference signal frame and at least one microphone signal frame, the steps include: Each microphone signal frame and / or reference signal frame is pre-emphasized and subjected to a short-time Fourier transform. The step of obtaining the transmitted signal based on the residual signal frame of the optimal microphone includes: Each residual signal frame of the optimal microphone is subjected to de-emphasis processing and inverse short-time Fourier transform, and then the processed residual signal frames are spliced ​​together to form a continuous transmission signal.

6. The adaptive dynamic audio processing method according to claim 5, characterized in that, After the step of selecting the optimal microphone in the microphone array based on the index score of each microphone, the method includes: Based on the microphone signal frame, the residual signal frame, and the reference signal frame of each microphone in the microphone array, and the preset performance indicators set for the microphone, the optimal scene classification is determined based on the preset feature threshold and the machine learning classifier. The optimal scene classification includes any one of the following: near-end speech dominance, far-end speech dominance, two-way talk, steady-state noise, and burst noise. Based on the optimal scene classification, a corresponding optimization algorithm is selected, and the residual signal frame is optimized using the optimization algorithm.

7. The adaptive dynamic audio processing method according to claim 6, characterized in that, The step of splicing the processed residual signal frames into a continuous transmission signal includes: The processed residual signal frame is subjected to residual echo suppression to obtain an echo-suppressed frame; The noise energy is estimated using spectral subtraction or a deep learning-based method, and noise suppression is performed on the echo-suppressed frame based on the noise energy to obtain a noise-suppressed frame. According to the application scenario, gain compensation is performed on the noise suppression frame to obtain the gain compensation frame; The gain-compensated frames are spliced ​​together to form the transmission signal.

8. The adaptive dynamic audio processing method according to claim 6, characterized in that, After the step of selecting the optimal microphone in the microphone array based on the index score of each microphone, the method further includes: The gain of the other microphones in the microphone array is adjusted based on the residual signal energy of the optimal microphone; Based on the pre-stored microphone frequency response compensation coefficient table, the frequency response deviation of each microphone is compensated by an FIR filter; The delay difference between the other microphones and the optimal microphone is calculated by cross-correlation analysis, and a delay compensation frame is inserted into the residual signal frame.

9. A storage medium, characterized in that, The device stores a computer program that, when executed by a processor, causes the processor to perform the steps of the method as described in any one of claims 1 to 8.

10. An audio transmission and playback terminal, characterized in that, It includes a memory and a processor, the memory storing a computer program that, when executed by the processor, causes the processor to perform the steps of the method as described in any one of claims 1 to 8.