Microphone far-field pickup amplitude dynamic range control method
Through frame processing and artificial intelligence collaborative analysis of bone conduction microphone and air conduction microphone signals, a gain control strategy is generated, which solves the problem of distinguishing far-field target signals from near-field interference sound sources and improves the pickup quality and robustness of the microphone system.
Patent Information
- Application Number
- CN202511139569.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-14
- Publication Date
- 2025-09-12
- Estimated Expiration
- 2045-08-14
AI Technical Summary
Existing technologies cannot effectively distinguish between weak signals from far-field targets and strong interfering sound sources from near-field devices, resulting in improper gain control, signal clipping, or poor quality of target voice capture.
The bone conduction microphone signal and the air conduction microphone signal are framed and processed to extract the feature vectors of near-field interference and far-field human voice. The artificial intelligence decision engine is used for collaborative analysis to generate a gain control strategy package for refined gain adjustment.
It achieves accurate distinction of mixed signals, reduces gain misadjustment, improves the pickup quality and clarity of far-field target signals, suppresses noise and distortion, and improves the robustness of the microphone system in complex acoustic environments.
Smart Images

Figure CN120640179A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of audio signal processing, and in particular to a method for controlling the dynamic range of far-field sound pickup amplitude of a microphone. Background Art
[0002] The field of audio capture technology, particularly portable recording devices (such as AI recorders) that work in conjunction with smart devices, often face the challenge of coping with complex sound pickup environments. These devices typically use air conduction microphones to capture distant ambient sounds, such as the target person's voice. However, a common technical challenge is that while these air conduction microphones capture weak far-field signals, they inevitably also capture strong signals from the host device's own speakers, which are very close by.
[0003] The significant difference in physical location and signal amplitude between these two sound sources results in an extremely wide dynamic range in the mixed signal received by the air conduction microphone. This significant amplitude disparity poses a severe challenge to traditional gain control methods. Existing technical solutions, such as dynamic range control (DRC) or automatic gain control (AGC), typically analyze and respond to the overall amplitude of the mixed signal. These methods are unable to effectively distinguish between different sound source components in the signal. When a strong near-field signal appears, the input signal may produce irreversible clipping distortion before the gain control unit can respond. If the gain is maintained at a low level to avoid clipping, the weak target signal in the far field will be submerged in the system noise floor, resulting in information loss.
[0004] Therefore, existing technologies generally lack a mechanism that can distinguish and perceive signals from different sources in complex acoustic environments and perform forward-looking and refined gain adjustment accordingly. It is difficult to simultaneously take into account the pickup quality and clarity of far-field target signals without generating clipping distortion.
[0005] Therefore, the present invention proposes a method for controlling the dynamic range of far-field sound pickup amplitude of a microphone to address the deficiencies of the prior art. Summary of the Invention
[0006] In response to the shortcomings of the existing technology, the present invention provides a method for controlling the dynamic range of the far-field sound pickup amplitude of a microphone, which solves the problem of improper gain control caused by the inability to effectively distinguish between the weak far-field target human voice and the strong interference sound of the near-field device in the mixed signal, thereby causing signal clipping or poor quality of target human voice collection.
[0007] To achieve the above objectives, the present invention is implemented through the following technical solutions: a method for controlling the dynamic range of far-field sound pickup amplitude of a microphone, the method comprising the following steps: S1. performing frame processing on the synchronously collected bone conduction microphone signal and the air conduction microphone signal as a mixed signal to obtain corresponding bone conduction signal frames and air conduction signal frames; S2. Extracting a first eigenvector for characterizing near-field interference characteristics based on the bone conduction signal frame; S3. Extracting a second eigenvector for characterizing a far-field human voice state based on the air conduction signal frame; S4. Inputting the first feature vector and the second feature vector into a preset artificial intelligence decision engine, and having the artificial intelligence decision engine perform collaborative analysis to determine a gain control strategy package that can respond to the near-field interference characteristics and take into account the far-field human voice state; S5. Perform gain adjustment on the air conduction signal frame according to the gain control strategy package, and reconstruct an output audio signal with controlled dynamic range.
[0008] Preferably, in step S1, the step of performing frame processing on the synchronously collected bone conduction microphone signal and the air conduction microphone signal as a mixed signal to obtain corresponding bone conduction signal frames and air conduction signal frames includes: The bone conduction microphone signal and the air conduction microphone signal are segmented according to a preset frame length and a preset frame shift to obtain bone conduction signal frames and air conduction signal frames that partially overlap. multiplying the sampling point sequences of the bone conduction signal frame and the air conduction signal frame by a coefficient sequence of a smoothing window function point by point, respectively, to suppress spectrum leakage during spectrum analysis; The bone conduction signal frame and the air conduction signal frame after point-by-point multiplication are respectively subjected to short-time Fourier transform to generate bone conduction spectrum data and air conduction spectrum data containing amplitude spectrum information and phase spectrum information.
[0009] Preferably, in step S2, the step of extracting a first eigenvector for characterizing near-field interference characteristics based on the bone conduction signal frame includes: Determine, based on the amplitude spectrum of the bone conduction spectrum data, a spectrum centroid in the spectrum morphological feature of the near-field interference sound type by accumulating the product of each frequency index and the amplitude value corresponding to the frequency index, and dividing the accumulated sum by the sum of all amplitude values of the amplitude spectrum; Dividing the frequency domain covered by the bone conduction spectrum data into a plurality of preset sub-bands, and accumulating the energy values of all frequency points in each of the plurality of preset sub-bands to obtain a frequency-band energy feature representing the energy frequency domain distribution of the near-field interference; Calculating a time domain envelope of the bone conduction signal frame, and extracting at least one statistic based on the time domain envelope as a time domain envelope feature characterizing a transient change characteristic of the near-field interference in the time domain; The determined spectrum centroid, the obtained frequency segment energy features, and the extracted time domain envelope features are combined into the first feature vector.
[0010] Preferably, in step S3, the step of extracting a second eigenvector for characterizing the far-field human voice state based on the air conduction signal frame includes: Applying a voice activity detection algorithm to process the air conduction signal frame to generate a voice activity detection state indicating whether a far-field human voice component is present in the air conduction signal frame; When the voice activity detection state determines that a far-field human voice component exists, applying a noise power spectrum estimation algorithm to process the air conduction spectrum data to obtain an estimated noise power spectrum; The total energy of the air conduction signal is obtained by accumulating the energy values of all frequency points in the air conduction spectrum data; The estimated total noise energy is obtained by accumulating the energy values of all frequency points in the estimated noise power spectrum; Dividing the total energy of the air conduction signal by the estimated total energy of the noise to obtain a ratio, and performing a logarithmic operation on the ratio to determine an estimated signal-to-noise ratio for quantifying far-field vocal clarity; The generated voice activity detection state and the determined estimated signal-to-noise ratio are combined into the second feature vector.
[0011] Preferably, the gain control strategy package includes: a multi-band target gain vector, wherein the multi-band target gain vector includes target gain values for a plurality of preset frequency bands; Adaptive attack time; Adaptive release time.
[0012] Preferably, in step S4, the first feature vector and the second feature vector are input into a preset artificial intelligence decision engine, and the step of collaborative analysis by the artificial intelligence decision engine includes: When the voice activity detection status in the second eigenvector is that there is no far-field human voice, the artificial intelligence decision engine uses the near-field interference characteristics represented by the first eigenvector as the sole basis to generate a gain control strategy package with the priority goal of preventing signal clipping from occurring in the air conduction signal frame.
[0013] Preferably, in step S4, the step of inputting the first feature vector and the second feature vector into a preset artificial intelligence decision engine, and the step of performing collaborative analysis by the artificial intelligence decision engine further includes: When the voice activity detection status in the second feature vector indicates the presence of far-field human voice and the estimated signal-to-noise ratio is low, if the near-field interference intensity represented by the first feature vector is higher than a preset threshold, the artificial intelligence decision engine generates a gain control strategy package with the priority goals of maintaining gain stability and suppressing output distortion, so as to avoid the introduction of additional noise due to drastic gain adjustment.
[0014] Preferably, in step S5, the step of performing gain adjustment on the air conduction signal frame according to the gain control strategy package and reconstructing an output audio signal with controlled dynamic range includes: Decomposing the air conduction signal frame into a plurality of sub-band signals by an analysis filter bank, wherein the sub-band signals correspond to a plurality of preset frequency bands in the gain control strategy package; Applying a smoothed gain to each of the plurality of sub-band signals; All sub-band signals to which gains have been applied are combined by a synthesis filter bank matched with the analysis filter bank to reconstruct the output audio signal.
[0015] Preferably, the step of applying a smoothed gain to each of the plurality of sub-band signals comprises: Obtaining a target gain value corresponding to a currently processed subband signal from the gain control strategy package; Determining a change trend of the target gain value relative to a smoothed gain value applied at a previous moment, and determining an adaptive time constant based on the change trend, wherein when the change trend is a gain increase, the adaptive time constant is set to an adaptive attack time in the gain control strategy package; and when the change trend is a gain decrease or gain retention, the adaptive time constant is set to an adaptive release time in the gain control strategy package; A smoothing coefficient is calculated by the following formula : ; Where, Represents the processing time interval of the signal frame, represents the adaptive time constant; By the following recursive formula, using the smoothing coefficient Update the smoothing gain value: ; Where, Represents the smoothing gain value at the current moment, Represents the smoothing gain value of the previous moment, represents the target gain value; The calculated smoothing gain value at the current moment Applied to the sub-band signal.
[0016] The present invention also provides a microphone far-field sound pickup amplitude dynamic range control system, the system comprising: A signal preprocessing module is used to perform frame processing on the synchronously collected bone conduction microphone signal and the air conduction microphone signal as a mixed signal to obtain corresponding bone conduction signal frames and air conduction signal frames; A first feature extraction module is configured to extract a first feature vector for characterizing near-field interference characteristics based on the bone conduction signal frame; A second feature extraction module is used to extract a second feature vector for characterizing a far-field human voice state based on the air conduction signal frame; an artificial intelligence decision engine, configured to receive the first feature vector and the second feature vector for collaborative analysis to determine a gain control strategy package that can respond to the near-field interference characteristics and take into account the far-field human voice state; The gain adjustment and reconstruction module is used to perform gain adjustment on the air conduction signal frame according to the gain control strategy package, and reconstruct an output audio signal with controlled dynamic range.
[0017] The present invention provides a method for controlling the dynamic range of far-field sound pickup amplitude of a microphone. It has the following beneficial effects: 1. This invention uses the bone conduction microphone signal as an independent reference for near-field interference, enabling the system to accurately distinguish far-field human voices from near-field interference in mixed signals. This dual-path signal collaborative analysis mechanism resolves the ambiguity of a single air conduction microphone signal source, making dynamic range control decisions more clear and reliable, and reducing gain misadjustments caused by misidentified sound sources.
[0018] 2. This invention utilizes an artificial intelligence decision engine to collaboratively analyze the extracted first and second eigenvectors, dynamically generating an optimal gain control strategy package based on the changing acoustic environment. This approach not only considers signal energy but also comprehensively analyzes the type of near-field interfering sound, the clarity of far-field human voices, and the presence of speech activity. This enables scenario-adaptive gain adjustment, enabling it to cope with even more complex practical application environments.
[0019] 3. The gain control strategy package output by this invention includes multi-dimensional control parameters, such as multi-band target gain vectors and adaptive attack and release times. This refined control strategy eliminates the need for global, crude gain adjustments and instead enables differentiated processing across frequency bands. Furthermore, by applying smoothly varying gain, it effectively suppresses the pumping and noise breathing effects common in traditional dynamic range control, improving the naturalness and listening quality of the final output audio.
[0020] 4. The technical solution provided by this invention, in the extreme case of strong near-field interference superimposed on weak far-field human voices, can generate a control strategy that prioritizes maintaining gain stability and suppressing output distortion through specific collaborative analysis logic. This design avoids excessively increasing overall gain in an attempt to amplify weak human voices, thereby preventing catastrophic amplification of noise and distortion, and ensuring the robustness of the entire microphone far-field pickup system in harsh acoustic conditions. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] Figure 1 is a flow chart of the method of the present invention; Figure 2 This is a schematic diagram of the working principle of the AI recorder of the present invention; Figure 3 This is a schematic diagram of the far-field sound pickup algorithm framework of the present invention; Figure 4 This is a schematic diagram of the working principle of the AI recorder algorithm of the present invention; Figure 5 This is a system architecture diagram of the present invention.
[0022] Among them, 110, signal preprocessing module; 120, first feature extraction module; 130, second feature extraction module; 140, artificial intelligence decision engine; 150, gain adjustment and reconstruction module. DETAILED DESCRIPTION
[0023] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the drawings in the present specification. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0024] Reference Figure 1 The present invention provides a method for controlling the dynamic range of a microphone's far-field sound pickup amplitude. In a specific embodiment, the method may include the following steps: S1. Perform frame processing on the synchronously collected bone conduction microphone signal and the air conduction microphone signal as a mixed signal to obtain corresponding bone conduction signal frames and air conduction signal frames.
[0025] S2. Extracting a first eigenvector for characterizing near-field interference characteristics based on the bone conduction signal frame.
[0026] S3. Extract a second eigenvector for characterizing the far-field human voice state based on the air conduction signal frame.
[0027] S4. Input the first eigenvector and the second eigenvector into a preset artificial intelligence decision engine, which performs collaborative analysis to determine a gain control strategy package that can respond to near-field interference characteristics and take into account far-field human voice status.
[0028] S5. According to the gain control strategy package, the air conduction signal frame is gain-adjusted, and an output audio signal with a controlled dynamic range is reconstructed.
[0029] Reference Figure 5 To implement the above method, the present invention also provides a microphone far-field sound pickup amplitude dynamic range control system. The system includes: A signal preprocessing module 110 , a first feature extraction module 120 , a second feature extraction module 130 , an artificial intelligence decision engine 140 , and a gain adjustment and reconstruction module 150 .
[0030] The signal pre-processing module 110 is responsible for receiving and processing the synchronized bone conduction and air conduction microphone signals.
[0031] The first feature extraction module 120 and the second feature extraction module 130 extract feature vectors from the processed bone conduction signals and air conduction signals, respectively.
[0032] The artificial intelligence decision engine 140 receives these two feature vectors, performs collaborative analysis and outputs a gain control strategy.
[0033] The gain adjustment and reconstruction module 150 adjusts the air conduction signal according to the strategy and outputs the result.
[0034] Reference Figure 1 The method flow shown, and Figure 5 The system structure shown in FIG. 1 illustrates the specific implementation steps of the method of the present invention.
[0035] First, in step S1, a signal preprocessing operation is performed. This operation is performed by Figure 5 The signal preprocessing module 110 shown is completed. The input of the signal preprocessing module 110 is two parallel, synchronously collected time domain digital signals, namely the bone conduction microphone signal and the air conduction microphone signal.
[0036] To facilitate time-frequency analysis, the signal preprocessing module 110 first performs frame segmentation on the two continuous signal streams. This process segments the continuous signals into a series of partially overlapping data frames based on a preset frame length and a preset frame shift. This results in a discrete sequence of bone conduction signal frames and air conduction signal frames.
[0037] Before performing spectrum transformation on the signal frame, a smoothing window function is applied to each bone conduction signal frame and air conduction signal frame to suppress spectral leakage inevitably introduced by signal frame truncation. In one specific embodiment, a Hanning window function can be used. This application process involves point-by-point multiplication of the window function coefficient sequence with the signal frame's sampling point sequence.
[0038] Subsequently, the signal preprocessing module 110 performs a short-time Fourier transform (STFT) on each frame of the bone conduction signal and the air conduction signal after applying the smoothing window function. The STFT maps each frame of the time-domain signal to the complex frequency domain. The result contains not only the amplitude spectrum information of the signal within that time period, but also the phase spectrum information. This transformation is typically implemented using the fast Fourier transform (FFT) algorithm to improve computational efficiency.
[0039] After the short-time Fourier transform is completed, the output is the bone conduction spectrum data and the air conduction spectrum data. These two spectrum data provide the necessary frequency domain information basis for feature extraction in the subsequent steps and are transmitted to the first feature extraction module 120 and the second feature extraction module 130 respectively.
[0040] In step S2, the first feature vector is extracted. Figure 5 The first feature extraction module 120 shown in FIG. 1 is executed with the bone conduction signal frames and bone conduction spectrum data from the signal preprocessing module 110 as input. This step aims to calculate and combine these data into a first feature vector that comprehensively characterizes near-field interference. In one specific embodiment, this first feature vector is composed of the following features of different dimensions.
[0041] The first type of feature is spectral morphology, which is used to characterize the sound type attributes of near-field interference. This type of feature is calculated based on bone conduction spectrum data. In one embodiment, it may include spectral centroid and spectral flatness. The spectral centroid is used to quantify the center of gravity frequency position of the spectral energy. Specifically, the first feature extraction module 120 calculates the spectral centroid based on the amplitude spectrum of the input bone conduction spectrum data using the following formula: ; Where: Represents the spectrum centroid value of the current bone conduction signal frame; Represents bone conduction spectrum data in The amplitude value at each frequency point; is the frequency index, and its value range is from 1 to ; is the total number of frequency points, which is determined by the number of points of the short-time Fourier transform in step S1.
[0042] Additionally, spectral flatness can be calculated, which measures how closely the signal spectrum resembles a tonal sound or noise.
[0043] The second type of feature is the frequency-band energy feature, which characterizes the energy distribution of near-field interference in the frequency domain. To calculate this feature, the first feature extraction module 120 first divides the entire frequency domain into multiple subbands based on preset frequency boundaries. Then, for each subband, the total energy of the subband is calculated by accumulating the energy values (i.e., the square of the amplitude values) of the bone conduction spectrum data at all frequencies within that subband. The energy values of all subbands together constitute the frequency-band energy feature.
[0044] The third type of feature is the time-domain envelope feature, which is used to characterize the transient variation characteristics of near-field interference in the time domain. This feature is calculated based on the bone conduction signal frame in the time domain. The first feature extraction module 120 first calculates the time-domain envelope of the bone conduction signal frame and then extracts one or more statistics based on this envelope that can reflect its transient variation, such as the peak value, mean value, or zero-crossing rate of the envelope.
[0045] The extracted spectral morphology features, sub-band energy features, and time-domain envelope features are combined to form a first feature vector, which is then transmitted to the artificial intelligence decision engine 140 as an important input for its collaborative analysis.
[0046] In step S3, the second feature vector is extracted. Figure 5 The second feature extraction module 130 shown is executed, and its input is the air conduction signal frame and air conduction spectrum data from the signal preprocessing module 110. The purpose of this step is to calculate and combine these mixed signal data into a second feature vector for characterizing the far-field human voice state.
[0047] In a specific embodiment, extracting the second feature vector includes the following process. First, the second feature extraction module 130 applies a voice activity detection (VAD) algorithm to the input air conduction signal frame. This algorithm analyzes multiple acoustic characteristics of the signal to generate a binary or multivariate voice activity detection state. This state is used to determine whether a far-field vocal component is present in the current air conduction signal frame.
[0048] Secondly, if and only if the voice activity detection result indicates the presence of far-field human voice, the second feature extraction module 130 initiates subsequent calculations. Based on the air conduction spectrum data, this module applies a noise power spectrum estimation algorithm to estimate the power spectrum of the noise component in the current signal frame. This estimation process can be performed over multiple consecutive frames, using statistical analysis to identify the minimum value or the portion with the most gradual change in signal energy, thereby continuously tracking the background noise level.
[0049] After obtaining the estimated noise power spectrum, the second feature extraction module 130 calculates an estimated signal-to-noise ratio (SNR) for quantifying the far-field vocal clarity. The estimated SNR is calculated using the following formula: ; Where: Represents the estimated signal-to-noise ratio of the current air conduction signal frame, in decibels (dB); Representative air conduction spectrum data Complex value at frequency points; Represents the air conduction signal at The energy at each frequency point; Represents the noise power spectrum obtained by the noise power spectrum estimation algorithm. The complex value of the noise spectrum estimated at the frequency point; Represents the estimated noise in The energy at each frequency point; is the frequency index, and its value range is from 1 to ; is the total number of frequency points, which is determined by the number of points of the short-time Fourier transform in step S1.
[0050] Finally, the second feature extraction module 130 combines the generated voice activity detection status and the estimated signal-to-noise ratio calculated under specific conditions into a second feature vector. This vector is then transmitted to the artificial intelligence decision engine 140 as another input for collaborative analysis. If the voice activity detection status of the current frame indicates that far-field human voice is absent, the estimated signal-to-noise ratio may be set to a preset invalid value or not calculated or transmitted.
[0051] In step S4, artificial intelligence decision making and gain strategy generation are performed. Figure 5 The artificial intelligence decision engine 140 is shown as complete. The input of the artificial intelligence decision engine 140 is connected to the output of the first feature extraction module 120 and the second feature extraction module 130, respectively, to receive the first feature vector and the second feature vector. Its function is to perform a collaborative analysis based on these two input feature vectors and output a gain control strategy package.
[0052] In one specific embodiment, the gain control strategy package is a data structure consisting of three core components: a multi-band target gain vector, an adaptive attack time, and an adaptive release time. The multi-band target gain vector contains target gain values for multiple preset frequency bands, enabling frequency-selective gain adjustment of the signal. The adaptive attack time and adaptive release time are used to control the speed of subsequent gain smoothing.
[0053] The AI decision engine 140 executes a rule-based or pre-trained model-based decision process to generate the most appropriate gain control strategy package based on the real-time acoustic scenario reflected by the input feature vector. This collaborative analysis process includes processing the following typical logical scenarios: A logical scenario is when the voice activity detection status in the second eigenvector indicates the absence of far-field human voice, indicating that the dominant component in the current air conduction signal is near-field interference or ambient noise. In this case, the decision-making process of the artificial intelligence decision engine 140 will be based solely on the near-field interference characteristics represented by the first eigenvector. The generated gain control strategy package will prioritize preventing signal clipping in the air conduction signal frame. For example, a low or suppressive target gain value is set based on the energy characteristics of the first eigenvector.
[0054] Another logical scenario is that when the voice activity detection state in the second eigenvector is that there is a far-field human voice and the estimated signal-to-noise ratio it contains is low, if the near-field interference intensity represented by the first eigenvector is higher than a preset threshold at this time, it indicates that the current scene is a strong near-field interference superimposed on a weak far-field human voice. In this case, if the gain is greatly increased in an attempt to enhance the far-field human voice, the stronger near-field interference will be disproportionately amplified, resulting in severe distortion of the output signal. Therefore, the artificial intelligence decision engine 140 generates a gain control strategy package with the priority goal of maintaining gain stability and suppressing output distortion. The target gain value in the strategy package will tend to maintain the level of the previous frame or slightly attenuate, while setting longer adaptive attack and release times to avoid the introduction of additional noise or auditory pumping effects due to drastic gain adjustments.
[0055] For other scenario combinations, such as when there is no near-field interference but clear far-field human voice, the artificial intelligence decision engine 140 will generate a gain control strategy package aimed at increasing the amplitude of the far-field human voice signal.
[0056] Finally, the artificial intelligence decision engine 140 transmits the gain control strategy package generated according to the results of each collaborative analysis to the gain adjustment and reconstruction module 150 to guide subsequent gain adjustment operations.
[0057] In step S5, gain adjustment and signal reconstruction are performed. Figure 5 The gain adjustment and reconstruction module 150 shown is completed. The input end of the module is connected to the artificial intelligence decision engine 140 and the signal preprocessing module 110 respectively to receive the gain control strategy package and the original air conduction signal frame.
[0058] In a specific embodiment, the gain adjustment and reconstruction module 150 first decomposes the air conduction signal frame into multiple sub-band signals using an analysis filter bank. The frequency band division of the analysis filter bank corresponds to the frequency band division of the multi-band target gain vector in the gain control strategy package, ensuring that each sub-band signal has a corresponding target gain value.
[0059] For each subband signal, the gain adjustment and reconstruction module 150 obtains the target gain value corresponding to the subband from the gain control strategy package. To prevent the sudden gain change and unnatural auditory perception that may be introduced by directly applying this target gain value to the subband signal, the target gain value needs to be smoothed.
[0060] The smoothing process first needs to determine the change trend of the target gain value relative to the smoothing gain value applied at the previous moment. Based on the change trend, a smoothing coefficient is calculated. .
[0061] If the current target gain value is greater than the smoothing gain value at the previous moment, it is determined to be a gain increase trend, and the smoothing coefficient The calculation of will be based on the adaptive attack time in the gain control strategy package.
[0062] If the current target gain value is less than or equal to the smoothing gain value at the previous moment, it is determined that the gain is decreasing or maintaining the trend, and the smoothing coefficient The calculation of will be based on the adaptive release time in the gain control strategy package.
[0063] Smoothing coefficient The specific calculation method is as follows: ; Where: represents the calculated smoothing coefficient; Represents the processing time interval of the signal frame, that is, the number of seconds corresponding to the frame shift; Represents the adaptive time constant. When in the gain increasing trend, The value is the adaptive attack time in the gain control strategy package; when the gain is decreasing or maintaining, The value is the adaptive release time.
[0064] In calculating the smoothing coefficient Then, the gain adjustment and reconstruction module 150 uses the coefficient to smooth the target gain value through the following recursive formula to generate a smoothly changing gain: ; Where: represents the smoothing gain value calculated at the current moment and to be applied to the sub-band signal; Represents the smoothing gain value that has been calculated and applied at the previous moment; Represents the current target gain value obtained from the gain control strategy package; is the smoothing coefficient calculated according to the above method.
[0065] The gain adjustment and reconstruction module 150 calculates the smoothed gain of each sub-band , which are applied to the corresponding subband signals. After the gains are applied to all subbands, all gain-adjusted subband signals are combined through a synthesis filter bank that matches the analysis filter bank to reconstruct an output audio frame with controlled dynamic range.
[0066] Finally, by splicing consecutive output audio frames (for example, using an overlap-add method), a final, continuous, and dynamically range-controlled output audio signal can be obtained.
[0067] In order to further illustrate the implementation process and technical effects of the technical solution of the present invention, a specific application scenario is provided below as an example.
[0068] Reference Figure 2-Figure 5 Imagine a scenario where a user is using an AI recorder in a conference room that incorporates the proposed microphone far-field pickup amplitude dynamic range control system. The recorder is placed on a conference table or in contact with the back panel of a mobile phone, aiming to record the voice of a speaker three meters away (i.e., the far-field human voice). During the recording process, the speaker of the host device (a smartphone) housing the recorder suddenly plays a short but loud beep (i.e., near-field interference sound).
[0069] In this scenario, the complete workflow of the system of the present invention is as follows: Before the prompt tone plays, the system only captures far-field vocals. The signal preprocessing module 110 performs frame segmentation and short-time Fourier transform (SFT) on the synchronized air conduction and bone conduction microphone signals. Since there is no speaker vibration at this time, the first feature vector extracted from the bone conduction signal by the first feature extraction module 120 reflects very low near-field interference energy. Simultaneously, the second feature extraction module 130 detects the presence of far-field vocals from the air conduction signal (the voice activity detection state is true) and calculates a moderate estimated signal-to-noise ratio. The artificial intelligence decision engine 140 receives these two sets of feature vectors and, after collaborative analysis, determines that the current scene is free of interference and presents far-field vocals. It then generates a gain control strategy package aimed at improving vocal clarity, including a high multi-band target gain value. Based on this strategy package, the gain adjustment and reconstruction module 150 smoothly applies a high gain to the air conduction signal, ensuring the speaker's voice is clearly audible.
[0070] When the alert sound suddenly sounds, the bone conduction microphone picks up a strong housing vibration signal, while the air conduction microphone picks up a mixed signal of a strong alert sound and a weak human voice. At this point, the first feature vector extracted by the first feature extraction module 120 immediately reflects the high intensity of near-field interference. While the second feature extraction module 130 can still detect vocal activity, the estimated signal-to-noise ratio (SNR) calculated drops dramatically. The artificial intelligence decision engine 140 receives these two sets of significantly altered feature vectors and, after collaborative analysis, immediately identifies the scenario as weak far-field vocal interference within the context of strong near-field interference. To prevent signal clipping, the artificial intelligence decision engine 140 immediately generates a new gain control strategy package. This strategy package sets the multi-band target gain values to a low, suppressive value, and may include a short adaptive attack time for rapid response. Upon receiving this new strategy, the gain adjustment and reconstruction module 150's internal gain smoothing mechanism quickly reduces the applied smoothing gain, effectively suppressing the alert sound's amplitude and preventing clipping distortion in the output signal.
[0071] After the prompt tone ends, the system status is restored. The first feature extraction module 120 once again outputs a low-interference energy feature vector. The artificial intelligence decision engine 140 then regenerates the target gain strategy for enhancing the human voice. The gain adjustment and reconstruction module 150 then smoothly restores the gain to a higher level, continuing to clearly pick up the far-field human voice.
[0072] While embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions, and variations may be made to these embodiments without departing from the principles and spirit of the invention, and that the scope of the invention is defined by the appended claims and their equivalents.
Claims
1. A method for controlling the dynamic range of far-field sound pickup amplitude of a microphone, characterized in that: The method comprises the following steps: S1. performing frame processing on the synchronously collected bone conduction microphone signal and the air conduction microphone signal as a mixed signal to obtain corresponding bone conduction signal frames and air conduction signal frames; S2. Extracting a first eigenvector for characterizing near-field interference characteristics based on the bone conduction signal frame; S3. Extracting a second eigenvector for characterizing a far-field human voice state based on the air conduction signal frame; S4. Inputting the first feature vector and the second feature vector into a preset artificial intelligence decision engine, and having the artificial intelligence decision engine perform collaborative analysis to determine a gain control strategy package that can respond to the near-field interference characteristics and take into account the far-field human voice state; S5. Perform gain adjustment on the air conduction signal frame according to the gain control strategy package, and reconstruct an output audio signal with controlled dynamic range.
2. A method for controlling the dynamic range of far-field sound pickup amplitude of a microphone according to claim 1, characterized in that: In step S1, the steps of performing frame processing on the synchronously collected bone conduction microphone signal and the air conduction microphone signal as a mixed signal to obtain corresponding bone conduction signal frames and air conduction signal frames include: The bone conduction microphone signal and the air conduction microphone signal are segmented according to a preset frame length and a preset frame shift to obtain bone conduction signal frames and air conduction signal frames that partially overlap. multiplying the sampling point sequences of the bone conduction signal frame and the air conduction signal frame by a coefficient sequence of a smoothing window function point by point, respectively, to suppress spectrum leakage during spectrum analysis; The bone conduction signal frame and the air conduction signal frame after point-by-point multiplication are respectively subjected to short-time Fourier transform to generate bone conduction spectrum data and air conduction spectrum data containing amplitude spectrum information and phase spectrum information.
3. The method for controlling the dynamic range of far-field sound pickup amplitude of a microphone according to claim 2, wherein: In step S2, the step of extracting a first eigenvector for characterizing near-field interference characteristics based on the bone conduction signal frame includes: Determine, based on the amplitude spectrum of the bone conduction spectrum data, a spectrum centroid in the spectrum morphological feature of the near-field interference sound type by accumulating the product of each frequency index and the amplitude value corresponding to the frequency index, and dividing the accumulated sum by the sum of all amplitude values of the amplitude spectrum; Dividing the frequency domain covered by the bone conduction spectrum data into a plurality of preset sub-bands, and accumulating the energy values of all frequency points in each of the plurality of preset sub-bands to obtain a frequency-band energy feature representing the energy frequency domain distribution of the near-field interference; Calculating a time domain envelope of the bone conduction signal frame, and extracting at least one statistic based on the time domain envelope as a time domain envelope feature characterizing a transient change characteristic of the near-field interference in the time domain; The determined spectrum centroid, the obtained frequency segment energy features, and the extracted time domain envelope features are combined into the first feature vector.
4. The method for controlling the dynamic range of far-field sound pickup amplitude of a microphone according to claim 2, wherein: In step S3, the step of extracting a second eigenvector for characterizing the far-field human voice state based on the air conduction signal frame includes: Applying a voice activity detection algorithm to process the air conduction signal frame to generate a voice activity detection state indicating whether a far-field human voice component is present in the air conduction signal frame; When the voice activity detection state determines that a far-field human voice component exists, applying a noise power spectrum estimation algorithm to process the air conduction spectrum data to obtain an estimated noise power spectrum; The total energy of the air conduction signal is obtained by accumulating the energy values of all frequency points in the air conduction spectrum data; The estimated total noise energy is obtained by accumulating the energy values of all frequency points in the estimated noise power spectrum; Dividing the total energy of the air conduction signal by the estimated total energy of the noise to obtain a ratio, and performing a logarithmic operation on the ratio to determine an estimated signal-to-noise ratio for quantifying far-field vocal clarity; The generated voice activity detection state and the determined estimated signal-to-noise ratio are combined into the second feature vector.
5. The method for controlling the dynamic range of far-field sound pickup amplitude of a microphone according to claim 1, characterized in that: The gain control strategy package includes: a multi-band target gain vector, wherein the multi-band target gain vector includes target gain values for a plurality of preset frequency bands; Adaptive attack time; Adaptive release time.
6. The method for controlling the dynamic range of far-field sound pickup amplitude of a microphone according to claim 5, characterized in that: In step S4, the first feature vector and the second feature vector are input into a preset artificial intelligence decision engine, and the artificial intelligence decision engine performs collaborative analysis, including: When the voice activity detection status in the second eigenvector is that there is no far-field human voice, the artificial intelligence decision engine uses the near-field interference characteristics represented by the first eigenvector as the sole basis to generate a gain control strategy package with the priority goal of preventing signal clipping from occurring in the air conduction signal frame.
7. The method for controlling the dynamic range of far-field sound pickup amplitude of a microphone according to claim 5, characterized in that: In step S4, the first feature vector and the second feature vector are input into a preset artificial intelligence decision engine, and the step of collaborative analysis by the artificial intelligence decision engine further includes: When the voice activity detection status in the second feature vector indicates the presence of far-field human voice and the estimated signal-to-noise ratio is lower than a preset signal-to-noise ratio threshold, if the near-field interference intensity represented by the first feature vector is higher than a preset threshold, the artificial intelligence decision engine generates a gain control strategy package with the priority goals of maintaining gain stability and suppressing output distortion, so as to avoid the introduction of additional noise due to drastic gain adjustment.
8. The method for controlling the dynamic range of far-field sound pickup amplitude of a microphone according to claim 1, wherein: In step S5, the steps of performing gain adjustment on the air conduction signal frame according to the gain control strategy package and reconstructing an output audio signal with controlled dynamic range include: Decomposing the air conduction signal frame into a plurality of sub-band signals by an analysis filter bank, wherein the sub-band signals correspond to a plurality of preset frequency bands in the gain control strategy package; Applying a smoothed gain to each of the plurality of sub-band signals; All sub-band signals to which gains have been applied are combined by a synthesis filter bank matched with the analysis filter bank to reconstruct the output audio signal.
9. The method for controlling the dynamic range of far-field sound pickup amplitude of a microphone according to claim 8, characterized in that: The step of applying a smoothed gain to each of the plurality of sub-band signals comprises: Obtaining a target gain value corresponding to a currently processed subband signal from the gain control strategy package; Determining a change trend of the target gain value relative to a smoothed gain value applied at a previous moment, and determining an adaptive time constant based on the change trend, wherein when the change trend is a gain increase, the adaptive time constant is set to an adaptive attack time in the gain control strategy package; and when the change trend is a gain decrease or gain retention, the adaptive time constant is set to an adaptive release time in the gain control strategy package; A smoothing coefficient is calculated by the following formula : ; Where, Represents the processing time interval of the signal frame, represents the adaptive time constant; By the following recursive formula, using the smoothing coefficient Update the smoothing gain value: ; Where, Represents the smoothing gain value at the current moment, Represents the smoothing gain value at the previous moment, represents the target gain value; The calculated smoothing gain value at the current moment Applied to the sub-band signal.
10. A far-field sound pickup amplitude dynamic range control system for a microphone, applied to the method according to any one of claims 1 to 9, characterized in that: The system comprises: A signal preprocessing module is used to perform frame processing on the synchronously collected bone conduction microphone signal and the air conduction microphone signal as a mixed signal to obtain corresponding bone conduction signal frames and air conduction signal frames; A first feature extraction module is configured to extract a first feature vector for characterizing near-field interference characteristics based on the bone conduction signal frame; A second feature extraction module is used to extract a second feature vector for characterizing a far-field human voice state based on the air conduction signal frame; an artificial intelligence decision engine, configured to receive the first feature vector and the second feature vector for collaborative analysis to determine a gain control strategy package that can respond to the near-field interference characteristics and take into account the far-field human voice state; The gain adjustment and reconstruction module is used to perform gain adjustment on the air conduction signal frame according to the gain control strategy package, and reconstruct an output audio signal with controlled dynamic range.
Citation Information
Patent Citations
Deep learning speech extraction and noise reduction method fusing bone vibration sensor and microphone signals
CN110931031A
Speech enhancement method and device, electronic equipment, chip and storage medium
CN116403592A
Voice signal processing method and related equipment
CN117953912A
Acoustic output device
WO2025123359A1
Cited By
Microphone far-field pickup tri-state DRC control method
CN121547712A