A hybrid signal processing method and system integrating MEMS speaker and microphone

Through the mixed signal processing method of integrating MEMS speakers and microphones, preprocessing, hybrid computing and closed-loop feedback systems are used to solve the problems of acoustic feedback, echo suppression and noise suppression, achieving high-quality audio equipment performance and user experience.

CN118660262BActive Publication Date: 2025-08-26SHENZHEN SHENGJIALI ELECTRONICS CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410614599.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-05-17
Publication Date
2025-08-26
Estimated Expiration
2044-05-17

AI Technical Summary

Technical Problem

The prior art has difficulties in audio feedback and echo suppression, ambient noise suppression and sound pressure level control in audio devices integrating MEMS speakers and microphones, especially in complex environments, it is difficult to maintain high-quality audio playback and sound pickup performance.

Method used

Using preprocessing, mixing calculation, signal mixing, feedback control and error analysis, dynamic balance and noise suppression between the speaker and the microphone are achieved by calculating the mixing coefficient and closed-loop feedback system in real time, combining the noise suppression algorithm and sound pressure level control.

Benefits of technology

Effectively suppress sound feedback and echo, improve signal-to-noise ratio, ensure the clarity of voice communication, and realize flexible sound pressure level control, adapt to the sound pressure needs of different application scenarios, and ensure user comfort and equipment safety.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118660262B_ABST
    Figure CN118660262B_ABST
Patent Text Reader

Abstract

The present invention discloses a hybrid signal processing method and system integrating a MEMS speaker and a microphone. The present invention realizes effective suppression of acoustic feedback and echo through the steps of preprocessing, hybrid calculation, signal mixing, feedback control, error analysis and adjustment: by calculating the mixing coefficient in real time, linearly mixing the speaker output signal and the microphone acquisition signal, a dynamic balance is achieved between the speaker and the microphone, and the risk of acoustic feedback is greatly reduced. At the same time, the closed-loop feedback system can capture the echo signal in time, and effectively suppress the echo by accurately adjusting the speaker output power to ensure the clarity of voice communication. The ambient sound signal collected by the microphone is subjected to preprocessing operations such as noise suppression, gain control and automatic level adjustment. The advanced noise suppression algorithm is used to significantly reduce the impact of background noise on the voice signal, significantly improve the signal-to-noise ratio, and ensure that high-quality voice interaction can still be achieved in a noisy environment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of MEMS loudspeakers, and in particular relates to a mixed signal processing method and system integrating a MEMS loudspeaker and a microphone. Background Art

[0002] With the rapid development of micro-electromechanical systems (MEMS) technology, audio devices integrating MEMS speakers and microphones are becoming increasingly popular and are widely used in smart speakers, mobile communication devices, voice recognition systems, and other fields. Such devices generally require high-quality audio playback and pickup functions, and the ability to maintain stable acoustic performance in complex environments. However, existing technologies face the following technical problems when processing mixed signals from MEMS speakers and microphones:

[0003] Acoustic feedback and echo suppression: Because the MEMS speaker and microphone are housed in the same package, the audio signal played by the speaker is easily picked up by the microphone, forming an acoustic feedback loop, leading to problems such as howling and distortion. Furthermore, in scenarios such as duplex communication or remote conferencing, the device must effectively suppress the echo of the speaker playback signal at the microphone end to ensure clear voice communication.

[0004] Environmental noise suppression: MEMS microphones are often subject to interference from background noise and mechanical vibration when collecting ambient sound signals, affecting the accuracy and clarity of speech recognition and voice communication. Using software algorithms to efficiently suppress noise and improve the signal-to-noise ratio within hardware constraints is a pressing technical challenge.

[0005] Sound pressure level control: Devices integrating MEMS speakers and microphones need to have flexible sound pressure level control capabilities to address different application scenarios and user needs. This ensures adequate volume output while preventing excessive sound pressure from causing user discomfort or device damage. Furthermore, sound pressure level control must be able to respond quickly and accurately to accommodate changes in ambient sound pressure. Summary of the Invention

[0006] The object of the present invention is to provide a mixed signal processing method and system integrating a MEMS speaker and a microphone, so as to solve the problems in the prior art mentioned in the above background technology.

[0007] To achieve the above objectives, the present invention adopts the following technical solution: a mixed signal processing method integrating a MEMS speaker and a microphone, comprising the following steps:

[0008] Signal acquisition: The MEMS speaker and MEMS microphone integrated in the same package receive and output audio signals and collect ambient sound signals respectively;

[0009] Preprocessing: Perform power adjustment, frequency equalization, and distortion correction on the audio signal from the MEMS speaker, and perform noise suppression, gain control, and automatic level adjustment on the ambient sound signal from the MEMS microphone;

[0010] Mixing calculation: The mixing coefficient is calculated using the following formula (1) to achieve real-time mixing of the speaker output signal and the microphone acquisition signal:

[0011]

[0012] Among them, P target is the target total sound pressure level, P mic is the ambient sound pressure level collected by the current microphone, P spk The sound pressure level of the current speaker output;

[0013] Signal mixing: Using the calculated mixing coefficient k, the pre-processed loudspeaker output signal S is mixed according to formula (2) out and microphone to collect signal M in Perform linear blending:

[0014] S mix =(1-k)·S out +k·M in

[0015] Feedback control: The mixed signal S mix The data is sent back to the MEMS speaker for playback and collected again by the MEMS microphone to form a closed-loop feedback system.

[0016] Error analysis and adjustment: Calculate the error between the currently collected ambient sound pressure level and the target sound pressure level. If the error exceeds the preset threshold, adjust the speaker output power according to formula (3):

[0017] P spk,new =P spk,old +α·(P target +P curent )

[0018] Among them, P spk,new is the output power of the speaker after adjustment, P spk,old To adjust the front speaker output power, α is the power adjustment coefficient, P curent is the current ambient sound pressure level;

[0019] Iterative loop: Repeat the mixing calculation, signal mixing, feedback control, and error analysis and adjustment process until the difference between the ambient sound pressure level and the target sound pressure level is lower than the preset threshold, completing the mixed signal processing.

[0020] Preferably, the power adjustment in the pre-processing step comprises the following steps:

[0021] Signal framing: Divide the audio signal from the MEMS speaker into continuous time frames and perform instantaneous amplitude calculation, including calculating the instantaneous amplitude value of each frame signal;

[0022] Threshold determination: Determine the upper threshold value T of dynamic range compression based on the pre-set compression ratio and inflection point upper and the lower threshold value T lower , where the compression ratio formula is:

[0023]

[0024] in, and are the upper and lower threshold values ​​corresponding to the uncompressed state respectively;

[0025] Compression processing: For the instantaneous amplitude value of each frame signal, if it is less than the lower limit threshold value T lower , it remains unchanged; if it is located at T lower With T upper Between, nonlinear compression is performed according to formula (4):

[0026]

[0027] Among them, A in is the original instantaneous amplitude value, A comp is the instantaneous amplitude value after compression;

[0028] Reconstruct the signal: Reconstruct the audio signal frame based on the compressed instantaneous amplitude value, and splice all frames to restore the complete audio signal to complete the power adjustment.

[0029] Preferably, the noise suppression in the preprocessing step comprises the following steps:

[0030] Noise reference signal acquisition: A MEMS microphone is used to collect ambient noise during non-speech activity as a noise reference signal;

[0031] Noise spectrum estimation: Perform fast Fourier transform on the noise reference signal and calculate its frequency domain power spectrum as the noise spectrum estimation;

[0032] Filter initialization: construct an adaptive filter whose initial parameters are set according to the noise spectrum estimate;

[0033] Voice activity detection: During real-time processing, a voice activity detection algorithm is used to determine whether the current period is voice active.

[0034] Noise cancellation update: During periods of non-speech activity, the noise reference signal is used to update the adaptive filter parameters so that the filter output signal is consistent with the noise reference signal, achieving noise tracking;

[0035] Noise cancellation processing: During periods of speech activity, an adaptive filter is applied to the real-time collected ambient sound signal to filter out the parts that are similar to the noise spectrum estimate, generating a noise-canceled signal.

[0036] Inverse superposition: The noise-canceled signal is superimposed on the original ambient sound signal in an inverse phase to eliminate residual noise and complete noise suppression.

[0037] Preferably, the automatic level adjustment in the pre-processing step is based on the signal-to-noise ratio monitored in real time, and the microphone gain is adjusted according to a preset signal-to-noise ratio target value.

[0038] Preferably, the target total sound pressure level P in the mixed calculation step is target It can be set according to the actual application scenario.

[0039] Preferably, in the feedback control step, the signal collected by the MEMS microphone is subjected to delay compensation processing and then used for closed-loop feedback.

[0040] Preferably, in the error analysis and adjustment step, the power adjustment coefficient α can be dynamically adjusted according to the actual application scenario.

[0041] Preferably, during the iterative cycle, when the difference between the ambient sound pressure level and the target sound pressure level is still higher than a preset threshold after several consecutive iterations, the target sound pressure level is automatically lowered.

[0042] Preferably, the method further includes: performing digital sound effect processing on the mixed signal.

[0043] In another aspect, the present invention provides a hybrid signal processing system integrating a MEMS speaker and a microphone, comprising:

[0044] A signal acquisition module is used to receive and output audio signals and collect ambient sound signals through a MEMS speaker and a MEMS microphone integrated in the same package;

[0045] A pre-processing module, configured to perform power adjustment, frequency equalization, and distortion correction pre-processing operations on the audio signal from the MEMS speaker, and to perform noise suppression, gain control, and automatic level adjustment pre-processing operations on the ambient sound signal from the MEMS microphone;

[0046] A mixing calculation module is used to calculate the mixing coefficient to achieve real-time mixing of the speaker output signal and the microphone acquisition signal;

[0047] The feedback control module is used to send the mixed signal back to the MEMS speaker for playback and collect it again through the MEMS microphone to form a closed-loop feedback system;

[0048] The error analysis and adjustment module is used to calculate the error between the currently collected ambient sound pressure level and the target sound pressure level. If the error exceeds a preset threshold, the speaker output power is adjusted:

[0049] The loop iteration module is used to repeatedly perform the mixing calculation, signal mixing, feedback control and error analysis and adjustment process until the difference between the ambient sound pressure level and the target sound pressure level is lower than the preset threshold, thus completing the mixed signal processing.

[0050] Technical effects and advantages of the present invention: The hybrid signal processing method and system for integrating a MEMS speaker and a microphone proposed in the present invention have the following advantages over the prior art:

[0051] 1. This invention effectively suppresses acoustic feedback and echo through preprocessing, hybrid calculation, signal mixing, feedback control, error analysis, and adjustment. By calculating the mixing coefficient in real time and linearly mixing the speaker output signal with the microphone signal, a dynamic balance is achieved between the speaker and microphone, significantly reducing the risk of acoustic feedback. Furthermore, the closed-loop feedback system can promptly capture echo signals and precisely adjust the speaker output power to effectively suppress echoes and ensure the clarity of voice communications.

[0052] 2. Perform pre-processing operations such as noise suppression, gain control, and automatic level adjustment on the ambient sound signals collected by the microphone. Utilizing advanced noise suppression algorithms, the impact of background noise on voice signals is significantly reduced, significantly improving the signal-to-noise ratio and ensuring high-quality voice interaction even in noisy environments.

[0053] 3. Based on the set target total sound pressure level, the system monitors the current ambient sound pressure level in real time and iteratively adjusts the speaker output power to achieve precise control of the sound pressure level. This system can not only adapt to the sound pressure requirements of various application scenarios, but also quickly respond to changes in ambient sound pressure, avoiding problems caused by excessively high or low sound pressure, thereby ensuring user comfort and equipment safety. BRIEF DESCRIPTION OF THE DRAWINGS

[0054] Figure 1 This is a flow chart of a mixed signal processing method integrating a MEMS speaker and a microphone according to the present invention;

[0055] Figure 2 This is a block diagram of a mixed signal processing system integrating a MEMS speaker and a microphone according to the present invention. DETAILED DESCRIPTION

[0056] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the embodiments described are only part of the embodiments of the present invention, not all of the embodiments. The specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention.

[0057] This invention provides a hybrid signal processing method integrating a MEMS speaker and microphone. It aims to address key issues in existing technologies, such as acoustic feedback, echo suppression, ambient noise suppression, and sound pressure level control, significantly improving the acoustic performance and user experience of audio devices. The details are as follows.

[0058] like Figure 1 As shown, Figure 1 This is a flow chart of a mixed signal processing method integrating a MEMS speaker and a microphone according to the present invention.

[0059] In this embodiment, the mixed signal processing method for integrating a MEMS speaker and a microphone includes the following steps:

[0060] Step 1: Signal acquisition: Receive and output audio signals and collect ambient sound signals through the MEMS speaker and MEMS microphone integrated in the same package. The specific implementation is as follows:

[0061] Hardware selection and principles

[0062] MEMS speakers, models: Currently, MEMS speaker models include Knowles' IS2010 / IS2011 series, TDK's NPM0100 series, and USound's uSpeaker series. These products are based on piezoelectric or electrostatic drive principles, achieving miniaturized, low-power audio output.

[0063] Principle: Take the Knowles IS2010 as an example. It uses piezoelectric film technology. When an electrical signal is applied to the piezoelectric film, the film generates mechanical vibrations, which in turn drive the air to produce sound waves. The speaker incorporates key components such as the diaphragm, piezoelectric film, and back cavity. Through precise structural design and material selection, it ensures wide frequency response, high sound pressure level, and low distortion within a very small volume.

[0064] MEMS microphones, models: Mainstream MEMS microphone models on the market include Knowles SiSonic series, Bosch Sensortec BM series, InvenSense ICS series, etc. These products are mainly capacitive MEMS microphones with high sensitivity, low noise, and excellent signal-to-noise ratio.

[0065] Principle: Taking the Knowles SiSonic series as an example, it uses a capacitor structure formed by a backplate electrode and a diaphragm. When sound waves cause the diaphragm to vibrate, the distance between the capacitors changes, thereby changing the capacitance value. The built-in ASIC (Application Specific Integrated Circuit) converts the capacitance change into an electrical signal output. This type of microphone has extremely high stability and consistency, and can maintain excellent audio acquisition performance in various environments.

[0066] 2. Signal Acquisition and Environmental Sound Collection Process

[0067] Audio signal output and driver circuit: MEMS speakers require a suitable driver circuit to provide adequate voltage and current. Common driver circuits include Class D (digital switching amplifiers) and Class AB (linear amplifiers). For example, the TPA2012D1 is a high-efficiency Class D audio amplifier suitable for MEMS speakers. It uses PWM (pulse width modulation) technology to convert digital audio signals into highly efficient speaker drive signals.

[0068] Signal transmission: Audio signals are typically generated or decoded by a host control chip (such as a DSP or MCU) or an audio codec (CODEC) and then transmitted to the speaker driver circuit via interfaces such as I2S and PDM (Pulse Density Modulation). For example, the Cirrus Logic CS42L52 is an audio CODEC with an integrated I2S interface that supports stereo line input, microphone input, and speaker output, making it ideal for use with MEMS speakers.

[0069] Ambient sound acquisition and front-end amplification: The output signal of a MEMS microphone is typically weak and requires gain adjustment through a preamplifier (PGA). For example, the Maxim MAX98600 is a low-noise, high-gain PGA designed specifically for MEMS microphones. Its built-in anti-aliasing filter helps improve microphone signal quality.

[0070] Analog-to-digital conversion: The amplified analog microphone signal must be converted to a digital signal using an ADC (Analog-to-Digital Converter). For example, the Texas Instruments TLV320ADC3101 is a high-performance, low-power audio ADC with a sampling rate of up to 96kHz, meeting the requirements for high-quality audio acquisition.

[0071] Signal processing: The digitized microphone signal is usually sent to the main control chip or DSP for further processing, such as noise suppression, echo cancellation, automatic gain control (AGC), etc., to optimize audio quality. For example, the Qualcomm QCC512x series Bluetooth audio SoC integrates a high-performance audio DSP that can implement the above-mentioned multiple audio processing functions.

[0072] 3. Packaging Design and System Integration

[0073] The packaging design of integrated MEMS speakers and microphones aims to minimize space occupation while ensuring acoustic isolation between the two to avoid mutual interference. Common packaging forms include the following:

[0074] Coplanar packaging: The speaker and microphone are located on the same plane, and acoustic coupling is reduced through sophisticated acoustic isolation materials and structural design, such as acoustic foam and silicone insulation. For example, Knowles' SiSonic MEMS microphone and IS2010 MEMS speaker can be coplanar packaged to achieve high integration.

[0075] Stacked packaging: A speaker and microphone are stacked vertically, using physical distance and an intermediate medium (such as a PCB or acoustic isolation material) to achieve acoustic isolation. For example, TDK's NPM0100 MEMS speaker and InvenSense ICS-40730 MEMS microphone can be stacked to create an ultra-thin audio module.

[0076] Integrated packaging: This integrates the speaker, microphone, and some driver, amplifier, and ADC circuitry into a single package, creating a true single-chip solution. This type of package is typically developed in collaboration between chip manufacturers and packaging houses, such as Bosch's BNO080, which integrates the BMA400 accelerometer and BMM150 magnetometer.

[0077] Step 2: Preprocessing: Perform power adjustment, frequency equalization, and distortion correction preprocessing operations on the audio signal from the MEMS speaker, and perform noise suppression, gain control, and automatic level adjustment preprocessing operations on the ambient sound signal from the MEMS microphone;

[0078] Exemplarily, the power adjustment in the preprocessing step includes the following steps:

[0079] Signal framing: Divide the audio signal from the MEMS speaker into continuous time frames and perform instantaneous amplitude calculation, including calculating the instantaneous amplitude value of each frame signal;

[0080] Threshold determination: Determine the upper threshold value T of dynamic range compression based on the pre-set compression ratio and inflection point upper and the lower threshold value Tlower , where the compression ratio formula is:

[0081]

[0082] in, and are the upper and lower threshold values ​​corresponding to the uncompressed state respectively;

[0083] Compression processing: For the instantaneous amplitude value of each frame signal, if it is less than the lower limit threshold value T lower , it remains unchanged; if it is located at T lower With T upper Between, nonlinear compression is performed according to formula (4):

[0084]

[0085] Among them, A in is the original instantaneous amplitude value, A comp is the instantaneous amplitude value after compression;

[0086] Reconstruct the signal: Reconstruct the audio signal frame based on the compressed instantaneous amplitude value, and splice all frames to restore the complete audio signal to complete the power adjustment.

[0087] Specifically, signal framing involves dividing the continuous audio signal from the MEMS speaker into frames of fixed length (e.g., 20 milliseconds), forming a series of temporally continuous but non-overlapping audio signal frames. This framing operation facilitates the subsequent independent instantaneous amplitude calculation and compression processing of each frame, ensuring real-time algorithm performance and processing efficiency. For each frame, the instantaneous amplitude value is calculated using either the effective mean square (RMS) or peak detection method. For example, the RMS value can be calculated for discrete audio signal frames.

[0088] DSP-based dynamic range compression of MEMS speaker audio signals is achieved. This process includes signal framing, instantaneous amplitude calculation, threshold determination, compression processing, and signal reconstruction. It effectively adjusts the power of the audio signal, preventing distortion or equipment damage caused by excessive signals while maintaining the overall dynamic range of the audio signal and improving the listening experience.

[0089] In addition, noise suppression includes key steps such as noise reference signal acquisition, noise spectrum estimation, filter initialization, voice activity detection, noise cancellation update and processing, and inverse superposition, aiming to significantly reduce the interference of environmental noise on voice signals and improve the performance of voice communication or voice recognition systems. The following are specific implementation details:

[0090] 1. Noise reference signal acquisition

[0091] Purpose: To collect environmental noise samples and provide a basis for subsequent noise spectrum estimation and adaptive filter initialization.

[0092] Implementation: Use a MEMS microphone integrated into the system to continuously collect ambient noise during periods of insignificant speech activity (such as periods of silence or non-speech periods during user confirmation). The acquisition period should be long enough to obtain representative noise samples, typically ranging from a few seconds to tens of seconds, depending on the specific application scenario and noise characteristics.

[0093] Note: Ensure that there is no significant speech interference during noise reference signal acquisition to avoid inadvertently incorporating speech components into the noise model. A dual-microphone array or multi-microphone system can be used for spatial filtering to further eliminate mixed near-field speech signals.

[0094] 2. Noise spectrum estimation

[0095] Purpose: To convert the collected noise reference signal into frequency domain representation for subsequent filter design and noise cancellation processing.

[0096] Implementation: Perform a Fast Fourier Transform (FFT) on the noise reference signal to convert it from the time domain to the frequency domain. Next, calculate the square of the complex amplitude at each frequency point to obtain the power spectrum of the noise reference signal. To improve computational efficiency and reduce storage requirements, a short-term window (such as a Hamming window or Hann window) can be used, followed by a segmented FFT, to obtain an estimate of the noise power spectrum in the time-frequency distribution.

[0097] 3. Filter initialization

[0098] Purpose: Construct and initialize an adaptive filter for subsequent noise cancellation processing.

[0099] Implementation: Select an adaptive filter structure suitable for noise suppression, such as a minimum mean square error (MMSE) filter, a recursive least squares (RLS) filter, or a Kalman filter. Set the initial filter parameters, such as filter coefficients and step size, based on the noise spectrum estimation results. The initial filter should closely mimic the spectral characteristics of the noise reference signal to lay the foundation for subsequent noise tracking and cancellation.

[0100] 4. Voice Activity Detection (VAD)

[0101] Purpose: To determine in real time whether there is voice activity in the current audio stream, so as to perform noise cancellation processing during voice activity periods and perform noise tracking updates during non-voice activity periods.

[0102] Implementation: Use one or more voice activity detection algorithms, such as energy thresholding, zero-crossing rate detection, spectral flatness analysis, and deep learning-based VAD. Combine the outputs of these algorithms and set appropriate decision logic (such as majority voting or weighted fusion) to determine whether the current audio frame contains speech. The VAD algorithm should be robust and adaptable to varying signal-to-noise ratios, speaking styles, and background noise conditions.

[0103] 5. Noise cancellation update

[0104] Objective: To update the adaptive filter parameters using a noise reference signal during periods of non-speech activity, so that it can accurately track changes in ambient noise.

[0105] Implementation: When the VAD determines that there is no voice activity, it uses the real-time ambient sound signal as the error signal and adjusts the filter coefficients using the adaptive filter update rule (such as the update steps of the LMS algorithm, NLMS algorithm, or RLS algorithm) to make the filter output signal consistent with the noise reference signal. This process continues until voice activity is detected.

[0106] 6. Noise cancellation processing

[0107] Purpose: During periods of speech activity, an adaptive filter is applied to process the real-time collected ambient sound signal, filtering out the parts similar to the noise spectrum estimate and generating a noise-canceled signal.

[0108] Implementation: When the VAD determines that the current period is speech activity, the real-time ambient sound signal is used as the input for the adaptive filter. The filter output is the speech signal after preliminary noise reduction. At this time, the filter parameters remain unchanged, and noise suppression is performed using the optimal filter state obtained in the previous noise tracking stage.

[0109] 7. Inverse superposition

[0110] Purpose: To further eliminate the residual noise components after noise cancellation processing and improve the signal-to-noise ratio of the speech signal.

[0111] Implementation: The noise-canceled signal is then inverted and superimposed with the original ambient sound signal. Specifically, the difference between the original and noise-reduced signals is calculated. Since the noise components are roughly the same in both signals, the noise components cancel each other out after subtraction. However, the randomness of the speech components does not completely cancel each other out, thus achieving further noise reduction. This step is optional and depends on the specific application's noise reduction requirements and system resource constraints.

[0112] Through the above-described implementation, a MEMS microphone-based noise suppression system has been constructed. Through steps such as noise reference signal acquisition, noise spectrum estimation, filter initialization, voice activity detection, noise cancellation update and processing, and inverted superposition, this system effectively suppresses ambient noise, significantly improving the performance of voice communication or speech recognition systems. In practical applications, each step can be optimized and adjusted based on specific scenarios and device characteristics to ensure a balance between noise suppression effectiveness, system real-time performance, and resource consumption.

[0113] In addition, the automatic level adjustment in the pre-processing step is based on the real-time monitored signal-to-noise ratio and adjusts the microphone gain according to a preset signal-to-noise ratio target value. The specific embodiment is as follows:

[0114] 1. Real-time signal-to-noise ratio (SNR) monitoring

[0115] Purpose: To evaluate the signal-to-noise ratio of the current speech signal in real time and provide a basis for automatic level adjustment.

[0116] Implementation: During the preprocessing phase of the speech signal, a signal-to-noise ratio estimation algorithm is used to calculate the signal-to-noise ratio of the audio signal collected by the microphone in real time. Commonly used signal-to-noise ratio estimation algorithms include but are not limited to:

[0117] Based on the short-term energy ratio: The ratio of the short-term energy of the speech frame to the short-term energy of the noise frame is calculated, which is approximately the signal-to-noise ratio. The noise frame is usually taken from the signal during the period determined by the VAD to be non-speech activity.

[0118] Based on spectral entropy: Analyze the spectral distribution of the speech frame, calculate its entropy value, and compare it with the noise entropy value to reflect the signal-to-noise ratio level.

[0119] Based on deep learning: Train a specialized neural network model to directly predict the signal-to-noise ratio of the input audio clip.

[0120] Note: The signal-to-noise ratio estimation algorithm should be real-time and able to quickly respond to noise changes without affecting the overall system latency.

[0121] 2. Preset SNR target value setting

[0122] Purpose: Defines the desired speech signal-to-noise ratio target to guide automatic level adjustment.

[0123] Implementation: Predetermine one or more SNR target values ​​based on the performance requirements, application scenarios, and user preferences of the voice communication or speech recognition system. For example, for high-definition calls, the SNR target value could be set to above 20dB; for voice interaction in noisy environments, the requirement could be appropriately lowered, such as to 15dB. The target values ​​should have a dynamic adjustment mechanism to accommodate different operating modes or environmental changes.

[0124] 3. Microphone gain control

[0125] Purpose: To dynamically adjust the microphone gain based on the real-time monitored SNR and the preset SNR target value to optimize the voice signal quality.

[0126] Implementation: a) Real-time comparison: Compare the signal-to-noise ratio calculated in real time with the preset SNR target value. b) Gain adjustment decision: If the real-time signal-to-noise ratio is lower than the preset target value, increase the microphone gain to increase the signal strength and relatively improve the signal-to-noise ratio; conversely, if the real-time signal-to-noise ratio is higher than the preset target value and close to saturation (to prevent overload), appropriately reduce the microphone gain to maintain it within an appropriate signal-to-noise ratio range. c) Gain adjustment amplitude: The gain adjustment amplitude should follow the principle of smoothness to avoid sudden large-scale changes that cause signal distortion or auditory discomfort. Control theory methods such as proportional integral controller (PID) and adaptive controller can be used to design a reasonable gain adjustment strategy. d) Gain upper and lower limits: Set the upper and lower limits of the microphone gain to prevent excessive amplification from introducing too much noise or excessive attenuation from causing loss of voice signals. The specific setting of the gain range should take into account hardware characteristics, application scenarios and user comfort.

[0127] 4. Feedback and Adaptive Adjustment

[0128] Purpose: To continuously optimize the gain adjustment strategy based on the speech signal quality after automatic level adjustment and system feedback.

[0129] Implementation: Evaluate the effectiveness of automatic level control periodically or when system conditions change (e.g., changes in ambient noise characteristics, user location, etc.), using methods such as voice quality evaluation indicators (PESQ, MOS, etc.), false recognition rates, and user satisfaction surveys. Based on the evaluation results, dynamically adjust the preset SNR target value or gain adjustment strategy to adapt to the new conditions and ensure optimal voice signal quality.

[0130] In summary, automatic level adjustment is achieved by dynamically adjusting the microphone gain to optimize the voice signal quality in a mixed-signal processing approach that integrates a MEMS speaker and a microphone.

[0131] Step 3: Mixing calculation: Use the following formula (1) to calculate the mixing coefficient to achieve real-time mixing of the speaker output signal and the microphone acquisition signal:

[0132]

[0133] Among them, P target is the target total sound pressure level, P mic is the ambient sound pressure level collected by the current microphone, P spkis the sound pressure level output by the current speaker; further, the target total sound pressure level P target It can be set according to the actual application scenario. It aims to provide users with flexible and intuitive operation methods to adapt to the sound pressure level requirements of different application scenarios. The specific implementation method is as follows:

[0134] 1. User Interface Design

[0135] Purpose: To design a simple and easy-to-use user interface for users to set the target total sound pressure level.

[0136] Implementation: a) Sound pressure level adjustment control: Add a sound pressure level adjustment control, such as a slider, numeric input box, knob, etc., to the device's settings menu or dedicated audio adjustment interface. The default value of the control can be set to the recommended commonly used sound pressure level, and the range should cover the device's available sound pressure level range. b) Unit display: Clearly mark the unit (such as dBSPL) next to the sound pressure level adjustment control to help users understand the meaning of the setting value. c) Real-time feedback: When the user adjusts the sound pressure level, the current setting value should be displayed on the interface in real time to provide visual feedback. If supported by the device, audio examples corresponding to the set sound pressure level can also be played in real time to provide auditory feedback. d) Save and apply: Provide a "Save" or "Apply" button. After the user confirms the setting, the system will use the newly set target total sound pressure level for subsequent mixing calculations.

[0137] 2. User setup process

[0138] Purpose: To guide the user through the user interface to complete the setting of the target total sound pressure level.

[0139] Implementation: a) Access the settings interface: The user enters the sound pressure level settings interface through the device menu, shortcut keys, or voice commands. b) Adjust the sound pressure level: The user sets the target total sound pressure level through the sound pressure level adjustment control based on the actual application scenario and personal preferences. For example, the user may prefer a lower sound pressure level in a quiet environment to maintain quietness, while in a noisy environment, the user may need to increase the sound pressure level to ensure that the speaker output is clearly heard. c) Confirm and save: After the user confirms that the set sound pressure level is correct, click the "Save" or "Apply" button, and the system will save the newly set sound pressure level. target Save to device configuration and immediately use in subsequent hybrid calculations.

[0140] 3. Device Response and Hybrid Computing Updates

[0141] Purpose: After receiving the target total sound pressure level set by the user, the device updates the P in the mixing calculation process. target value.

[0142] Implementation: a) Receiving setting value: The device main control unit receives the new setting P from the user interface targetb) Update the mixing coefficient calculation: In the subsequent mixing calculation process, use the newly set P target Replace the original value and calculate the mixing coefficient k according to formula (1) to achieve real-time mixing of the speaker output signal and the microphone acquisition signal. c) Real-time effect: The device immediately applies the updated mixing coefficient to mix the signals, and the user can immediately feel the change in sound pressure level.

[0143] Through user interface design, user setup process, device response and hybrid calculation updates, user-defined setting of the target total sound pressure level in the hybrid signal processing method integrating MEMS speakers and microphones is achieved. This gives users the ability to flexibly adjust the sound pressure level according to actual application scenarios, improving the personalized experience and adaptability of the device.

[0144] Step 4: Signal mixing: Using the calculated mixing coefficient k, the pre-processed loudspeaker output signal S is mixed according to formula (2). out and microphone to collect signal M in Perform linear blending:

[0145] S mix =(1-k)·S out +k·M in ;

[0146] It aims to achieve effective integration of speaker output and microphone acquisition signals to meet the needs of different application scenarios. The details are as follows:

[0147] 1. Mixing coefficient calculation

[0148] Purpose: Calculate the mixing coefficient k according to formula (1). According to the set target total sound pressure level P target , the ambient sound pressure level P collected by the current microphone mic And the sound pressure level P output by the current speaker spk , use formula (1) to calculate the mixing coefficient k:

[0149] 2. Signal preprocessing, output signal S to the speaker out and microphone to collect signal M in Perform necessary preprocessing, such as power adjustment, frequency equalization, noise suppression, gain control, etc.

[0150] Implementation: Based on the specific algorithm and parameters in the preprocessing step, out and M in Perform corresponding processing to ensure that the quality of the pre-processed signal meets the requirements of hybrid computing.

[0151] 3. Linear mixing, using the calculated mixing coefficient k and the preprocessed speaker output signal S out , microphone collects signal Min , perform linear mixing according to formula (2) to generate a mixed signal S mix .

[0152] Implementation: a) Calculate the mixed signal: According to formula (2), the mixing coefficient k, the pre-processed loudspeaker output signal S out and microphone to collect signal M in Combined, calculate the mixed signal S mix :

[0153] b) Data format conversion: Ensure that the signal S involved in the mixing out and M in The data formats (such as sampling rate, quantization bit number, data type, etc.) must be consistent to facilitate mathematical operations.

[0154] c) Floating-point and fixed-point operations: Based on the hardware characteristics of the device, select the appropriate floating-point or fixed-point operation method for mixed calculations. If fixed-point operations are used, the impact of quantization error on the mixed results must be considered to ensure the quality of the mixed signal.

[0155] d) Overflow processing: During the hybrid calculation process, monitor whether the data exceeds the device data representation range. If overflow occurs, adopt strategies such as saturation truncation, rounding or wraparound to prevent signal distortion.

[0156] 4. Mixed signal output

[0157] Purpose: To convert the mixed signal S mix Output to subsequent processing links or speakers for playback.

[0158] Implementation: Mixed signal S mix The mixed signal is then sent to the device's main control unit or audio processor for further post-processing (such as digital sound effects processing and echo cancellation) or directly to the speaker for playback. This ensures that there is no data loss or significant delay during the transmission of the mixed signal, maintaining the continuity and real-time nature of the sound quality.

[0159] Step 5: Feedback control: The mixed signal S mix The signal is sent back to the MEMS speaker for playback and collected again by the MEMS microphone to form a closed-loop feedback system. Furthermore, the signal collected by the MEMS microphone is processed with delay compensation and then used for closed-loop feedback.

[0160] Step 6: Error analysis and adjustment: Calculate the error between the currently collected ambient sound pressure level and the target sound pressure level. If the error exceeds the preset threshold, adjust the speaker output power according to formula (3):

[0161] P spk,new =P spk,old +α·(P target+P curent )

[0162] Among them, P spk,new is the output power of the speaker after adjustment, P spk,old To adjust the front speaker output power, α is the power adjustment coefficient, P curent is the current ambient sound pressure level; further, the power adjustment coefficient α can be dynamically adjusted according to the actual application scenario.

[0163] Step 7: Iterate: Repeat the mixing calculation, signal mixing, feedback control, and error analysis and adjustment process until the difference between the ambient sound pressure level and the target sound pressure level falls below a preset threshold, completing the mixed signal processing. Furthermore, if the difference between the ambient sound pressure level and the target sound pressure level remains above the preset threshold after several consecutive iterations, the target sound pressure level is automatically lowered.

[0164] In some other embodiments, the above-mentioned method for processing mixed signals integrating a MEMS speaker and a microphone further includes: performing digital sound effect processing on the mixed signal. Specific implementation methods are as follows:

[0165] 1. Digital sound processor module integration: integrating a digital sound processor (Digital Signal Processor, DSP) or a hardware unit with similar functions at the hardware level to perform real-time digital sound processing on the mixed signal.

[0166] Implementation: a) Add a dedicated digital audio processor chip to the existing hardware design or select a mixed signal processor with built-in audio processing capabilities (such as an MCU with a DSP core). b) Ensure that there is a high-speed, low-latency data channel between the processor and the speaker output interface, microphone input interface and mixed signal processing module, such as I 2 c) Configure necessary power management and clock signal distribution to ensure stable operation of DSP.

[0167] 2. Audio effect algorithm selection and implementation: determine the required digital audio effect type, write or select the corresponding digital signal processing algorithm, and deploy it in the digital audio processor.

[0168] Implementation: a) Based on the application scenario requirements, select appropriate sound effects, such as equalizer (EQ), dynamic range compression (DRC), noise suppression (NS), acoustic echo cancellation (AEC), reverberation (Reverb), and voice changer. b) Obtain or develop corresponding digital signal processing algorithms, which can be based on filter design (such as FIR / IIR filters), frequency domain transforms (such as FFT / IFFT), and adaptive algorithms (such as LMS / RLS). c) Compile the algorithms into code suitable for the target DSP architecture, such as assembly instructions or C / C++ code, and optimize performance to meet real-time processing requirements.

[0169] 3. Expand the mixed signal processing process by adding digital sound processing to the original mixed signal processing process.

[0170] Implementation: Step 1: Mixing calculation and signal mixing, keeping the original process unchanged, generating a preliminary mixed signal. Step 2: Digital sound effect processing, a) sending the preliminary mixed signal to the digital sound effect processor. b) The digital sound effect processor processes the sound effect according to the preset sound effect parameters and algorithms to generate a sound effect processed signal. Step 3: Feedback control, error analysis and adjustment (keeping the original process unchanged); perform feedback control, error analysis and adjustment on the sound effect processed signal to determine whether the preset threshold of the difference between the ambient sound pressure level and the target sound pressure level is reached, and perform corresponding iterations or target sound pressure level adjustments. Step 4: Output processing, output the final adjusted signal (i.e., the signal that has undergone sound effect processing and sound pressure level control) to the speaker for playback.

[0171] Sound effect parameter configuration and real-time control, providing a user interface or API interface to flexibly adjust sound effect parameters or dynamically update sound effect settings according to environmental changes.

[0172] Implementation: a) Develop a user interface (e.g., a graphical software panel, mobile app) that allows users to select sound effect types and adjust parameters (e.g., gain, cutoff frequency, reverberation time, etc.). b) For automated or intelligent applications, design a sensor data fusion mechanism that dynamically adjusts sound effect parameters based on factors such as ambient noise levels and spatial characteristics to improve sound quality or adaptability. c) Implement a communication interface with the digital sound effects processor to ensure that parameter changes are transmitted to the processor in real time and take effect.

[0173] By performing digital sound processing on the mixed signal, the functional diversity and sound quality of mixed-signal processing systems integrating MEMS speakers and microphones are enhanced. Through hardware integration, algorithm implementation, process expansion, and sound parameter configuration and real-time control, seamless conversion from mixed signals to sound-rich output signals is achieved, improving user experience and adapting to complex application scenarios.

[0174] In addition, this embodiment also proposes a hybrid signal processing system integrating a MEMS speaker and a microphone, such as Figure 2 As shown, it includes: a signal acquisition module, a preprocessing module, a hybrid calculation module, a feedback control module, an error analysis and adjustment module, and a loop iteration module.

[0175] Specifically, the signal acquisition module is used to receive and output audio signals and collect ambient sound signals respectively through the MEMS speaker and MEMS microphone integrated in the same package;

[0176] Specifically, the preprocessing module is used to perform power adjustment, frequency equalization, and distortion correction preprocessing operations on the audio signal from the MEMS speaker, and to perform noise suppression, gain control, and automatic level adjustment preprocessing operations on the ambient sound signal from the MEMS microphone;

[0177] Specifically, the mixing calculation module is used to calculate the mixing coefficient to achieve real-time mixing of the loudspeaker output signal and the microphone acquisition signal;

[0178] Specifically, the feedback control module is used to send the mixed signal back to the MEMS speaker for playback and collect it again through the MEMS microphone to form a closed-loop feedback system;

[0179] Specifically, the error analysis and adjustment module is used to calculate the error between the currently collected ambient sound pressure level and the target sound pressure level. If the error exceeds a preset threshold, the speaker output power is adjusted:

[0180] Specifically, the loop iteration module is used to repeatedly perform the mixing calculation, signal mixing, feedback control, and error analysis and adjustment process until the difference between the ambient sound pressure level and the target sound pressure level is lower than a preset threshold, completing the mixed signal processing.

[0181] In addition, the above-mentioned signal acquisition module, preprocessing module, hybrid calculation module, feedback control module, error analysis and adjustment module, and loop iteration module are also used to implement other steps of the above-mentioned hybrid signal processing method integrating MEMS speakers and microphones during execution, which will not be described one by one here.

[0182] In summary, the present invention achieves effective acoustic feedback and echo suppression through preprocessing, hybrid calculation, signal mixing, feedback control, error analysis, and adjustment. By calculating the mixing coefficient in real time and linearly mixing the speaker output signal with the microphone signal, a dynamic balance is achieved between the speaker and microphone, significantly reducing the risk of acoustic feedback. Furthermore, the closed-loop feedback system can promptly capture echo signals and precisely adjust the speaker output power to effectively suppress echoes, ensuring the clarity of voice communications.

[0183] The ambient sound signals collected by the microphone are pre-processed with noise suppression, gain control, and automatic level adjustment. Advanced noise suppression algorithms are used to significantly reduce the impact of background noise on voice signals and significantly improve the signal-to-noise ratio, ensuring high-quality voice interaction even in noisy environments.

[0184] Based on the set target total sound pressure level, the current ambient sound pressure level is monitored in real time, and the speaker output power is iteratively adjusted to achieve precise control of the sound pressure level. This can not only adapt to the sound pressure requirements of various application scenarios, but also quickly respond to changes in ambient sound pressure, avoiding problems caused by excessively high or low sound pressure, and ensuring user comfort and equipment safety.

[0185] In addition, the present invention also provides a terminal device. The mixed signal processing method of the integrated MEMS speaker and microphone involved in this embodiment is mainly used in the terminal device, which can be a PC, portable computer, mobile terminal and other devices with display and processing functions.

[0186] Specifically, a terminal device may include a processor (e.g., a CPU), a communication bus, a user interface, a network interface, and a memory. The communication bus is used to enable communication between these components; the user interface may include a display and an input unit such as a keyboard; the network interface may optionally include a standard wired interface or a wireless interface (e.g., a Wi-Fi interface); and the memory may be high-speed RAM or non-volatile memory, such as a disk drive. The memory may also be a storage device independent of the processor.

[0187] The memory stores a readable storage medium, and the readable storage medium stores a mixed signal processing program. The processor can call the mixed signal processing program stored in the memory and execute the mixed signal processing method of the integrated MEMS speaker and microphone provided in the embodiment of the present invention.

[0188] It will be understood that a computer-readable storage medium can be a tangible device that can hold and store instructions used by an instruction execution device. A computer-readable storage medium can be, for example, but not limited to, an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination thereof. More specific examples (a non-exhaustive list) of computer-readable storage media include: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disk (DVD), a memory stick, a floppy disk, a mechanical encoding device, such as a punch card or a raised structure in a groove on which instructions are stored, and any suitable combination thereof. As used herein, a computer-readable storage medium is not to be construed as a transient signal per se, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagating through a waveguide or other transmission medium (e.g., a light pulse through a fiber optic cable), or an electrical signal transmitted through an electrical wire.

[0189] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to each computing / processing device, or downloaded to an external computer or external storage device via a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network can include copper transmission cables, fiber optic transmission, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. The network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions to be stored in the computer-readable storage medium in each computing / processing device.

[0190] The computer program instructions for performing the operations of the present disclosure may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages ​​such as Smalltalk, C++, and conventional procedural programming languages ​​such as "C" language or similar programming languages. Computer-readable program instructions may be executed entirely on a user's computer, partially on a user's computer, as an independent software package, partially on a user's computer, partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., utilizing an Internet service provider to connect via the Internet). In some embodiments, an electronic circuit, such as a programmable logic circuit, a field programmable gate array (FPGA), or a programmable logic array (PLA), may be personalized by utilizing the state information of the computer-readable program instructions. The electronic circuit may execute the computer-readable program instructions, thereby realizing various aspects of the present disclosure.

[0191] Finally, it should be noted that the above is only a preferred embodiment of the present invention and is not intended to limit the present invention. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art can still modify the technical solutions described in the aforementioned embodiments or make equivalent substitutions for some of the technical features therein. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.

Claims

1. A mixed signal processing method integrating a MEMS speaker and a microphone, characterized in that: The steps include: Signal acquisition: The MEMS speaker and MEMS microphone integrated in the same package receive and output audio signals and collect ambient sound signals respectively; Preprocessing: Perform power adjustment, frequency equalization, and distortion correction on the audio signal from the MEMS speaker, and perform noise suppression, gain control, and automatic level adjustment on the ambient sound signal from the MEMS microphone; Mixing calculation: The mixing coefficient is calculated using the following formula (1) to achieve real-time mixing of the speaker output signal and the microphone acquisition signal: in, is the target total sound pressure level, is the ambient sound pressure level collected by the current microphone, The sound pressure level of the current speaker output; Signal mixing: Using the calculated mixing coefficient k, the pre-processed loudspeaker output signal is mixed according to formula (2) and microphone to collect signals Perform linear blending: ; Feedback control: The mixed signal The data is sent back to the MEMS speaker for playback and collected again by the MEMS microphone to form a closed-loop feedback system. Error analysis and adjustment: Calculate the error between the currently collected ambient sound pressure level and the target sound pressure level. If the error exceeds the preset threshold, adjust the speaker output power according to formula (3): in, To adjust the speaker output power, To adjust the front speaker output power, is the power adjustment factor, is the current ambient sound pressure level; Iterative loop: Repeat the mixing calculation, signal mixing, feedback control, and error analysis and adjustment process until the difference between the ambient sound pressure level and the target sound pressure level is lower than the preset threshold, completing the mixed signal processing.

2. The mixed signal processing method for integrating a MEMS speaker and a microphone according to claim 1, characterized in that: The power adjustment in the pre-processing step includes the following steps: Signal framing: Divide the audio signal from the MEMS speaker into continuous time frames and perform instantaneous amplitude calculation, including calculating the instantaneous amplitude value of each frame signal; Threshold determination: Determine the upper threshold value of dynamic range compression based on the pre-set compression ratio and inflection point and lower threshold , where the compression ratio formula is: in, and are the upper and lower threshold values ​​corresponding to the uncompressed state respectively; Compression processing: For the instantaneous amplitude value of each frame signal, if it is less than the lower limit threshold , it remains unchanged; if it is located and Between, nonlinear compression is performed according to formula (4): in, is the original instantaneous amplitude value, is the instantaneous amplitude value after compression; Reconstruct the signal: Reconstruct the audio signal frame based on the compressed instantaneous amplitude value, and splice all frames to restore the complete audio signal to complete the power adjustment.

3. The mixed signal processing method for integrating a MEMS speaker and a microphone according to claim 1, characterized in that: The noise suppression in the preprocessing step comprises the following steps: Noise reference signal acquisition: A MEMS microphone is used to collect ambient noise during non-speech activity as a noise reference signal; Noise spectrum estimation: Perform fast Fourier transform on the noise reference signal and calculate its frequency domain power spectrum as the noise spectrum estimation; Filter initialization: construct an adaptive filter whose initial parameters are set according to the noise spectrum estimate; Voice activity detection: During real-time processing, a voice activity detection algorithm is used to determine whether the current period is voice active. Noise cancellation update: During periods of non-speech activity, the noise reference signal is used to update the adaptive filter parameters so that the filter output signal is consistent with the noise reference signal; Noise cancellation processing: During periods of speech activity, an adaptive filter is applied to the real-time collected ambient sound signal to filter out the parts that are similar to the noise spectrum estimate, generating a noise-canceled signal. Inverse superposition: The noise-canceled signal is superimposed on the original ambient sound signal in an inverse phase to eliminate residual noise and complete noise suppression.

4. The mixed signal processing method for integrating a MEMS speaker and a microphone according to claim 1, characterized in that: The automatic level adjustment in the pre-processing step is based on the real-time monitored signal-to-noise ratio and adjusts the microphone gain according to a preset signal-to-noise ratio target value.

5. The mixed signal processing method for integrating a MEMS speaker and a microphone according to claim 1, characterized in that: The target total sound pressure level in the mixing calculation step It can be set according to the actual application scenario.

6. The mixed signal processing method for integrating a MEMS speaker and a microphone according to claim 1, characterized in that: In the feedback control step, the signal collected by the MEMS microphone is processed with delay compensation before being used for closed-loop feedback.

7. The mixed signal processing method for integrating a MEMS speaker and a microphone according to claim 1, characterized in that: In the error analysis and adjustment steps, the power adjustment coefficient It can be dynamically adjusted according to the actual application scenario.

8. The mixed signal processing method for integrating a MEMS speaker and a microphone according to claim 1, characterized in that: During the iterative cycle, if the difference between the ambient sound pressure level and the target sound pressure level is still higher than a preset threshold after several consecutive iterations, the target sound pressure level is automatically lowered.

9. The mixed signal processing method for integrating a MEMS speaker and a microphone according to claim 1, characterized in that: Also includes: Perform digital sound processing on the mixed signal.

10. A mixed signal processing system integrating a MEMS speaker and a microphone for implementing the method according to any one of claims 1 to 9, characterized in that: include: A signal acquisition module is used to receive and output audio signals and collect ambient sound signals through a MEMS speaker and a MEMS microphone integrated in the same package; A pre-processing module, configured to perform power adjustment, frequency equalization, and distortion correction pre-processing operations on the audio signal from the MEMS speaker, and to perform noise suppression, gain control, and automatic level adjustment pre-processing operations on the ambient sound signal from the MEMS microphone; A mixing calculation module is used to calculate the mixing coefficient to achieve real-time mixing of the speaker output signal and the microphone acquisition signal; The feedback control module is used to send the mixed signal back to the MEMS speaker for playback and collect it again through the MEMS microphone to form a closed-loop feedback system; The error analysis and adjustment module is used to calculate the error between the currently collected ambient sound pressure level and the target sound pressure level. If the error exceeds a preset threshold, the speaker output power is adjusted: The loop iteration module is used to repeatedly perform the mixing calculation, signal mixing, feedback control and error analysis and adjustment process until the difference between the ambient sound pressure level and the target sound pressure level is lower than the preset threshold, thus completing the mixed signal processing.

Citation Information

Patent Citations

  • Feedback multichannel active control system and method before mixing of in-vehicle road noise

    CN116189649A

  • Adaptive feedback control for earbuds, headphones, and handsets

    US20160300562A1