Voice noise reduction method and device based on vibration sensor, equipment and medium
Through a multi-step processing method based on vibration sensors, combined with adaptive filtering and deep learning, the problem of insufficient accuracy and stability of traditional speech recognition systems in vibrating noise environments is solved, and higher speech recognition accuracy and environmental adaptability are achieved.
Patent Information
- Application Number
- CN202510622207.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-15
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2045-05-15
AI Technical Summary
When traditional speech recognition systems face vibration noise in complex environments, their accuracy and stability are seriously affected, especially in industrial environments, cities with busy traffic and noisy outdoor scenes, background noise and vibration lead to distortion of voice signals. The existing technology lacks an effective noise reduction solution for vibration noise.
The speech noise reduction method based on vibration sensor is adopted to collect mixed signals through the microphone and synchronously acquire the reference signal of the vibration sensor. The linear echo estimation is generated using an adaptive filter. Combined with the deep learning compensation network and the Kalman filter module, multi-step processing is performed to reduce noise, including time-frequency feature extraction, nonlinear echo estimation and dual-talk state decision-making, and output pure near-end speech.
It effectively reduces noise in speech signals, improves the accuracy and clarity of speech recognition, enhances the adaptability to various environmental characteristics and responds to dynamic environmental changes, and improves robustness.
Smart Images

Figure CN120148539A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of speech noise reduction, and particularly to a speech noise reduction method, device, equipment and medium based on a vibration sensor. Background Art
[0002] With the development of technology, speech recognition technology has been widely used in many fields, such as intelligent assistants, autonomous driving, telephone customer service, etc. However, in the face of vibration noise in complex environments, which usually originates from mechanical equipment, transportation vehicles, etc. and has strong low-frequency characteristics, the accuracy and stability of traditional speech recognition systems are often seriously affected. Especially in noisy scenarios such as industrial environments, busy cities, and outdoors, background noise and vibration will cause distortion of speech signals, thus reducing the performance of the recognition system. In the prior art, although there are various noise reduction solutions, there is no technical solution for vibration noise. Therefore, how to improve the anti-interference ability of speech recognition in these noisy environments has become an important topic in the current research of speech recognition technology. Summary of the Invention
[0003] Based on this, in view of the existing speech noise reduction problem based on a vibration sensor, a speech noise reduction method, device, equipment and medium based on a vibration sensor are proposed.
[0004] A speech noise reduction method based on a vibration sensor, the method is applied to a vibration sensor circuit, the vibration sensor circuit includes a vibration sensor, and the vibration sensor is in contact connection with a vibrating object. The method includes: Collect a mixed signal through a microphone and simultaneously collect a reference signal of the vibration sensor; Generate a linear echo estimate through an adaptive filter based on the reference signal and the mixed signal, and calculate an initial residual; Extract the time-frequency features of the initial residual and the sensor reference signal and input them into a pre-trained deep learning compensation network to obtain a non-linear echo estimate; Adjust the initial residual according to a double-talk state decision function and the non-linear echo estimate to obtain a target residual; Process the target residual through a post-Kalman filter noise reduction module and output a pure proximal speech.
[0005] Further, the step of generating a linear echo estimate through an adaptive filter based on the reference signal and calculating an initial residual includes: Iteratively adjust the coefficient of the reference signal through a preset algorithm ; where represents the coefficient of the k-th impulse response; According to the formula Calculate the linear echo estimation; wherein, represents the linear echo estimation, represents the length of the impulse response in the reference signal, represents the (n - k)-th impulse response in the reference signal; According to the formula Minimize the error signal to obtain the initial residual; wherein, represents the mixed signal, represents the initial residual.
[0006] Furthermore, the preset algorithm is one of the least mean square algorithm, the normalization algorithm, and the recursive least squares method.
[0007] Furthermore, the steps of extracting the time-frequency features of the initial residual and the sensor reference signal and inputting them into a pre-trained deep learning compensation network to obtain the non-linear echo estimation include: Convert the initial residual and the sensor reference signal into frequency signals respectively through short-time Fourier transform, and correspondingly obtain the initial residual frequency signal and the sensor reference frequency signal; Input the initial residual frequency signal and the sensor reference frequency signal into the pre-trained deep learning compensation network to obtain the non-linear echo time-frequency mask; Apply the non-linear echo time-frequency mask to the initial residual frequency signal and perform inverse short-time Fourier transform to obtain the non-linear echo estimation.
[0008] Furthermore, the steps of adjusting the initial residual according to the double-talk state decision function and the non-linear echo estimation to obtain the target residual include: Calculate the short-time energy of the reference signal and the initial residual respectively to obtain the reference signal short-time energy and the initial residual short-time energy; Calculate the ratio of the initial residual short-time energy to the reference signal short-time energy and determine whether the ratio is greater than a preset threshold; If the ratio is greater than the threshold, it is recorded as the double-talk state, and then according to the formula Adjust the residual output; wherein, is the target residual, is the initial residual, is the non-linear echo estimation; If the ratio is less than or equal to the threshold, it is recorded as the non-double-talk state, and then according to the formula Adjust the residual output.
[0009] Furthermore, the steps of processing the target residual through a post-Kalman filter noise reduction module and outputting the pure proximal speech include: Perform a short-time Fourier transform on the target residual to obtain a residual time-frequency signal; Model each frequency point of the residual time-frequency signal to obtain a state vector for each frequency point , where represents the state vector of the t-th frequency point, represents the clean speech amplitude spectrum coefficient of the t-th frequency point, represents the noise amplitude spectrum coefficient of the t-th frequency point; Assume that the speech and noise follow a first-order autoregressive model ; where , are first-order autoregressive coefficients, which are iteratively updated through Kalman filtering, respectively represent process noise; Perform Kalman filter iteration on the state vector of each frequency point to obtain a target state vector; Extract clean proximal speech from the target state vector.
[0010] The present invention also provides a vibration sensor circuit for collecting the above-mentioned reference signal, including: a power supply filtering module, a vibration sensor module, a positive feedback amplification module, a two-stage filtering module, and a chip; The power supply filtering module is connected to the vibration sensor module. The power supply filtering module is used to filter and supply power to the vibration sensor. The vibration sensor module is in contact connection with the vibrating object to obtain a vibration signal; The vibration sensor module is connected to the positive feedback amplification module. The positive feedback amplification module is used to amplify the vibration signal; The positive feedback amplification module is connected to the two-stage filtering module. The two-stage filtering module is used to filter the amplified vibration signal to obtain a reference signal; The two-stage filtering module is connected to the chip and is used to transmit the reference signal to the chip for processing.
[0011] The present invention also provides a voice noise reduction device based on a vibration sensor. The device is applied to a vibration sensor circuit. The vibration sensor circuit includes a vibration sensor, and the vibration sensor is in contact connection with a vibrating object. The device includes: An acquisition module, configured to acquire a mixed signal through a microphone and synchronously acquire a reference signal of the vibration sensor; A generation module, configured to generate a linear echo estimate through an adaptive filter based on the reference signal and the mixed signal, and calculate an initial residual; An extraction module, configured to extract the time-frequency features of the initial residual and the sensor reference signal, and input them into a pre-trained deep learning compensation network to obtain a non-linear echo estimation; An adjustment module, configured to adjust the initial residual according to a double-talk state decision function and the non-linear echo estimation to obtain a target residual; An output module, configured to process the target residual through a post Kalman filtering noise reduction module and output a pure near-end speech.
[0012] A computer device, including a memory and a processor, where the memory stores a computer program, and when the computer program is executed by the processor, the processor performs the following steps: Collect a mixed signal through a microphone and synchronously collect a reference signal of a vibration sensor; Generate a linear echo estimation based on the reference signal and the mixed signal through an adaptive filter and calculate an initial residual; Extract the time-frequency features of the initial residual and the sensor reference signal, and input them into a pre-trained deep learning compensation network to obtain a non-linear echo estimation; Adjust the initial residual according to a double-talk state decision function and the non-linear echo estimation to obtain a target residual; Process the target residual through a post Kalman filtering noise reduction module and output a pure near-end speech.
[0013] A computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the processor performs the following steps: Collect a mixed signal through a microphone and synchronously collect a reference signal of a vibration sensor; Generate a linear echo estimation based on the reference signal and the mixed signal through an adaptive filter and calculate an initial residual; Extract the time-frequency features of the initial residual and the sensor reference signal, and input them into a pre-trained deep learning compensation network to obtain a non-linear echo estimation; Adjust the initial residual according to a double-talk state decision function and the non-linear echo estimation to obtain a target residual; Process the target residual through a post Kalman filtering noise reduction module and output a pure near-end speech.
[0014] The beneficial effects of the present invention: The voice noise reduction method based on a vibration sensor effectively reduces the noise in the voice signal through multiple processing steps, improving the accuracy and clarity of voice recognition. Combining the advantages of adaptive filtering and deep learning enables the method to have good adaptability to various environmental characteristics, improving the response ability and robustness to dynamic environmental changes. Brief Description of the Drawings
[0015] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0016] Among them: Figure 1 It is an application environment diagram of a voice noise reduction method based on a vibration sensor in an embodiment; Figure 2 It is a flowchart of a voice noise reduction method based on a vibration sensor in an embodiment; Figure 3 It is a circuit diagram of a vibration sensor in an embodiment; Figure 4 It is a structural block diagram of a voice noise reduction device based on a vibration sensor in an embodiment; Figure 5 It is a structural block diagram of a computer device in an embodiment. Detailed Embodiments
[0017] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts belong to the scope of protection of the present invention.
[0018] Figure 1 It is an application environment diagram of a voice noise reduction based on a vibration sensor in an embodiment. Refer to Figure 1 , the voice noise reduction method based on a vibration sensor is applied to a voice noise reduction system based on a vibration sensor. The voice noise reduction system based on a vibration sensor includes a terminal 110 and a server 120. The terminal 110 and the server 120 are connected through a network. The terminal 110 may specifically be a desktop terminal or a mobile terminal, and the mobile terminal may specifically be at least one of a mobile phone, a tablet computer, a laptop computer, etc. The server 120 may be implemented by an independent server or a server cluster composed of multiple servers. The terminal 110 is used to collect signals, and the server 120 is used to process signals.
[0019] Such as Figure 2As shown, in one embodiment, a voice noise reduction method based on a vibration sensor is provided. The method is applied to a vibration sensor circuit, which includes a vibration sensor that is in contact connection with a vibrating object. This method can be applied to both terminals and servers. In this embodiment, it is exemplified by being applied to a terminal. The voice noise reduction method based on the vibration sensor specifically includes the following steps: S1: Collect a mixed signal through a microphone and synchronously collect a reference signal of the vibration sensor; S2: Generate a linear echo estimate through an adaptive filter based on the reference signal and the mixed signal, and calculate an initial residual; S3: Extract the time-frequency features of the initial residual and the sensor reference signal, and input them into a pre-trained deep learning compensation network to obtain a non-linear echo estimate; S4: Adjust the initial residual according to a double-talk state decision function and the non-linear echo estimate to obtain a target residual; S5: Process the target residual through a post-Kalman filter noise reduction module and output a pure proximal voice.
[0020] As described in step S1 above, collect a mixed signal through a microphone and synchronously collect a reference signal of the vibration sensor. Among them, the voice noise reduction method based on the vibration sensor is applicable to scenarios such as smart homes, in-vehicle systems, mobile terminals, AI conversations, etc. The vibrating object can be an object with vibration noise such as a car or an electric bicycle. Collect a mixed signal through a microphone (that is, a signal containing target voice and background noise). The method of collecting the mixed signal by the microphone is the same as the prior art and will not be elaborated here. At the same time, use the vibration sensor to synchronously collect the reference signal. The reference signal mainly reflects the characteristics of vibration noise, which can help better distinguish the voice signal from the noise part in subsequent processing. It should be noted that it is necessary to ensure that the collected data stream maintains good time synchronization to avoid the impact of delay on subsequent processing.
[0021] As described in step S2 above, generate a linear echo estimate through an adaptive filter based on the reference signal and the mixed signal, and calculate an initial residual. Based on the collected reference signal and mixed signal, use an adaptive filter (such as the least mean square error (LMS) algorithm or the recursive least squares (RLS) method) to generate a linear echo estimate. This linear echo estimate is actually a prediction of the noise in the voice signal based on the reference signal. Then, calculate the initial residual. The calculation method is to subtract the reference signal from the mixed signal. Since the reference signal is mainly a noise signal, the initial residual can be regarded as a preliminarily noise-reduced voice signal.
[0022] As described in step S3 above, extract the time-frequency features of the initial residual and the sensor reference signal, and input them into a pre-trained deep learning compensation network to obtain a non-linear echo estimate; extract the time-frequency features of the initial residual and the vibration sensor reference signal, which can be achieved by short-time Fourier transform (STFT) or wavelet transform. In this way, the changing characteristics of the signal in time and frequency can be captured. Input the extracted features into a pre-trained deep learning compensation network (such as convolutional neural network CNN or long short-term memory network LSTM). The goal of this network is to generate a more accurate non-linear echo estimate based on the input features. Among them, the training method of the network is to collect multiple groups of pure far-end signals and record the echoes played through real hardware. Use a linear adaptive filter to generate a linear echo estimate and calculate the initial residual. Construct input feature pairs, the form of the feature pairs is frequency and time, and the label of the feature pairs is the corresponding real non-linear echo. The loss function can be any of the following loss functions, time-frequency domain loss functions: mean squared error (MSE) and mean absolute error (MAE), time domain loss functions: SI-SDR (scale-invariant signal-to-noise ratio) or waveform MSE, and joint loss: combine time-frequency and time domain losses. During the training process, training optimization can be carried out in the following ways, data augmentation: adding noise, simulating different non-linear distortion levels, and teacher forcing: using the real residual instead of the network output during training. Finally, a pre-trained deep learning compensation network is obtained through training.
[0023] As described in step S4 above, the initial residual is adjusted according to the double-talk state decision function and the nonlinear echo estimation to obtain the target residual, wherein double-talk means that the voice signals of the two parties in the conversation (such as the user and the AI) appear at the same time in time, that is, both parties speak at the same time. Scenario examples: A. When the user asks a question, the AI fails to stop responding in time due to misjudgment or delay, resulting in overlapping voices of both parties. B. Common interruptions and snatching phenomena in natural conversations (such as users interrupting AI to supplement information). Non-double-talk means that only one party (user or AI) is speaking, and the other party remains silent or listening. Scenario example: The user asks a one-way question, and the AI waits for the user to finish speaking before responding. When the AI broadcasts information, the user does not interrupt. Double-talk detection (DTD): Determine whether double-talk occurs through algorithms such as energy detection and spectrum analysis. In addition, cross-correlation delay, spectrum similarity and deep learning features are used as the basis for deciding whether it is double talk. Cross-correlation delay is to calculate the short-time energy ratio of the reference signal and the initial difference. When the ratio exceeds a preset threshold (for example, 0.8), it is determined to be a double talk state, otherwise it is determined to be a non-double talk state. The deep learning feature converts the reference signal and the initial difference into a spectrum signal, and then inputs the spectrum signal into the deep learning network to output the double talk probability, thereby judging whether it is double talk. When the probability is greater than 0.5, it is determined to be a double talk state, otherwise it is determined to be a non-double talk state. The double talk state decision function is specifically that when double talk occurs, the initial residual needs to be processed based on the nonlinear echo, and when double talk does not occur, the initial residual is directly used as the target residual.
[0024] As described in step S5 above, the dynamically adjusted target residual will be input into the post-Kalman filter module for noise reduction processing. The Kalman filter can effectively reduce the residual noise and output a pure near-end speech signal. The speech noise reduction method based on the vibration sensor effectively reduces the noise in the speech signal through multiple processing steps, thereby improving the accuracy and clarity of speech recognition. Combining the advantages of adaptive filtering and deep learning, this method has good adaptability to various environmental characteristics and improves the responsiveness and robustness to dynamic environmental changes.
[0025] In one embodiment, the step S2 of generating a linear echo estimate through an adaptive filter based on the reference signal and calculating an initial residual comprises: S201: Iteratively adjust the coefficient of the reference signal using a preset algorithm ;in, represents the coefficient of the k-th impulse response; S202: According to the formula Compute a linear echo estimate; where, represents the linear echo estimate, represents the length of the impulse response in the reference signal, Denote the (n - k)-th impulse response in the reference signal; S203: According to the formula Minimize the error signal to obtain the initial residual; where, Denote the mixed signal, Denote the initial residual.
[0026] As described in the above steps S201 - S203, iteratively adjust the coefficients of the reference signal through a preset algorithm , where the preset algorithm is one of the least mean square algorithm, normalized least mean square algorithm, and recursive least squares method. Then calculate its linear echo estimation. The filter length L is determined according to the echo path delay (usually 10 - 500 ms). In this application, the linear echo estimation is a noise signal. After calculating the linear echo estimation here, subtract the linear echo estimation from the mixed signal to obtain the initial residual. Since this noise is a vibration signal and generally has a vibration period, it is the linear echo estimation.
[0027] In one embodiment, the preset algorithm is one of the least mean square algorithm, normalization algorithm, and recursive least squares method. The least mean square algorithm aims to minimize the mean square error between the output signal and the reference signal. This algorithm iteratively updates the coefficients of the filter to adjust the output to achieve the optimal effect. The normalized least mean square algorithm is an improved version of the least mean square algorithm. Its adaptive step is normalized with the energy of the input signal, thereby improving the convergence speed and robustness, especially when the amplitude of the input signal changes greatly. The recursive least squares method directly calculates the filter coefficients that minimize the mean square error by weighting all past input signals and errors. The recursive least squares method can quickly adapt to the changes of the signal. Specifically, the filter coefficients , the filter length L is determined according to the echo path delay (usually 10 - 500 ms). Calculate the linear echo estimation according to the formula and continuously update the filter coefficients. It should be noted that the current filter coefficients are related to the previous filter coefficients, and finally the initial residual is calculated.
[0028] In one embodiment, the step S3 of extracting the time - frequency features of the initial residual and the sensor reference signal and inputting them into a pre - trained deep - learning compensation network to obtain the non - linear echo estimation includes: S301: Respectively transform the initial residual and the sensor reference signal into frequency signals through short - time Fourier transform, and correspondingly obtain the initial residual frequency signal and the sensor reference frequency signal; S302: Input the initial residual frequency signal and the sensor reference frequency signal into the pre - trained deep - learning compensation network to obtain the non - linear echo time - frequency mask; S303: Apply the non-linear echo time-frequency mask to the initial residual frequency signal, and perform inverse short-time Fourier transform to obtain the non-linear echo estimate.
[0029] As described in the above steps S301 - S303, the pre-trained deep learning compensation network usually consists of multiple convolutional layers, pooling layers, and fully connected layers, aiming to learn how to extract features from the input time-frequency signal to generate an estimate of the non-linear echo. That is, first, the initial residual and the sensor reference signal are respectively transformed into frequency signals through short-time Fourier transform, corresponding to obtaining the initial residual frequency signal and the sensor reference frequency signal, so as to be used as the input of the subsequent deep learning model. Taking the initial residual frequency signal and the sensor reference frequency signal as inputs, usually, the two need to be combined into a composite feature graph. For example, they can be stacked, and the combined feature map is input into the deep learning compensation network for processing to obtain the non-linear echo time-frequency mask. Apply the non-linear echo time-frequency mask to the initial residual frequency signal, that is, multiply the non-linear echo time-frequency mask by the initial residual frequency signal, and perform inverse short-time Fourier transform to obtain the non-linear echo estimate.
[0030] In one embodiment, step S4 of adjusting the initial residual according to the double-talk state decision function and the non-linear echo estimate to obtain the target residual includes: S401: Calculate the short-time energy of the reference signal and the initial residual respectively to obtain the reference signal short-time energy and the initial residual short-time energy; S402: Calculate the ratio of the initial residual short-time energy to the reference signal short-time energy, and determine whether the ratio is greater than a preset threshold; S403: If the ratio is greater than the threshold, it is recorded as the double-talk state, and then adjust the residual output according to the formula where, is the target residual, is the initial residual, is the non-linear echo estimate; S404: If the ratio is less than or equal to the threshold, it is recorded as the non-double-talk state, and then adjust the residual output according to the formula Adjust the residual output.
[0031] As described in the above steps S401 - S404, the double - talk state is detected using energy detection. Specifically, the calculation formula for short - time energy usually sums the squares of the signals within each time window, calculates the ratio between the short - time energy of the initial residual and the short - time energy of the reference signal, and compares the calculated ratio with a preset threshold. If it is greater than the set threshold, it is recorded as the double - talk state (i.e., there are multiple people speaking or there is no obvious difference between the main speech and background noise). Otherwise, it is recorded as the non - double - talk state. This can effectively and dynamically determine the double - talk state and adjust the residual output according to different states. This mechanism can improve the adaptability of the system to complex speech performances in actual application scenarios, ensuring that the system can accurately extract the target speech signal even in multi - person conversations or noisy environments, thereby improving the accuracy of speech recognition.
[0032] In one embodiment, step S5 of processing the target residual through the post - Kalman filtering noise reduction module to output pure proximal speech includes: S501: Perform a short - time Fourier transform on the target residual to obtain a residual time - frequency signal; S502: Model each frequency point of the residual time - frequency signal to obtain a state vector for each frequency point , where represents the state vector of the t - th frequency point, represents the pure speech amplitude spectrum coefficient of the t - th frequency point, represents the noise amplitude spectrum coefficient of the t - th frequency point; S503: Assume that the speech and noise follow a first - order autoregressive model ; where , are the first - order autoregressive coefficients, updated iteratively through Kalman filtering, respectively represent the process noise; S504: Perform Kalman filtering iteration on the state vector of each frequency point to obtain a target state vector; S505: Extract pure proximal speech from the target state vector.
[0033] As described in the above steps S501 - S503, the target residual is processed through the post - Kalman filtering noise reduction module. Specifically, a short - time Fourier transform is performed on the target residual to obtain a residual time - frequency signal, and each frequency point is modeled. The state vector includes information such as the amplitude, phase, and their change rates of the current frequency point. Initialize the state vector of each frequency point, which can generally be set to zero or initialized by some method (e.g., using the data of the first frame), and set the initial state and covariance of the system. The initial covariance matrix is usually set to a relatively large value to represent uncertainty, such as , where is a large number, i.e., a number greater than a preset value. is the identity matrix. The process noise covariance is usually selected according to the characteristics of the system noise, and a suitable value can be obtained through experiments. The observation noise covariance R is estimated based on the noise characteristics of the observed signal. Assume that the speech and background noise follow a first-order autoregressive (AR) model. The autoregressive model indicates that the current state can be represented by a linear combination of the previous state. Perform a prediction step on the state vector at each frequency point, update the state prediction according to the autoregressive model, update the state vector using the observed time-frequency signal, adjust the state estimate by calculating the Kalman gain, and update the state estimate. Iteratively process each frequency point and continuously update the state vector to obtain more accurate clean speech features. Extract the clean speech signal information from the processed target state vector, which can be reconstructed based on the amplitude and phase information. Restore the obtained target state vector signal to the time domain and obtain the final clean proximal speech through the inverse Fourier transform.
[0034] Refer to Figure 3 , the present invention also provides a vibration sensor circuit for collecting a reference signal, including: a power supply filtering module, a vibration sensor module, a positive feedback amplification module, a two-stage filtering module, and a chip; The power supply filtering module is connected to the vibration sensor module. The power supply filtering module is used to filter and supply power to the vibration sensor. The vibration sensor module is in contact connection with the vibrating object to obtain a vibration signal; The vibration sensor module is connected to the positive feedback amplification module. The positive feedback amplification module is used to amplify the vibration signal; The positive feedback amplification module is connected to the two-stage filtering module. The two-stage filtering module is used to filter the amplified vibration signal to obtain a reference signal; The two-stage filtering module is connected to the chip and is used to transmit the reference signal to the chip for processing.
[0035] Among them, the power supply filtering module includes capacitor C1, capacitor C2, capacitor C3, resistor R1, and resistor R2. The power supply is respectively connected to the first end of capacitor C1, and the first ends of resistor R2 and resistor R1. The second end of capacitor C1 is respectively grounded and connected to the first end of capacitor C2. The second end of capacitor C2 is respectively connected to the second end of resistor R2 and one end of the vibration sensor module. The second end of resistor R1 is respectively connected to one end of the positive feedback amplification module and the first end of capacitor C3. The second end of capacitor C3 is grounded.
[0036] The vibration sensor module includes a vibration sensor, resistor R6, vibration sensor SW2, and resistor R3. Among them, the first end of resistor R6 is connected to the second end of resistor R2, the first end of vibration sensor SW2 is connected to one end of resistor R6, the second end of vibration sensor SW2 is respectively connected to the first end of resistor R3 and one end of the positive feedback amplification module, and the second end of resistor R3 is grounded.
[0037] The positive feedback amplification module includes capacitor C5, capacitor C7, capacitor C8, capacitor C9, triode Q1, resistor R5, resistor R4, and resistor R9. Among them, the first end of capacitor C5 is connected to the second end of vibration sensor SW2, the second end of capacitor C5 is respectively connected to the first end of capacitor C7, the base of triode Q1, and the first end of resistor R5. The second end of resistor R5 is respectively connected to the second end of resistor R4, the collector of triode Q1, the first end of capacitor C9, and one end of the two-stage filtering module. The emitter of triode Q1 is respectively connected to the first end of resistor R9 and the first end of capacitor C8. The second ends of resistor R9, capacitor C8, and capacitor C9 are all grounded.
[0038] The two-stage filtering module includes capacitor C4, resistor R7, capacitor C10, resistor R8, capacitor C11, and capacitor C6. Among them, the first end of capacitor C4 is connected to the second end of resistor R5, the second end of capacitor C4 is connected to the first end of resistor R7, the second end of resistor R7 is respectively connected to the first end of resistor R8 and the first end of capacitor C10. The second end of capacitor C10 is grounded. The second end of resistor R8 is connected to the first end of capacitor C11 and the first end of capacitor C6. The second end of capacitor C11 is grounded. The second end of capacitor C6 is connected to the chip.
[0039] Refer to Figure 4 , the present invention also provides a voice noise reduction device based on a vibration sensor. The device is applied to a vibration sensor circuit. The vibration sensor circuit includes a vibration sensor, and the vibration sensor is in contact connection with a vibrating object. The device includes: An acquisition module 10, configured to acquire a mixed signal through a microphone and synchronously acquire a reference signal of the vibration sensor; A generation module 20, configured to generate a linear echo estimate through an adaptive filter based on the reference signal and the mixed signal, and calculate an initial residual; An extraction module 30, configured to extract the time-frequency features of the initial residual and the sensor reference signal, and input them into a pre-trained deep learning compensation network to obtain a non-linear echo estimate; An adjustment module 40, configured to adjust the initial residual according to a double-talk state decision function and the non-linear echo estimate to obtain a target residual; The output module 50 is configured to process the target residual through a post Kalman filter noise reduction module and output a pure proximal voice.
[0040] In one embodiment, the generation module 20 includes: A coefficient adjustment sub-module for iteratively adjusting the coefficients of the reference signal through a preset algorithm ; where represents the coefficient of the k-th impulse response; A linear echo estimation calculation sub-module for calculating the linear echo estimation according to the formula ; where represents the linear echo estimation, represents the length of the impulse response in the reference signal, represents the (n - k)-th impulse response in the reference signal; An initial residual calculation sub-module for minimizing the error signal according to the formula to obtain the initial residual; where represents the mixed signal, represents the initial residual.
[0041] In one embodiment, the preset algorithm is one of the least mean square algorithm, the normalization algorithm, and the recursive least squares method.
[0042] In one embodiment, the extraction module 30 includes: A frequency signal conversion sub-module for converting the initial residual and the sensor reference signal into frequency signals respectively through short-time Fourier transform, and correspondingly obtaining the initial residual frequency signal and the sensor reference frequency signal; A frequency signal input sub-module for inputting the initial residual frequency signal and the sensor reference frequency signal into a pre-trained deep learning compensation network to obtain a non-linear echo time-frequency mask; A short-time Fourier transform sub-module for applying the non-linear echo time-frequency mask to the initial residual frequency signal and performing inverse short-time Fourier transform to obtain the non-linear echo estimation.
[0043] In one embodiment, the adjustment module 40 includes: A short-time energy calculation sub-module for calculating the short-time energy of the reference signal and the initial residual respectively to obtain the reference signal short-time energy and the initial residual short-time energy; A ratio calculation sub-module for calculating the ratio of the initial residual short-time energy to the reference signal short-time energy and determining whether the ratio is greater than a preset threshold; A double-talk state marking sub-module for, if the ratio is greater than the threshold, marking it as the double-talk state, and then according to the formula Adjust the residual output; where, is the target residual, is the initial residual, is the non-linear echo estimation; The non-dual-talking state marking sub-module is used to record the non-dual-talking state if the ratio is less than or equal to the threshold, and then adjust the residual output according to the formula Adjust the residual output.
[0044] In one embodiment, the output module 50 includes: The transformation sub-module is used to perform short-time Fourier transform on the target residual to obtain the residual time-frequency signal; The modeling sub-module is used to model each frequency point of the residual time-frequency signal to obtain the state vector of each frequency point , where, represents the state vector of the t-th frequency point, represents the clean speech amplitude spectrum coefficient of the t-th frequency point, represents the noise amplitude spectrum coefficient of the t-th frequency point; The hypothesis sub-module is used to assume that the speech and noise follow a first-order autoregressive model ; where , are the first-order autoregressive coefficients, which are updated iteratively through Kalman filtering, respectively represent the process noise; The iteration sub-module is used to perform Kalman filtering iteration on the state vector of each frequency point to obtain the target state vector; The extraction sub-module is used to extract the clean proximal speech from the target state vector.
[0045] Figure 5 Fig. shows the internal structure diagram of a computer device in one embodiment. The computer device can specifically be a terminal or a server. As Figure 5 shown, the computer device includes a processor, a memory, and a network interface connected through a system bus. Among them, the memory includes a non-volatile storage medium and an internal memory. The non-volatile storage medium of the computer device stores an operating system and can also store a computer program. When the computer program is executed by the processor, the processor can be implemented. The internal memory can also store a computer program. When the computer program is executed by the processor, the processor can execute the voice noise reduction method based on the vibration sensor. Those skilled in the art can understand that Figure 5 the structure shown in
[0046] In one embodiment, a computer device is provided, including a memory and a processor. The memory stores a computer program. When the computer program is executed by the processor, the processor performs the following steps: Collect a mixed signal through a microphone and synchronously collect a reference signal of a vibration sensor; Generate a linear echo estimate through an adaptive filter based on the reference signal and the mixed signal, and calculate an initial residual; Extract the time-frequency features of the initial residual and the sensor reference signal, and input them into a pre-trained deep learning compensation network to obtain a non-linear echo estimate; Adjust the initial residual according to a double-talk state decision function and the non-linear echo estimate to obtain a target residual; Process the target residual through a post-Kalman filtering noise reduction module to output a pure proximal speech.
[0047] The speech noise reduction method based on a vibration sensor effectively reduces the noise in a speech signal through multiple processing steps, improving the accuracy and clarity of speech recognition. By combining the advantages of adaptive filtering and deep learning, the method has good adaptability to various environmental characteristics, improving the response ability and robustness to dynamic environmental changes.
[0048] In one embodiment, a computer-readable storage medium is provided, storing a computer program. When the computer program is executed by a processor, the processor performs the following steps: Collect a mixed signal through a microphone and synchronously collect a reference signal of a vibration sensor; Generate a linear echo estimate through an adaptive filter based on the reference signal and the mixed signal, and calculate an initial residual; Extract the time-frequency features of the initial residual and the sensor reference signal, and input them into a pre-trained deep learning compensation network to obtain a non-linear echo estimate; Adjust the initial residual according to a double-talk state decision function and the non-linear echo estimate to obtain a target residual; Process the target residual through a post-Kalman filtering noise reduction module to output a pure proximal speech.
[0049] The speech noise reduction method based on a vibration sensor effectively reduces the noise in a speech signal through multiple processing steps, improving the accuracy and clarity of speech recognition. By combining the advantages of adaptive filtering and deep learning, the method has good adaptability to various environmental characteristics, improving the response ability and robustness to dynamic environmental changes.
[0050] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The program can be stored in a non-volatile computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the embodiments provided in the present application can include non-volatile and / or volatile memories. Non-volatile memories can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memories can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.
[0051] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.
[0052] The above-described embodiments merely represent several implementation manners of the present application. The description is relatively specific and detailed, but it should not be construed as a limitation on the patent scope of the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent of the present application shall be subject to the appended claims.
Claims
1. A speech noise reduction method based on a vibration sensor, characterized in that: The method is applied to a vibration sensor circuit, the vibration sensor circuit includes a vibration sensor, the vibration sensor is in contact with a vibrating object, and the method includes: The mixed signal is collected through the microphone, and the reference signal of the vibration sensor is collected synchronously; Generate a linear echo estimate based on the reference signal and the mixed signal through an adaptive filter, and calculate an initial residual; Extracting time-frequency features of the initial residual and the sensor reference signal, and inputting them into a pre-trained deep learning compensation network to obtain a nonlinear echo estimate; Adjusting the initial residual according to the double-talk state decision function and the nonlinear echo estimation to obtain a target residual; The target residual is processed by a post-Kalman filter noise reduction module to output pure near-end speech.
2. The method for speech noise reduction based on a vibration sensor according to claim 1, characterized in that: The step of generating a linear echo estimate through an adaptive filter based on the reference signal and calculating an initial residual comprises: The reference signal is iteratively adjusted using a preset algorithm. ;in, represents the coefficient of the k-th impulse response; According to the formula Compute a linear echo estimate; where, represents the linear echo estimate, represents the length of the impulse response in the reference signal, represents the nkth impulse response in the reference signal; According to the formula Minimize the error signal to get the initial residual; where, represents a mixed signal, represents the initial residual.
3. The method for speech noise reduction based on a vibration sensor according to claim 2, characterized in that: The preset algorithm is one of a least mean square algorithm, a normalization algorithm and a recursive least squares method.
4. The method for speech noise reduction based on a vibration sensor according to claim 1, characterized in that: The step of extracting the time-frequency features of the initial residual and the sensor reference signal and inputting them into a pre-trained deep learning compensation network to obtain a nonlinear echo estimation includes: The initial residual and the sensor reference signal are respectively converted into frequency signals by short-time Fourier transform, and an initial residual frequency signal and a sensor reference frequency signal are correspondingly obtained; Inputting the initial residual frequency signal and the sensor reference frequency signal into a pre-trained deep learning compensation network to obtain a nonlinear echo time-frequency mask; The nonlinear echo time-frequency mask is applied to the initial residual frequency signal, and an inverse short-time Fourier transform is performed to obtain a nonlinear echo estimate.
5. The method for speech noise reduction based on a vibration sensor according to claim 1, characterized in that: The step of adjusting the initial residual according to the double-talk state decision function and the nonlinear echo estimation to obtain a target residual comprises: Calculating the short-time energy of the reference signal and the initial residual to obtain the reference signal short-time energy and the initial residual short-time energy respectively; Calculating a ratio of the initial residual short-time energy to the reference signal short-time energy, and determining whether the ratio is greater than a preset threshold; If the ratio is greater than the threshold, it is recorded as a double talk state, then according to the formula Adjust the residual output; where, is the target residual, is the initial residual, is the nonlinear echo estimation; If the ratio is less than or equal to the threshold, it is recorded as a non-dual talk state, then according to the formula Adjust the residual output.
6. The method for speech noise reduction based on a vibration sensor according to claim 1, characterized in that: The step of processing the target residual by a post-Kalman filter noise reduction module to output a clean near-end speech comprises: Performing short-time Fourier transform on the target residual to obtain a residual time-frequency signal; Model each frequency point of the residual time-frequency signal to obtain the state vector of each frequency point ,in, represents the state vector of the tth frequency point, represents the pure speech amplitude spectrum coefficient of the tth frequency point, Represents the noise amplitude spectrum coefficient of the tth frequency point; Assume that speech and noise obey the first-order autoregressive model ;in , is the first-order autoregressive coefficient, which is updated iteratively through the Karl filter. represent process noise respectively; Perform Kalman filter iteration on the state vector of each frequency point to obtain the target state vector; A clean near-end speech is extracted from the target state vector.
7. A vibration sensor circuit, used for collecting a reference signal in the speech noise reduction method based on a vibration sensor according to any one of claims 1 to 6, characterized in that: include: Power supply filter module, vibration sensor module, positive feedback amplifier module, two-stage filter module and chip; The power supply filter module is connected to the vibration sensor module, and the power supply filter module is used to filter and supply power to the vibration sensor. The vibration sensor module is in contact with the vibrating object to obtain a vibration signal. The vibration sensor module is connected to the positive feedback amplification module, and the positive feedback amplification module is used to amplify the vibration signal; The positive feedback amplification module is connected to the two-stage filtering module, and the two-stage filtering module is used to filter the amplified vibration signal to obtain a reference signal; The two-stage filtering module is connected to the chip and is used to transmit the reference signal to the chip for processing.
8. A speech noise reduction device based on a vibration sensor, characterized in that: The device is applied to a vibration sensor circuit, the vibration sensor circuit includes a vibration sensor, the vibration sensor is in contact with a vibrating object, and the device includes: An acquisition module, used for acquiring the mixed signal through a microphone and synchronously acquiring a reference signal from a vibration sensor; A generating module, configured to generate a linear echo estimate through an adaptive filter based on the reference signal and the mixed signal, and calculate an initial residual; An extraction module, used to extract the time-frequency features of the initial residual and the sensor reference signal, and input them into a pre-trained deep learning compensation network to obtain a nonlinear echo estimate; An adjustment module, configured to adjust the initial residual according to the double-talk state decision function and the nonlinear echo estimation to obtain a target residual; The output module is used to process the target residual through a post-Kalman filter noise reduction module to output pure near-end speech.
9. A computer-readable storage medium, characterized in that: A computer program is stored, and when the computer program is executed by a processor, the processor executes the steps of the speech noise reduction method based on a vibration sensor as claimed in any one of claims 1 to 6.
10. A computer device, characterized in that: The device comprises a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, the processor executes the steps of the speech noise reduction method based on a vibration sensor as claimed in any one of claims 1 to 6.
Citation Information
Patent Citations
Use of vibration sensor in acoustic echo cancellation
CN104243732A
Echo elimination method and device, electronic equipment and computer readable medium
CN115083431A
Echo cancellation method and system based on neural network double-talk detection
CN115457928A
Nonlinear echo suppression method and device, electronic equipment and storage medium
CN118486317A
Use of vibration sensor in acoustic echo cancellation
US20140363008A1
Cited By
Digital audio noise reduction method based on time-frequency mask separation
CN120636429A
Digital audio noise reduction method based on time-frequency mask separation
CN120636429B