Vibration Sensor-Based Speech Noise Reduction Method, Device, Equipment, and Medium

Through a multi-step processing method based on vibration sensors, combined with adaptive filtering and deep learning, the problem of insufficient accuracy and stability of traditional speech recognition systems in vibrating noise environments is solved, and higher speech recognition accuracy and environmental adaptability are achieved.

CN120148539BActive Publication Date: 2025-08-01SHENZHEN WAYTRONIC ELECTRONICS CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510622207.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-15
Publication Date
2025-08-01
Estimated Expiration
2045-05-15

AI Technical Summary

Technical Problem

The accuracy and stability of traditional speech recognition systems are affected in the face of vibration noise environments, and the prior art lacks effective vibration noise reduction schemes.

Method used

The speech noise reduction method based on vibration sensor is adopted, mixed signals and vibration sensor reference signals are collected through the microphone, and linear echo estimation is generated using an adaptive filter. Combined with deep learning compensation network and Kalman filter noise reduction, multi-step processing of the initial residuals is realized to output pure near-end speech.

Benefits of technology

Effectively reduce noise in speech signals, improve the accuracy and clarity of speech recognition, and enhance adaptability and robustness to dynamic environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120148539B_ABST
    Figure CN120148539B_ABST
Patent Text Reader

Abstract

This application relates to the field of artificial intelligence technology, and discloses a voice noise reduction method, device, equipment and medium based on a vibration sensor. The method includes: synchronously collecting a mixed signal and a reference signal of the vibration sensor, generating a linear echo estimate, and calculating an initial residual; extracting the time-frequency features of the initial residual and the sensor reference signal to obtain a non-linear echo estimate; adjusting the initial residual according to a double-talk state decision function and the non-linear echo estimate to obtain a target residual; processing the target residual through a post-Kalman filter noise reduction module to output a pure proximal voice. The beneficial effects of the present invention are as follows: effectively reducing the noise in the voice signal, improving the accuracy and clarity of voice recognition, combining the advantages of adaptive filtering and deep learning, having good adaptability to various environmental characteristics, and improving the response ability and robustness to dynamic environmental changes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of speech noise reduction, and particularly to a speech noise reduction method, device, equipment and medium based on a vibration sensor. Background Art

[0002] With the development of technology, speech recognition technology has been widely used in many fields, such as intelligent assistants, autonomous driving, telephone customer service, etc. However, in the face of vibration noise in complex environments, which usually originates from mechanical equipment, transportation tools, etc. and has strong low-frequency characteristics, the accuracy and stability of traditional speech recognition systems are often severely affected. Especially in noisy scenarios such as industrial environments, busy cities, and outdoors, background noise and vibration can cause distortion of speech signals, thereby reducing the performance of the recognition system. In the prior art, although there are various noise reduction solutions, there is no technical solution for vibration noise. Therefore, how to improve the anti-interference ability of speech recognition in these noise environments has become an important topic in the current research of speech recognition technology. Summary of the Invention

[0003] Based on this, in view of the existing speech noise reduction problem based on a vibration sensor, a speech noise reduction method, device, equipment and medium based on a vibration sensor are proposed.

[0004] A speech noise reduction method based on a vibration sensor, the method is applied to a vibration sensor circuit, the vibration sensor circuit includes a vibration sensor, and the vibration sensor is in contact connection with a vibrating object. The method includes:

[0005] Collect a mixed signal through a microphone and synchronously collect a reference signal of the vibration sensor;

[0006] Generate a linear echo estimate through an adaptive filter based on the reference signal and the mixed signal, and calculate an initial residual;

[0007] Extract the time-frequency features of the initial residual and the sensor reference signal, and input them into a pre-trained deep learning compensation network to obtain a non-linear echo estimate;

[0008] Adjust the initial residual according to a double-talk state decision function and the non-linear echo estimate to obtain a target residual;

[0009] Process the target residual through a post-Kalman filter noise reduction module and output a pure proximal speech.

[0010] Further, the step of generating a linear echo estimate through an adaptive filter based on the reference signal and calculating an initial residual includes:

[0011] Iteratively adjust the coefficients of the reference signal through a preset algorithm ; where represents the coefficient of the k-th impulse response;

[0012] Calculate the linear echo estimation according to the formula ; where represents the linear echo estimation, represents the length of the impulse response in the reference signal, represents the (n - k)-th impulse response in the reference signal;

[0013] Minimize the error signal according to the formula to obtain the initial residual; where represents the mixed signal, represents the initial residual.

[0014] Further, the preset algorithm is one of the least mean square algorithm, the normalization algorithm, and the recursive least squares method.

[0015] Further, the step of extracting the time-frequency features of the initial residual and the sensor reference signal and inputting them into a pre-trained deep learning compensation network to obtain the non-linear echo estimation includes:

[0016] Convert the initial residual and the sensor reference signal into frequency signals through short-time Fourier transform respectively, and correspondingly obtain the initial residual frequency signal and the sensor reference frequency signal;

[0017] Input the initial residual frequency signal and the sensor reference frequency signal into the pre-trained deep learning compensation network to obtain the non-linear echo time-frequency mask;

[0018] Apply the non-linear echo time-frequency mask to the initial residual frequency signal and perform inverse short-time Fourier transform to obtain the non-linear echo estimation.

[0019] Further, the step of adjusting the initial residual according to the double-talk state decision function and the non-linear echo estimation to obtain the target residual includes:

[0020] Calculate the short-time energy of the reference signal and the initial residual respectively to obtain the reference signal short-time energy and the initial residual short-time energy;

[0021] Calculate the ratio of the initial residual short-time energy to the reference signal short-time energy and determine whether the ratio is greater than a preset threshold;

[0022] If the ratio is greater than the threshold, it is recorded as the double-talk state, and then adjust the residual output according to the formula ; where is the target residual, is the initial residual, is the non - linear echo estimation;

[0023] If the ratio is less than or equal to the threshold, it is recorded as the non - double - talk state, and then according to the formula adjust the residual output.

[0024] Further, the step of processing the target residual by the post - Kalman filtering and noise reduction module to output the pure proximal speech includes:

[0025] Perform a short - time Fourier transform on the target residual to obtain a residual time - frequency signal;

[0026] Model each frequency point of the residual time - frequency signal to obtain the state vector of each frequency point , where represents the state vector of the t - th frequency point, represents the pure speech amplitude spectrum coefficient of the t - th frequency point, represents the noise amplitude spectrum coefficient of the t - th frequency point;

[0027] Assume that the speech and noise follow a first - order autoregressive model ; where , are the first - order autoregressive coefficients, which are updated iteratively through Kalman filtering, respectively represent the process noise;

[0028] Perform Kalman filtering iteration on the state vector of each frequency point to obtain the target state vector;

[0029] Extract the pure proximal speech from the target state vector.

[0030] The present invention also provides a vibration sensor circuit for collecting the above - mentioned reference signal, including: a power supply filtering module, a vibration sensor module, a positive - feedback amplification module, a two - stage filtering module, and a chip;

[0031] The power supply filtering module is connected to the vibration sensor module. The power supply filtering module is used to filter and supply power to the vibration sensor. The vibration sensor module is in contact connection with the vibrating object to obtain a vibration signal;

[0032] The vibration sensor module is connected to the positive - feedback amplification module. The positive - feedback amplification module is used to amplify the vibration signal;

[0033] The positive - feedback amplification module is connected to the two - stage filtering module. The two - stage filtering module is used to filter the amplified vibration signal to obtain a reference signal;

[0034] The two-stage filtering module is connected to the chip and is used to transfer the reference signal to the chip for processing.

[0035] The present invention also provides a voice noise reduction device based on a vibration sensor. The device is applied to a vibration sensor circuit, and the vibration sensor circuit includes a vibration sensor which is in contact connection with a vibrating object. The device includes:

[0036] An acquisition module, configured to acquire a mixed signal through a microphone and synchronously acquire a reference signal of the vibration sensor;

[0037] A generation module, configured to generate a linear echo estimate through an adaptive filter based on the reference signal and the mixed signal, and calculate an initial residual;

[0038] An extraction module, configured to extract time-frequency features of the initial residual and the sensor reference signal and input them into a pre-trained deep learning compensation network to obtain a non-linear echo estimate;

[0039] An adjustment module, configured to adjust the initial residual according to a double-talk state decision function and the non-linear echo estimate to obtain a target residual;

[0040] An output module, configured to process the target residual through a post-Kalman filtering noise reduction module and output a pure near-end voice.

[0041] A computer device includes a memory and a processor. When a computer program stored in the memory is executed by the processor, the processor performs the following steps:

[0042] Acquire a mixed signal through a microphone and synchronously acquire a reference signal of the vibration sensor;

[0043] Generate a linear echo estimate through an adaptive filter based on the reference signal and the mixed signal, and calculate an initial residual;

[0044] Extract time-frequency features of the initial residual and the sensor reference signal and input them into a pre-trained deep learning compensation network to obtain a non-linear echo estimate;

[0045] Adjust the initial residual according to a double-talk state decision function and the non-linear echo estimate to obtain a target residual;

[0046] Process the target residual through a post-Kalman filtering noise reduction module and output a pure near-end voice.

[0047] A computer-readable storage medium stores a computer program. When the computer program is executed by a processor, the processor performs the following steps:

[0048] Collect the mixed signal through a microphone and synchronously collect the reference signal of the vibration sensor;

[0049] Generate a linear echo estimate through an adaptive filter based on the reference signal and the mixed signal, and calculate the initial residual;

[0050] Extract the time-frequency features of the initial residual and the sensor reference signal, and input them into a pre-trained deep learning compensation network to obtain a non-linear echo estimate;

[0051] Adjust the initial residual according to the double-talk state decision function and the non-linear echo estimate to obtain the target residual;

[0052] Process the target residual through a post-Kalman filter noise reduction module and output the clean near-end speech.

[0053] Advantages of the present invention: The speech noise reduction method based on a vibration sensor effectively reduces the noise in the speech signal through multiple processing steps, improving the accuracy and clarity of speech recognition. Combining the advantages of adaptive filtering and deep learning enables the method to have good adaptability to various environmental characteristics, improving the response ability and robustness to dynamic environmental changes. BRIEF DESCRIPTION OF THE DRAWINGS

[0054] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0055] Among them:

[0056] Figure 1 It is an application environment diagram of the speech noise reduction method based on a vibration sensor in an embodiment;

[0057] Figure 2 It is a flowchart of the speech noise reduction method based on a vibration sensor in an embodiment;

[0058] Figure 3 It is a circuit diagram of a vibration sensor in an embodiment;

[0059] Figure 4 It is a structural block diagram of the speech noise reduction device based on a vibration sensor in an embodiment;

[0060] Figure 5 It is a structural block diagram of a computer device in an embodiment. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0061] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts belong to the scope of protection of the present invention.

[0062] Figure 1 It is an application environment diagram of voice noise reduction based on a vibration sensor in an embodiment. Refer to Figure 1 , the voice noise reduction method based on a vibration sensor is applied to a voice noise reduction system based on a vibration sensor. The voice noise reduction system based on a vibration sensor includes a terminal 110 and a server 120. The terminal 110 and the server 120 are connected through a network. The terminal 110 may specifically be a desktop terminal or a mobile terminal, and the mobile terminal may specifically be at least one of a mobile phone, a tablet computer, a laptop computer, etc. The server 120 may be implemented by an independent server or a server cluster composed of multiple servers. The terminal 110 is used to collect signals, and the server 120 is used to process signals.

[0063] As Figure 2 shown, in an embodiment, a voice noise reduction method based on a vibration sensor is provided. The method is applied to a vibration sensor circuit, and the vibration sensor circuit includes a vibration sensor, and the vibration sensor is in contact connection with a vibrating object. This method can be applied to both a terminal and a server. This embodiment takes the application to a terminal as an example for illustration. The voice noise reduction method based on a vibration sensor specifically includes the following steps:

[0064] S1: Collect a mixed signal through a microphone and synchronously collect a reference signal of the vibration sensor;

[0065] S2: Generate a linear echo estimate through an adaptive filter based on the reference signal and the mixed signal, and calculate an initial residual;

[0066] S3: Extract the time-frequency features of the initial residual and the sensor reference signal, and input them into a pre-trained deep learning compensation network to obtain a non-linear echo estimate;

[0067] S4: Adjust the initial residual according to a double-talk state decision function and the non-linear echo estimate to obtain a target residual;

[0068] S5: Process the target residual through a post-Kalman filtering noise reduction module and output a pure proximal voice.

[0069] As described in step S1 above, a mixed signal is collected through a microphone, and a reference signal of a vibration sensor is collected synchronously. Among them, the voice noise reduction method based on the vibration sensor is applicable to scenarios such as smart homes, in-vehicle systems, mobile terminals, AI conversations, etc. The vibrating object can be an object with vibration noise such as a car or an electric bicycle. The mixed signal (i.e., the signal containing the target voice and background noise) is collected through the microphone. The method of collecting the mixed signal by the microphone is the same as that in the prior art and will not be elaborated here. At the same time, the vibration sensor is used to collect the reference signal synchronously. The reference signal mainly reflects the characteristics of the vibration noise, which can help better distinguish the voice signal from the noise part in subsequent processing. It should be noted that it is necessary to ensure that the collected data stream maintains good time synchronization to avoid the impact of delay on subsequent processing.

[0070] As described in step S2 above, a linear echo estimate is generated through an adaptive filter based on the reference signal and the mixed signal, and an initial residual is calculated. Based on the collected reference signal and mixed signal, an adaptive filter (such as the least mean square error (LMS) algorithm or the recursive least squares (RLS) method) is used to generate a linear echo estimate. This linear echo estimate is actually a prediction of the noise in the voice signal based on the reference signal. Then, the initial residual is calculated. The calculation method is to subtract the reference signal from the mixed signal. Since the reference signal is mainly a noise signal, the initial residual can be regarded as a preliminarily noise-reduced voice signal.

[0071] As described in step S3 above, extract the time-frequency features of the initial residual and the sensor reference signal, and input them into a pre-trained deep learning compensation network to obtain a non-linear echo estimate; extract the time-frequency features of the initial residual and the vibration sensor reference signal, which can be achieved by short-time Fourier transform (STFT) or wavelet transform. In this way, the changing features of the signal in time and frequency can be captured. Input the extracted features into a pre-trained deep learning compensation network (such as convolutional neural network CNN or long short-term memory network LSTM). The goal of this network is to generate a more accurate non-linear echo estimate based on the input features. Among them, the network is trained by collecting multiple groups of pure far-end signals and playing and recording the echoes through real hardware. Use a linear adaptive filter to generate a linear echo estimate and calculate the initial residual. Construct input feature pairs, the form of the feature pairs is frequency and time, and the label of the feature pairs is the corresponding true non-linear echo. The loss function can be any of the following loss functions, time-frequency domain loss functions: mean squared error (MSE) and mean absolute error (MAE), time domain loss functions: SI-SDR (scale-invariant signal-to-noise ratio) or waveform MSE, and joint loss: combining time-frequency and time domain losses. During the training process, training optimization can be carried out in the following ways, data augmentation: adding noise, simulating different non-linear distortion levels, and teacher forcing: using the true residual instead of the network output during training. Finally, a pre-trained deep learning compensation network is obtained through training.

[0072] As described in step S4 above, the initial residual is adjusted according to the double-talk state decision function and the non-linear echo estimation to obtain the target residual. Here, double-talk means that the voice signals of both parties in the conversation (such as the user and the AI) appear simultaneously in time, that is, both parties speak at the same time. Scenario examples: A. When the user asks a question, the AI fails to stop responding in time due to misjudgment or delay, resulting in overlapping voices of both parties. B. The common phenomena of interrupting and talking over in natural conversations (such as the user interrupting the AI to supplement information). Non-double-talk means that only one party (the user or the AI) is speaking, and the other party remains silent or in a listening state. Scenario examples: The user asks a one-way question, and the AI waits for the user to finish speaking before responding. When the AI broadcasts information, the user does not interrupt. Double-talk detection (DTD): It is determined whether double-talk occurs through algorithms such as energy detection and spectrum analysis. In addition, cross-correlation delay, spectral similarity, and deep learning features are used as decision bases for double-talk. The cross-correlation delay is to calculate the short-time energy ratio of the reference signal and the initial difference. When the ratio exceeds a preset threshold (such as 0.8), it is determined as the double-talk state, otherwise it is determined as the non-double-talk state. For the deep learning features, the reference signal and the initial difference are converted into spectral signals, and then the spectral signals are input into the deep learning network to output the double-talk probability, so as to determine whether it is double-talk. When the probability is greater than 0.5, it is determined as the double-talk state, otherwise it is determined as the non-double-talk state. The double-talk state decision function is specifically that when double-talk occurs, the initial residual needs to be processed based on the non-linear echo, and when double-talk does not occur, the initial residual is directly used as the target residual.

[0073] As described in step S5 above, the target residual after dynamic adjustment is input into the post-Kalman filtering module for noise reduction processing. The Kalman filter can effectively reduce the remaining noise and output a pure proximal voice signal. The voice noise reduction method based on the vibration sensor effectively reduces the noise in the voice signal through multiple processing steps, improving the accuracy and clarity of speech recognition. Combining the advantages of adaptive filtering and deep learning makes this method have good adaptability to various environmental characteristics, improving the response ability and robustness to dynamic environmental changes.

[0074] In one embodiment, step S2 of generating a linear echo estimation based on the reference signal through an adaptive filter and calculating the initial residual includes:

[0075] S201: Iteratively adjust the coefficients of the reference signal through a preset algorithm ; where represents the coefficient of the k-th impulse response;

[0076] S202: Calculate the linear echo estimation according to the formula ; where represents the linear echo estimation, represents the length of the impulse response in the reference signal, Represents the (n-k)-th impulse response in the reference signal;

[0077] S203: According to the formula Minimize the error signal to obtain the initial residual; where, represents the mixed signal, represents the initial residual.

[0078] As described in the above steps S201-S203, the coefficients of the reference signal are iteratively adjusted through a preset algorithm , where the preset algorithm is one of the least mean square algorithm, the normalized least mean square algorithm, and the recursive least squares method. Then calculate its linear echo estimation. The filter length L is determined according to the echo path delay (usually 10 - 500 ms). In this application, the linear echo estimation is a noise signal. After calculating the linear echo estimation here, subtracting the linear echo estimation from the mixed signal can obtain the initial residual. Since this noise is a vibration signal and generally has a vibration period, it is the linear echo estimation.

[0079] In one embodiment, the preset algorithm is one of the least mean square algorithm, the normalized algorithm, and the recursive least squares method. The least mean square algorithm aims to minimize the mean square error between the output signal and the reference signal. This algorithm iteratively updates the coefficients of the filter to adjust the output to achieve the optimal effect. The normalized least mean square algorithm is an improved version of the least mean square algorithm. Its adaptive step is normalized with the energy of the input signal, thereby improving the convergence speed and robustness, especially when the amplitude of the input signal changes greatly. The recursive least squares method directly calculates the filter coefficients that minimize the mean square error by weighting all past input signals and errors. The recursive least squares method can quickly adapt to the changes of the signal. Specifically, the filter coefficients , the filter length L is determined according to the echo path delay (usually 10 - 500 ms). Calculate the linear echo estimation according to the formula and continuously update the filter coefficients. It should be noted that the current filter coefficients are related to the previous filter coefficients, and finally the initial residual is calculated.

[0080] In one embodiment, step S3 of extracting the time-frequency features of the initial residual and the sensor reference signal and inputting them into a pre-trained deep learning compensation network to obtain the non-linear echo estimation includes:

[0081] S301: Respectively transform the initial residual and the sensor reference signal into frequency signals through short-time Fourier transform to obtain the initial residual frequency signal and the sensor reference frequency signal;

[0082] S302: Input the initial residual frequency signal and the sensor reference frequency signal into a pre-trained deep learning compensation network to obtain a non-linear echo time-frequency mask;

[0083] S303: Apply the non-linear echo time-frequency mask to the initial residual frequency signal and perform an inverse short-time Fourier transform to obtain a non-linear echo estimate.

[0084] As described in the above steps S301 - S303, the pre-trained deep learning compensation network usually consists of multiple convolutional layers, pooling layers, and fully connected layers, aiming to learn how to extract features from the input time-frequency signal to generate an estimate of the non-linear echo. That is, first, the initial residual and the sensor reference signal are respectively transformed into frequency signals through short-time Fourier transform, corresponding to obtaining the initial residual frequency signal and the sensor reference frequency signal, so as to be used as the input of the subsequent deep learning model. Taking the initial residual frequency signal and the sensor reference frequency signal as inputs, usually, the two need to be combined into a composite feature graph. For example, they can be stacked, and the combined feature map is input into the deep learning compensation network for processing to obtain a non-linear echo time-frequency mask. Applying the non-linear echo time-frequency mask to the initial residual frequency signal means multiplying the non-linear echo time-frequency mask by the initial residual frequency signal and performing an inverse short-time Fourier transform to obtain a non-linear echo estimate.

[0085] In one embodiment, step S4 of adjusting the initial residual according to the double-talk state decision function and the non-linear echo estimate to obtain a target residual includes:

[0086] S401: Calculate the short-time energy of the reference signal and the initial residual respectively to obtain the short-time energy of the reference signal and the short-time energy of the initial residual;

[0087] S402: Calculate the ratio of the short-time energy of the initial residual to the short-time energy of the reference signal and determine whether the ratio is greater than a preset threshold;

[0088] S403: If the ratio is greater than the threshold, it is recorded as the double-talk state, and then adjust the residual output according to the formula where, is the target residual, is the initial residual, is the non-linear echo estimate;

[0089] S404: If the ratio is less than or equal to the threshold, it is recorded as the non-double-talk state, and then adjust the residual output according to the formula to adjust the residual output.

[0090] As described in the above steps S401 - S404, the energy detection method is used to detect the double - talk state. Specifically, the calculation formula of short - time energy usually sums the squares of the signals within each time window, calculates the ratio between the short - time energy of the initial residual and the short - time energy of the reference signal, and compares the calculated ratio with a preset threshold. If it is greater than the set threshold, it is recorded as the double - talk state (that is, there are multiple people speaking or there is no obvious difference between the main speech and the background noise). Otherwise, it is recorded as the non - double - talk state. It can effectively and dynamically judge the double - talk state and adjust the residual output according to different states. This mechanism can improve the adaptability of the system to complex speech performances in actual application scenarios, ensuring that even in multi - person conversations or noisy environments, the system can accurately extract the target speech signal, thereby improving the accuracy of speech recognition.

[0091] In one embodiment, step S5 of processing the target residual through the post - Kalman filtering noise reduction module and outputting the pure proximal speech includes:

[0092] S501: Perform a short - time Fourier transform on the target residual to obtain a residual time - frequency signal;

[0093] S502: Model each frequency point of the residual time - frequency signal to obtain the state vector of each frequency point , where, represents the state vector of the t - th frequency point, represents the pure speech amplitude spectrum coefficient of the t - th frequency point, represents the noise amplitude spectrum coefficient of the t - th frequency point;

[0094] S503: Assume that the speech and noise follow a first - order autoregressive model ; where , are the first - order autoregressive coefficients, which are updated iteratively through Kalman filtering, respectively represent the process noise;

[0095] S504: Perform Kalman filtering iteration on the state vector of each frequency point to obtain the target state vector;

[0096] S505: Extract the pure proximal speech from the target state vector.

[0097] As described in the above steps S501 - S503, the target residual is processed by a post - Kalman filter noise reduction module. Specifically, the target residual is subjected to a short - time Fourier transform to obtain a residual time - frequency signal, and modeling is performed for each frequency point. The state vector includes information such as the amplitude, phase, and their change rates of the current frequency point. The state vectors of each frequency point are initialized, which can generally be set to zero or initialized by a certain method (for example, using the data of the first frame), and the initial state and covariance of the system are set. The initial covariance matrix is usually set to a relatively large value to represent uncertainty. For example , where is a large number, that is, a number greater than a preset value, is the identity matrix. The process noise covariance is usually selected according to the characteristics of the system noise, and a suitable value can be obtained through experiments. The observation noise covariance R is estimated based on the noise characteristics of the observed signal. Assume that the speech and background noise follow a first - order autoregressive (AR) model. The autoregressive model indicates that the current state can be represented by a linear combination of the previous state. The prediction step is performed on the state vectors of each frequency point, the state prediction is updated according to the autoregressive model, the state vectors are updated using the observed time - frequency signal, the state estimate is adjusted by calculating the Kalman gain, and the state estimate is updated. Iterative processing is performed on each frequency point, and the state vectors are continuously updated to obtain more accurate clean speech features. The clean speech signal information is extracted from the processed target state vectors, and can be reconstructed according to the amplitude and phase information. The obtained target state vector signal is restored to the time domain, and the final clean proximal speech is obtained through the inverse Fourier transform.

[0098] Referring to Figure 3 , the present invention also provides a vibration sensor circuit for collecting reference signals, including: a power supply filtering module, a vibration sensor module, a positive - feedback amplification module, a two - stage filtering module, and a chip;

[0099] The power supply filtering module is connected to the vibration sensor module. The power supply filtering module is used to filter and supply power to the vibration sensor. The vibration sensor module is in contact connection with the vibrating object to obtain vibration signals;

[0100] The vibration sensor module is connected to the positive - feedback amplification module. The positive - feedback amplification module is used to amplify the vibration signals;

[0101] The positive - feedback amplification module is connected to the two - stage filtering module. The two - stage filtering module is used to filter the amplified vibration signals to obtain reference signals;

[0102] The two - stage filtering module is connected to the chip and is used to transmit the reference signal to the chip for processing.

[0103] Among them, the power supply filtering module includes capacitor C1, capacitor C2, capacitor C3, resistor R1 and resistor R2. The power supply is respectively connected to the first end of capacitor C1, and the first ends of resistor R2 and resistor R1. The second end of capacitor C1 is respectively grounded and connected to the first end of capacitor C2. The second end of capacitor C2 is respectively connected to the second end of resistor R2 and one end of the vibration sensor module. The second end of resistor R1 is respectively connected to one end of the positive feedback amplification module and the first end of capacitor C3. The second end of capacitor C3 is grounded.

[0104] The vibration sensor module includes a vibration sensor, resistor R6, vibration sensor SW2, resistor R3. Among them, the first end of resistor R6 is connected to the second end of resistor R2. The first end of vibration sensor SW2 is connected to one end of resistor R6. The second end of vibration sensor SW2 is respectively connected to the first end of resistor R3 and one end of the positive feedback amplification module. The second end of resistor R3 is grounded.

[0105] The positive feedback amplification module includes capacitor C5, capacitor C7, capacitor C8, capacitor C9, triode Q1, resistor R5, resistor R4 and resistor R9. Among them, the first end of capacitor C5 is connected to the second end of vibration sensor SW2. The second end of capacitor C5 is respectively connected to the first end of capacitor C7, the base of triode Q1 and the first end of resistor R5. The second end of resistor R5 is respectively connected to the second end of resistor R4, the collector of triode Q1, the first end of capacitor C9 and one end of the two-stage filtering module. The emitter of triode Q1 is respectively connected to the first end of resistor R9 and the first end of capacitor C8. The second ends of resistor R9, capacitor C8 and capacitor C9 are all grounded.

[0106] The two-stage filtering module includes capacitor C4, resistor R7, capacitor C10, resistor R8, capacitor C11 and capacitor C6. Among them, the first end of capacitor C4 is connected to the second end of resistor R5. The second end of capacitor C4 is connected to the first end of resistor R7. The second end of resistor R7 is respectively connected to the first end of resistor R8 and the first end of capacitor C10. The second end of capacitor C10 is grounded. The second end of resistor R8 is connected to the first end of capacitor C11 and the first end of capacitor C6. The second end of capacitor C11 is grounded. The second end of capacitor C6 is connected to the chip.

[0107] Referring to Figure 4 , the present invention also provides a voice noise reduction device based on a vibration sensor. The device is applied to a vibration sensor circuit. The vibration sensor circuit includes a vibration sensor. The vibration sensor is in contact connection with a vibrating object. The device includes:

[0108] An acquisition module 10, configured to acquire a mixed signal through a microphone and synchronously acquire a reference signal of the vibration sensor;

[0109] A generating module 20, configured to generate a linear echo estimation through an adaptive filter based on the reference signal and the mixed signal, and calculate an initial residual;

[0110] An extraction module 30, configured to extract the time-frequency features of the initial residual and the sensor reference signal, and input them into a pre-trained deep learning compensation network to obtain a non-linear echo estimation;

[0111] An adjustment module 40, configured to adjust the initial residual according to a double-talk state decision function and the non-linear echo estimation to obtain a target residual;

[0112] An output module 50, configured to process the target residual through a post Kalman filter noise reduction module and output a pure near-end speech.

[0113] In one embodiment, the generating module 20 includes:

[0114] A coefficient adjustment sub-module, configured to iteratively adjust the coefficients of the reference signal through a preset algorithm ; where represents the coefficient of the k-th impulse response;

[0115] A linear echo estimation calculation sub-module, configured to calculate a linear echo estimation according to the formula ; where represents the linear echo estimation, represents the length of the impulse response in the reference signal, represents the (n-k)-th impulse response in the reference signal;

[0116] An initial residual calculation sub-module, configured to minimize the error signal according to the formula to obtain an initial residual ; where represents the mixed signal, represents the initial residual.

[0117] In one embodiment, the preset algorithm is one of a least mean square algorithm, a normalization algorithm, and a recursive least squares method.

[0118] In one embodiment, the extraction module 30 includes:

[0119] A frequency signal conversion sub-module, configured to respectively convert the initial residual and the sensor reference signal into frequency signals through a short-time Fourier transform, and correspondingly obtain an initial residual frequency signal and a sensor reference frequency signal;

[0120] A frequency signal input sub-module, configured to input the initial residual frequency signal and the sensor reference frequency signal into a pre-trained deep learning compensation network to obtain a non-linear echo time-frequency mask;

[0121] A short-time Fourier transform sub-module, which is used to apply the non-linear echo time-frequency mask to the initial residual frequency signal and perform an inverse short-time Fourier transform to obtain a non-linear echo estimation.

[0122] In one embodiment, the adjustment module 40 includes:

[0123] A short-time energy calculation sub-module, which is used to calculate the short-time energy of the reference signal and the initial residual, and obtain the short-time energy of the reference signal and the short-time energy of the initial residual respectively;

[0124] A ratio calculation sub-module, which is used to calculate the ratio of the short-time energy of the initial residual to the short-time energy of the reference signal, and judge whether the ratio is greater than a preset threshold;

[0125] A double-talk state marking sub-module, which is used to record it as a double-talk state if the ratio is greater than the threshold, and then adjust the residual output according to the formula wherein, is the target residual, is the initial residual, is the non-linear echo estimation;

[0126] A non-double-talk state marking sub-module, which is used to record it as a non-double-talk state if the ratio is less than or equal to the threshold, and then adjust the residual output according to the formula Adjust the residual output.

[0127] In one embodiment, the output module 50 includes:

[0128] A transform sub-module, which is used to perform a short-time Fourier transform on the target residual to obtain a residual time-frequency signal;

[0129] A modeling sub-module, which is used to model each frequency point of the residual time-frequency signal to obtain a state vector of each frequency point , wherein, represents the state vector of the t-th frequency point, represents the clean speech amplitude spectrum coefficient of the t-th frequency point, represents the noise amplitude spectrum coefficient of the t-th frequency point;

[0130] An assumption sub-module, which is used to assume that the speech and noise follow a first-order autoregressive model ; wherein , are first-order autoregressive coefficients, which are updated iteratively through Kalman filtering, respectively represent process noise;

[0131] An iteration sub-module, which is used to perform Kalman filter iteration on the state vector of each frequency point to obtain a target state vector;

[0132] An extraction sub-module, configured to extract clean proximal speech from the target state vector.

[0133] Figure 5 The internal structure diagram of a computer device in an embodiment is shown. The computer device may specifically be a terminal or a server. As Figure 5 shown, the computer device includes a processor, a memory, and a network interface connected through a system bus. Among them, the memory includes a non-volatile storage medium and an internal memory. The non-volatile storage medium of the computer device stores an operating system and may also store a computer program. When the computer program is executed by the processor, the processor can be implemented. The internal memory may also store a computer program. When the computer program is executed by the processor, the processor can execute the speech noise reduction method based on a vibration sensor. Those skilled in the art can understand that Figure 5 the structure shown in is only a block diagram of some structures related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.

[0134] In one embodiment, a computer device is provided, including a memory and a processor. The memory stores a computer program. When the computer program is executed by the processor, the processor performs the following steps:

[0135] Collect a mixed signal through a microphone and synchronously collect a reference signal of a vibration sensor;

[0136] Generate a linear echo estimate through an adaptive filter based on the reference signal and the mixed signal, and calculate an initial residual;

[0137] Extract the time-frequency features of the initial residual and the sensor reference signal, and input them into a pre-trained deep learning compensation network to obtain a non-linear echo estimate;

[0138] Adjust the initial residual according to a double-talk state decision function and the non-linear echo estimate to obtain a target residual;

[0139] Process the target residual through a post-Kalman filter noise reduction module and output clean proximal speech.

[0140] The speech noise reduction method based on a vibration sensor effectively reduces the noise in the speech signal through multiple processing steps, improving the accuracy and clarity of speech recognition. Combining the advantages of adaptive filtering and deep learning, the method has good adaptability to various environmental characteristics, improving the response ability and robustness to dynamic environmental changes.

[0141] In one embodiment, a computer-readable storage medium is provided, storing a computer program. When the computer program is executed by a processor, the processor is caused to perform the following steps:

[0142] Collect a mixed signal through a microphone and synchronously collect a reference signal of a vibration sensor;

[0143] Generate a linear echo estimate through an adaptive filter based on the reference signal and the mixed signal, and calculate an initial residual;

[0144] Extract the time-frequency features of the initial residual and the sensor reference signal, and input them into a pre-trained deep learning compensation network to obtain a non-linear echo estimate;

[0145] Adjust the initial residual according to a double-talk state decision function and the non-linear echo estimate to obtain a target residual;

[0146] Process the target residual through a post Kalman filter noise reduction module and output a clean near-end speech.

[0147] The speech noise reduction method based on a vibration sensor effectively reduces the noise in a speech signal through multiple processing steps, improving the accuracy and clarity of speech recognition. By combining the advantages of adaptive filtering and deep learning, the method has good adaptability to various environmental characteristics, improving the response ability and robustness to dynamic environmental changes.

[0148] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The program can be stored in a non-volatile computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the embodiments provided in the present application can include non-volatile and / or volatile memories. Non-volatile memories can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memories can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.

[0149] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.

[0150] The above-described embodiments merely represent several implementation manners of the present application. The description thereof is relatively specific and detailed, but it should not be construed as a limitation on the patent scope of the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent of the present application shall be subject to the appended claims.

Claims

1. A voice noise reduction method based on a vibration sensor, characterized in that, The method is applied to a vibration sensor circuit, which includes a vibration sensor. The vibration sensor is in contact connection with a vibrating object. The method includes: Collecting a mixed signal through a microphone and synchronously collecting a reference signal of the vibration sensor; Generating a linear echo estimate through an adaptive filter based on the reference signal and the mixed signal, and calculating an initial residual; Extracting the time-frequency features of the initial residual and the sensor reference signal and inputting them into a pre-trained deep learning compensation network to obtain a non-linear echo estimate; Adjusting the initial residual according to a double-talk state decision function and the non-linear echo estimate to obtain a target residual; Processing the target residual through a post-Kalman filter noise reduction module and outputting a pure near-end speech; The step of adjusting the initial residual according to a double-talk state decision function and the non-linear echo estimate to obtain a target residual includes: Calculating the short-time energy of the reference signal and the initial residual, respectively obtaining the short-time energy of the reference signal and the short-time energy of the initial residual; Calculating the ratio of the short-time energy of the initial residual to the short-time energy of the reference signal and determining whether the ratio is greater than a preset threshold; If the ratio is greater than the threshold, it is recorded as the double-talk state, and the residual output is adjusted according to the formula ; where is the target residual is the initial residual is the non-linear echo estimation If the ratio is less than or equal to the threshold, it is recorded as a non-dual talk state, and then the residual output is adjusted according to the formula Adjust the residual output.

2. The voice noise reduction method based on a vibration sensor according to claim 1, wherein, The step of generating a linear echo estimate through an adaptive filter based on the reference signal and calculating an initial residual includes: Iteratively adjust the coefficients of the reference signal through a preset algorithm ; wherein represents the coefficient of the k-th impulse response Calculate the linear echo estimation according to the formula wherein represents the linear echo estimation represents the length of the impulse response in the reference signal represents the (n-k)-th impulse response in the reference signal According to the formula minimize the error signal to obtain an initial residual; where represents the mixed signal represents the initial residual 3. The voice noise reduction method based on a vibration sensor according to claim 2, wherein The preset algorithm is one of the least mean square algorithm, the normalization algorithm, and the recursive least squares algorithm.

4. The voice noise reduction method based on a vibration sensor according to claim 1, wherein The step of extracting the time-frequency features of the initial residual and the sensor reference signal and inputting them into a pre-trained deep learning compensation network to obtain a non-linear echo estimate includes: Respectively transforming the initial residual and the sensor reference signal into frequency signals through short-time Fourier transform, correspondingly obtaining an initial residual frequency signal and a sensor reference frequency signal; Inputting the initial residual frequency signal and the sensor reference frequency signal into a pre-trained deep learning compensation network to obtain a non-linear echo time-frequency mask; Applying the non-linear echo time-frequency mask to the initial residual frequency signal and performing inverse short-time Fourier transform to obtain a non-linear echo estimate.

5. The method for voice noise reduction based on a vibration sensor according to claim 1, wherein The step of processing the target residual through a post-Kalman filter noise reduction module and outputting a pure near-end speech includes: Performing short-time Fourier transform on the target residual to obtain a residual time-frequency signal; Model each frequency point of the residual time-frequency signal to obtain the state vector of each frequency point , where represents the state vector of the \(t\)-th frequency point, represents the clean speech amplitude spectrum coefficient of the \(t\)-th frequency point, represents the noise amplitude spectrum coefficient of the \(t\)-th frequency point; Assume that the speech and noise follow a first-order autoregressive model ; where , are the first-order autoregressive coefficients, which are updated iteratively through Kalman filtering, and represent the process noise respectively; Performing Kalman filter iteration on the state vector of each frequency point to obtain a target state vector; Extracting a pure near-end speech from the target state vector.

6. A vibration sensor circuit for collecting a reference signal in the vibration sensor-based voice noise reduction method according to any one of claims 1-5, characterized in that, Including: A power supply filtering module, a vibration sensor module, a positive feedback amplification module, a two-stage filtering module, and a chip; The power supply filtering module is connected to the vibration sensor module. The power supply filtering module is used to filter and supply power to the vibration sensor. The vibration sensor module is in contact connection with a vibrating object to obtain a vibration signal; The vibration sensor module is connected to the positive feedback amplification module. The positive feedback amplification module is used to amplify the vibration signal; The positive feedback amplification module is connected to the two-stage filtering module. The two-stage filtering module is used to filter the amplified vibration signal to obtain a reference signal; The two-stage filtering module is connected to the chip and is used to transfer the reference signal to the chip for processing.

7. A voice noise reduction device based on a vibration sensor, characterized in that, The device is applied to a vibration sensor circuit. The vibration sensor circuit includes a vibration sensor, and the vibration sensor is in contact connection with a vibrating object. The device includes: An acquisition module, configured to acquire a mixed signal through a microphone and synchronously acquire a reference signal of the vibration sensor; A generation module, configured to generate a linear echo estimate through an adaptive filter based on the reference signal and the mixed signal, and calculate an initial residual; An extraction module, configured to extract the time-frequency features of the initial residual and the sensor reference signal and input them into a pre-trained deep learning compensation network to obtain a non-linear echo estimate; An adjustment module, configured to adjust the initial residual according to a double-talk state decision function and the non-linear echo estimate to obtain a target residual; An output module, configured to process the target residual through a post-Kalman filtering noise reduction module and output a pure proximal speech; The adjustment module 40 includes: A short-time energy calculation sub-module, configured to calculate the short-time energy of the reference signal and the initial residual, respectively obtaining the short-time energy of the reference signal and the short-time energy of the initial residual; A ratio calculation sub-module, configured to calculate the ratio of the short-time energy of the initial residual to the short-time energy of the reference signal and determine whether the ratio is greater than a preset threshold; The double-talking state marking sub-module is used to record the double-talking state if the ratio is greater than the threshold, and then adjust the residual output according to the formula ; where is the target residual, is the initial residual, is the non-linear echo estimation; The non-dual talk state marking sub-module is used to record the non-dual talk state if the ratio is less than or equal to the threshold, and then adjust the residual output according to the formula Adjust the residual output.

8. A computer-readable storage medium, characterized in that, A computer program is stored. When the computer program is executed by a processor, the processor is caused to execute the steps of the vibration sensor-based speech noise reduction method according to any one of claims 1 to 5.

9. A computer device, characterized in that, The device includes a memory and a processor. The memory stores a computer program. When the computer program is executed by the processor, the processor is caused to execute the steps of the vibration sensor-based speech noise reduction method according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • Use of vibration sensor in acoustic echo cancellation

    CN104243732A

  • Echo elimination method and device, electronic equipment and computer readable medium

    CN115083431A