A Noise-Strict Speech Recognition Optimization Method and System Based on Ultrasonic Micro-Motion Detection
By using ultrasonic micro-motion detection and signal processing technology, and utilizing a speaker and microphone to detect the user's lip movements, the accuracy and robustness issues of speech recognition under noise conditions are solved, achieving efficient speech recognition optimization.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- SUZHOU QIMENGZHE NETWORK TECH CO LTD
- Filing Date
- 2023-06-20
- Publication Date
- 2026-05-26
AI Technical Summary
Under noisy conditions, speech recognition performance degrades. Existing technologies struggle to accurately determine the start and end points of speech, and the recognition rate for noisy speech is low. Multimodal fusion technology is limited by cost, power consumption, and size.
The ultrasonic micro-motion detection method is adopted, which uses a speaker to emit ultrasonic waves and a microphone to receive them. The micro-movement of the user's lips is detected to indicate the presence or absence of speech. The difference frequency signal is extracted using signal processing technology to determine the micro-movement, and the speech recognition is optimized by combining MVDR and NLMS algorithms.
It improves the accuracy and robustness of speech recognition under noise, reduces speech distortion, is suitable for industrial application, and is low in cost and does not require additional equipment.
Smart Images

Figure CN116705021B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of speech recognition technology, specifically relating to a speech recognition optimization method and system under noise based on ultrasonic micro-motion detection. Background Technology
[0002] More and more smart electronic devices support voice interaction, but under noisy conditions, voice recognition performance deteriorates sharply, resulting in an unsatisfactory interactive experience and significantly reducing the value of voice interaction. The most fundamental problem under noise is that machines cannot determine the start and end points of speech, followed by a sharp drop in the recognition rate of noisy speech. To improve speech recognition performance under noise, speech denoising techniques have been employed, including single-channel and multi-channel denoising techniques. While denoising can eliminate or suppress most noise, it inevitably introduces distortion of the speech signal. This distortion requires corresponding adaptation of the speech recognition model; otherwise, recognition performance will also significantly decrease. Using multimodal fusion technology is another approach to solving speech recognition under noise. The most typical solution is visual, one approach using video to detect the speaker's lip movements as a speech activity indicator, and another using lip-reading technology to assist speech recognition. These methods are helpful for speech recognition under noise. However, the cost, power consumption, and size of video equipment limit the widespread application of this solution. Summary of the Invention
[0003] To address the technical problems existing in the prior art, the purpose of this invention is to provide a method and system for optimizing speech recognition under noise based on ultrasonic micro-motion detection.
[0004] To achieve the above objectives and technical effects, the technical solution adopted by this invention is as follows:
[0005] A noise-based speech recognition optimization method based on ultrasonic micro-motion detection includes the following steps:
[0006] S1: The ultrasonic transmitting module continuously sends ultrasonic signals;
[0007] S2: The ultrasound receiving module continuously receives ultrasound signals;
[0008] S3: Preprocess the received ultrasonic signal to obtain the difference frequency signal of the reflected wave from the moving object;
[0009] S4: Based on the obtained difference frequency signal, detect whether there is micro-movement within a specific distance;
[0010] S5: If there is a micro-motion, the output will have voice; otherwise, the output will not have voice.
[0011] Furthermore, the ultrasonic transmitting module is a speaker, the ultrasonic receiving module is a microphone, the speaker emits ultrasonic waves toward the user's lip area, and the microphone must be located at a position that can receive the reflected waves directly from the user's lip area.
[0012] Furthermore, the ultrasonic signal is an orthogonal frequency-modulated continuous wave signal, with each chirp having a starting frequency of f. c =20kHz, bandwidth is B=4kHz, duration is T c =0.05s.
[0013] Furthermore, in step S3, the step of preprocessing the received ultrasonic signal to obtain the difference frequency signal of the reflected wave from the moving object includes:
[0014] The received ultrasonic signal is processed by removing direct waves and reflected waves from static objects, calculating the mixed signal, and performing low-pass filtering to finally obtain the difference frequency signal that reflects the movement of the object at a distance of r.
[0015] Furthermore, the difference frequency signal is a complex exponential signal, expressed as:
[0016]
[0017] in: c is the speed of sound, and λ is the initial frequency f. c The corresponding sound wave wavelength, r is the distance from the target object to the microphone.
[0018] Furthermore, in step S4, the step of detecting whether there is micro-movement within a specific distance based on the obtained difference frequency signal includes:
[0019] First, calculate the difference frequency signal of the echo corresponding to each transmitted chirp signal;
[0020] Subsequently, an FFT is performed on the difference frequency signal to calculate the phase at each frequency point;
[0021] Estimate the distance from the target activity area to the microphone. And from this, the corresponding frequency point on the difference frequency signal is estimated to be...
[0022] Finally, calculate the intermediate frequency of the difference frequency signal of the echo corresponding to N consecutive chirp signals. The variance of the corresponding phase is used to determine whether there are any minor movements in the current target area.
[0023] Furthermore, if the variance is greater than a preset threshold, it is considered that there is slight movement in the current target area, which means that the user is speaking; otherwise, it is considered that there is no slight movement in the target area, which means that the user is not speaking.
[0024] Furthermore, the distance from the target activity area to the microphone It ranges from 10cm to 25cm.
[0025] This invention also discloses a speech recognition optimization system under noise based on ultrasonic micro-motion detection, employing the speech recognition optimization method under noise based on ultrasonic micro-motion detection as described above. The system includes:
[0026] An ultrasonic transmitting module is used to continuously transmit ultrasonic signals;
[0027] An ultrasonic receiving module, adapted to the ultrasonic transmitting module, is used to continuously receive ultrasonic signals;
[0028] The signal processing module, connected to the ultrasonic receiving module, is used to preprocess the received ultrasonic signal to obtain the difference frequency signal of the reflected wave from the moving object.
[0029] The micro-motion detection module, connected to the signal processing module, is used to detect whether there is micro-motion within a specific distance based on the obtained difference frequency signal.
[0030] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0031] This invention discloses a speech recognition optimization method and system based on ultrasonic micro-motion detection under noise. It indicates the presence or absence of target speech by detecting lip micro-movements, and has good noise robustness. It uses the built-in speaker and microphone of the smart device to emit and receive ultrasonic waves to detect whether the lips are micro-moving, without adding extra costs. It has high accuracy, strong practicality, and is suitable for industrial promotion and use. Attached Figure Description
[0032] Figure 1 This is a flowchart of the present invention. Detailed Implementation
[0033] The present invention will now be described in detail so that its advantages and features can be more easily understood by those skilled in the art, thereby providing a clearer and more explicit definition of the scope of protection of the present invention.
[0034] The following provides a brief overview of one or more aspects to offer a basic understanding of them. This overview is not an exhaustive summary of all conceived aspects, nor is it intended to identify key or decisive elements of all aspects, nor to define the scope of any or all aspects. Its sole purpose is to present some concepts of one or more aspects in a simplified form to prepare for the more detailed descriptions that follow.
[0035] like Figure 1 As shown, a noise-based speech recognition optimization method based on ultrasonic micro-motion detection includes the following steps:
[0036] S1: The ultrasonic transmitting module continuously sends ultrasonic signals. T (t);
[0037] S2: The ultrasound receiving module continuously receives ultrasound signals. R (t);
[0038] S3: Preprocess the received ultrasonic signal to obtain the difference frequency signal of the reflected wave from the moving object;
[0039] S4: Based on the obtained difference frequency signal, detect whether there is micro-movement within a specific distance;
[0040] S5: If there is a micro-motion, the output will have voice; otherwise, the output will not have voice.
[0041] Ultrasonic signals are orthogonal frequency modulated continuous wave (FMCW) signals.
[0042] In step S3, the preprocessing of the ultrasonic signal includes removing direct waves and reflected waves from static objects, calculating the mixed signal, and performing low-pass filtering to finally obtain the difference frequency signal of the object's movement at a reaction distance of r. This difference frequency signal is a complex exponential signal, which can be expressed as follows:
[0043]
[0044] in: c is the speed of sound, and λ is the initial frequency f. c The corresponding sound wave wavelength, r is the distance from the target object to the microphone.
[0045] In step S4, the step of detecting whether there is micro-motion within a specific distance based on the obtained difference frequency signal includes:
[0046] First, calculate the difference frequency signal of the echo corresponding to each transmitted chirp signal;
[0047] Subsequently, an FFT is performed on the difference frequency signal to calculate the phase at each frequency point;
[0048] Estimate the distance from the target activity area to the microphone. And from this, the corresponding frequency point on the difference frequency signal is estimated to be...
[0049] Finally, calculate the intermediate frequency of the difference frequency signal of the echo corresponding to N consecutive chirp signals. The variance of the corresponding phase; if the variance is greater than a preset threshold, the current target area is considered to have slight movement; otherwise, the target area is considered not to have slight movement. Slight movement in the target area means the user is speaking, while no slight movement means the user is not speaking.
[0050] In a typical embodiment of this patent, the speaker of the smart device emits ultrasound toward the user's lip area, and the microphone of the smart device must be located in a position that can receive the reflected waves directly from the user's lip area.
[0051] Typically, after detecting a slight movement of the user's mouth, a message can be sent to the speech recognition engine (which can be built into the smart device and deployed locally, or deployed in the cloud and accessed by the smart device via the network) to start speech recognition. After detecting no further movement of the user's mouth, a message can be sent to the speech recognition engine to end speech recognition.
[0052] Typically, detecting slight movements or stillness of the user's mouth can help the voice front-end signal processing module (this module is a built-in function module of smart devices, but not a standard module; for example, mobile phones perform single-channel or multi-channel voice enhancement and echo cancellation to better capture target voice in noisy scenarios) to better suppress noise and reduce voice distortion.
[0053] In a typical embodiment of this patent, for array signal processing technology employing the Minimum Variance Distortionless Response (MVDR) algorithm, when the system detects no micro-movement of the user's mouth, it uses the data collected by the microphone to update the noise covariance matrix; when it detects micro-movement of the user's mouth, it does not update the noise covariance matrix.
[0054] In a typical embodiment of this patent, in echo cancellation scenarios using Normalized Least Mean Square Error (NLMS) or Recursive Least Squares (RLS) algorithms, when the system detects no micro-movement of the user's mouth, it can perform deep cancellation of far-end echoes and near-end noise and update the adaptive filter parameters; when it detects micro-movement of the user's mouth, it does not update the adaptive filter parameters, and the suppression of far-end echoes and near-end noise is appropriately relaxed to prevent near-end speech distortion.
[0055] In a typical embodiment of this patent, for a smartphone assistant application, the distance from the target activity area to the microphone is set. For smart glasses, setting the distance from the target activity area to the microphone is important.
[0056] This invention also discloses a noise-based speech recognition optimization system based on ultrasonic micro-motion detection, comprising:
[0057] An ultrasonic transmitting module is used to continuously transmit ultrasonic signals;
[0058] An ultrasonic receiving module, adapted to the ultrasonic transmitting module, is used to continuously receive ultrasonic signals;
[0059] The signal processing module, connected to the ultrasonic receiving module, is used to preprocess the received ultrasonic signal to obtain the difference frequency signal of the reflected wave from the moving object.
[0060] The micro-motion detection module, connected to the signal processing module, is used to detect whether there is micro-motion within a specific distance based on the obtained difference frequency signal.
[0061] Currently, smart hardware such as mobile phone assistants, TWS earphones, and smart glasses are beginning to support voice interaction. Users can easily control devices, obtain information, or input information using voice commands. These devices are typically equipped with microphones or microphone arrays, and some also have speakers. For devices without speakers, an additional speaker needs to be added.
[0062] Example 1
[0063] like Figure 1 As shown, a noise-based speech recognition optimization method based on ultrasonic micro-motion detection includes the following steps:
[0064] S1: The speaker continuously emits ultrasonic signals. T (t);
[0065] S2: The microphone continuously receives ultrasonic signals. R (t);
[0066] S3: The signal processing module preprocesses the received ultrasound signal, including removing direct waves and static object reflection waves, calculating the mixed signal, low-pass filtering, and finally obtaining the difference frequency signal of the object's movement at a distance of r.
[0067] S4: The micro-motion detection module detects whether there is micro-motion within a specific distance based on the obtained difference frequency signal;
[0068] S5: If there is a micro-motion, the output will have voice; otherwise, the output will not have voice.
[0069] In this embodiment, the ultrasonic signal is an orthogonal frequency-modulated continuous wave (FMCW) signal, and the starting frequency of each chirp is f. c =20kHz, bandwidth is B=4kHz, duration is T c =0.05s.
[0070] The smart device's speaker emits ultrasound towards the user's lip area, and the smart device's microphone must be located in a position that can receive the reflected waves directly from the user's lip area.
[0071] The difference frequency signal in step S3 is a complex exponential signal, which can be expressed as follows:
[0072]
[0073] in: c is the speed of sound, and λ is the initial frequency f. c The corresponding sound wave wavelength, r is the distance from the target object to the microphone.
[0074] In step S4, the step of detecting whether there is micro-motion within a specific distance based on the obtained difference frequency signal includes:
[0075] First, calculate the difference frequency signal of the echo corresponding to each transmitted chirp signal;
[0076] Subsequently, an FFT is performed on the difference frequency signal to calculate the phase at each frequency point;
[0077] Estimate the distance from the target activity area to the microphone. And from this, the corresponding frequency point on the difference frequency signal is estimated to be...
[0078] Finally, calculate the intermediate frequency of the difference frequency signal of the echo corresponding to N consecutive chirp signals. The variance of the corresponding phase; if the variance is greater than a preset threshold, the current target area is considered to have slight movement; otherwise, the target area is considered not to have slight movement. Slight movement in the target area means the user is speaking, while no slight movement means the user is not speaking.
[0079] After detecting a slight movement of the user's mouth, a message can be sent to the speech recognition engine to initiate speech recognition. After detecting no further movement of the user's mouth, a message can be sent to the speech recognition engine to terminate speech recognition.
[0080] Detecting slight movements or stillness of the user's mouth can help the voice front-end signal processing module better suppress noise and reduce voice distortion.
[0081] For array signal processing technology that uses the Minimum Variance Distortionless Response (MVDR) algorithm, when the system detects no micro-movement of the user's mouth, it uses the data collected by the microphone to update the noise covariance matrix; when it detects micro-movement of the user's mouth, it does not update the noise covariance matrix.
[0082] In far-field noisy speech acquisition or interactive scenarios, multi-channel speech signal processing techniques are typically used to enhance target speech from a specific direction and suppress ambient noise. Minimum Variance Distortionless Response (MVDR) is a classic multi-channel array signal enhancement technique. This algorithm first estimates the optimal channel weights. Then, speech enhancement in the target direction is achieved using the following formula.
[0083]
[0084] Among them, the optimal channel weight It can be estimated using the following formula:
[0085]
[0086] in: R is the covariance matrix of the interference noise signal. n The inverse matrix, θ is the array steering vector for the azimuth of the sound source.
[0087] The covariance matrix R of the interference noise signal here n The calculation needs to be performed on signals without target speech and needs to be continuously updated over time. Therefore, the algorithm needs to know in advance precisely which moments are purely interference noise signals. If data with target speech is used in the calculation of the noise covariance matrix, it will lead to inaccurate weight vector estimation and severe distortion of the target speech. Using the method disclosed in this patent, the time segments containing only pure interference noise signals can be identified more accurately. When the system detects that the target user's mouth does not move, it uses the data collected by the microphone to update the noise covariance matrix; when it detects that the user's mouth moves, it does not update the noise covariance matrix.
[0088] In echo cancellation scenarios using Normalized Least Mean Square Error (NLMS) or Recursive Least Squares (RLS) algorithms, when the system detects no micro-movements in the user's mouth, it can perform deep cancellation of far-end echoes and near-end noise and update the adaptive filter parameters. However, when it detects micro-movements in the user's mouth, it does not update the adaptive filter parameters and appropriately relaxes the suppression of far-end echoes and near-end noise to prevent near-end speech distortion.
[0089] Echo cancellation is a classic problem in voice communication and human-computer interaction applications. During audio playback from a speaker, echo cancellation is necessary to achieve clean target speech acquisition via a microphone; this means removing the speaker's signal from the microphone's acquired signal. A typical echo cancellation algorithm uses an adaptive filter to estimate the signal after the speaker's signal propagates to the microphone, and then subtracts it from the microphone's acquired signal. A typical echo cancellation workflow includes: two-way detection, delay estimation, adaptive filtering, and residual echo cancellation. The parameter estimation algorithm for the adaptive filter typically uses Normalized Least Mean Square Error (NLMS) or Recursive Least Squares (RLS) algorithms. When estimating the adaptive filter parameters, it is best to use a time period where only the speaker signal exists. In echo cancellation technology, the state where only the speaker signal exists is called single-way, and other states are called two-way. The step of determining whether only the speaker signal exists is called two-way detection. During echo cancellation, in the single-way state, the filter parameters are estimated and updated; while in the two-way state, methods such as adjusting the step size factor are needed to pause or slow down the filter parameter updates. Traditional two-way speech detection methods primarily rely on a combination of energy and the coherence of far and near-end signals for judgment. However, this method is highly unreliable in situations with strong noise or echo. The method disclosed in this patent can accurately detect the lip movements of the target speaker, thus enabling a more accurate determination of single-way and two-way speech states. When the system detects no lip movements (single-way state), it iteratively updates the adaptive filter parameters, even performing deep cancellation of far-end echo and near-end noise (e.g., directly lowering the microphone recording gain or enhancing residual echo cancellation). Conversely, when lip movements are detected (two-way speech state), the adaptive filter parameters are not updated, and the suppression of far-end echo and near-end noise is appropriately relaxed (e.g., residual echo cancellation is not performed) to prevent near-end speech distortion.
[0090] This embodiment discloses a noise-based speech recognition optimization system based on ultrasonic micro-motion detection, including:
[0091] An ultrasonic transmitting module is used to continuously transmit ultrasonic signals; the ultrasonic transmitting module is the speaker built into the smartphone.
[0092] An ultrasonic receiving module, adapted to the ultrasonic transmitting module, is used to continuously receive ultrasonic signals; the ultrasonic receiving module is a microphone built into a smartphone.
[0093] The signal processing module, connected to the ultrasonic receiving module, is used to preprocess the received ultrasonic signal to obtain the difference frequency signal of the reflected wave from the moving object.
[0094] The micro-motion detection module, connected to the signal processing module, is used to detect whether there is micro-motion within a specific distance based on the obtained difference frequency signal.
[0095] Any parts or structures not specifically described in this invention can be made using existing technologies or products, and will not be elaborated upon here.
[0096] The above description is merely an embodiment of the present invention and does not limit the patent scope of the present invention. Any equivalent structural or procedural transformations made based on the content of the present invention specification, or direct or indirect applications in other related technical fields, are similarly included within the patent protection scope of the present invention.
Claims
1. A method for optimizing speech recognition under noise based on ultrasonic micro-motion detection, characterized in that, Includes the following steps: S1: The ultrasonic transmitting module continuously sends ultrasonic signals; S2: The ultrasound receiving module continuously receives ultrasound signals; S3: Preprocess the received ultrasonic signal to obtain the difference frequency signal of the reflected wave from the moving object; S4: Based on the obtained difference frequency signal, detect whether there is micro-movement within a specific distance; S5: If there is a micro-motion, the output will have voice; otherwise, the output will not have voice. Step S3, which involves preprocessing the received ultrasonic signal to obtain the difference frequency signal of the reflected wave from the moving object, includes: The received ultrasonic signal is processed by removing direct waves and reflected waves from static objects, calculating the mixed signal, and low-pass filtering to finally obtain the difference frequency signal of the object's movement at a distance of r. The difference frequency signal is a complex exponential signal, expressed as: ; in: , c is the speed of sound. The starting frequency f c The corresponding sound wave wavelength, r is the distance from the target object to the microphone; In step S4, the step of detecting whether there is micro-motion within a specific distance based on the obtained difference frequency signal includes: First, calculate the difference frequency signal of the echo corresponding to each transmitted chirp signal; Subsequently, an FFT is performed on the difference frequency signal to calculate the phase at each frequency point; Estimate the distance from the target activity area to the microphone. And from this, the corresponding frequency point on the difference frequency signal is estimated to be... ; Finally, calculate the intermediate frequency of the difference frequency signal of the echo corresponding to N consecutive chirp signals. The variance of the corresponding phase is used to determine whether there is any micro-movement in the current target area; If the variance is greater than the preset threshold, it is considered that there is slight movement in the current target area, which means that the user is speaking; otherwise, it is considered that there is no slight movement in the target area, which means that the user is not speaking. The distance from the target activity area to the microphone It ranges from 10cm to 25cm.
2. The noise-based speech recognition optimization method based on ultrasonic micro-motion detection according to claim 1, characterized in that, The ultrasonic transmitting module is a speaker, and the ultrasonic receiving module is a microphone. The speaker emits ultrasonic waves toward the user's lip area, and the microphone must be located in a position that can receive the reflected waves directly from the user's lip area.
3. The noise-based speech recognition optimization method based on ultrasonic micro-motion detection according to claim 1, characterized in that, The ultrasonic signal is an orthogonal frequency-modulated continuous wave signal, with each chirp having a starting frequency of f. c =20kHz, bandwidth is B=4kHz, duration is T c =0.05s.
4. A speech recognition optimization system under noise based on ultrasonic micro-motion detection, characterized in that, The system employs a noise-based speech recognition optimization method based on ultrasonic micro-motion detection as described in any one of claims 1-3, the system comprising: An ultrasonic transmitting module is used to continuously transmit ultrasonic signals; An ultrasonic receiving module, adapted to the ultrasonic transmitting module, is used to continuously receive ultrasonic signals; The signal processing module, connected to the ultrasonic receiving module, is used to preprocess the received ultrasonic signal to obtain the difference frequency signal of the reflected wave from the moving object. The micro-motion detection module, connected to the signal processing module, is used to detect whether there is micro-motion within a specific distance based on the obtained difference frequency signal.