A signal denoising method for translation pens

By combining millimeter-wave radar and microphone array Doppler frequency offset and acoustic signal processing in the translation pen, noise sources can be classified and differentially suppressed, solving the problem of insufficient noise suppression in traditional translation pens in complex environments, and improving translation accuracy and real-time performance.

CN120148536BActive Publication Date: 2025-10-31SHANDONG PETROCHEMICAL INST
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510403129.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-01
Publication Date
2025-10-31
Estimated Expiration
2045-04-01

AI Technical Summary

Technical Problem

Traditional translation pens struggle to effectively distinguish between high-speed moving noise, low-speed moving noise, and static noise, resulting in a lack of targeted noise reduction strategies. Furthermore, single-modal acoustic signal processing cannot accurately track the spatial motion characteristics of noise sources, thus affecting translation accuracy.

Method used

The Doppler frequency shift and frequency change rate of the noise source are analyzed by using millimeter-wave radar to detect the signal. Combined with microphone array beam pattern adjustment and machine learning model, the motion classification and differential suppression of the noise source are realized. Speech and noise in the same frequency band are separated by fusing the Doppler motion trajectory with the time-frequency characteristics of the acoustic signal.

Benefits of technology

The translation pen achieves high-precision noise suppression and speech enhancement in complex environments, significantly improving the accuracy and real-time performance of real-time cross-language translation. Through the deep integration of multimodal perception and intelligent processing technologies, it improves noise suppression efficiency and speech signal-to-noise ratio.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120148536B_ABST
    Figure CN120148536B_ABST
Patent Text Reader

Abstract

This invention discloses a signal denoising method for translation pens, relating to the field of noise reduction in translation pens. The method includes: transmitting millimeter-wave detection signals to the environment, receiving reflected signals, and analyzing the Doppler frequency offset and frequency change rate of noise sources; classifying noise sources by motion based on the frequency offset and change rate; performing differentiated suppression on the classified noise sources; adjusting the beam pattern of the microphone array by combining the spatial orientation data of the noise sources; fusing the time-frequency characteristics of the Doppler motion trajectory and acoustic signals, and separating speech from noise in the same frequency band using a machine learning model; and outputting the enhanced speech signal processed by multi-dimensional noise reduction to the translation engine. The advantages of this invention are: based on millimeter-wave Doppler frequency offset and change rate, this invention classifies noise sources into high-speed movement, low-speed movement, and static noise, and performs differentiated processing using corresponding methods for each, effectively improving noise suppression efficiency and enhancing the speech signal-to-noise ratio.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of noise reduction for translation pens, specifically to a signal denoising method for translation pens. Background Technology

[0002] Traditional translation pens primarily rely on single-mode signal processing from acoustic sensors (such as microphone arrays) to suppress background noise, using methods like spectral subtraction and beamforming. However, these methods have significant limitations in complex dynamic environments: firstly, traditional techniques struggle to effectively distinguish between high-speed moving noise (such as vehicles), low-speed moving noise (such as rotating fans), and static noise (such as air conditioner sounds), resulting in a lack of targeted noise reduction strategies; secondly, simply relying on acoustic signals makes it difficult to accurately track the spatial motion characteristics of noise sources, hindering proactive noise suppression before noise reaches the microphone. Furthermore, existing technologies have limited ability to separate speech and noise within the same frequency band, especially when noise and speech spectra overlap, easily causing speech distortion and severely impacting translation accuracy.

[0003] In recent years, millimeter-wave radar and Doppler frequency shift analysis technology have been increasingly applied to moving target detection, such as recognizing pedestrian movements using micro-Doppler features in autonomous driving. However, the application of these technologies in speech denoising is still in the exploratory stage, especially the deep integration with acoustic signals and machine learning models, which is not yet mature. Therefore, there is an urgent need for a technical solution that integrates multimodal perception and intelligent processing to improve the robustness of noise reduction in dynamic environments. Summary of the Invention

[0004] To address the aforementioned technical problems, a signal denoising method for translation pens is provided. This technical solution solves the problem that traditional techniques are unable to effectively distinguish between high-speed moving noise, low-speed moving noise, and static noise, resulting in a lack of targeted denoising strategies.

[0005] To achieve the above objectives, the technical solution adopted by the present invention is as follows:

[0006] A signal denoising method for a translation pen includes:

[0007] It emits millimeter-wave detection signals into the environment, receives reflected signals, and analyzes the Doppler frequency shift and frequency change rate of noise sources.

[0008] The noise sources are classified according to their motion based on the frequency offset and the rate of change. The noise sources are classified according to motion as: high-speed moving noise, low-speed moving noise, and static noise.

[0009] Differential suppression is performed on the classified noise sources;

[0010] By combining the spatial orientation data of the noise source, the beam pattern of the microphone array is adjusted to generate a suppression zone in the direction of the noise source and an enhancement beam in the direction of the user's voice.

[0011] By integrating the time-frequency features of Doppler motion trajectories and acoustic signals, speech is separated from noise in the same frequency band through a machine learning model;

[0012] The enhanced speech signal, after multi-dimensional noise reduction processing, is output to the translation engine. After translation by the translation engine, the translated content is output through the audio-visual output module.

[0013] As a preferred embodiment of this solution, the high-speed moving noise is environmental noise with a frequency offset greater than or equal to 20Hz and a frequency change rate exceeding 5Hz / s.

[0014] The low-speed moving noise is environmental noise with a frequency offset between 5Hz and 20Hz and a frequency change rate of less than or equal to 2Hz / s.

[0015] The static noise is environmental noise with a frequency offset of less than 5 Hz and a fluctuation amplitude of less than 1 Hz.

[0016] More preferably, the differential suppression of the classified noise sources specifically includes:

[0017] The approach velocity of the noise source is calculated based on the Doppler effect principle for high-speed moving noise, the time window of its arrival at the microphone is predicted, and a lead cancellation signal is generated.

[0018] A dynamic notch filter is constructed at the noise characteristic frequency point for low-speed moving noise to suppress interference in specific frequency bands;

[0019] Static noise is eliminated by spectral subtraction while preserving speech harmonic components.

[0020] A further preferred embodiment: adjusting the beam pattern of the microphone array by combining the spatial orientation data of the noise source specifically includes:

[0021] The horizontal and vertical azimuth angles of the noise source relative to the microphone array are obtained using millimeter-wave radar or ultrasonic positioning modules.

[0022] Based on the azimuth data of the noise source, the beamforming weighting coefficient of the microphone array is adjusted to generate a suppression zone in the direction of the noise source. In the suppression zone, the noise suppression method described above is applied based on the noise source type.

[0023] The user's lip area is determined by the optical acquisition module. Based on the user's lip area and the horizontal and vertical azimuth angles relative to the microphone array, a beam enhancement area is constructed. Selective gain enhancement is performed in the frequency band of the beam enhancement area, prioritizing the enhancement of the main voice frequency band from 300Hz to 4kHz.

[0024] Further preferred embodiment: The method of fusing the time-frequency features of Doppler motion trajectories and acoustic signals, and separating speech from noise in the same frequency band using a machine learning model, specifically includes:

[0025] Mel frequency cepstral coefficients of acoustic signals are extracted as speech time-frequency features. Doppler velocity sequences and azimuth data of noise sources are synchronously input from millimeter-wave radar. The Mel frequency cepstral coefficients, Doppler velocity sequences and azimuth data are encoded into a three-dimensional feature tensor.

[0026] The correlation weights between speech feature channels and noise feature channels are calculated using a cross-modal attention mechanism.

[0027] Multimodal data is fused based on attention weights, and a speech separation model is trained based on a machine learning model.

[0028] Generate noise-suppressed speech feature vectors based on a speech separation model;

[0029] Frequency domain analysis is performed on the noise-suppressed speech feature vectors to identify the high-frequency bands corresponding to consonants. A dynamic equalizer is then used to compensate the gain of the high-frequency bands corresponding to consonants, preserving the transient details of consonant plosives.

[0030] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0031] This invention achieves high-precision noise suppression and speech enhancement in complex environments through the deep integration of multimodal perception and intelligent processing technologies. Based on millimeter-wave Doppler frequency offset and change rate, noise sources are classified into high-speed movement, low-speed movement, and static noise, and differentiated processing is performed using lead cancellation, dynamic notch filtering, and spectral subtraction respectively, effectively improving noise suppression efficiency and speech signal-to-noise ratio. Furthermore, the microphone beam pattern is adjusted by combining noise source location data to generate a suppression zone in the direction of the noise source and a gain beam in the direction of the user's speech, effectively separating spatial interference between speech and multi-source noise. By fusing acoustic Mel features, Doppler motion trajectories, and location data through a machine learning model, and utilizing a cross-modal attention mechanism to enhance the weights of speech-related channels, noise in the same frequency band is suppressed, and the consonant detail retention rate is improved. This significantly enhances the accuracy and real-time performance of real-time cross-language translation. Attached Figure Description

[0032] Figure 1 This is a flowchart of the signal denoising method for a translation pen proposed in Example 1;

[0033] Figure 2 This is a flowchart of the method for noise reduction of high-speed mobile devices proposed in Example 2;

[0034] Figure 3This is a flowchart of the method for reducing low-speed moving noise proposed in Example 3;

[0035] Figure 4 This is a flowchart of the static noise reduction method proposed in Example 4;

[0036] Figure 5 This is a flowchart of the method for adjusting the beam pattern of a microphone array as proposed in Example 5;

[0037] Figure 6 This is a flowchart of the method for separating speech from noise in the same frequency band proposed in Example 6;

[0038] Figure 7 This is a schematic diagram of the electronic device structure of the present invention;

[0039] Figure 8 This is a schematic diagram of the computer-readable storage medium structure of the present invention. Detailed Implementation

[0040] The following description is intended to disclose the invention and enable those skilled in the art to implement it. The preferred embodiments described below are merely examples, and other obvious variations will occur to those skilled in the art.

[0041] Example 1:

[0042] Reference Figure 1 As shown, a signal denoising method for a translation pen includes:

[0043] It emits millimeter-wave detection signals into the environment, receives reflected signals, and analyzes the Doppler frequency shift and frequency change rate of noise sources.

[0044] The noise sources are classified according to their motion based on the frequency offset and the rate of change. The noise sources are classified according to motion as: high-speed moving noise, low-speed moving noise, and static noise.

[0045] Differential suppression is performed on the classified noise sources;

[0046] By combining the spatial orientation data of the noise source, the beam pattern of the microphone array is adjusted to generate a suppression zone in the direction of the noise source and an enhancement beam in the direction of the user's voice.

[0047] By integrating the time-frequency features of Doppler motion trajectories and acoustic signals, speech is separated from noise in the same frequency band through a machine learning model;

[0048] The enhanced speech signal, after multi-dimensional noise reduction processing, is output to the translation engine. After translation by the translation engine, the translated content is output through the audio-visual output module.

[0049] The method employs a dual-frequency alternating transmission mode to transmit millimeter-wave detection signals into the environment. Static objects in the environment (such as walls and furniture) will produce fixed reflections of the millimeter-wave signals, resulting in interference from non-target noise sources in the spectrum. By alternately transmitting two millimeter-wave signals of different frequencies, the frequency shift of the Doppler effect is related to the target's velocity, while the frequency shift of static reflectors is zero or constant, thus separating the frequency shift characteristics of dynamic targets.

[0050] Millimeter-wave signals with transmission frequencies f1 and f2 are transmitted, and the reflected signals are subjected to spectral analysis to extract the frequency offsets Δf1 and Δf2 of f1 and f2.

[0051] The frequency shift of a static reflector should satisfy Δf1≈Δf2, while the frequency shift of a dynamic target should satisfy Δf1≠Δf2.

[0052] Calculate the frequency shift difference: Δf 动态 =Δf1-Δf2, using differential signals to eliminate static reflection interference and extract the clean frequency shift characteristics of dynamic targets;

[0053] Therefore, dynamic noise sources and static noise sources can be distinguished based on frequency offset;

[0054] The movement of dynamic noise sources (such as pedestrians and vehicles) typically causes periodic Doppler frequency shifts (micro-Doppler effect). By jointly analyzing the time and frequency domains of the received signal, the vibration or movement characteristics of the noise source can be extracted.

[0055] Perform a short-time Fourier transform on the received signal to generate a time-frequency distribution map;

[0056] In the time-frequency distribution diagram, the frequency shift of a dynamic target is represented by a sloping line or curve that changes with time;

[0057] Use spectral ridge detection algorithms (such as short-time energy maximum tracking) to extract the periodic components of frequency variation over time in the time-frequency graph.

[0058] For non-periodic noise (such as white noise), its frequency shift characteristics are randomly distributed and can be eliminated by setting a threshold.

[0059] Based on the rate and amplitude of frequency shift change, the type of noise source can be further distinguished;

[0060] Based on the above theory, in this embodiment, environmental noise with a frequency offset greater than or equal to 20Hz and a frequency change rate exceeding 5Hz / s is classified as high-speed moving noise.

[0061] Low-speed moving noise is defined as environmental noise with a frequency offset between 5 Hz and 20 Hz and a frequency change rate of less than or equal to 2 Hz / s.

[0062] Environmental noise with a frequency offset of less than 5 Hz and a fluctuation amplitude of less than 1 Hz is classified as static noise.

[0063] Example 2:

[0064] In this embodiment, based on Embodiment 1, the approach velocity of the noise source is calculated according to the Doppler effect principle for high-speed moving noise, the time window for its arrival at the microphone is predicted, and a lead cancellation signal is generated, referring to... Figure 2 As shown, the specific steps are as follows:

[0065] The approach velocity of the noise source is calculated based on the Doppler effect principle, and the time window for its arrival at the microphone is predicted.

[0066] Obtain the phase characteristics of high-speed moving noise;

[0067] During the time window of its arrival at the microphone, a cancellation signal with the opposite phase characteristics to the high-speed moving noise is generated and superimposed on the acoustic signal acquired in real time.

[0068] The system monitors the movement status changes of high-speed moving noise sources in real time, updates the predicted time window for their arrival at the microphone, and adjusts the parameters of the generated cancellation signal to ensure that the coherence of the cancellation signal is greater than the cancellation coherence threshold.

[0069] Specifically, the distance d of the noise source is calculated based on the transmission and reception times of the millimeter-wave signal and the signal propagation speed. Where c is the signal propagation speed, t is the speed of light, and t is the time for the millimeter wave signal to be transmitted and received.

[0070] Based on the Doppler frequency shift, the moving speed v of the noise source in the dual-frequency alternating transmission mode is:

[0071] Based on the calculated moving speed v of the noise source and the distance d of the noise source, combined with the speed of sound, the time window for the noise emitted by the noise source to reach the microphone is dynamically predicted.

[0072] Within the prediction time window Δt, an advance cancellation signal with the opposite phase to the noise signal is generated to cancel the noise that is about to reach the microphone. Based on the spectral characteristics of the noise source (obtained by Doppler frequency shift analysis), an adaptive inverse filter is designed, and the inverse filter signal is generated in advance on the time axis according to the prediction time window.

[0073] Example 3: Based on Example 1, a dynamic notch filter is constructed at the noise characteristic frequency point for low-speed moving noise to suppress interference in specific frequency bands, referring to... Figure 3 As shown, it specifically includes:

[0074] Perform spectrum analysis on the received low-speed moving noise to extract the noise characteristic frequency points;

[0075] The center frequency and initial bandwidth of the notch filter are initialized based on the noise characteristic frequency points, and the drift of the noise spectrum is monitored in real time.

[0076] Based on the drift of the noise frequency, the center frequency and bandwidth of the notch filter are adaptively adjusted to suppress interference in specific frequency bands;

[0077] Set a guard band in the main frequency band of the voice signal to prevent the filtering range of the notch filter from avoiding the main frequency band of the voice signal.

[0078] Spectral analysis is performed on the received low-speed moving noise signal to extract the noise characteristic frequency point F. noise ;

[0079] Construct a dynamic notch filter and initialize the notch center frequency F. center =f noise and initial bandwidth B initial ;

[0080] Real-time monitoring of noise spectrum changes and dynamic adjustment of notch filter center frequency f center (t)=f noise (t)+Δf 偏移 and bandwidth B dynamic (t)=B initial (t)+0.5|Δf 偏移 | represents the frequency shift of the noise.

[0081] In f center (t)±B dynamic A dynamic notch filter is constructed within the range of (t), and a guard band is set in the main frequency band of the speech signal. For example, if the main frequency range of speech is 300Hz to 4kHz, and the guard band range is set to 50Hz, then the filtering range of the notch filter avoids 350Hz to 4.05kHz.

[0082] Example 4:

[0083] Based on Example 1, background noise is eliminated by spectral subtraction of static noise while preserving speech harmonic components, referring to... Figure 4 As shown, the specific steps include:

[0084] The average power spectrum of static noise is calculated using Fourier transform.

[0085] The acquired speech with static noise is subjected to Fourier transform to obtain the noisy speech spectrum. The average power spectrum of static noise is subtracted from the noisy speech spectrum to obtain the preliminary enhanced spectrum.

[0086] Based on the speech harmonic detection results, the frequency positions of the speech harmonic structure are identified;

[0087] The frequency position of the harmonic structure is compensated in the initial enhanced spectrum to obtain the secondary enhanced spectrum;

[0088] Phase recovery is performed on the secondary enhanced spectrum, and the time-frequency speech signal is reconstructed through inverse Fourier transform.

[0089] Specifically:

[0090] The speech silence segment is determined by the energy threshold and the zero-crossing rate. For example, when the signal energy of 10 consecutive frames is below -40dB and the zero-crossing rate is <5 times / frame, it is determined to be a silence segment. The silence segment signal is subjected to STFT, and the average power spectrum of multiple frames is calculated as the average power spectrum of static noise.

[0091] The noisy speech is framed and windowed (e.g., Hamming window), and then Fourier transform is performed to obtain the amplitude spectrum and phase spectrum.

[0092] Subtract the average power frequency for each static noise level;

[0093] The fundamental frequency F0 of speech can be extracted by autocorrelation or cepstral method. For example, the fundamental frequency of male speech is about 80-150Hz and that of female speech is about 150-300Hz.

[0094] At integer multiples of the fundamental frequency (F0, 2F0, 3F0, ...), the energy is boosted to 80%-90% of the original noisy speech;

[0095] The time-frequency speech signal is reconstructed using inverse Fourier transform.

[0096] Example 5:

[0097] Reference Figure 5 As shown, in this embodiment, adjusting the beam pattern of the microphone array by combining the spatial orientation data of the noise source specifically includes:

[0098] The horizontal and vertical azimuth angles of the noise source relative to the microphone array are obtained using millimeter-wave radar or ultrasonic positioning modules.

[0099] Based on the azimuth data of the noise source, the beamforming weighting coefficient of the microphone array is adjusted to generate a suppression zone in the direction of the noise source. In the suppression zone, the noise suppression method as in Examples 2-4 is adopted based on the noise source type.

[0100] The user's lip area is determined by the optical acquisition module. Based on the user's lip area and the horizontal and vertical azimuth angles relative to the microphone array, a beam enhancement area is constructed. Selective gain enhancement is performed in the frequency band of the beam enhancement area, prioritizing the enhancement of the main voice frequency band from 300Hz to 4kHz.

[0101] Transmitting 60GHz-64GHz frequency modulated continuous wave signals, receiving reflected signals and resolving the azimuth angle of noise sources, in some preferred embodiments, the horizontal azimuth accuracy of ±0.5° is achieved through the spatial diversity capability of the MIMO antenna array;

[0102] In the beam pattern, a null region with a width of ±10° (covering 20° to 40°) is generated at the noise source location (e.g., 30° to the left front) in the noise source pattern. The noise signal is attenuated by more than 15dB through an adaptive algorithm. If the noise intensity is high (e.g., mechanical noise), the null depth is extended to 30dB.

[0103] Based on the user's lip position located by the camera, the main lobe beam is aligned with the area within 0°±2° directly in front, resulting in a 10dB gain increase, and an additional 5dB gain increase in the 300Hz-4kHz frequency band.

[0104] Example 6:

[0105] Reference Figure 6 As shown, in this embodiment, the separation of speech and noise in the same frequency band by fusing the time-frequency features of Doppler motion trajectory and acoustic signal and using a machine learning model specifically includes...

[0106] Mel frequency cepstral coefficients of acoustic signals are extracted as speech time-frequency features. Doppler velocity sequences and azimuth data of noise sources are synchronously input from millimeter-wave radar. The Mel frequency cepstral coefficients, Doppler velocity sequences and azimuth data are encoded into a three-dimensional feature tensor.

[0107] The correlation weights between speech feature channels and noise feature channels are calculated using a cross-modal attention mechanism.

[0108] Multimodal data is fused based on attention weights, and a speech separation model is trained based on a machine learning model.

[0109] Generate noise-suppressed speech feature vectors based on a speech separation model;

[0110] Frequency domain analysis is performed on the noise-suppressed speech feature vectors to identify the high-frequency bands corresponding to consonants. A dynamic equalizer is then used to compensate the gain of the high-frequency bands corresponding to consonants, preserving the transient details of consonant plosives.

[0111] After frame segmentation and windowing, the Mel frequency cepstral coefficients are calculated. The first 20 coefficients are retained to represent the short-term characteristics of the speech. The speech is sliced ​​according to time windows (e.g., 50ms) to generate a Doppler velocity-time series. The location of the noise source is represented in polar coordinates.

[0112] Speech features (MFCC) are used as queries, and noise features (Mel frequency cepstral coefficients, Doppler velocity, and azimuth angle) are used as key-value pairs. Attention scores are calculated and normalized to weights of 0-1. When the Doppler velocity is highly correlated with the short-time energy of the speech (correlation coefficient > 0.8), the speech channel is assigned a weight of 0.9. If the noise azimuth coincides with the user's direction (e.g., the angle < 10°), the noise channel is assigned a weight of 0.2. Based on the short-time zero-crossing rate and the proportion of high-frequency energy in the speech signal, the consonant band is located.

[0113] Furthermore, the method according to the embodiments of this application can also be achieved by means of... Figure 7 The architecture of the electronic device shown is used to implement this. For example... Figure 7 As shown, the electronic device 500 may include a bus 501, one or more CPUs 502, a read-only memory (ROM) 503, a random access memory (RAM) 504, a communication port 505 connected to a network, an input / output component 506, a hard disk 507, etc. The storage device in the electronic device 500, such as the ROM 503 or the hard disk 507, may store a signal denoising method for a translation pen provided in this application. The electronic device 500 may also include a user interface 508. Of course, Figure 7 The architecture shown is merely exemplary and can be omitted as needed when implementing different devices. Figure 7 One or more components in the illustrated electronic device.

[0114] Figure 8 This is a schematic diagram of a computer-readable storage medium structure provided in one embodiment of this application. Figure 8 The diagram illustrates a computer-readable storage medium 600 according to one embodiment of this application. The computer-readable storage medium 600 stores computer-readable instructions. When executed by a processor, the computer-readable instructions can perform a signal denoising method for a translation pen according to an embodiment of this application, as described with reference to the above figures. The storage medium 600 includes, but is not limited to, volatile memory and / or non-volatile memory. Volatile memory may include, for example, random access memory (RAM) and cache memory. Non-volatile memory may include, for example, read-only memory (ROM), hard disk, flash memory, etc.

[0115] In summary, the advantages of this invention are as follows: By deeply integrating multimodal perception and intelligent processing technologies, this invention achieves high-precision noise suppression and speech enhancement for the translation pen in complex environments. Based on millimeter-wave Doppler frequency offset and change rate, noise sources are classified into high-speed movement, low-speed movement, and static noise, and differentiated processing is performed using lead cancellation, dynamic notch filtering, and spectral subtraction respectively, effectively improving noise suppression efficiency and speech signal-to-noise ratio. Furthermore, by combining noise source location data with microphone beam pattern adjustment, a suppression zone is generated in the direction of the noise source, while a gain beam is formed in the direction of the user's speech, effectively separating spatial interference between speech and multi-source noise. Through machine learning models that fuse acoustic Mel features, Doppler motion trajectories, and location data, a cross-modal attention mechanism is used to enhance the weights of speech-related channels, suppressing noise in the same frequency band and improving consonant detail retention. This significantly improves the accuracy and real-time performance of real-time cross-language translation.

[0116] The foregoing has shown and described the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The embodiments and descriptions in the specification are merely principles of the invention. Various changes and modifications can be made to the invention without departing from its spirit and scope, and all such changes and modifications fall within the scope of the claimed invention. The scope of protection claimed by the appended claims and their equivalents is defined.

Claims

1. A signal denoising method for a translation pen, characterized in that, include: It emits millimeter-wave detection signals into the environment, receives reflected signals, and analyzes the Doppler frequency shift and frequency change rate of noise sources. The noise sources are classified according to their motion based on the frequency offset and the rate of change. The noise sources are classified according to motion as: high-speed moving noise, low-speed moving noise, and static noise. Differential suppression is performed on the classified noise sources; By combining the spatial orientation data of the noise source, the beam pattern of the microphone array is adjusted to generate a suppression zone in the direction of the noise source and an enhancement beam in the direction of the user's voice. By integrating the time-frequency features of Doppler motion trajectories and acoustic signals, speech is separated from noise in the same frequency band through a machine learning model; The enhanced speech signal, after multi-dimensional noise reduction processing, is output to the translation engine. After translation by the translation engine, the translated content is output through the audio-visual output module. The method of fusing the time-frequency features of Doppler motion trajectories and acoustic signals to separate speech from noise in the same frequency band using a machine learning model specifically includes: Mel frequency cepstral coefficients of acoustic signals are extracted as speech time-frequency features. Doppler velocity sequences and azimuth data of noise sources are synchronously input from millimeter-wave radar. The Mel frequency cepstral coefficients, Doppler velocity sequences and azimuth data are encoded into a three-dimensional feature tensor. The correlation weights between speech feature channels and noise feature channels are calculated using a cross-modal attention mechanism. Multimodal data is fused based on attention weights, and a speech separation model is trained based on a machine learning model. Generate noise-suppressed speech feature vectors based on a speech separation model; Frequency domain analysis is performed on the noise-suppressed speech feature vectors to identify the high-frequency bands corresponding to consonants. A dynamic equalizer is then used to compensate the gain of the high-frequency bands corresponding to consonants, preserving the transient details of consonant plosives.

2. The signal denoising method for a translation pen according to claim 1, characterized in that, The high-speed moving noise is environmental noise with a frequency offset greater than or equal to 20 Hz and a frequency change rate exceeding 5 Hz / s; The low-speed moving noise is environmental noise with a frequency offset between 5 Hz and 20 Hz and a frequency change rate of less than or equal to 2 Hz / s; The static noise is environmental noise with a frequency offset of less than 5 Hz and a fluctuation amplitude of less than 1 Hz.

3. The signal denoising method for a translation pen according to claim 2, characterized in that, The differential suppression of the classified noise sources specifically includes: The approach velocity of the noise source is calculated based on the Doppler effect principle for high-speed moving noise, the time window of its arrival at the microphone is predicted, and a lead cancellation signal is generated. A dynamic notch filter is constructed at the noise characteristic frequency point for low-speed moving noise to suppress interference in specific frequency bands; Static noise is eliminated by spectral subtraction while preserving speech harmonic components.

4. The signal denoising method for a translation pen according to claim 3, characterized in that, The process of calculating the approach velocity of the noise source based on the Doppler effect principle, predicting its arrival time window at the microphone, and generating a lead cancellation signal specifically includes: The approach velocity of the noise source is calculated based on the Doppler effect principle, and the time window for its arrival at the microphone is predicted. Obtain the phase characteristics of high-speed moving noise; During the time window of its arrival at the microphone, a cancellation signal with the opposite phase characteristics to the high-speed moving noise is generated and superimposed on the acoustic signal acquired in real time. The system monitors the movement status changes of high-speed moving noise sources in real time, updates the predicted time window for their arrival at the microphone, and adjusts the parameters of the generated cancellation signal to ensure that the coherence of the cancellation signal is greater than the cancellation coherence threshold.

5. The signal denoising method for a translation pen according to claim 3, characterized in that, The method of constructing a dynamic notch filter at the characteristic frequency point of low-speed moving noise to suppress interference in specific frequency bands specifically includes: Perform spectrum analysis on the received low-speed moving noise to extract the noise characteristic frequency points; The center frequency and initial bandwidth of the notch filter are initialized based on the noise characteristic frequency points, and the drift of the noise spectrum is monitored in real time. Based on the drift of the noise frequency, the center frequency and bandwidth of the notch filter are adaptively adjusted to suppress interference in specific frequency bands; Set a guard band in the main frequency band of the voice signal to prevent the filtering range of the notch filter from avoiding the main frequency band of the voice signal.

6. The signal denoising method for a translation pen according to claim 3, characterized in that, The method of eliminating background noise by spectral subtraction while preserving speech harmonic components specifically includes: The average power spectrum of static noise is calculated using Fourier transform. The acquired speech with static noise is subjected to Fourier transform to obtain the noisy speech spectrum. The average power spectrum of static noise is subtracted from the noisy speech spectrum to obtain the preliminary enhanced spectrum. Based on the speech harmonic detection results, the frequency positions of the speech harmonic structure are identified; The frequency position of the harmonic structure is compensated in the initial enhanced spectrum to obtain the secondary enhanced spectrum; Phase recovery is performed on the secondary enhanced spectrum, and the time-frequency speech signal is reconstructed through inverse Fourier transform.

7. A signal denoising method for a translation pen according to any one of claims 4-6, characterized in that, The adjustment of the microphone array beam pattern by combining the spatial orientation data of the noise source specifically includes: The horizontal and vertical azimuth angles of the noise source relative to the microphone array are obtained using millimeter-wave radar or ultrasonic positioning modules. Based on the azimuth data of the noise source, the beamforming weighting coefficient of the microphone array is adjusted to generate a suppression zone in the direction of the noise source. In the suppression zone, the noise suppression method described in any one of claims 4-6 is applied based on the noise source type. The user's lip area is determined by an optical acquisition module. Based on the user's lip area and the horizontal and vertical azimuth angles relative to the microphone array, a beam enhancement area is constructed. Selective gain enhancement is performed in the frequency band of the beam enhancement area, prioritizing the enhancement of the main voice frequency band from 300 Hz to 4 kHz.

8. An electronic device, characterized in that, include: At least one processor; and a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform a signal denoising method for a translation pen as described in any one of claims 1-7.

9. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements a signal denoising method for a translation pen according to any one of claims 1-7.

Citation Information

Patent Citations

  • Sound processing method and device, electronic equipment and readable storage medium

    CN112185406A

  • Device module for preventing auditory fatigue and auditory impairment of children

    CN115359802A

  • Method and apparatus for adjusting voice recognition processing based on noise characteristics

    WO2015017303A1