Signal denoising method for translation pen
By combining millimeter wave Doppler frequency analysis and microphone array beam adjustment in the translation pen, dynamic classification and differentiated suppression of noise sources are achieved, solving the shortcomings of traditional technology in complex environments, and significantly improving the noise suppression and speech enhancement capabilities of the translation pen.
Patent Information
- Application Number
- CN202510403129.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-01
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2045-04-01
AI Technical Summary
Traditional translation pens are difficult to effectively distinguish high-speed moving noise, low-speed moving noise from static noise in complex dynamic environments, resulting in a lack of targeted noise reduction strategy, and single-mode signal processing is difficult to accurately track the spatial motion characteristics of the noise source, and it is impossible to perform forward-looking suppression before the noise reaches the microphone.
By transmitting millimeter wave detection signals to the environment, analyzing the Doppler frequency offset and frequency change rate of the noise source, classifying the motion of the noise source, and adjusting the beam pattern of the microphone array based on the spatial orientation data of the noise source. Using a differentiated suppression method, the time-frequency characteristics of the Doppler motion trajectory and the acoustic signal are fused to separate speech and noise from the same frequency band through machine learning models.
It realizes high-precision noise suppression and speech enhancement in complex environments, improves noise suppression efficiency, improves voice signal-to-noise ratio, and significantly improves the accuracy and real-time real-time cross-language translation.
Smart Images

Figure CN120148536A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of noise reduction for translation pens, and specifically to a signal denoising method for translation pens. Background Art
[0002] The noise suppression technology of traditional translation pens mainly relies on the single-modal signal processing of acoustic sensors (such as microphone arrays), for example, suppressing background noise through methods such as spectral subtraction and beamforming. However, such methods have significant limitations in complex dynamic environments: on the one hand, traditional technologies are difficult to effectively distinguish high-speed moving noises (such as moving vehicles), low-speed moving noises (such as rotating fans), and static noises (such as air conditioner sounds), resulting in a lack of pertinence in the noise reduction strategy; on the other hand, simply relying on acoustic signals is difficult to accurately track the spatial movement characteristics of noise sources and cannot perform prospective suppression before the noise reaches the microphone. In addition, the existing technology has limited ability to separate co-frequency band speech and noise, especially when the noise and speech spectra overlap, it is easy to cause speech distortion, seriously affecting the translation accuracy.
[0003] In recent years, millimeter-wave radar and Doppler frequency shift analysis technology have gradually been applied to the field of moving target detection, such as identifying pedestrian actions through micro-Doppler features in autonomous driving. However, the application of such technology in the speech noise reduction scenario is still in the exploratory stage, especially the deep integration with acoustic signals and machine learning models is not yet mature. Therefore, there is an urgent need for a technical solution that integrates multi-modal perception and intelligent processing to improve the noise reduction robustness in dynamic environments. Summary of the Invention
[0004] To solve the above technical problems, a signal denoising method for a translation pen is provided. This technical solution solves the problem that the above traditional technology is difficult to effectively distinguish high-speed moving noise, low-speed moving noise, and static noise, resulting in a lack of pertinence in the noise reduction strategy.
[0005] To achieve the above objectives, the technical solution adopted by the present invention is as follows:
[0006] A signal denoising method for a translation pen, comprising:
[0007] Transmitting a millimeter-wave detection signal to the environment, receiving the reflected signal and analyzing the Doppler frequency offset and frequency change rate of the noise source;
[0008] Classifying the motion of the noise source according to the frequency offset and change rate, and the noise source includes, according to the motion classification: high-speed moving noise, low-speed moving noise, and static noise;
[0009] Performing differential suppression on the classified noise sources;
[0010] Adjust the beam pattern of the microphone array in combination with the spatial azimuth data of the noise source to generate a suppression zone in the direction of the noise source and form an enhanced beam in the direction of the user's speech;
[0011] Fuse the Doppler motion trajectory and the time-frequency characteristics of the acoustic signal, and separate the speech and the co-frequency band noise through a machine learning model;
[0012] Output the enhanced speech signal processed by multi-dimensional noise reduction to the translation engine. After being translated by the translation engine, output the translated content through the audio-visual output module.
[0013] As an optimization of this solution, the high-speed moving noise is environmental noise with a frequency offset greater than or equal to 20 Hz and a frequency change rate exceeding 5 Hz / s;
[0014] The low-speed moving noise is environmental noise with a frequency offset between 5 Hz and 20 Hz and a frequency change rate less than or equal to 2 Hz / s;
[0015] The static noise is environmental noise with a frequency offset less than 5 Hz and a fluctuation amplitude lower than 1 Hz.
[0016] Further preferably, the differential suppression of the classified noise sources specifically includes:
[0017] For high-speed moving noise, calculate the approaching speed of the noise source according to the Doppler effect principle, predict the time window for it to reach the microphone, and generate an advanced cancellation signal;
[0018] For low-speed moving noise, construct a dynamic notch filter at the noise characteristic frequency point to suppress interference in a specific frequency band;
[0019] For static noise, use spectral subtraction to eliminate background noise and retain the speech harmonic components.
[0020] Further preferably: The adjustment of the beam pattern of the microphone array in combination with the spatial azimuth data of the noise source specifically includes:
[0021] Obtain the horizontal azimuth angle and vertical azimuth angle of the noise source relative to the microphone array through a millimeter-wave radar or ultrasonic positioning module;
[0022] Based on the azimuth angle data of the noise source, adjust the beamforming weight coefficient of the microphone array to generate a suppression zone in the direction of the noise source, and adopt the above noise suppression method based on the noise source type in the suppression zone;
[0023] Determine the user's lip region through an optical acquisition module. Based on the user's lip region and the horizontal azimuth angle and vertical azimuth angle relative to the microphone array, construct a beam enhancement region, and perform selective gain enhancement in the frequency band of the beam enhancement region, preferentially enhancing the main speech frequency band of 300 Hz to 4 kHz.
[0024] Further preferably: separating speech and in-band noise through the time-frequency features of the fused Doppler motion trajectory and acoustic signal specifically includes:
[0025] Extracting the Mel-frequency cepstral coefficients of the acoustic signal as speech time-frequency features, synchronously inputting the Doppler velocity sequence parsed by the millimeter-wave radar and the azimuth angle data of the noise source, and encoding the Mel-frequency cepstral coefficients, Doppler velocity sequence and azimuth angle data into a three-dimensional feature tensor;
[0026] Calculating the correlation weights between the speech feature channel and the noise feature channel through a cross-modal attention mechanism;
[0027] Fusing multi-modal data based on the attention weights and training a speech separation model based on a machine learning model;
[0028] Generating a speech feature vector after noise suppression based on the speech separation model;
[0029] Performing frequency-domain analysis on the speech feature vector after noise suppression, identifying the high-frequency band corresponding to the consonant, and compensating the gain of the high-frequency band corresponding to the consonant through a dynamic equalizer to retain the transient details of the consonant plosive.
[0030] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0031] Through the deep integration of multi-modal perception and intelligent processing technologies, the present invention realizes high-precision noise suppression and speech enhancement of the translation pen in complex environments. Based on the millimeter-wave Doppler frequency offset and change rate, the noise sources are classified into high-speed moving, low-speed moving and static noises, and are respectively processed differently by using forward cancellation, dynamic notch filtering and spectral subtraction, effectively improving the noise suppression efficiency and improving the speech signal-to-noise ratio. And combining the noise source azimuth data to adjust the microphone beam pattern, generating a suppression area in the direction of the noise source, and at the same time forming a gain beam in the direction of the user's speech, effectively separating the spatial interference between the speech and multi-source noises. By fusing acoustic Mel features, Doppler motion trajectories and azimuth data through a machine learning model, using a cross-modal attention mechanism to enhance the weights of speech-related channels, suppressing in-band noise, and improving the retention rate of consonant details. Significantly improving the accuracy and real-time performance of real-time cross-language translation. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] Figure 1 It is a flowchart of a signal denoising method for a translation pen proposed in Embodiment 1;
[0033] Figure 2 It is a flowchart of a method for reducing noise of high-speed moving noise proposed in Embodiment 2;
[0034] Figure 3Flowchart of the method for reducing low-speed moving noise proposed in Embodiment III;
[0035] Figure 4 Flowchart of the method for reducing static noise proposed in Embodiment IV;
[0036] Figure 5 Flowchart of the method for adjusting the beam pattern of the microphone array proposed in Embodiment V;
[0037] Figure 6 Flowchart of the method for separating speech from co-frequency band noise proposed in Embodiment VI;
[0038] Figure 7 Schematic diagram of the structure of the electronic device of the present invention;
[0039] Figure 8 Schematic diagram of the structure of the computer-readable storage medium of the present invention. Detailed implementation manners
[0040] The following description is used to disclose the present invention so that those skilled in the art can implement the present invention. The preferred embodiments in the following description are only examples, and those skilled in the art can think of other obvious variations.
[0041] Embodiment I:
[0042] Referring to Figure 1 as shown, a signal denoising method for a translation pen includes:
[0043] Transmitting a millimeter-wave detection signal to the environment, receiving the reflected signal and analyzing the Doppler frequency offset and frequency change rate of the noise source;
[0044] Classifying the noise source according to the frequency offset and change rate, and the noise source includes, according to the motion classification: high-speed moving noise, low-speed moving noise, and static noise;
[0045] Performing differential suppression on the classified noise source;
[0046] Combining the spatial orientation data of the noise source, adjusting the beam pattern of the microphone array, generating a suppression area in the direction of the noise source, and forming an enhanced beam in the direction of the user's speech;
[0047] Fusing the Doppler motion trajectory and the time-frequency characteristics of the acoustic signal, and separating speech from co-frequency band noise through a machine learning model;
[0048] Outputting the enhanced speech signal processed by multi-dimensional noise reduction to the translation engine, and after being translated by the translation engine, outputting the translated content through the audio-visual output module.
[0049] Among them, when emitting millimeter-wave detection signals to the environment, a dual-frequency alternating emission mode is adopted. Static objects in the environment (such as walls and furniture) will produce fixed reflections on the millimeter-wave signals, resulting in interference from non-target noise sources in the spectrum. By alternately emitting millimeter-wave signals of two different frequencies, using the characteristic that the frequency shift of the Doppler effect is related to the target movement speed, while the frequency shift of static reflectors is zero or constant, the frequency shift characteristics of dynamic targets are separated.
[0050] Transmission frequency f 1 and f 2 of the millimeter-wave signals, perform spectrum analysis on the reflected signals, and extract the frequency offsets Δf 1 and f 2 of. 1 and Δf 1 .
[0051] The frequency offset of the static reflector should satisfy Δf 1 ≈Δf 2 , while the frequency offset of the dynamic target satisfies Δf 1 ≠Δf 2
[0052] Calculate the frequency shift difference: Δf 动态 =Δf 1 -Δf 2 , eliminate the static reflection interference through the differential signal, and extract the pure frequency shift characteristics of the dynamic target;
[0053] Therefore, dynamic noise sources and static noise sources can be distinguished based on the frequency offset;
[0054] The movement of dynamic noise sources (such as pedestrians and vehicles) usually causes periodic Doppler frequency shift changes (micro-Doppler effect). Through the joint time-frequency domain analysis of the received signals, the vibration or movement characteristics of the noise sources can be extracted;
[0055] Perform short-time Fourier transform on the received signals to generate a time-frequency distribution diagram;
[0056] In the time-frequency distribution diagram, the frequency shift of the dynamic target is manifested as an oblique line or curve that changes with time;
[0057] Use the spectrum ridge detection algorithm (such as short-time energy maximum tracking) to extract the periodic components of the frequency changing with time in the time-frequency diagram.
[0058] For non-periodic noise (such as white noise), its frequency shift characteristics are manifested as random distributions and are eliminated by setting a threshold;
[0059] According to the rate and amplitude of the frequency shift change, the types of noise sources can be further distinguished;
[0060] Based on the above theory, in this embodiment, environmental noise with a frequency offset greater than or equal to 20 Hz and a frequency change rate exceeding 5 Hz / s is classified as high-speed moving noise;
[0061] Environmental noise with a frequency offset between 5 Hz and 20 Hz and a frequency change rate less than or equal to 2 Hz / s is classified as low-speed moving noise;
[0062] Environmental noise with a frequency offset less than 5 Hz and a fluctuation amplitude lower than 1 Hz is classified as static noise.
[0063] Embodiment 2:
[0064] In this embodiment, based on Embodiment 1, for high-speed moving noise, the approaching speed of the noise source is calculated according to the Doppler effect principle, the time window for it to reach the microphone is predicted, and an advanced cancellation signal is generated. Refer to Figure 2 As shown, the specific steps are as follows:
[0065] Calculate the approaching speed of the noise source according to the Doppler effect principle and predict the time window for it to reach the microphone;
[0066] Obtain the phase characteristics of the high-speed moving noise;
[0067] Generate a cancellation signal with a phase opposite to that of the high-speed moving noise within the time window for it to reach the microphone, and superimpose it on the acoustical signal collected in real time;
[0068] Real-time monitor the change in the moving state of the high-speed moving noise source, update the predicted time window for it to reach the microphone, and adjust the parameters of the generated cancellation signal to ensure that the coherence of the cancellation signal is greater than the cancellation coherence threshold.
[0069] Specifically, based on the time of millimeter-wave signal transmission and reception, according to the signal propagation speed, calculate the distance d of the noise source, where c is the signal propagation speed, the speed of light, and t is the time of millimeter-wave signal transmission and reception;
[0070] Based on the Doppler frequency shift, the moving speed v of the noise source in the dual-frequency alternating transmission mode is:
[0071] Based on the calculated moving speed v of the noise source and the distance d of the noise source, combined with the speed of sound, dynamically predict the time window for the noise emitted by the noise source to reach the microphone;
[0072] Within the predicted time window Δt, generate an advanced cancellation signal with a phase opposite to that of the noise signal to cancel the noise about to reach the microphone. According to the spectral characteristics of the noise source (obtained by Doppler frequency shift analysis), design an adaptive inverse filter, and generate an inverse filtering signal in advance on the time axis according to the predicted time window.
[0073] Example 3: On the basis of Example 1, a dynamic notch filter is constructed at the noise characteristic frequency points of the low-speed movement noise to suppress interference in a specific frequency band. Refer to Figure 3 as shown, and specifically includes:
[0074] Perform spectrum analysis on the received low-speed movement noise to extract the noise characteristic frequency points;
[0075] Initialize the center frequency and initial bandwidth of the notch filter according to the noise characteristic frequency points, and monitor the drift amount of the noise spectrum in real time;
[0076] Based on the drift amount of the noise frequency, adaptively adjust the center frequency and bandwidth of the notch filter to suppress interference in a specific frequency band;
[0077] Set a guard band in the main frequency band of the speech signal to avoid the filtering range of the notch filter from avoiding the main frequency band of the speech signal.
[0078] Perform spectrum analysis on the received low-speed movement noise signal to extract the noise characteristic frequency point F noise ;
[0079] Construct a dynamic notch filter and initialize the notch center frequency F center = f noise and the initial bandwidth B initial ;
[0080] Monitor the change of the noise spectrum in real time and dynamically adjust the notch center frequency f center (t) = f noise (t)+Δf 偏移 and the bandwidth B dynamic (t) = B initial (t)+0.5|Δf 偏移 |, which is the drift amount of the noise frequency;
[0081] Construct a dynamic notch filter within the range of f center (t)±B dynamic (t), and at the same time set a guard band in the main frequency band of the speech signal. For example, if the main frequency spectrum range of the speech is from 300 Hz to 4 kHz and the set guard band range is 50 Hz, then the filtering range of the notch filter avoids 350 Hz to 4.05 kHz.
[0082] Example 4:
[0083] On the basis of Example 1, for static noise, use spectral subtraction to eliminate background noise and retain the speech harmonic components. Refer to Figure 4 as shown, and the specific steps include:
[0084] The Fourier transform is used to calculate the average power spectrum of the static noise for the static noise;
[0085] The Fourier transform is performed on the collected speech with static noise to obtain the noisy speech spectrum, and the average power spectrum of the static noise is subtracted from the noisy speech spectrum to obtain a preliminary enhanced spectrum;
[0086] Based on the speech quality detection result, the frequency positions of the speech harmonic structure are identified;
[0087] Compensation is performed on the frequency positions of the harmonic structure in the preliminary enhanced spectrum to obtain a secondary enhanced spectrum;
[0088] Phase recovery is performed on the secondary enhanced spectrum, and the time-frequency speech signal is reconstructed through the inverse Fourier transform.
[0089] Specifically:
[0090] The speech silent segment is judged through the energy threshold and the zero-crossing rate. For example, when the energy of 10 consecutive frames of signals is lower than -40 dB and the zero-crossing rate < 5 times / frame, it is judged as a silent segment, and the STFT is performed on the silent segment signal, and the average power spectrum of multiple frames is calculated as the average power spectrum of the static noise;
[0091] The noisy speech is framed and windowed (such as a Hamming window), and the Fourier transform is performed to obtain the amplitude spectrum and the phase spectrum;
[0092] Subtraction is performed on each average power frequency point of the static noise;
[0093] The fundamental frequency F0 of the speech is extracted by the autocorrelation method or the cepstrum method. For example: the fundamental frequency of male speech is about 80 - 150 Hz, and that of female is about 150 - 300 Hz;
[0094] At the integer multiple frequency points of the fundamental frequency (F0, 2F0, 3F0,...), the energy is increased to 80% - 90% of the original noisy speech;
[0095] The time-frequency speech signal is reconstructed through the inverse Fourier transform.
[0096] Example 5:
[0097] Refer to Figure 5 As shown, in this embodiment, in combination with the spatial orientation data of the noise source, adjusting the beam pattern of the microphone array specifically includes:
[0098] Through a millimeter-wave radar or an ultrasonic positioning module, the horizontal azimuth angle and the vertical azimuth angle of the noise source relative to the microphone array are obtained;
[0099] Based on the azimuth angle data of the noise source, the beamforming weight coefficients of the microphone array are adjusted to generate a suppression area in the direction of the noise source, and the noise suppression method in Embodiments 2 - 4 is adopted in the suppression area based on the type of the noise source;
[0100] Determine the user's lip region through the optical acquisition module, construct a beam enhancement region based on the user's lip region and the horizontal and vertical azimuth angles relative to the microphone array, perform selective gain enhancement in the frequency band of the beam enhancement region, and preferentially enhance the main speech frequency band from 300 Hz to 4 kHz.
[0101] Transmit a 60 GHz - 64 GHz frequency modulated continuous wave signal, receive the reflected signal and analyze the azimuth angle of the noise source. In some preferred embodiments, through the spatial diversity ability of the MIMO antenna array, the horizontal azimuth angle accuracy is achieved at ±0.5°;
[0102] Generate a null region with a width of ±10° (covering 20° to 40°) at the azimuth of the noise source in the beam pattern (such as 30° in the front left), attenuate the noise signal by more than 15 dB through an adaptive algorithm, and if the noise intensity is high (such as mechanical noise), extend the null depth to 30 dB;
[0103] According to the position of the user's lips located by the camera, align the main lobe beam within the range of 0° ± 2° directly in front, increase the gain by 10 dB, and additionally increase the gain by 5 dB in the 300 Hz - 4 kHz frequency band.
[0104] Embodiment Six:
[0105] Refer to Figure 6 As shown, in this embodiment, fusing the Doppler motion trajectory and the time - frequency characteristics of the acoustic signal, separating speech and in - band noise through a machine learning model specifically includes.
[0106] Extract the Mel - Frequency Cepstral Coefficients of the acoustic signal as the speech time - frequency characteristics, synchronously input the Doppler velocity sequence analyzed by the millimeter - wave radar and the azimuth angle data of the noise source, and encode the Mel - Frequency Cepstral Coefficients, Doppler velocity sequence and azimuth angle data into a three - dimensional feature tensor;
[0107] Calculate the correlation weights between the speech feature channel and the noise feature channel through a cross - modal attention mechanism;
[0108] Fuse the multi - modal data based on the attention weights and train a speech separation model based on a machine learning model;
[0109] Generate a speech feature vector after noise suppression based on the speech separation model;
[0110] Perform frequency - domain analysis on the speech feature vector after noise suppression, identify the high - frequency band corresponding to the consonant, and perform gain compensation on the high - frequency band corresponding to the consonant through a dynamic equalizer to retain the transient details of the consonant plosive sound.
[0111] Calculate the Mel Frequency Cepstral Coefficients (MFCCs) after frame windowing, retain the first 20 dimensions of the coefficients to characterize the short-term characteristics of speech, slice them according to a time window (such as 50 ms), generate a Doppler velocity-time series, and represent the azimuth of the noise source in a polar coordinate system.
[0112] Use the speech features (MFCCs) as the Query, and the noise features (Mel Frequency Cepstral Coefficients, Doppler velocity, azimuth angle) as the Key-Value to calculate the attention score and normalize it to a weight value between 0 and 1. When the Doppler velocity is highly correlated with the short-term energy of the speech (correlation coefficient > 0.8), assign a speech channel weight of 0.9; if the noise azimuth coincides with the user's direction (such as an angle < 10°), assign a noise channel weight of 0.2; based on the short-term zero-crossing rate and the high-frequency energy ratio of the speech signal, locate the auxiliary audio segment.
[0113] Furthermore, the method according to the embodiment of the present application can also be implemented with the aid of Figure 7 the architecture of the electronic device shown. As Figure 7 shown, the electronic device 500 may include a bus 501, one or more CPUs 502, a read-only memory (ROM) 503, a random access memory (RAM) 504, a communication port 505 connected to the network, an input / output component 506, a hard disk 507, etc. The storage device in the electronic device 500, such as the ROM 503 or the hard disk 507, can store a signal denoising method for a translation pen provided by the present application. The electronic device 500 may also include a user interface 508. Of course, Figure 7 the architecture shown is only exemplary. When implementing different devices, one or more components shown in the Figure 7 electronic device may be omitted according to actual needs.
[0114] Figure 8 is a schematic diagram of the structure of a computer-readable storage medium provided by an embodiment of the present application. As Figure 8 shown, it is a computer-readable storage medium 600 according to an embodiment of the present application. A computer-readable instruction is stored on the computer-readable storage medium 600. When the computer-readable instruction is run by a processor, it can execute a signal denoising method for a translation pen according to the embodiment of the present application described with reference to the above drawings. The storage medium 600 includes, but is not limited to, for example, volatile memory and / or non-volatile memory. Volatile memory may include, for example, random access memory (RAM) and cache memory, etc. Non-volatile memory may include, for example, read-only memory (ROM), hard disk, flash memory, etc.
[0115] In summary, the advantages of the present invention are as follows: Through the deep integration of multi-modal perception and intelligent processing technologies, the present invention realizes high-precision noise suppression and speech enhancement of the translation pen in complex environments. Based on the millimeter-wave Doppler frequency offset and change rate, the noise sources are classified into high-speed moving, low-speed moving, and static noises, and are respectively processed differentially by using feedforward cancellation, dynamic notch filtering, and spectral subtraction, effectively improving the noise suppression efficiency and improving the speech signal-to-noise ratio. And by combining the azimuth data of the noise source, the microphone beam pattern is adjusted to generate a suppression area in the direction of the noise source, and at the same time, a gain beam is formed in the direction of the user's speech, effectively separating the spatial interference between the speech and multi-source noises. By fusing acoustic Mel features, Doppler motion trajectories, and azimuth data through a machine learning model, the cross-modal attention mechanism is used to enhance the weights of speech-related channels, suppress noises in the same frequency band, and improve the retention rate of consonant details. It significantly improves the accuracy and real-time performance of real-time cross-language translation.
[0116] The above shows and describes the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited by the above embodiments. What is described in the above embodiments and the specification is only the principle of the present invention. Without departing from the spirit and scope of the present invention, the present invention will have various changes and improvements, and these changes and improvements all fall within the scope of the present invention claimed. The scope of protection required by the present invention is defined by the appended claims and their equivalents.
Claims
1. A signal denoising method for a translation pen, characterized in that: include: Transmit millimeter wave detection signals to the environment, receive reflected signals and analyze the Doppler frequency offset and frequency change rate of the noise source; Classifying the noise source by movement according to the frequency offset and the change rate, wherein the noise source includes: high-speed moving noise, low-speed moving noise and static noise according to the movement classification; performing differential suppression on the classified noise sources; Combined with the spatial orientation data of the noise source, the beam pattern of the microphone array is adjusted to generate a suppression zone in the direction of the noise source and an enhanced beam in the direction of the user's voice. The Doppler motion trajectory is integrated with the time-frequency characteristics of the acoustic signal, and the speech and the noise in the same frequency band are separated through the machine learning model; The enhanced speech signal after multi-dimensional noise reduction processing is output to the translation engine. After being translated by the translation engine, the translated content is output through the audio-visual output module.
2. The signal denoising method for a translation pen according to claim 1, characterized in that: The high-speed moving noise is the environmental noise with a frequency offset greater than or equal to 20 Hz and a frequency change rate exceeding 5 Hz / s; The low-speed moving noise is environmental noise with a frequency offset between 5 Hz and 20 Hz and a frequency change rate of less than or equal to 2 Hz / s; The static noise is environmental noise with a frequency offset less than 5 Hz and a fluctuation amplitude less than 1 Hz.
3. The signal denoising method for a translation pen according to claim 2, characterized in that: The performing differential suppression on the classified noise sources specifically includes: For high-speed moving noise, the approach speed of the noise source is calculated based on the Doppler effect principle, the time window of its arrival at the microphone is predicted, and an advance cancellation signal is generated; For low-speed moving noise, a dynamic notch filter is constructed at the noise characteristic frequency point to suppress interference in specific frequency bands; Spectral subtraction is used to eliminate background noise and retain the harmonic components of speech.
4. The signal denoising method for a translation pen according to claim 3, characterized in that: The method of calculating the approach speed of the noise source according to the Doppler effect principle, predicting the time window of the noise source arriving at the microphone, and generating the advance cancellation signal specifically includes: Calculate the approaching speed of the noise source based on the Doppler effect principle and predict the time window when it arrives at the microphone; Obtain the phase characteristics of high-speed moving noise; A cancellation signal with a phase characteristic opposite to that of the high-speed moving noise is generated within the time window when the noise arrives at the microphone, and is superimposed with the acoustic signal collected in real time; The movement state changes of the high-speed moving noise source are monitored in real time, the predicted time window of its arrival at the microphone is updated, and the parameters of the generated cancellation signal are adjusted to ensure that the coherence of the cancellation signal is greater than the cancellation coherence threshold.
5. The signal denoising method for a translation pen according to claim 3, characterized in that: The method of constructing a dynamic notch filter at a noise characteristic frequency point for the low-speed moving noise to suppress interference in a specific frequency band specifically includes: Perform spectrum analysis on the received low-speed mobile noise and extract the characteristic frequency points of the noise; Initializing the center frequency and initial bandwidth of the notch filter according to the noise characteristic frequency point, and monitoring the drift of the noise spectrum in real time; Based on the drift of the noise frequency, the center frequency and bandwidth of the notch filter are adaptively adjusted to suppress interference in specific frequency bands; A protection band is set in the main frequency band of the speech signal to prevent the filtering range of the notch filter from avoiding the main frequency band of the speech signal.
6. The signal denoising method for a translation pen according to claim 3, characterized in that: The method of eliminating background noise by spectrum subtraction and retaining speech harmonic components specifically includes: The average power spectrum of static noise is calculated by Fourier transform. Performing Fourier transform on the collected speech with static noise to obtain the spectrum of the speech with noise, and subtracting the average power spectrum of the static noise from the spectrum of the speech with noise to obtain the preliminary enhanced spectrum; Based on the results of speech harmonic detection, identify the frequency position of the speech harmonic structure; Compensating the frequency position of the harmonic structure in the initial enhanced spectrum to obtain a secondary enhanced spectrum; The phase of the secondary enhanced spectrum is restored and the time-frequency speech signal is reconstructed by inverse Fourier transform.
7. A signal denoising method for a translation pen according to any one of claims 4 to 6, characterized in that: The step of adjusting the beam pattern of the microphone array in combination with the spatial orientation data of the noise source specifically includes: The horizontal azimuth and vertical azimuth of the noise source relative to the microphone array are obtained through a millimeter wave radar or ultrasonic positioning module; Based on the azimuth data of the noise source, adjusting the beamforming weight coefficient of the microphone array, generating a suppression zone in the direction of the noise source, and adopting the noise suppression method according to any one of claims 4 to 6 in the suppression zone based on the type of the noise source; The user's lip area is determined through the optical acquisition module, and a beam enhancement area is constructed based on the user's lip area and the horizontal azimuth and vertical azimuth relative to the microphone array. Selective gain enhancement is performed in the frequency band of the beam enhancement area, with the main voice frequency band from 300Hz to 4kHz being enhanced first.
8. A signal denoising method for a translation pen according to any one of claim 1, characterized in that: The method of fusing the Doppler motion trajectory with the time-frequency characteristics of the acoustic signal and separating the speech and the noise in the same frequency band through the machine learning model specifically includes: Extracting Mel frequency cepstral coefficients of the acoustic signal as speech time-frequency features, synchronously inputting Doppler velocity sequence analyzed by millimeter wave radar and azimuth angle data of the noise source, and encoding the Mel frequency cepstral coefficients, Doppler velocity sequence and azimuth angle data into a three-dimensional feature tensor; The correlation weights between speech feature channels and noise feature channels are calculated through a cross-modal attention mechanism. Fusion of multimodal data based on attention weights, training of speech separation model based on machine learning model; Generate a noise-suppressed speech feature vector based on a speech separation model; The speech feature vector after noise suppression is analyzed in the frequency domain to identify the high frequency band corresponding to the consonant, and the dynamic equalizer is used to perform gain compensation on the high frequency band corresponding to the consonant to retain the transient details of the consonant plosive sound.
9. An electronic device, characterized in that: include: at least one processor; and a memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the signal denoising method for a translation pen as described in any one of claims 1-8.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, a signal denoising method for a translation pen according to any one of claims 1 to 8 is implemented.
Citation Information
Patent Citations
Sound processing method and device, electronic equipment and readable storage medium
CN112185406A
Device module for preventing auditory fatigue and auditory impairment of children
CN115359802A
Audio noise reduction method and system based on deep learning
CN119541515A
Static and low speed moving object detecting device and method
KR1019980075587A
Method and apparatus for adjusting voice recognition processing based on noise characteristics
WO2015017303A1