Cross-language hearing aid with speech translation function
Through cross-language hearing aids that integrate multilingual speech acquisition, intelligent noise processing and real-time speech translation functions, traditional hearing aids have solved the problems of communication difficulties and cross-lingual disorders in noisy environments, and achieved clear speech acquisition and real-time language translation in complex noise environments, improving the communication ability and quality of life of people with hearing impairment.
Patent Information
- Application Number
- CN202510416095.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-03
- Publication Date
- 2025-07-04
AI Technical Summary
Traditional hearing aids are difficult to clearly understand dialogue and cross-language communication in noisy environments, and cannot effectively deal with complex noise and language barriers, affecting daily communication and international communication among hearing impaired people.
Design a cross-language hearing aid with voice translation functions, integrate multilingual voice acquisition, intelligent noise processing, real-time voice translation and parameter adjustment units, and use microphone arrays, convolutional neural networks and Bluetooth connections to achieve background noise suppression, echo cancellation, voice enhancement and language translation.
Clearly hear surrounding sounds in various noise environments and translate different languages in real time, improving the communication ability of hearing impaired people, enhancing social interaction and quality of life, and providing customized hearing compensation and convenient operation.
Smart Images

Figure CN120264208A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of speech translation, and more specifically, to a cross-language hearing aid with speech translation function. Background Art
[0002] Hearing loss affects a large number of people globally, making it challenging for them to communicate with others in daily life. Although traditional hearing aids can help wearers improve their ability to perceive sounds, they have limited support for clearly understanding conversations in noisy environments or for cross-language communication. Against the backdrop of increasing globalization, the need for cross-language communication is growing. For people with hearing impairments, they not only need to solve hearing problems but also have to deal with the additional challenges brought about by language barriers, especially in international travel, business communication, or multicultural environments, where language barriers can be a significant issue. Even in the same language environment, background noise in the environment (such as street noise, conversations in a restaurant, etc.) can seriously interfere with the normal communication of people with hearing impairments. Traditional noise reduction technologies may not be sufficient to effectively handle complex noise environments. Therefore, a cross-language hearing aid with speech translation function is designed. Summary of the Invention
[0003] The purpose of the present invention is to provide a cross-language hearing aid with speech translation function to solve the problems of hearing impairment, communication difficulties, cross-language communication barriers, and environmental noise interference mentioned in the above background art.
[0004] To achieve the above object, the present invention provides a cross-language hearing aid with speech translation function, including:
[0005] A multi-lingual speech acquisition unit that directionally acquires ambient speech and converts it into a digital signal, supporting multi-language input;
[0006] An intelligent processing unit under noise that suppresses background noise, eliminates echoes, enhances the clarity of ambient speech based on a convolutional neural network model, and introduces an adaptive step size for optimization during the process of echo cancellation;
[0007] A real-time speech translation unit that real-time converts ambient speech into text, translates it into the target language, and then synthesizes it into the target speech;
[0008] A parameter adjustment unit that dynamically adjusts sound parameters according to the user's hearing profile and outputs the translation result through a speaker;
[0009] A Bluetooth connection unit that transmits the translated speech stream and text data to the mobile phone through low-power Bluetooth and supports two-way control instructions.
[0010] As a further improvement of this technical solution, the multilingual speech acquisition unit includes a speech acquisition module and a speech processing module;
[0011] Among them, the speech acquisition module collects a speech data set containing various noise environments through a microphone array, and simultaneously records the corresponding noise-free speech as a training label;
[0012] The speech processing module divides the audio signal of the speech data set into short-time frames, reduces spectral leakage through a Hamming window, performs a fast Fourier transform on each frame of the signal to generate an amplitude spectrum and a phase spectrum, converts the amplitude spectrum into a Mel-scale filter bank to generate a Mel spectrogram, compresses high-frequency information and retains the key features of the speech.
[0013] As a further improvement of this technical solution, the intelligent processing unit under noise includes a noise reduction module, an echo cancellation module and a speech enhancement module;
[0014] Among them, the noise reduction module receives the Mel spectrogram and uses a convolutional neural network model to suppress background noise;
[0015] The echo cancellation module eliminates the hearing aid feedback noise;
[0016] The speech enhancement module uses beamforming technology to enhance the clarity of ambient speech.
[0017] As a further improvement of this technical solution, the noise reduction module receives the Mel spectrogram and uses a convolutional neural network model to suppress background noise, including the following steps:
[0018] S1.1, Receive the Mel spectrogram of the noisy speech as input;
[0019] S1.2, Design a convolutional neural network architecture for the noise reduction task;
[0020] S1.3, Output the predicted Mel spectrogram after noise reduction;
[0021] S1.4, Use the mean square error to measure the difference between the predicted Mel spectrogram and the noise-free speech;
[0022] S1.5, Use the speech data set to train the convolutional neural network. In each iteration, the model receives the spectrogram of the noisy audio and tries to predict the corresponding noise-free spectrogram;
[0023] S1.6, Real-time collect the audio stream of the ambient speech, segment it into small segments, extract the Mel spectrogram features, and transmit them to the trained convolutional neural network model for prediction.
[0024] As a further improvement of this technical solution, the echo cancellation module eliminates the hearing aid feedback noise, including the following steps:
[0025] S1.7. Collect audio data in the hearing aid usage scenario, and label the audio data to distinguish the target speech, background noise, and echo signal;
[0026] S1.8. Calculate the time difference between the signals received by the microphone array and the signals output by the speaker by comparing them;
[0027] S1.9. Analyze the amplitude characteristics of the echo signal to determine its intensity relative to the target speech;
[0028] S1.10. Use the speaker output signal as the reference signal and the signal collected by the microphone array as the input signal;
[0029] S1.11. Output the filtered signal through the normalized least mean square error. To address the problem of slow convergence speed of echo cancellation caused by the fixed step size in the traditional normalized least mean square error algorithm, an adaptive step size μ(n) is introduced in the normalized least mean square error for optimization.
[0030] As a further improvement of this technical solution, in S1.11, the normalized least mean square error is:
[0031]
[0032] where represents the filtered output signal; w k (n) represents the weight coefficient of the adaptive filter; k represents the weight index; n represents the time step; x(n - k) represents the delayed version of the input signal; k represents the time delay of the signal; N represents the length of the filter;
[0033] To address the problem of slow convergence speed of echo cancellation caused by the fixed step size in the traditional normalized least mean square error algorithm, an adaptive step size μ(n) is introduced in the normalized least mean square error for optimization:
[0034]
[0035] where represents the optimized output signal; e(n) represents the difference between the target signal and the filter output; β represents the learning rate (taking 0 < β ≤ 1); ε represents a positive number.
[0036] As a further improvement of this technical solution, the speech enhancement module uses beamforming technology to improve the clarity of ambient speech, including the following steps:
[0037] S1.12. Use the direction-of-arrival estimation algorithm to determine the direction of the target speech source;
[0038] S1.13. Design an adaptive beamformer to enhance the sound in the direction of the target voice;
[0039] S1.14. For each collection window, apply a beamforming algorithm to process the data of the microphone array;
[0040] S1.15. Evaluate the effect of beamforming and make adjustments according to the feedback.
[0041] As a further improvement of this technical solution, the real-time speech translation unit includes a speech recognition module, a machine translation module, and a speech synthesis module;
[0042] Among them, the speech recognition module preprocesses the speech processed by the intelligent processing unit under noise, extracts features and converts the speech into text using a long short-term memory network;
[0043] After receiving the text generated by the speech recognition module, the machine translation module translates the text from the source language to the target language through a neural machine translation model;
[0044] The speech synthesis module uses natural language processing and speech synthesis technologies to synthesize the target language text output by the machine translation module into speech.
[0045] As a further improvement of this technical solution, the parameter adjustment unit dynamically adjusts the sound parameters according to the user's hearing profile and outputs the translation result through a speaker, including the following steps:
[0046] S2.1. Read the user's hearing profile and apply gains in different frequency ranges to compensate for the user's hearing loss;
[0047] S2.2. Calculate the loudness of the current audio. If it exceeds the user's comfort range, use dynamic range compression technology for dynamic compression;
[0048] S2.3. Adjust the frequency response curve of the audio signal to better meet the user's hearing needs.
[0049] As a further improvement of this technical solution, the Bluetooth connection unit includes a Bluetooth pairing module and a mobile phone control module;
[0050] Among them, the Bluetooth pairing module establishes a wireless connection between the hearing aid and devices such as mobile phones, and uses low-power Bluetooth technology for device search, authentication, and automatic reconnection;
[0051] The mobile phone control module allows the user to remotely manage the functions of the hearing aid through a mobile phone APP, supports real-time voice stream transmission, sends the translated voice to the mobile phone side, and at the same time receives the voice input from the mobile phone side for hearing aid playback.
[0052] Advantages of the present invention compared with the prior art:
[0053] 1. In this cross - language hearing aid with voice translation function, by integrating functions such as multi - language voice collection, intelligent noise processing, and real - time voice translation, the hearing aid can not only help users clearly hear the surrounding sounds in various noisy environments, but also convert voices in different languages into the target language in real - time. This greatly improves the communication ability of hearing - impaired people with those using different languages, enabling them to communicate more freely in both daily life and international travel, enhancing their social interaction ability and quality of life.
[0054] 2. In this cross - language hearing aid with voice translation function, it can dynamically adjust the sound parameters according to the user's hearing loss situation, and enhance the voice or translation result through the speaker or bone conduction output, providing a customized hearing compensation plan for users. At the same time, it supports wireless connection with devices such as mobile phones, allowing users to conveniently perform various personalized settings and controls through the mobile phone APP, such as volume adjustment, noise reduction level selection, language translation mode switching, etc. The operation is simple and the functions are powerful, further improving the user's convenience and satisfaction. In addition, the secure pairing and data transmission encryption provided by the Bluetooth module also ensure the user's privacy and data security. BRIEF DESCRIPTION OF THE DRAWINGS
[0055] Figure 1 is the overall flow block diagram of the present invention;
[0056] The meanings of each label in the figure are as follows:
[0057] 1. Multi - language voice collection unit; 11. Voice collection module; 12. Voice processing module; 2. Intelligent processing unit under noise; 21. Noise reduction module; 22. Echo cancellation module; 23. Voice enhancement module; 3. Real - time voice translation unit; 31. Voice recognition module; 32. Machine translation module; 33. Voice synthesis module; 4. Parameter adjustment unit; 5. Bluetooth connection unit; 51. Bluetooth pairing module; 52. Mobile phone control module. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0058] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0059] Embodiment: Please refer to Figure 1 as shown, a cross - language hearing aid with voice translation function is provided, including.
[0060] The multilingual voice acquisition unit 1 directionally acquires ambient voice and converts it into a digital signal, supporting multilingual input;
[0061] In this embodiment, the multilingual voice acquisition unit 1 includes a voice acquisition module 11 and a voice processing module 12;
[0062] Among them, the voice acquisition module 11 collects a voice data set containing various noise environments (such as streets, restaurants, traffic) through a microphone array, and simultaneously records the corresponding noise-free voice as a training label;
[0063] The voice processing module 12 divides the audio signal of the voice data set into short-time frames, reduces spectral leakage through a Hamming window (the Hamming window is a window function used in signal processing, which improves the frequency-domain representation when the signal is truncated by reducing spectral leakage, and is commonly used for frame-by-frame operations on signals in speech and audio processing), performs a fast Fourier transform on each frame of the signal (the Fourier transform is a mathematical tool used to convert a time-domain signal into a frequency-domain representation for analyzing the frequency components of the signal), generates an amplitude spectrum and a phase spectrum, converts the amplitude spectrum into a Mel-scale filter bank, generates a Mel spectrogram (the Mel spectrogram is a method of representing the frequency of sound converted based on the characteristics of human auditory perception, which converts linearly spaced frequencies into a non-uniform Mel scale to better reflect the perceptual sensitivity of the human ear to sounds of different frequencies), compresses high-frequency information and retains the key features of the voice.
[0064] The intelligent processing unit 2 under noise suppresses background noise, eliminates echoes based on a convolutional neural network model, enhances the clarity of ambient voice, and introduces an adaptive step size for optimization during the process of eliminating echoes;
[0065] In this embodiment, the intelligent processing unit 2 under noise includes a noise reduction module 21, an echo cancellation module 22, and a voice enhancement module 23;
[0066] Among them, the noise reduction module 21 receives the Mel spectrogram and suppresses background noise using a convolutional neural network model. It can effectively distinguish various sounds in the environment and reduce or eliminate those sounds that are not desired to be heard;
[0067] The echo cancellation module 22 eliminates the hearing aid feedback noise, thereby providing a more natural and smooth auditory experience;
[0068] The voice enhancement module 23 uses beamforming technology to enhance the clarity of ambient voice.
[0069] Among them, the noise reduction module 21 receives the Mel spectrogram and suppresses background noise using a convolutional neural network model, including the following steps:
[0070] S1.1, Receive the Mel spectrogram of the noisy voice as input;
[0071] S1.2. Design a convolutional neural network architecture for noise reduction tasks, including multiple convolutional layers (shallow convolution: using small convolutional kernels (3×3) to extract local time-frequency features (speech formants, noise impulses), deep convolution: capturing global noise patterns (continuous background noise, burst noise) through deep networks), activation functions (using ReLU to enhance non-linear expression ability), pooling layers (interspersed with max pooling to reduce dimensions and retain significant features), and fully connected layers;
[0072] S1.3. Output the predicted Mel spectrogram after noise reduction;
[0073] S1.4. Use the mean squared error (MSE) to measure the difference between the predicted Mel spectrogram and the noise-free speech. The mean squared error (MSE) quantifies the difference between the two by calculating the sum of the squares of the differences between the predicted Mel spectrogram and the Mel spectrogram of the noise-free speech and then averaging, so as to evaluate the prediction accuracy of the model;
[0074] S1.5. Use the speech dataset to train the convolutional neural network. In each iteration, the model receives the spectrogram of the noisy audio and attempts to predict the corresponding noise-free spectrogram;
[0075] S1.6. Real-time collect the audio stream of the ambient speech, segment it into small segments (each segment is 20 ms), extract the Mel spectrogram features, and transmit them to the trained convolutional neural network model for prediction.
[0076] Furthermore, the echo cancellation module 22 cancels the hearing aid feedback noise, including the following steps:
[0077] S1.7. Collect the audio data in the hearing aid usage scenario, analyze the possible echo sources, and label the audio data to distinguish the target speech, background noise, and echo signals;
[0078] S1.8. Calculate the time difference between the signals received by the microphone array and the signals output by the speaker by comparing them;
[0079] S1.9. Analyze the amplitude characteristics of the echo signal to determine its intensity relative to the target speech;
[0080] S1.10. Use the speaker output signal as the reference signal and the microphone array collected signal as the input signal;
[0081] S1.11. Output the filtered signal through normalized least mean square error, where the echo component is significantly weakened. To address the problem of slow convergence speed of echo cancellation in the traditional normalized least mean square error algorithm caused by a fixed step size, an adaptive step size μ(n) is introduced in the normalized least mean square error for optimization;
[0082] The normalized least mean square error is as follows:
[0083]
[0084] Wherein, represents the output signal after filtering, which is the final output result of the echo cancellation module 22, and the echo component has been weakened; w k (n) represents the weight coefficient of the adaptive filter; k represents the weight index, indicating the position of the filter tap; n represents the time step (current moment); x(n - k) represents the delayed version of the input signal; k represents the time delay of the signal; N represents the length of the filter (i.e., the number of filter taps);
[0085] The fixed step size of the traditional least mean square error (LMS) algorithm may lead to a slow convergence speed, resulting in a long time for the hearing aid to effectively cancel echoes. By introducing an adaptive step size and dynamically adjusting the step size according to the power of the input signal, the echo cancellation can converge faster and the real-time performance can be enhanced; the power of environmental noise and speech signals will change continuously. The algorithm with a fixed step size may adjust too fast at low-energy inputs (causing the filter to diverge) or too slow at high-energy inputs (inadequate echo cancellation); the normalized step size is normalized by the input signal energy, automatically adapting to signal amplitude changes and preventing the algorithm from failing;
[0086] To address the problem of slow echo cancellation convergence speed caused by the fixed step size in the traditional normalized least mean square error algorithm, an adaptive step size μ(n) is introduced into the normalized least mean square error for optimization:
[0087]
[0088] Wherein, represents the optimized output signal; e(n) represents the difference between the target signal and the filter output; β represents the learning rate (0 < β ≤ 1); ε represents a small positive number (used to avoid numerical instability caused by too small a denominator).
[0089] Furthermore, the speech enhancement module 23 uses beamforming technology to improve the clarity of ambient speech, including the following steps:
[0090] S1.12. Use a direction-of-arrival estimation algorithm to determine the direction of the target speech source;
[0091] S1.13. Design an adaptive beamformer to enhance the sound in the direction of the target speech (An adaptive beamformer is a technology used to enhance signals in a specific direction and is commonly applied in microphone array systems. Its main purpose is to extract the sound source of interest from the mixed signals (such as various sounds in the environment) while suppressing interference noise and echoes from other directions. This technology creates a "virtual auditory direction" by adjusting the weights and delays of the signals received by each microphone, as if forming a "beam" in space that can receive sounds directionally. When this "beam" is directed towards the direction of the target sound source, the sound of that source is amplified while the sounds from other directions are attenuated);
[0092] S1.14. For each collection window (usually from a few milliseconds to dozens of milliseconds), apply the beamforming algorithm to process the data of the microphone array. The beamforming algorithm adjusts the amplitudes and phases of the signals received by each microphone in the microphone array to enhance the sound signals from a specific direction while suppressing interference noise from other directions;
[0093] S1.15. Evaluate the effect of beamforming and make adjustments according to the feedback (Regularly check the performance of the beamformer and evaluate it using the signal-to-noise ratio (SNR); Adjust the design parameters of the beamformer according to the evaluation results to continuously improve the system performance).
[0094] The real-time speech translation unit 3 converts the ambient speech into text in real time, translates it into the target language, and then synthesizes it into the target speech;
[0095] In this embodiment, the real-time speech translation unit 3 includes a speech recognition module 31, a machine translation module 32, and a speech synthesis module 33;
[0096] Among them, the speech recognition module 31 preprocesses the speech processed by the intelligent processing unit 2 under noise, including noise reduction and enhancement, extracts features (features include Mel-frequency cepstral coefficients, spectral features, etc.) and converts the speech into text using a long short-term memory network;
[0097] After receiving the text generated by the speech recognition module 31, the machine translation module 32 translates the text from the source language to the target language through a neural machine translation model. The neural machine translation model uses deep learning technology to convert the source language text into the target language text through an encoder-decoder architecture to achieve automatic translation;
[0098] The speech synthesis module 33 adopts natural language processing and speech synthesis technologies (natural language processing technology involves methods for enabling computers to understand, interpret, and generate human language, while speech synthesis technology is a tool for converting text information into audible speech output. The combination of the two can achieve automatic conversion from text to speech). It synthesizes the target language text output by the machine translation module 32 into speech. By analyzing the emotion and rhythm of the text and using high-quality vocoder technology, this module generates clear and easy-to-understand target speech, ultimately realizing the real-time speech-to-speech translation function.
[0099] The parameter adjustment unit 4 dynamically adjusts the sound parameters according to the user's hearing profile and outputs the translation result through the speaker.
[0100] In this embodiment, the parameter adjustment unit 4 dynamically adjusts the sound parameters according to the user's hearing profile and outputs the translation result through the speaker, including the following steps:
[0101] S2.1. Read the user's hearing profile and apply gain in different frequency ranges to compensate for the user's hearing loss.
[0102] S2.2. Calculate the loudness of the current audio. If it exceeds the user's comfortable range, use dynamic range compression technology for dynamic compression to make weak sounds more audible without making strong sounds too harsh, which is particularly important for users with a small dynamic range of hearing.
[0103] S2.3. Adjust the frequency response curve of the audio signal to make it more in line with the user's hearing needs. Some users may require sounds in specific frequency ranges to be amplified or attenuated.
[0104] The Bluetooth connection unit 5 transmits the translated speech stream and text data to the mobile phone via low-power Bluetooth and supports two-way control commands.
[0105] In this embodiment, the Bluetooth connection unit 5 includes a Bluetooth pairing module 51 and a mobile phone control module 52.
[0106] Among them, the Bluetooth pairing module 51 establishes a wireless connection between the hearing aid and devices such as mobile phones. It uses low-power Bluetooth technology for device search, authentication, and automatic reconnection. When using it for the first time, the user can discover the hearing aid device through Bluetooth scanning, input the pairing code or directly confirm the connection request to complete the pairing. This module supports multi-device management, allows the user to switch between different devices, and provides encrypted transmission to ensure data security and prevent unauthorized device connections.
[0107] The mobile phone control module 52 allows users to remotely manage the functions of the hearing aid through the mobile phone APP, including personalized settings such as volume adjustment, noise reduction level selection, and language translation mode switching. This module also supports real-time voice stream transmission, sending the translated voice or optimized ambient sound to the mobile phone side, while receiving the voice input from the mobile phone side for hearing aid playback. In addition, users can view the battery status and connection status through the APP and enable the lost reminder function to prevent the hearing aid from being lost.
[0108] The above shows and describes the basic principles, main features and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited by the above embodiments. The above embodiments and descriptions in the specification are only preferred examples of the present invention and are not used to limit the present invention. Without departing from the spirit and scope of the present invention, the present invention will have various changes and improvements, and these changes and improvements all fall within the scope of the present invention claimed.
Claims
1. A cross-language hearing aid with a voice translation function, characterized in that Comprising: A multilingual voice collection unit (1), which directionally collects ambient voice and converts it into a digital signal, supporting multilingual input; An intelligent processing unit under noise (2), which suppresses background noise and eliminates echo based on a convolutional neural network model, enhances the clarity of ambient voice, and introduces an adaptive step size for optimization during the process of eliminating echo; A real-time voice translation unit (3), which converts ambient voice into text in real time, translates it into the target language, and then synthesizes it into the target voice; A parameter adjustment unit (4), which dynamically adjusts sound parameters according to the user's hearing profile and outputs the translation result through a speaker; A Bluetooth connection unit (5), which transmits the translated voice stream and text data to the mobile phone side through low-power Bluetooth and supports two-way control instructions.
2. The cross-language hearing aid with voice translation function according to claim 1, wherein: The multilingual voice collection unit (1) includes a voice collection module (11) and a voice processing module (12); Among them, the voice collection module (11) collects a voice data set containing various noise environments through a microphone array, and simultaneously records the corresponding noise-free voice as a training label; The voice processing module (12) divides the audio signal of the voice data set into short-time frames, reduces spectral leakage through a Hamming window, performs a fast Fourier transform on each frame of signal to generate an amplitude spectrum and a phase spectrum, converts the amplitude spectrum into a Mel-scale filter bank to generate a Mel spectrogram, compresses high-frequency information and retains the key features of the voice.
3. The cross-language hearing aid with voice translation function according to claim 2, wherein: The intelligent processing unit under noise (2) includes a noise reduction module (21), an echo cancellation module (22) and a voice enhancement module (23); Among them, the noise reduction module (21) receives the Mel spectrogram and suppresses background noise using a convolutional neural network model; The echo cancellation module (22) eliminates the hearing aid feedback noise; The voice enhancement module (23) uses beamforming technology to enhance the clarity of ambient voice.
4. The cross-language hearing aid with voice translation function according to claim 3, wherein: The noise reduction module (21) receives the Mel spectrogram and suppresses background noise using a convolutional neural network model, including the following steps: S1.1, Receive the Mel spectrogram of the noisy voice as input; S1.2, Design a convolutional neural network architecture for the noise reduction task; S1.3, Output the predicted Mel spectrogram after noise reduction; S1.4, Use the mean square error to measure the difference between the predicted Mel spectrogram and the noise-free voice; S1.5, Use the voice data set to train the convolutional neural network. In each iteration, the model receives the spectrogram of the noisy audio and tries to predict the corresponding noise-free spectrogram; S1.6, Real-time collect the audio stream of the ambient voice, divide it into small segments, extract the Mel spectrogram features, and transmit them to the trained convolutional neural network model for prediction.
5. The cross-language hearing aid with voice translation function according to claim 4, characterized in that: The echo cancellation module (22) eliminates the hearing aid feedback noise, including the following steps: S1.7, Collect audio data in the hearing aid usage scenario, and label the audio data to distinguish the target voice, background noise and echo signal; S1.
8. Calculate the time difference between the signals received by the microphone array and the signals output by the speaker by comparing them. S1.
9. Analyze the amplitude characteristics of the echo signal to determine its intensity relative to the target speech. S1.
10. Use the speaker output signal as the reference signal and the signals collected by the microphone array as the input signals. S1.
11. Output the filtered signal through the normalized least mean square error. To address the problem of slow convergence speed of echo cancellation caused by the fixed step size in the traditional normalized least mean square error algorithm, an adaptive step size μ(n) is introduced in the normalized least mean square error for optimization.
6. The cross-language hearing aid with voice translation function according to claim 5, characterized in that: In S1.11, the normalized least mean square error is: Among them, represents the output signal after filtering; w k (n) represents the weight coefficient of the adaptive filter; k represents the weight index; n represents the time step; x(n - k) represents the delayed version of the input signal; k represents the time delay of the signal; N represents the length of the filter; To address the problem of slow convergence speed of echo cancellation caused by the fixed step size in the traditional normalized least mean square error algorithm, an adaptive step size μ(n) is introduced in the normalized least mean square error for optimization: Among them, represents the optimized output signal; e(n) represents the difference between the target signal and the filter output; β represents the learning rate (0 < β ≤ 1); ε represents a positive number.
7. The cross-language hearing aid with voice translation function according to claim 6, characterized in that: The voice enhancement module (23) uses beamforming technology to enhance the clarity of ambient speech, including the following steps: S1.
12. Use the direction of arrival estimation algorithm to determine the direction of the target speech source. S1.
13. Design an adaptive beamformer to enhance the sound in the direction of the target speech. S1.
14. For each collection window, apply the beamforming algorithm to process the data of the microphone array. S1.
15. Evaluate the effect of beamforming and make adjustments according to the feedback.
8. The cross-language hearing aid with voice translation function according to claim 7, characterized in that: The real-time speech translation unit (3) includes a speech recognition module (31), a machine translation module (32), and a speech synthesis module (33). Among them, the speech recognition module (31) preprocesses the speech processed by the intelligent processing unit (2) under noise, extracts features, and converts the speech into text using a long short-term memory network. After receiving the text generated by the speech recognition module (31), the machine translation module (32) translates the text from the source language to the target language through a neural machine translation model. The speech synthesis module (33) uses natural language processing and speech synthesis technologies to synthesize the target language text output by the machine translation module (32) into speech.
9. The cross-language hearing aid with voice translation function according to claim 8, characterized in that: The parameter adjustment unit (4) dynamically adjusts the sound parameters according to the user's hearing profile and outputs the translation result through the speaker, including the following steps: S2.
1. Read the user's hearing profile and apply gains in different frequency ranges to compensate for the user's hearing loss. S2.
2. Calculate the loudness of the current audio. If it exceeds the user's comfort range, use dynamic range compression technology for dynamic compression. S2.
3. Adjust the frequency response curve of the audio signal to better meet the user's hearing needs.
10. The cross-language hearing aid with voice translation function according to claim 9, characterized in that: The Bluetooth connection unit (5) includes a Bluetooth pairing module (51) and a mobile phone control module (52). Among them, the Bluetooth pairing module (51) establishes a wireless connection between the hearing aid and devices such as mobile phones, and uses low-power Bluetooth technology for device search, authentication, and automatic reconnection. The mobile phone control module (52) allows users to remotely manage the hearing aid functions through a mobile phone APP, supports real-time voice stream transmission, sends the translated speech to the mobile phone side, and at the same time receives the voice input from the mobile phone side for hearing aid playback.