Hearing aid enhancement method, device and system
Hearing aid enhancement devices, by receiving, storing, and replaying speech information, combined with artificial intelligence and microphone array technology, solve the problem of inaccurate speech pickup in noisy environments, improve speech clarity and user experience, and provide multifunctional hearing aid devices.
Patent Information
- Application Number
- CN202511291622.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-10
- Publication Date
- 2026-01-23
AI Technical Summary
Hearing aids have difficulty accurately picking up and providing the required sound in noisy environments, resulting in inaccurate speech pickup during interactions between hearing-impaired individuals and failing to meet their needs.
The hearing aid enhancement device receives and stores voice information, responds to the user's re-hearing command, replays and enhances the voice information, optimizes and adjusts parameters using artificial intelligence models, and provides directional sound pickup and voice positioning in combination with microphone arrays and displays, so as to achieve clear playback and personalized optimization of voice information.
It improves the accuracy of voice interaction and user experience, meets the voice needs of different scenarios, enhances voice clarity and auditory experience, and provides a multifunctional hearing aid device.
Smart Images

Figure CN121397443A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of computer technology, and more particularly to hearing aid enhancement methods, devices, and systems. Background Technology
[0002] Hearing aids are indispensable devices for people with hearing impairments. They enable them to hear sounds from the outside world just like everyone else.
[0003] However, real-world human interaction environments are often complex. For example, people wearing hearing aids might go to noisy public places like markets or train stations. In such situations, it's difficult for people to accurately pick up sound during interaction, which in turn prevents hearing aids from providing the necessary sound accurately. Summary of the Invention
[0004] This disclosure provides hearing aid enhancement methods, devices, and systems.
[0005] According to a first aspect of this disclosure, a hearing aid enhancement method is provided. The method specifically includes: receiving and storing speech information; in response to a hearing enhancement command triggered by a user through headphones or a hearing aid enhancement device, determining at least a portion of the speech information to be replayed; and replaying the at least portion of the speech information in an agreed manner.
[0006] Based on the above, it is clear that the replayed audio information reflects the user's interests or hearing difficulties. It encompasses speech features in specific scenarios (such as speech in noisy environments or specific pronunciation patterns), thus providing a high-quality data foundation for subsequent voiced / unvoiced sound differentiation, formant optimization, and personalized parameter adjustments. In this way, the system can dynamically learn the user's preferences and hearing characteristics, continuously optimizing speech enhancement strategies and improving the user experience.
[0007] According to at least one embodiment of this disclosure, at least a portion of the voice information is replayed in a pre-agreed manner, including: replaying the target voice corresponding to the voice information within a pre-agreed time range.
[0008] Based on the above, it can be seen that if a user did not hear the audio clearly, the target audio can be replayed, thus meeting the user's needs.
[0009] According to at least one embodiment of this disclosure, replaying at least a portion of voice information in a pre-agreed manner includes: performing sentence segmentation recognition on stored voice information; and determining and replaying at least one statement that last responded to the user based on the sentence segmentation recognition result.
[0010] Based on the above, it can be seen that the system can accurately identify the audio content that the user wants to replay, which can not only meet the user's need to listen again, but also avoid replaying too much content and causing the user to miss the latest content of the speaker and cause the inconvenience of replaying it again.
[0011] According to at least one embodiment of this disclosure, replaying at least a portion of voice information in a pre-agreed manner includes: performing sentence segmentation recognition on stored voice information; determining at least one statement from the last responding user based on the sentence segmentation recognition result; determining the statement type of the at least one statement from the last responding user; if the statement type is an interrogative sentence, starting a countdown; and if the user does not reply after the countdown ends, replaying the at least one statement from the last responding user.
[0012] Based on the above, it can be seen that the system can determine the sentence type of the last sentence in the captured audio information without requiring any replay selection from the user. If it is a question, it usually indicates that a user response is needed. If the hearing aid enhancement device does not capture the user's response, it may be because the user did not hear the question clearly. Therefore, by combining the judgment of sentence type and the capture of the user's response, the system can automatically replay the audio information, effectively improving the user experience.
[0013] According to at least one embodiment of this disclosure, after receiving and storing voice information, the method further includes: parsing the received voice information to determine the voiced information contained in the voice information; enhancing the voiced information using initial adjustment parameters to generate target voice carrying enhanced voice; and sending the target voice to the user's headphone device, or playing it through a speaker integrated with the microphone.
[0014] Based on the aforementioned publicly available information, after receiving speech information using a hearing aid enhancement device, the speech information can be further analyzed to obtain the voiced information contained within the speech information. Subsequently, initial adjustment parameters can be used to enhance the voiced information in the received speech information, resulting in target speech carrying enhanced voiced sounds. This target speech is then sent to the user's headset device. In this solution, the hearing aid enhancement device performs targeted enhancement processing on voiced sounds. The hearing aid enhancement device has a stronger speech enhancement processing capability than hearing aid headsets (i.e., hearing aids in conventional solutions that do not have hearing aid enhancement devices), thus improving the effectiveness of hearing aid use.
[0015] According to at least one embodiment of this disclosure, after receiving and storing speech information, the method further includes: selecting training samples from the stored speech information; distinguishing between voiced and unvoiced sounds in the training samples to determine the voiced portion contained in the speech information; inputting initial adjustment parameters as training samples into the optimization model, and fine-tuning the optimization model based on user feedback to generate target adjustment parameters for optimizing the formants of the voiced portion.
[0016] Based on the above, it is clear that using artificial intelligence models (i.e., optimization models) can enable rapid optimization and adjustment of parameters, thereby generating target adjustment parameters that meet the user's current needs. This not only effectively improves the efficiency of parameter optimization and adjustment but also meets the parameter adjustment needs of users in diverse scenarios.
[0017] According to at least one embodiment of this disclosure, enhancing voiced information to generate target speech carrying enhanced voiced sounds includes: using a model to perform text recognition on the speech information carrying enhanced voiced sounds; correcting the voiced text corresponding to the voiced information based on the context content in the text recognition result; and converting the corrected text recognition result into target speech.
[0018] Based on the above, it can be seen that after first performing voiced enhancement processing on the received speech information, text recognition and correction are then performed. The corrected text recognition result is then converted into target semantics, thereby effectively improving the playback accuracy of the hearing aid. Another approach is to directly utilize the context content obtained from the speech information conversion to correct the voiced text, and then perform enhancement processing, which can also effectively improve the playback accuracy of the hearing aid.
[0019] According to at least one embodiment of the present disclosure, the hearing aid enhancement device is configured with a display screen; after generating target speech carrying enhanced voiced sounds, the device further includes: converting the target speech into text and displaying it on the display screen.
[0020] Based on the above, it is clear that this allows users to intuitively see the content of voice interactions and related prompts. For example, if a project is mentioned during a voice interaction, information about that project's meetings and progress will be displayed.
[0021] According to at least one embodiment of this disclosure, the hearing aid enhancement device is equipped with calendar and alarm clock functions; calendar reminder information and / or alarm clock reminder information are set according to the user's voice commands.
[0022] Based on the above, it can be seen that this hearing aid enhancement device with a display screen can also display a calendar, allowing users to record to-do items and receive reminders. It can also be used as an alarm clock. Because the hearing aid enhancement device has voice pickup capabilities, when used as a calendar, users can control it via voice to set relevant tasks, reminder content, and reminder times. This better meets the diverse needs of users. Therefore, the hearing aid enhancement device not only helps users hear speech clearly but also serves as a notebook, calendar, and other tools to improve users' daily work efficiency.
[0023] According to at least one embodiment of the present disclosure, the hearing aid enhancement device includes a microphone array; upon receiving voice information, the microphone array is used to locate the target object emitting the voice information.
[0024] Based on the above, it can be seen that this positioning capability not only helps hearing aids better separate target speech from background noise, but also adaptively adjusts the pickup direction according to the user's needs, thereby significantly improving speech clarity and the user's auditory experience, especially in noisy environments or multi-person conversation scenarios.
[0025] According to at least one embodiment of this disclosure, the method further includes: parsing the replayed audio information during replay to determine voiced information contained in the audio information; enhancing the voiced information using initial adjustment parameters to generate target audio carrying enhanced voiced information; and sending the target audio to the user's headphone device or playing it through a speaker integrated with the microphone.
[0026] Based on the aforementioned publicly available information, after receiving the voice information to be replayed using a hearing aid enhancement device, the voice information can be further analyzed to obtain the voiced information contained within the voice information. Subsequently, initial adjustment parameters can be used to enhance the voiced information in the received voiced information, resulting in a target voiced speech being generated after enhancement. This target voiced speech is then sent to the user's headset. In this solution, the hearing aid enhancement device performs targeted enhancement of voiced information, giving it a stronger voice enhancement capability than hearing aid headsets (i.e., headset-type hearing aids in conventional solutions that do not have a hearing aid enhancement device), thus improving the effectiveness of hearing aid replay.
[0027] According to a second aspect of this disclosure, a hearing aid device is provided, comprising: a memory storing execution instructions; and a processor executing the execution instructions stored in the memory, such that the processor performs the method described in the first aspect of any embodiment of this disclosure.
[0028] According to a third aspect of this disclosure, a pair of glasses with hearing aid functionality is provided, the glasses being equipped with: a hearing aid enhancement device and headphones; wherein the hearing aid enhancement device is used to perform the method described in any of the first aspects, and the headphones are used to play voice information.
[0029] According to a fourth aspect of this disclosure, a hearing aid enhancement system is provided, comprising: a hearing aid enhancement device with a microphone, an earphone, a cloud computer and / or a mobile phone, wherein the microphone is used to acquire speech information, and the earphone is used to play speech information; the hearing aid enhancement device is connected to the earphone and / or the mobile phone via Bluetooth communication, and the cloud computer or the mobile phone or the hearing aid enhancement device is used to perform the method of any of the first aspects. Attached Figure Description
[0030] The accompanying drawings illustrate exemplary embodiments of the present disclosure and, together with the description thereof, serve to explain the principles of the present disclosure. These drawings are included to provide a further understanding of the present disclosure and are incorporated in and constitute a part of this specification.
[0031] Figure 1 This is a flowchart illustrating a hearing aid enhancement method provided in this disclosure.
[0032] Figure 2 This is a flowchart illustrating the enhancement method exemplified in this disclosure.
[0033] Figure 3 The hearing aid enhancement system provided in this disclosure.
[0034] Figure 4 This is a schematic diagram illustrating how a hearing aid enhancement device is worn, as exemplified in this disclosure.
[0035] Figure 5 This is a schematic diagram of the real-time audio graph provided in this disclosure.
[0036] Figure 6 A schematic diagram of resonance peaks provided in an embodiment of this disclosure.
[0037] Figure 7 This is a schematic block diagram of the structure of a hearing aid enhancement device according to one embodiment of the present disclosure.
[0038] Figure 8 This is a schematic diagram illustrating the structure of eyeglasses with hearing aid function as an example of this disclosure.
[0039] Figure 9 This is a schematic block diagram of an electronic device according to one embodiment of the present disclosure. Detailed Implementation
[0040] The present disclosure will now be described in further detail with reference to the accompanying drawings and examples. It should be understood that the specific examples described herein are for illustrative purposes only and are not intended to limit the scope of the disclosure. Furthermore, it should be noted that, for ease of description, only the parts relevant to the present disclosure are shown in the accompanying drawings.
[0041] It should be noted that, where there is no conflict, the embodiments and features described in this disclosure can be combined with each other. The technical solutions of this disclosure will now be described in detail with reference to the accompanying drawings and embodiments.
[0042] In hearing aid applications, real-world interaction environments are often complex. For example, hearing aid wearers might be in noisy public places like markets or train stations. In such situations, accurate sound pickup is difficult, preventing the hearing aid from accurately providing the desired sound. Furthermore, there are differences in hearing abilities among hearing aid wearers and subtle variations between different batches of hearing aids. Therefore, a solution that accurately provides sound to users is urgently needed to better meet the diverse needs of hearing aid users.
[0043] To facilitate description and make the technical solutions of the specific embodiments of this disclosure easier to understand, the technical terms involved in the specific embodiments of this disclosure will be explained as follows before describing the hearing aid enhancement method implemented in this disclosure.
[0044] Initial adjustment parameters: These refer to the parameters required to enhance the output sound of the hearing aid enhancement device. By setting appropriate adjustment parameters, the picked-up sound is enhanced, allowing the user to hear the sound more clearly. Enhancement parameters can include equalizer parameters, filter parameters, dynamic range compression parameters, etc.
[0045] Formants are key acoustic features in speech signals, reflecting the resonance characteristics of the vocal tract and playing a crucial role in speech perception and understanding. Extracting and analyzing formants enables various applications such as speech recognition, speech synthesis, and hearing aid optimization. In practical applications, combining advanced signal processing techniques (such as LPC and STFT) with machine learning algorithms allows for more accurate formant extraction and leverage of their characteristics to improve the performance of speech processing systems.
[0046] Figure 1 This is a flowchart illustrating a hearing aid enhancement method provided in this disclosure. Figure 1The method shown includes steps 101 to 103. This method can be performed by a hearing aid enhancement device. This hearing aid enhancement device can be worn on the user's waist or chest. The hearing aid enhancement device can establish a connection with the user's headphones via wireless network or Bluetooth communication methods. Then, the hearing aid enhancement device can transmit the picked-up sound to the user's headphones.
[0047] Specifically, Figure 1 The method shown includes: Step 101: Receiving and storing voice information.
[0048] This hearing aid enhancement device includes a memory. This memory stores received speech information. This speech information is the most original speech information.
[0049] The speech information mentioned here can be understood as the speech content of other people or loudspeakers in the surrounding environment picked up by hearing aid enhancement devices.
[0050] Hearing aid enhancement devices can selectively pick up speech information. In practical applications, directional pickup can be achieved. These devices are not worn directly on the user's ears but can be worn on the chest or waist, allowing for a larger size. Correspondingly, microphone arrays (e.g., circular or linear arrays) can be incorporated into the device to achieve directional speech information pickup and suppress interference from other directions.
[0051] In addition, hearing aid enhancement devices can also have headphone charging functions. When the headphones run out of power, they can be placed in the charging case of the hearing aid enhancement device to charge.
[0052] Step 102: In response to a hearing recovery command triggered by the user through headphones or a hearing aid enhancement device, determine at least a portion of the voice information to be replayed.
[0053] In the optional solution, the hearing aid enhancement device can be separate from the headphones (or it can be an integrated device, meaning the headphones have hearing aid enhancement functionality).
[0054] In practical applications, even when a user wears a hearing aid, they may still experience difficulty hearing clearly. In this case, the user can replay the previously heard audio information through headphones or a hearing aid enhancement device. For example, the user can directly touch the hearing aid enhancement device, or the user can press the replay button on the headphones. The specific implementation method for replaying will be explained in detail in subsequent embodiments, and will not be repeated here.
[0055] Generally, if a user requests repeated playback, it indicates that there are certain parts of the audio message that the user cannot hear clearly or has difficulty hearing. Future hearing aid enhancement devices should specifically address these issues. Therefore, the portion of the audio message that the user chooses to replay can be used as a training sample.
[0056] When storing received voice information, the voice data should be segmented according to time sequence or based on the results of voice activity detection (VAD), and each segment should be added with a timestamp and contextual information (such as volume, frequency distribution, ambient noise level, etc.) for subsequent retrieval and processing.
[0057] When a user triggers a hearing-enhancing command through headphones or a hearing aid enhancement device, the system will identify the speech segment that needs to be replayed based on the user's operation record, and replay at least part of the speech information in an agreed manner (such as slow playback, high-frequency enhancement, or playback after noise suppression) to meet the user's hearing needs.
[0058] Step 103: Replay at least a portion of the voice information according to the agreed method.
[0059] The agreed-upon methods mentioned here include: replaying according to a pre-agreed time, replaying the last sentence based on sentence segmentation, and replaying according to sentence type. These will be explained in detail in the following embodiments, and will not be repeated here.
[0060] Based on the above, it is clear that the replayed audio information reflects the user's interests or hearing difficulties. This replayed audio information not only includes content of interest to the user but may also encompass speech features in specific scenarios (such as speech in noisy environments or specific pronunciation patterns), thus providing a high-quality data foundation for subsequent voiced / unvoiced sound differentiation, formant optimization, and personalized parameter adjustments. In this way, the system can dynamically learn the user's preferences and hearing characteristics, continuously optimize speech enhancement strategies, and improve the user experience.
[0061] In one or more embodiments of this disclosure, at least a portion of the voice information is replayed in a pre-agreed manner, including: replaying the target voice corresponding to the voice information within a pre-agreed time range.
[0062] In practical applications, the time range for replaying voice messages can be set as needed. This time range is typically the length of a single sentence, such as 10 seconds. This time range is usually set according to the user's requirements. The target voice message referred to here is the voice message that needs to be replayed within the specified time range.
[0063] The time range mentioned here can be understood as the time range that needs to be replayed, set by the developers or users. The starting time of the agreed time range is the moment the user sends the replay request, and the audio information within the length of the time range starting from the starting time is the target audio for replay.
[0064] Based on the above solution, if the user did not hear clearly, the target audio can be replayed, thus meeting the user's needs.
[0065] In one or more embodiments of this disclosure, at least a portion of the voice information is replayed in a pre-agreed manner, including: performing sentence segmentation recognition on the stored voice information; and determining and replaying at least one statement that last responded to the user based on the sentence segmentation recognition result.
[0066] When performing sentence segmentation recognition on stored speech information, acoustic features and language models can be combined. First, by analyzing acoustic features such as short-time energy, fundamental frequency changes, and pause duration in the speech signal, sentence boundaries are initially determined. Then, speech-to-text (ASR) technology is used to convert speech into text, and punctuation prediction or semantic segmentation algorithms are used to further refine the sentence boundaries. Based on the sentence segmentation recognition results, the system can divide the speech information into independent sentence segments and record the timestamp and context information of each sentence. When a user needs to recall information, the system will prioritize locating and replaying at least one complete sentence that last responded to the user, and can adjust the playback method according to user needs (such as reducing the speech rate, enhancing specific frequencies, or adding background noise reduction). This approach not only improves the efficiency of voice interaction but also ensures that users can quickly obtain key information, significantly enhancing the user experience, especially in complex scenarios (such as noisy environments or multi-turn dialogues).
[0067] In practical applications, if a user has difficulty hearing something, they often replay the audio immediately upon noticing the problem. This means they might only need to hear the most recent 1 or 3 seconds of audio, without playing excessive amounts of additional content. Playing for too long could cause the user to miss the speaker's latest remarks. Therefore, accurately identifying the audio content the user wants to replay satisfies their need to listen again without causing them to miss the speaker's latest remarks and inconvenience by replaying too much content.
[0068] In one or more embodiments of this disclosure, at least a portion of the voice information is replayed in a pre-agreed manner, specifically including: performing sentence segmentation recognition on the stored voice information; determining at least one statement from the last responding user based on the sentence segmentation recognition result; determining the statement type of the at least one statement from the last responding user; if the statement type is an interrogative sentence, then starting a countdown; if the user does not reply after the countdown ends, then replaying the at least one statement from the last responding user.
[0069] When performing sentence segmentation recognition on stored speech information, the system analyzes pauses, energy changes, and fundamental frequency features in the speech signal, combining speech-to-text (ASR) technology and natural language processing (NLP) algorithms to accurately locate sentence boundaries and extract independent sentences. Based on the sentence segmentation recognition results, the system can locate at least one complete sentence that last responded to the user and further determine its sentence type (such as declarative, interrogative, or command sentence) through semantic analysis. If the sentence is identified as an interrogative sentence, specific logic is triggered, and a countdown is started to wait for the user's response. During the countdown, the system can monitor whether the user provides feedback through voice or other means. If the countdown ends and no user response is received, the interrogative sentence is automatically replayed, and playback parameters can be adjusted according to the user's hearing needs (such as enhancing high-frequency components, reducing speech rate, or adding prompts). This mechanism not only improves the intelligence level of voice interaction but also ensures that users do not miss key information, significantly enhancing system usability and user experience, especially in multi-tasking scenarios or noisy environments.
[0070] In practical applications, if a user encounters a situation where they cannot hear something clearly during communication, they often need to manually select to replay the audio. If the replay selection is not timely, it may be overwritten by newly acquired audio information. Alternatively, there may be many audio messages stored simultaneously, requiring the user to go through a tedious selection process to find the previously missed audio, increasing the user's workload. Therefore, the above solution can determine the sentence type of the last sentence in the acquired audio information without requiring any replay selection from the user. If it is a question, it usually indicates that a response is required from the user. If the hearing aid enhancement device does not pick up the user's response, it may be because the user did not hear the question clearly. Therefore, by comprehensively judging the sentence type and picking up the user's response, automatic replay of the audio information can be achieved, effectively improving the user experience.
[0071] In one or more embodiments of this disclosure, during replay, the replayed voice information is parsed to determine the voiced information contained in the voice information; the voiced information is enhanced using initial adjustment parameters to generate target voiced speech; the target voiced speech is sent to the user's headphone device, or played through a speaker integrated with the microphone.
[0072] In practical applications, after identifying the voice information that needs to be replayed, the replayed voice can be further enhanced. That is, the voiced information to be replayed is parsed to identify any voiced consonants, and these voiced consonants are enhanced. This ensures that the enhanced voice is clearly audible to the user.
[0073] In one or more embodiments of this disclosure, after receiving and storing the speech information, the method further includes: selecting training samples from the stored speech information; distinguishing between voiced and unvoiced sounds in the training samples to determine the voiced portion contained in the speech information; inputting the initial adjustment parameters as training samples into the optimization model, and fine-tuning the optimization model based on user feedback to generate target adjustment parameters for optimizing the formants of the voiced portion.
[0074] In practical applications, due to differences in hearing ability among users, the break-in process between users and hearing aids varies. Therefore, users need to make adjustments according to their actual needs during use. Current technology typically requires professional tuners and equipment for professional tuning, usually conducted in a specialized tuning room. However, users' actual usage environments are complex and diverse, and the tuning results from a tuning room may not meet their needs in some scenarios (such as noisy environments like a farmers' market).
[0075] Therefore, this disclosure proposes a scheme for real-time optimization and adjustment of hearing aids using an artificial intelligence model. Through the AI model, rapid optimization of adjustment parameters can be achieved, even in real-time. In one optional scheme, after obtaining the target adjustment parameters through optimization, these target adjustment parameters can be stored according to different scenarios. For example, adjustment parameters for a farmers' market, a meeting room, or a dinner party can be stored. In subsequent applications, by performing semantic analysis on the received voice information, the corresponding adjustment parameters of the hearing aid enhancement device can be quickly adjusted.
[0076] Specifically, when storing received speech information, a segmented storage approach can be adopted. The speech signal is divided into multiple segments according to a fixed time window or based on the results of speech activity detection (VAD), and each segment is added with a timestamp and metadata (such as volume, frequency distribution, etc.) for subsequent processing. When selecting training samples from the stored speech information, speech segments containing rich voiced and unvoiced features should be prioritized, while ensuring sample diversity to cover different speakers, speech rates, intonations, and environmental noise conditions. When distinguishing between voiced and unvoiced sounds in the training samples, acoustic features (such as fundamental frequency, short-time energy, zero-crossing rate, etc.) and machine learning classifiers (such as support vector machines or deep neural networks) can be combined to extract the voiced portion, and its spectral characteristics can be further refined through formant analysis. Initial adjustment parameters (such as equalizer gain, filter cutoff frequency, etc.) are used as input, and optimization models (such as genetic algorithms or reinforcement learning models) are used to iteratively optimize them. At the same time, a user feedback mechanism is introduced, such as obtaining user preference information through questionnaires or real-time listening scores, thereby dynamically adjusting the model weights. The final target adjustment parameters can be precisely optimized for the formant frequencies of the voiced parts, improving speech clarity and intelligibility while meeting the personalized needs of users.
[0077] Based on the above approach, it is evident that utilizing artificial intelligence models (i.e., optimization models) enables rapid optimization and adjustment of parameters, thereby generating target adjustment parameters that meet the user's current needs. This not only effectively improves the efficiency of parameter optimization and adjustment but also satisfies the diverse parameter adjustment needs of users in various scenarios.
[0078] When selecting training samples, replayed audio can be chosen. Replayed audio information, reflecting the user's interests or hearing difficulties, can be considered representative training samples. These samples not only contain content of interest to the user but may also cover speech features in specific scenarios (such as speech in noisy environments or specific pronunciation patterns), thus providing a high-quality data foundation for subsequent voiced / unvoiced sound differentiation, formant optimization, and personalized parameter adjustments. In this way, the system can dynamically learn user preferences and hearing characteristics, continuously optimize speech enhancement strategies, and improve the user experience.
[0079] In one or more embodiments of this disclosure, such as Figure 2 This is a schematic flowchart illustrating an embodiment of the enhancement method. After receiving and storing the voice information, the method further includes: Step 201: Parsing the received voice information to determine the voiced information contained in the voice information.
[0080] After the hearing aid enhancement device picks up the sound, it analyzes the speech information. In practical applications, the analysis methods can include fundamental frequency detection, spectral feature analysis, or classification algorithms based on machine learning models (such as support vector machines, decision trees, random forests, etc.). After analyzing the speech information using the above methods, the unvoiced and voiced spectra corresponding to the speech information are obtained.
[0081] In practical applications, unvoiced sounds have a more uniform spectrum, while voiced sounds have a more concentrated spectrum and are prone to formants. This makes them susceptible to interference, preventing users from clearly hearing the voiced portion of the speech. In this proposed solution, the obtained voiced spectrum needs to be enhanced. Therefore, during parsing, the focus can be on extracting the voiced information from the speech data.
[0082] Step 202: Enhance the voiced information using the initial adjustment parameters to generate target speech carrying enhanced voiced sounds.
[0083] As mentioned earlier, after parsing the speech information, the voiced components are obtained. Since different voiced components correspond to different formant frequencies, it is necessary to enhance the voiced components using corresponding adjustment parameters.
[0084] After enhancement processing, the voiced consonants in the final target speech are made clearer. Furthermore, the enhanced voiced consonants can be integrated with the unvoiced consonants to obtain the target speech. This allows users to hear the actual content of the speech information more clearly.
[0085] Step 203: Send the target voice to the user's headset device.
[0086] After obtaining the target speech, the hearing aid enhancement device can further transmit the target speech to the user's headset device. In this disclosed solution, the hearing aid enhancement device and the headset device are designed separately. This significantly improves the speech pickup and enhancement processing capabilities of the headset device without increasing its weight and size, thereby better meeting the diverse needs of users.
[0087] Based on the aforementioned publicly available solution, the hearing aid enhancement device performs targeted enhancement processing on voiced sounds. The hearing aid enhancement device has a stronger speech enhancement processing capability than hearing aid headphones, which is beneficial to improving the effect of hearing aid use.
[0088] In one or more embodiments of this disclosure, the voiced information is enhanced to generate target speech carrying enhanced voiced sounds, specifically including: using a model to perform text recognition on the speech information carrying enhanced voiced sounds; correcting the voiced text corresponding to the voiced information based on the context content in the text recognition result; and converting the corrected text recognition result into target speech.
[0089] In practical applications, even after enhancing voiced information by adjusting parameters, some voiced information may still remain unclear. In such cases, a speech recognition model can be used to perform text recognition on the enhanced voiced speech information, and then the text content can be corrected based on the text recognition results. After completing the text content correction, the text recognition results can be further converted into target speech and played back.
[0090] In practical applications, instead of first enhancing the voiced information by adjusting parameters, a speech recognition model can be used to perform text recognition on the voiced speech information. Then, the text content can be corrected based on the text recognition results. After text content correction, the text recognition results can be converted into target speech, and the target speech can be enhanced with voiced information before being played back.
[0091] When enhancing voiced information, the acoustic features of the voiced portion (such as fundamental frequency, formant frequencies, and harmonic structure) are first analyzed. Signal processing techniques (such as dynamic equalization, harmonic enhancement, or noise suppression) are then used to optimize the voiced portion, generating target speech carrying the enhanced voiced sounds. Subsequently, the enhanced speech information is input into a speech recognition model (such as a deep learning-based ASR model) for text recognition to extract the corresponding text content. Since enhancement processing may introduce some distortion or error, the voiced text in the recognition result needs to be corrected based on contextual information. For example, language models or semantic analysis algorithms can be used to correct potential errors (such as homonym ambiguity or polysemous words). Finally, the corrected text recognition result is converted into target speech using high-quality speech synthesis technology (such as TTS, Text-to-Speech), ensuring that the output speech achieves a high level of clarity, naturalness, and semantic accuracy. This process not only improves the intelligibility and listening quality of the speech information but also provides more reliable data support for subsequent voice interaction or hearing assistance applications, thereby better meeting users' personalized needs.
[0092] Based on the above scheme, the received speech information is first subjected to voice enhancement processing, and then text recognition and correction are performed. In turn, the corrected text recognition result is converted into target semantics, which can effectively improve the playback accuracy of hearing aids.
[0093] In one or more embodiments of this disclosure, the hearing aid enhancement device is configured with a display screen; after generating target speech carrying enhanced voiced sounds, the device further includes: converting the target speech into text and displaying it on the display screen.
[0094] In practical applications, a display screen can also be configured for hearing aid enhancement devices. This screen can then display the text content converted from the speech information picked up by the hearing aid enhancement device.
[0095] In addition, this hearing aid enhancement device can also record speech. For example, during a meeting, the device can be used as a voice recorder to record conversations and store them in a memo.
[0096] Based on the above, it is clear that this allows users to intuitively see the content of voice interactions and related prompts. For example, if a project is mentioned during a voice interaction, information about that project's meetings and progress will be displayed.
[0097] Hearing aid enhancement devices are equipped with calendar and alarm clock functions; calendar reminders and / or alarm clock reminders can be set according to the user's voice commands.
[0098] This hearing aid enhancement device with a display screen can also display a calendar, allowing users to record to-do items and receive reminders. It can also be used as an alarm clock. Because the device has voice pickup capabilities, when used as a calendar, users can control it via voice to set relevant tasks, reminder content, and reminder times, thus better meeting diverse user needs.
[0099] In one alternative, the hearing aid enhancement device can be made into the form of a name tag or badge, which can be hung around the neck with a strap. This allows the hearing aid with a display screen to be worn around the neck at all times, enabling better acquisition of the speech content of the speaker in front of the user (e.g., in face-to-face communication).
[0100] In addition, this hearing aid enhancement device can also take the form of a charging case, worn on the user's chest or hung around the neck via a strap. This case can also include a display screen. When the earphones run out of power, they can be placed inside the case to charge.
[0101] Based on the aforementioned publicly available solutions, hearing aid enhancement devices not only help users hear speech information clearly, but can also serve as tools such as notebooks and calendars to help users improve their daily work efficiency.
[0102] In one or more embodiments of this disclosure, the hearing aid enhancement device includes a microphone array; after receiving voice information, the microphone array is used to locate the target object that emitted the voice information.
[0103] Microphone arrays in hearing aid enhancement devices, through the coordinated operation of multiple microphones, can accurately capture sound signals in space and determine the location of the target object emitting the voice information using sound source localization technology. Specifically, after receiving sound signals from different directions, the microphone array calculates the time difference of arrival (TDOA) or phase difference of the signals received by each microphone, and combines this with beamforming technology to focus on the direction of the target sound source while suppressing environmental noise and interference from other directions. Furthermore, to improve localization accuracy, algorithms based on DOA (Direction of Arrival) or multi-channel signal processing techniques can be used to dynamically track the positional changes of moving sound sources.
[0104] This localization capability not only helps hearing aids better separate target speech from background noise, but also adaptively adjusts the pickup direction according to the user's needs, thereby significantly improving speech clarity and the user's auditory experience, especially in noisy environments or multi-person conversation scenarios.
[0105] Based on the same idea, this disclosure also provides a hearing aid enhancement system. For example... Figure 3 This disclosure provides a hearing aid enhancement system. The system includes: a hearing aid enhancement device 31, used to parse received speech information to determine voiced information contained in the speech information; to enhance the voiced information using initial adjustment parameters to generate target speech carrying enhanced voiced information; and to send the target speech to a user's headphone device. The headphone device 32 is used to receive and play the target speech.
[0106] Furthermore, the connection between the headset device 32 and the mobile phone 33 can be selected via Bluetooth as needed. If the headset device 32 and the mobile phone 33 establish a connection via Bluetooth, the headset device 32 can be used as a regular Bluetooth headset to receive audio information transmitted by the mobile phone 33 via Bluetooth.
[0107] Even without establishing a connection between the hearing aid enhancement device 31 and the mobile phone 33, the hearing aid enhancement device 31 can directly connect to the server via its built-in network capabilities (such as WLAN or LTE) to achieve cloud storage and data processing functions. The mobile phone 33 can synchronize relevant audio data or text data converted from audio from the cloud server.
[0108] For example, Figure 4This is a schematic diagram illustrating how a hearing aid enhancement device is worn, as exemplified in this disclosure. From Figure 4 As seen in the image, the user is wearing Bluetooth headphones and has the hearing aid accessory (HAA) worn in their chest pocket. This hearing aid accessory can take many forms, such as a name tag, badge, or even a charging case.
[0109] Hearing aid enhancement devices have a built-in array of multiple microphones, which can be triggered by a button to support directional or omnidirectional sound pickup facing the speaker. Through noise reduction and beam forming, ambient noise is suppressed, and only sound from the speaker's direction is picked up. This sound is then transmitted to a Bluetooth headset (or a hearing aid with Bluetooth functionality) via Bluetooth A2DP profile.
[0110] In addition, the hearing aid enhancement device features a one-button repeat function. This allows users to listen to the past 10 seconds of audio with a single button press. It also allows for loop recording, playing back the audio of the last 10 seconds of voice information.
[0111] In addition, the system also includes: a mobile phone 33, and a hearing aid enhancement device with a microphone; the hearing aid enhancement device is connected to headphones and / or the mobile phone via Bluetooth communication.
[0112] In practical applications, artificial intelligence techniques can be used to enhance voiced sounds. For example, an optimized model trained on the system can be used to enhance voiced sounds. To improve the performance of the optimized model, suitable training samples can be selected for training.
[0113] For example, misheard or missed words and phrases can be marked, and the system can use statistical analysis to find the commonalities and characteristics of these words and phrases, such as their voiced or unvoiced qualities. Figure 5 This is a schematic diagram of the real-time audio spectrogram provided in this disclosure. The audio contains the speech information of "real-time audio and video interaction." The unvoiced sounds "real," "time," and "video" exhibit a more uniform spectrum, while the voiced sounds corresponding to "sound," "interaction," and "movement" show a more concentrated spectrum. Linguistically, voiced sounds contain more information. Therefore, the following analysis will focus on voiced sounds.
[0114] Voiced sounds produce a resonance peak during pronunciation. For example... Figure 6 A schematic diagram of resonance peaks provided for embodiments of this disclosure. From Figure 6As can be seen, different voiced sounds correspond to different formants, meaning that different words correspond to different formant frequencies. For example, the formants of "a" are 750Hz, 1200Hz, and 2600Hz. By using information from a large number of words that were missed, we can analyze which frequencies need to be enhanced. Based on the analysis results, we can find the initial adjustment parameters, which means we can set the initial frequency enhancement parameters.
[0115] During user interaction, the focus is on collecting and storing speech information that users mishear or miss, especially voiced sounds. This speech information can be used as training samples. To improve model training efficiency and ensure the training results better match actual user needs, AI reinforcement learning from human feedback (RLHF) can be used. This involves replaying the speech information during training and having users rate the quality of the enhanced speech processed by the optimized model. The ratings are then fed back to the training model for further optimization. This results in an optimized model that better suits the user's needs. In other words, each user can train an optimized model that meets their specific requirements.
[0116] Based on any of the above embodiments, this disclosure also provides a hearing aid enhancement device. Figure 7 This is a schematic block diagram of a hearing aid enhancement device according to one embodiment of this disclosure. Figure 7 As shown, the hearing aid enhancement method apparatus includes: a receiving module 71 for receiving and storing speech information; a determining module 72 for determining at least a portion of the speech information to be replayed in response to a hearing enhancement command triggered by a user through headphones or a hearing aid enhancement device; and a playback module 73 for replaying at least a portion of the speech information according to a pre-agreed method.
[0117] Optionally, the playback module 73 is used to replay the target speech corresponding to the speech information within the agreed time range.
[0118] Optionally, the playback module 73 is used to perform sentence segmentation recognition on the stored voice information; and based on the sentence segmentation recognition result, to determine and replay at least one statement that last responded to the user.
[0119] Optionally, the playback module 73 is used to perform sentence segmentation recognition on the stored voice information; determine at least one statement that last responded to the user based on the sentence segmentation recognition result; determine the statement type of the at least one statement that last responded to the user; if the statement type is an interrogative sentence, start a countdown; if the user does not reply after the countdown ends, replay the at least one statement that last responded to the user.
[0120] Optionally, it also includes an enhancement module 74, which is used to parse the received speech information, determine the voiced information contained in the speech information; enhance the voiced information using initial adjustment parameters to generate target speech carrying enhanced voiced information; and send the target speech to the user's headphone device, or play it through a speaker integrated with the microphone.
[0121] Optionally, it also includes a training module 75, which is used to select training samples from stored speech information; distinguish between voiced and unvoiced sounds in the training samples to determine the voiced parts contained in the speech information; input the initial adjustment parameters as training samples into the optimization model, and fine-tune the optimization model according to user feedback to generate target adjustment parameters for optimizing the formants of the voiced parts.
[0122] Optionally, the enhancement module 74 is used to perform text recognition on speech information carrying enhanced voiced consonants using the model; to correct the voiced text corresponding to the voiced consonant information based on the context content in the text recognition result; and to convert the corrected text recognition result into target speech.
[0123] Optionally, the hearing aid enhancement device is equipped with a display screen; after generating target speech carrying enhanced voiced sounds, it further includes: converting the target speech into text and displaying it on the display screen.
[0124] Optionally, the hearing aid enhancement device is equipped with calendar and alarm clock functions; calendar reminders and / or alarm clock reminders can be set according to the user's voice commands.
[0125] Optionally, the hearing aid enhancement device includes a microphone array; after receiving voice information, the microphone array is used to locate the target object emitting the voice information.
[0126] The enhancement module 74 is used to parse the replayed voice information during replay to determine the voiced information contained in the voice information; enhance the voiced information using initial adjustment parameters to generate target voiced speech; and send the target voiced speech to the user's headphone device, or play it through a speaker integrated with the microphone.
[0127] The specific implementation process of the functions and roles of each module in the above device can be found in the implementation process of the corresponding steps in the above method, and will not be repeated here.
[0128] The subject executing the hearing aid enhancement method in this specific embodiment can be a hearing aid device. The hearing aid device includes a hearing aid enhancement device and an earphone, wherein the hearing aid enhancement device and the earphone are either an integral or separate structure; wherein the hearing aid enhancement device is used to execute the hearing aid enhancement method.
[0129] Therefore, based on any of the above embodiments, this disclosure also provides an electronic device that can perform the hearing aid enhancement method of any of the embodiments described above.
[0130] In one alternative embodiment, the electronic device can be a pair of glasses with hearing aid functionality, on which are mounted a hearing aid enhancement device and headphones; wherein the hearing aid enhancement device is used to perform a hearing aid enhancement method, and the headphones are used to play speech information. The hearing aid enhancement device is capable of picking up ambient sounds.
[0131] Furthermore, the electronic device can also consist of headphones and a box-shaped hearing aid enhancement device, which can be used to charge the headphones. The box can be worn on the chest during use.
[0132] like Figure 8 This is a schematic diagram illustrating the structure of eyeglasses with hearing aid function as an example of this disclosure.
[0133] from Figure 8 As can be seen, the glasses 81 are equipped with a hearing aid enhancement device 82, an earphone 83, and an audio receiver 84. This gives the glasses the function of a hearing aid, and also enables them to achieve… Figure 1 Enhancements to the embodiments described. Figure 8 The headphones shown are bone conduction headphones, mounted on the temples of glasses. Of course, they can also be configured as earbuds. Users can make corresponding modifications based on the inventive concept of this solution according to their own needs.
[0134] Furthermore, the mounting positions of the hearing aid enhancement device 82, earphone 83, and audio receiver 84 on the glasses can be adjusted and changed according to user design requirements and functional needs. For example, the left and right positions of the hearing aid enhancement device 82 and the audio receiver 84 can be interchanged. Figure 8 The illustrated glasses structure with hearing aid function is for illustrative purposes only and does not constitute a limitation on the technical solution disclosed herein. Those skilled in the art can adjust and modify the inventive concept as needed. The hearing aid enhancement device 82 can also be disposed at the connection position between the two lenses (i.e., above the nose pad of the glasses).
[0135] The glasses 8 can be smart glasses (such as glasses that support AI, AR, or VR functions) or ordinary glasses.
[0136] Figure 9 This is a schematic block diagram of an electronic device according to one embodiment of the present disclosure.
[0137] The hardware architecture of the electronic device 1000 can be implemented using a bus architecture. The bus architecture can include any number of interconnect buses and bridges, depending on the specific application of the hardware and overall design constraints. Bus 1100 connects various circuits, including one or more processors 1200, memory 1300, and / or hardware modules. Bus 1100 can also connect various other circuits 1400, such as peripheral devices, voltage regulators, power management circuits, external antennas, etc.
[0138] Bus 1100 can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Component (EISA) bus, etc. Buses can be categorized as address buses, data buses, control buses, etc. For ease of representation, only one connection line is used in this diagram, but this does not imply that there is only one bus or only one type of bus.
[0139] This disclosure also provides a readable storage medium storing a computer program that, when executed by a processor, is used to implement the methods described above. A "readable storage medium" can be any means capable of containing, storing, communicating, propagating, or transmitting a program for use by or in conjunction with an instruction execution system, apparatus, or device. More specific examples of a readable storage medium include: an electrical connection with one or more wires (electronic device), a portable computer disk drive (magnetic device), random access memory (RAM), read-only memory (ROM), erasable and programmable read-only memory (EPROM or flash memory), fiber optic devices, and portable read-only memory (CDROM), etc.
[0140] This disclosure also provides a computer program product, the methods of which can be implemented wholly or partially through software, hardware, firmware, or any combination thereof. When implemented in software, it can be implemented wholly or partially as a computer program product. The computer program product includes one or more computer programs or instructions. When the computer program or instructions are loaded and executed, all or part of the processes or functions of this disclosure are performed.
[0141] Computer programs or instructions can be stored in a readable storage medium or transferred from one readable storage medium to another. For example, the computer program or instructions can be transferred from one website, computer, server, or data center to another website, computer, server, or data center via wired or wireless means. The readable storage medium can be any available medium capable of access, or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium, such as a floppy disk, hard disk, or magnetic tape; an optical medium, such as a digital video optical disc; or a semiconductor medium, such as a solid-state drive. The computer-readable storage medium can be a volatile or non-volatile storage medium, or it can include both volatile and non-volatile types of storage media.
[0142] Those skilled in the art will understand that embodiments of this disclosure can be provided as methods, systems, or computer program products. Therefore, this disclosure can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this disclosure can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0143] This disclosure is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to this disclosure. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0144] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing apparatus to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0145] These computer program instructions may also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable apparatus for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0146] In the description of this specification, the references to terms such as "one embodiment / mode," "some embodiments / modes," "example," "specific example," or "some examples," etc., refer to specific features, structures, or characteristics described in connection with that embodiment / mode or example, which are included in at least one embodiment / mode or example of this disclosure. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment / mode or example. Moreover, the specific features, structures, or characteristics described may be combined in any suitable manner in one or more embodiments / modes or examples. Furthermore, without contradiction, those skilled in the art can combine and integrate the different embodiments / modes or examples described in this specification, as well as the features of different embodiments / modes or examples.
[0147] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of that feature. In the description of this disclosure, "a plurality of" means at least two, such as two, three, etc., unless otherwise explicitly specified.
[0148] Those skilled in the art should understand that the above embodiments are merely for illustrating the present disclosure and are not intended to limit the scope of the disclosure. Those skilled in the art can make other changes or modifications based on the above disclosure, and these changes or modifications still fall within the scope of the present disclosure.
Claims
1. A method for playing voice in a hearing aid, characterized in that, Applied to hearing aid enhancement devices, the method includes: Receive and store voice information; In response to a hearing recovery command triggered by the user via headphones or the hearing aid enhancement device, at least a portion of the voice information to be replayed is determined; At least a portion of the voice information will be replayed in accordance with the agreed method.
2. The method according to claim 1, characterized in that, The replaying of at least a portion of the voice information in accordance with the agreed method includes: According to the agreed time range, the target voice corresponding to the voice information within the time range is replayed.
3. The method according to claim 1, characterized in that, The replaying of at least a portion of the voice information in accordance with the agreed method includes: Sentence segmentation and recognition are performed on the stored voice information; Based on the sentence segmentation recognition results, at least one of the last statements that responded to the user is identified and replayed.
4. The method according to claim 1, characterized in that, The replaying of at least a portion of the voice information in accordance with the agreed method includes: Sentence segmentation and recognition are performed on the stored voice information; Based on the sentence segmentation recognition results, at least one statement that last responded to the user is determined; Determine the statement type of the last statement that responded to the user; If the statement type is an interrogative sentence, then start the countdown; If the user does not respond after the countdown ends, at least one statement that last responded to the user will be replayed.
5. The method according to claim 1, characterized in that, After receiving and storing the voice information, it also includes: The received voice information is parsed to determine the voiced information contained in the voice information; The voiced information is enhanced using initial adjustment parameters to generate target speech carrying enhanced voiced sounds; The target voice is sent to the user's headset device, or played through a speaker integrated with the microphone.
6. The method according to claim 1, characterized in that, After receiving and storing the voice information, it also includes: Select training samples from the stored speech information; The training samples are subjected to voiced / unvoiced sound differentiation to determine the voiced components contained in the speech information; The initial adjustment parameters are input into the optimization model as training samples, and the optimization model is fine-tuned based on user feedback to generate target adjustment parameters for optimizing the formants of the voiced part.
7. The method according to claim 5, characterized in that, The enhancement processing of the voiced information to generate target speech carrying enhanced voiced sounds includes: The model is used to perform text recognition on speech information carrying enhanced voiced sounds; The voiced text corresponding to the voiced information is corrected based on the context content in the text recognition results; The corrected text recognition results are converted into target speech.
8. The method according to any one of claims 2 to 4, characterized in that, Also includes: During replay, the replayed audio information is parsed to determine the voiced information contained in the audio information; The voiced information is enhanced using initial adjustment parameters to generate target speech carrying enhanced voiced sounds; The target voice is sent to the user's headset device, or played through a speaker integrated with the microphone.
9. A hearing aid device, characterized in that, include: A hearing aid enhancement device and headphones, wherein the hearing aid enhancement device and the headphones are an integral structure or separate structures; wherein the hearing aid enhancement device is used to perform the method according to any one of claims 1 to 8.
10. A hearing aid system, characterized in that, include: A hearing aid enhancement device with a microphone, headphones, a cloud computer and / or a mobile phone, wherein the microphone is used to acquire speech information and the headphones are used to play speech information; The hearing aid enhancement device is connected to the headphones and / or the mobile phone via Bluetooth communication, and the cloud computer, the mobile phone, or the hearing aid enhancement device is used to perform the method of any one of claims 1 to 8.