A method and device for speech enhancement in a multi-mode digital hearing aid
By analyzing environmental sound information and using a multi-gain method to select the main sound source for differentiated amplification, the problem of insufficient sound discrimination ability of digital hearing aids in complex environments is solved, achieving a more natural auditory experience.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- ZUODIAN IND (HUBEI) CO LTD
- Filing Date
- 2024-03-16
- Publication Date
- 2026-06-02
AI Technical Summary
The gain technology of existing digital hearing aids lacks coordination in complex environments, resulting in noisy output audio and difficulty in effectively distinguishing important sound information.
By acquiring environmental sound information, performing noise reduction processing, and analyzing the frequency, magnitude, location, and distance of the sound source, a multi-gain method is used to select the main sound source and perform differentiated amplification output, avoiding hearing impairment caused by a single gain method.
It improves the ability of hearing-impaired patients to distinguish important sounds in complex environments, closely approximating the hearing mechanism of the human ear, compensating for hearing loss, and enabling patients to work and live like ordinary people.
Smart Images

Figure CN122138110A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of digital hearing aid technology, specifically to a speech enhancement method and apparatus for a multi-mode digital hearing aid. Background Technology
[0002] Digital hearing aids automatically collect information about the types of sound signals, signal-to-noise ratio, and intensity differences between the front and rear microphones in the surrounding environment. They define different environments and automatically adjust noise reduction, directionality, and compression ratio to adapt to constantly changing conditions. This avoids the problem of analog hearing aid users experiencing difficulty hearing soft sounds or discomfort from loud sounds. Digital hearing aids employ frequency segmentation and multi-channel technology, resulting in more refined sound signal processing. In terms of speech fidelity and sound quality, they more closely resemble the natural perception of the human ear. Different hearing programs can be set according to the user's different usage environments and automatically switched. For example, Chinese Patent Publication No. CN113938804A discloses a range... A hearing aid method and device are disclosed. By filtering and preserving sound signals of the same frequency, it can distinguish sound signals from different sources. Through microphone array sound localization technology, it can obtain the source location information of each sound signal, enabling the hearing aid to amplify and process sound signals within a close range. It is less affected by distant sounds, which is beneficial for hearing-impaired individuals' near-field responsiveness and communication. It significantly improves the hearing-impaired individual's ability to distinguish and differentiate surrounding sounds, making it particularly suitable for small-scale communication scenarios such as classroom teaching, office meetings, long-distance travel, and family life. Chinese Patent Publication No. CN115396799A discloses a self-... Oriented hearing aids, with their directional sound-to-electricity conversion modules dividing the auditory range, allow the hearing aid to simultaneously receive sound signals. Based on the decibel level, frequency, and acquisition time of the sound signal, the effective auditory range is identified, enabling the hearing-impaired individual to focus their attention on a specific area. This reduces the signal-to-noise ratio, alleviates the workload on the brain in speech discrimination, and enhances the individual's sound discrimination ability. When hearing aids are worn in both ears, the wireless signal shares real-time hearing aid operation information, ensuring a consistently symmetrical division of the auditory range for both ears. This allows for more precise analysis of the sound source's location by the brain, making the auditory mechanism of the hearing-impaired individual essentially the same as that of sighted individuals. This is essentially the same as traditional hearing aids. Chinese Patent Publication No. CN103811020B discloses an intelligent speech processing method that, by establishing a speech model library for interlocutors, can intelligently identify the identities of multiple interlocutors in a multi-person speech environment, while separating mixed speech to obtain the independent speech of each interlocutor. According to the user's needs, it amplifies the speech of the interlocutor the user wants to hear while eliminating the speech of interlocutors not requested by the user. Unlike traditional hearing aids, this method can automatically provide the user with the sound they need based on their individual needs, reducing interference from non-target human voices other than noise, and reflecting the personalization, interactivity and intelligence of the method. The various gain technologies mentioned above often have significant limitations and are difficult to cope with complex work and life environments. In order to improve the performance of digital hearing aids, more and more manufacturers are trying to add various gain technologies to hearing aids. However, each gain technology amplifies its own gain source independently and lacks coordination, resulting in noisy output audio. This has created new technical problems. Therefore, this invention provides a speech enhancement method and device for multi-mode digital hearing aids to solve the above problems.
[0003] Invention Patent Content To address the shortcomings of existing technologies, this invention provides a speech enhancement method and apparatus for multi-mode digital hearing aids, thereby improving the overall performance of digital hearing aids.
[0004] According to a first aspect of the present disclosure, a preferred embodiment of the present invention provides a speech enhancement method for a multi-mode digital hearing aid, the method comprising: Acquire natural sound information from the external environment and perform noise reduction processing on the natural sound information to obtain comprehensive sound information; The comprehensive sound information is analyzed to obtain a set of sound source information, wherein the sound source information includes frequency information, magnitude information, azimuth information and distance information; The most suitable sound source information is selected as the main sound source based on a multi-gain method, wherein different gain methods have different weights; and The main sound source and the other sound sources are amplified and output differently according to a preset gain method.
[0005] In one embodiment, the synthesized sound information is analyzed to obtain a set of sound source information, wherein the sound source information includes frequency information, magnitude information, azimuth information, and distance information, including: The acoustic model is compared with the waveform of the integrated sound information, and different frequency information is divided as sound source information; The average intensity of the different sound source information per unit time is calculated as the magnitude information; Based on the time difference in acquiring different sound source information, the directional and distance information of different sound source information are calculated.
[0006] In one embodiment, the most suitable sound source information is selected as the main sound source based on a multi-gain method, wherein the weights of different gain methods are different, including: Different gain methods are used to select and record the gain sound sources within the sound source information set. The gain source that is selected the most times will be used as the main sound source.
[0007] In one embodiment, if the highest gain source is not unique among the gain sources selected most frequently, the gain source with the largest sum of gain mode weights is selected as the main source.
[0008] According to a second aspect of the present disclosure, the present invention provides a speech enhancement device for a multi-mode digital hearing aid, the device comprising: The preprocessing module is used to acquire natural sound information from the external environment and to perform noise reduction processing on the natural sound information to obtain comprehensive sound information. The analysis module is used to analyze the comprehensive sound information to obtain a set of sound source information, wherein the sound source information includes frequency information, magnitude information, azimuth information and distance information; The determination module is used to select the most suitable sound source information as the main sound source based on multiple gain methods, wherein different gain methods have different weights; and The output module is used to amplify and output the main sound source and other sound sources differently according to a preset gain method.
[0009] In one embodiment, the analysis module includes: The identification module is used to compare the waveform of the acoustic model with the waveform of the integrated sound information and divide different frequency information as sound source information; The first calculation module is used to calculate the average intensity of different sound source information per unit time as magnitude information; The second calculation module is used to calculate the directional and distance information of different sound sources based on the time difference in acquiring the different sound source information.
[0010] In one embodiment, the determination module includes: The simulation module is used to select and record the gain sound sources within the sound source information set using different gain methods; The third calculation module is used to select the gain sound source that has been selected the most times as the main sound source.
[0011] In one embodiment, if the selection of the highest gain sound source in the third calculation module is not unique, the gain sound source with the largest sum of gain mode weights is selected as the main sound source.
[0012] According to a third aspect of the present disclosure, the present invention provides a speech enhancement device for a multi-mode digital hearing aid, comprising: processor; Memory used to store the processor's executable instructions; The processor is configured to perform the steps of the above method.
[0013] According to a fourth aspect of the present disclosure, the present invention provides a computer-readable storage medium having a computer program stored thereon, the computer program being executed by a processor of the steps of the above-described method.
[0014] As can be seen from the above technical solution, the speech enhancement method and device for a multi-mode digital hearing aid provided by this invention patent can obtain the frequency information, magnitude information, orientation information and distance information of different sound sources based on the environmental sound waveform diagram, acoustic model and the time difference of the same environmental sound. Different gain methods are used to select the gain sound source, and the gain sound source selected most often is used as the main sound source for output. This avoids the situation where a single gain method will cause hearing impairment in hearing-impaired patients. It can help users effectively distinguish important sound information in the environment, which is close to the hearing mechanism of the human ear. To a certain extent, it makes up for the hearing loss of hearing-impaired patients, enabling them to work and live like ordinary people.
[0015] It should be understood that the above general description and the following detailed description are merely exemplary and do not limit this disclosure. Attached Figure Description
[0016] To more clearly illustrate the specific embodiments of this invention, the accompanying drawings used in the description of the specific embodiments or prior art will be briefly introduced below. In all the drawings, the elements or parts are not necessarily drawn to scale.
[0017] Figure 1 A flowchart of a speech enhancement method for a multi-mode digital hearing aid provided for this invention patent; Figure 2 A flowchart of step S20 in a speech enhancement method for a multi-mode digital hearing aid provided by this invention patent; Figure 3 A flowchart of step S30 in a speech enhancement method for a multi-mode digital hearing aid provided by this invention patent; Figure 4 A block diagram of a speech enhancement device for a multi-mode digital hearing aid provided by this invention patent; Figure 5 A block diagram of another speech enhancement device for a multi-mode digital hearing aid provided by this invention. Detailed Implementation
[0018] The embodiments of the technical solution of this invention will now be described in detail with reference to the accompanying drawings. These embodiments are merely illustrative of the technical solution of this invention and are therefore intended to limit the scope of protection of this invention.
[0019] Figure 1This invention provides a flowchart of a speech enhancement method for a multi-mode digital hearing aid. The method is applied to an electronic nebulizer terminal, which can display images, videos, text messages, WeChat messages, etc. The terminal can be equipped with any device with a display screen, such as a mobile phone, computer, digital broadcasting terminal, messaging device, game console, tablet, medical device, fitness equipment, or personal digital assistant. This embodiment provides a speech enhancement method for a multi-mode digital hearing aid, such as... Figure 1 As shown, the method includes the following steps S10-S40: In step S10, natural sound information of the external environment is acquired, and the natural sound information is subjected to noise reduction processing to obtain comprehensive sound information; In this implementation, at least two sets of microphones are provided. The two sets of microphones convert the sound signal into an electrical signal, and the electrical signal is then processed by a filter to reduce noise and avoid noise interference.
[0020] In step S20, the integrated sound information is analyzed to obtain a set of sound source information, wherein the sound source information includes frequency information, magnitude information, azimuth information and distance information; In step S30, the most suitable sound source information is selected as the main sound source based on a multi-gain method, wherein different gain methods have different weights; and In step S40, the main sound source and the other sound sources are amplified and output differently according to a preset gain method; In this implementation, the main sound source and the other sound sources are amplified by amplifiers and then amplified and output by loudspeakers. The output intensity of the main sound source is much greater than that of the other sound sources, which can highlight the main sound source and help hearing-impaired patients better distinguish the main sound source.
[0021] Among them, such as Figure 2 As shown, in step S20, the integrated sound information is analyzed to obtain a set of sound source information, wherein the sound source information includes frequency information, magnitude information, azimuth information, and distance information, including the following steps S21-S23: In step S21, the acoustic model is compared with the waveform of the integrated sound information, and different frequency information is divided as sound source information; In this implementation, the acoustic model is one of the most important parts of the speech recognition system. With the help of the acoustic model, sound can be divided into speech, wind, whistle, flowing water and other sound source information.
[0022] In step S22, the average intensity of different sound source information per unit time is calculated as magnitude information; In this implementation, since the sound intensity of the sound source information changes constantly, using the average value as the magnitude information can improve the accuracy of the data.
[0023] In step S23, the azimuth and distance information of different sound sources are calculated based on the time difference in acquiring the different sound source information; In this implementation, the sound source localization system can quickly identify the location and distance information of the sound source.
[0024] In one embodiment, such as Figure 3 As shown, in step S30, the most suitable sound source information is selected as the main sound source based on multiple gain methods, wherein different gain methods have different weights, including the following steps S31-S32: In step S31, different gain methods are used to select and record the gain sound sources in the sound source information set; In this implementation, the gain methods include range-based hearing aids, directional hearing aids, speech enhancement methods, etc. Organically combining the above gain methods can overcome the application limitations of different types of hearing aids. For example, traditional hearing aids use sound intensity as gain, amplifying the sound source information proportionally, which can easily cause speech to be ignored; speech hearing aids often cause environmental sounds such as wind, whistles, and flowing water to be ignored; directional hearing aids can amplify the sound in front of the hearing-impaired patient by using directional information, which can easily ignore nearby sounds; range-based hearing aids can amplify the sound close to the hearing-impaired patient by using distance information, which can easily ignore other sounds.
[0025] In step S32, the gain sound source that is selected the most times is taken as the main sound source; In this implementation, based on human hearing habits, louder speech sounds that are close to the hearing-impaired patient are given priority attention. When the above conditions cannot be met simultaneously, the gain sound source that meets the most of the above conditions is selected as the main sound source. This can overcome gain defects and eliminate the need to switch gain modes according to different application scenarios. It can help users effectively distinguish important sound information in the environment, which is close to the human hearing mechanism. To a certain extent, it can make up for the hearing loss of hearing-impaired patients, enabling them to work and live like ordinary people, and be more flexible.
[0026] In one embodiment, if the highest gain source is not unique among the gain sources selected most frequently, the gain source with the largest sum of gain mode weights is selected as the main source. In this implementation, each gain method is scored differently based on the behavioral habits of hearing-impaired patients. The higher the score, the greater the weight, in order to ensure the uniqueness of the main sound source.
[0027] The following are embodiments of the apparatus disclosed herein, which can be used to execute embodiments of the method disclosed herein.
[0028] Figure 4 This invention patent provides a block diagram of a speech enhancement device for a multi-mode digital hearing aid. This device can be implemented as part or all of an electronic device through software, hardware, or a combination of both. Figure 4 As shown, the device, used in a digital hearing aid, includes: The preprocessing module 100 is used to acquire natural sound information from the external environment and perform noise reduction processing on the natural sound information to obtain comprehensive sound information. The analysis module 200 is used to analyze the comprehensive sound information to obtain a set of sound source information, wherein the sound source information includes frequency information, magnitude information, azimuth information and distance information; The determination module 300 is used to select the most suitable sound source information as the main sound source based on multiple gain methods, wherein different gain methods have different weights; and The output module 400 is used to amplify and output the main sound source and other sound sources differently according to a preset gain method.
[0029] This disclosure, based on environmental sound waveform diagrams, acoustic models, and time differences of the same environmental sound, can derive frequency, magnitude, location, and distance information of different sound sources. It employs different gain methods to select the gain sound source, and the gain sound source selected most frequently is used as the main sound source for output. This avoids the situation where a single gain method causes hearing impairment in hearing-impaired patients, and can help users effectively distinguish important sound information in the environment. It is close to the human hearing mechanism and, to a certain extent, compensates for the hearing loss of hearing-impaired patients, enabling them to work and live like ordinary people.
[0030] In one embodiment, such as Figure 4 As shown, the analysis module 200 includes: The identification module 201 is used to compare the waveform of the acoustic model with the waveform of the integrated sound information and divide different frequency information as sound source information; The first calculation module 202 is used to calculate the average intensity of different sound source information per unit time as magnitude information; The second calculation module 203 is used to calculate the directional and distance information of different sound sources based on the time difference of acquiring different sound source information.
[0031] In one embodiment, such as Figure 4 As shown, the determination module 300 includes: The simulation module 301 is used to select and record the gain sound sources in the sound source information set using different gain methods; The third calculation module 302 is used to take the gain sound source that is selected the most times as the main sound source.
[0032] In one embodiment, if the selection of the highest gain sound source is not unique in the third calculation module 302, the gain sound source with the largest sum of gain mode weights is selected as the main sound source.
[0033] Regarding the apparatus in the above embodiments, the specific manner in which each module performs its operation has been described in detail in the embodiments related to the method, and will not be elaborated upon here.
[0034] This disclosure also provides another speech enhancement device for a multi-mode digital hearing aid: Figure 5 This is a block diagram illustrating a voice enhancement device 800 for a multi-mode digital hearing aid according to an exemplary embodiment. For example, device 800 may be a mobile phone, computer, digital broadcasting terminal, messaging device, game console, tablet device, medical device, fitness equipment, personal digital assistant, etc.
[0035] Reference Figure 5 The device 800 may include one or more of the following components: a processing component 802, a memory 804, a power supply component 806, a multimedia component 808, an audio component 810, an input / output (I / O) interface 812, a sensor component 814, and a communication component 816.
[0036] Processing component 802 typically controls the overall operation of device 800, such as operations associated with display, telephone calls, data communication, camera operation, and recording. Processing component 802 may include one or more processors 820 to execute instructions to perform all or part of the steps of the methods described above. Furthermore, processing component 802 may include one or more modules to facilitate interaction between processing component 802 and other components. For example, processing component 802 may include a multimedia module to facilitate interaction between multimedia component 808 and processing component 802.
[0037] Memory 804 is configured to store various types of data to support the operation of device 800. Examples of such data include instructions for any application or method operating on device 800, contact data, phonebook data, messages, pictures, videos, etc. Memory 804 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk.
[0038] Power supply component 806 provides power to various components of device 800. Power supply component 806 may include a power management system, one or more power sources, and other components associated with generating, managing, and distributing power to device 800.
[0039] Multimedia component 808 includes a screen that provides an output interface between the device 800 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen may be implemented as a touchscreen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The touch sensors may sense not only the boundaries of the touch or swipe action but also the duration and pressure associated with the touch or swipe operation. In some embodiments, multimedia component 808 includes a front-facing camera and / or a rear-facing camera. When the device 800 is in an operating mode, such as a shooting mode or a video mode, the front-facing camera and / or the rear-facing camera may receive external multimedia data. Each front-facing camera and rear-facing camera may be a fixed optical lens system or have focal length and optical zoom capabilities.
[0040] Audio component 810 is configured to output and / or input audio signals. For example, audio component 810 includes a microphone (MIC) configured to receive external audio signals when device 800 is in an operating mode, such as call mode, recording mode, and voice recognition mode. The received audio signals may be further stored in memory 804 or transmitted via communication component 816. In some embodiments, audio component 810 also includes a speaker for outputting audio signals.
[0041] I / O interface 812 provides an interface between processing component 802 and peripheral interface modules, such as keyboards, click wheels, buttons, etc. These buttons may include, but are not limited to, home buttons, volume buttons, power buttons, and lock buttons.
[0042] Sensor assembly 814 includes one or more sensors for providing status assessments of various aspects of device 800. For example, sensor assembly 814 may detect the on / off state of device 800, the relative positioning of components such as the display and keypad of device 800, changes in the position of device 800 or a component of device 800, the presence or absence of user contact with device 800, the orientation or acceleration / deceleration of device 800, and temperature changes of device 800. Sensor assembly 814 may include a proximity sensor configured to detect the presence of nearby objects without any physical contact. Sensor assembly 814 may also include a light sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, sensor assembly 814 may also include an accelerometer, a gyroscope, a magnetometer, a pressure sensor, or a temperature sensor.
[0043] The communication component 816 is configured to facilitate wired or wireless communication between the device 800 and other devices. The device 800 can access wireless networks based on communication standards, such as WiFi, 2G or 3G, or combinations thereof.
[0044] In one exemplary embodiment, the communication component 816 receives broadcast signals or broadcast-related information from an external broadcast management system via a broadcast channel. In another exemplary embodiment, the communication component 816 further includes a near-field communication (NFC) module to facilitate short-range communication. For example, the NFC module may be implemented based on radio frequency identification (RFID) technology, Infrared Data Association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology, and other technologies.
[0045] In an exemplary embodiment, the apparatus 800 may be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components to perform the methods described above.
[0046] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions is also provided, such as a memory 804 including instructions, which can be executed by a processor 820 of the device 800 to perform the above-described method. For example, the non-transitory computer-readable storage medium may be a ROM, random access memory (RAM), CD-ROM, magnetic tape, floppy disk, and optical data storage device, etc.
[0047] Other embodiments of this disclosure will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of this disclosure that follow the general principles of this disclosure and include common knowledge or customary techniques in the art not disclosed herein. The specification and examples are to be considered exemplary only, and the true scope and spirit of this disclosure are indicated by the following claims.
[0048] It should be understood that this disclosure is not limited to the precise structures described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this disclosure is limited only by the appended claims.
Claims
1. A speech enhancement method for a multi-mode digital hearing aid, characterized in that, For use in digital hearing aids, the method includes: Acquire natural sound information from the external environment and perform noise reduction processing on the natural sound information to obtain comprehensive sound information; The comprehensive sound information is analyzed to obtain a set of sound source information, wherein the sound source information includes frequency information, magnitude information, azimuth information and distance information; The most suitable sound source information is selected as the main sound source based on a multi-gain method, wherein different gain methods have different weights; and The main sound source and the other sound sources are amplified and output differently according to a preset gain method.
2. The method according to claim 1, characterized in that, The comprehensive sound information is analyzed to obtain a set of sound source information, wherein the sound source information includes frequency information, amplitude information, azimuth information, and distance information, including: The acoustic model is compared with the waveform of the integrated sound information, and different frequency information is divided as sound source information; The average intensity of the different sound source information per unit time is calculated as the magnitude information; Based on the time difference in acquiring different sound source information, the directional and distance information of different sound source information are calculated.
3. The method according to claim 1, characterized in that, The most suitable sound source information is selected as the main sound source based on a multi-gain method, wherein different gain methods have different weights, including: Different gain methods are used to select and record the gain sound sources within the sound source information set. The gain source that is selected the most times will be used as the main sound source.
4. The method according to claim 3, characterized in that, If the highest gain source is not unique among the selected sources, the gain source with the largest sum of gain mode weights is selected as the main source.
5. A speech enhancement device for a multi-mode digital hearing aid, characterized in that, For a digital hearing aid, the device includes: The preprocessing module is used to acquire natural sound information from the external environment and to perform noise reduction processing on the natural sound information to obtain comprehensive sound information. The analysis module is used to analyze the comprehensive sound information to obtain a set of sound source information, wherein the sound source information includes frequency information, magnitude information, azimuth information and distance information; The determination module is used to select the most suitable sound source information as the main sound source based on multiple gain methods, wherein different gain methods have different weights; and The output module is used to amplify and output the main sound source and other sound sources differently according to a preset gain method.
6. The apparatus according to claim 5, characterized in that, The analysis module includes: The identification module is used to compare the waveform of the acoustic model with the waveform of the integrated sound information and divide different frequency information as sound source information; The first calculation module is used to calculate the average intensity of different sound source information per unit time as magnitude information; The second calculation module is used to calculate the directional and distance information of different sound sources based on the time difference in acquiring the different sound source information.
7. The apparatus according to claim 5, characterized in that, The determination module includes: The simulation module is used to select and record the gain sound sources within the sound source information set using different gain methods; The third calculation module is used to select the gain sound source that has been selected the most times as the main sound source.
8. The apparatus according to claim 7, characterized in that, If the selection of the highest gain sound source in the third calculation module is not unique, then the gain sound source with the largest sum of gain mode weights is selected as the main sound source.
9. A speech enhancement device for a multi-mode digital hearing aid, characterized in that, include: processor; Memory used to store the processor's executable instructions; The processor is configured to perform the steps of the method of any one of claims 1 to 4.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of any one of claims 1 to 4.
Citation Information
Patent Citations
A kind of intelligent voice processing method
CN103811020B
Range hearing aid method and device
CN113938804A
Direction-adaptive hearing aid
CN115396799A