Hearing protection device and method for improving fidelity of speech
Patent Information
- Application Number
- PCT/IB2026/051510
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2025-02-18
- Filing Date
- 2026-02-17
- Publication Date
- 2026-08-27
Smart Images

Figure IB2026051510_27082026_PF_FP_ABST
Abstract
Description
[0001] HEARING PROTECTION DEVICE AND METHOD FOR IMPROVING FIDELITY OF SPEECH
[0002] Technical Field
[0003] The present disclosure generally relates to a hearing protection device and a method for improving fidelity of speech.
[0004] Background
[0005] Hearing protection devices are widely used for protecting users against environmental noise. Some hearing protection devices typically include one or more microphones that receive a speech of a user or an ambient sound from the surroundings of the user for communication purposes. For example, some hearing protection devices include inner earbud microphones (IEMS) designed to pick up speech from the user.
[0006] Summary
[0007] In a first aspect, the present disclosure provides a method for improving fidelity of speech in a hearing protection device. The method includes receiving a sound signal from a sound receiver. The method further includes processing, by a sound filter, the sound signal to obtain a processed sound signal. The processing includes applying a voice activity detector (VAD) to the sound signal. The processing further includes filtering the sound signal into a plurality of frequency band. The processing further includes detecting a noise level in at least one frequency band of the plurality of frequency bands. The processing further includes applying a frequency gain rule to each of the plurality of frequency bands. The processing further includes reconstmcting the plurality of frequency bands to generate the processed sound signal. The method further includes broadcasting the processed sound signal.
[0008] In a second aspect, the present disclosure provides a method for improving fidelity of speech in a hearing protection device. The method includes receiving a sound signal from a sound receiver. The method further includes processing the sound signal to obtain a processed sound signal. The processing includes applying a voice activity detector (VAD) to the sound signal. The method further includes delaying broadcasting the processed sound signal by a time delay. The method further includes detecting, based on the processed sound signal, that the sound signal includes a speech. The method further includes retrieving a previous sound signal captured prior to the received sound signal based on the detection that the sound signal includes the speech. The method further includes generating a modified processed sound signal including the previous sound signal and the processed sound signal. The method further includes broadcasting the modified processed sound signal.
[0009] In a third aspect, the present disclosure provides a method for improving fidelity of speech in a hearing device. The method includes receiving a sound signal from a sound receiver. The method furtherincludes processing, by a sound filter, the sound signal to obtain a processed sound signal. The processing includes applying a voice activity detector (VAD) to the sound signal. The processing includes applying a square-factor crossfader. The method further includes broadcasting the processed sound signal.
[0010] In a fourth aspect, the present disclosure provides an in-ear hearing protection device. The in-ear hearing protection device includes an in-ear microphone configured to be positioned on an ear-canal side of the in-ear hearing protection device. The in-ear hearing protection device further includes an external microphone configured to capture an ambient sound external to the in-ear hearing protection device. The in-ear hearing protection device further includes a speaker configured to broadcast a processed sound signal. The in-ear hearing protection device further includes a sound processing unit configured to receive a sound signal from either the in-ear microphone or the external microphone and generate the processed sound signal. The sound processing unit includes a sound filter configured to separate the received sound signal into a first plurality of filter bands. The sound processing unit further includes a voice activity detector (VAD) configured to detect a speech in one or more of the first plurality of filter bands. The sound processing unit further includes a frequency gain modifier configured to apply a frequency gain mle to one or more of the first plurality of filter bands, thereby generating a second plurality of filter bands. The sound processing unit further includes a reconstructor configured to recombine the second plurality of filter bands to obtain the processed sound signal. The sound processing unit further includes a crossfader configured to crossfade the processed sound signal.
[0011] Brief Description of Drawings
[0012] Exemplary embodiments disclosed herein may be more completely understood in consideration of the following detailed description in connection with the following figures. The figures are not necessarily drawn to scale. Like numbers used in the figures refer to like components. However, it will be understood that the use of a number to refer to a component in a given figure is not intended to limit the component in another figure labeled with the same number.
[0013] FIG. 1 is schematic perspective view of an example in-ear hearing protection device, in accordance with embodiments herein;
[0014] FIG. 2 is a block diagram of an example in-ear hearing protection device, in accordance with embodiments herein;
[0015] FIG. 3 is a block diagram of an example in-ear hearing protection device, in accordance with another embodiment of the present disclosure;
[0016] FIGS. 4-6 are block diagrams illustrating various methods for improving fidelity of speech in a hearing protection device, in accordance with embodiments herein;
[0017] FIG. 7 is a block diagram of an example remote server architecture that can be used in embodiments shown in previous Figures; andFIG. 8 is a block diagram of a computing environment that can be used in embodiments shown in previous figures.
[0018] Detailed Description
[0019] In the following description, reference is made to the accompanying figures that form a part thereof and in which various embodiments are shown by way of illustration. It is to be understood that other embodiments are contemplated and may be made without departing from the scope or spirit of the present disclosure. The following detailed description, therefore, is not to be taken in a limiting sense.
[0020] In the following disclosure, the following definitions are adopted.
[0021] As used herein, all numbers should be considered modified by the term “about” . As used herein, “a,” “an,” “the,” “at least one,” and “one or more” are used interchangeably.
[0022] As used herein as a modifier to a property or attribute, the term “generally”, unless otherwise specifically defined, means that the property or attribute would be readily recognizable by a person of ordinary skill but without requiring absolute precision or a perfect match (e.g., within + / - 20 % for quantifiable properties).
[0023] The term “substantially”, unless otherwise specifically defined, means to a high degree of approximation (e.g., within + / - 10% for quantifiable properties) but again without requiring absolute precision or a perfect match.
[0024] As used herein, the term “configured to,” and the like, is at least as restrictive as the term “adapted to” and requires actual design intention to perform the specified function rather than mere physical capability of performing such a function.
[0025] The term “about”, unless otherwise specifically defined, means to a high degree of approximation (e.g., within + / - 5% for quantifiable properties) but again without requiring absolute precision or a perfect match.
[0026] As used herein, the terms “first” and “second” are used as identifiers. Therefore, such terms should not be construed as limiting of this disclosure. The terms “first” and “second” when used in conjunction with a feature or an element can be interchanged throughout the embodiments of this disclosure.
[0027] As used herein, “at least one of A and B” should be understood to mean “only A, only B, or both A and B”.
[0028] As used herein, the term “hearing protection device” generally refers to a personal protective equipment (PPE) article worn to reduce harmful auditory or annoying subjective effects of sound. The term “hearing protection device” may refer either to passive hearing protection or active hearing protection devices.
[0029] As used herein, the term “microphone” generally refers to an acoustic-to-electric transducer or a sensor that converts sound or acoustic vibrations into an electrical signal. Microphone can be anelectromagnetic induction or dynamic microphone, a capacitance charge or condenser microphone, a piezoelectric generation microphone adapted to produce an electrical signal from air pressure variations, or other types of sensors that convert sound or acoustic vibrations into an electrical signal.
[0030] As used herein, the term “earbud” generally refers to an electrical device consisting of an earphone, or a very small headphone, worn inside the ear of a user. The earbud can be wireless or wirebased to provide audio signals from an audio signal source.
[0031] As used herein, the term “speaker” or “loudspeaker” generally refers to an electro-acoustic transducer that produces sound in response to an electrical audio signal input.
[0032] Hearing protection devices are widely used for protecting users against environmental noise. Some hearing protection devices typically include one or more microphones that receive a speech of a user or an ambient sound from the surroundings of the user for communication purposes. For example, some hearing protection devices include inner earbud microphones (IEMS) designed to pick up the speech of the user. Such IEMs suffer from inherent bandwidth limitations due to their placement within an ear canal of the user, resulting in reduced intelligibility in communications to far-end receivers or voice recognition modules. In some cases, hearing protection devices include outer earbud microphones (OEMs) designed for ambient listening. However, OEMs are not typically used for communication purposes due to high static and / or variable ambient noise intrusion even if their bandwidth is much higher than that of the IEMs.
[0033] The present disclosure provides a method for improving fidelity of speech in a hearing protection device. The method includes receiving a sound signal from a sound receiver. The method further processing, by a sound filter, the sound signal to obtain a processed sound signal. The processing includes applying a voice activity detector (VAD) to the sound signal. The processing further includes filtering the sound signal into a plurality of frequency band. The processing further includes detecting a noise level in at least one frequency band of the plurality of frequency bands. The processing further includes applying a frequency gain rule to each of the plurality of frequency bands. The processing further includes reconstmcting the plurality of frequency bands to generate the processed sound signal. The method further includes broadcasting the processed sound signal.
[0034] The method of the present disclosure utilizes the sound filter to process the received sound signal. The processing includes filtering the sound signal in the plurality of frequency bands and detecting the noise level in the at least one frequency band of the plurality of frequency bands. This may allow the sound filter to determine one or more frequency bands of the plurality of frequency bands that can be used to generate the processed sound signal. Additionally, the processing of the sound signal includes applying the VAD to the sound signal. This may enable precise detection of speech starts and stops for reducing a noise in the processed sound signal, e.g., by applying a gate during speech pauses. In some cases, the processing further includes delaying broadcasting the processed sound signal and retrieving a previous sound signal to allow high-frequency fricatives (e.g., s, sh, f, th) at beginnings andendings of speech sentences to be included in the broadcast, that otherwise would have missed due to VAD bandwidth limitation.
[0035] FIG. 1 is a schematic perspective view of an exemplary in-ear hearing protection device 100, in accordance with embodiments herein. The in-ear hearing protection device 100 may be suitable for fitting into the concha of a human ear of a user (not shown). In some embodiments, the in-ear hearing protection device 100 is a device that substantially prevents an ambient sound 108 from directly entering the ear canal of the user. The in-ear hearing protection device 100 may include electronic components that receive the ambient sound 108, convert the ambient sound 108 to electronic signals, process the electronic signals, convert the processed electronic signals into a processed sound, and then emit the processed sound through a speaker, as described in detail later herein. The term “in-ear hearing protection device 100” is interchangeably used herein with the term “hearing protection device 100”.
[0036] It should be noted that a shape and a design of the in-ear hearing protection device 100 as shown in FIG. 1 is used for exemplary illustration only, and other shapes and designs for the hearing protection device 100 may also be contemplated without limiting the scope of the present disclosure.
[0037] The in-ear hearing protection device 100 is in the form of an earbud. In some embodiments, the in-ear hearing protection 100 includes an ear tip 112 and an earpiece body 114. The earpiece body 114 is configured (i.e., shaped and sized) to fit into the concha of the ear of the user. At least a portion of the ear tip 112 is configured (i.e., shaped and sized and is comprised of a material of suitable softness) to fit into at least a portion of the ear canal of the user. In some embodiments, the ear tip 112 is detachably attached to the earpiece body 114, such that the ear tip 112 can be removed and cleaned or replaced if desired. The ear tip 112 may include a through-passage that allows passage of sound therethrough. Fitting of at least a portion of the ear tip 112 into at least a portion of the ear canal externally occludes the ear canal, thereby preventing the ambient sound 108 from traveling along the ear canal so as to reach the inner ear.
[0038] In some embodiments, the ear tip 112 may include a body of which at least major portions are resiliently compressible and / or deformable at least in a radially inward direction, such that when the ear tip 112 is inserted into the ear canal, at least some portions of the ear tip 112 are resiliently biased radially outward so that at least some radially outward surfaces of the ear tip 112 are held against portions of the ear canal walls to substantially or completely eliminate any air gap therebetween. Thus, the ear tip 112 substantially prevents the ambient sound 108 from traveling down the ear canal in a space between the ear tip 112 and the ear canal walls.
[0039] In some embodiments, the ear tip 112 may consist of a single (e.g., molded) piece of organic polymeric material, e.g., a resiliently deformable and / or compressible material. However, it is expressly contemplated that it may not be necessary that all, or even any, of the material of which the ear tip 112 is made must be significantly compressible, as long as at least certain components of the ear tip 112 are resiliently deformable and are provided in geometric shapes that allow such deformation to provide the desired resilient biasing of surfaces of such components against the ear canal walls.The in-ear hearing protection device 100 further includes a flange 110. Specifically, the ear tip 112 may include one or more radially -outward-protmding flanges 110 made of a resiliently deformable material. Insertion of the ear tip 112 into the ear canal may result in the flange 110 being deformed, thereby resiliently biasing surfaces of the flange 110 against the ear canal walls. In some embodiments, the flange 110 may be at least generally semi-hemispherical in shape. In some embodiments, the flange 110 may be made of rubber, silicon, or a polymer, such as polyurethane.
[0040] In some embodiments, the earpiece body 114 includes a housing comprised of a molded polymeric material. In some embodiments, the housing may be formed by mating together two major housing portions. In some embodiments, the housing is hollow to at least partially define an interior space, which may contain any suitable electronic components, one or more internal batteries, and so on.
[0041] The in-ear hearing protection 100 further includes an in-ear microphone 102 configured to be positioned on an ear-canal side 104 of the in-ear hearing protection device 100. In some embodiments, the in-ear microphone 102 is configured to pick up a speech of the user. In some embodiments, the in-ear microphone 102 may be positioned on or within the ear tip 112 of the hearing protection device 100. However, it is expressly contemplated that the in-ear microphone 102 may be positioned at any other location on the hearing protection device 100 based on application requirements. The flange 110 is configured to provide a seal between the in-ear microphone 102 and the ambient sound 108. The in-ear hearing protection device 100 further includes an external microphone 106 configured to capture the ambient sound 108 external to the in-ear hearing protection device 100. In some embodiments, the external microphone 106 is positioned on the earpiece body 114.
[0042] FIG. 2 is a block diagram of an example in-ear hearing protection device 200, in accordance with embodiments herein. In some embodiments, the in-ear hearing protection device 200 is functionally similar to the in-ear hearing protection device 100 shown in FIG. 1. The term in-ear hearing protection device 200 is interchangeably referred to herein as “the hearing protection device 200”.
[0043] The in-ear hearing protection device 200 includes an in-ear microphone 202 configured to be positioned on an ear-canal side (e.g., the ear canal side 104 shown in FIG. 1) of the in-ear hearing protection device 200. In some embodiments, the in-ear microphone 202 is configured to pick up a speech of a user of the hearing protection device 200. In some embodiments, the in-ear microphone 202 is configured to generate an in-ear sound signal 260 based on the speech of the user.
[0044] The in-ear hearing protection device 200 further includes an external microphone 206 configured to capture an ambient sound 208 external to the in-ear hearing protection device 200. In some embodiments, the external microphone 206 is configured to generate an external sound signal 262 based on the ambient sound 208 external to the in-ear hearing protection device 200. In some embodiments, the in-ear microphone 202 or the external microphone 206 includes a voice-pick up (VPU) sensor 209. The VPU sensor 209 may be an air-borne vibration sensor or a solid-borne vibration sensor capable of detecting the ambient sound 208. In some embodiments, the VPU bone sensor 209 may capture the speech of the user using bone conduction (i.e., vibrations transmitted through the bonesin the skull). In the illustrated embodiment of FIG. 2, both the in-ear microphone 202 and the external microphone 206 includes the VPU bone sensor 209.
[0045] The in-ear hearing protection device 200 further includes a sound processing unit 216 configured to receive a sound signal 218 from either the in-ear microphone 202 or the external microphone 206 and generate a processed sound signal 220. In some embodiments, the sound signal 218 includes a mix of the in-ear sound signal 260 and the external sound signal 262. Alternatively, in some embodiments, the received sound signal 218 is received from the in-ear microphone 202 only. Alternatively, in some embodiments, the received sound signal 218 is received from the external microphone 206 only. In some embodiments, the sound signal 218 may include the speech of the user and / or loud noise from the surroundings. The in-ear hearing protection device 200 further includes a speaker 232 configured to broadcast the processed sound signal 220.
[0046] The sound processing unit 216 includes a sound filter 221 configured to filter the received sound signal 218 into a first plurality of frequency bands 222 (i.e., frequency bins). For example, the sound filter 221 may utilize time-domain techniques (e.g., a parallel array of biquad filters) or frequency -domain techniques (e.g., a Fast Fourier Transform or FFT) for generating the first plurality of frequency bands 222. In the case of time-domain, the first plurality of frequency bands 222 may begin with a 100 Hertz-centered band with a logarithmic progression of widths. In the case of frequencydomain, a first band may start at 0 Hertz (Hz) and the widths are all the same. The width may depend on processing resources available. In some embodiments, the first plurality of frequency bands 222 may be generated corresponding to each of the in-ear sound signal 260 and the external sound signal 262 received from the in-ear microphone 202 and the external microphone 206, respectively. The sound processing unit 216 further includes a voice activity detector (VAD) 228 configured to detect a speech 230 in one or more of the first plurality of frequency bands 222. Specifically, the VAD 228 may detect a beginning and an end of the speech 230. This may allow the sound processing unit 216 to apply agate to reduce noise in the sound signal 218 during speech pauses.
[0047] The sound processing unit 216 further includes a noise module 213 configured to detect a noise level 215 of one or more of the first plurality of filter bands 222. In some embodiments, the noise module 213 may determine the noise level 215 of the one or more of the first plurality of filter bands 222 corresponding to each of the in-ear sound signal 260 and the external sound signal 262. In some embodiments, the noise module 213 may determine the noise level 215 of the one or more of the first plurality of filter bands 222 based on the detection of the speech 230 by the VAD 228, e.g., in between the speech pauses.
[0048] In some embodiments, the noise module 213 is further configured to determine if the noise level 215 of the one or more of the first plurality of filter bands 222 is above a noise threshold 217. For example, the noise module 213 may determine if the noise level 215 of the one or more of the first plurality of filter bands 222 corresponding to each of the in-ear sound signal 260 and the external sound signal 262 is above the noise threshold 217. Accordingly, the sound processing unit 216 may determinethe contribution of the in-ear sound signal 260 and the external sound signal 262 in the processed sound signal 220.
[0049] In some embodiments, the noise threshold 217 may be determined as a percentage of the expected speech energy. Thus, if the noise module 213 finds that the noise energy in the external sound signal 262 is, for example, above 15% of what can be expected clean speech energy to be, then the noise module 213 may progressively back off the contribution of the external sound signal 262 (or proportionally throttle the external sound signal 262 above the 15% threshold) in the processed sound signal 220, and instead rely more on the in-ear sound signal 260.
[0050] For example, in some cases where the noise level 215 of the one or more of the first plurality of filter bands 222 of the external sound signal 262 is above the noise threshold 217, the sound processing unit 216 may utilize only the in-ear sound signal 260 for generating the processed sound signal 220. In some embodiments, the noise threshold 217 may be determined based on a characterization of the speech of the user during an earlier fit test, or even learned over time. For example, the noise threshold 217 may be determined based on historic speech characterization information 219 of the user. The historic speech characterization information 219 may be stored in a memory or a remote server accessible to the hearing protection device 200. Alternatively, the noise threshold 217 may have a default value based on generalization of human speech.
[0051] In some embodiments, the sound processing unit 216 may mix sound from the in-ear sound signal 260 and the external sound signal 262 in the processed sound signal 220 based on the detected noise level 215. For example, the sound processing unit 216 is configured to determine a ratio of the noise level 215 of the one or more of the first plurality of filter bands 222 for the in-ear sound signal 260 to the external sound signal 262 to determine a contribution of the in-ear sound signal 260 and the external sound signal 262 in the processed sound signal 220.
[0052] In some cases where the one or more of the first plurality of filter bands 222 of the in-ear sound signal 260 is above a predetermined threshold (i.e., the user speaking loudly), the external sound signal 262 is progressively mixed out of the processed sound signal 220 because presumably intelligibility will be acceptable if the speech 230 is loud. This may reduce a variable noise in the processed sound signal 220.
[0053] The sound processing unit 216 further includes a frequency gain modifier 236 configured to apply a frequency gain rule 240 to one or more of the first plurality of filter bands 222, thereby generating a second plurality of filter bands 246. For example, the frequency gain modifier 236 may amplify the one or more of the first plurality of filter bands 222 for spectral balancing. In some embodiments, the frequency gain rule 240 for a first frequency band 224 of the first plurality of filter bands 222 is selected based on the detected noise level 215 of the first frequency band 224. In some embodiments, the frequency gain modifier 236 applies a first gain rule 242 to the first frequency band 224 and a second gain rule 244 to a second frequency band 226 of the first plurality of filter bands 222. In some embodiments, the first gain rule 242 is different from the second gain rule 244.In some embodiments, the frequency gain modifier 236 includes a level dependent function 238. In some embodiments, the level dependent function 238 may filter out undesired sound content, for example, actively reducing the construction noises while providing human speech at substantially unchanged levels.
[0054] The sound processing unit 216 further includes a reconstructor 248 configured to recombine the second plurality of filter bands 246 to obtain the processed sound signal 220. For example, in case of time domain, the reconstructor 248 may sum the bands, and in case of frequency domain, the reconstructor 248 may perform an Inverse Fast Fourier Transform. The sound processing unit 216 further includes a crossfader 250 configured to crossfade the processed sound signal 220. In some embodiments, the crossfader 250 is a square-factor crossfader. Square factor crossfading is closer to the way hearing perception of human beings work as compared to linear fading, and thus, the processed sound signal 220 may appear more natural. Additionally, while a classic perceptual model of hearing follows an exponential curve, a square factor curve is a computationally -efficient substitute. However, it is expressly contemplated that other types of crossfaders (e.g., linear, exponential, S-curve, etc.) may also be utilized based on application requirements. In some examples, the crossfader 250 is configured to smooth the mixing of the external sound signal 262 into the processed sound signals 220 while the speech 230 and the noise are varying. In some embodiments, the sound processing unit 216 may further utilize active noise cancellation (ANC) and noise estimations to intelligently mix the external sound signal 262 into the processed sound signals 220. This may reduce a static noise and enhance intelligibility.
[0055] In some embodiments, if the noise level 215 of the one or more of the first plurality of filter bands 222 is above the noise threshold 217, the crossfader 250 reduces a contribution of the external sound signal 262 in the processed sound signal 220, thereby enabling the sound processing unit 216 to reduce noise in the processed sound signal 220. If no speech 230 is detected by the VAD 228, no processed sound is broadcast by the speaker 232.
[0056] In some embodiments, the in-ear hearing protection device 200 further includes a delay module 252 configured to delay broadcasting the processed sound signal 220 by a time delay 256. For example, the delay module 252 is configured to delay broadcasting the processed sound signal 220 by a predetermined time interval, e.g., 1 second. Delaying the processed sound signal 220 may allow high-frequency fricatives (e.g., s, sh, f, th) at endings of speech sentences to pass, which otherwise would have missed due to application of the VAD 228.
[0057] In some embodiments, the delay module 252 is further configured to, when the speech 230 is detected by the VAD 228, retrieve a previous sound signal 258 captured prior to the received sound signal 218. For example, the delay module 252 is configured to retrieve the previous sound signal 258 for a predetermined time interval, e.g., 1 second. The previous sound signal 258 may be stored in a memory buffer 234.The delay module 252 is further configured to generate a modified processed sound signal 254 including the previously captured sound signal 258 and the processed sound signal 220. This may allow high-frequency fricatives (e.g., s, sh, f, th) at beginning of speech sentences to pass (or be detected and provided audibly to the user), which otherwise would have missed due to application of the VAD 228. In some embodiments, the delay module 252 provides the modified processed sound signal 254 within 1 millisecond (ms) of the sound processing unit 216 receiving the sound signal 218. In some embodiments, the delay module 252 provides the modified processed sound signal 254 within 10 ms, 20 ms, 30 ms, 50 ms, 70 ms, or 100 ms of the sound processing unit 216 receiving the sound signal 218. In some embodiments, the delay module 252 is further configured to provide the modified processed sound signal 254 to the crossfader 250.
[0058] FIG. 3 is a block diagram of an in-ear hearing protection device 300, according to another embodiment of the present disclosure. In some embodiments, the in-ear hearing protection device 300 is functionally similar to the in-ear hearing protection devices 100, 200 shown in FIGS. 1 and 2, respectively. The term in-ear hearing protection device 300 is interchangeably referred to herein as “the hearing protection device 300”.
[0059] The in-ear hearing protection device 300 includes a sound receiver 301. The sound receiver 301 is configured to generate a sound signal 318 based on detection of a speech of a user of the in-ear hearing protection device 300 or an ambient sound 308 external to the hearing protection device 300. In some embodiments, the sound receiver 301 is at least one of an in-ear microphone, an out-of-ear microphone, and a voice pickup (VPU). Specifically, the sound receiver 301 is an in-ear microphone disposed on the in-ear hearing protection device 300 and configured to capture the speech of the user.
[0060] In some embodiments, the sound receiver 301 is a first microphone 302 and the sound signal 318 is a first sound signal 360. In some embodiments, the in-ear hearing protection device 300 further includes a second microphone 306 configured to generate a second sound signal 362. In some embodiments, the second microphone 306 is an external microphone configured to capture the ambient sound 308 external to the in-ear hearing protection device 300.
[0061] The in-ear hearing protection device 300 further includes a sound filter 321 configured to process the received sound signal 318 to obtain a processed sound signal 320. In some embodiments, the in-ear hearing protection device 300 further includes a voice activity detector (VAD) 328. In some embodiments, the VAD 328 is configured to detect a speech 330 of the user. Specifically, the sound filter 321 includes the VAD 328 that may detect a beginning and an end of the speech 330. This may allow the sound fdter 321 to apply agate to reduce noise in the sound signal 318 during speech pauses.
[0062] In some embodiments, the sound filter 321 is configured to filter the sound signal 318 into a plurality of frequency bands 322 (i.e., frequency bins). Specifically, the sound filter 321 is configured to filter the first sound signal 360 into the plurality of frequency bands 322 (i.e., frequency bins). In some embodiments, the sound filter 321 applies the VAD 328 separately to each of the plurality of frequency bands 322. In some embodiments, the sound filter 321 is further configured to classify, basedon the application of the VAD 328, a sound type 364 in the at least one frequency band 322 of the plurality of frequency bands 322. In some embodiments, the sound type 364 includes the speech 330 or an alarm 366 (e.g., a noise in the surroundings).
[0063] Similarly, in some embodiments, the sound filter 321 is further configured to filter the second sound signal 362 into a plurality of frequency bands 323 (i.e., frequency bins) corresponding to the plurality of frequency bands 322 of the first sound signal 360. In some embodiments, the sound filter 321 applies the VAD 328 separately to each of the plurality of frequency bands 323.
[0064] In some embodiments, the hearing protection device 300 (or the sound filter 321) further includes a noise module 313 configured to detect a noise level 315 in at least one frequency band 322 of the plurality of frequency bands 322. In some embodiments, the noise module 313 is further configured to detect the noise level 315 in each of the plurality of frequency bands 322. In some embodiments, the noise module 313 may determine the noise level 315 in each of the plurality of frequency bands 322 based on the detection of the speech 330 by the VAD 328, e.g., in between the speech pauses.
[0065] In some embodiments, the noise module 313 is further configured to determine that the noise level 315 in at least one frequency band 322 of the plurality of frequency bands 322 is above a noise threshold 317. In some embodiments, the noise module 313 is further configured to throttle a contribution of the sound receiver 301 (i.e., the first microphone 302) in the processed sound signal 320 based on the determination that the noise level 315 in the at least one frequency band 322 is above the noise threshold 317.
[0066] In some embodiments, the noise module 313 is further configured to compare one frequency band 322, 323 of the plurality of frequency bands 322, 323 of the first sound signal 360 and the second sound signal 362. In some embodiments, the noise module 313 is further configured to determine that a microphone noise threshold 368 of the first microphone 302 is above a predetermined threshold 370 based on the comparison of the one frequency band 322, 323. In some embodiments, the noise module 313 is further configured to throttle a contribution of the second microphone 306 to the processed sound signal 320 based on the determination that the microphone noise threshold 368 of the first microphone 302 is above the predetermined threshold 370.
[0067] Accordingly, the sound filter 321 may determine the contribution of the first microphone 302 and the second microphone 306 in the processed sound signal 320. For example, in some cases where the microphone noise threshold 368 of the first microphone 302 is above the predetermined threshold 370, the sound filter 321 may mix the contribution of the second microphone 306 (i.e., the second sound signal 362) in the processed sound signal 320. In some embodiments, the predetermined threshold 370 may be determined based on a characterization of a speech of the user during an earlier fit test, or even learned over time. Alternatively, the predetermined threshold 370 may have a default value based on generalization of human speech.In some embodiments, the hearing protection device 300 (or the sound filter 321) further includes a frequency gain modifier 336 configured to apply a frequency gain rule 340 to each of the plurality of frequency bands 322. In some examples, the frequency gain rule 340 includes at least one of increasing a gain, reducing a gain, and applying a limiter. For example, the frequency gain modifier 336 may amplify each of the plurality of frequency bands 322 for spectral balancing. In some embodiments, the frequency gain mle 340 is based on the noise level 315 in the at least one frequency band 322 of the plurality of frequency bands 322.
[0068] In some embodiments, the frequency gain modifier 336 is configured to apply a level-dependent function 338 to the sound signal 318. In some embodiments, the level-dependent function 338 may filter out undesired, or unsafe, sound content, for example, actively reducing the construction noises while providing human speech at substantially unchanged levels.
[0069] In some embodiments, the hearing protection device 300 (or the sound filter 321) further includes a reconstructor 348 configured to reconstruct the plurality of frequency bands 322 to generate the processed sound signal 320. In some embodiments, the hearing protection device 300 (or the sound filter 321) further includes a speaker 332 configured to broadcast the processed sound signal 320. In some embodiments, the hearing protection device 300 (or the sound filter 321) further includes a crossfader 350 configured to crossfade the processed sound signal 320. In some embodiments, the crossfader 350 is a square-factor crossfader.
[0070] The in-ear hearing protection device 300 further includes a delay module 352 configured to delay the broadcasting of the processed sound signal 320 by a time delay 356. In some embodiments, the time delay 356 is less than 1 second. For example, the delay module 352 is configured to delay broadcasting the processed sound signal 320 by a predetermined time interval, e.g., 1 second. Delaying the processed sound signal 320 may allow high-frequency fricatives (e.g., s, sh, f, th) at endings of speech sentences to be provided, which, without the VAD 328, may have been unintentionally cut off.
[0071] In some embodiments, the delay module 352 is further configured to retrieve a previous sound signal 358 captured prior to the received sound signal 318 based on the detection that the sound signal 318 includes the speech 330. In some embodiments, the previous sound signal 358 is captured within 1 second of the received sound signal 318. The previous sound signal 258 may be stored in a memory buffer 334.
[0072] The delay module 352 is further configured to generate a modified processed sound signal 354 including the previous sound signal 358 and the processed sound signal 320. This may allow high-frequency fricatives (e.g., s, sh, f th) at beginning of speech sentences to be captured and processed, which otherwise would have missed due to application of the VAD 328. The delay module 352 is further configured to provide the modified processed sound signal 354 to the crossfader 350. Subsequently, the speaker 332 may broadcast the modified processed sound signal 354.FIG. 4 illustrates a block diagram of a method 400 for improving fidelity of speech in a hearing protection device (e.g., any of the hearing protection devices 100, 200, 300 shown in FIGS. 1-3, or another suitable hearing protection device), according to embodiments herein.
[0073] At block 402, the method 400 includes receiving a sound signal from a sound receiver. In some embodiments, the sound receiver is at least one of an in-ear microphone, an out-of-ear microphone, and a voice pickup (VPU). In some embodiments, the sound receiver is an in-ear microphone disposed on the in-ear hearing protection device. In some embodiments, the sound receiver is a first microphone and the sound signal is a first sound signal. For example, as shown in FIG. 3, the in-ear hearing protection device 300 may include the sound receiver 301 for receiving the sound signal 318, where the sound receiver 301 is the first microphone 302 and the sound signal 318 is the first sound signal 360.
[0074] At block 404, the method 400 further includes processing, by a sound filter, the sound signal to obtain a processed sound signal. For example, with respect to the in-ear hearing protection device 300 illustrated in FIG. 3, the sound filter processes the received sound signal to obtain the processed sound signal.
[0075] At block 406, the processing includes applying a voice activity detector (VAD) to the sound signal. For example, as shown in FIG. 3, the sound filter 321 includes the VAD 328.
[0076] At block 408, the processing further includes filtering the sound signal into a plurality of frequency bands. For example, with respect to the in-ear hearing protection device 300 illustrated in FIG. 3, the sound filter is filters the sound signal into the plurality of frequency bands (i.e., frequency bins).
[0077] In some embodiments, the method 400 further includes applying the VAD separately to each of the plurality of frequency bands. In some embodiments, the method 400 further includes classifying, based on the application of the VAD, a sound type in at least one frequency band of the plurality of frequency bands. The sound type could include, for example, detected speech or a detected alarm, or other noise. For example, with respect to the in-ear hearing protection device 300 illustrated in FIG. 3, the sound filter applies the VAD separately to each of the plurality of frequency bands, where the sound filter is may classify, based on the application of the VAD, the sound type in the at least one frequency band of the plurality of frequency bands.
[0078] At block 410, the processing further includes detecting a noise level in at least one frequency band of the plurality of frequency bands. In some embodiments, the method 400 further includes detecting the noise level in each of the plurality of frequency bands. For example, with respect to the in-ear hearing protection device 300 illustrated in FIG. 3, the noise module detects the noise level in the at least one frequency band of the plurality of frequency bands. In some embodiments, the noise module detects the noise level in each of the plurality of frequency bands.
[0079] In some embodiments, the method 400 further includes determining that the noise level in the at least one frequency band is above a noise threshold. The method 400 further includes throttling a contribution of the second receiver in the processed sound signal based on the determination that thenoise level in the at least one frequency band is above the noise threshold. For example, with respect to the in-ear hearing protection device 300 illustrated in FIG. 3, the noise module determines that the noise level in the at least one frequency band is above the noise threshold. The noise module may also throttle a contribution of the sound receiver in the processed sound signal based on the determination that the noise level in the at least one frequency band is above the noise threshold.
[0080] In some embodiments, the method 400 further includes receiving a second sound signal from a second microphone. The method 400 further includes filtering the second sound signal into a plurality of frequency bands corresponding to the plurality of frequency bands of the first sound signal. For example, with respect to the in-ear hearing protection device 300 illustrated in FIG. 3, the in-ear hearing protection device includes the second microphone that generates the second sound signal. The sound filter filters the second sound signal into the plurality of frequency bands.
[0081] In some embodiments, the method 400 further includes comparing one frequency band of the plurality of frequency bands of the first sound signal and the second sound signal. The method 400 further includes determining that a microphone noise threshold of the first microphone is above a predetermined threshold based on the comparison of the one frequency band. The method 400 further includes throttling a contribution of the second microphone to the processed sound signal based on the determination that the microphone noise threshold of the first microphone is above the predetermined threshold.
[0082] For example, with respect to the in-ear hearing protection device 300 illustrated in FIG. 3, the noise module may compare the one frequency band of the plurality of frequency bands of the first sound signal and the second sound signal. The noise module may also determine that the microphone noise threshold of the first microphone is above the predetermined threshold based on the comparison of the one frequency band. The noise module may throttle a contribution of the second microphone to the processed sound signal based on the determination that the microphone noise threshold of the first microphone is above the predetermined threshold.
[0083] At block 412, the processing further includes applying a frequency gain rule to each of the plurality of frequency bands. In some embodiments, the frequency gain rule is based on a detection that the noise level in the at least one frequency band of the plurality of frequency bands is above a noise threshold. In some embodiments, the frequency gain mle includes at least one of increasing a gain, reducing a gain, and applying a limiter. In some embodiments, processing the sound signal further includes applying a level-dependent function to the sound signal.
[0084] For example, with respect to the in-ear hearing protection device 300 illustrated in FIG. 3, the frequency gain modifier may apply the frequency gain rule to each of the plurality of frequency bands. The frequency gain rule is based on the noise level in the at least one frequency band of the plurality of frequency bands. The frequency gain modifier may also apply the level-dependent function to the sound signal.At block 414, the processing further includes reconstructing the plurality of frequency bands to generate the processed sound signal. In some embodiments, the processing further includes applying a square-factor crossfader. For example, with respect to the in-ear hearing protection device 300 illustrated in FIG. 3, the reconstructor may reconstruct the plurality of frequency bands to generate the processed sound signal. The crossfader may crossfade the processed sound signal.
[0085] At block 416, the method 400 further includes broadcasting the processed sound signal. For example, with respect to the in-ear hearing protection device 300 of FIG. 3, the speaker may broadcast the processed sound signal 320.
[0086] In some embodiments, broadcasting the processed sound signal, as illustrated in block 416, further includes delaying broadcasting the processed sound signal by a time delay. In some embodiments, the time delay is less than 1 second. Broadcasting the processed sound signal further includes detecting, based on the processed sound signal, that the sound signal includes a speech. Broadcasting the processed sound signal further includes retrieving a previous sound signal captured prior to the received sound signal based on the detection that the sound signal includes the speech. In some embodiments, the previous sound signal is captured within 1 second of the received sound signal. Broadcasting the processed sound signal further includes generating a modified processed sound signal including the previous sound signal and the processed sound signal. Broadcasting the processed sound signal further includes broadcasting the modified processed sound signal.
[0087] For example, with respect to the in-ear hearing protection device 300 of FIG. 3, the delay module may delay the broadcasting of the processed sound signal by the time delay. The VAD may detect that the sound signal includes the speech. The delay module may retrieve the previous sound signal captured prior to the received sound signal based on the detection that the sound signal includes the speech. The delay module may generate the modified processed sound signal including the previous sound signal and the processed sound signal. The delay module may provide the modified processed sound signal to the crossfader. Subsequently, the speaker may broadcast the modified processed sound signal.
[0088] While the method 400 is described in the context of FIG. 3, it is expressly contemplated that the in-ear hearing protection device 300 shown in FIG. 3 may operate in a process other than that of the method 400. It is further contemplated that the method 400 may be practiced with other suitable devices and / or systems.
[0089] FIG. 5 is a block diagram illustrating a method 500 for improving fidelity of speech in a hearing protection device (e.g., the hearing protection device 100, 200, 300 shown in FIGS. 1-3, or another suitable hearing protection device), according to embodiments herein.
[0090] At block 502, the method 500 includes receiving a sound signal from a sound receiver. In some embodiments, the sound receiver is at least one of an in-ear microphone, an out-of-ear microphone, and a voice pickup (VPU). In some embodiments, the sound receiver is an in-ear microphone. In some embodiments, the sound receiver is a first microphone and the sound signal is a first sound signal. Forexample, as shown in FIG. 3, the in-ear hearing protection device 300 may include the sound receiver 301 for receiving the sound signal 318, where the sound receiver 301 is the first microphone 302 and the sound signal 318 is the first sound signal 360.
[0091] At block 504, the method 500 further includes processing the sound signal to obtain a processed sound signal. Processing may include applying a voice activity detector (VAD) to the sound signal. For example, with respect to the in-ear hearing protection device 300 illustrated in FIG. 3, the sound filter may process the received sound signal to obtain the processed sound signal. The sound filter may include the VAD. Other processing steps may occur as well.
[0092] At block 506, the method 500 further includes delaying broadcasting the processed sound signal by a time delay. In some embodiments, the time delay is less than 1 second. For example, with respect to the in-ear hearing protection device 300 illustrated in FIG. 3, the delay module may delay the broadcasting of the processed sound signal by the time delay.
[0093] At block 508, the method 500 further includes detecting, based on the processed sound signal, that the sound signal includes a speech. For example, with respect to the in-ear hearing protection device 300 illustrated in FIG. 3, the VAD may detect the speech.
[0094] At block 510, the method 500 further includes retrieving a previous sound signal captured prior to the received sound signal based on the detection that the sound signal includes the speech. In some embodiments, the previous sound signal was captured within 1 second of the received sound signal. For example, with respect to the in-ear hearing protection device 300 illustrated in FIG. 3, the delay module may retrieve the previous sound signal captured prior to the received sound signal based on the detection that the sound signal includes the speech.
[0095] At block 512, the method 500 further includes generating a modified processed sound signal including the previous sound signal and the processed sound signal. For example, with respect to the in-ear hearing protection device 300 illustrated in FIG. 3, the delay module may generate the modified processed sound signal including the previous sound signal and the processed sound signal.
[0096] At block 514, the method 500 further includes broadcasting the modified processed sound signal. For example, with respect to the in-ear hearing protection device 300 illustrated in FIG. 3, the speaker may broadcast the modified processed sound signal.
[0097] While the method 500 is described in the context of FIG. 3, it is expressly contemplated that the in-ear hearing protection device 300 shown in FIG. 3 may operate in a process other than that of the method 500. It is further contemplated that the method 500 may be practiced with other suitable devices and / or systems.
[0098] FIG. 6 illustrates a block diagram of a method 600 for improving fidelity of speech in a hearing protection device (e.g., any of the hearing protection devices 100, 200, 300 shown in FIGS. 1-3, or another suitable hearing protection device), according to embodiments herein.
[0099] At block 602, the method 600 includes receiving a sound signal from a sound receiver. In some embodiments, the sound receiver is at least one of an in-ear microphone, an out-of-ear microphone, anda voice pickup (VPU). In some embodiments, the sound receiver is an in-ear microphone. In some embodiments, the sound receiver is a first microphone and the sound signal is a first sound signal. For example, as shown in FIG. 3, the in-ear hearing protection device 300 may include the sound receiver 301 for receiving the sound signal 318, where the sound receiver 301 is the first microphone 302 and the sound signal 318 is the first sound signal 360.
[0100] At block 604, the method 600 further includes processing, by a sound filter, the sound signal to obtain a processed sound signal. The processing includes applying a voice activity detector (VAD) to the sound signal. The processing further includes applying a square-factor crossfader. For example, with respect to the in-ear hearing protection device 300 illustrated in FIG. 3, the sound filter may process the received sound signal to obtain the processed sound signal. The sound filter includes the VAD. The crossfader may crossfade the processed sound signal.
[0101] At block 606, the method 600 further includes broadcasting the processed sound signal. For example, with respect to the in-ear hearing protection device 300 illustrated in FIG. 3, the speaker may broadcast the processed sound signal.
[0102] While the method 600 is described in the context of FIG. 3, it is expressly contemplated that the in-ear hearing protection device 300 shown in FIG. 3 may operate in a process other than that of the method 600. It is further contemplated that the method 600 may be practiced with other suitable devices and / or systems.
[0103] FIG. 7 illustrates a remote server architecture 700 for a sound processing unit 710 of an in-ear hearing protection device (e.g., any of the in-ear hearing protection devices 100, 200, 300 shown in FIGS. 1, 2, and 3, respectively) when hosted in a cloud-based architecture. The remote server architecture 700 illustrates one embodiment of an implementation of the sound processing unit 710. As an example, the remote server architecture 700 may provide computation, software, data access, and storage services that do not require end-user knowledge of a physical location or configuration of the sound processing unit 710 that delivers the services. In various embodiments, the remote server architecture 700 may deliver services over a wide area network, such as the internet, using appropriate protocols. For instance, the remote server architecture 700 may deliver applications over a wide area network and they can be accessed through a web browser or any other computing component.
[0104] Software or components shown or described in FIGS. 1-6 as well as the corresponding data may be stored on servers at a remote location. The computing resources in a remote server environment may be consolidated at a remote data center location or they may be dispersed. The remote server architecture 700 may deliver services through shared data centers, even though they appear as a single point of access for a user. Thus, components and functions described herein may be provided from a remote server at a remote location using the remote server architecture 700. Alternatively, they can be provided by a conventional server, installed on client devices directly, or in other ways.
[0105] In the example shown in FIG. 7, some items are similar to those shown in earlier figures. FIG.
[0106] 7 specifically shows that the sound processing unit 710 may be located at a remote server location 702.Therefore, a computing device 720 accesses the sound processing unit 710 through the remote server location 702. A user 750 can use the computing device 720 to access a user interface 722 as well. Embodiments described herein have focused on systems and methods that automatically retrieve data to process a sound signal, based on that processing, generate a processed sound signal, all in-situ. However, it is expressly contemplated that the retrieved data, the processing, or other information may be presented on the user interface 722 for action or approval from the user 750. .
[0107] FIG. 7 shows that it is also contemplated that some elements of systems described herein are disposed at the remote server location 702 while others are not. By way of example, historic speech characterization information 740, previous sound signal 730 can be disposed at a location separate from remote server location 702 and accessed through the remote server location 702. Regardless of where they are located, they can be accessed directly by the computing device 720, through a network (either a wide area network or a local area network), hosted at a remote site by a service, provided as a service, or accessed by a connection service that resides in a remote location. Also, data can be stored in substantially any location and intermittently accessed by, or forwarded to, interested parties. For instance, physical carriers may be used instead of, or in addition to, electromagnetic wave carriers.
[0108] It will also be noted that elements of systems described herein, or portions of them, can be disposed on a wide variety of different devices. Some of those devices include servers, desktop computers, laptop computers, imbedded computer, industrial controllers, tablet computers, or other mobile devices, such as palm top computers, cell phones, smart phones, multimedia players, personal digital assistants, etc. Any suitable computing device with a display may be able to service as the computing device 720 with the user interface 722.
[0109] FIG. 8 is a block diagram of a computing environment 800 that may be used in embodiments shown in previous Figures. Specifically, FIG. 8 is one example of the computing environment 800 in which elements of systems and methods described herein, or parts of them, may be deployed. With reference to FIG. 8, an example system for implementing some embodiments includes a general-purpose computing device in the form of a computer 810. Components of the computer 810 may include, but are not limited to, a processing unit 820 (which may include a processor), a system memory 830, and a system bus 821 that couples various system components including the system memory 830 to the processing unit 820. The system bus 821 may be any of several types of bus stmctures, including a memory bus or memory controller, a peripheral bus, and a local bus using any of a variety of bus architectures. Memory and programs described with respect to systems and methods described herein may be deployed in corresponding portions of FIG. 8.
[0110] The computer 810 typically includes a variety of computer readable media. Computer readable media may be any available media that may be accessed by the computer 810 and includes both volatile / nonvolatile media and removable / non-removable media. By way of example, and not limitation, computer readable media may include computer storage media and communication media. Computer storage media is different from, and does not include, a modulated data signal or a carrierwave. The term “modulated data signal” means a signal that has one or more of its characteristics set or changed in such a manner as to encode information in the signal. Computer storage media includes hardware storage media, including both volatile / nonvolatile and removable / non-removable media implemented in any method or technology for storage of information, such as computer readable instructions, data stmctures, program modules or other data.
[0111] Computer storage media includes, but is not limited to, random access memory (RAM), read only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory, or other memory technology, CD-ROM, digital versatile disks (DVD), or other optical disk storage, magnetic cassettes, magnetic tape, magnetic disk storage, or other magnetic storage devices, or any other medium which may be used to store the desired information and which may be accessed by the computer 810. Communication media may embody computer readable instructions, data structures, program modules or other data in a transport mechanism and includes any information delivery media.
[0112] The system memory 830 includes computer storage media in the form of volatile and / or nonvolatile memory, such as read only memory (ROM) 831 and random access memory (RAM) 832. A basic input / output system 833 (BIOS) containing basic routines that helps to transfer information between elements within the computer 810, such as during start-up, is typically stored in the ROM 831. The RAM 832 typically contains data and / or program modules that are immediately accessible to and / or presently being operated on by the processing unit 820. By way of example, and not limitation, FIG. 8 illustrates an operating system 834, application programs 835, other program modules 836, and program data 837.
[0113] The computer 810 may also include other removable / non-removable and volatile / nonvolatile computer storage media. By way of example only, FIG. 8 illustrates a hard disk drive 841 that reads from or writes to non-removable, nonvolatile magnetic media, a nonvolatile magnetic disk, an optical disk drive 855, and a nonvolatile optical disk 856. The hard disk drive 841 is typically connected to the system bus 821 through a non-removable memory interface, such as an interface 840, and the optical disk drive 855 is typically connected to the system bus 821 by a removable memory interface, such as an interface 850.
[0114] Alternatively, or in addition, the functionality described herein may be performed, at least in part, by one or more hardware logic components. For example, and without limitation, illustrative types of hardware logic components that may be used include Field-programmable Gate Arrays (FPGAs), Application-specific Integrated Circuits (ASICs), Application-specific Standard Products (ASSPs), System-on-a-chip systems (SOCs), Complex Programmable Logic Devices (CPLDs), etc.
[0115] The drives and their associated computer storage media discussed above and illustrated in FIG.
[0116] 8, provide storage of computer readable instructions, data structures, program modules, and other data for the computer 810. In FIG. 8, e.g., the hard disk drive 841 is illustrated as storing an operating system 844, application programs 845, other program modules 846, and program data 847. Note that thesecomponents can either be the same as or different from operating system 834, application programs 835, the other program modules 836, and the program data 837.
[0117] Ausermay enter commands and information into the computer 810 through input devices, such as a keyboard 862, a microphone 863, and a pointing device 861, such as a mouse, trackball or touch pad. Other input devices (not shown) may include a joystick, a game pad, a satellite receiver, a scanner, or the like. These and other input devices are often connected to the processing unit 820 through a user input interface 860 that is coupled to the system bus 821, but may be connected by other interface and bus structures. A visual display 891 or other type of display device is also connected to the system bus 821 via an interface, such as a video interface 890. In addition to the monitor, the computer 810 may also include other peripheral output devices, such as speaker 897 and printer 896, which may be connected through an output peripheral interface 895.
[0118] The computer 810 is operated in a networked environment using logical connections, such as a Local Area Network (LAN) or a Wide Area Network (WAN) to one or more remote computers, such as a remote computer 880.
[0119] When used in a LAN networking environment, the computer 810 is connected to the LAN 871 through a network interface or adapter 870. When used in a WAN networking environment, the computer 810 typically includes a modem 872 or other means for establishing communications over a WAN 873, such as the Internet. In a networked environment, program modules may be stored in a remote memory storage device. FIG. 8 illustrates, e.g., that remote application programs 885 can reside on the remote computer 880.
[0120] The in-ear hearing protection device and the method of the present disclosure utilizes the sound filter to process the received sound signal. The processing includes filtering the sound signal in the plurality of frequency bands, and detecting the noise level in the at least one frequency band of the plurality of frequency bands. This may allow the sound filter to determine one or more frequency bands of the plurality of frequency bands that can be used to generate the processed sound signal. Additionally, processing of the sound signal includes applying the VAD to the sound signal. This may enable precise detection of speech starts and stops for reducing a noise in the processed sound signal, e.g., by applying a gate during speech pauses. In some cases, processing the sound signal further includes delaying broadcasting the processed sound signal and retrieving a previous sound signal to allow high-frequency fricatives (e.g., s, sh,f th) at beginnings and endings of speech sentences to be included in the broadcast, that otherwise would have missed due to VAD bandwidth limitation.
[0121] Unless otherwise indicated, all numbers expressing feature sizes, amounts, and physical properties used in the specification and claims are to be understood as being modified by the term “about”. Accordingly, unless indicated to the contrary, the numerical parameters set forth in the foregoing specification and attached claims are approximations that can vary depending upon the desired properties sought to be obtained by those skilled in the art utilizing the teachings disclosed herein.As used in this specification and the appended claims, the singular forms “a,” “an,” and “the” encompass embodiments having plural referents, unless the content clearly dictates otherwise. As used in this specification and the appended claims, the term “or” is generally employed in its sense including “and / or” unless the content clearly dictates otherwise.
[0122] Spatially related terms, including but not limited to, “proximate,” “distal,” “lower,” “upper,” “beneath,” “below,” “above,” and “on top,” if used herein, are utilized for ease of description to describe spatial relationships of an element(s) to another. Such spatially related terms encompass different orientations of the device in use or operation in addition to the particular orientations depicted in the figures and described herein. For example, if an object depicted in the figures is turned over or flipped over, portions previously described as below, or beneath other elements would then be above or on top of those other elements.
[0123] As used herein, when an element, component, or layer for example is described as forming a “coincident interface” with, or being “on,” “connected to,” “coupled with,” “stacked on” or “in contact with” another element, component, or layer, it can be directly on, directly connected to, directly coupled with, directly stacked on, in direct contact with, or intervening elements, components or layers may be on, connected, coupled or in contact with the particular element, component, or layer, for example. When an element, component, or layer for example is referred to as being “directly on,” “directly connected to,” “directly coupled with,” or “directly in contact with” another element, there are no intervening elements, components or layers for example.
[0124] Various examples have been described. These and other examples are within the scope of the following claims.
[0125] List of Illustrative embodiments:
[0126] Embodiment 1. A method for improving fidelity of speech in a hearing protection device, comprising receiving a sound signal from a sound receiver; processing the sound signal using a sound filter by applying a voice activity detector, filtering into multiple frequency bands, detecting a noise level in at least one band, applying a frequency gain rule to each band, reconstructing the bands to generate a processed sound signal; and broadcasting the processed sound signal.
[0127] Embodiment 2. The method of Embodiment 1, further comprising determining that the noise level in at least one frequency band is above a noise threshold and throttling the contribution of the sound receiver based on that determination.
[0128] Embodiment 3. The method of any of Embodiments 1-2, wherein the sound receiver is an in-ear microphone disposed on the hearing protection device, and the frequency gain rule is based on detecting that the noise level in at least one frequency band is above a noise threshold.
[0129] Embodiment 4. The method of any of Embodiments 1-3, wherein broadcasting the processed sound signal comprises delaying broadcasting by a time delay, detecting that the sound signal comprises speech, retrieving a previous sound signal captured prior to the received sound signal, generating amodified processed sound signal comprising both signals, and broadcasting the modified processed sound signal.
[0130] Embodiment 5. The method of any of Embodiments 1-4, wherein the time delay is less than 1 second. Embodiment 6. The method of any of Embodiments 1-5, wherein the previous sound signal is captured within 1 second of the received sound signal.
[0131] Embodiment 7. The method of any of Embodiments 1-6, wherein the sound signal is a first sound signal and the sound receiver is a first microphone, and further comprising receiving a second sound signal from a second microphone, filtering the second sound signal into corresponding bands, comparing a band of the first and second signals, determining that a microphone noise threshold of the first microphone is above a predetermined threshold based on the comparison, and throttling a contribution of the second microphone based on that determination.
[0132] Embodiment 8. The method of any of Embodiments 1-7, wherein the sound receiver is an in-ear microphone, an out-of-ear microphone, or a voice pickup.
[0133] Embodiment 9. The method of any of Embodiments 1-8, further comprising applying the VAD separately to each frequency band.
[0134] Embodiment 10. The method of any of Embodiments 1-9, further comprising classifying a sound type in at least one frequency band based on the VAD.
[0135] Embodiment 11. The method of any of Embodiments 1-10, wherein the sound type comprises speech or an alarm.
[0136] Embodiment 12. The method of any of Embodiments 1-11, further comprising detecting the noise level in each frequency band.
[0137] Embodiment 13. The method of any of Embodiments 1-12, wherein the frequency gain rule comprises increasing a gain, reducing a gain, or applying a limiter.
[0138] Embodiment 14. The method of any of Embodiments 1-13, wherein processing further comprises applying a level-dependent function.
[0139] Embodiment 15. The method of any of Embodiments 1-14, wherein processing further comprises applying a square-factor crossfader.
[0140] Embodiment 16. A method for improving fidelity of speech in a hearing protection device, comprising receiving a sound signal from a sound receiver, processing the sound signal by applying a VAD, delaying broadcasting by a time delay, detecting that the sound signal comprises speech, retrieving a previous sound signal captured prior to the received sound signal, generating a modified processed sound signal comprising both signals, and broadcasting the modified processed sound signal.
[0141] Embodiment 17. The method of Embodiment 16, wherein the time delay is less than 1 second.
[0142] Embodiment 18. The method of any of Embodiments 16-17, wherein the previous sound signal was captured within 1 second of the received sound signal.
[0143] Embodiment 19. The method of any of Embodiments 16-18, further comprising determining that the sound signal has a noise level above a noise threshold and throttling a contribution of the sound receiverbased on that determination.
[0144] Embodiment 20. The method of any of Embodiments 16-19, wherein processing the sound signal comprises filtering it into frequency bands, detecting noise in at least one band, applying a frequency gain rule to each band, and reconstructing the bands.
[0145] Embodiment 21. The method of any of Embodiments 16-20, wherein the sound receiver is an in-ear microphone and the frequency gain rule is based on detecting noise above a threshold in one band. Embodiment 22. The method of any of Embodiments 16-21, further comprising detecting noise in each frequency band.
[0146] Embodiment 23. The method of any of Embodiments 16-22, wherein the frequency gain rule comprises increasing a gain, reducing a gain, or applying a limiter.
[0147] Embodiment 24. The method of any of Embodiments 16-23, further comprising receiving a second sound signal from a second microphone, filtering it into corresponding bands, comparing a band of the first and second signals, determining that a microphone noise threshold of the first microphone is above a predetermined threshold, and throttling a contribution of the second microphone accordingly.
[0148] Embodiment 25. The method of any of Embodiments 16-24, wherein the sound receiver is an in-ear microphone, an out-of-ear microphone, or a voice pickup.
[0149] Embodiment 26. The method of any of Embodiments 16-25, further comprising applying the VAD separately to each frequency band.
[0150] Embodiment 27. The method of any of Embodiments 16-26, further comprising classifying a sound type in a frequency band based on the VAD.
[0151] Embodiment 28. The method of any of Embodiments 16-27, wherein the sound type is speech or an alarm.
[0152] Embodiment 29. The method of any of Embodiments 16-28, wherein processing further comprises applying a level-dependent function.
[0153] Embodiment 30. The method of any of Embodiments 16-29, wherein processing further comprises applying a square-factor crossfader.
[0154] Embodiment 31. A method for improving fidelity of speech in a hearing protection device, comprising receiving a sound signal, processing it by applying a VAD and a square-factor crossfader, and broadcasting the processed sound signal.
[0155] Embodiment 32. The method of Embodiment 31, further comprising determining that the sound signal has noise above a threshold and throttling a contribution of the sound receiver accordingly.
[0156] Embodiment 33. The method of any of Embodiments 31-32, wherein broadcasting comprises delaying broadcasting by a time delay, detecting that the sound signal comprises speech, retrieving a previous sound signal captured prior to the received sound signal, generating a modified processed sound signal comprising both signals, and broadcasting it.
[0157] Embodiment 34. The method of any of Embodiments 31-33, wherein the time delay is less than 1 second.Embodiment 35. The method of any of Embodiments 31-34, wherein the previous sound signal was captured within 1 second of the received sound signal.
[0158] Embodiment 36. The method of any of Embodiments 31-35, wherein processing further comprises filtering the sound signal into frequency bands, detecting noise in at least one band, applying a frequency gain rule to each band, and reconstructing the bands.
[0159] Embodiment 37. The method of any of Embodiments 31-36, wherein the sound receiver is an in-ear microphone and the frequency gain rule is based on detecting noise above a threshold.
[0160] Embodiment 38. The method of any of Embodiments 31-37, further comprising detecting noise in each frequency band.
[0161] Embodiment 39. The method of any of Embodiments 31-38, wherein the frequency gain rule comprises increasing gain, reducing gain, or applying a limiter.
[0162] Embodiment 40. The method of any of Embodiments 31-39, further comprising applying the VAD separately to each frequency band.
[0163] Embodiment 41. The method of any of Embodiments 31-40, further comprising classifying a sound type in at least one frequency band based on the VAD.
[0164] Embodiment 42. The method of any of Embodiments 31-41, wherein the sound type is speech or an alarm.
[0165] Embodiment 43. The method of any of Embodiments 31-42, further comprising receiving a second sound signal, filtering it into corresponding bands, comparing a band of the first and second signals, determining that a microphone noise threshold of the first microphone is above a predetermined threshold, and throttling a contribution of the second microphone.
[0166] Embodiment 44. The method of any of Embodiments 31-43, wherein the sound receiver is an in-ear microphone, an out-of-ear microphone, or a voice pickup.
[0167] Embodiment 45. The method of any of Embodiments 31-44, wherein processing further comprises applying a level-dependent function.
[0168] Embodiment 46. An in-ear hearing protection device comprising an in-ear microphone, an external microphone, a speaker, and a sound processing unit configured to filter a received sound signal into multiple bands, apply a VAD, apply a frequency gain rule to at least one band, reconstruct the bands, and crossfade the processed sound signal.
[0169] Embodiment 47. The device of Embodiment 46, wherein the in-ear microphone or external microphone comprises a voice-pickup bone sensor.
[0170] Embodiment 48. The device of any of Embodiments 46-47. further comprising a delay module configured to delay broadcasting by a time delay, retrieve a previous sound signal upon speech detection, generate a modified processed sound signal comprising both signals, and provide it to the crossfader. Embodiment 49. The device of any of Embodiments 46-48, wherein the delay module provides the modified processed sound signal within 1 second of receiving the sound signal.
[0171] Embodiment 50. The device of any of Embodiments 46-49. wherein the crossfader is a square-factorcrossfader.
[0172] Embodiment 51. The device of any of Embodiments 46-50, wherein the received sound signal is from the in-ear microphone and the device includes a noise module configured to detect noise levels in one or more bands.
[0173] Embodiment 52. The device of any of Embodiments 46-51, wherein the frequency gain rule for a first band is selected based on the detected noise level in that band.
[0174] Embodiment 53. The device of any of Embodiments 46-52, wherein the sound signal comprises a mix of an in-ear sound signal and an external sound signal.
[0175] Embodiment 54. The device of any of Embodiments 46-53, wherein when noise exceeds a threshold, the crossfader reduces a contribution of the external sound signal.
[0176] Embodiment 55. The device of any of Embodiments 46-54, wherein the frequency gain modifier comprises a level-dependent function.
[0177] Embodiment 56. The device of any of Embodiments 46-55, wherein the frequency gain modifier applies different gain mles to different frequency bands.
[0178] Embodiment 57. The device of any of Embodiments 46-56, further comprising a flange configured to provide a seal between the in-ear microphone and ambient sound.
[0179] Embodiment 58. The device of any of Embodiments 46-57, wherein if no speech is detected by the VAD, no processed sound is broadcast.
Claims
CLAIMS:
1. A method for improving fidelity of speech in a hearing protection device, the method comprising:receiving a sound signal from a sound receiver;processing, by a sound filter, the sound signal to obtain a processed sound signal, wherein processing comprises:applying a voice activity detector (VAD) to the sound signal;filtering the sound signal into a plurality of frequency bands;detecting a noise level in at least one frequency band of the plurality of frequency bands; applying a frequency gain rule to each of the plurality of frequency bands; and reconstructing the plurality of frequency bands to generate the processed sound signal; and broadcasting the processed sound signal.
2. The method of claim 1, wherein broadcasting the processed sound signal further comprises:delaying broadcasting the processed sound signal by a time delay;detecting, based on the processed sound signal, that the sound signal comprises a speech; retrieving a previous sound signal captured prior to the received sound signal based on the detection that the sound signal comprises the speech;generating a modified processed sound signal comprising the previous sound signal and the processed sound signal; andbroadcasting the modified processed sound signal.
3. The method of claim 2, wherein the time delay is less than 1 second.
4. The method of claim 2, wherein the previous sound signal is captured within 1 second of the received sound signal.
5. The method of any of claims 1 through 4, wherein the sound signal is a first sound signal and the sound receiver is a first microphone, and wherein the method further comprises:receiving a second sound signal from a second microphone;filtering the second sound signal into a plurality of frequency bands corresponding to the plurality of frequency bands of the first sound signal;comparing one frequency band of the plurality of frequency bands of the first sound signal and the second sound signal;determining that a microphone noise threshold of the first microphone is above a predetermined threshold based on the comparison of the one frequency band; andthrottling a contribution of the second microphone to the processed sound signal based on the determination that the microphone noise threshold of the first microphone is above the predetermined threshold.
6. The method of claims 1 though 5, further comprising applying the VAD separately to each of the plurality of frequency bands.
7. The method of claim 6, further comprising classifying, based on the application of the VAD, a sound type in the at least one frequency band of the plurality of frequency bands.
8. The method of any of claims 1 through 7, wherein processing the sound signal further comprises applying a level-dependent function to the sound signal.
9. The method of any of claims 1 through 8, wherein processing the sound signal further comprises applying a square-factor crossfader.
10. A method for improving fidelity of speech in a hearing protection device, the method comprising:receiving a sound signal from a sound receiver;processing the sound signal to obtain a processed sound signal, wherein processing comprises applying a voice activity detector (VAD) to the sound signal;delaying broadcasting the processed sound signal by a time delay;detecting, based on the processed sound signal, that the sound signal comprises a speech; retrieving a previous sound signal captured prior to the received sound signal based on the detection that the sound signal comprises the speech;generating a modified processed sound signal comprising the previous sound signal and the processed sound signal; andbroadcasting the modified processed sound signal.
11. The method of claim 10, wherein the time delay is less than 1 second.
12. The method of claim 10, wherein the previous sound signal was captured within 1 second of the received sound signal.
13. The method of any of claims 10 through 12, further comprising:determining that the sound signal has a noise level above a noise threshold; andthrottling a contribution of the sound receiver based on the determination that the sound signal has the noise level above the noise threshold.
14. The method of any of claims 10 through 13, wherein processing the sound signal further comprises:filtering the sound signal into a plurality of frequency bands;detecting a noise level in at least one frequency band of the plurality of frequency bands; applying a frequency gain rule to each of the plurality of frequency bands; and reconstructing the plurality of frequency bands into the processed sound signal;15. The method of claim any of claims 10 through 14, wherein the sound receiver is an in-ear microphone, and wherein the frequency gain rule is based on a detection that the noise level is above a noise threshold in one frequency band of the plurality of frequency bands.
16. A method for improving fidelity of speech in a hearing protection device, the method comprising:receiving a sound signal from a sound receiver;processing, by a sound filter, the sound signal to obtain a processed sound signal, wherein processing comprises:applying a voice activity detector (VAD) to the sound signal; andapplying a square-factor crossfader; andbroadcasting the processed sound signal.
17. The method of claim 16, further comprising:determining that the sound signal has a noise level above a noise threshold; and throttling a contribution of the sound receiver based on the determination that the sound signal has the noise level above the noise threshold.
18. The method of claim 16 or 17, wherein broadcasting the processed sound signal further comprises:delaying broadcasting the processed sound signal by a time delay;detecting, based on the processed sound signal, that the sound signal comprises a speech; retrieving a previous sound signal captured prior to the received sound signal based on the detection that the sound signal comprises the speech;generating a modified processed sound signal comprising the previous sound signal and the processed sound signal; andbroadcasting the modified processed sound signal.
19. The method of any of claims 16 through 18, wherein the time delay is less than 1 second.
20. The method of any of claims 16 through 18, wherein the previous sound signal was captured within 1 second of the received sound signal.