Speech enhancement system and method
The hearing protection device enhances audio fidelity in noisy environments by using an individualized speech enhancement model for in-ear microphones, addressing the issues of attenuated sound and discomfort in existing PPE systems.
Patent Information
- Application Number
- PCT/IB2025/055010
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-05-30
- Filing Date
- 2025-05-13
- Publication Date
- 2025-12-04
AI Technical Summary
In noisy environments, existing personal protective equipment (PPE) with in-ear microphones suffer from attenuated audio fidelity due to sound transmission through human tissue, leading to distorted and unnatural sounding utterances, and existing solutions like boom mics are cumbersome and uncomfortable.
A hearing protection device with an in-ear microphone and a microprocessor applies an individualized speech enhancement model to attenuated audio signals, enhancing fidelity before transmission to other users.
The system provides natural-sounding speech communication among workers by applying a user-specific speech enhancement model, improving audio fidelity and maintaining comfort and convenience.
Smart Images

Figure IB2025055010_04122025_PF_FP_ABST
Abstract
Description
SPEECH ENHANCEMENT SYSTEM AND METHODTECHNICAL FIELD
[0001] The present disclosure generally relates to the field of personal protection equipment, and more specifically, to personal protection equipment that provides hearing protection along with a communication component for sensing and wirelessly transmitting voice to other workers in an area.BACKGROUND
[0002] In many scenarios, multiple workers may be deployed to a given work environment. The workers often need to communicate with one another via spoken communications while deployed to the work environment. For safety reasons, the workers may wear personal protective equipment, including but not limited to hearing protection.SUMMARY
[0003] Verbal communication between workers in noisy environments, such as industrial settings, raises myriad issues. For example, workers typically need to be wearing ear protection in such environments, which means a solution typically needs to involve embedded speakers and microphones. Microphones can be deployed, but unless they are positioned at the end of a boom mic, they may not pick up a clear enough signal to be able to cancel out ambient noise. Boom mics are also cumbersome and not well tolerated in many work environments. However, when they are used, an utterance may be received, and ambient noise (sensed from a different microphone not facing the wearer) may be subtracted from the utterance, to provide a relatively clean and intelligible utterance to other workers in an area, usually by radio.
[0004] In-ear hearing protection and communications devices to some degree address the issue of the unsuitability of boom microphones in some environments. Instead of the microphone being positioned in front of a wearer’s mouth, an “in-ear” microphone is embedded in an earpiece which is positioned in a user’s ear. The earpiece occludes the ear canal, thereby providing hearing protection to the user and isolating the microphone from outside noise. The microphone picks up utterances made by a wearer and transmitted to the in-ear microphone through human tissue (rather than through the air as typical sound would be conducted). The received utterance, however, loses fidelity and is attenuated during its transmission through the human tissue. What arrives at the in-ear microphone is bandwidth-limited due to the transmission path, and generally sounds muted and lacking. Generic speech enhancement algorithms and / or artificial bandwidth extension algorithms may be used to try to add fidelity to the muted utterance, which is then transmitted to communications devices of other users in the area.
[0005] In one example, a hearing protection device includes a sound muffling portion and a processor. The sound muffling portion includes a sound dampening material configured to fit at least partially within a user’s ear canal when donned, to define a sound attenuation area that includes the user’s eardrum. Thesound muffling portion further includes a microphone disposed to sense audio signals within the sound attenuation area. The processor is communicatively coupled to the microphone. The processor is configured to receive, upon the user generating a first utterance, attenuated audio signals associated with the first utterance, and to apply an individualized speech enhancement model to the audio signals to yield transformed audio signals associated with the first utterance. The individualized speech enhancement model comprises a model that is pre-trained model specifically for the user.
[0006] In another example, a method includes defining a sound attenuation area that includes an eardrum of a user, and sensing audio signals within the sound attenuation area. The method further includes receiving attenuated audio signals associated with the first utterance, and applying an individualized speech enhancement model to the audio signals to yield transformed audio signals associated with the first utterance. The individualized speech enhancement model comprises a model that is pre-trained model specifically for the user.
[0007] In another example, a hearing protection apparatus includes means for defining a sound attenuation area that includes an eardrum of a user, means for sensing audio signals within the sound attenuation area; means for receiving attenuated audio signals associated with the first utterance, and means for applying an individualized speech enhancement model to the audio signals to yield transformed audio signals associated with the first utterance, where the individualized speech enhancement model comprises a model that is pre-trained model specifically for the user.
[0008] In another example, a non-transitory computer-readable storage medium is encoded with instructions that, when executed, cause processing circuitry of a hearing protection device to perform operations. The operations include defining a sound attenuation area that includes an eardrum of a user, and sensing audio signals within the sound attenuation area. The operations further include receiving attenuated audio signals associated with the first utterance, and applying an individualized speech enhancement model to the audio signals to yield transformed audio signals associated with the first utterance. The individualized speech enhancement model comprises a model that is pre-trained model specifically for the user.
[0009] The details of one or more examples of the disclosure are set forth in the accompanying drawings and the description below. Other features, objects, and advantages of the disclosure will be apparent from the description and drawings, and from the claims.BRIEF DESCRIPTION OF THE DRAWINGS
[0010] FIG. 1 is a drawing of a plurality of workers in a noisy work environment.
[0011] FIG. 2A is a top perspective view of an earpiece and wire according to an embodiment.
[0012] FIG. 2B is a bottom view of the earpiece of FIG. 2A.
[0013] FIG. 3 is a top perspective view of an earpiece with an eartip according to an embodiment.
[0014] FIG. 4A is a side view of the earpiece of FIG. 3 according to an embodiment.
[0015] FIG. 4B is a cross-sectional view of the earpiece of FIG. 3 without an eartip according to anembodiment.
[0016] FIG. 5 is an exploded view of the earpiece and cable of FIG. 2A according to an embodiment.
[0017] FIG. 6 A is a perspective view of a lower part of the earpiece shell and internal components of the earpiece of FIG. 3 according to an embodiment.
[0018] FIG. 6B is a cross-sectional view of the lower part of FIG. 6A according to an embodiment.
[0019] FIG. 7 is a perspective view of an insert of the earpiece of FIG. 3 according to an embodiment.
[0020] FIG. 8 A is a top view of the insert of FIG. 7 according to an embodiment.
[0021] FIG. 8B is a bottom view of the insert of FIG. 7 according to an embodiment.
[0022] FIGS. 8C and 8D are side and front views, respectively, of the insert of FIG. 7 according to an embodiment.
[0023] FIG. 8E is a top perspective view of the insert of FIG. 7 according to an embodiment.
[0024] FIG. 8F is a bottom perspective view of the insert of FIG. 7 according to an embodiment.
[0025] FIG. 8G is a cross-sectional view of the insert of FIG. 7 according to an embodiment.
[0026] FIG. 8H is a cross-sectional view of the insert of FIG. 7 according to an embodiment.
[0027] FIG. 9 is a top view of the lower part of the earpiece shell of the earpiece of FIG. 3 according to an embodiment.
[0028] FIG. 10 A is a schematic view of the groove in the earpiece shell of the earpiece of FIG. 3 according to an embodiment.
[0029] FIG. 10B is a schematic, partial cross-sectional view of the groove in the earpiece shell of the earpiece of FIG. 3 according to an embodiment.
[0030] FIGS. 11 A and 11B are schematic cross-sectional views of an earpiece showing an alternative acoustic channel according to an embodiment.
[0031] FIGS. 12A, 12B, and 12C are schematic cross-sectional views of an earpiece showing alternative acoustic channels according to an embodiment.
[0032] FIG. 13 is a flowchart showing a process by which an individualized speech enhancement model is generated.
[0033] FIG. 14 is a high-level system diagram showing certain functional and logical components of a communications device as described herein.
[0034] FIG. 15 is a flow chart illustrating an exemplary process that a communications device of this disclosure and its processors would use to apply an individualized speech enhancement model to a digitized utterance received from a wearer.
[0035] FIG. 16 is a flowchart showing how an individualized speech enhancement model may be applied to utterances, real-time, as part of a communications system involving a plurality of users, each wearing a compatible communications device.
[0036] It is to be understood that the embodiments may be utilized, and structural changes may be made without departing from the scope of the invention. The figures are not necessarily to scale. Like numbers used in the figures refer to like components. However, it will be understood that the use of a number torefer to a component in a given figure is not intended to limit the component in another figure labeled with the same number.DETAILED DESCRIPTION
[0037] In the following detailed description, reference is made to the accompanying drawings, which form a part hereof, and in which is shown by way of illustration embodiments in which the inventions may be practiced. These embodiments are described in sufficient detail to enable those skilled in the art to practice the invention, and it is to be understood that other embodiments may be utilized and mechanical changes may be made without departing from the spirit and scope of the present invention. The following detailed description is, therefore, not to be taken in a limiting sense, and the scope of the present invention is defined by the claims and equivalents thereof.
[0038] Many types of personal protective equipment (PPE) include a speaker, a microphone, or both. The speakers may provide a received audio transmission to a user, while the microphones may capture audio from the wearer. Different forms of PPE have different quality microphones and speakers. Additionally, wearing different PPE interferes with the ability of microphones to pick up audio, and for speakers to transmit audio.
[0039] Some hearing protection devices include active hearing protection as well as passive hearing protection. Active hearing protection includes one or more microphones that receive ambient sound from a user’s surroundings, a processor to evaluate the sound and limit it to a safe level, and one or more speakers to play it back, at the safe level, to a user. Active hearing protection devices use electronic circuitry to pick up ambient sound through a microphone and convert it to safe levels before playing it back to the user through a speaker. Additionally, active hearing protection may comprise filtering out sound above a given sound pressure level, for example actively reducing the sound of a gunshot while providing human speech at substantially unchanged levels. At least some embodiments herein are applicable to active hearing devices.
[0040] In an active hearing protection system, a sound signal is first received by a microphone. The received sound signal is converted to an electronic signal for processing. After processing the sound signal such that all frequencies are at safe levels for a user, the sound signal is reproduced and played back to a user through a speaker of the hearing protection device.
[0041] Some active hearing protection units are level dependent, such that an electronic circuit adapts the sound pressure level. Level dependent hearing protection units help to filter out impulse noises, such as gunshots from surrounding noises, and / or continuously adapt all ambient sound received to an appropriate level before it is reproduced to a user. Active hearing protection units, specifically level-dependent active hearing protection units, may be necessary to facilitate communication in noisy environments, or environments where noise levels can vary significantly, or where high impulse sounds may cause hearing damage. A user may need to hear nearby ambient sounds, such as machine sounds or speech, while alsobeing protected from harmful noise levels. Additionally, a user may need to communicate with another user in such an environment. For instance, a user may wish to communicate through the ambient listening microphones. However, speech enhancement may not be necessary for this scenario, as the microphone is receiving full-fidelity audio. It will be appreciated that speech enhancement differs from noise reduction, the latter of which may aid in communication precision in this scenario.
[0042] Communication between users in noisy environments (such as those that necessitate donning hearing protection type PPE) presents a number of challenges. The environment necessitates that users wear over-ear hearing protection (e.g., earmuffs) and / or in-ear hearing protection (e.g., earplugs). In some noisy environments, such as helicopters, earmuffs are fitted with boom mics that pick up a user’s utterances (as well as high levels of ambient noise), and then filter out the ambient background noise before presenting the now fdtered utterance to another user in the cockpit. Often, the boom mics include two microphones, one directed toward a wearer’s mouth to pick up utterances, and another directed outwardly from the user’s mouth, to pick up ambient noise. Earmuffs fitted with boom mics, while effective at picking up utterances, are bulky and not well tolerated in, for example, industrial or manufacturing environments.
[0043] While boom mics perform well in noisy environments, boom mics may diminish the user experience. Boom mics are generally associated with circumaural hearing protection (e.g., as provided by over-the-ear headsets), which worker feedback often describes as being hot, uncomfortable, bulky, and having other qualities that diminish the user experience, especially when worn for extended periods of time such as a ten-hour or twelve-hour shift as is often the case in industrial environments. Moreover, due to compliance with regulatory requirements in various jurisdictions, many industrial workers are accustomed to wearing earplugs during their shift. For these workers, in-ear communication is a relatively smooth transition because such workers tend to be familiar with and proficient at properly fitting an earplug and are accustomed to having in-ear equipment in place all shift long. Earplug wearers may be less accustomed to the user experience associated with using an over-the-ear communications headset. In such environments, in-ear hearing protection, which is relatively lightweight and non-bulky, is preferred by most users. Additionally, in such environments, it is often impractical to use a boom-type microphone to pick up utterances from a worker. Instead, worker utterances are sensed, in some devices, using an “in-ear” microphone which is shielded from ambient sounds, as it picks up sound from within the user’s ear canal and is insulated from external noise in the same or similar way the wearer’s ear drum would be.
[0044] An in-ear microphone thus picks up utterances emanating from a wearer, transmitted through body tissue directly into the ear canal. These utterances, however, lose fidelity due to attenuation by passage through body tissue. The in-ear microphone is coupled to a microprocessor, which is typically housed on a small circuit board and worn in or near the ear, in an earpiece or “earbud” type form factor. The microprocessor may receive a digital rendering of the attenuated utterance, apply generic speech enhancement algorithms to try to restore some of the lost fidelity, then may use a radio to transmit theutterances to other users in the area, which wear a similar device and which receives the radio signals, converts them into analog sounds which are presented to a wearer. The algorithms may be configured to attempt to enhance the signals to restore some of the lost fidelity, but they aren’t perfect and lead to distorted and unnatural sounding utterances. In such manner, it is possible for a plurality of workers in a noisy industrial setting to communicate with one another, while still having protection against loud and potentially damaging sound levels. One such system that provides this ability is manufactured by 3M Company of St. Paul, MN under the trade name “PELTOR™ Professional In-Ear Communication Headset, PIC-100.”
[0045] It has been discovered that more natural-sounding speech may be generated from captured attenuated sound by applying, via a microprocessor, an individualized speech enhancement model to the attenuated audio signal containing the utterance. The individualized, or user-specific, model would have been created previously, and is, in one embodiment, deployed into the memory of the communications device, to be applied by the processor therein. By individualized, in one embodiment, it is meant that the model is specific to a wearer of the communication device (the wearer being also the speaker, and also sometimes referred to as the user herein).
[0046] FIG. 1 is a drawing of an industrial setting 10, which includes loud commercial or industrial equipment (not shown in FIG. 1) that interferes with verbal communication between inhabitants. For example, industrial setting 10 may comprise a manufacturing or construction environment. A plurality of workers, shown as worker 18A, 18B, and 18C each don a communication device 12, which is communicatively connected with earpieces 14 A and 14B. The earpieces will be described further with respect to subsequent figures, but in general are hearing protection devices that include one or more speakers and one or more microphones on the protected side of the hearing protection device (otherwise referred to as in-ear, though this term is to be read broadly to include configurations where the microphone is not technically inside of the ear canal, but are on the wearer’s side of the hearing protection apparatus). As such, it will be appreciated that microphones referred to herein as “in-ear” microphones need not necessarily be physically positioned inside the ear canal while in use. In some such examples, an in-ear microphone may have an unobstructed acoustic path from the inner ear to the sound capturing hardware of the microphone. This type of in-ear microphone is acoustically isolated from outside noise either totally or as much as possible such that it is considered to be behind the hearing protection the worker is wearing, thereby being effectively in-ear.
[0047] Communications device 12 is shown to be tethered to earpieces 14A and 14B by a wire, but in various embodiments the earpieces may be self-contained and communicate with each other via radio (e.g., in a wireless manner). In some wireless communication scenarios, a first step may be to establish a wireless link between the earpieces and the communications device. Eventually in the process, it is desirable to have one earpiece include the communication device and also communicate with the second earpiece, although currently available systems do not provide this functionality. Worker 18A may make an utterance, which is picked up by the in-ear microphone of the communications device worker 18A iswearing. In some examples in accordance with the aspects of this disclosure the sound is captured in digital format using one or more digital microphones and / or digital microphone arrays, thereby eliminating the resource consumption associated with transmitting low-voltage analog signals along the entire length of a wire or the cumulative length of a series of wires. In other examples, the captured sound can be converted to digital by an analog-to-digital chip in communications device 12 to produce a digitized utterance. A processor also within communications device 12 may then apply an individualized speech enhancement model to the digitized utterance, which yields an enhanced digitized utterance. The process of creating the enhancement model and ultimately the enhanced digitized utterance will be the subject of later discussion herein. The enhanced digitized utterance is broadcast, wirelessly, to communication devices associated with and worn by other workers 18B and 18C, which receive the enhanced digitized utterance, and use a processor and digital to analog converter (DAC) to output an analog version of the enhanced digitized utterance to workers 18B and 18C. Workers 18B and 18C can similarly make utterances which are picked up by their in-ear microphones, digitized, enhanced, and broadcast to the other respective workers, and thereby communicate with each other even with hearing protection devices being worn.
[0048] FIGS. 2A and 2B are close-ups of earpiece 14a and 14b shown in FIG. 1. Earpiece 20 includes an outer shell 100 that houses the interior components of the earpiece 20. One difference between wired and wireless earpieces is that the electrical components of wired earpieces are connected to a wire or cable that may connect the earpiece components to a device, such as a communications or audio device, whereas a wireless earpiece connects wirelessly to the device and may include a rechargeable battery. The exemplary wired earpiece 20 includes a cable extension 112 protruding from the shell 100. Wire 160 may extend through the cable extension 112. Wire 160 may contain a cable 162 and an electrical wire 161 (shown in FIG. 5), connected at one end to one or more of the inner components of the earpiece 20. The wire 160 may form an earhook. The earpiece 20 is shown without the wire 160 in most subsequent drawings.
[0049] The earpiece 20 includes a speaker port 124, shown in FIG. 2B. The speaker port 124 is connectable with an eartip 126 as shown in FIG. 3. The eartip 126 may be inserted into the ear of a user. The eartip 126 may be made from an elastomeric material that allows for the formation of an acoustic seal between the earpiece 1 and the ear canal. The eartip 126 may be removable and replaceable.
[0050] As shown in FIGS. 4A and 4B and the exploded view in FIG. 5, the shell 100 of the earpiece 1 may include two parts: a first (or lower) part 101 and a second (or upper) part 102. The first and second parts 101, 102 may be coupled together, forming the shell 100 and defining an interior 110. The interior110 may house the various internal components of the earpiece 1. The first part 101 may further include the speaker port 124. The first part 101 and the speaker port 124 may define a sound channel 140 extending through the speaker port.
[0051] The interior 110 may be divided into a first cavity 111 and a second cavity 112. The first cavity111 may be mainly or completely housed in the first part 101. The second cavity 112 may be mainly orcompletely housed in the second part 102. The interior 110 may be divided into the two cavities by a circuit board, such as a printed circuit board 314, housed in the interior 110.
[0052] Referring now to FIGS. 4B, 5, and FIGS. 6A and 6B (showing the earpiece 1 with the upper part 102 removed), the shell 100 (e.g., the first part 101) forms a cavity 111 (e.g., the first cavity 111) and a sound channel 140. The sound channel 140 extends from the cavity 111 from a first end 141 of the sound channel 140 to an opposing second end 142. The second end 142 is an open end. According to an embodiment, the sound channel 140 forms a single pass-through cavity extending from the first end 141 to the second end 142. That is, the sound channel 140 is undivided and is not divided into co -extending channels by a wall or other structure.
[0053] An insert 200 is disposed within the cavity 111. The insert 200 may be a molded element. In some embodiments, the insert 200 may define a single integral mass of elastomeric material. According to an embodiment, the insert 200 is a single integral molded element. For example, the insert 200 may be injection molded as a single integral piece. Alternatively, the insert 200 may be formed from two or more molded pieces.
[0054] The insert 200 may be injection molded from elastomeric material as a single integral piece. The insert 200 may be injection molded from elastomeric material as two or more pieces. In some embodiments, the insert 200 consists of elastomeric material. For example, the insert 200 may be free of adhesives.
[0055] In some embodiments the insert 200 is made from another (non-elastomeric) material. For example, the insert 200 may be injection molded from a polymeric material. The insert 200 may be injection molded from polymeric material as a single integral piece or as two or more pieces.
[0056] The earpiece 1 further includes a speaker 320. In some embodiments, the speaker 320 is partially embedded in the insert 200, as shown in FIGS. 4B and 6B. In some alternative embodiments, the speaker 320 is not embedded in the insert 200. For example, the speaker 320 may be disposed on the outside of the insert 200. The speaker 320 may be attached to the insert 200, for example, by a glue. The speaker 320 is constructed to project sound through the sound channel 140 and eartip 126 and into the user’s ear canal. The speaker 320 may be positioned such that the sound-projecting end (e.g., first end 321) is oriented toward the sound channel 140. The first end 321 may extend into the sound channel 140.
[0057] A circuit board assembly 310 is mounted onto the insert 200. The circuit board assembly 310 may be seated on and seal against a rim 213 or ledge formed by the wall 212 of the insert 200. The circuit board assembly 310 includes a printed circuit board 314. The printed circuit board 314 defines a first major side 311 and a second major side 312 opposite of the first major side 311. The first major side 311 faces the first cavity 111 and the first part 101 of the shell 100. The second major side 312 faces the second cavity 312 and the second part 102 of the shell 100. A first microphone 330 is disposed on the first major side 311 of the printed circuit board 314. The first microphone 330 may be an in-ear microphone. An in-ear microphone may be used to read a sound pressure level in the ear canal of a user (e.g., at the junction of the earpiece and the ear canal). A second microphone 340 may be disposed on thesecond major side 312 of the printed circuit board 314. Second microphone 340 is referred to herein as an “ambient listening microphone” because it is exposed to outside sounds / noise and enables the microprocessor to digitize the external sounds and mix in a desired amount to improve situational awareness. The first and second microphones 330, 340 may be independently selected from any suitable microphones, such as MEMS (Micro-Electro-Mechanical System) microphones.
[0058] The circuit board assembly 310 may also include a controller 315. The controller 315 may include one or more processors such as, e.g., one or more microprocessors, digital signal processors (DSPs), application-specific integrated circuit (ASICs), field programmable gate arrays (FPGAs), complex programmable logic device (CPLDs), microcontrollers, digital-to-analog converters (DACs), analog-to-digital converters (ADCs), or any other equivalent integrated or discrete logic circuitry. The controller 315 may be operatively coupled to transducers of the earpiece 1 such as, for example, the speaker 320, the first microphone 330, and the second microphone 340. The controller 315 may be operatively coupled to such transducers via a wired or wireless connection that allows the controller 315 to receive audio signals from the microphones 330, 340, or transmit audio signals to the speaker 320. Audio signals received from the microphones 330, 340 may converted to digital signals for digital signal processing. Audio signals transmitted to the speaker 320 may be digital or analog depending on the type of speaker 320. The audio signals transmitted to the speaker 320 may include components of the audio signals received from one or both of the microphones 330, 340.
[0059] An acoustic chamber 130 is formed between the insert 200 and the circuit board assembly 310. The first microphone 330 is disposed within the acoustic chamber 130. An acoustic channel 150 extends from the acoustic chamber 130 to the sound channel 140. The acoustic channel 150 may extend from the acoustic chamber to the first end 141 of the sound channel 140.
[0060] Referring now to FIGS. 6B, 7, and 8A-8H, the insert 200 has a bottom 210 surrounded by a wall 212. The outer surface 211 of the bottom 210 of the insert 200 may be disposed against the inner surface 104 of the shell 100. The acoustic channel 150 may be formed between the outer surface 211 of the bottom 210 and the inner surface 104 of the shell 100. One or both of the bottom 210 and the inner surface 104 may include an indentation or groove to facilitate formation of the acoustic channel 150. In the embodiment shown, the groove 105 is formed on the inner surface 104 of the shell 100. The acoustic channel 150 may be formed by the groove 105 and the bottom 210 of the insert 200. The groove 105, and thus the acoustic channel 150, may extend from a first end 151 below the first microphone 330 to a second end 152 at the sound channel 140 (see FIG. 9). The insert 200 has an aperture 250 that connects the acoustic chamber 130 to the acoustic channel 150. The shell 100 may include a protrusion 106 extending from the inner surface 104 at the first end 151 of the acoustic channel 150. The protrusion 106 may protrude through the aperture 250 in the insert 200. The protrusion 106 may form a first end of the groove 105.
[0061] The insert 200 is constructed to acoustically isolate the first microphone 330 from the speaker 320. The insert 200 may have a through opening 220 in which the speaker 320 is received. The insert200 includes a partition wall 214. The partition wall 214 may form part of the wall surrounding the through opening 220. The partition wall 214 separates the through opening 220, in which the speaker 320 is disposed, from the acoustic chamber 130, in which the first microphone 330 is disposed. According to an embodiment, the first microphone 330 is in communication with the speaker 320 only via the acoustic channel 150 and the aperture 250.
[0062] The acoustic channel 150 is configured and constructed to reduce or dampen sounds emanating from the speaker at certain frequencies. By adjusting the length and the transverse cross-sectional area of the channel, the dampened frequencies and the amount of dampening can be tailored to the specific needs of the earpiece. For example, the acoustic channel 150 can be designed as a “low pass” channel that cuts off high frequency sounds and lets low frequency sounds pass. The acoustic channel 150 may be formed as a groove 105 in the first part 101 of the shell 100, as shown in FIG. 10. Alternatively, the acoustic channel 1150 may be formed as a groove or channel in the insert 1200, as shown in FIGS. 11A-11B, or using a pre-formed tube 2150, as shown in FIGS. 12A, 12B, and 12C. For example, as schematically shown in FIGS. 11 A-11B, the acoustic channel 1150 may be formed as a micro channel in the insert 1200. The acoustic channel 1150 (micro channel) may be molded into the insert 1200. The acoustic channel 1150 has a first end 1151 at the chamber 1130 and a second end 1152 near the sound channel 1140. The sound channel 1140 may be defined by the speaker port 1124. The speaker port 1124 may be included in a shell 1100. In another example, shown schematically in FIGS. 12A, 12B, and 12C, the acoustic channel 2150 is constructed by including a preformed tube 2155. The preformed tube 2155 may be disposed inside the insert 2200. The acoustic channel 2150 extends from a first end 2151 at the acoustic chamber 2130 to a second end 2152 near the sound channel 2140. The sound channel 2140 may be defined by the speaker port 2124. The speaker port 2124 may be included in a shell 2100. In another embodiment, the preformed tube 2155 is disposed on the outside of the insert 3200, extending from a first end 2151 at the acoustic chamber 3130 to a second end 2152 near the sound channel 2140. Alternatively, the acoustic channel 150 may be formed by corresponding grooves in both the shell 100 and the insert 200.
[0063] Reducing or dampening sounds emanating from the speaker 320 at certain frequencies using the acoustic channel 150 may allow the microphones 330, 340 to be utilized even when the speaker 320 is producing or generating sound. Accordingly, the controller 315 can receive audio signals from one or both of the microphones 330 (an example of an in-ear microphone) when the speaker 320 generates sound. In other words, the microphones 330, 340 may not be muted when the speaker 320 generates sound in earpieces that include one or more acoustic channels such as, for example, acoustic channel 150. The acoustic channel 150, in cooperation with the acoustic chamber 130, can prevent acoustical overload of the microphones that otherwise could cause clipping and distortion of the acoustic signals provided by the microphones. In contrast, earpieces that do not include means to dampen sounds emanating from a speaker at certain frequencies typically are configured to mute inner ear microphones to prevent such acoustical overload.
[0064] According to an embodiment, the acoustic chamber 130 has a volume VI 30 and the acousticchannel 150 has a length L150 and a transverse cross-sectional area A150, as shown in FIGS. 10A and 10B. The volume V130, length L150, and cross-sectional area A150 may be adjusted such that the acoustic channel reduces sounds at a desired cut-off frequency. The cut-off frequency may be 20 Hz or greater, 50 Hz or greater, 100 Hz or greater, 200 Hz or greater, 400 Hz or greater, 500 Hz or greater, 600 Hz or greater, 800 Hz or greater, 1000 Hz or greater, 2000 Hz or greater, 5000 Hz or greater, 10000 Hz or greater, 15000 Hz or greater, or 20000 Hz or greater. The cut-off frequency may be 10000 Hz or lower, 5000 Hz or lower, 3000 Hz or lower, 2500 Hz or lower, 2000 Hz or lower, 1500 Hz or lower, or 1200 Hz or lower. The cut-off frequency may be chosen in a range of 20 Hz to 20000 Hz, 200 Hz to 10000 Hz, 500 Hz to 5000 Hz, or 800 Hz to 2000 Hz. Non-limiting examples of suitable cut-off frequencies are 800 Hz, 1000 Hz, 1200 Hz, and 1500 Hz. In some cases, it is desired to adjust the volume V130, length L150, and cross-sectional area A150 to achieve a cut-off frequency in a range of 500 Hz to 2000 Hz, 800 Hz to 1200 Hz, or about 1000 Hz. The length L150 and cross-sectional area A150 may be adjusted such that the acoustic channel reduces sounds at frequencies of 20 Hz or greater, 50 Hz or greater, 100 Hz or greater, 200 Hz or greater, 400 Hz or greater, 500 Hz or greater, 600 Hz or greater, 800 Hz or greater, 1000 Hz or greater, 2000 Hz or greater, 5000 Hz or greater, 10000 Hz or greater, 15000 Hz or greater, or 20000 Hz or greater. The length L150 and cross-sectional area A150 may be adjusted such that the acoustic channel reduces sounds at frequencies of 40000 Hz or below, 20000 Hz or below, or 10000 Hz or below. The length L150 and cross-sectional area A150 may be adjusted such that the acoustic channel reduces sounds at frequencies between 600 Hz and 10000 Hz, between 800 Hz and 10000 Hz, or between 1000 Hz and 10000 Hz. The volume V130, length L150, and cross-sectional area A150 may be adjusted such that the system (made up of the acoustic chamber 130 and acoustic channel 150) acts as a second order filter. That is, the system may reduce sounds above the cut-off frequency from the speaker to the first microphone by about 12 dB per octave.
[0065] By adding a second (or further) acoustic chamber and a second (or further) acoustic channel in series with the first acoustic chamber 130 and acoustic channel 150, the sound reduction may be increased by multiples of 12 dB per octave. In one embodiment, the earpiece includes a first acoustic chamber and a first acoustic channel, and a second acoustic chamber and a second acoustic channel, the first acoustic channel extending from the first acoustic chamber to the second acoustic chamber, and the second acoustic channel extending from the second acoustic chamber to the sound channel. The first microphone (the in-ear microphone) may be disposed within the first acoustic chamber. In this embodiment, the two sets of acoustic chambers and acoustic channels reduce the sound from the speaker to the microphone by 24 dB per octave.
[0066] When tuning the cut-off frequency of the acoustic channel 150, the length LI 50 and cross- sectional area A150 are considered in correlation with one another. The resonance frequency of the system made up of the acoustic chamber 130 and the acoustic channel 150 can be described by the following equation:
[0067] f_i=v / 27H(A / Vl) ,where fr is the resonance frequency, v is the speed of sound, A is the cross-sectional area (A150) of the acoustic channel 150, V is the volume (V130) of the acoustic chamber 130, and 1 is the length (L150) of the acoustic channel 150.
[0068] For practical reasons, the minimum and maximum size of the acoustic chamber 130 and the acoustic channel 150 may be limited. For example, the acoustic chamber 130 may have a volume V130 of 1 mm3 or greater, 5 mm3 or greater, 10 mm3 or greater, 20 mm3 or greater, 30 mm3 or greater, 40 mm3 or greater, 50 mm3 or greater, or 60 mm3 or greater. The acoustic chamber 130 may have a volume V130 of 400 mm3 or less, 300 mm3 or less, 200 mm3 or less, 150 mm3 or less, 125 mm3 or less, or 100 mm3 or less. The acoustic chamber 130 may have a volume V130 ranging from 1 mm3 to 400 mm3, from 30 mm3 to 150 mm3, from 40 mm3 to 125 mm3, or from 70 mm3 to 85 mm3. The acoustic channel 150 may have a length L150 is 1 mm or greater, 3 mm or greater, 5 mm or greater, 6 mm or greater, 7 mm or greater, 8 mm or greater or 9 mm or greater. The length L 150 may be 50 mm or less, 40 mm or less, 30 mm mor less, 20 mm or less, 15 mm or less, 14 mm or less, 13 mm or less, or 12 mm or less. In some embodiments, the length LI 50 ranges from 1 mm to 50 mm, 2 mm to 25 mm, 5 mm to 15 mm, or from 9 mm to 12 mm. The cross-sectional area A150 may be 0.03 mm2 or greater, 0.08 mm2 or greater, 0.1 mm2 or greater, 0.2 mm2 or greater, 0.3 mm2 or greater, 0.4 mm2 or greater, or 0.5 mm2 or greater. The cross-sectional area A150 may be 3.2 mm2 or less, 3.0 mm2 or less, 2.5 mm2 or less, 2.0 mm2 or less, 1.5 mm2 or less, 1.2 mm2 or less, 1.0 mm2 or less, or 0.8 mm2 or less. The cross-sectional area A150 may range from 0.03 mm2 to 3.2 mm2, 0.1 mm2 to 2.0 mm2, or 0.3 mm2 to 1.0 mm2. In one embodiment, the volume V130, length L150, and cross-sectional area A150 are adjusted such that above-discussed cutoff frequency is reached.
[0069] According to an embodiment, the insert provides several benefits for ease of manufacturing. For example, by using the insert 200, the internal components of the earpiece 1 can be mounted, secured in place, and sealed without the use of an adhesive. As discussed above, the insert 200 isolates the speaker 320 from the first microphone 330 and forms an acoustic network that includes an acoustic chamber 130 and an acoustic channel 150 that dampens certain sound frequencies from the speaker 320. The insert 200 also provides a supporting structure for the speaker 320, as well as a surface for mounting the circuit board assembly 310.
[0070] According to an embodiment, the insert 200 has a through opening 220 for receiving the speaker 320. The through opening 220 has a longitudinal center axis A220. The longitudinal center axis A220 may be disposed at an angle within the insert 200. That is, the longitudinal center axis A220 may be non- orthogonal relative to a plane defined by an outer rim 213 or ledge formed by the wall 212 of the insert 200, as shown in FIG. 8G. The longitudinal center axis A220 may be disposed at an angle a relative to the plane of the outer rim 213. When the speaker 320 is received within the through opening 220, the speaker 320 is disposed at an angle relative to the printed circuit board 314, which is mounted on the rim 213. The speaker 320 has a first end 321 and an opposing second end 322, and a longitudinal axis A320 extending from the first end 321 to the second end 322. When the speaker 320 is received within thethrough opening 220, the longitudinal axis A320 of the speaker 320 aligns with the longitudinal center axis A220 of the through opening 220. The through opening 220 may be sized so that the speaker 320 fits snugly within the through opening 220 (that is, the speaker 320 is supported on all sides by the material of the insert 200). As the speaker 320 is supported by the insert 200, the speaker 320 may be directly coupled with the printed circuit board 314. For example, the second end 322 of the speaker 320 may be soldered directly onto the printed circuit board 314 without the use of wires. Alternatively, the speaker 320 may be connected electrically to the printed circuit board 314 by using a flex part or one or more wires. The first end 321 of the speaker 320 may extend into the sound channel 140.
[0071] As noted, the internal components of the earpiece 1 can be mounted, secured in place, and sealed without the use of an adhesive. The lack of adhesive eliminates manufacturing complexities and potential messes that may result from the use of adhesives. Further, the components may be dismantled, if necessary, for repairs or adjustments. According to an embodiment, the earpiece 1 is free of an adhesive between the shell 100 and the circuit board assembly 310. In particular, the earpiece may be free of an adhesive between the first part 101 of the shell 100 and the circuit board assembly 310. The two parts (first part 101 and second part 102) of the shell 100 may be adhered together by an adhesive, snap fit, friction fit, or a fastener (such as a screw, clip, or the like).
[0072] The insert 200 has a bottom 210 outer surface 211 that is disposed against the inner surface 104 of the shell 100. The outer surface 211 may be shaped so that it is in contact with the inner surface 104 of the shell 100 along the entire bottom 210 with the exception of the groove 105. The outer surface 211 may have a continuously convex surface in a transverse cross section as shown in FIG. 8H.
[0073] The circuit board assembly 310 may be attached to the shell 100 via one or more fasteners 316 extending through the insert 200. Any suitable fastener may be used, such as a screw, clip, pin, bayonetstyle fastener, or the like. In one embodiment, the fasteners are screws. In one embodiment, the circuit board assembly 310 is attached to the first part 101 of the shell 100 by two screws. The shell 100 may include one or more supports 107 constructed to receive one or more fasteners 316 (e.g., screws). The insert 200 may include a corresponding protrusion 216 extending inwardly (into the acoustic chamber 130) from the wall 212 that mates with and / or covers the support 107. The protrusion 216 may have a through hole 217 for the fastener 316 (e.g., screw) to extend through. The printed circuit board 314 fastened to the shell 100 via the fasteners 316 may apply a permanent compressive force on the insert 200. The compressive force may improve the acoustical sealing provided by the insert 200.
[0074] The insert 200 may be constructed from any suitable material, such as an elastomer. In some embodiments, the insert 200 is made of an elastomeric material having a Shore A hardness of 20 or greater, 30 or greater, 40 or greater, 50 or greater, 60 or greater, or 65 or greater. The insert 200 may be made of an elastomeric material having a Shore A hardness of 90 or less, 85 or less, 80 or less, or 75 or less. The insert 200 may be made of an elastomeric material having a Shore A hardness from 20 to 90, from 50 to 85, or from 65 to 75. In one embodiment, the insert 200 is made of an elastomeric material having a Shore A hardness of about 70. Examples of suitable elastomeric materials include, for example,silicones, thermoplastic elastomers, thermoplastic polyurethanes, and the like. Preferably the elastomeric material is selected to provide sufficient support for the circuit board assembly 310, be able withstand compression over time, and be sufficiently soft to seal against the circuit board assembly 310 and around the speaker 320. In one embodiment, the insert 200 is made of a silicone having a Shore A hardness from 65 to 75 or about 70. The insert 200 may be a single integral piece molded from the elastomeric material. For example, the insert 200 may be a single integral piece molded from silicone having a Shore A hardness from 65 to 75 or about 70.
[0075] The earpiece 1 may include additional parts as shown in FIG. 5, such as a microphone seal 342 constructed to seal around the second microphone 340, a wind screen 170 mounted on the outside of the second part 102 of the shell 100, and an anchoring screw 362 constructed to anchor the cable 162 to the earpiece 1.
[0076] Assembling an earpiece configured according to the present disclosure may include first inserting the insert into the shell 100 (e.g., the first part 101). The insert 200 may be placed into the shell 100 without the use of an adhesive. According to an embodiment, placing the insert 200 into the shell 100 forms the acoustic channel 150. The speaker 320 may then be placed into the opening 220. The circuit board assembly 310, which may include the first microphone 330 and the second microphone 340, may be placed onto the insert 200 and attached to the first part 101 of the shell 100 by one or more fasteners (e.g., screws). The speaker 320 may then be soldered to the printed circuit board 314. Preferably, the speaker 320 is soldered directly onto the printed circuit board 314 without intervening wires. The microphone seal 342 may be placed on the second microphone 340. The circuit board assembly 310 may then be soldered onto wires (in the case that the earpiece is a wired earpiece). The second part 102 of the shell 100 may then be placed onto the first part 101. The second part 102 may be adhered to the first part 101 by an adhesive.
[0077] The techniques described in this disclosure, including those attributed to the systems, or various constituent components, may be implemented, at least in part, in hardware, software, firmware, or any combination thereof. For example, various aspects of the techniques may be implemented by the controller 315, which may use one or more processors such as, e.g., one or more microprocessors, DSPs, ASICs, FPGAs, CPLDs, microcontrollers, or any other equivalent integrated or discrete logic circuitry, as well as any combinations of such components, sound processing devices, or other devices. The term “processing apparatus,” “processor,” or “processing circuitry” may generally refer to any of the foregoing logic circuitry, alone or in combination with other logic circuitry, or any other equivalent circuitry. Additionally, the use of the word “processor” may not be limited to the use of a single processor but is intended to connote that at least one processor may be used to perform the exemplary techniques and processes described herein.
[0078] Such hardware, software, and / or firmware may be implemented within the same device or within separate devices to support the various operations and functions described in this disclosure. In addition, any of the described components may be implemented together or separately as discrete but interoperablelogic devices. Depiction of different features, e.g., using block diagrams, etc., is intended to highlight different functional aspects and does not necessarily imply that such features must be realized by separate hardware or software components. Rather, functionality may be performed by separate hardware or software components, or integrated within common or separate hardware or software components.
[0079] When implemented in software, the functionality ascribed to the systems, devices and techniques described in this disclosure may be embodied as instructions on a computer-readable medium such as RAM, ROM, NVRAM, EEPROM, FLASH memory, magnetic data storage media, optical data storage media, or the like. The instructions may be executed by the controller 315 to support one or more aspects of the functionality described in this disclosure.
[0080] As mentioned earlier, an utterance from the wearer is picked up by the in-ear microphone 330. A typical utterance, as one might hear in casual person-to-person conversation, is relatively wide band, from the low-frequency fundamental of the user's voice, to the voice's natural harmonics to the higher frequency modulation products that influence intelligibility. However, as received by the in-ear microphone, the utterance is distorted and may sound muffled, thereby resulting in loss of fidelity of the utterance. This is due in large degree to the manner in which the sound travels through human tissue to get to the in-ear microphone. Existing products, such as the aforementioned PIC-100, certain methods are employed to try to restore some of the lost fidelity. These methods include artificial bandwidth extension - essentially, an algorithm modifies a digitized version of the utterance to try to enhance it by applying a non-linear algorithm that we license from EERS. The original method is explained in “In-ear Microphone Speech Quality Enhancement via Adaptive Filtering and Artificial Bandwidth Extension” by Bouserhal, Falk, and Voix the entire disclosure of which is incorporated herein by reference. The result is an enhanced digitized utterance which is closer to the original than having not applied the artificial bandwidth extension, but still different enough from the original wideband speech that better approaches are needed. Importantly, artificial bandwidth extension algorithms deployed are generic to the user population - i.e., they apply the same basic approach on utterances from all wearers. Because of this, they work acceptably well, but are incapable of fully “restoring” speech fidelity, due to the loss of restoration data individualized on a per-user basis.
[0081] In one embodiment, as described herein, instead of applying a generic artificial bandwidth extension algorithm to a digitized copy of the received utterance, an individualized speech enhancement model that is specific to the wearer is applied to the digitized utterance. This approach yields, in some embodiments, far superior enhanced digitized utterances than applied generic algorithms.
[0082] INDIVIDUALIZED SPEECH ENHANCEMENT MODEL CREATION
[0083] To create an individualized speech enhancement model, a training recording is made of a wearer reading a script. In some examples, when the script is recited, the user is wearing the in-ear microphone(s), with an external source of truth being used as well. Other examples may use the ambient listening microphones on the in-ear headset to do the same thing as long as the user in in a relatively quiet and non-reverberant room when reading the training script. The script is fairly short, and in someembodiments only takes about 15 seconds to read. The script is designed to elicit a broad range of fidelity as would typically be included in a wearer’s utterances. In an industrial environment where there will be multiple users or wearers of communications devices, ideally each wearer would create a training recording, but if a subset of wearers do not engage in the creating of an individualized speech enhancement model, the system would default back to simply apply the generic bandwidth extension algorithm mentioned above.
[0084] The training recording is then used to train a base speech enhancement model using a machine learning algorithm. The training (and subsequent refinement retraining) may be accomplished using one or more learning paradigms, including but not limited to, transfer learning. For instance, systems of this disclosure may use transfer learning to pretrain a model with a large dataset and then use that model as a starting point for training the user-specific model(s). In accordance with some aspects of this disclosure, systems may use feature extraction where the pretrained model is used to extract the features from the recorded script audio. Those extracted features are then used to train the user specific model. Training and / or retraining may be conducted in a cloud computing environment, but local computing environments would also work (though they may take longer).
[0085] Reports indicate that certain speech-generation models may extract all the necessary features to be able to duplicate a person’s voice using less than five seconds of a ground truth sampling of the person’s voice. Given the shortcomings in the state of the art of in-ear audio technology, models of this disclosure, instead of recreating a person’s voice (e.g., as is done in so-called “deepfake” audio generation), are trained to rebuild or reconstruct lost information in the audio stream from the bandwidthlimited in-ear data that is indeed available.
[0086] In one example, the training phase of a base speech enhancement model of this disclosure started by using several thousand recordings available from voice command training and applying a lowpass filter to approximate how these recordings would be heard if picked up from within the ear. Using the resulting data as a set of inputs (as simulated bandwidth limited audio) and using the original full-fidelity recordings as outputs, for each recording, mel-frequency cepstral coefficients (MFCCs) were calculated, and the speech was marked on the input audio to simulate the voice activity detection (VAD) input. These data points were used as the original training data set. In turn, in this example, TensorFlow was used to set up and train a hitherto untrained model to take the provided inputs and generate, as execution phase output, data resembling the known outputs. Upon reaching convergence, this trained model was used as the pretrained model in a transfer learning scenario. In turn, this pretrained model was used as a starting point in generating multiple individualized models (in accordance with a transfer learning paradigm) with some additional recordings of an individual user’s voice in each case to generate a final model individualized for each user.
[0087] The end result is one or more individualized speech enhancement model(s) trained according to a transfer learning paradigm using the base pretrained model. Each individualized speech enhancement model represents a collection of weights for various inputs and their lags for each layer such as a VADoutput, one or more MFCCs, and / or one or more additional statistics of the audio stream that are computed in real time or near-real time from the incoming audio stream. The model output in its execution phase is an updated set of MFCCs that are then used as augmentation to generate the enhanced audio. In some examples, a dedicated program may be mn as a process that applies the model to the pre- processed inputs and generate the audio from the model output. Such a model may be of any appropriate size, but in some embodiments is about 10 megabytes or less. This model is uploaded to the individual wearer’s communications unit. If multiple users are sharing communications units, an initial onboarding of the user to the particular unit would accommodate the uploading of a user profde that could include such an individualized speech enhancement model to the communications unit.
[0088] FIG. 13 is a flowchart showing a process 72 by which an individualized speech enhancement model is generated according to aspects of this disclosure. Process 72 is an example of a training and deployment process of this disclosure. It will be appreciated that, in some use case scenarios, process 72 may be run iteratively, e.g., in multiple passes. By executing process 72 in multiple passes, systems of this disclosure may refine the training of a base model to provide a more accurate launching point for transfer learning purposes, may train and aid in the deployment of multiple “child” models for individualized speech enhancement, and may provide other data precision-related technical improvements. As such, it will be appreciated that the description below pertains to a single pass of process 72, while the described process may be performed potentially multiple times in practice.
[0089] Process 72 may begin with obtaining speech samples (74). For instance, the speech samples may be recordings available from voice command training or from other voice data available to the training system. The systems of this disclosure may apply lowpass filtering to the speech samples to simulate in- ear pickup of the sound associated with the speech samples (76). By applying the lowpass filter, the systems of this disclosure may approximate how the speech samples would be heard if picked up from within the ear, e.g., through hard tissue. In turn, the systems of this disclosure may train a base speech enhancement model using the lowpass-filtered samples (as a set of inputs) and the original full-fidelity recordings (as a set of outputs) (78). For instance, for each recording, the systems of this disclosure may calculate MFCCs and mark the MFCCs on the input audio to simulate the VAD input.
[0090] The systems of this disclosure may determine whether the base speech enhancement model has achieved convergence (decision block 82). If the base speech enhancement model has not yet achieved convergence (NO branch of decision block 82), the systems of this disclosure may continue to train the base speech enhancement model using the inputs and outputs described above (e.g., by iterating operation 78). In this way, the systems of this disclosure may refine the performance of the base speech enhancement model by retraining the model until it achieves convergence. If the systems of this disclosure determine that the base speech enhancement model has achieved convergence (YES branch of decision block 82), then the systems of this disclosure may implement a transfer learning paradigm to train one or more individualized speech enhancement models for deployment (84). For instance, the systems of this disclosure may use the training of the base speech enhancement model as a basis, inaddition to user-specific speech samples, to train each individualized speech enhancement model on a user-by-user basis. In this way, the systems of this disclosure may apply the transfer learning paradigm to transfer speech enhancement training to individual voice samples to train and deploy speech enhancement models tailored for each individual user for in-ear microphone scenarios.
[0091] APPLICATION OF INDIVIDUALIZED MODEL TO UTTERANCE(S)
[0092] The individualized models of this disclosure can be applied to a digitized utterance by a general- purpose processors, or, to conserve clock cycles, can be applied by a dedicated Al accelerator, or alternatively, by a digital signal processor (DSP) to handle all mel-calculations separately. Thus, in some preferred embodiments, the individualized speech enhancement model applied by a chip that is called an artificial intelligence accelerator chip (or “Al accelerator”). An Al accelerator is a class of specialized hardware accelerator or computer system designed to accelerate artificial intelligence and machine learning applications, including artificial neural networks and machine vision. They are often manycore designs and generally focus on low-precision arithmetic, novel dataflow architectures or in-memory computing capability. The Al accelerator is included in the communications unit, ideally on the same circuit board of the general-purpose processor. Some Al accelerators are already included within system- on-a-chip (SoC) architectures, and integrated within other chipsets. One non-limiting example of such an Al accelerator is the Syntiant NDP120 (or similar hardware). The NDP120 is a low power device and can natively run a variety of model architectures. The NDP120 also has a DSP core for preprocessing the audio and generating the enhanced audio stream from the model output. Further, the NDP120 can handle four digital microphone inputs and has all the peripherals needed to easily interface to Twin or Anvil.
[0093] FIG. 14 shows a high-level system diagram 550 of the communications device 12 (which was shown in FIG. 1). Processor 552 may be a general-purpose processor, an application-specific integrated circuit, a field programmable gate arrays (FPGAs), a complex programmable logic device (CPLDs), microcontrollers, digital-to-analog converters (DACs), analog-to-digital converters (ADCs), or any other equivalent processing circuitry, such as fixed-function circuitry, programmable processing circuitry, and / or integrated or discrete logic circuitry. It is communicatively coupled to the other modules shown in system diagram 550, by for example one or more communications busses. An Al accelerator 554, described above, is tasked with applying an individualized speech enhancement model to digitized utterance in the form of a data stream, yielding an enhanced digitized utterance. A memory 556 is at the disposal of processor 552 and Al accelerator 554, and may be any form of digital memory. Analog-to- digital converter 558 works in conjunction with digital-to-analog converter to 560 to convert incoming and outgoing audio to analog or digital and vice-versa, as needed. Radio 548 communicates with other communication devices in the area, and / or earpieces depending on implementation. Generic model database 562 contains one or more generic speech enhancement models that would be applied by processor 552 and / or Al accelerator 554 in the absence of a user-specific speech enhancement model. User-specific model database 564 contains one or a plurality of user-specific individualized speech enhancement models. The data components 562 and 564 Communications device 12 includes one or twoearpieces 568 communicatively coupled via tether 566. Tether 556 may be, for example, communication wires, or could be wireless, e.g. may conform to Bluetooth™ or another wireless protocol. The databases mentioned may be implemented in memory 556. The module shown in system diagram 550 may be implemented on one or more circuit boards, as is known in the art. As described above, earpieces 568 may include their own circuit boards and processors, which may supersede the need for some of the components shown in system diagram 550, depending on implementation. For example, it is possible that the earpieces include some on-board real-time processing capabilities using on-board processors and converters, microphones, and speakers. Other system diagrams are possible, as would be known by the artisan, depending on implementation. The design shown in FIG. 1, for example, uses tethered earpieces and a communications unit that is worn on the hip. It is possible to implement this same architecture in a wireless pure earpiece form factor, where one primary earpiece handles most common processing, and is wirelessly communicatively coupled to the secondary earpiece.
[0094] FIG. 15 shows a flow chart illustrating an exemplary process 500 that communications device 12 and its processors would use to apply an individualized speech enhancement model to a digitized utterance received from a wearer. At step 502, an utterance is received from a wearer. As mentioned above, the utterance is received by a microphone after it has traversed human tissue of the wearer, to an in-ear microphone. It is thus received in an attenuated form by the in-ear microphones. The utterance is digitized using an analog-to -digital converter, to produce a digitized utterance, in step 504. In step 506, the individualized speech enhancement model is applied to the digitized utterance to create an individually enhanced digitized utterance, which is then further processed as needed and ultimately communicated via radio to the communication units of other users in a work area (step 508), where it is received and converted to analog and output to an in-ear speaker of a different wearer.
[0095] FIG. 16 is a flowchart 600 showing an exemplary process whereby a user may check out a communications device which may or may not have an individualized speech enhancement model on it. In step 602, the user authenticates to a system using, for example, an app on a smart phone. The app is communicatively coupled to a communications device 12, as described earlier, by known communications protocols like Bluetooth™. Upon authentication, a back-and-forth communications pathway is opened (step 604) whereby information about what individualized speech enhancement models are present and ready to be applied on the communications device. Under such arrangement, each individualized speech enhancement model has a unique model identifier, and a user’s profile is associated with the unique model identifier. When the app communicates with communications device (step 604) and determines if the unique model identifier associated with the user is present and ready on the communications device (step 606). If yes, the device is ready for communications and may present the user an indication as such (using auditory or visual queues, depending on what user interface is present on the communications device), and proceed to apply the individualized speech enhancement model to the wearer’s utterances in real-time or near real time, then using a radio to broadcast individually enhanced digitized utterances to other users’ communications devices in the area, as described earlier (step 610). Ifthe individualized speech enhancement model is not present on the communications device, it is uploaded from the app, or the app may retrieve it from elsewhere (or cause the communications device to do the same) (step 608). If no individualized speech enhancement model exists for the user, the communications device is directed to use a generic model.
[0096] In the present detailed description of the preferred embodiments, reference is made to the accompanying drawings, which illustrate specific embodiments in which the invention may be practiced. The illustrated embodiments are not intended to be exhaustive of all embodiments according to the invention. It is to be understood that other embodiments may be utilized, and structural or logical changes may be made without departing from the scope of the present invention. The following detailed description, therefore, is not to be taken in a limiting sense, and the scope of the present invention is defined by the appended claims.
[0200] Unless otherwise indicated, all numbers expressing feature sizes, amounts, and physical properties used in the specification and claims are to be understood as being modified in all instances by the term “about.” Accordingly, unless indicated to the contrary, the numerical parameters set forth in the foregoing specification and attached claims are approximations that can vary depending upon the desired properties sought to be obtained by those skilled in the art utilizing the teachings disclosed herein.
[0201] As used in this specification and the appended claims, the singular forms “a,” “an,” and “the” encompass embodiments having plural referents, unless the content clearly dictates otherwise. As used in this specification and the appended claims, the term “or” is generally employed in its sense including "and / or" unless the content clearly dictates otherwise.
[0202] Spatially related terms, including but not limited to, “proximate,” “distal,” “lower,” “upper,” “beneath,” “below,” “above,” and “on top,” if used herein, are utilized for ease of description to describe spatial relationships of an element(s) to another. Such spatially related terms encompass different orientations of the device in use or operation in addition to the particular orientations depicted in the figures and described herein. For example, if an object depicted in the figures is turned over or flipped over, portions previously described as below, or beneath other elements would then be above or on top of those other elements.
[0203] As used herein, when an element, component, or layer for example is described as forming a “coincident interface” with, or being “on,” “connected to,” “coupled with,” “stacked on” or “in contact with” another element, component, or layer, it can be directly on, directly connected to, directly coupled with, directly stacked on, in direct contact with, or intervening elements, components or layers may be on, connected, coupled or in contact with the particular element, component, or layer, for example. When an element, component, or layer for example is referred to as being “directly on,” “directly connected to,” “directly coupled with,” or “directly in contact with” another element, there are no intervening elements, components or layers for example. The techniques of this disclosure may be implemented in a wide variety of computer devices, such as servers, laptop computers, desktop computers, notebook computers, tablet computers, hand-held computers, smart phones, and the like. Any components, modules or unitshave been described to emphasize functional aspects and do not necessarily require realization by different hardware units. The techniques described herein may also be implemented in hardware, software, firmware, or any combination thereof. Any features described as modules, units or components may be implemented together in an integrated logic device or separately as discrete but interoperable logic devices. In some cases, various features may be implemented as an integrated circuit device, such as an integrated circuit chip or chipset. Additionally, although a number of distinct modules have been described throughout this description, many of which perform unique functions, all the functions of all of the modules may be combined into a single module, or even split into further additional modules. The modules described herein are only exemplary and have been described as such for better ease of understanding.
[0204] If implemented in software, the techniques may be realized at least in part by a computer-readable medium comprising instructions that, when executed in a processor, performs one or more of the methods described above. The computer-readable medium may comprise a tangible computer-readable storage medium and may form part of a computer program product, which may include packaging materials. The computer-readable storage medium may comprise random access memory (RAM) such as synchronous dynamic random-access memory (SDRAM), read-only memory (ROM), non-volatile random access memory (NVRAM), electrically erasable programmable read-only memory (EEPROM), FLASH memory, magnetic or optical data storage media, and the like. The computer-readable storage medium may also comprise a non-volatile storage device, such as a hard-disk, magnetic tape, a compact disk (CD), digital versatile disk (DVD), Blu-ray™ disk, holographic data storage media, or other non-volatile storage device.
[0205] The term “processor,” as used herein may refer to any of the foregoing structure or any other structure suitable for implementation of the techniques described herein. In addition, in some aspects, the functionality described herein may be provided within dedicated software modules or hardware modules configured for performing the techniques of this disclosure. Even if implemented in software, the techniques may use hardware such as a processor to execute the software, and a memory to store the software. In any such cases, the computers described herein may define a specific machine that is capable of executing the specific functions described herein. Also, the techniques could be fully implemented in one or more circuits or logic elements, which could also be considered a processor.
[0206] In one or more examples, the functions described may be implemented in hardware, software, firmware, or any combination thereof. If implemented in software, the functions may be stored on or transmitted over, as one or more instructions or code, a computer-readable medium and executed by a hardware-based processing unit. Computer-readable media may include computer-readable storage media, which corresponds to a tangible medium such as data storage media, or communication media including any medium that facilitates transfer of a computer program from one place to another, e.g., according to a communication protocol. In this manner, computer-readable media generally may correspond to (1) tangible computer-readable storage media, which is non-transitory or (2) acommunication medium such as a signal or carrier wave. Data storage media may be any available media that can be accessed by one or more computers or one or more processors to retrieve instructions, code and / or data stmctures for implementation of the techniques described in this disclosure. A computer program product may include a computer-readable medium.
[0207] By way of example, and not limitation, such computer-readable storage media can comprise RAM, ROM, EEPROM, CD-ROM or other optical disk storage, magnetic disk storage, or other magnetic storage devices, flash memory, or any other medium that can be used to store desired program code in the form of instructions or data stmctures and that can be accessed by a computer. Also, any connection is properly termed a computer-readable medium. For example, if instructions are transmitted from a website, server, or other remote source using a coaxial cable, fiber optic cable, twisted pair, digital subscriber line (DSL), or wireless technologies such as infrared, radio, and microwave, then the coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technologies such as infrared, radio, and microwave are included in the definition of medium. It should be understood, however, that computer- readable storage media and data storage media do not include connections, carrier waves, signals, or other transient media, but are instead directed to non-transient, tangible storage media. Disk and disc, as used, includes compact disc (CD), laser disc, optical disc, digital versatile disc (DVD), floppy disk and Blu-ray disc, where disks usually reproduce data magnetically, while discs reproduce data optically with lasers. Combinations of the above should also be included within the scope of computer-readable media.
[0208] Instructions may be executed by one or more processors, such as one or more digital signal processors (DSPs), general purpose microprocessors, application specific integrated circuits (ASICs), field programmable logic arrays (FPGAs), programmable processing circuitry, fixed-function circuitry, or other equivalent integrated or discrete logic circuitry. Accordingly, the term “processor”, as used may refer to any of the foregoing structure or any other structure suitable for implementation of the techniques described. In addition, in some aspects, the functionality described may be provided within dedicated hardware and / or software modules. Also, the techniques could be fully implemented in one or more circuits or logic elements.
[0209] The techniques of this disclosure may be implemented in a wide variety of devices or apparatuses, including a wireless handset, an integrated circuit (IC) or a set of ICs (e.g., a chip set). Various components, modules, or units are described in this disclosure to emphasize functional aspects of devices configured to perform the disclosed techniques, but do not necessarily require realization by different hardware units. Rather, as described above, various units may be combined in a hardware unit or provided by a collection of interoperative hardware units, including one or more processors as described above, in conjunction with suitable software and / or firmware.
[0210] It is to be recognized that depending on the example, certain acts or events of any of the methods described herein can be performed in a different sequence, may be added, merged, or left out altogether (e.g., not all described acts or events are necessary for the practice of the method). Moreover, in certain examples, acts or events may be performed concurrently, e.g., through multi-threaded processing,interrupt processing, or multiple processors, rather than sequentially.
[0211] In some examples, a computer-readable storage medium includes a non-transitory medium. The term “non-transitory” indicates, in some examples, that the storage medium is not embodied in a carrier wave or a propagated signal. In certain examples, a non-transitory storage medium stores data that can, over time, change (e.g., in RAM or cache).
[0212] Various examples have been described. These and other examples are within the scope of the following claims.
Claims
WHAT IS CLAIMED IS:
1. A hearing protection device comprising: a sound muffling portion comprising a sound dampening material configured to fit at least partially within a user’s ear canal when donned, to define a sound attenuation area that includes the user’s eardrum, and further comprising a microphone disposed to sense audio signals within the sound attenuation area; and a processor communicatively coupled to the microphone, the processor being configured to: upon the user generating a first utterance, receive attenuated audio signals associated with the first utterance; and apply an individualized speech enhancement model to the audio signals to yield transformed audio signals associated with the first utterance, wherein the individualized speech enhancement model comprises a model that is pre-trained model specifically for the user.
2. The hearing protection device of claim 1, wherein the attenuated audio signals are attenuated by transmission through the user’s human tissue to the sound attenuation area.
3. The hearing protection device of claim 1, wherein the individualized speech enhancement model comprises an audio transformation model.
4. The hearing protection device of claim 1, wherein the microphone comprises an in- ear microphone.
5. The hearing protection device of claim 1, wherein the individualized speech enhancement model is pre -trained according to a transfer learning paradigm.
6. The hearing protection device of claim 5, wherein the individualized speech enhancement model is pre-trained according to the transfer learning paradigm by: obtaining speech samples associated with one or more base users different from the user;applying a filter to the speech samples to form a training dataset; training a base speech enhancement model using the training dataset; and transferring a training of the base speech enhancement model to the individualized speech enhancement model.
7. The hearing protection device of claim 6, wherein the training dataset comprises filtered samples that simulate one or more in-ear pickups of the speech samples.
8. The hearing protection device of claim 7, wherein the training dataset further fullfidelity recordings associated with the speech samples.
9. The hearing protection device of claim 6, wherein the training dataset further includes speech samples associated with the user.
10. The hearing protection device of claim 6, further comprising training the individualized speech enhancement model using one or more speech samples associated with the user.
11. The hearing protection device of claim 6, further comprising deploying the individualized speech enhancement model in an execution phase upon determining that the individualized speech enhancement model has achieved convergence.
12. The hearing protection device of claim 1, further comprising a radio configured to broadcast a data stream indicative of the transformed audio signals.
13. The hearing protection device of claim 1, wherein the processor comprises an artificial intelligence (Al) accelerator in combination with a digital signal processing (DSP) core.
14. A method comprising: defining a sound attenuation area that includes an eardrum of a user; sensing audio signals within the sound attenuation area;receiving atenuated audio signals associated with the first uterance; and applying an individualized speech enhancement model to the audio signals to yield transformed audio signals associated with the first uterance, wherein the individualized speech enhancement model comprises a model that is pre-trained model specifically for the user.
15. A hearing protection apparatus comprising: means for defining a sound atenuation area that includes an eardrum of a user; means for sensing audio signals within the sound atenuation area; means for receiving atenuated audio signals associated with the first uterance; and means for applying an individualized speech enhancement model to the audio signals to yield transformed audio signals associated with the first uterance, wherein the individualized speech enhancement model comprises a model that is pre-trained model specifically for the user.
Citation Information
Patent Citations
In-Ear Utility Device Having Dual Microphones
US20170347183A1
Multi-part eardrum-contact hearing aid placed deep in the ear canal
US20210211811A1
Transforming speech signals to attenuate speech of competing individuals and other noise
US20240161765A1