Headphone device, method for operating the headphone device and computer-readable storage medium
The headphone device with a detachable microphone and array switches modes to enhance speech capture and suppress noise, addressing call quality issues in noisy environments by using cardioid patterns to filter out interference.
Patent Information
- Application Number
- PCT/CN2023/141878
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-12-26
- Publication Date
- 2025-07-03
AI Technical Summary
Existing headphones with a single pluggable boom microphone struggle to maintain clear phone call quality in noisy environments due to capturing both user speech and environmental noise, and the microphone is often removed for convenience, leading to unsatisfactory call quality.
A headphone device with a detachable first microphone and a microphone array that switches operating modes based on the microphone's connection state, using forward and backward cardioid patterns to enhance user speech capture and suppress noise and interference.
The device effectively filters out environmental noise and interference, maintaining clear voice capture whether the boom microphone is plugged in or unplugged, improving call quality in noisy conditions.
Smart Images

Figure CN2023141878_03072025_PF_FP_ABST
Abstract
Description
HEADPHONE DEVICE, METHOD FOR OPERATING THE HEADPHONE DEVICE AND COMPUTER-READABLE STORAGE MEDIUMTECHNICAL FIELD
[0001] The present disclosure relates to sound signal processing. In particular, it relates to a headphone device, method for operating the headphone device and computer-readable storage medium.BACKGROUND
[0002] Headphones are widely used in industrial and consumer products. For phone call purpose, these headphones generally contain one or more microphones. Users may use this type of headphones in a noisy environment, such as in a running metro or a heavily noisy industry factory. Also, those headphones are used in some scenarios which contain strong interferences, such as a user (player) wears a gaming headphone, while other players or audiences next to the user talk, cheer or yell. Most current designs are not able to cancel or diminish those extremely strong noises and interferences well.
[0003] Pluggable boom microphone can be placed or bended near a user’s mouth and capture the user’s speech with high signal to noise ratio (SNR) , especially with a cardioid beam towards a user’s mouth. However, it still encounters a number of challenges such as noises from the environment and interferences from other sources. In the case that only a single boom microphone is provided in a headphone, the phone call quality can be unsatisfactory sometimes. For example, headphones with a single pluggable microphone are widely used by operators or customer service representatives from call centers, but the call quality are often very poor as the microphone not only captures the user’s speech, but also captures environmental noise and voices from others in the same room.
[0004] Sometimes, users wear the headphone may take the pluggable boom microphone off for reasons of appearance and convenience. For instance, people prefer to bring a smaller and more compact headphone when they are travelling. There is a demand that the headphone can still capture the user’s speech even when the pluggable boom microphone is unplugged.
[0005] SUMMARY OF THE DISCLOSURE
[0006] In order to solve at least one of the above problems, embodiments of the present disclosure propose a headphone device, method for operating the headphone device and computer-readable storage medium, which can switch the working mode depending on the connecting state of a pluggable boom microphone. Besides, embodiments of the present disclosure can also enhance voice signal from a user and suppress environmental noise and interferences from other directions, which could improve the call quality when the headphone device is used to make a phone call in an extremally noisy environment.
[0007] In one representative aspect, there is provided a headphone device, comprising: two earpieces configured to cover a user’s ears when the headphone device is mounted on the user’s head; a main strap configured to connect to the two earpieces; a first microphone detachably connected to one of the two earpieces; a microphone array of at least two microphones arranged on one of the two earpieces; and a control unit configured to switch an operating mode of the headphone device between a first mode and a second mode depending on a connection state of the first microphone. In a case that the first microphone is plugged in, the headphone device is switched to the first mode and a first signal captured by the first microphone is used as a main signal. In a case that the first microphone is unplugged, the headphone device is switched to the second mode and the main signal is determined based on signals captured by the microphone array.
[0008] In some embodiments, the headphone device further comprises a detection unit configured to detect whether the first microphone is plugged into the one of the two earpieces.
[0009] In some embodiments, the headphone device further comprises a noise and interference determining unit configured to determine noise and interference based on the signals captured by the microphone array.
[0010] In some embodiments, the headphone device further comprises a user speech extractor configured to extract speech signal of the user by eliminating the noise and interference from the main signal in each of the first mode or the second mode.
[0011] In some embodiments, in the first mode, the main signal is captured from the first microphone has a polar mode of a forward cardioid pattern towards a mouth of the user, and the noise and interference determining unit is configured to determine the noise and interference based on a backward cardioid pattern among the microphone array.
[0012] In some embodiments, in the second mode, a forward cardioid beamforming among the microphone array is carried out to determine the main signal, and the noise and interference determining unit is configured to determine the noise and interference based on a backward cardioid pattern among the microphone array.
[0013] In some embodiments, the forward cardioid pattern of the first microphone maximizes a capture of speech signal of the user and minimizes a capture of the noise and interference in an opposite direction of the mouth of the user in the first mode.
[0014] In some embodiments, the backward cardioid pattern among the microphone array minimizes a capture of speech signal of the user and maximizes a capture of the noise and interference in an opposite direction of the mouth of the user in the first or the second mode.
[0015] In some embodiments, the forward cardioid beamforming among the microphone array maximizes a capture of speech signal of the user and minimize a capture of the noise and interference in an opposite direction of the mouth of the user in the second mode.
[0016] In some embodiments, each microphone of the microphone array is an omnidirectional microphone.
[0017] In another representative aspect, there is provided a method for operating a headphone device, the headphone device comprises two earpieces configured to cover a user’s ears when the headphone device is mounted on the user’s head; a main strap configured to connect to the two earpieces; a first microphone detachably connected to one of the two earpieces; and a microphone array of at least two microphones arranged on one of the two earpieces. The method comprising: switching, by using a control unit, an operating mode of the headphone device between a first mode and a second mode depending on a connection state of the first microphone. In a case that the first microphone is plugged in, the headphone device is switched to the first mode and a first signal captured by the first microphone is used as a main signal. In a case that the first microphone is unplugged, the headphone device is switched to the second mode and the main signal is determined based on signals captured by the microphone array.
[0018] In some embodiments, the method further comprising: detecting, by a detection unit, whether the first microphone is plugged into the one of the two earpieces.
[0019] In some embodiments, the method further comprising: determining, by a noise and interference determining unit, noise and interference based on the signals captured by the microphone array.
[0020] In some embodiments, the method further comprising: extracting, by a user speech extractor, speech signal of the user by eliminating the noise and interference from the main signal in each of the first mode or the second mode.
[0021] In some embodiments, in the first mode, the main signal is captured from the first microphone has a polar mode of a forward cardioid pattern towards a mouth of the user, and the noise and interference determining unit is configured to determine the noise and interference based on a backward cardioid pattern among the microphone array.
[0022] In some embodiments, in the second mode, a forward cardioid beamforming among the microphone array is carried out to determine the main signal, and the noise and interference determining unit is configured to determine the noise and interference based on a backward cardioid pattern among the microphone array.
[0023] In some embodiments, the forward cardioid pattern of the first microphone maximizes a capture of speech signal of the user and minimizes a capture of the noise and interference in an opposite direction of the mouth of the user in the first mode.
[0024] In some embodiments, the backward cardioid pattern among the microphone array minimizes a capture of speech signal of the user and maximizes a capture of the noise and interference in an opposite direction of the mouth of the user in the first or the second mode.
[0025] In some embodiments, the forward cardioid beamforming among the microphone array maximizes a capture of speech signal of the user and minimize a capture of the noise and interference in an opposite direction of the mouth of the user in the second mode.
[0026] In yet another representative aspect, there is provided a non-transitory computer readable storage medium with instruction stored thereon, when executed, the instructions cause a processor of a headphone device to execute any one of the methods described above.
[0027] This summary is provided to introduce a selection of concepts in a simplified form that are further described below in the description. This summary is not intended to identify key features or essential features of the claimed subject matter, nor is it intended to be used to limit the scope of the claimed subject matter.
[0028] These and other aspects are described in the present disclosure.BRIEF DESCRIPTION OF THE DRAWINGS
[0029] Various features of examples and embodiments in accordance with the principles described herein may be more readily understood with reference to the following detailed description taken in conjunction with the accompanying drawings, where like reference numerals designate like structural elements, and in which:
[0030] FIG. 1 illustrates an example headphone device working in a first mode according to one of a number of embodiments of the present disclosure;
[0031] FIG. 2 illustrates an example headphone device working in a second mode according to one of a number of embodiments of the present disclosure;
[0032] FIG. 3 illustrates an example diagram of the switching operation of the headphone device according to one of a number of embodiments of the present disclosure.
[0033] FIG. 4A illustrates an example diagram of backward cardioid pattern according to one of a number of embodiments of the present disclosure.
[0034] FIG. 4B illustrates an example diagram of forward cardioid pattern according to one of a number of embodiments of the present disclosure.
[0035] FIG. 5A illustrates an example diagram of signal processing of the headphone device in the first mode according to one of a number of embodiments of the present disclosure.
[0036] FIG. 5B illustrates an example of signal spectrums before and after processing in the first mode according to one of a number of embodiments of the present disclosure;
[0037] FIG. 6 illustrates an example diagram of signal processing of the headphone device in the second mode according to one of a number of embodiments of the present disclosure;
[0038] FIG. 7 illustrates an example flowchart of a method for operating the headphone device according to one of a number of embodiments of the present disclosure;
[0039] FIG. 8 illustrates an example non-transitory computer-readable medium for operating the headphone device according to one of a number of embodiments of the present disclosure.
[0040] These and other features are detailed below with reference to the above-referenced figures.
[0041] DESCRIPTION OF THE EMBODIMENTS
[0042] The preferred embodiments of the present disclosure will now be described with reference to the drawings. Identical elements in the various figures are identified with the same reference numerals.
[0043] Reference will now be made in detail to each embodiment of the present disclosure. Such embodiments are provided by way of explanation of the present disclosure, which is not intended to be limited thereto. In fact, those of ordinary skill in the art may appreciate upon reading the present specification and viewing the present drawings that various modifications and variations may be made thereto.
[0044] Although some embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms, and should not be interpreted as limited to the embodiments described herein.
[0045] It also should be noted that the steps described in the method embodiments of the present disclosure can be executed in different order and / or in parallel, unless it is obviously unsuitable or explicitly stated to the contrary. Further, the method embodiments may include more steps and / or omit certain steps.
[0046] The term “comprising” and its variants used herein refers to “including but not limited to” . The term “based on” or “depending on” refers to “at least based on” or “at least depending on” , respectively. It should also be understood that the term “first” , “second” or “third” mentioned in the present disclosure are only used to distinguish different devices, modules or elements, instead of to define the order or interdependence of functions performed by these devices, modules or units.
[0047] In order to simplify the description, various examples of headphone devices and methods for operating the headphone device according to a number of embodiments of the present disclosure are described below.
[0048] FIG. 1 illustrates an example headphone device 1000 working in a first mode (i.e. MODE 1) according to one of a number of embodiments of the present disclosure.
[0049] For the convenience of description, the headphone device 1000 shown in FIG. 1 is in a state of being worn on the head of an artificial head, which can be treated as an avatar of a user.
[0050] As shown in FIG. 1, the headphone device 1000 comprises two earpieces 1100, a main strap 1200 configured to connect to the two earpieces 1100, and a first microphone Mic1 detachably connected to one of the two earpieces 1100, for example, the left earpiece 1100 as shown in FIG. 1.
[0051] The two earpieces 1100 are configured to cover a user’s ears when the headphone device 1000 is mounted on the user’s head.
[0052] For example, as shown in FIG. 1, the headphone device 1000 is mounted on the head of an artificial head and the two earpieces 1100 cover both ears of the artificial head.
[0053] FIG. 1 illustrates the two earpieces 1100 as two earphone cups that can fully cover the ears of the artificial head, which is only an example.
[0054] In other embodiments, the earpieces 1100 can also be in the form of a headphone pad that completely covers the user’s ear, or in the form of an earplug that partially covers the user’s ear, or any other possible shape or size.
[0055] Similarly, the main strap 1200 can be a flexible or rigid strap of any material and any shape, which is not limited to the shape or size of the main strap 1200 shown in FIG. 1.
[0056] Generally, the first microphone Mic1 is a microphone with a slender rod, one end of which can be detachably connected to one of the two earpieces 1100 (e.g. left one or the right one) , and the other end is located near the user’s mouth when the headphone device 1000 is mounted on the user’s head.
[0057] The pole of the first microphone Mic1, for example, the end that contains the sound capturing elements, is located close to the user’s mouth when the headphone device 1000 is mounted on the user’s head, which helps to capture the user’s speech with relative high signal-to-noise ratio (SNR) .
[0058] In addition, the first microphone Mic1 as shown is preferred to be rotatable relative to the connection end, and the slender rod of the first microphone Mic1 is preferred to be bendable, which helps users adjust the posture of the first microphone Mic1 to fit their head size when one end of the first microphone Mic1 is connected to the one of the two earpieces 1100, thereby enabling the first microphone Mic1 to capture the user’s speech in the optimal position.
[0059] For example, the first microphone Mic1 can be a boom type microphone, which is adjustable and controlled by the user when it is connected to the one of the two earpieces 1100.
[0060] In addition to the first microphone Mic1, the headphone device 1000 further comprises a microphone array 1300, which contains at least two microphones.
[0061] For simplicity, FIG. 1 shows an example of the microphone array 1300 with two microphones, i.e., a second and third microphone Mic2, Mic3.
[0062] In some embodiments, each microphone of the microphone array 1300 can be an omnidirectional microphone. For example, the second and third microphone Mic2, Mic3 are omnidirectional microphones, which have equal sensitivity and response to sounds in all directions.
[0063] As shown in FIG. 1, the second and the third microphone Mic2, Mic3 as an example of the microphone array 1300 are arranged on one of the two earpieces 1100, for example, the left earpiece 1100 connected with the first microphone Mic1.
[0064] In addition, compared with the second microphone Mic2, the third microphone Mic3 is further away from the mouth of the user.
[0065] Please note that other arrangements of the first microphone Mic1 and the microphone array 1300 are also possible.
[0066] For example, the first microphone Mic1 can be connected to the right earpiece 1100 when plugged in, while the microphone array 1300 of the second and the third microphone Mic2, Mic3 can be arranged on a different earpiece 1100, for example, on the right earpiece 1100, which is not shown in FIG. 1.
[0067] The headphone device 1000 further comprises a control unit (not shown) , which is configured to switch an operating mode of the headphone device 1000 between a first mode and a second mode depending on a connection state of the first microphone Mic1.
[0068] Specifically, in a case that the first microphone Mic1 is plugged in, as shown in FIG. 1, the headphone device 1000 works in the first mode. The first mode is a mode in which a first signal captured by the first microphone Mic1 is used as a main signal.
[0069] Herein, the term “main signal” refers to the dominant signal containing the desired user’s voice or speech to be captured, and the main signal is generally used to extract the user’s voice or speech by eliminating noise and interference signals from it.
[0070] In addition, the term “noise” or “environmental noise” herein generally refers to the random and irregular sound generated in the environment where the headphone device 1000 is located, such as white noise produced by various devices in the room, noise from colleagues typing on mice or keyboards in the office, or sound generated by wind, rain, and lightning in nature. This type of noise is usually not produced from a single fixed source and has a certain degree of randomness and directionless.
[0071] Besides, the term “interference” herein typically refers to the sound produced from a specific direction by a specific source in the environment where the headphone device 1000 is located. For example, the sound produced by a player playing music in a room, or the voice of other colleagues making phone calls in the office. This type of interference typically comes from a specific source and has a specific direction.
[0072] In general, signals captured by any of the three microphones Mic1, Mic2 and Mic3 not only include the user’s speech, but also include environmental noise and directional interference from other sources as described above.
[0073] The purpose of this disclosure is to switch the working mode of the headphone device 1000 based on the connection status of the first microphone Mic1, in order to select an appropriate signal as the main signal to extract the user’s speech.
[0074] Meanwhile, when multiple microphones are arranged in the headphone device, if the signal captured by one or a group of microphones is used as the main signal, the signals captured by the other microphones can be used as auxiliary signals to remove noise and interference.
[0075] As shown in FIG. 1, if the first microphone Mic1 is plugged in, the headphone device 1000 will be switched into the first mode. Then, the first signal captured by the first microphone Mic1 will be used as the main signal and will be processed to remove noise and interference signals, thereby extracting signals representing user’s voice or speech.
[0076] Herein, the extracted user speech signal can also be referred to as estimated user speech signal.
[0077] In another scenario, for example, if the user unplugs the first microphone Mic1, the headphone device 1000 will be switched to work in the second mode different from the first mode. In the second mode, instead of using the first signal captured by the first microphone Mic1 as the main signal, the main signal will be determined based on signals captured by the microphone array 1300.
[0078] FIG. 2 illustrates an example headphone device 1000 working in a second mode (i.e. MODE 2) according to one of a number of embodiments of the present disclosure;
[0079] Different from the headphone device 1000 shown in FIG. 1, in FIG. 2, the first microphone Mic1 of the headphone device 1000 is unplugged, which means that the first microphone Mic1 will no longer capture speech of the user, only the microphone array 1300 arranged on the headphone 1200 will capture the speech of the user, as well as environmental noise and interferences.
[0080] The “switching” of the working mode for the headphone device 1000 is implemented by a control unit (not shown) , which can be implemented through any suitable hardware or software, such as a processor embedded in the headphone device 1000, hardware circuits in the headphone device 1000, or software instructions stored in the memory of the headphone device 1000.
[0081] In addition, to detect whether the first microphone Mic1 is plugged in or not, the headphone device 1000 may also comprise a detection unit 1400, as shown in FIG. 3, which illustrates an example diagram of the switching operation of the headphone device according to one of a number of embodiments of the present disclosure.
[0082] As shown in FIG. 3, the detection unit 1400 is configured to detect whether the first microphone Mic1 is plugged into the one of the two earpieces 1100.
[0083] If “Yes” , i.e., if the detection unit 1400 determines that the first microphone Mic1 is plugged in, the headphone device 1000 will be switched into MODE 1, in which a first signal captured by the first microphone Mic1 will be used as the main signal.
[0084] Otherwise, i.e., if the detection unit 1400 determines that the first microphone Mic1 is unplugged, the headphone device 1000 will be switched into MODE 2, in which signals captured by the microphone array 1300 will be used to determine the main signal.
[0085] In some embodiments, the detection unit 1400 can be implemented by a resistor detection circuit, and the resistor detection circuit can be embedded in the headphone device 1000.
[0086] For example, the resistor detection circuit may determine whether the first microphone Mic1 is in a connection state by detecting resistance between predefined two points.
[0087] For example, if the value of the resistance is detected to be greater than a first threshold, such as detecting an open circuit between the predefined two points, it can be determined that the first microphone Mic1 has been unplugged. Otherwise, it can be determined that the first microphone Mic 1 is plugged in.
[0088] Other implementation methods of the detection unit 1400 are also possible, such as capacitor based detecting or detecting through pressure sensors, and so on.
[0089] So far, the basic concept of the disclosed headphone device 1000 has been introduced in conjunction with FIG 1 to FIG. 3, which benefits from multiple microphones switching between different working modes, making it convenient for users to use, so that even when the first microphone Mic1 is unplugged, the microphone array arranged on the headphone device can still be used to capture sound.
[0090] Next, by referring to FIG. 4A to FIG. 6, as well as FIG. 1 and FIG. 2, we will further describe how the headphone device proposed in this disclosure filters out environmental noise and interference and estimates the user speech signal.
[0091] As described above, in general, signals captured by a microphone not only includes the user’s speech, but also may include environmental noise and / or directional interference from other sources.
[0092] In order to remove the environmental noise and / or interference, the headphone device proposed in the present disclosure further comprises a noise and interference determining unit (not shown) , which can be configured to determine noise and interference to be removed from the main signal.
[0093] As described above, when multiple microphones are arranged in a headphone device, if the signal captured by one or a group of microphones is used as the main signal, the signals captured by the other microphones can be used as auxiliary signals to remove noise and interference.
[0094] In a preferred embodiments, the noise and interference determining unit is configured to determine noise and interference based on the signals captured by the microphone array, for example, by the microphone array 1300 shown in FIG. 1 or FIG. 2.
[0095] Similar to the control unit which is configured to switch an operating mode of the headphone device 1000 between a first mode and a second mode depending on a connection state of the first microphone Mic1, the noise and interference determining unit as described can also be implemented with any suitable hardware or software, such as a processor embedded in the headphone device 1000, hardware circuits in the headphone device 1000, or software instructions stored in the memory of the headphone device 1000.
[0096] Referring back to FIG. 1, when the headphone device 1000 is working in the first mode, i.e. MODE 1, the main signal is captured from the first microphone Mic1. Preferably, the first microphone Mic1 is provided to have a polar mode of a forward cardioid pattern towards a mouth of the user, as shown in FIG . 1.
[0097] A polar mode of a microphone generally indicates the pickup range of a microphone. If the polar mode of a microphone has a forward cardioid pattern, it means that in front of the microphone, its sensitivity to audio signals is very high and the capture of sound will be maximized. In contrast, on the side of the microphone (e.g., at 90 degrees) , its sensitivity is also good, but it is 6 dB lower than the front end. Another important aspect is that the microphone has a very good shielding effect on the sound coming from behind the microphone, which may minimize the capture of the sounds from this direction.
[0098] For example, taking the cardioid pattern H1 surrounding the first microphone Mic1 in FIG. 1 as an example, the front of the cardioid pattern H1 directly points towards a mouth of the user, so that the speech of the user will be captured by the first microphone Mic1 without attenuation or with minimum attenuation. Meanwhile, the sound from the back of the cardioid pattern H1 will be completely shielded or shielded as much as possible.
[0099] Therefore, ideally, the first signal captured by the first microphone Mic1 will contain user’s speech without any attenuation. In addition, the first signal may also contain partially attenuated noise or interference from other directions.
[0100] As described above, in the first mode, the first signal captured by the first microphone Mic1 will be used as the main signal.
[0101] For example, if we denote the user speech signal as S and the mixed noise as n and interference as I, the captured signal by the first microphone Mic1 at the time t is a mixture of those signals and can be denoted as x1 (t) =α1S (t) +β1 n (t) +γ1I (t+τI1) , (1)
[0102] wherein α1, β1 and γ1 represent the weight of each component signal, respectively, and τI1 represents time delay from the source that generated the interference I.
[0103] As the first microphone Mic1 is close to the mouth of the user and the forward cardioid pattern H1 attenuates the noises and interferences from other directions, the signal to noise ratio is very high and α1, β1 and γ1 generally satisfy the below condition. α1≈1, β1<< α1, γ1<<α1 (2)
[0104] Similarly, the signal captured by the microphone array 1300 will also contain the user’s sound, as well as noise and / or interference.
[0105] For example, the captured signal by the microphone array 1300 (e.g., Mic2 and Mic2 as shown in FIG. 1) can be denoted as: x2 (t) =α2S (t+τs2) +β2 n (t) +γ2I (t+τI2) , (3) x3 (t) =α3S (t+τs3) +β3 n (t) +γ3I (t+τI3) . (4)
[0106] wherein x2 (t) and x3 (t) represents the signal captured by the second and third microphone Mic2, Mic3, respectively. τs2 and τs3 represents the time delay from the user’s mouth to the second and third microphone Mic2, Mic3, respectively. τI2 and τI3 represents the time delay from the source of the interference to the second and third microphone Mic2, Mic3, respectively. For example, suppose there is a specific source behind the user (e.g. person A) speaking, τI2 and τI3 represents the time delay of the speech signal of person A to the second and third microphone Mic2, Mic3, respectively.
[0107] In MODE 1, the captured signal x1 by the first microphone Mic1 is a mixture of the user speech signal S, the noise n and the interference I, as described in equation (1) . In the cased that the captured signal x1 by the first microphone Mic1 is used as the main signal, solutions are required to eliminate the noise n and the interference I from the main signal x1, so as to only extract the desired user speech signal S.
[0108] To eliminate the noise n and the interference I from the main signal, the present disclosure further proposes that the headphone device may comprises a noise and interference determining unit, as mentioned above, and the noise and interference determining unit can be configured to determine the noise and interference based on a backward cardioid pattern among the microphone array 1300.
[0109] Specifically, as shown in FIG. 1, a backward cardioid pattern H2 can be formed among the second and third microphone Mic1, Mic3 of the microphone array 1300.
[0110] Please note that the “backward” cardioid pattern discussed herein refers to a cardioid pattern which will attenuate or shield speech signals from the direction of the user’s mouth and generally maintain signals from the opposite direction.
[0111] The opposite of the backward cardioid pattern is a “forward” cardioid pattern, which will not attenuate the speech signal from the user’s mouth direction but shield signals from the opposite direction.
[0112] FIG. 4A illustrates an example diagram of backward cardioid pattern according to one of a number of embodiments of the present disclosure. FIG. 4B illustrates an example diagram of forward cardioid pattern according to one of a number of embodiments of the present disclosure.
[0113] As shown in FIG. 4A, a source S1 is located at the “back” of the cardioid pattern, i.e., at the direction of 180°, the sound emitted in this specific direction by the source S1 will be completely shielded by the microphone array.
[0114] Therefore, herein, the cardioid pattern that has a 180° directionality relative to the source S1 is referred to as a backward cardioid pattern.
[0115] FIG. 4A shows the backward cardioid patterns of a dual microphone array and a three microphone array, respectively. The inner backward cardioid pattern is for a microphone array including 3 microphones, and the outer backward cardioid pattern is for a microphone array including 2 microphones.
[0116] The cardioid pattern of the microphone array can also be referred to as the response diagram of the array, where the coordinate value of the cardioid pattern in each direction represent the attenuation of sound signals in this direction.
[0117] For example, in FIG. 4A, the attenuation in the direction of 0° equals to 0dB, which means the signal in the direction of 0° will be completely captured by the microphone array without any attenuation.
[0118] In contrast, the attenuation in the direction of 180° resides at the center of the circle, which may indicate the biggest attenuation (e.g., ∞) , and it means the signal in the direction of 180° will be completely shileded by the microphone array.
[0119] Contrary to FIG. 4A, FIG. 4B illustrates a forward cardioid pattern with the source S1 located at the front of the cardioid pattern , i.e., at the direction of 0°.
[0120] In this case, the sound emitted by the source S1 will be completely captured by the microphone array without any attenuation, and the signals from the opposite direction will be completely shielded by the microphone array.
[0121] Therefore, herein, the cardioid pattern that has a 0° directionality relative to the source S1 is referred to as a forward cardioid pattern.
[0122] Please note that, when we discusse the forward or backward cardioid pattern in this disclosure, the directionality of the cardioid pattern is actually defined based on the relative location of the user’s mouth. For ease of understanding, the user’s mouth can be considered as the source S1 in FIG. 4A and FIG. 4B.
[0123] Now referring back to FIG. 1, the cardioid pattern H1 has a 0° directionality relative to the mouth of the user. Therefore, the cardioid pattern H1 is a forward cardioid pattern.
[0124] In contrast, the cardioid pattern H2 has a 180° directionality relative to the mouth of the user. Therefore, the cardioid pattern H2 is a backward cardioid pattern.
[0125] It should also be noted that the speech signal of the user will not be fully shielded by the microphone array 1300 even with a backward cardioid pattern beamforming, because the speech signal of the user is not strictly unidirectional, i.e., the speech signal of the user not only contains component signals in the direction of 180°, but also contains component signals in other directions beyond 180°. This is why the captured signal of the second and third microphone Mic2, Mic3 also contain the component of the user speech signal S, as indicated in the above equation (3) and (4) .
[0126] In the case that the noise and interference are determined based on a backward cardioid pattern among the microphone array 1300, a reference of the noise and interference can be determined as: xref (t) =x3 (t) -hb*x2 (t) (5)
[0127] wherein hb is a coefficient to characterize the ratio of user speech signals captured by Mic 3 and Mic 2.
[0128] The purpose of equation (5) is to remove user speech signal S as much as possible from the signal captured by Mic3, so that xref only contains noise and interference signals, which can be used as the reference of the noise and interference.
[0129] Once the reference of the noise and interference xref is determined, a noise and interference cancellation operation can be executed to extract user speech signal.
[0130] For example, the user speech signal can be extracted from the main signal by eliminating the noise and interference from the main signal, e.g., by using the following equation (6) :
[0131] Wherein w1 is a coefficient of a filter, and is the extracted user speech signal.
[0132] FIG. 5A illustrates an example block diagram of the above signal processing and functional units of the headphone device 1000 in MODE 1.
[0133] As shown in FIG. 5A, the block 5100 refers to a backward cardioid beamforming process of the microphone array, with the captured signal x2 and x3 by the second and third microphone Mic2, Mic3 as the inputs, and the reference of the noise and interference xref as output.
[0134] For example, the backward cardioid beamforming process in block 5100 can be mathematically implemented by using equation (5) as described above.
[0135] Then, the reference of the noise and interference xref will be input into the filter 5200 to determine the noise and interference to be filtered out from the main signal. For example, the noise and interference to be filtered out from the main signal can be indicated by w1*xref, which is similar to those in the above equation (6) .
[0136] Please note that an adaptive algorithm 5300 is used to calculated the coefficient w1 of the filter. For example, the adaptive algorithm 5300 may calculate the coefficient w1 of the filter based on the reference of the noise and interference xref (t1) of time t1 and the extracted user speech signal of the previous time t0.
[0137] Next, the calculated noise and interference w1*xref to be filtered out will be feed into the user speech extractor 5400, in which the calculated noise and interference w1*xref will be filtered out by subtracting it from the main signal (i.e., the captured signal x1 by the first microphone Mic1 in MODE 1) , thereby the estimated user speech signal is generated.
[0138] The subtracting process executed by the user speech extractor 5400 can be indicated by the above equation (6) mathematically.
[0139] FIG. 5B illustrates an example of signal spectrums before and after processing in the first mode according to one of a number of embodiments of the present disclosure.
[0140] As shown in FIG. 5B, the subgraph (a) shows the spectrum of the first signal captured by the first microphone Mic1 in MODE 1, in which some spectrum sections are marked in ellipses to represent noise and / or interference in the first signal.
[0141] The subgraph (b) and (c) shows the spectrum of the second and third signal captured by the second and third microphone Mic2, Mic3 in MODE 1, respectively.
[0142] The subgraph (d) shows the spectrum of the processed result, i.e., the estimated user speech signal as described above.
[0143] Compared with subgraph (a) , it’s clear that the noise and / or interference marked in ellipses are almost diminished after processing, as shown in subgraph (d) .
[0144] Various embodiments have been described above to elaborate how to remove noise and / or interference when the first microphone Mic1 is plugged in in MODE 1.
[0145] Embodiments of how to remove noise and / or interference when the first microphone Mic1 is unplugged in MODE 2 will be described in conjunction with FIG. 6 and FIG. 2.
[0146] Referring back to FIG. 2, the first microphone Mic1 of the headphone device 1000 is unplugged, which means that the first microphone Mic1 will no longer capture speech of the user, and only the microphone array 1300 (e.g., Mic2 and Mic3) arranged on the headphone 1200 will capture speech of the user, as well as environmental noise and interferences.
[0147] In MODE 2, a forward cardioid beamforming among the microphone array 1300 will be carried out to determine the main signal.
[0148] For example, as shown in FIG. 2, the second and third microphone Mic2 and Mic3 was surrounded by a forward cardioid pattern H3, which means that the mouth of the user is located in the front of the forward cardioid pattern H3, i.e., at the direction of 0°.
[0149] In this case, the speech emitted by the mouth of the user will be completely captured by the microphone array 1300 with the lowest attenuation, and the signals from the opposite direction will be shielded by the microphone array 1300 as much as possible.
[0150] Specifically, the main signal in MODE 2 can be determined by the following equation (7) : xf (t) =x2 (t) -hf*x3 (t) , (7)
[0151] Wherein x2 (t) and x3 (t) represents the signal captured by the second and third microphone Mic2, Mic3, respectively, which are similar to those of equations (3) and (4) . hf is a coefficient to characterize the ratio of noise and / or interference signals captured by Mic 2 and Mic 3.
[0152] Once the main signal xf is determined, a noise and interference cancellation operation can be executed to extract user speech signal.
[0153] For example, the estimated user speech signal can be extracted from the main signal by eliminating the noise and interference from the main signal, e.g., by using the following equation (8) :
[0154] Wherein w2 is a coefficient of a filter, and is the extracted user speech signal in MODE 2.
[0155] Please note that the reference of noise and interference xref can be similarly determined based on a backward cardioid pattern among the microphone array 1300, as indicated by the above equation (5) .
[0156] FIG. 6 illustrates an example block diagram of the above signal processing and functional units of the headphone device 1000 in MODE 2.
[0157] As shown in FIG. 6, the block 6100 refers to a forward cardioid beamforming process of the microphone array 1300, with the captured signal x2 and x3 by the second and third microphone Mic2, Mic3 as the inputs, and the calculated main signal xf as output.
[0158] For example, the forward cardioid beamforming process in block 6100 can be mathematically implemented by using equation (7) as described above.
[0159] In addition, the block 6200 refers to a backward cardioid beamforming process of the microphone array 1300, also with the captured signal x2 and x3 by the second and third microphone Mic2, Mic3 as the inputs, and the reference of the noise and interference xref as output.
[0160] For example, the backward cardioid beamforming process in block 6200 can be mathematically implemented by using equation (5) as described above.
[0161] Then, the reference of the noise and interference xref will be input into the filter 6300 to determine the noise and interference to be filtered out from the main signal xf.
[0162] For example, the noise and interference to be filtered out from the main signal can be indicated by w2*xref, which is similar to those in the above equation (8) .
[0163] Please note that the adaptive algorithm 6400 is similar to the adaptive algorithm 5300 shown in FIG. 5A, which is used to calculated the coefficient w2 of the filter.
[0164] For example, the adaptive algorithm 6400 may calculate the coefficient w2 of the filter based on the reference of the noise and interference xref (t1) of time t1, as well as the estimated user speech signal of the previous time t0.
[0165] Next, the calculated noise and interference w2*xref to be filtered out will be feed into the user speech extractor 6500, in which the calculated noise and interference w2* xref will be filtered out by subtracting it from the main signal (i.e., the calculated signal xf) , thereby the estimated user speech signal is generated.
[0166] The subtracting process executed by the extractor 6500 can be indicated by the above equation (8) mathematically.
[0167] Various embodiments have been described in conjunction with FIG. 6 and FIG. 2 to elaborate how to remove noise and / or interference when the first microphone Mic1 is unplugged in in MODE 2.
[0168] In general, the proposed headphone device can switch between two different working modes. When the pluggable microphone is plugged in, signal received by the pluggable microphone is used as the main signal and the other microphones in the array (mounted in an earpiece) form a backward cardioid beam which contains very little speech signal will be used to determine the noise and interference. When the pluggable microphone is unplugged, the adaptive differential beamforming algorithm is performed based on the signals captured by the microphones in the array, to enhance the speech signal of the user.
[0169] The proposed headphone device is able to provide a better speech enhancement and noise suppression.
[0170] Please note that in the above description, the microphone array 1300 is described to have two microphone Mic2 and Mic3, which is only an example, rather than a limitation.
[0171] In other embodiments, other numbers of microphones are also possible to form the microphone array 1300, for example, depending on the desired level of noise reduction or the desired level of interference reduction.
[0172] For example, the microphone array 1300 can include 3, 4, or even more microphones.
[0173] For example, in the case that the microphone array 1300 include 3 microphones, the captured signal by the microphone array 1300 (e.g., Mic2 and Mic3 as shown in FIG. 1, and an additional Mic4, not shown) can be denoted as: x2 (t) =α2S (t+τs2) +β2 n (t) +γ2I (t+τI2) , (3*) x3 (t) =α3S (t+τs3) +β3 n (t) +γ3I (t+τI3) . (4*) x4 (t) =α4S (t+τs4) +β4 n (t) +γ4I (t+τI4) . (5*)
[0174] In this case, the noise and interference can be similarly determined based on a backward cardioid pattern among the microphone array 1300.
[0175] For example, assuming the additional Mic 4 is located even further away from the user’s mouth than Mic2 and Mic3, then a reference of the noise and interference can be determined as: xref (t) =x4 (t) -h3x3 (t) -hb*x2 (t) (5**)
[0176] wherein h3 is a coefficient to characterize the ratio of user speech signals captured by Mic 4 and Mic 3.
[0177] Similarly, in the case that the microphone array 1300 include 3 microphones, when the headphone device 100 works in MODE 2, a forward cardioid beamforming among the microphone array 1300 will be carried out to determine the main signal.
[0178] For example, the second and third microphone Mic2 and Mic3, as well as the fourth microphone Mic4 (not shown) will be surrounded by a forward cardioid pattern, and the main signal in MODE 2 can be determined by the following equation (7*) : xf (t) =x2 (t) -hf*x3 (t) -h4*x4 (t) , (7*)
[0179] Wherein h4 is a coefficient to characterize the ratio of noise and / or interference signals captured by Mic 2 and Mic 4.
[0180] In the case that the microphone array 1300 include 3 microphones, the same noise and interference cancellation operation will be carried out based on the previous equation (6) and equation (8) in the first mode or in the second mode, which will not be repeated here.
[0181] The present disclosure also provides a method for operating the headphone device 1000 as described in various examples above.
[0182] For example, the headphone device 1000 comprises two earpieces 1100 as shown in FIG. 1, and the two earpieces 1100 are configured to cover a user’s ears when the headphone device 1000 is mounted on the user’s head.
[0183] The headphone device 1000 further comprises a main strap 1200, which is configured to connect to the two earpieces 1100.
[0184] As shown in FIG. 1, a first microphone Mic 1 is detachably connected to one of the two earpieces 1100, and a microphone array 1300 of at least two microphones is arranged on one of the two earpieces.
[0185] Specifically, the proposed method step of switching, by using a control unit, an operating mode of the headphone device 1000 between a first mode and a second mode depending on a connection state of the first microphone Mic.
[0186] FIG. 7 illustrates an example flowchart of a method 7000 for operating the headphone device according to one of a number of embodiments of the present disclosure.
[0187] As shown in FIG. 7, in step 7100, it is determined whether the first microphone is plugged in or not.
[0188] In a case that the first microphone Mic1 is plugged in (Yes) , the headphone device will work in the first mode and a first signal captured by the first microphone is used as a main signal.
[0189] In a case that the first microphone Mic2 is unplugged (No) , the headphone device will work in the second mode and signals captured by the microphone array will be used to determine the main signal.
[0190] In addition, the method 7000 for operating the headphone device may further comprise: determining, by a noise and interference determining unit, noise and interference based on the signals captured by the microphone array.
[0191] In some embodiments, the method 7000 further comprising: detecting, by a detection unit, whether the first microphone is plugged into the one of the two earpieces.
[0192] In some embodiments, the method 7000 further comprising: determining, by a noise and interference determining unit, noise and interference based on the signals captured by the microphone array.
[0193] In some embodiments, the method 7000 further comprising: extracting, by a user speech extractor, speech signal of the user by eliminating the noise and interference from the main signal in each of the first mode or the second mode.
[0194] In some embodiments, in the first mode, the main signal is captured from the first microphone has a polar mode of a forward cardioid pattern towards a mouth of the user, and the noise and interference determining unit is configured to determine the noise and interference based on a backward cardioid pattern among the microphone array.
[0195] In some embodiments, in the second mode, a forward cardioid beamforming among the microphone array is carried out to determine the main signal, and the noise and interference determining unit is configured to determine the noise and interference based on a backward cardioid pattern among the microphone array.
[0196] In some embodiments, the forward cardioid pattern of the first microphone maximizes a capture of speech signal of the user and minimizes a capture of the noise and interference in an opposite direction of the mouth of the user in the first mode.
[0197] In some embodiments, the backward cardioid pattern among the microphone array minimizes a capture of speech signal of the user and maximizes a capture of the noise and interference in an opposite direction of the mouth of the user in the first or the second mode.
[0198] In some embodiments, the forward cardioid beamforming among the microphone array maximizes a capture of speech signal of the user and minimize a capture of the noise and interference in an opposite direction of the mouth of the user in the second mode.
[0199] Besides, the various operations or functions performed by the various components or elements of the headphone device 1000 described above in conjunction with FIG. 1 to FIG. 6 are also applicable to the above methods, unless it is obviously inappropriate from the context, and will not be repeated here.
[0200] The present disclosure also provides a non-transitory computer-readable storage medium for operating the headphone device.
[0201] FIG. 8 illustrates an example non-transitory computer-readable medium 8000 for operating the headphone device according to one of a number of embodiments of the present disclosure.
[0202] As shown in FIG. 8, the computer-readable storage medium 8000 has instructions 8100 stored thereon and the instructions 8100 can cause a processor of a headphone device to implement any of methods or the functional procedures as described above with respect to FIG. 1 to FIG. 6, which will not be reproduced herein.
[0203] Various methods and features described above with respect to FIGS. 1 to 6 are also applicable to the computer-readable storage medium 8000, unless it is obviously inappropriate from the context.
[0204] The foregoing detailed description of the present disclosure has been presented for purposes of illustration and description. It is not intended to be exhaustive or to limit the technology to the precise form disclosed. Many modifications and variations are possible in light of the above teaching. The described embodiments were chosen to best explain the principles of the technology and its practical application to thereby enable others skilled in the art to best utilize the technology in various embodiments and with various modifications as are suited to the particular use contemplated. It is intended that the scope of the technology be defined by the claims appended hereto.
Claims
1.A headphone device, comprising:two earpieces configured to cover a user’s ears when the headphone device is mounted on the user’s head;a main strap configured to connect to the two earpieces;a first microphone detachably connected to one of the two earpieces;a microphone array of at least two microphones arranged on one of the two earpieces; anda control unit configured to switch an operating mode of the headphone device between a first mode and a second mode depending on a connection state of the first microphone,wherein in a case that the first microphone is plugged in, the headphone device is switched to the first mode and a first signal captured by the first microphone is used as a main signal, andwherein in a case that the first microphone is unplugged, the headphone device is switched to the second mode and the main signal is determined based on signals captured by the microphone array.2.The headphone device of claim 1, further comprising:a detection unit configured to detect whether the first microphone is plugged into the one of the two earpieces.3.The headphone device of claim 1, further comprising:a noise and interference determining unit configured to determine noise and interference based on the signals captured by the microphone array.4.The headphone device of claim 3, further comprising:a user speech extractor configured to extract speech signal of the user by eliminating the noise and interference from the main signal in each of the first mode or the second mode.5.The headphone device of claim 3, wherein in the first mode, the main signal is captured from the first microphone has a polar mode of a forward cardioid pattern towards a mouth of the user, and the noise and interference determining unit is configured to determine the noise and interference based on a backward cardioid pattern among the microphone array.6.The headphone device of claim 3, wherein in the second mode, a forward cardioid beamforming among the microphone array is carried out to determine the main signal, and the noise and interference determining unit is configured to determine the noise and interference based on a backward cardioid pattern among the microphone array.7.The headphone device of claim 5, wherein the forward cardioid pattern of the first microphone maximizes a capture of speech signal of the user and minimizes a capture of the noise and interference in an opposite direction of the mouth of the user in the first mode.8.The headphone device of claim 5, wherein the backward cardioid pattern among the microphone array minimizes a capture of speech signal of the user and maximizes a capture of the noise and interference in an opposite direction of the mouth of the user in the first or the second mode.9.The headphone device of claim 6, wherein the forward cardioid beamforming among the microphone array maximizes a capture of speech signal of the user and minimize a capture of the noise and interference in an opposite direction of the mouth of the user in the second mode.10.The headphone device of claim 1, wherein each microphone of the microphone array is an omnidirectional microphone.11.A method for operating a headphone device, the headphone device comprises two earpieces configured to cover a user’s ears when the headphone device is mounted on the user’s head; a main strap configured to connect to the two earpieces; a first microphone detachably connected to one of the two earpieces; and a microphone array of at least two microphones arranged on one of the two earpieces, the method comprising:switching, by using a control unit, an operating mode of the headphone device between a first mode and a second mode depending on a connection state of the first microphone,wherein in a case that the first microphone is plugged in, the headphone device is switched to the first mode and a first signal captured by the first microphone is used as a main signal, andwherein in a case that the first microphone is unplugged, the headphone device is switched to the second mode and the main signal is determined based on signals captured by the microphone array.12.The method of claim 11, further comprising:detecting, by a detection unit, whether the first microphone is plugged into the one of the two earpieces.13.The method of claim 11, further comprising:determining, by a noise and interference determining unit, noise and interference based on the signals captured by the microphone array.14.The method of claim 13, further comprising:extracting, by a user speech extractor, speech signal of the user by eliminating the noise and interference from the main signal in each of the first mode or the second mode.15.The method of claim 13, wherein in the first mode, the main signal is captured from the first microphone has a polar mode of a forward cardioid pattern towards a mouth of the user, and the noise and interference determining unit is configured to determine the noise and interference based on a backward cardioid pattern among the microphone array.16.The method of claim 13, wherein in the second mode, a forward cardioid beamforming among the microphone array is carried out to determine the main signal, and the noise and interference determining unit is configured to determine the noise and interference based on a backward cardioid pattern among the microphone array.17.The method of claim 15, wherein the forward cardioid pattern of the first microphone maximizes a capture of speech signal of the user and minimizes a capture of the noise and interference in an opposite direction of the mouth of the user in the first mode.18.The method of claim 15, wherein the backward cardioid pattern among the microphone array minimizes a capture of speech signal of the user and maximizes a capture of the noise and interference in an opposite direction of the mouth of the user in the first or the second mode.19.The method of claim 16, wherein the forward cardioid beamforming among the microphone array maximizes a capture of speech signal of the user and minimize a capture of the noise and interference in an opposite direction of the mouth of the user in the second mode.20.A non-transitory computer readable storage medium with instruction stored thereon, when executed, the instructions cause a processor of a headphone device to execute the method recited in any one of claims 11 to 19.
Citation Information
Patent Citations
Microphone arrangement
US20180091882A1
Multi-microphone Headset
US20220417639A1