System, method, and computer-readable medium for automatically switching an active microphone
By automatically switching active microphones between wireless earbud headphones, using voice input detection and sensor detection, the problem of microphones not being shared seamlessly in the prior art is solved, improving user experience and reducing resource consumption.
Patent Information
- Application Number
- CN202210589911.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2018-11-27
- Filing Date
- 2019-11-27
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2039-11-27
AI Technical Summary
Existing wireless earbud headphones can only serve as active microphones during phone calls, resulting in poor user experience and inability to share microphones seamlessly. There is a problem that microphones are not switched in time in manually specifying microphone mode.
Provides a system and method to automatically switch the active microphone between multiple devices, use voice input detection endpoints to determine the timing of microphone switching, including pauses, keywords or word end changes, etc., and combines sensors to detect user behavior to achieve a seamless microphone switching and notification mechanism.
It realizes seamless sharing of active microphones between wireless earbud headphones, improves user experience, reduces user interaction needs, saves bandwidth and resources, and ensures immediacy and accuracy of microphone switching.
Smart Images

Figure CN115150705B_ABST
Abstract
Description
[0001] Division Statement
[0002] This application is a divisional application of Chinese Patent Application No. 201980077007.4 with a filing date of November 27, 2019.
[0003] Cross - Reference to Related Applications
[0004] This application claims the benefit of U.S. Provisional Patent Application No. 62 / 771,747, filed on November 27, 2018, the disclosure of which is incorporated herein by reference. Technical Field
[0005] This disclosure relates to automatically switching an active microphone. Background Art
[0006] Due to the limitations of short - range wireless communication standards, wireless earbuds only support one of the earbuds in the earbud pair that acts as the active microphone during a phone call. This creates a user experience problem for users because they cannot easily share their earbuds with a friend during a phone call for a "three - way" call that should work seamlessly. One possible solution is that the user can manually specify which earbud to use the microphone on. The manually - specified earbud acts as the active microphone whether or not it is worn on the user's head. In this mode, the earbud with the active microphone can hold for some time in this case until the user discovers that voice data is not being captured by the other earbud being worn. Summary of the Invention
[0007] This disclosure provides a system, particularly an audio playback system suitable for automatically switching an active microphone back and forth between two or more devices. For example, in a case where the system includes a pair of earbuds, when each earbud is worn by a separate user, the system can switch the active microphone to the device worn by the user who is speaking at a given time. Although the device holds the active microphone, the other device can wait until a specific event that frees the microphone occurs, such as if the user wearing the device with the active microphone stops speaking. Such an event can trigger the active microphone to become free, at which time the other device can claim the active microphone. According to some examples, sidetone or comfort noise or another notification can be provided by one or more of the devices in the system to let the user know, for example, that he does not have the active microphone, the active microphone is free, the active microphone has been switched, etc.
[0008] One aspect of the present disclosure provides a system including a first device in wireless communication with a second device. The first device includes: a speaker configured to operate in an active mode in which it captures audio input for sending to a computing device and an inactive mode in which it does not capture audio input; and one or more processors. When the microphone of the first device is in the active mode and the microphone of the second device is in the inactive mode, the one or more processors of the first device are configured to receive voice input through the microphone of the first device, detect endpoints in the received voice input, and provide the microphone of the second device with an opportunity to switch to the active mode. Detecting endpoints may include, for example, detecting at least one of the following: a pause, a keyword, or an inflection. The first device and the second device may be audio playback devices, such as earbuds. However, in other examples, the first device and the second device may be other types of devices, which may be of the same type or different types. For example, the first device may be an in-ear speaker / microphone, while the second device is a smartwatch or a head-mounted display device.
[0009] Providing the microphone of the second device with an opportunity to switch to the active mode may include, for example, switching the microphone of the first device to the inactive mode. According to some examples, when the microphone of the first device is in the inactive mode, it listens for audio input without capturing audio for transmission. When the first device is in the inactive mode, the one or more processors of the first device may also determine whether to switch the microphone of the first device to the active mode based at least on the listening. In some examples, each of the first device and the second device may have a microphone, and only one of the microphones is in the active mode at a given time. The first device and / or the second device may then be configured to determine endpoints in the received voice input, for example when the user stops speaking. For example, the device with the active microphone can thus detect whether the user has reached an endpoint in the voice received by the active microphone, and in response to detecting such an endpoint, can automatically release the active microphone, thereby providing the microphone of the other device with an opportunity to switch to the active mode.
[0010] According to some examples, the one or more processors of the first device may also be configured to receive a notification when the microphone of the second device switches to the active mode. For example, the notification may be a sound emitted from the speaker of the second device, such as sidetone or comfort noise.
[0011] The one or more processors of the first device may also be configured to determine whether the microphone of the first device is in the active mode, detect whether the user of the first device is providing audio input, and provide a notification to the user of the first device when the microphone of the first device is in the inactive mode and audio input is detected.
[0012] According to some examples, one or more processors of the second device may be configured to determine an opportunity for the second device to switch to an active mode, and thus determine, for example, that an active microphone has become available. The one or more processors may also be configured to then fix the active microphone by switching the microphone of the second device to an inactive mode.
[0013] Another aspect of the present disclosure provides a method including: receiving voice input via a first device microphone of a first wireless device, where the first wireless device operates in an active microphone mode and communicates with a second wireless device operating in an inactive microphone mode; detecting, by one or more processors of the first device, endpoints in the received voice input; and providing, by one or more processors of the first device, an opportunity for the microphone of the second device to switch to an active mode. Providing the opportunity for the second device microphone to switch to an active mode may include switching the microphone of the first device to an inactive mode.
[0014] According to some examples, the method may further include: determining whether the microphone of the first device is in an active mode; detecting whether a user of the first device is providing audio input; and providing a notification via the first device when the microphone of the first device is in an inactive mode and audio input is detected.
[0015] Yet another aspect of the present disclosure provides a computer-readable medium storing instructions executable by one or more processors of a first device wirelessly communicating with a second device to perform a method including the steps of: receiving voice input via a first device microphone of the first device, where the first device operates in an active microphone mode and communicates with a second wireless device operating in an inactive microphone mode; detecting endpoints in the received voice input; and providing an opportunity for the microphone of the second device to switch to an active mode. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] Figure 1 is a schematic diagram illustrating an example use of an auxiliary device according to various aspects of the present disclosure.
[0017] Figure 2 is a schematic diagram illustrating another example use of an auxiliary device according to various aspects of the present disclosure.
[0018] Figure 3 is a functional block diagram of an example system according to various aspects of the present disclosure.
[0019] Figure 4 is a table for indicating various possible operation modes of an auxiliary device according to various aspects of the present disclosure.
[0020] Figure 5It is a flowchart illustrating an example method performed by an audio device having an active microphone according to various aspects of the present disclosure. Detailed Description
[0021] Overview:
[0022] The present disclosure provides for seamless sharing of an active microphone source among multiple user devices, such as earbuds worn by two different people, without user input. Each user device can be configured to determine which device may need the active microphone. For example, the system can detect endpoints in a user's voice input to detect when another user may be inputting a response voice. Examples of such endpoints can be pauses, keywords, inflections, or another factor. The device that may need the active microphone can be switched to the active microphone. In some examples, a particular device can request or otherwise claim the active microphone. Such a device can continue to hold the active microphone until its user temporarily stops providing audio input.
[0023] In some examples, each device can detect whether the user wearing the device is speaking. If the user is speaking but the device does not have the active microphone, a notification can be provided. For example, sidetone, comfort sound, or another audible notification can be provided. In other examples, the notification can be tactile, such as vibration of the device.
[0024] In some examples, it may be beneficial to indicate to the users of the system whether the device they are using has the active microphone. Such an indication can be provided, for example, by sidetone from the active microphone through the speaker of the device in the inactive mode. In this regard, when the user hears the sidetone, they will know that they do not have the active microphone. As another example, comfort noise can be provided when the active microphone is idle and either device can claim it. In yet another example, the volume in the active and / or inactive device can be adjusted in a way that indicates whether the device is active or inactive. Any of various other possible indications can be implemented.
[0025] One advantage of this automatic switching of the active microphone compared to explicit manual switching is that it is seamless and does not require any user interaction. It provides an "expected" behavior for the devices without any training, thus providing an improved user experience. Also, the solution does not significantly consume bandwidth and other resources.
[0026] Example System
[0027] Figure 1Illustrated is an example of a user 101 wearing a first audio playback device 180 and a second user 102 wearing a second audio playback device 190. In this example, the first device 180 and the second device 190 are earbuds. However, in other examples, the first device and the second device can be other types of devices, which can be of the same type or different types. For example, the first device can be an in-ear speaker / microphone, while the second device is a smartwatch or a head-mounted display device.
[0028] As Figure 1 shown, the first device 180 is coupled to the second device 190 via a connection 185. The connection 185 can include a standard short-range wireless coupling, such as a Bluetooth connection.
[0029] Each of the first device 180 and the second device 190 has a microphone, and only one of the microphones is "active" at a given time. The active microphone can capture the user's voice and send it to a computing device 170, which can be, for example, a mobile phone or other mobile computing device. In Figure 1 the example, the second device 180 worn by the second user 102 maintains the active microphone and thus captures the voice input of the second user 102.
[0030] For example, the inactive microphone on the first device 180 can capture the user's voice to determine whether an attempt is made to fix the active microphone, or to notify the user that no voice is captured to be sent to the computing device 170.
[0031] Each of the first device 180 and the second device 190 can be configured to determine when the user starts or stops speaking. For example, the device with the active microphone can determine whether its user has reached an endpoint in the speech received by the active microphone. The endpoint can be based on, for example, inflection, speech rate, keywords, pauses, or other characteristics of the audio input. According to some examples, other information can also be used in the determination, such as speech recognition, movement of the device, change in interference level, etc. For example, the device can determine whether audio is received from the user wearing the active microphone or from another user via the active microphone based on speech recognition, detected movement consistent with the user's jaw movement, volume of the received audio, etc.
[0032] The endpoint can serve as an indication that the user of the inactive device will likely provide audio input next. Similarly, a device with an inactive microphone can listen for audio without capturing it for transmission. Thus, a device with an inactive microphone can determine whether its user is speaking. If so, it can attempt to fix the active microphone and / or notify its user that its microphone is inactive. A device with an inactive microphone can make a similar determination based on accelerometer movement or other sensor information.
[0033] When the user of a device with an active microphone stops speaking, the active microphone can be released. For example, both the first device 180 and the second device 190 can enter a mode where the active microphone is available. According to some examples, releasing the active microphone can include indicating to the computing device 170 that the microphone is entering the inactive mode. For example, the device that releases the active microphone can send a signal for indicating its release of the active microphone. In this mode, either device can fix the active microphone. For example, a device that may need the microphone can fix it. A device that may need the microphone can be, for example, a device that moves in a specific way, a previously inactive device, a device that the user starts speaking to, or any combination of these factors or other factors.
[0034] According to some examples, a machine learning algorithm can be implemented to determine which device should switch to the active mode. For example, the machine learning algorithm can use a training data set that includes voice input parameters such as speaking time, pause time, keywords (such as proper nouns or pronouns), volume, etc. Other parameters in the training data can include movement measured by an accelerometer or other device, signal strength, battery level, interference, or any other information. Based on one or more of such parameters, the system can determine which device should switch to the active mode to capture the user's voice input.
[0035] According to some examples, one or more sensors on the device can be used to determine whether the user of a device with an active microphone has stopped speaking, or whether the user of a device with an inactive microphone has started speaking. For example, in addition to the microphone, the device can include a capacitance sensor, a thermal sensor, or other sensors to detect whether the electronic device 180 is in contact with the skin, thereby indicating whether the electronic device 180 is being worn. In other examples, the sensor can include an accelerometer for detecting user movement consistent with speaking. For example, when the user wearing the electronic device 180 starts speaking, his mouth, jaw, and other parts of his body move. Such movement can indicate that he is speaking.
[0036] Figure 2Another example is illustrated, where the first device 180 has been switched to operate as an active microphone, and the second device 190 has been switched to operate as an inactive microphone. Thus, the first device 180 can capture the voice of the first user 101 and send it to the computing device 170. The second device 190 can wait for the active microphone to become available, such as when the first user 101 stops speaking. If the second user 102 starts speaking before the active microphone becomes available, a notification can be provided by the second device 190. For example, the second device can play a sound, such as a bell, it can play sidetone or comfort noise, it can vibrate, illuminate a light-emitting diode, or provide some other type of notification.
[0037] Figure 3 An example block diagram of the first (auxiliary) device 180 and the (second) auxiliary device 190 is provided. The auxiliary devices 180, 190 can be any of a variety of types of devices, such as earbuds, head-mounted devices, smartwatches, etc. Each device includes one or more processors 391, 381, memories 392, 382, and other components typically present in audio playback devices and auxiliary devices. Although multiple components are shown, it should be understood that such components are merely non-limiting examples and may additionally or alternatively include other components.
[0038] One or more processors 391, 381 can be any conventional processor, such as a commercially available microprocessor. Alternatively, one or more processors can be a specialized device, such as an application-specific integrated circuit (ASIC) or other hardware-based processor. Although Figure 3 Functionally, the processor, memory, and other elements of the auxiliary devices 180, 190 are illustrated within the same respective boxes, but those of ordinary skill in the art will understand that the processor or memory can actually include multiple processors or memories, which may or may not be stored within the same physical housing. Similarly, the memory can be a hard disk drive or other storage medium, which is located in a housing different from the auxiliary devices 180, 190. Thus, a reference to a processor or computing device will be understood to include a reference to a collection of processors or computing devices or memories, which may or may not operate in parallel.
[0039] The memory 382 can store information accessible by the processor 381, including instructions 383 and data 384 that can be executed by the processor 381. The memory 382 can be a type of memory operable to store information accessible by the processor 381, including non-transitory computer-readable media, or other media for storing data that can be read by an electronic device, such as a hard disk drive, a memory card, read-only memory (“ROM”), random access memory (“RAM”), optical discs, and other writable and read-only memories. The subject matter disclosed herein can include different combinations of the foregoing, whereby different portions of the instructions 383 and data 384 are stored on different types of media.
[0040] According to the instructions 384, data 384 can be retrieved, stored, or modified by the processor 381. For example, although the present disclosure is not limited to a particular data structure, the data 384 can be stored in a computing register, stored in a relational database as a table with multiple different fields and records, an XML document, or a flat file. The data 384 can also be formatted in a computer-readable format, such as, but not limited to, binary values, ASCII, or Unicode. Further by way of example only, the data 384 can be stored as a bitmap stored in compressed or uncompressed form including pixels, or stored in various image formats (e.g., JPEG), vector-based formats (e.g., SVG), or computer instructions for rendering graphics. Moreover, the data 384 can include information sufficient to identify related information, such as numbers, descriptive text, proprietary codes, pointers, references to data stored in other memories (including other network locations), or information used by a function for computing related data.
[0041] The instructions 383 can be executed to improve the user experience during a three-way call, where one user wears a first auxiliary device 180 and another user wears a second auxiliary device 190. For example, the instructions 383 can provide endpoints in the voice of the user waiting for the active device, determine that the active microphone has become available, and fix the active microphone.
[0042] When the first auxiliary device 180 is executing instruction 383, the second auxiliary device 190 may also be executing instruction 393 stored in the memory 392 together with data 394. For example, similar to the auxiliary device 180, the auxiliary device 190 may also include a memory 392 for storing data 394 and instructions 393 executable by one or more processors 391. The memory 392 may be any of various types, and the data 394 may be any of various formats, similar to the memory 382 and data 384 of the auxiliary device 180. When the auxiliary device 180 is receiving and encoding speech from a user wearing the auxiliary device 180, the second auxiliary device 190 may also listen for and receive speech through a microphone 398. The instruction 393 may provide for holding an active microphone, capturing and transmitting the voice of the user of the second device 190, detecting endpoints in the speech of the second user, and automatically releasing the active microphone when the endpoints are detected. Thus, the first device 180 and the second device 190 may be configured to switch back and forth between operating as an inactive microphone device and an active microphone device. Thus, although Figure 3 the example of
[0043] illustrates a particular set of operations in each instruction set, it should be understood that either device may be capable of executing either instruction set as well as additional or other instructions. By way of example only, instructions 383, 393 may be executed to determine whether the first device 180 and the second device 190 are worn by the same user, determine which user is providing audio input, etc.
[0044] It should be understood that the auxiliary device 180 and the mobile device 190 may respectively include other components not shown, such as a charging input for a battery, signal processing components, etc. Such components may also be used to execute instructions 383, 393.
[0045] Figure 4 A diagram is provided that illustrates some example operating modes of the first auxiliary device 180 and the second auxiliary device 190. In a first example mode, the first device holds the active microphone, and the second device waits for the active microphone to be released. For example, the second device may wait for an endpoint in the speech of the first user of the first device.
[0046] In a second example mode, the active microphone is available. In this mode, the active microphone has been released from its previous device but has not yet been claimed by another device. In fact, the device typically only operates in this mode for a very short period of time, such as a fraction of a second or a millisecond. In this regard, when neither device is capturing voice input, there will be no unpleasant latency.
[0047] In a third example mode, the second device has claimed the active microphone and the first device is waiting for an endpoint. In some examples, if the user of the inactive device provides voice input in a particular manner (such as above a threshold decibel level or above a particular rate), for example, the active microphone can switch devices rather than wait for an endpoint.
[0048] Example Method
[0049] In addition to the operations illustrated above and in the figures, various operations will now be described. It should be understood that the following operations need not be performed in the exact order described below. Rather, the various steps can be disposed of in a different order or simultaneously, and steps can also be added or omitted.
[0050] Figure 5 is a flowchart illustrating an example method performed by an audio system such as a pair of earbuds, where one device in the system is the "active" device and holds the active microphone, while one or more other devices in the system operate as "inactive" devices such that their microphones do not capture audio input.
[0051] In block 410, the inactive device waits for the active microphone to become available. Meanwhile, in block 510, the active device captures the user's voice input and in block 520 sends the voice to a computing device.
[0052] In block 530, the active device determines whether an endpoint has been reached, such as whether the user of the active device has stopped speaking. The endpoint can serve as an indication that the user of the inactive device will likely provide audio input next. The endpoint can be based on, for example, inflection at the end of the audio input, speaking rate, keywords, pauses, or other characteristics. According to some examples, other information such as voice recognition, movement of the device, change in interference level, etc. can also be used in the determination. If the endpoint has not been reached, the device continues to capture input in block 510. However, if the endpoint is reached, in block 540, the active device can release the active microphone.
[0053] In block 420, the inactive device determines whether an active microphone is available. The inactive device will continue to wait until it is available. If in block 425 the inactive device detects that its user is speaking, the inactive device may provide a notification that the device does not have an active microphone. However, if the active microphone is available, the inactive device secures the active microphone (block 430), thus switching modes. Accordingly, it will capture the user's voice (block 440) and send it to the computing device (block 450). Meanwhile, the active device, which has also been switched to the inactive mode and is now operating as an inactive device, waits for the active microphone to become available (block 550).
[0054] Although the above examples mainly describe two devices sharing an active microphone, in other examples, three or more devices may share an active microphone. For example, two inactive devices will wait for the active microphone to become available. When it becomes available, it may be secured by one of the inactive devices, such as whichever device first detects an audio input from its user or movement of the user, such as movement consistent with the user speaking of the user's jaw or mouth.
[0055] Unless otherwise specified, the foregoing alternative examples are not mutually exclusive, but may be implemented in various combinations to achieve unique advantages. Since these and other variations and combinations of the features discussed above may be utilized without departing from the subject matter defined by the claims, the foregoing description of the embodiments should be by way of illustration rather than by way of limitation of the subject matter defined by the claims. Additionally, the provision of the examples described herein and the clauses phrased as "such as", "including", etc. should not be construed as limiting the subject matter of the claims to the specific examples; rather, the examples are not intended to illustrate only one of many possible embodiments. Further, the same reference numerals in different figures may identify the same or similar elements.
Claims
1. A method of using a first microphone of a first wireless device to communicate with a second microphone of a second wireless device, comprising: Operating the first wireless device in an inactive microphone mode, in which the first microphone is inactive; Identifying, by the first wireless device, that the second microphone of the second wireless device has become inactive; Detecting, by the first wireless device, a signal that a first user wearing the first wireless device is starting to speak; And Switching, by the first wireless device, to an active microphone mode based on the signal, Wherein a microphone in the active microphone mode captures audio input, and a microphone in the inactive microphone mode does not capture audio input.
2. The method according to claim 1, wherein, The signal includes one of a voice input or a jaw movement.
3. The method according to claim 1, further comprising receiving a first notification that the active microphone mode is available.
4. The method according to claim 3, wherein, The first notification includes one of a sound, a light, or a vibration.
5. The method according to claim 1, further comprising: Receiving, by the first wireless device, a voice input before switching the first wireless device to the active microphone mode; Providing, to the first wireless device, a second notification that the first wireless device does not have the active microphone mode.
6. The method according to claim 1, further comprising, when the first microphone is in the active microphone mode: Receiving a voice input through the first microphone; and Transmitting the received voice input to a computing device.
7. The method according to claim 6, further comprising, when the first microphone is in the active microphone mode, detecting a second endpoint in the received voice input.
8. The method according to claim 7, further comprising providing an opportunity for the second microphone of the second wireless device to switch to the active microphone mode based on the detected second endpoint.
9. The method according to claim 7, wherein Detecting the second endpoint includes detecting at least one of a pause, a keyword, and an end-of-word change.
10. A system, comprising: A first wireless device wirelessly communicating with a second wireless device, the first wireless device comprising: A speaker; A first microphone; One or more processors; Wherein, when the first microphone of the first wireless device is in an inactive microphone mode and the second microphone of the second wireless device is in an active microphone mode, the one or more processors of the first wireless device are configured to perform operations, the operations including: Identifying that the second microphone of the second wireless device has become inactive; Detecting a signal that a first user wearing the first wireless device is starting to speak; and Switching to an active microphone mode based on the signal, Wherein a microphone in the active microphone mode captures audio input, And a microphone in the inactive microphone mode does not capture audio input.
11. The system according to claim 10, wherein The signal includes one of a voice input or a jaw movement.
12. The system according to claim 10, wherein The operations further include: receiving a first notification that the active microphone mode is available.
13. The system according to claim 12, wherein, The first notification includes one of a sound, a light, or a vibration.
14. The system according to claim 10, wherein The operations further include: Before switching the first wireless device to the active microphone mode, receive voice input by the first wireless device; Provide a second notification to the first wireless device that the first wireless device does not have the active microphone mode.
15. The system according to claim 10, wherein, The operation further includes, when the first microphone is in the active microphone mode: Receive voice input through the first microphone; and Transmit the received voice input to a computing device.
16. The system according to claim 15, wherein, The operation further includes, when the first microphone is in the active microphone mode, detecting a second endpoint in the received voice input.
17. The system according to claim 16, wherein The operation further includes: Based on the detected second endpoint, provide an opportunity for the second wireless device microphone to switch to the active mode.
18. The system according to claim 16, wherein Detecting the second endpoint includes detecting at least one of a pause, a keyword, and an inflection.
19. A non-transitory computer-readable medium storing instructions that can be executed by one or more processors of a first wireless device wirelessly communicating with a second wireless device to perform a method, the method including: Operate in an inactive microphone mode, in which the first microphone of the first wireless device is inactive; Identify that the second microphone of the second wireless device has become inactive; Detect a signal that a first user wearing the first wireless device is starting to speak; and Based on the signal, switch to the active microphone mode, wherein a microphone in the active microphone mode captures audio input, and a microphone in the inactive microphone mode does not capture audio input.
20. The non-transitory computer-readable medium according to claim 19, wherein, The signal includes one of voice input or jaw movement.
Citation Information
Patent Citations
Microphone share method and device, computer equipment and storage medium
CN107957908A
Speech enhancement using multiple microphones on multiple devices
US20090238377A1