Automatic acoustic switching
Through acoustic analysis, near- and far-field voice signals are processed, and the communication mode of wearable audio output devices is automatically switched using microphone arrays and direct RF links, which solves the problem of inconvenience in switching when the acoustic environment changes, and improves the automation and user experience of the device.
Patent Information
- Application Number
- CN202210183695.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2021-02-26
- Filing Date
- 2022-02-25
- Publication Date
- 2025-09-02
- Estimated Expiration
- 2042-02-25
AI Technical Summary
Existing wearable audio output devices require manual operation when switching between different communication modes, resulting in poor user experience, especially when the acoustic environment changes.
The near- and far-field voice signals are processed through acoustic analysis, the acoustic environment parameters are estimated and the communication mode is automatically switched, and the signal is enhanced by microphone arrays and direct RF links to achieve seamless conversion.
It realizes automatic adjustment of communication mode in different acoustic environments, improves user experience, ensures clear transmission of voice signals and signal-to-noise ratio, and reduces manual intervention.
Smart Images

Figure CN114979880B_ABST
Abstract
Description
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS
[0002] This application claims the benefit of U.S. Provisional Application No. 63 / 154,651, filed on February 26, 2021, the disclosure of which is incorporated herein by reference in its entirety. Technical Field
[0003] The present disclosure relates to the field of audio communications, including digital signal processing methods designed to automatically identify and switch between various electroacoustic communication modes that adapt to changing acoustic environments. Other aspects are also described. Background Art
[0004] Audio output devices (including wearable audio output devices such as headphones, earbuds, and earphones) are widely used to provide audio output to users using various electroacoustic communication modes. Wearable audio output devices can be paired with a phone in phone mode or operate in a transparent mode that allows the user to hear ambient sound through the audio output device, thereby facilitating communication with nearby speakers without removing the audio output device. Summary of the Invention
[0005] Aspects of methods and systems for automatically switching between communication modes of a wearable audio output device based solely on acoustic analysis are disclosed. When worn by a user using the audio output device to communicate, the audio output device can operate in one of three electroacoustic modes. In transparent mode, the audio output device can transmit voice signals of nearby users. In peer-to-peer mode, the audio output device can establish a direct, low-latency radio frequency (RF) link to another audio output device within the communication range of the RF link. In phone mode, the audio output device can communicate with another audio output device using a networked phone. The disclosed methods and systems perform acoustic analysis on the near-field voice signals of a local wearer of the audio output device and the far-field voice signals of a remote caller to determine the optimal mode for the audio output device, and seamlessly switch between modes as the acoustic environment between the local wearer of the audio output device and the remote caller changes.
[0006] In one aspect, the method can process near-field and far-field speech signals captured by one or more microphones of an audio output device to estimate parameters of an acoustic environment. In one aspect, the audio output device of the local wearer and the audio output device of the remote caller can estimate the acoustic parameters of the environment back and forth. The two audio output devices can each estimate the acoustic parameters and their rate of change based on their respective near-field and far-field speech signals. The two audio output devices can exchange the estimated acoustic parameters, for example, via a direct RF link in peer-to-peer mode to increase the confidence level of the estimated acoustic parameters. In practice, the two audio output devices can act as a distributed, non-phase-locked microphone array to perform back-and-forth estimation of acoustic parameters to determine an electroacoustic pattern for communication between the two wearers of the audio output devices. In one aspect, if the other audio output device has no processing power, has processing limitations, or wants to save power, only one audio output device can estimate the acoustic parameters and their rate of change.
[0007] The method may process the estimated acoustic parameters to determine whether it is possible to allow the wearers of the audio output devices to communicate in a transparent mode, such as when the wearers are within audible range of each other for a face-to-face conversation. The method may further process the estimated acoustic parameters to generate spatialized metadata for the remote talker. In one aspect, when the far-field voice signal is too weak, such as when the distance between the two wearers exceeds the audible communication range, the local wearer's audio output device may establish a direct low-latency RF link with the remote talker's audio output device in a peer-to-peer mode to electromagnetically receive the far-field voice signal. The method may use the spatialized metadata to re-spatialize the far-field voice signal received via the direct RF link to spatially simulate the level and perceived direction of arrival of the remote talker. The spatialized far-field voice signal from the direct RF link may be used to enhance the far-field voice signal acoustically received by the microphone. In one aspect, the method may add the far-field voice signal acoustically received by the microphone with the spatialized far-field voice signal from the RF link to improve the signal-to-noise ratio (SNR) of the far-field voice signal. In one aspect, the local wearer's audio output device may output enhanced far-field voice signals to the user via the audio output device's speakers in peer-to-peer mode.
[0008] The method can estimate the power spectrum of an acoustic far-field speech signal, a spatialized far-field speech signal, or an enhanced far-field speech signal, such as by generating a running power spectral density (PSD) estimate of the far-field speech signal in transparent mode or peer-to-peer mode. In one aspect, the method can process the estimated acoustic parameters to determine that the distance between the two talkers exceeds the communication range of a direct RF link. The local wearer's audio output device can switch from peer-to-peer mode to phone mode to receive the far-field speech signal from the remote talker's audio output device via a networked phone. The method can balance the far-field speech signal received via phone mode with the running power spectral density estimate to smooth the transition from peer-to-peer mode to phone mode. In one aspect, the method can sum the balanced far-field speech signal in phone mode with the spatialized far-field speech signal or enhanced far-field speech signal in peer-to-peer mode. In one aspect, the method can estimate the power spectrum of the acoustic near-field signal by generating a running PSD estimate of the near-field speech signal in transparent mode or peer-to-peer mode. The method can process the estimated acoustic parameters, the PSD estimate of the far-field speech signal, and the PSD estimate of the near-field speech signal to estimate the distance between the two callers and switch between transparent mode, peer-to-peer mode, and telephony mode. In one aspect, if one of the audio output devices does not have direct link RF capability for peer-to-peer mode, the method can directly switch between transparent mode and telephony mode.
[0009] In one aspect, a method for communicating between a local caller wearing a local headset and a remote caller wearing a remote headset is disclosed. The method processes a near-field voice signal of the local caller and a far-field voice signal of the remote caller received by the local headset to estimate acoustic parameters. The method also processes the estimated acoustic parameters to determine a communication mode between the local headset and the remote headset. The communication mode includes an acoustic transparency mode, a peer-to-peer RF mode, or a telephony mode. If it is determined that the communication mode is peer-to-peer mode, the method processes the far-field voice signal received via the peer-to-peer mode to generate a spatialized voice signal. If it is determined that the communication mode is telephony mode, the method processes the far-field voice signal received via the telephony mode to generate a telephony voice signal. The method outputs the far-field voice signal received via the acoustic transparency mode, the spatialized voice signal in the peer-to-peer mode, or the telephony voice signal in the telephony mode to a speaker of the local headset.
[0010] In one aspect, a method for communicating between a local talker wearing a local headset and a remote talker wearing a remote headset is disclosed. The method processes a near-field speech signal of the local talker and a far-field speech signal of the remote talker to estimate acoustic parameters. The far-field speech signal is captured as an acoustic signal using a microphone of the local headset. The method processes the estimated acoustic parameters to determine whether to enhance the acoustic signal using an RF transmission received by the local headset from the remote headset. The RF transmission is used to electromagnetically carry the far-field speech signal. If enhancement is determined, the method processes the acoustic signal and the far-field speech signal received via the RF transmission to generate an enhanced acoustic signal. The method outputs the enhanced acoustic signal, or outputs the acoustic signal to a speaker of the local headset if it is not enhanced.
[0011] The above summary does not include an exhaustive list of all aspects of the present invention. It is contemplated that the present invention includes all systems and methods that can be practiced from all suitable combinations of the various aspects summarized above and disclosed in the following detailed description and particularly pointed out in the claims filed with this patent application. Such combinations have particular advantages not specifically recited in the above summary. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] Various aspects of the present disclosure are described by way of example and not by way of limitation in the drawings, in which similar reference numerals indicate similar elements. It should be noted that references to "one" or "an" aspect in the present disclosure are not necessarily to the same aspect, and are intended to refer to at least one aspect. In addition, for the sake of brevity and to reduce the total number of drawings, a given drawing may be used to illustrate features of more than one aspect of the present disclosure, and not all elements in the drawing may be required for a given aspect.
[0013] Figure 1 Depicted are two wearers of an audio output device communicating with each other using a transparent mode, a peer-to-peer mode, or a telephony mode of the audio output device according to one aspect of the present disclosure.
[0014] Figure 2 Depicted is a wearable audio output device and perceived ambient sound according to one aspect of the present disclosure.
[0015] Figure 3 A functional block diagram depicts a system for processing ambient sound, including speech signals acoustically captured by a microphone array of a local wearable audio output device and speech signals electromagnetically received from a remote wearable audio output device, to determine a communication mode between audio output devices based solely on acoustic analysis and to switch between communication modes, in accordance with one aspect of the present disclosure.
[0016] Figure 4Depicted is a functional block diagram of a feature extractor module that processes near-field speech and far-field speech signals to estimate parameters of an acoustic environment for determining a communication mode of a wearable audio output device, according to one aspect of the present disclosure.
[0017] Figure 5 Depicted is a functional block diagram of a classifier and parameter estimator module that processes estimated parameters to determine a communication mode and spatialization metadata for respatializing a far-end voice signal received in peer-to-peer mode, according to one aspect of the present disclosure.
[0018] Figure 6 Depicted is a functional block diagram of a spatial filter module that uses spatialization metadata to respatialize far-end voice signals received in peer-to-peer mode and generates power spectrum metadata for equalizing far-end voice signals received in telephony mode, according to one aspect of the present disclosure.
[0019] Figure 7 is a flow chart of a method for determining a communication mode based solely on acoustic analysis and switching between communication modes of a wearable audio output device according to one aspect of the present disclosure.
[0020] Figure 8 is a flow chart of a method for enhancing an acoustic signal of far-field speech captured by a microphone of a wearable audio output device with a far-field speech signal carried over an RF transmission based solely on acoustic analysis, according to one aspect of the present disclosure. DETAILED DESCRIPTION
[0021] Wearable audio output devices can operate in a transparent mode that allows a user to hear ambient sound without requiring the user to remove the audio output device. In some cases, ambient sound, including the speech of nearby speakers perceived by the user, can be attenuated by the physical obstruction presented by the audio output device. In one mode, the audio output device can deliver the attenuated ambient sound to the user's ear, or alternatively, can amplify the ambient sound by capturing the ambient sound using a microphone and playing back the captured acoustic signal.
[0022] In another mode, when the audio output device is paired with a phone, the audio output device can actively cancel ambient sound to allow the user to make a traditional phone call. The two communication modes are conventionally handled separately. When a user wishes to switch between modes, they may have to do so manually. For example, if a user wishes to interrupt a conversation with a nearby speaker in Transparency mode to make a phone call, the user may have to deactivate Transparency mode to make the phone call. After the phone call, the user may have to reactivate Transparency mode to continue the conversation with the nearby speaker.
[0023] In another scenario, a user may wish to continue a conversation with a nearby speaker even when the user or the nearby speaker is out of audible range for the conversation. When the voice signal from the nearby speaker becomes too weak to be heard due to increased distance, the user may have to manually turn off transparency mode to call the speaker, potentially interrupting the conversation. Therefore, requiring the user to manually switch between operating modes of a wearable audio output device can be inconvenient and may diminish the user's overall audio experience.
[0024] It is desirable to automatically switch between communication modes of wearable audio output devices based solely on acoustic analysis, without requiring manual user intervention or commands. For example, when two wearers of headphones, earbuds, earphones, etc. are in close proximity for a face-to-face conversation, each audio output device can operate in a transparent mode to acoustically capture the speech signal from the other speaker using a microphone array that preserves the spatial characteristics of the speech signal. Each audio output device can process the acoustic signal captured by the microphone array to extract acoustic parameters and their rates of change to determine whether it is feasible to continue the conversation in transparent mode as the distance between the two speakers or the acoustic environment changes. In one aspect, the acoustic parameters may include the level difference between the far-field speech of the remote talker and the near-field speech of the local talker, the direct-to-reverberation ratio of the far-field speech signal, a measure of the energy distribution of the far-field speech signal, the Lombard effect or level change of the near-field speech signal, the direction of arrival of the far-field speech signal, a measure of the intelligibility of the far-field speech signal, etc.
[0025] The audio output device may process the extracted acoustic parameters to determine that continuing the conversation using transparent mode may no longer be feasible due to an increased distance between the talkers or due to a noise source. The audio output device may enhance the acoustic signal in transparent mode by electromagnetically receiving the far-field voice signal through a direct low-latency RF link by switching the two devices to operate in peer-to-peer mode. The audio output device may estimate a desired level and direction for re-spatializing the far-field voice signal received through the RF link based on the extracted acoustic parameters. The audio output device may re-spatialize the far-field voice signal received through the RF link so that it is consistent with the spatial position of the remote speaker so that the acoustic signal can be enhanced in a seamless manner. In one aspect, the audio device may add the far-field voice signal acoustically received through the microphone to the spatialized far-field voice signal received through the RF link to improve the SNR of the far-field voice in the enhanced signal.
[0026] In one aspect, when the audio output device determines that the RF link is exceeding its operating range, the audio output device may switch to operating in telephony mode with other audio output devices. The audio output device may equalize the far-field voice signal carried by the telephony signal to have a power spectrum similar to that of the spatialized far-field voice signal. In one aspect, the audio output device may estimate running statistics of the power spectral density (PSD) of the spatialized far-field voice signal in transparent mode or peer-to-peer mode. The audio output device may use the running PSD estimate to equalize the far-field voice signal carried by the telephony signal to smooth the transition to telephony mode. The original acoustic signal in transparent mode, the enhanced far-field voice signal in peer-to-peer mode, or the equalized far-field voice signal in telephony mode may be output to the user via a speaker of the audio output device. In one aspect, the audio output device may estimate the PSD of the near-field voice signal in transparent mode or peer-to-peer mode. The method may compare the PSD estimate of the far-field voice signal with the PSD estimate of the near-field voice signal, or their relative rate of change, to estimate the distance between the two callers or changes in the acoustic environment. The audio output device can use this information to determine when to switch between transparent mode, peer-to-peer mode, and telephony mode.
[0027] The following description shows many specific details. However, it should be understood that aspects of the present disclosure can be practiced here without these specific details. In other cases, well-known circuits, structures, and technologies are not shown in detail to avoid obscuring the understanding of this description.
[0028] The terms used herein are only for describing specific aspects and are not intended to limit the present invention. Spatially relative terms, such as "under...", "below...", "below...", "above...", "on...", etc., may be used herein for convenience of description to describe the relationship between an element or feature and one or more other elements or one or more feature parts, as shown in the accompanying drawings. It should be understood that spatially relative terms are intended to cover different orientations of elements or features during use or operation other than the orientation shown in the drawings. For example, if a device containing multiple elements in the figure is flipped, the elements described as "below" or "below" other elements or features can then be oriented to be "above" other elements or features. Therefore, the exemplary term "below..." can cover both orientations of "above..." and "below...". The device can be oriented in other ways (e.g., rotated 90 degrees or in other orientations), and the spatially relative descriptors used in this article are interpreted accordingly.
[0029] As used herein, the singular forms "a," "an," and "the" are intended to include the plural forms as well, unless the context indicates otherwise. It should be further understood that the terms "include" and "comprise" specify the presence of stated features, steps, operations, elements, or components, but do not preclude the presence or addition of one or more other features, steps, operations, elements, components, or groups thereof.
[0030] As used herein, the terms "or" and "and / or" should be interpreted as inclusive or meaning any one or any combination. Thus, "A, B, or C" or "A, B, and / or C" means "any of the following: A; B; C; A and B; A and C; B and C; A, B, and C." An exception to this definition occurs only when a combination of elements, functions, steps, or actions are inherently mutually exclusive in some way.
[0031] Figure 1 Two wearers of audio output devices communicating with each other using the audio output devices in transparent mode, peer-to-peer mode, or telephone mode according to one aspect of the present disclosure are depicted. To simplify the description, the wearer of the audio output device receiving the voice signal from the other talker is referred to as the local talker. The audio output device worn by the local talker is referred to as the local audio output device. The signal representing the local talker's voice is referred to as the near-field voice signal. Conversely, the other talker is referred to as the remote talker, the audio output device worn by the remote talker is referred to as the remote audio output device, and the signal representing the remote talker's voice is referred to as the far-field voice signal.
[0032] In a sub-mode of transparency mode, the local audio output device can output one or more audio components, such as ambient sound, including the far-field voice signal of the remote talker. The local audio output device can capture the far-field voice signal using one or more microphones facing the surrounding acoustic environment. The local audio output device can amplify and play the captured far-field voice signal to the local talker via a speaker of the local audio output device. In this sub-mode of transparency mode involving active sound reproduction, the local talker can hear a greater amount of ambient sound from the surrounding physical environment than would otherwise be audible through passive attenuation of the ambient sound due to the physical obstruction of the local audio output device in the local talker's ear. In one aspect, if the two talkers are sufficiently close, the local audio output device can turn off active sound reproduction, such that any amount of ambient sound perceived by the local talker is due to passive attenuation by the local audio output device. This passive acoustic leakage sub-mode of transparency mode can be referred to as a pass-through sub-mode or an "off" sub-mode. Aspects of the present disclosure relating to transparency mode can be applied to the active sound reproduction sub-mode or pass-through sub-mode of transparency mode, or any other mode that allows the local talker to hear the natural world through the local audio output device. Similarly, references to an acoustic signal captured in transparent mode may refer to either the amplified signal or the passive leakage signal captured by the microphone without active amplification.
[0033] Figure 2 Depicts a wearable audio output device and perceived ambient sound according to one aspect of the present disclosure. Wearable audio output device 301 includes earbud 303, stem 305, and eartip 314. A user wears wearable audio output device 301 such that earbud 303 and eartip 314 are in the user's left ear. Eartip 314 extends at least partially into the user's ear canal. In one use case, when earbud 303 and eartip 314 are inserted into the user's ear, a seal may be formed between eartip 314 and the user's ear to isolate the user's ear canal from the surrounding physical environment. In other use cases, earbud 303 and eartip 314 together block some, but not necessarily all, ambient sound from the surrounding physical environment from reaching the user's ear.
[0034] A first microphone or first microphone array 302-1 is located on the wearable audio output device 301 to capture ambient sound, represented by waveform 322 in region 316 of the user's physical environment. A second microphone or second microphone array 302-2 is located on the wearable audio output device 301 to capture any ambient sound that is not completely blocked by the earbud 303 and eartip 314 and can be heard in region 318 within the user's ear canal, represented by waveform 324. In one aspect, the second microphone 302-2 can be used to capture the user's near-field voice signal.
[0035] Return Reference Figure 1 If the remote speaker moves away from the local speaker, the far-field voice signal weakens with the distance between the two speakers. The local audio output device can analyze the far-field voice signal and the near-field voice signal to estimate the acoustic parameters of the local environment and the rate of change of the estimated acoustic parameters. In one aspect, the local audio output device and the remote audio output device can each estimate the acoustic parameters of their respective environments and their rate of change based on their respective near-field voice signals and far-field voice signals. The two audio output devices can exchange the estimated acoustic parameters, for example, via a direct RF link in peer-to-peer mode, to increase the confidence level of the estimated acoustic parameters. For example, the local audio output device can analyze acoustic parameters received from the remote audio output device, where the acoustic parameters are estimated by the remote audio output device based on the local speaker's near-field voice signal acoustically received by the remote audio output device. The local audio output device can reciprocally estimate the acoustic parameters associated with the remote speaker's far-field voice signal acoustically received by the local audio output device. In effect, the two audio output devices can act as a distributed, non-phase-locked microphone array to perform reciprocal estimation of the acoustic parameters. In one aspect, if the other audio output device does not have processing capabilities, has processing limitations, or wants to save power, only one of the two audio output devices can estimate the acoustic parameters and their rate of change. The audio output device that estimates the acoustic parameters can transmit the estimated acoustic parameters to the other audio output device via an RF link.
[0036] The local audio output device can analyze the estimated acoustic parameters to determine whether it is possible to continue the conversation in transparent mode. If the analysis of the acoustic parameters indicates that the far-field voice signal is sufficiently attenuated that it is likely unintelligible, the local audio output device can establish a direct low-latency RF link with the remote audio output device in peer-to-peer mode to electromagnetically receive the far-field voice signal via the direct RF link. To achieve a smooth transition, the local audio output device can process the estimated acoustic parameters to generate spatialized metadata for the remote speaker.
[0037] The local audio output device may use the spatialization metadata to re-spatialize the far-field speech signal received over the direct RF link to have a spatially simulated level and perceived direction of arrival of the remote talker. The spatialized far-field speech signal from the direct RF link may be used to enhance the far-field speech signal acoustically received by the microphone in transparent mode. In one aspect, the local audio output device may time-align the far-field speech signal from the microphone with the spatialized far-field speech signal from the direct RF link to generate an enhanced far-field speech signal. In one aspect, the local audio output device may sum the far-field speech signal received from the microphone with the spatialized far-field speech signal from the RF link to improve the SNR of the enhanced far-field speech. In one aspect, the local audio output device may switch to peer-to-peer mode to output the spatialized far-field speech signal to the speakers of the local audio output device without enhancing the acoustic far-field speech signal of the transparent mode.
[0038] If the remote caller moves further away from the local caller, the local audio output device can analyze the estimated acoustic parameters to determine that the direct RF link is exceeding its operating range. The local audio output device can switch to operate in telephone mode with the remote audio output device. The local audio output device can equalize the far-field voice signal received via the telephone signal to have a power spectrum similar to that of the spatialized far-field voice signal. In one aspect, the local audio output device can estimate the running statistics of the power spectral density (PSD) of the spatialized far-field voice signal in transparent mode or peer-to-peer mode. The local audio output device can use the running PSD estimate to equalize the far-field voice signal received via the telephone signal to smooth the transition to telephone mode. The local audio output device can output the equalized far-field voice signal to the speaker of the local audio output device in telephone mode. In one aspect, if the remote caller is not wearing an audio output device, or the remote audio output device does not have the capability of a direct link RF in peer-to-peer mode, the local audio output device can directly switch between transparent mode and telephone mode. For example, the local audio output device may analyze the estimated acoustic parameters in the transparent mode to determine that the far-field voice signal acoustically received from the microphone is sufficiently attenuated such that the communication mode should be switched from the transparent mode to the telephony mode.
[0039] Figure 3 A functional block diagram depicts a system 300 for processing ambient sound, including voice signals acoustically captured by a microphone array of a local wearable audio output device and voice signals electromagnetically received from a remote wearable audio output device, to determine a communication mode between audio output devices based solely on acoustic analysis and to switch between communication modes, according to one aspect of the present disclosure. The system 300 can be located in the local audio output device or in a mobile device paired with the local audio output device.
[0040] The microphone array 340 may include Figure 2 3. The first microphone / microphone array 302-1 and the second microphone / microphone array 302-2 of the wearable audio output device 301 are depicted in FIG. The microphone array 340 can capture far-field speech signals of a remote talker and near-field speech signals of a local talker. In one aspect, the microphones in the microphone array 340 can have directional sensitivity to enable the system 300 to estimate the direction of arrival of the far-field speech signals.
[0041] The feature extractor module 350 can process acoustic signals of far-field and near-field speech signals to estimate parameters of the acoustic environment and the rate of change of the acoustic parameters. In one aspect, the feature extractor module 350 can receive acoustic parameters estimated by a remote audio output device. The local and remote audio output devices can exchange the estimated acoustic parameters via a direct RF link in peer-to-peer mode to increase the confidence level of the estimated acoustic parameters. In one aspect, the local audio output device can use the acoustic parameters estimated by the remote audio output device to estimate its round-trip acoustic parameters. For example, the estimated acoustic parameters received from the remote audio output device may indicate that the far-field speech signal from the local speaker is received by the remote audio output device at a certain direction of arrival at a certain speech level, while the near-field speech signal from the remote speaker is captured by the remote audio output device at another level. Based on the round-trip relationship between the two audio output devices, the feature extractor module 350 can use this information and information about the estimated speech level of the local speaker's near-end speech signal to estimate the direction of arrival and speech level of the remote speaker's far-field speech signal. In one aspect, the local audio output device may estimate acoustic parameters without assistance from the remote audio output device, and then use the acoustic parameters estimated by the remote audio output device to validate or refine the acoustic parameters estimated by the local audio output device.
[0042] Figure 4 Depicted is a functional block diagram of a feature extractor module 350 that processes near-field speech and far-field speech signals to estimate parameters of the acoustic environment for determining a communication mode of a local audio output device, according to one aspect of the present disclosure.
[0043] The filtering module 351 can filter the acoustic signals captured by the microphone array 340 to detect far-field speech signals and near-field speech signals. Figure 2302-1 and the second microphone / microphone array 302-2 of the wearable audio output device 301 depicted in FIG. 302-1 capture acoustic signals to detect far-field speech signals and near-field speech signals, respectively. In one aspect, the filtering module 351 can filter signals received via a direct RF link in peer-to-peer mode to detect far-field speech signals or acoustic parameters estimated by a remote audio output device. Various modules can process the far-field and near-field speech signals to estimate various acoustic parameters.
[0044] For example, the near-field level variation estimation module 352 can process the near-field speech signal to estimate the level of the near-field speech signal over time. For example, the near-field level variation estimation module 352 can measure the Lombard effect, which is the involuntary tendency of a local talker to increase the acoustic effect to enhance the audibility of speech when speaking in louder noise or when the distance from the remote talker increases. Such acoustic effects can include increased loudness, higher pitch, slower rate, or longer syllable duration, etc.
[0045] The far-field to near-field level difference estimation module 353 can process the near-field and far-field speech signals to estimate the level or volume difference between the near-field and far-field speech signals and the change in the level difference. For example, when the remote talker is far away from the local talker, the level difference between the near-field and far-field speech signals can be large. In one aspect, the far-field to near-field level difference estimation module 353 can estimate the PSD of the near-field and far-field speech signals. The PSDs of the near-field and far-field speech signals can be compared, or their relative rates of change can be analyzed to estimate the distance between the local talker and the remote talker, or to estimate changes in the acoustic environment.
[0046] The far-field direct-to-reverberation ratio (DRR) estimation module 354 can process the far-field speech signal to estimate the DRR and the change in DRR of the far-field speech signal. In one aspect, the voice activity detector and the near-field / far-field classifier can detect the far-field speech signal and estimate the direct component and the reverberation component of the far-field speech signal to estimate the DRR. In one aspect, the voice activity detector and the near-field / far-field classifier can apply machine learning methods, such as using a convolutional neural network (CNN), a recurrent neural network (RNN), etc. In one aspect, the voice activity detector can detect speech on the near-field speech signal. The local audio output device can transmit a signal to the remote audio output device indicating the detection of the local talker's speech so that the remote audio output device can estimate the acoustic parameters of the speech signal received from the local talker. Reciprocally, the feature extractor module 350 of the local audio output device can receive a signal from the remote audio output device indicating the detection of the speech from the remote talker so that the feature extractor module 350 can estimate the acoustic parameters of the far-field speech signal.
[0047] The far-field dominance estimation module 355 can process the far-field speech signal to estimate its energy distribution and the variation of the energy distribution, such as by estimating the spatial covariance matrix and the temporal variance of the spatial covariance matrix. The far-field dominance estimation module 355 can measure whether the energy of the far-field speech signal is dominated by a dense source (such as when the far talker has a clear acoustic signature) or diffuse energy (such as when the far talker is too far away to have a meaningful acoustic signature).
[0048] The far-field direction of arrival and localization module 356 can process far-field speech signals to estimate their direction of arrival and changes in direction of arrival. In one aspect, the microphone array 340 can have directional sensitivity to enable the far-field direction of arrival and localization module 356 to estimate the direction of arrival of the far-field speech signals. In one aspect, the direction of arrival of the far-field speech signals from the local talker estimated by the remote audio output device can be used as an aid to the local audio output device to estimate the direction of arrival of the far-field speech signals from the remote talker based on the reciprocity of the spatial relationship between the two audio output devices.
[0049] The far-field speech intelligibility index module 357 can process the far-field speech signal to estimate intelligibility parameters and changes in intelligibility parameters of the far-field speech. In one aspect, the far-field speech intelligibility index module 357 can apply machine learning methods, such as using CNN, RNN, etc.
[0050] Return Reference Figure 3 The classifier and parameter estimator module 360 can analyze the estimated acoustic parameters to determine the optimal communication mode for the local and remote audio output devices for the local and remote talkers to use in their conversation. In one aspect, the optimal communication mode can be a function of the intelligibility, directionality, DRR, energy distribution, etc. of the far-field speech signal.
[0051] If the analysis of the acoustic parameters by the classifier and parameter estimator module 360 indicates that the current communication mode may no longer support a conversation between the local talker and the remote talker, the classifier and parameter estimator module 360 may request the local audio output device to switch to a different communication mode. For example, when the signal captured by the microphone array 340 may no longer support acoustic communication between the local talker and the remote talker in transparent mode due to increased distance or due to noise sources, the local audio output device may enhance the acoustic signal in transparent mode with a far-field voice signal received via a direct low-latency RF link in peer-to-peer mode. The classifier and parameter estimator module 360 may estimate the required horizontal and directional metadata for re-spatializing the far-field voice signal received over the RF link based on the extracted acoustic parameters. The far-field voice signal received over the RF link may be re-spatialized so that it is consistent with the spatial position of the remote speaker so that the acoustic signal can be enhanced in a seamless manner.
[0052] In one aspect, the communication mode used on both the local and remote audio output devices can be the same. The local audio output device can synchronize the switching of the communication mode with the remote audio output device. In one aspect, the communication mode used on the local and remote audio output devices can be different. This asymmetric mode may occur when local noise or interference sources only affect the local or remote audio output device.
[0053] Figure 5 Depicted is a functional block diagram of a classifier and parameter estimator module 360 that processes estimated parameters to determine a communication mode and spatialization metadata for respatializing far-end voice signals received in peer-to-peer mode, according to one aspect of the present disclosure.
[0054] The speech mode determination module 361 can process the estimated acoustic parameters, such as a near-field level variation parameter, a far-field to near-field level difference parameter, a far-field DRR parameter, a far-field dominance parameter, a far-field direction of arrival and localization parameter, a far-field speech intelligibility parameter, etc., to determine the optimal communication mode. In one aspect, the speech mode determination module 361 can determine a composite intelligibility index of the far-field speech signal from the estimated acoustic parameters. If the composite intelligibility index is above a first threshold, the speech mode determination module 361 can determine that the optimal communication mode is transparent mode. If the composite intelligibility index drops below the first threshold but above a second threshold, the speech mode determination module 361 can determine that the optimal communication mode is to use the far-field speech signal received via the direct RF link to enhance the acoustic signal of the transparent mode. If the composite intelligibility index drops below a second threshold, the speech mode determination module 361 can determine that the optimal communication mode is telephony mode.
[0055] To enhance the acoustic signal in transparent mode using the far-field speech signal received via direct low-latency RF, the spatial parameter estimator 362 may estimate spatialized metadata to be applied to the far-field speech signal received via direct low-latency RF. For example, the speech mode determination module 361 may provide the spatial parameter estimator 362 with far-field to near-field level difference parameters, far-field direction of arrival and localization parameters, far-field speech intelligibility parameters, etc., so that the spatial parameter estimator 362 generates spatialized metadata of the remote talker, such as horizontal spatial metadata and directional spatial metadata.
[0056] Return Reference Figure 3 The spatial filter 370 can use the spatialization metadata to re-spatialize the far-field voice signal received via the direct RF link to have a level and perceived direction of arrival that spatially simulates the remote talker. The spatial filter 370 can also generate a PSD of the spatialized far-field voice signal for equalizing the far-field voice signal received via the telephony mode when the communication mode is switched to the telephony mode.
[0057] Figure 6 Depicted is a functional block diagram of a spatial filter 370 that uses spatialization metadata to respatialize far-end voice signals received in peer-to-peer mode and to generate power spectrum metadata for equalizing far-end voice signals received in telephony mode, in accordance with one aspect of the present disclosure.
[0058] Speech spatialization filter 371 applies the horizontal spatialization metadata and directional spatialization metadata generated by classifier and parameter estimator module 360 to the far-end speech signal received from the direct RF link in peer-to-peer mode to generate a spatialized speech signal. The spatialized far-field speech signal from the direct RF link can be used to enhance the far-field speech signal acoustically received by microphone array 340 in transparent mode. In one aspect, speech spatialization filter 371 can add the far-field speech signal received from the microphone to the spatialized far-field speech signal from the RF link to improve the SNR of the enhanced far-field speech in transparent mode or peer-to-peer mode.
[0059] The time alignment / mixer module 372 can time-align and mix the far-field speech signal from the microphone array 340 with the spatialized far-field speech signal from the direct RF link to generate an enhanced far-field speech signal. In one aspect, if the far-field speech signal from the microphone array 340 has a shorter delay than the spatialized far-field speech signal from the direct RF link due to the long processing delay of the speech spatialization filter 371, the frames of the far-field speech signal from the microphone array 340 can be delayed by the delay buffer to time-align with the frames of the spatialized far-field speech signal. In one aspect, if the spatialized far-field speech signal from the direct RF link has a shorter delay than the far-field speech signal from the microphone array 340, the frames of the spatialized far-field speech signal can be delayed by the delay buffer to time-align with the frames of the far-field speech signal from the microphone array 340.
[0060] The power spectrum estimation module 374 can estimate running statistics of the PSD of the spatialized far-field speech signal or enhanced far-field speech signal in transparent mode or peer-to-peer mode to generate power spectrum metadata. The power spectrum metadata can be used to equalize the far-end speech signal received in telephony mode to have a power spectrum similar to that of the spatialized far-field speech signal or enhanced far-field speech signal to smooth the transition to telephony mode. In one aspect, the power spectrum estimation module 374 can estimate running statistics of the PSD of the near-field speech signal in transparent mode to generate power spectrum metadata. When the communication mode is directly transitioned from transparent mode to telephony mode, the power spectrum metadata can be used to equalize the far-end speech signal received in telephony mode.
[0061] Return Reference Figure 3, the summing module 380 can use the power spectrum metadata to equalize the far-end voice signal received in telephony mode. The summing module 380 can sum the equalized far-end voice signal in telephony mode and the spatialized far-field voice signal or enhanced far-field voice signal in transparent mode or peer-to-peer mode to generate a processed far-field voice signal to drive the speaker 390 of the local audio output device. Alternatively, in transparent mode or peer-to-peer mode, the acoustic signal, spatialized far-field voice signal, or enhanced far-field voice signal from the microphone array 340 can be driven to the speaker 390.
[0062] Figure 7 is a flow chart of a method 700 for determining a communication mode based solely on acoustic analysis and switching between communication modes of a wearable audio output device such as headphones according to one aspect of the present disclosure. The method 700 may be performed by Figure 3 System 300 Practice.
[0063] In operation 701, method 700 processes a near-field speech signal and a far-field speech signal received by a local earphone to estimate acoustic parameters of an acoustic environment. The near-field speech signal is received from a local user of the local earphone, and the far-field speech signal is received from a remote user of the remote earphone.
[0064] In operation 703, the method 700 processes the estimated acoustic parameters to determine a communication mode between the local headset and the remote headset. The communication mode includes an acoustic transparency mode, a peer-to-peer RF mode, or a telephony mode.
[0065] In operation 705 , the method 700 determines whether the communication mode is transparent mode. If it is transparent mode, operation 709 outputs the far-field voice signal to the local user of the local headset.
[0066] If the communication mode is not transparent mode, operation 707 determines whether the communication mode is in RF peer-to-peer mode. If it is in RF peer-to-peer mode, operation 709 outputs a spatialized voice signal based on the far-field voice signal to the local headset. In one aspect, method 700 may process the far-field voice signal received in RF peer-to-peer mode to generate a spatialized voice signal based on the perceived direction of the remote user determined from the estimated acoustic parameters.
[0067] Otherwise, if the communication mode is neither the transparent mode nor the RF peer-to-peer mode, operation 709 outputs a telephone voice signal based on the far-field voice signal to the local headset.
[0068] Figure 8 is a flow chart of a method 800 for enhancing an acoustic signal of far-field speech captured by a microphone of a wearable audio output device, such as a headset, with a far-field speech signal carried over an RF transmission based solely on acoustic analysis according to one aspect of the present disclosure. The method 800 may be performed by Figure 3 System 300 Practice.
[0069] In operation 801 , the method 800 processes a near-field speech signal and a far-field speech signal received as acoustic signals by a microphone to estimate acoustic parameters of an acoustic environment.
[0070] In operation 803, method 800 processes the estimated acoustic parameters to determine whether to enhance the acoustic signal using the far-field speech signal carried by the RF transmission.
[0071] In operation 805, method 800 checks whether the decision is to enhance the acoustic signal. If not, operation 811 outputs the original far-field speech signal to the speaker of the headset.
[0072] If the decision is to enhance the acoustic signal, then in operation 807 the method notifies the remote headset to switch to peer-to-peer mode.
[0073] In operation 809 , the method 800 processes the far-field voice signal received by the microphone and enhances the far-field voice signal received by the microphone by the peer to peer RF signal.
[0074] In operation 811 , the method 800 outputs the enhanced far-field speech signal to a speaker of a headset.
[0075] The embodiments of the stereo signal identifier and audio signal identifier described herein can be implemented in a data processing system, for example, via a network computer, network server, tablet computer, smartphone, laptop computer, desktop computer, other consumer electronic device, or other data processing system. Specifically, the operations described for determining the optimal communication mode for use by the wearable audio output device are digital signal processing operations performed by a processor executing instructions stored in one or more memories. The processor can read the stored instructions from the memory and execute the instructions to perform the described operations. These memories represent examples of machine-readable, non-transitory storage media that can store or contain computer program instructions that, when executed, cause a data processing system to perform one or more methods described herein. The processor can be a processor in a local device such as a smartphone, a processor in a remote server, or a distributed processing system of multiple processors in local devices and remote servers, wherein their respective memories contain portions of the instructions necessary to perform the described operations.
[0076] The processes and blocks described herein are not limited to the specific examples described, and are not limited to the specific order used as examples herein. Instead, any processing blocks may be reordered, combined or removed, executed in parallel or serially, as needed to achieve the above results. The processing blocks associated with implementing the audio processing system may be executed by one or more programmable processors executing one or more computer programs stored on a non-transitory computer-readable storage medium to perform the functions of the system. All or part of the audio processing system may be implemented as a dedicated logic circuit (e.g., an FPGA (field programmable gate array) and / or an ASIC (application-specific integrated circuit)). All or part of the audio system may be implemented using an electronic hardware circuit comprising at least one of an electronic device such as a processor, a memory, a programmable logic device, or a logic gate. In addition, the process may be implemented in any combination of hardware devices and software components.
[0077] Although certain illustrative examples are described and shown in the drawings, it is to be understood that these examples are merely illustrative and not restrictive of the broader invention, and that the invention is not limited to the exact construction and arrangements shown and described, since various other modifications may be made by those skilled in the art. Accordingly, the description is to be regarded as illustrative rather than restrictive.
[0078] To assist the Patent Office and any reader of any patent issuing in this application in interpreting the appended claims, applicants wish to note that they do not intend that any of the appended claims or claim elements invoke 35 U.S.C. § 112(f) unless the words “means for” or “step for” are expressly used in a particular claim.
[0079] As described above, one aspect of the present technology is the transmission and use of voice or data from specific and legitimate sources to audio output devices using different communication modes. The present disclosure contemplates that, in some cases, the voice or data may include personal information data that uniquely identifies or can be used to identify a specific person. Such personal information data may include demographic data, location-based data, online identifiers, phone numbers, email addresses, home addresses, data or records related to the user's health or fitness level (e.g., vital signs measurements, medication information, exercise information), date of birth, or any other personal information. The present disclosure recognizes that the use of such personal information data in the present technology can be used to benefit users.
[0080] This disclosure contemplates that entities responsible for collecting, analyzing, disclosing, transmitting, storing, or otherwise using such personal information will adhere to established privacy policies and / or practices. Specifically, such entities are expected to implement and consistently apply privacy practices generally recognized as meeting or exceeding industry or government requirements for safeguarding user privacy. Such information regarding the use of personal data should be prominently displayed and easily accessible to users and updated as the collection and / or use of data changes. Users' personal information should be collected only for lawful uses. Furthermore, such collection / sharing should occur only after receiving user consent or other lawful basis as provided in applicable law. Furthermore, such entities should consider taking any necessary steps to safeguard and secure access to such personal information and ensure that others with access to such personal information adhere to their privacy policies and procedures. Furthermore, such entities may subject themselves to third-party assessments to demonstrate compliance with widely accepted privacy policies and practices. Furthermore, policies and practices should be tailored to the specific types of personal information being collected and / or accessed and adapted to applicable laws and standards, including jurisdiction-specific considerations that may impose higher standards. For example, in the United States, the collection or access of certain health data may be governed by federal and / or state laws, such as the Health Insurance Portability and Accountability Act (HIPAA); while health data in other countries may be subject to other regulations and policies and should be handled accordingly.
[0081] Regardless of the foregoing, the present disclosure also contemplates implementations in which users selectively block the use or access of personal information data. That is, the present disclosure contemplates providing hardware components and / or software components to prevent or block access to such personal information data.
[0082] Furthermore, it is an object of the present disclosure that personal information data should be managed and processed to minimize the risk of unintentional or unauthorized access or use. Risk can be minimized by limiting data collection and deleting data once it is no longer needed. In addition, and when applicable, including in certain health-related applications, data de-identification can be used to protect the privacy of users. De-identification can be facilitated where appropriate by removing identifiers, controlling the amount or specificity of stored data (e.g., collecting location data at a city level rather than an address level), controlling how data is stored (e.g., aggregating data across users), and / or other methods such as differential privacy.
[0083] Thus, while the present disclosure broadly covers the use of transmission of personal information data to implement one or more of the various disclosed embodiments, the present disclosure also contemplates that various embodiments may be implemented without access to such personal information data. That is, various embodiments of the present technology will not be unable to function properly due to the absence of all or a portion of such personal information data. For example, content may be selected and delivered to a user based on aggregated non-personal information data or an absolute minimum amount of personal information, such as content processed only on the user's device or other non-personal information available to the content delivery service.
Claims
1. A method for communicating between a local headset and a remote headset, the method comprising: Processing, by the local earphone, a near-field speech signal of a local talker wearing the local earphone and a far-field speech signal of a remote talker received by the local earphone to estimate acoustic parameters; determining a communication mode between the local headset and the remote headset worn by the remote caller based on the estimated acoustic parameters, the communication mode comprising one of the following: an acoustic transparency mode in which the far-field voice signal is captured by a microphone of the local headset, a peer-to-peer radio frequency (RF) mode, or a telephony mode, wherein the peer-to-peer RF mode or the telephony mode uses RF signals for communication between the local headset and the remote headset; as well as The far-field voice signal in the acoustic transparency mode, the enhanced voice signal based on the far-field voice signal in the peer RF mode, or the telephone voice signal based on the far-field voice signal in the telephone mode is output to the speaker of the local headset. The method of claim 1 , wherein the microphone comprises a microphone array.
3. The method of claim 1 , wherein determining the communication mode comprises: generating an intelligibility index of an acoustic signal carrying the far-field speech signal based on the estimated acoustic parameters, the far-field speech signal being captured by a microphone of the local headset; determining whether the comprehensibility index exceeds a first comprehensibility threshold; In response to the intelligibility index exceeding the first intelligibility threshold, determining the acoustic transparency mode as the communication mode; as well as The acoustic signal carrying the far-field speech signal is output to the speaker of the local headset in the acoustic transparency mode.
4. The method according to claim 3, further comprising: In response to the comprehensibility index not exceeding the first comprehensibility threshold, determining whether the comprehensibility index exceeds a second comprehensibility threshold; In response to the intelligibility index exceeding the second intelligibility threshold, determining the peer-to-peer RF mode as the communication mode, wherein in the peer-to-peer RF mode, the local headset receives the RF signal carrying the far-field speech signal via a peer-to-peer RF link with the remote headset; as well as In response to the intelligibility index not exceeding the second intelligibility threshold, the telephony mode is determined as the communication mode, wherein in the telephony mode, the local headset receives an RF signal carrying the far-field speech signal through a network link with the remote headset.
5. The method according to claim 4, further comprising: When the peer to peer RF mode is determined as the communication mode, the enhanced voice signal is generated based on enhancing the acoustic signal carrying the far-field voice signal with the RF signal carrying the far-field voice signal.
6. The method of claim 5, wherein generating the enhanced speech signal comprises: generating spatialized metadata for the remote talker using the estimated acoustic parameters; generating, based on the far-field speech signal carried by the RF signal and the spatialization metadata, a spatialized far-field speech signal having a level and a direction of arrival that spatially simulates the remote talker; as well as The enhanced speech signal is generated based on enhancing the acoustic signal carrying the far-field speech signal with the spatialized far-field speech signal to increase a signal-to-noise ratio (SNR) of the far-field speech signal.
7. The method of claim 6, wherein the spatialized far-field speech signal is spatially consistent with the acoustic signal carrying the far-field speech signal, and wherein the method further comprises: The enhanced speech signal is generated based on temporally aligning the acoustic signal with the spatialized far-field speech signal.
8. The method according to claim 1, further comprising: estimating a power spectrum of the far-field speech signal; as well as When the telephone mode is determined as the communication mode, the telephone voice signal equalized by the power spectrum of the far-field voice signal is generated.
9. The method of claim 1 , wherein processing the near-field voice signal of the local talker and the far-field voice signal of the remote talker comprises: An RF signal received by the local earphone from the remote earphone is processed to estimate the acoustic parameters, wherein the RF signal includes information about a reciprocating far-field speech signal of the local talker acoustically received by the remote earphone.
10. The method of claim 9, wherein the information about the reciprocating far-field speech signal comprises reciprocating acoustic parameters estimated by the remote earphone.
11. The method according to claim 1 , further comprising: The estimated acoustic parameters are transmitted by the local earphone to the remote earphone to assist the remote earphone in determining a communication mode between the local earphone and the remote earphone.
12. The method of claim 1 , wherein the estimated acoustic parameters include one or more of: a speech level difference between the near-field speech signal and the far-field speech signal or a rate of change of the speech level difference; a direct-to-reverberation ratio (DRR) of speech levels of a direct component and a reverberation component of the far-field speech signal; The rate of change of the DRR; an energy distribution measure of the far-field speech signal; the rate of change of said energy distribution measure; a change in the speech level of the near-field speech signal; a rate of change of the speech level of the near-field speech signal; an estimated direction of arrival of the far-field speech signal; a rate of change of the estimated direction of arrival; an intelligibility measure of the far-field speech signal; and The rate of change of the intelligibility measure.
13. An earphone, comprising: At least one processor configured to perform operations comprising: processing a near-field speech signal of a local talker wearing the headset and a far-field speech signal of a remote talker received by the headset to estimate acoustic parameters; determining a communication mode between the headset and a remote headset worn by the remote caller based on the estimated acoustic parameters, wherein the communication mode comprises one of the following: an acoustically transparent mode in which the far-field voice signal is captured by a microphone of the headset, a peer-to-peer radio frequency (RF) mode, or a telephony mode, the peer-to-peer RF mode or the telephony mode being configured to communicate between the local headset and the remote headset using RF signals; and The far-field voice signal in the acoustic transparency mode, the enhanced voice signal based on the far-field voice signal in the peer RF mode, or the telephone voice signal based on the far-field voice signal in the telephone mode is output to a speaker of the headset.
14. The headset of claim 13, wherein the operation of determining the communication mode comprises an operation for: generating an intelligibility index of an acoustic signal carrying the far-field speech signal based on the estimated acoustic parameters, the far-field speech signal being captured by a microphone of the headset; determining whether the comprehensibility index exceeds a first comprehensibility threshold; In response to the intelligibility index exceeding the first intelligibility threshold, determining the acoustic transparency mode as the communication mode; outputting the acoustic signal carrying the far-field speech signal to the speaker of the headset in the acoustic transparency mode; In response to the comprehensibility index not exceeding the first comprehensibility threshold, determining whether the comprehensibility index exceeds a second comprehensibility threshold; In response to the intelligibility index exceeding the second intelligibility threshold, determining the peer-to-peer RF mode as the communication mode, wherein in the peer-to-peer RF mode, the headset receives an RF signal carrying the far-field speech signal over a peer-to-peer RF link with the remote headset; as well as In response to the intelligibility index not exceeding the second intelligibility threshold, the telephony mode is determined as the communication mode, wherein in the telephony mode, the headset receives an RF signal carrying the far-field speech signal over a network link with the remote headset.
15. The headset of claim 14, wherein the operations further comprise: When the peer to peer RF mode is determined as the communication mode, the enhanced voice signal is generated based on enhancing the acoustic signal carrying the far-field voice signal with the RF signal carrying the far-field voice signal to increase a signal-to-noise ratio (SNR) of the far-field voice signal, wherein the enhanced voice signal is spatially consistent and temporally aligned with the acoustic signal.
16. The headset of claim 13, wherein the operation of processing the near-field voice signal of the local talker and the far-field voice signal of the remote talker comprises an operation for: An RF signal received by the earphone from the remote earphone is processed to estimate the acoustic parameters, wherein the RF signal includes reciprocating acoustic parameters estimated by the remote earphone with respect to a reciprocating far-field speech signal of the local talker acoustically received by the remote earphone.
Citation Information
Patent Citations
Self-adaptive call volume control method and device
CN109994104A
Method and apparatus for processing speech signal adaptive to noise environment
CN110447069A