Echo detection device, echo detector, host device, and non-transitory computer readable medium

By detecting and echoing and echoing using a lookup table for normalizing volume settings, the echo and volume instability issues when switching audio devices are solved, and lossless audio output and seamless user experience are achieved.

CN120199261APending Publication Date: 2025-06-24INTEL CORP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411680272.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-12-22
Filing Date
2024-11-22
Publication Date
2025-06-24

AI Technical Summary

Technical Problem

In voice calls and conference calls, audio devices switch easily lead to instability in echo and volume, resulting in poor user experience, especially in latency-critical applications.

Method used

By detecting and echoing, using gain to normalize the volume settings of the lookup table for power outputs, and recording and presenting audio signals at lower sample rates when the audio device switches to avoid audio loss.

Benefits of technology

It realizes lossless audio output when switching audio devices, eliminates echo interference, provides a seamless user experience, and avoids delay problems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120199261A_ABST
    Figure CN120199261A_ABST
Patent Text Reader

Abstract

An echo detection apparatus is provided. For example, an echo detection device may include a processor configured to activate a transcription process that transcribes a word generated from an audio transfer between a host device and at least one further device into text; analyzing the text for identifying duplicated strings indicative of echoes; and if an echo is detected: decoding the audio transmission into a plurality of audio packets; forming an echo filtered audio stream by filtering out similar audio packets of the plurality of audio packets; and transmitting the echo filtered audio stream provision to the at least one further device and / or an audio device of the host device.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure generally relates to an echo detection device, an echo detector, a host device, and a non-transitory computer-readable medium. Background Art

[0002] Bluetooth usage for voice calls via headphones, car kits, earbuds, etc. has become a necessity for both work and personal use. Switching audio from an external speaker to headphones (and vice versa) is a common use case.

[0003] The situation and the desired and / or provided privacy level can determine which audio device to select.

[0004] In the case of voice call audio, it is important to provide a high-quality user experience without latency or data loss. Brief Description of the Drawings

[0005] In the drawings, the same reference numerals generally refer to the same parts throughout different views. These drawings are not necessarily to scale, but rather emphasize generally showing the principles of the present invention. In the following description, various aspects of the present invention are described with reference to the following drawings, in which:

[0006] Figure 1 Exemplary schematic diagrams showing an echo detection device according to various embodiments.

[0007] Figure 2 Exemplary schematic diagrams showing an echo detector according to various embodiments.

[0008] Figure 3 Exemplary flowcharts showing the operation of an echo detector according to various embodiments.

[0009] Figure 4 Exemplary schematic diagrams showing a host device according to various embodiments.

[0010] Figure 5 Exemplary schematic diagrams showing a host device according to various embodiments.

[0011] Figure 6 Exemplary lookup tables showing a host device according to various embodiments.

[0012] Figure 7 Exemplary flowcharts showing the operation of a host device for adjusting output power according to various embodiments.

[0013] Figure 8 Exemplary schematic diagrams showing a host device according to various embodiments.

[0014] Figure 9 Exemplary schematic diagrams showing a host device according to various embodiments.

[0015] Figure 10 Exemplary schematic diagrams showing a host device according to various embodiments.

[0016] Figure 11 Exemplary table showing default parameters of a host device according to various embodiments.

[0017] Figure 12 Showing the operation of a host device according to various embodiments; and

[0018] Figure 13 Exemplary flowchart showing audio transmission processing between a host device and at least one other device according to various embodiments. Detailed Description

[0019] The following detailed description refers to the accompanying drawings, which by way of illustration show specific details in which the present disclosure may be practiced. One or more examples are described in sufficient detail to enable those skilled in the art to practice the present disclosure. Other examples may be utilized and structural, logical, and electrical changes may be made without departing from the scope of the disclosure. The various examples described herein are not necessarily mutually exclusive, as some examples may be combined with one or more other examples to form new examples. Various examples are described in connection with methods and various examples are described in connection with devices. However, it will be understood that examples described in connection with methods may be similarly applicable to devices and vice versa. Throughout the drawings, it should be noted that like reference numerals are used to depict like or similar elements, features, and structures.

[0020] When receiving an audio transmission, for example, during a conference call, when there is an echo created for various reasons, remote users may experience a bothersome audio experience.

[0021] In various embodiments, the transcript created during a conference call can be used to detect and eliminate echoes, and can be used as a basis for techniques to provide a satisfactory (e.g., echo-free) user experience. Algorithms running on a host device (e.g., a PC sending an audio signal) or in the cloud can continuously monitor redundant words and identify possible sources of echoes.

[0022] In other words, in various embodiments, a method for echo detection and cancellation in an audio signal (and transcript) during a video conference call, for example, in a Windows system, is provided.

[0023] When a user switches from one (e.g., Bluetooth-) headset to another, a sudden difference in volume level may be experienced. This can result in an unpleasant, especially sudden, change in audio level and, thus, damage to the ears.

[0024] In various embodiments, volume settings for a Bluetooth headset can be normalized by creating a lookup table (LUT) of gain versus power output during an initial connection process. To facilitate seamless transitions between headset types, the power level of a newly connected BT device can be normalized based on a previous (e.g., last) used value in the lookup table. Lookup tables can be created for different users and volume settings can be applied accordingly.

[0025] Accordingly, in various embodiments, a method is provided for normalizing volume levels across different (Bluetooth-) headsets for seamless transitions.

[0026] During an active audio (e.g., voice) session, any audio application that uses the device's microphone and relies on real-time audio may experience audio loss when switching from a speaker (as an output device) to (e.g., Bluetooth) headphones.

[0027] Some applications are latency-critical. The term "latency-critical application" may refer to an application where even a small delay in audio transmission may have a negative impact on the user experience. Some examples of latency-critical applications using Bluetooth audio include Teams calls, video conferencing, gaming, listening to music podcasts, navigation widgets, voice assistants, audiobooks, etc.

[0028] As an example, during a Teams call, if the wireless headphones are newly connected, there will be a loss of audio for a few seconds (e.g., two to five seconds) due to the switch, and the audio may be unavailable on any device. In this scenario, the voice call audio that was currently / initially presented on a first device (e.g., a speaker) may be intentionally switched by the user to (Bluetooth-) headphones.

[0029] During a switching scenario, it may take a while to set up the audio path, and there is a possibility of missing at least some audio information as it may not appear on the device (e.g., Bluetooth-) headphones) that the user has most recently connected to.

[0030] Similar issues may occur in certain voice recordings or GPS devices where audio is being received from a remote source and presented on a playback device.

[0031] The above problem of voice sampling loss during audio device switching scenarios is an improvement consideration. This may be a challenging problem to solve because there are various factors that can cause voice sampling loss during audio device switching. The factors can include, for example, latency, buffering, and hardware limitations. For example, if voice is sent as an audio signal (e.g., as part of a conversation), latency may be critical. Attempting to recover old data will add latency, which is intolerable. Therefore, the user experience of losing audio during audio (e.g., Bluetooth) channel setup is considered a trade-off for avoiding latency.

[0032] In various embodiments, a systematic method is provided to address the problem of lost audio (e.g., voice) samples during audio device switching in an active call. Thereby, a lossless audio user experience is provided without a trade-off regarding latency.

[0033] In various embodiments, methods and devices are provided for preventing audio loss during device switching in an active connection.

[0034] In various embodiments, a method is provided for presenting lost audio to a playback device without loss or latency trade-off of the audio.

[0035] In various embodiments, an enhanced user experience is provided for seamless device switching during an active audio connection without audio loss.

[0036] Use a handshake between an audio digital signal processor (audio DSP) and a controller (e.g., a Bluetooth controller) for wireless transmission of audio signals. A framework is proposed in which, once a newly connected audio device is detected, audio information (e.g., voice data) is recorded in the last few seconds when the new audio playback device is still unavailable, and the recorded / stored audio signal is presented using an enhanced sampling algorithm, where "enhanced" in this case refers to a lower sampling rate so that the amount of the audio signal represented by the sampled data is enhanced. The last recorded audio can be sent to an (e.g., Bluetooth) audio playback device at an increased (e.g., double) rate. Since the sampling rate is lower, the recorded data will catch up with the current moment. Thereby, the lost audio will be presented to the audio playback device without loss or latency trade-off of the audio.

[0037] The audio quality and rate may vary slightly for a few seconds during the playback of the recorded audio, but this should not have a significant impact on the listening experience.

[0038] Various embodiments avoid audio / voice sampling loss in audio device switching.

[0039] Various embodiments enable easy switching between different audio playback devices (e.g., between an external speaker (also referred to as a loudspeaker) and headphones) without loss of any audio information. A smooth transition can be provided as the audio can continue to play without any interruptions or delays, which may be particularly important in latency-critical applications.

[0040] Various embodiments address the problem of sampling loss during audio playback device switching. Similar sampling loss problems may occur (e.g., due to network congestion, bandwidth congestion, interference, etc.) in cases where there are gaps in the scheduling, because if for some reason sufficient samples are not ready for a particular interval, there may be a loss of audio data for that interval as it is real-time audio. However, the recovery mechanisms of various embodiments as outlined above and described in more detail below can apply to the few samples that are missing and can still be presented, thus recovering the loss.

[0041] Figure 1 Exemplary schematic diagram showing an echo detection device 100 according to various embodiments.

[0042] The echo herein should be understood to refer to the repeated sounds / words created by the external speaker during a conference call. If the echo persists for too long, it may affect the quality of the call.

[0043] In a conference call scenario, it may be difficult to identify the source of the echo without manual user intervention. This may lead to unnecessary delays and confusion among the users.

[0044] During a conference call, echoes may occur in the following two scenarios:

[0045] The device of a participant located in the same conference room as the host device can activate its external speaker. In this scenario, the audio played back by the external speaker of the participant's device will be captured by the microphone of the host device, and this audio will be retransmitted, thus creating an echo.

[0046] Remote participants use wireless (e.g., Bluetooth) devices for their external speakers and the microphones of their playback devices (e.g., PC, laptop, tablet, or smartphone) for recording. Existing echo cancellation algorithms are limited to working with an external speaker / microphone that are part of the same device. The playback through a wireless / Bluetooth external speaker will be captured by the microphone of the device, thus creating an echo.

[0047] The transcript of an activated audio session (e.g., a conference call) can be used for echo detection as the created echo will be reflected in the transcript.

[0048] As an example, a transcript of a part of a conference call may look as follows:

[0049] Person 1: "Yes, yes, yes, yes"

[0050] Person 2: "This is about the problem you want to solve, right?"

[0051] Person 1: "Received, received, received, received no"

[0052] Therefore, from the visual inspection, it is easy to identify repeated words ("yes" repeated three times) and word combinations ("received" also repeated three times).

[0053] Hereinafter, the echoes detected in the transcript are used to trigger actions to exclude echoes in the audio playback (and transcript).

[0054] The echo detection device 100 may include a processor 102.

[0055] The processor 102 may be a single digital processor 102, or may be composed of different potentially distributed processor units. The processor 102 may be at least one digital processor 102 unit. The processor 102 may include one or more of the following: a microprocessor, a microcontroller, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), discrete logic circuitry, etc., appropriately programmed with software and / or computer code, or a combination of dedicated hardware and programmable circuitry. The processor 102 may also be configured to access programs, software, etc. that may be stored in a storage element of the echo detection device 100 or an external storage element (e.g., in a computer network (e.g., cloud)).

[0056] The processor 102 may be configured to activate a transcription process that transcribes words generated from an audio transmission between a host device and at least one other device into text.

[0057] The host device may be a session host that hosts an audio session (e.g., an audio call or a teleconference), and at least one other device may be other participants in the audio session.

[0058] The echo detection device 100 may be or be included in the host device, or may be a different device (e.g., a server to which the host device is connected and to which the host device can deliver the transcript) or a part thereof.

[0059] The transcription process of transcribing words into text may be performed as is well known in the art. For example, certain audio sessions that host software packages (e.g., Microsoft Teams) have an optional function for providing a transcript.

[0060] The processor 102 may be configured to analyze the text generated by the transcription process to identify repeated strings indicative of an echo.

[0061] If a first string (at least two characters long) is immediately followed by a second identical string, this may indicate an echo.

[0062] The length of the string that typically produces an echo may depend on the processing delay resulting from (e.g., wireless) audio signal transmission and / or the size of the meeting room combined with the finite speed of sound, how fast a person speaks, etc., but typically, the echo affects word parts, individual words, or combinations of a few words, rather than an entire sentence, etc.

[0063] The limit on the maximum number of characters in a string that can be analyzed by the processor may be set, for example, accordingly to any natural number between 2 and 10.

[0064] The processor 102 may also be configured to decode the audio transmission into a plurality of audio packets if an echo is detected.

[0065] Wireless communication that sends audio data typically provides an audio signal according to a corresponding wireless standard. These audio signals that have been modified to be wirelessly sent are called audio packets.

[0066] The processor 102 may also be configured to form an echo-filtered audio stream by filtering out duplicate audio packets from the plurality of audio packets.

[0067] In various embodiments, the filtering of duplicate audio packets may be performed in various embodiments substantially the same as the filtering of a signal received from the external speaker of the same device in a microphone (in other words, as is known in the art).

[0068] The duplicate audio packets may be removed, and the remaining audio packets may be provided as an echo-filtered audio stream.

[0069] The processor 102 may also be configured to send the echo-filtered audio stream to at least one other device and / or to an audio device of the host device.

[0070] The processor 102 is also configured to send an unfiltered audio stream to at least one other device and / or to an audio device of the host device if no echo is detected.

[0071] The processor 102 is also configured to obtain information representing the physical location of each other device if an echo is detected.

[0072] As explained above, two different scenarios may similarly produce echoes, and to rule out the cause of the echoes, the physical location of each of the host device and the at least one additional device may be relevant.

[0073] The processor 102 is further configured to: based on the physical locations of the host device and each additional device, determine whether each additional device is in the same room as the host device.

[0074] The processor 102 may also be configured to: obtain information representing the activation state of the external speaker of each additional device that has been determined to be in the same room as the host device, and mute the external speaker of each additional device having an activation state of "on".

[0075] The processor 102 is further configured to: if an echo is detected, obtain information representing the activation state of the microphone of each additional device that has been determined not to be in the same room as the host device, and mute the microphone of each additional device having an activation state of "on".

[0076] Figure 2 An exemplary schematic diagram of an echo detector 200 according to various embodiments is shown. The echo detector 200 may provide functions similar to or the same as those of the echo detector 100.

[0077] The echo detector 200 may include: a transcription activator 220 for activating a transcription process that transcribes words generated from an audio transmission between the host device and at least one additional device into text; a text analyzer 222 for analyzing the text to identify repeated strings indicating an echo; and an echo filter 224 for decoding the audio transmission into a plurality of audio packets, filtering out duplicate audio packets from the plurality of audio packets to generate an echo-filtered audio stream, and sending the echo-filtered audio stream to at least one additional device and / or to an audio device of the host device.

[0078] Figure 3 An exemplary flowchart 300 for operating an echo detector device 100 and / or an echo detector 200 according to various embodiments is shown.

[0079] In various embodiments, a server (e.g., a host device) may request the locations of all users (conference room or remote location) while joining a conference call (310). A table may be created based on the input from each user.

[0080] In various embodiments, during a conference call, the server starts a transcript in the background (320). It will continuously monitor for redundant words / sentences to identify echoes.

[0081] When redundant words / echoes (330) are detected, the server will initiate an algorithm to filter out redundant audio packets (350) and stop sending the repeated words / echoes (i.e., the server sends the filtered audio packets to all receivers (360)).

[0082] In addition, the server will request the MIC and speaker configuration (mute or unmute) of each device participating in the call (340). All devices start polling the status of the MIC and speaker.

[0083] Once the filtering (at 350 and 360) is applied, the server will monitor the transcript again for repeated words (370). If the originator of the echo is present in the same conference room, there will be no redundant words.

[0084] If detected echoes still exist, based on the status information received at 310 and 340, the algorithm identifies all users in the conference room with an unmuted speaker.

[0085] The users in the conference room with an unmuted speaker will be notified to mute their speakers (or, the server mutes the PCs according to their permissions respectively) (390).

[0086] If redundant words are still detected after filtering (at 350) and sending the filtered audio packets (at 360), which indicates that the originator of the echo is a remote PC, the application will mute the microphone of the remote PC using an external speaker (380).

[0087] In addition to echo cancellation, the algorithm can be modified / supplemented to clean up the transcripts / meeting records captured during the meeting.

[0088] In the case where different Bluetooth audio devices are connected to the host device, separate digital-to-analog converters (DACs) will exist on the host device side and on the audio device (e.g., headphones) side.

[0089] The DACs in different audio devices (e.g., headphones) can be tuned to different gain values based on manufacturer preferences. Due to this, users will experience a drastic difference in volume levels when switching from one audio device to another.

[0090] For example, based on the gain set by the manufacturer, the volume may be too low or too high. Each time the user switches devices, the gain needs to be adjusted manually by the host. This will create an unpleasant experience for the user, especially during a conference call.

[0091] In various embodiments, methods and (host) devices are provided for normalizing volume levels across multiple audio devices (e.g., Bluetooth audio devices) without user intervention.

[0092] Figure 4 Exemplary schematic diagram showing a host device 400 according to various embodiments.

[0093] The host device 400 may include a processor 440.

[0094] The basic characteristics of the processor 440 may be similar to the basic characteristics of the processor 102.

[0095] The processor 440 may be configured to: store at least one look-up table including value pairs of gain to output power for a plurality of audio devices.

[0096] The gain value that can be adjusted by a user using a volume adjustment device (e.g., as a slider, rotatable knob, button, etc.) may determine the output power of an audio device (e.g., headphones), and the relationship between the gain value and the power output may be device-dependent.

[0097] Figure 6 Exemplary look-up table of a host device according to various embodiments is shown. The values for volume (given in %) represent gain and are listed for four different exemplary audio devices. The corresponding (output) power values are listed in the corresponding rows of the first column and are provided in dBA.

[0098] The processor 440 may be configured to: when switching from a first audio device operated by the host device to a second audio device to be operated by the host device, determine from at least one look-up table the output power value corresponding to the currently set gain value for the first audio device.

[0099] The processor 440 may also be configured to: determine from at least one look-up table the gain value of the second device corresponding to the determined output power value.

[0100] Thus, in the above example, if the second audio device to be operated by the host device is device 1, the resulting gain value of the second device will be 80%.

[0101] The processor 440 may also be configured to: set the gain value to the gain value determined for the second device.

[0102] Thus, in the above example, before activating the second device, the gain value of the second audio device will be set to 80%, and the volume perceived by the user will be the same as the volume for the first audio device.

[0103] The processor 440 may also be configured to: when a new audio device is connected to the host device 400, populate at least one look-up table with value pairs of gain to output power for the new audio device.

[0104] To populate a lookup table with output power values at a gain for a new audio device, the processor 440 can be configured to: set the gain to a sequence of different gain values (e.g., from a minimum value to a maximum value), and receive output power values from a value indicator of the new audio device.

[0105] The output power corresponding to the set volume level (gain) can be calculated using feedback received from the headset. Substantially, all headsets currently on the market support a volume indicator (VI) feedback mechanism to measure power output with respect to a specific gain.

[0106] In various embodiments, the output power numbers corresponding to each gain (volume level) can be passed to the host device via the BT sideband interface. Any reserved bits / registers in the BT protocol can be modified to pass volume indicator (VI) feedback to the host device 400.

[0107] In an exemplary embodiment, the user can use a Bluetooth headset at 60% volume. Referring Figure 6 to the table in, since this is a new device, the lookup table will populate the power level corresponding to this gain as -40 dBA, which can be the value transmitted by the volume indicator. The host device 400 can store the gain and the corresponding power until the new Bluetooth headset is connected. Optionally, additional gain values can be set, and additional corresponding output power values can be obtained for populating the lookup table.

[0108] Whenever the user connects a different Bluetooth headset, the host device 400 can check whether a lookup table for the new audio device already exists in the database (lookup table). If not, a test sequence will be played on the headset, and the power levels corresponding to each gain will be populated accordingly.

[0109] If an entry in the lookup table exists for the second device, the algorithm identifies the gain required to maintain the same power level set by the first device. If in the above example, the second audio device is Figure 6 device 3 in the table of, the gain requirement is only 30%. The volume level for the new Bluetooth device will be pre-adjusted to provide the user with a seamless experience without any abrupt volume level changes.

[0110] The mediator algorithm can be offloaded to a visual processing unit such as an Intel GNA / VPU, which can also be utilized by the central processing unit (CPU) of the host device 500 for the same purpose.

[0111] Figure 5 Exemplary schematic diagrams of the host device 500 according to various embodiments are shown.

[0112] The host device 500 may include: a buffer 500 for storing at least one lookup table, the lookup table including value pairs of gain to output power for a plurality of audio devices in the host device 500; an audio device controller 552 for instructing a switch from a first audio device operated by the host device 500 to a second audio device to be operated by the host device 500; and an output power controller 554 for determining, from the at least one lookup table, an output power value corresponding to a currently set gain value for the first audio device, for determining, from the at least one lookup table, a gain value for the second device corresponding to the determined output power value, and for setting the gain value to the gain value determined for the second device.

[0113] In various embodiments, multiple users may use the same host devices 400, 500. To identify the correct user, a user detection block may be introduced into the algorithm. For example, a camera that may be part of the host device or connected to the host device may be used in any case to identify the user during the training process and / or during use of the host device 500 connected to the audio device. Separate LUTs may be created for different users.

[0114] Figure 7 An exemplary flowchart 700 for operating a host device for adjusting output power in accordance with various embodiments is shown. The host device 400 and / or the host device 500 may be configured to perform the processes indicated in flowchart 700.

[0115] In various embodiments, when using the host devices 400, 500 to operate an audio device, a new Bluetooth device connection may be detected (710). Optionally, user identification may be performed (720).

[0116] Subsequently, a lookup table may be consulted to determine whether the audio device (e.g., headphones) has been registered (730) (e.g., the MAC ID of the audio device may be used for identification, for example) and whether the lookup table has been populated for the device (760).

[0117] If the audio device is not registered or the lookup table has not been populated for the device, the lookup table will be populated (740) with value pairs of gain to output power and stored in the lookup table (750).

[0118] Regardless of whether the lookup table has just been populated or value entries previously existed, the power output of the previous Bluetooth device is retched from the lookup table (770).

[0119] Subsequently, the gain / volume for the new audio device corresponding to the previous output power is set (780).

[0120] When the audio device is disconnected (790), the host device can store the power level (799) before the disconnection.

[0121] The current state-of-the-art process for device switching during an active audio connection is described below:

[0122] During a voice call or a data call (e.g., Teams or Zoom, etc.), it is possible for the audio device to switch from one output audio device to another output audio device.

[0123] During the switch, the audio output presented on the speaker can stop, and when the connection is ready, the second audio device (e.g., a Bluetooth headset) can be selected as the output device.

[0124] There is a short duration for the channel setup for sending data over the Bluetooth link, which occurs when the audio switch happens, and during the channel setup, the audio signal still cannot be sent over the Bluetooth link.

[0125] This means that audio loss of about 2 seconds may typically occur during the switching process.

[0126] This means that there is a need to improve the processing of voice sampling to avoid or mitigate the lost audio samples during the audio device switch, especially in latency-critical applications.

[0127] In various embodiments, methods and devices are provided to achieve lossless audio output during output device switching.

[0128] In various embodiments, when the audio playback device is not yet available, audio (e.g., voice) data can be recorded in the last few seconds, and the audio data can be presented using an enhanced sampling algorithm, thereby enabling the maintenance of latency.

[0129] Figure 8 An exemplary schematic diagram of the host device 800 according to various embodiments is shown.

[0130] The host device 800 can include a processor 880. The basic characteristics of the processor 880 can be similar to the basic characteristics of the processor 102 and / or 880.

[0131] The processor 880 can be configured in various embodiments to buffer the incoming audio stream in the host device 80 (e.g., in a buffer that can be included in the host device 800).

[0132] The processor 880 can also be configured to: switch from a first audio device connected to the host device 800 to a second audio device (e.g., from a built-in speaker to a headset, or from a first headset to a second headset).

[0133] The processor 880 may also be configured in various embodiments to sample the buffered audio stream at a first sampling rate to form a first sampled audio stream; and to sample the real-time incoming audio stream at a second sampling rate higher than the first sampling rate.

[0134] The processor 880 may also be configured to send the first sampled audio stream to a second audio device.

[0135] Sampling the buffered audio stream at a sampling rate lower than the commonly used sampling rate may allow the transmitted audio data that has been generated from the stored data to be transmitted faster than the real-time audio data and ultimately catch up with the real-time audio data. This will also be explained in more detail in the context of the following with other drawings.

[0136] The processor 880 may also be configured in various embodiments to switch from sending the first sampled audio stream to sending the second sampled audio stream when the transmitted first sampled audio stream corresponds to the real-time incoming audio stream (in other words, when the stored audio stream has caught up with the real-time audio stream).

[0137] Buffering may be initiated in various embodiments only when needed (e.g., upon receiving an instruction to switch from a first audio device to a second audio device).

[0138] Alternatively, buffering may be performed continuously (or more precisely, continuously while providing the incoming audio stream). In this case, sampling the first audio stream at the first sampling rate may be performed only when needed (e.g., upon receiving an instruction to switch from a first audio device to a second audio device).

[0139] Buffering may include, for example, storing the audio stream as raw pulse code modulation (PCM) samples in a buffer.

[0140] For example, a suitable ratio of the first sampling rate to the second sampling rate may be 0.5. In other words, the buffered audio stream may be sampled at half the sampling rate of the real-time audio stream. As an example, the first sampling rate (also referred to as the sampling frequency) may be, for example, 16K, while the second sampling rate may be, for example, 32K.

[0141] More generally, the second sampling rate may be between approximately 1.33 times and approximately 4 times the first sampling rate. The first sampling rate and the second sampling rate may be user-configurable parameters or may be predefined by the manufacturer, for example.

[0142] The connection between the host device 800 and each of the first and second audio devices can be a synchronous audio connection (e.g., a Bluetooth connection) in various embodiments. In various embodiments, only the second audio device can be connected via a synchronous audio connection, while the first audio device can be connected via a different type of audio connection. For example, the first device can be integrated into the host device 800 or can be connected via a wired connection.

[0143] The synchronous audio connection can be configured to send audio packets of a predefined size at predefined intervals. Thus, if audio packets are generated from an audio stream sampled at a lower rate (e.g., half of the normal sampling rate), more audio data (e.g., twice as much) can be sent than normal.

[0144] Sending a buffered audio stream faster than a real-time audio stream would be sent can enable the buffered audio stream to "catch up" with the real-time audio stream. Thus, when the timestamp of the buffered audio stream coincides with the real time (which may occur within a few seconds (e.g., between 1 and 5 seconds) of the transmission of the buffered audio stream), the audio transmission can switch from the buffered audio stream to the real-time audio stream sampled at a second sampling rate that is the normal sampling rate.

[0145] Thereby, it can be achieved that all audio data can be sent even when a buffered audio stream is sent (in other words, no audio data is lost), and the quality of the audio transmission may degrade slightly. Further, after a short duration of sending the buffered audio stream, the transmission of the real-time audio stream is resumed, so as to avoid latency.

[0146] If the ratio of the first sampling rate to the second sampling rate is user-configurable, this can also allow the user to determine based on their needs whether to set their preference for reaching a latency-free state as early as possible (the lower the sampling rate, the faster the transmission of the buffered data is completed and thus the real-time audio stream transmission is resumed) or to maintain a higher audio transmission quality even when sending a buffered audio stream (the closer the first sampling rate is to the second sampling rate, the better the audio quality).

[0147] Figure 11 List parameters that can be predefined, for example, by the manufacturer of the host device 800 and / or can be user-configurable.

[0148] Figure 9 Exemplary schematic diagram of a host device 900 according to various embodiments.

[0149] The host device 900 can include a buffer 990 for buffering (e.g., temporarily storing) incoming audio streams in the host device 900.

[0150] The host device 900 may further include an audio device controller for switching from a first audio device connected to the host device to a second audio device. Each of the connected audio devices may be connected via a wireless connection (e.g., a synchronous wireless connection such as a Bluetooth connection).

[0151] The host device 900 may further include a sampling circuit for sampling a buffered audio stream at a first sampling rate to form a first sampled audio stream and for sampling a real-time incoming audio stream at a second sampling rate higher than the first sampling rate.

[0152] The sampling process itself may generally be performed as is known in the art.

[0153] The host device 900 may further include a transmitter for sending the first sampled audio stream to the second audio device and for switching from sending the first sampled audio stream to the second audio device to sending a second sampled audio stream when the sent first sampled audio stream corresponds to a real-time incoming audio stream (in other words, when the timestamp of the audio stream has reached real time).

[0154] Figure 10 Exemplary schematic diagram of a host device 1000 according to various embodiments showing the switching of an audio connection from a first audio device 1040 to a second audio device 1050.

[0155] Initially, a first audio device 1040 that may be connected to or integrated in the host device 1000 may be used for audio output. For example, the first audio device 1040 may be a built-in external speaker of a PC, a laptop device, or a smartphone.

[0156] The user may initiate a switch to a second audio device 1050 (e.g., wireless headphones to be connected to the host device 1000).

[0157] During the switching process, for example, audio data (e.g., an incoming audio stream sent to the host device 1000 via a microphone 1010 or received in the host device 1000) may be recorded / buffered during the last few seconds when the second audio device 1050 to be used for audio playback is still unavailable. The audio module 1020 (e.g., an audio digital signal processor (audio DSP)) may receive the incoming audio stream and optionally buffer the incoming audio stream.

[0158] Once an audio channel (e.g., a Bluetooth channel) is set up between the host device 1000 and the second audio device 1050, the Bluetooth controller 1030, which can be part of the host device 1000, can instruct the audio module 1020 to transmit the buffered / recorded audio stream through the channel (e.g., connection) between the Bluetooth controller 1030 and the audio module 1020. This channel can be, for example, a SoundWire, Slimbus, or I2S channel.

[0159] The audio module 1020 can provide the buffered / recorded audio stream through the channel together with the real-time audio stream. The audio data channel should be able to handle both audio streams.

[0160] From that point in time, and as long as the (portion of the) buffered audio stream exists in the Bluetooth controller 1030, the audio stream is presented at a lower quality (first sampling rate) to allow the audio data to be sent at a faster (e.g., double) rate.

[0161] For example, the Bluetooth channel can be configured for a 32K sampling frequency. During an initial period after the Bluetooth channel to the second audio device 1050 is established, the data (i.e., the buffered incoming audio stream, and thus the last recorded audio stream) will be sampled at, for example, 16K and sent to the second audio device 1050, which can act as a Bluetooth playback device, at an increased speed (e.g., double rate speed).

[0162] Due to the lower sampling rate, the buffered / recorded data will catch up to the current moment. Once all the buffered audio data is sent, the sampling rate switches back to the original sampling rate. In this example, it switches to 32K sampling. Thus, the audio data missed during real-time transmission can be presented to the second audio device (playback device) without a compromise in terms of audio data loss or latency.

[0163] The audio quality and rate may vary slightly for a few seconds during the playback of the buffered / recorded audio stream, but this should not have a significant impact on the listening experience. The audio playback speed of the recorded audio can be faster (e.g., twice the speed of the original audio), which allows it to catch up to the real-time audio.

[0164] Figure 12 The operation of the host device 1000 according to various embodiments is shown in the schematic illustration 1200.

[0165] As time progresses (from top to bottom), various stages of processing and / or interaction are indicated for the components of the host device 1000, particularly the audio module 1020 and the Bluetooth controller 1030, and the (second) audio device 1050, which can be, for example, a Bluetooth headset.

[0166] In the upper left, the detection stage is indicated.

[0167] During this stage, the audio module 1020 ("audio DSP") can detect a new device ((second) audio device 1050) based on a device control interrupt (DCI) made by the HLOS (High-Level OS).

[0168] The audio module 1020 can include a so-called "T-filter optimizer (abbreviation: T-filter)" 1216.

[0169] The T-filter 1216 can include a buffer for buffering incoming audio streams. The buffer / storage volume available in the T-filter 1216 can be required to be minimal as it should not affect the memory. For example, a buffer size of 2KB should be sufficient to store approximately 1 second of audio data (raw pcm data). Figure 11 The configurable parameters can be relevant for the configuration of the T-filter 1216.

[0170] The T-filter 1216 can be enabled to store the last few seconds of the incoming audio stream while delivering the data to the output audio device (e.g., the first audio device, before it is disabled regarding the switch to the second audio device 1050). The T-filter can be activated whenever an active audio connection exists. This storage can operate as a backup copy of the audio stream available for feedback to the Bluetooth controller 1030 upon request.

[0171] When an active connection to the audio device 1050 exists, the T-filter 1216 in the audio module 1020 can be operative to always store the last few seconds of the incoming audio stream as raw PCM samples.

[0172] This application can be not limited to incoming audio streams for playback in the audio device 1050, but rather can become a general implementation for many extensible use cases around the missing audio samples.

[0173] The subsequent stage can be related to presenting switched-timeframe audio.

[0174] This stage can handle the communication between the audio module 102 and the Bluetooth controller 1030 and the enhancement that can be applied to the (buffered) audio stream to present the missing data.

[0175] The following provides a detailed description of how the Bluetooth controller 1030 obtains the buffered / stored (also called "historical") audio stream from the T-filter 1216 and how the Bluetooth controller 1030 delivers the audio stream to the (second) audio device 1050 without audio loss.

[0176] Once the Bluetooth channel is set up and ready, the Bluetooth controller 1030 can request from the audio module 1020 (audio DSP) to send any audio samples that were missed before the channel was set up.

[0177] Now, the content of the history buffer, which is part of the Bluetooth controller and used for audio data transmission to the (second) audio device 1050 (e.g., headphones), can be requested to be transmitted at a faster rate to achieve a state where real-time audio data is again transmitted / broadcast by the audio device 1050 without latency.

[0178] However, since the Bluetooth connection is a synchronous channel, the history buffer cannot send audio samples to the audio device 1050 in batches.

[0179] To achieve a task with a fixed synchronization interval and size for the packets, the Bluetooth controller 1030 can compromise on the quality of the audio samples by downsampling the rate.

[0180] In the example, as Figure 12 indicated, the sampling rate is typically assumed to be 32K samples per second.

[0181] The buffered audio stream ("history packets") can be downsampled, for example, by a factor of 2, so that the buffered audio stream is sampled at a (first) sampling rate of 16K samples per second.

[0182] The buffered audio stream sampled at the first sampling rate can then be encoded.

[0183] This allows more real-time data (or: more buffered data that was once real-time data) to fit within the given packet size.

[0184] The rendering / broadcasting of the downsampled audio stream from the buffer can occur at a higher (e.g., double) speed compared to the normal real-time speed.

[0185] When the buffered audio stream is broadcast by the audio device 1050 at an increased speed, more audio data is added at the source. In other words, during the broadcast of the buffered audio stream, the real-time incoming audio stream can still be active.

[0186] However, due to the increased speed, the buffered audio stream will catch up with the real-time audio stream at some point (e.g., after a few seconds).

[0187] Once the history buffer is completed / emptied (e.g., the timestamp of the last sent audio packet corresponds to the real time), the algorithm for downsampling the audio stream can stop.

[0188] Therefore, the original audio quality is restored (without latency, or with only the normal latency).

[0189] Figure 13 Exemplary flowchart 1300 showing a method for handling audio transmission between a host device and at least one additional device according to various embodiments.

[0190] The method, which may be executed, for example, by processor 880, may include: buffering an incoming audio stream in the host device (1310); switching from a first audio device connected to the host device to a second audio device (1320); sampling the buffered audio stream at a first sampling rate to form a first sampled audio stream (1330); sampling a real-time incoming audio stream at a second sampling rate higher than the first sampling rate to form a second sampled audio stream (1340); and sending the first sampled audio stream to the second audio device (1350).

[0191] The above audio devices may include wireless audio devices and may be configured for short-range mobile radio communication. The audio devices (and corresponding host devices) may include a wireless interface (e.g., a Bluetooth interface (e.g., a Bluetooth Low Energy (LE) interface), ZigBee, Z-Wave, Wi-Fi HaLow / IEEE 802.11ah, etc.). By way of example, one or more of the following Bluetooth interfaces may be provided: Bluetooth V1.0A / 1.0B interface, Bluetooth V1.1 interface, Bluetooth V1.2 interface, Bluetooth V 2.0 interface (optionally, plus EDR (Enhanced Data Rate)), Bluetooth V 2.1 interface (optionally, plus EDR (Enhanced Data Rate)), Bluetooth V 3.0 interface, Bluetooth V 4.0 interface, Bluetooth V 4.1 interface, Bluetooth V 4.2 interface, Bluetooth V 5.0 interface, Bluetooth V 5.1 interface, Bluetooth V 5.2 interface, etc.

[0192] Generally, a computer-readable medium may be a floppy disk, a hard disk, a USB (Universal Serial Bus) storage device, a RAM (Random Access Memory), a ROM (Read-Only Memory), an EPROM (Erasable Programmable Read-Only Memory), or a flash memory. A computer-readable medium may also be a data communication network (e.g., the Internet) that permits downloading of program code. A computer-readable medium may be a non-transitory or transitory medium.

[0193] Examples

[0194] The examples set forth herein are illustrative and not exhaustive.

[0195] Example 1 is an echo detection device, comprising: a processor configured to: activate a transcription process that transcribes words generated from an audio transmission between a host device and at least one other device into text; analyze the text to identify a repeating string indicative of an echo; and if an echo is detected, then: decode the audio transmission into a plurality of audio packets; filter out duplicate audio packets from the plurality of audio packets to generate an echo-filtered audio stream; and send the echo-filtered audio stream to the audio device of the at least one other device and / or the host device; and optionally, a memory storing the text from the transcription process thereon.

[0196] In Example 2, the subject matter as described in Example 1 may optionally include: the processor is further configured to: if no echo is detected, send the audio transmission between the host device and the at least one other device without filtering (e.g., send an unfiltered audio stream to the audio device of the at least one other device and / or the host device).

[0197] In Example 3, the subject matter as described in Example 1 or 2 may optionally include: the repeating string indicative of an echo includes a second identical string of at least two characters immediately preceding the first string.

[0198] In Example 11, the subject matter as described in Example 10 may optionally include: the string includes a single word, a word part, or a combination of words.

[0199] In Example 5, the subject matter as described in any one of Examples 1 to 4 may optionally include: the processor is further configured to: if an echo is detected, obtain information representing the physical location of each other device (i.e., each of the at least one other device); based on the physical locations of the host device and each other device, determine whether each other device is in the same room as the host device; obtain information representing the activation state of the external speaker of each other device that has been determined to be in the same room as the host device; and mute the external speaker of each other device having an activation state of "on".

[0200] In Example 6, the subject matter as described in any one of Examples 1 to 5 may optionally further include: if the echo is detected, obtain information representing the physical location of each other device (i.e., each of the at least one other device); based on the physical locations of the host device and each other device, determine whether each other device is in the same room as the host device; obtain information representing the activation state of the microphone of each other device that has been determined not to be in the same room as the host device; and mute the microphone of each other device having an activation state of "on".

[0201] Example 7 is an echo detector. The echo detector may include: a transcription activator for activating a transcription process that transcribes words generated from an audio transmission between the host device and the at least one additional device into text; a text analyzer for analyzing the text to identify a repeating string indicative of an echo; and an echo filter for: decoding the audio transmission into a plurality of audio packets; filtering out duplicate audio packets from the plurality of audio packets to generate an echo-filtered audio stream, and sending the echo-filtered audio stream to the audio device of the at least one additional device and / or the host device.

[0202] Example 8 is a non-transitory computer-readable medium storing instructions that, when executed by a processor, cause a host device connected to at least one additional device via a communication interface to: activate a transcript process that transcribes words generated from an audio transmission between the host device and the at least one additional device into text; analyze the text to identify a repeating string indicative of an echo; if an echo is identified, decode the audio transmission into a plurality of audio packets; form an echo-filtered audio stream by filtering out duplicate audio packets from the plurality of audio packets, and send the echo-filtered audio stream to the audio device of the at least one additional device and / or the host device.

[0203] Example 9 is a host device. The host device may include: a processor configured to: store at least one lookup table including value pairs of gain to output power for a plurality of audio devices; when switching from a first audio device operated by the host device to a second audio device to be operated by the host device: (e.g., from system settings) determine a currently set gain value; determine, from the at least one lookup table, an output power value corresponding to the currently set gain value for the first audio device; determine, from the at least one lookup table, a gain value for the second audio device corresponding to the determined output power value; and set the gain value to the gain value determined for the second audio device; and optionally, a memory storing the at least one lookup table thereon.

[0204] In Example 10, the subject matter as described in Example 9 may optionally include: the processor further configured to: populate the at least one lookup table with value pairs of gain to output power for the new audio device when a new audio device is connected to the host device.

[0205] In Example 11, the subject matter as described in Example 10 may optionally include: To populate the at least one look-up table with output power values for the gain for the new audio device, the processor is configured to: set the gain to a sequence of different gain values and receive output power values from a value indicator of the new audio device.

[0206] In Example 12, the subject matter as described in Example 11 may optionally include: receiving the output power value from the new audio device via a Bluetooth sideband interface.

[0207] In Example 13, the subject matter as described in any one of Examples 9 to 12 may optionally include: the at least one look-up table is user-specific, and the processor is further configured to: identify a user of the second audio device before determining an output power value corresponding to a currently set gain value for the first audio device.

[0208] Example 14 is a host device. The host device may include: a buffer for storing at least one look-up table, the at least one look-up table including value pairs of gain to output power for a plurality of audio devices; an audio device controller for instructing a switch from a first audio device operated by the host device to a second audio device to be operated by the host device; and an output power controller for: determining a currently set gain value; determining, from the at least one look-up table, an output power value corresponding to the currently set gain value for the first audio device; determining, from the at least one look-up table, a gain value of the second device corresponding to the determined output power value; and setting the gain value to the gain value determined for the second audio device.

[0209] Example 15 is a non-transitory computer-readable medium including instructions that, when executed by a processor, cause a host device connected to an audio device to: store at least one look-up table, the at least one look-up table including value pairs of gain to output power for a plurality of audio devices; when switching from a first audio device operated by the host device to a second audio device to be operated by the host device: (e.g., from system settings) determine a currently set gain value; determine, from the at least one look-up table, an output power value corresponding to the currently set gain value for the first audio device; determine, from the at least one look-up table, a gain value of the second device corresponding to the determined output power value; and set the gain value to the gain value determined for the second audio device.

[0210] Example 16 is a host device. The host device may include: a processor configured to: buffer an incoming audio stream in the host device to form a buffered audio stream; switch from a first audio device connected to the host device to a second audio device; sample the buffered audio stream at a first sampling rate to form a first sampled audio stream; sample a real-time incoming audio stream at a second sampling rate higher than the first sampling rate to form a second sampled audio stream; and send the first sampled audio stream to the second audio device.

[0211] In Example 17, the subject matter as described in Example 16 may optionally include: the processor is further configured to: when the timestamp of the sent first sampled audio stream corresponds to real time, switch from sending the first sampled audio stream to the second audio device to sending the second sampled audio stream.

[0212] In Example 18, the subject matter as described in Example 16 or 17 may optionally include: the processor is further configured to: initiate the buffering upon receiving an instruction to switch from the first audio device to the second audio device.

[0213] In Example 19, the subject matter as described in Example 16 or 17 may optionally include: continuously (e.g., continuously while receiving the incoming audio stream) perform the buffering.

[0214] In Example 20, the subject matter as described in any one of Examples 16 to 19 may optionally include: the buffering includes: storing the incoming audio stream as raw pulse code modulation (PCM) samples.

[0215] In Example 21, the subject matter as described in any one of Examples 16 to 20 may optionally include: the second sampling rate is between approximately 1.33 and approximately 4 times the first sampling rate.

[0216] In Example 22, the subject matter as described in any one of Examples 16 to 21 may optionally include: the first sampling rate and the second sampling rate are user-configurable parameters.

[0217] In Example 23, the subject matter as described in any one of Examples 16 to 22 may optionally include: the connection is a synchronous audio connection.

[0218] Example 24 is a host device. The host device may include: a buffer for buffering an incoming audio stream in the host device; an audio device controller for switching from a first audio device connected to the host device to a second audio device; a sampling circuit for sampling the buffered audio stream at a first sampling rate to form a first sampled audio stream and for sampling a real-time incoming audio stream at a second sampling rate higher than the first sampling rate; and a transmitter for sending the first sampled audio stream to the second audio device and for switching from sending the first sampled audio stream to the second audio device to sending the second sampled audio stream when the first sampled audio stream being sent corresponds to the real-time incoming audio stream; and optionally, a memory on which the buffered audio stream is stored.

[0219] Example 25 is a non-transitory computer-readable medium storing instructions that, when executed by a processor, cause a host device connected to an audio device to: buffer an incoming audio stream in the host device; switch from a first audio device connected to the host device to a second audio device; sample the buffered audio stream at a first sampling rate to form a first sampled audio stream; sample a real-time incoming audio stream at a second sampling rate higher than the first sampling rate; send the first sampled audio stream to the second audio device; and switch from sending the first sampled audio stream to the second audio device to sending the second sampled audio stream when the first sampled audio stream being sent corresponds to the real-time transmitted audio stream.

[0220] Example 26 is a method for handling audio transmission between a host device and at least one other device, the method including: activating a transcription process that transcribes words generated from the audio transmission between the host device and the at least one other device into text; analyzing the text for identifying a repeating string indicative of an echo; and if an echo is detected, then: decoding the audio transmission into a plurality of audio packets; filtering out duplicate audio packets from the plurality of audio packets to generate an echo-filtered audio stream, and sending the echo-filtered audio stream to the at least one other device and / or an audio device of the host device.

[0221] Example 27 is a method for handling audio transmission between a host device and at least one other device, the method comprising: storing at least one lookup table, the at least one lookup table including value pairs of gain to output power for a plurality of audio devices; indicating a switch from a first audio device operated by the host device to a second audio device to be operated by the host device; determining, from the at least one lookup table, an output power value corresponding to a currently set gain value for the first audio device; determining, from the at least one lookup table, a gain value for the second device corresponding to the determined output power value; and setting the gain value to the gain value determined for the second audio device.

[0222] Example 27 is a method for handling audio transmission between a host device and at least one other device, the method comprising: buffering an incoming audio stream in the host device; switching from a first audio device connected to the host device to a second audio device; sampling the buffered audio stream at a first sampling rate to form a first sampled audio stream; sampling a real-time incoming audio stream at a second sampling rate higher than the first sampling rate; sending the first sampled audio stream to the second audio device; and when the sent first sampled audio stream corresponds to the real-time transmitted audio stream, switching from sending the first sampled audio stream to the second audio device to sending the second sampled audio stream.

[0223] As used herein, the term "exemplary" is used to mean "serving as an example, instance, or illustration". Any example or design described herein as "exemplary" is not necessarily to be construed as superior to or better than other examples or designs.

[0224] The words "plurality" and "several" in the specification and claims clearly refer to a number greater than 1. The terms "group", "[set]", "[collection]", "[series]", "[sequence]", "[grouping]", etc. in the specification or claims refer to a number equal to or greater than 1 (i.e., one or more). Any term in the plural form that does not explicitly state "plurality" or "several" also refers to a number equal to or greater than 1.

[0225] For example, the terms "processor" or "controller" as used herein can be understood as any kind of technical entity that allows for the processing of data. The data can be processed in accordance with one or more specific functions performed by the processor or controller. Additionally, the processor or controller as used herein can be understood as any kind of circuit (e.g., any kind of analog or digital circuit). Thus, a processor or controller can be or include an analog circuit, a digital circuit, a mixed-signal circuit, a logic circuit, a processor, a microprocessor, a central processing unit (CPU), a graphics processing unit (GPU), a digital signal processor (DSP), a field-programmable gate array (FPGA), an integrated circuit, an application-specific integrated circuit (ASIC), etc., or any combination thereof. Any other kind of implementation of the corresponding functions can also be understood as a processor, controller, or logic circuit. It should be understood that any two (or more) processors, controllers, or logic circuits detailed herein can be implemented as a single entity having equivalent or similar functions, and conversely, any single processor, controller, or logic circuit detailed herein can be implemented as two (or more) separate entities having equivalent or similar functions.

[0226] The term "connected" can be understood in the sense of (e.g., mechanical and / or electrical) connection and / or interaction, for example, directly or indirectly. For example, several elements can be mechanically connected together such that they are physically held (e.g., a plug connected to a socket), and electrically such that they have a conductive path (e.g., a signal path exists along a communication link).

[0227] Although the above description and the related drawings may depict the components of an electronic device as separate elements, those skilled in the art should understand the various possibilities of combining or integrating discrete elements into a single element. This can include: combining two or more components from a single component, mounting two or more components onto a common chassis to form an integrated component, executing discrete software components on a common processor core, etc. Conversely, those skilled in the art will recognize the possibility of separating a single element into two or more discrete elements (e.g., splitting a single component into two or more separate components, separating a chip or chassis into the discrete elements originally provided thereon, separating a software component into two or more segments and executing each segment on a separate processor core, etc.). Additionally, it should be understood that the specific implementation of the hardware and / or software components is merely illustrative, and other combinations of hardware and / or software that perform the methods described herein are within the scope of the present disclosure.

[0228] It should be understood that the implementation of the methods detailed herein is exemplary in nature and is thus understood to be capable of being implemented in a corresponding device. Similarly, it should be understood that the implementation of the devices detailed herein is understood to be capable of being implemented as a corresponding method. Thus, it should be understood that a device corresponding to the methods detailed herein can include one or more components configured to perform each aspect of the related methods.

[0229] All of the acronyms defined in the above description also apply to all claims included herein.

[0230] Although the present disclosure has been particularly shown and described with reference to specific embodiments, those skilled in the art will understand that various changes in form and detail may be made therein without departing from the spirit and scope of the present disclosure as defined by the appended claims. The scope of the present disclosure is thus indicated by the appended claims and is therefore intended to cover all changes coming within the meaning and range of equivalents of the claims.

Claims

1. An echo detection device, comprising: The processor is configured as: activating a transcription process that transcribes words generated from audio transmission between the host device and at least one additional device into text; analyzing the text for identifying repeated character strings indicative of echoes; as well as If an echo is detected, then: decoding the audio transmission into a plurality of audio packets; filtering out identical audio packets from the plurality of audio packets to generate an echo-filtered audio stream; as well as sending the echo-filtered audio stream to an audio device of the at least one further device and / or the host device; and A memory on which text from the transcription process is stored.

2. The echo detection device according to claim 1, wherein: The processor is further configured to: If no echo is detected, the audio transmission is sent between the host device and the at least one additional device without filtering.

3. The echo detection device according to claim 1, wherein: The repeated character string indicating the echo includes a second identical character string of at least two characters immediately preceding the first character string.

4. The echo detection device according to claim 1, wherein: A string of characters may consist of individual words, parts of words, or combinations of words.

5. The echo detection device according to any one of claims 1 to 4, wherein: The processor is further configured to: If an echo is detected, then: obtaining information representing the physical location of each of the additional devices; determining, based on the physical location of the host device and each additional device, whether each additional device is in the same room as the host device; obtaining information indicative of an activation state of an external speaker of each additional device that has been determined to be in the same room as the host device; as well as Mutes the external speakers of each additional device that has the active status "On".

6. The echo detection device according to any one of claims 1 to 4, further comprising: If an echo is detected, then: obtaining information representing the physical location of each of the additional devices; determining, based on the physical location of the host device and each additional device, whether each additional device is in the same room as the host device; obtaining information indicative of an activation state of a microphone of each additional device that has been determined not to be in the same room as the host device; as well as Mute the microphone of each additional device that has the activation status "On".

7. A host device, comprising: The processor is configured as: storing at least one lookup table including pairs of gains versus output powers for a plurality of audio devices; When switching from a first audio device operated by the host device to a second audio device to be operated by the host device: Determine the currently set gain value; determining, from the at least one lookup table, an output power value for the first audio device corresponding to a currently set gain value; determining, from the at least one lookup table, a gain value for the second audio device corresponding to the determined output power value; as well as setting the gain value to a gain value determined for the second audio device; and A memory on which the at least one look-up table is stored.

8. The host device according to claim 7, wherein: The processor is further configured to, upon connecting a new audio device to the host device, populate the at least one lookup table with a gain versus output power value pair for the new audio device.

9. The host device according to claim 8, wherein: To populate the at least one lookup table with value pairs of gain versus output power for the new audio device, the processor is configured to: The gain is set to a sequence of different gain values ​​and an output power value is received from the value indicator of the new audio device.

10. The host device of claim 9, wherein: The output power value is received from the new audio device via a Bluetooth sideband interface.

11. The host device according to any one of claims 7 to 10, wherein: The at least one lookup table is user-specific, and the processor is further configured to: Prior to determining an output power value for the first audio device corresponding to the currently set gain value, a user of the second audio device is identified.

12. A host device, comprising: The processor is configured as: buffering an incoming audio stream in the host device; switching from a first audio device connected to the host device to a second audio device; Sampling the buffered audio stream at a first sampling rate to form a first sampled audio stream; sampling the real-time incoming audio stream at a second sampling rate higher than the first sampling rate to form a second sampled audio stream; as well as Sending the first sampled audio stream to the second audio device; and A memory on which to store the buffered audio stream.

13. The host device of claim 12, wherein: The processor is further configured to: When the timestamp of the transmitted first sampled audio stream corresponds to real time, the transmission of the first sampled audio stream to the second audio device is switched to the transmission of the second sampled audio stream.

14. The host device of claim 12, wherein: The processor is further configured to: The buffering is initiated upon receiving an instruction to switch from the first audio device to the second audio device.

15. The host device of claim 12, wherein: The buffering is performed continuously.

16. The host device of claim 12, wherein: The buffer comprises: The incoming audio stream is stored as raw pulse code modulation (PCM) samples.

17. The host device of claim 12, wherein: The second sampling rate is between about 1.33 and about 4 times the first sampling rate.

18. The host device according to claim 12, in, The first sampling rate and the second sampling rate are user-configurable parameters.

19. The host device according to any one of claims 12 to 18, in, The connection is a synchronous audio connection.

20. The host device according to any one of claims 12 to 18, in, The connection is a Bluetooth connection.

21. A non-transitory computer readable medium storing instructions that, when executed by a processor, cause a host device connected to an audio device to: buffering an incoming audio stream in the host device; switching from a first audio device connected to the host device to a second audio device; Sampling the buffered audio stream at a first sampling rate to form a first sampled audio stream; sampling the real-time incoming audio stream at a second sampling rate higher than the first sampling rate; sending the first sampled audio stream to the second audio device; as well as When the transmitted first sampled audio stream corresponds to the real-time incoming audio stream, switching is performed from transmitting the first sampled audio stream to the second audio device to transmitting the second sampled audio stream.

22. A method of processing audio transmission between a host device and at least one further device, the method comprising: activating a transcription process that transcribes words generated from audio transmission between the host device and the at least one additional device into text; analyzing the text for identifying repeated character strings indicative of echoes; as well as If an echo is detected, then: decoding the audio transmission into a plurality of audio packets; filtering out identical audio packets from the plurality of audio packets to generate an echo-filtered audio stream; as well as The echo-filtered audio stream is sent to the at least one further device and / or an audio device of the host device.