Communication between Networked Audio Devices
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- SHURE ACQUISITION HLDG INC
- Filing Date
- 2023-07-21
- Publication Date
- 2026-07-23
AI Technical Summary
Existing audio systems requiring high-performance audio with low latency are costly due to complex circuitry, expensive components, and high network bandwidth usage, making them unsuitable for applications with relaxed quality and latency requirements, such as voice conferencing.
Implementing local asynchronous media clocks in audio devices using low-precision oscillators like crystal oscillators, MEMS, ceramic resonators, or SAW oscillators, which are independent of network clocks, to simplify audio processing and reduce costs while maintaining audio quality and latency within acceptable ranges.
This approach allows for cost-effective, high-quality audio transmission and reception in audio devices by reducing complexity and cost without sacrificing performance in applications like video conferencing and public address systems.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to communication between networked audio devices (Cross - reference to related applications) This application claims priority to U.S. Patent Application Serial No. 18 / 223,766, filed July 19, 2023, which claims priority to U.S. Provisional Patent Application Serial No. 63 / 391,061, filed July 21, 2022, and each of these applications is hereby incorporated by reference in its entirety.
Background Art
[0002] Some audio systems are designed to provide high - performance audio using Internet Protocol (IP) functionality. For example, the AES - 67 standard provides a way for audio device manufacturers to interoperate in IP audio solutions that transmit and receive professional - quality audio, such as 24 - bit uncompressed PCM audio sampled at 48 kilohertz (KHz), with extremely low latency. Such professional quality and low latency can be costly. In particular, such systems involve relatively complex circuitry and / or software, a number of expensive components, high network bandwidth usage, and other costs.
Summary of the Invention
[0003] The following summary simplifies certain features. The summary is not an extensive overview nor is it intended to identify key or decisive elements.
[0004] While the above costs of providing professional-quality audio with extremely low latency can be valuable in certain situations, in some market segments, the required audio quality and latency requirements may not be so high or may have relaxed restrictions. Therefore, the cost may be excessive relative to the perceived benefits. One example is in voice conferencing spaces. The audio needs in an average voice-grade video teleconference may be different from the more stringent needs in a major pop star's tour business. Off-the-shelf solutions such as standards-based Voice over IP (VoIP) may not provide the level of fidelity that users expect. These standards are often telephone-grade and may comply with IT or Internet standards. Additionally, precise phase accuracy and sample accuracy in audio transmission typically require designing high-precision clocking hardware modules in audio products. Due to various reasons (e.g., the ongoing current semiconductor shortage and / or supply chain or other economic conditions), these types of hardware modules can sometimes be extremely difficult to obtain or extremely expensive to acquire. It is desirable to provide an audio solution that is positioned midway between high-performance audio solutions and telephone-grade or IT-based solutions and does not incur inappropriate design and construction costs for the required levels of audio quality and latency. Such a solution may not utilize any precise clocking hardware modules while still meeting the needs of various audio contexts (e.g., video conferencing, video meetings, public address systems, etc.). For example, such a solution can selectively and carefully relax one or more of the bandwidth, quality, and / or latency requirements that existing high-performance audio over IP solutions require. Such a solution can potentially continue to offer the best-in-class audio performance while being done at a much lower implementation, complexity, and / or usage cost.
[0005] For example, in some high-performance systems, the network clock is closely synchronized with the global network clock, and other internal clocks are derived from the closely synchronized network clock. This is typically achieved using complex and expensive, accurate hardware clocking modules. Such a highly synchronized clocking mechanism is used to control analog-to-digital and digital-to-analog conversion processes, as well as packetization and depacketization of packets. This aids in giving the system high performance from a latency perspective. However, alternatively, as described herein, it is possible to simply run one or more local clocks in an asynchronous manner. For example, a local clock such as a local media clock may be asynchronous with all other clocks (e.g., may be generated independently) and can be used to drive one or more processes such as analog-to-digital audio conversion, digital-to-analog audio conversion, packetization, and / or depacketization. By using such one or more local asynchronous clocks, it may be possible to enable simpler and less expensive audio devices while still achieving audio quality and latency predictions suitable for certain types of audio applications.
[0006] For example, some aspects described herein may involve an audio system in which multiple audio devices can communicate with each other. For example, a first audio device can transmit and / or receive data (e.g., audio and / or other information) to and / or from a second audio device, and the second audio device can transmit and / or receive data (e.g., audio and / or other information) to and / or from the first audio device. The audio devices can be communicatively connected to each other via a communication medium, which can involve a direct connection between the audio devices, an indirect connection between the audio devices, and / or a communication network. The connection(s) between the audio devices can be, for example, IP-based. For example, one or more audio devices can transmit data of an audio device to another data and / or receive data from another data via a communication medium within a plurality of packets such as IP packets. One or more of the audio devices can operate according to multiple clocks. For example, the transmission and / or reception of packets between the audio devices can be performed according to a first clock (e.g., transmitted and / or received based on frequency and / or phase). The first clock can be based on, for example, a master clock shared by the audio devices, such as via a communication medium. One or more of the audio devices can further convert an analog signal (such as an analog audio signal) into digital data and / or convert received digital data (such as digital audio or other data) into an analog signal according to a second clock (e.g., transmitted and / or received based on frequency and / or phase). The second clock can be asynchronous with the first clock (e.g., generated independently of the first clock).One or more audio devices can further packetize digital data to be transmitted and / or depacketize received digital data according to a first clock or a second clock (e.g., based on frequency and / or phase, transmitted and / or received).
[0007] According to further aspects described herein, the method can be implemented by an audio device. The method can include receiving an analog audio signal based on a detected sound and generating a local synchronous media clock using a local oscillator such as a crystal oscillator, a microelectromechanical system (MEMS) oscillator, a ceramic resonator, a surface acoustic wave (SAW) oscillator, an inductor / capacitor (LC) oscillator, or another type of asynchronous clocking implementation. The local asynchronous media clock need not have high accuracy. For example, the local asynchronous media clock can have a frequency variation of at least 1 ppm, or at least 10 ppm, or at least 100 ppm. As an example, an off-the-shelf crystal oscillator typically has a frequency variation in the range of 10 ppm to 100 ppm at room temperature (e.g., about 20 °C). As discussed below, the local asynchronous media clock need not be very accurate, where "very accurate" can be expected to translate efficiently to "expensive" and / or "complex". For example, a very accurate chip atomic clock is not required to implement the local asynchronous media clock. Rather, a less expensive, less complex, and / or more readily available technology can be used to implement the local asynchronous media clock for both audio transmission and reception. The method can further include generating digital audio data based on the analog audio signal using the local asynchronous media clock. A network clock can be generated using a master clock of a network connected to the audio device. For example, the network clock can be synchronized with the master clock. The audio device can transmit digital audio data over the network based on the network clock.
[0008] According to further aspects described herein, the method can be implemented by an audio device. The method can include receiving digital audio data over a network based on a network clock synchronized with the master clock of the network. The method can further include generating a local asynchronous media clock using a local oscillator such as a crystal oscillator, MEMS, ceramic resonator, SAW oscillator, LC oscillator, or another type of asynchronous clocking implementation. As discussed above and further discussed herein, the local asynchronous media clock can have an accuracy such that it has a frequency variation of at least 1 ppm, or at least 10 ppm, or at least 100 ppm. The method can further include generating an analog audio signal based on the digital audio data using the local asynchronous media clock. The method can further include generating sound based on the analog audio signal, for example, by using a speaker.
[0009] Further aspects described herein relate to a plurality of audio devices that implement the above and other methods, a system including two or more of these audio devices, and computer-executable instructions (e.g., software and / or firmware) that, when executed, cause an audio device to implement the above or other methods.
[0010] These and other features and potential advantages are described in more detail below.
Brief Description of the Drawings
[0011] Some features are shown by way of example in the accompanying drawings and are not limiting. In the drawings, the same numbers refer to like elements.
[0012]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Best Mode for Carrying Out the Invention
[0013] The accompanying drawings, which form a part of the description of this specification, illustrate embodiments of the present disclosure. It should be understood that the embodiments shown in the drawings and / or the embodiments discussed in this specification are not exclusive, and there are other embodiments regarding the implementation methods of the disclosure.
[0014] FIG. 1 is a block diagram of an exemplary audio system 100. The audio system 100 may include a plurality of audio devices such as an audio device 101 and an audio device 102. The plurality of audio devices can be communicatively coupled to each other via a communication medium such as a communication network 103. An audio device can be any kind of device that can transmit, receive, and / or process audio (e.g., modify, store, and / or operate in response to the audio). Non-limiting examples of audio devices include microphones, speakers, conferencing devices, audio recorders, personal computers, servers, display devices (e.g., televisions, or computer displays), networking devices, audio mixers, and musical instruments, or devices including these. Thus, for example, the audio device 101 can be a microphone or, alternatively, can include this, and the audio device 102 can be a speaker or, alternatively, can include this. Audio data generated based on the sound detected by the microphone can be transmitted by the audio device 101 via the communication network 103 to at least the audio device 102. The audio device 102 can appropriately generate sound in its speaker based on the received audio data. However, this is an example, and as another example, each of the audio devices 101 and 102 can include both a microphone and a speaker. As a further example, the audio device 101 can include a microphone, and the audio device 102 can include a computer device configured to store the audio data received from the audio device 101. As a further example, each of the audio devices 101 and 102 can be an element of a videoconference or video conferencing system. As a further example, each of the audio devices 101 and 102 can be an element of a public address system.Although two audio devices are shown in FIG. 1, this is merely an example, and system 100 may include any plurality of audio devices interconnected by communication network 103, such as three audio devices, four audio devices, or more.
[0015] Communication network 103 can be any type of network (including a simple connection between audio devices) that uses any one or more protocols. For example, communication network 103 can use the Internet Protocol (IP) to carry data such as audio data in IP datagrams. Communication network 103 can transmit those IP datagrams using a specific data link layer protocol such as Ethernet. This combination of IP and Ethernet is known as IP Over Ethernet (IPoE), where data (such as audio data) is placed in IP datagrams, and the IP datagrams are encapsulated in Ethernet frames. The term "packet" is used herein to include various organized groupings of data, such as datagrams (e.g., User Datagram Protocol (UDP) datagrams) and frames, although not limited to them.
[0016] Each of the audio devices can be configured to transmit data to one or more other audio devices via the communication network 103. Each of the audio devices can be configured to receive data from one or more other audio devices via the network 103. Any of the audio devices can be configured to perform both transmission and reception of data, or to transmit data exclusively, or to receive data exclusively. For example, the audio device 101 can be configured to transmit data to, and / or receive data from, the audio device 102 via the network 103, and the audio device 102 can be configured to transmit data to, and / or receive data from, the audio device 101 via the network 103. The data transmitted between the audio devices can include audio data, video data, communication control data, system control data, audio processing parameter data, and / or other types of data.
[0017] FIG. 2 is a block diagram showing exemplary details of an audio device that can be part of an audio system such as the audio system 100 of FIG. 1. For example, the audio device shown in FIG. 2 can be the audio device 101 or the audio device 102. The audio device can include, in this embodiment, the media source 201, or, in another case, can be connected to the media source 201. The media source 201, which may be inside or outside the housing of the audio device, can be any type of media source, such as a microphone, musical instrument, storage device containing recorded audio, speakerphone, telephone, or any other device capable of generating or providing an audio signal, such as an analog audio signal provided to the ADC 202.
[0018] Media source 201 can generate an audio signal representing audio, which can be an analog audio signal. The analog audio signal can be sent to an analog-to-digital converter (ADC) 202. ADC 202 can convert the analog audio signal into a digital audio signal and send it to the transmitter buffer 204. ADC 202 can operate according to (e.g., be managed by) a local clock that is asynchronous with any other clock used by the audio device. For example, ADC 202 can sample the analog audio signal at a sampling rate based on (e.g., equal to) the clock speed of the asynchronous local clock. This asynchronous clock is referred to herein as the local asynchronous media clock 203. The local asynchronous media clock 203 can be a clock having a specific nominal frequency, e.g., a nominal frequency of about 32 kHz, or about 48 kHz, or any other nominal frequency. ADC 202 can generate digital data at the frequency of the local asynchronous media clock 203 based on the analog audio signal. For example, if the local asynchronous media clock 203 is a clock having a nominal frequency of F Hertz (Hz), ADC 202 can sample the analog audio signal at a nominal sampling rate of F Hz and generate data (e.g., bytes of data) for each sample. Thus, ADC 202 can generate a digital audio signal at the frequency of the local asynchronous media clock 203 by generating nominally F bytes of data per second (or some other amount of data).
[0019] The term "nominal" is used in the above discussion because the local asynchronous media clock 203 does not need to be a very accurate clock and does not need to be synchronized with any other accurate clock (or, in fact, with any other clock regardless of accuracy). For example, the local asynchronous media clock 203 can be implemented as, or include, a local oscillator such as a crystal oscillator, for example, a piezoelectric crystal, and / or a circuit for operating the crystal. As a further example, the crystal oscillator can be, or include, a temperature-compensated crystal oscillator (TCXO) or a voltage-controlled crystal oscillator (VCXO). When a voltage is applied to the crystal oscillator, the crystal can oscillate at a specific frequency F. The frequency F can be the frequency of the local asynchronous media clock 203, which can be represented as a time-varying voltage signal. The frequency F can be fixed (e.g., by using a TCXO or by using a VCXO with a constant voltage input, with a nominally stable frequency) or the frequency F can be adjustable, such as by implementing the local asynchronous media clock 203 as a VCXO with an adjustable voltage input. As a further example, the local asynchronous media clock 203 can be implemented as a microelectromechanical systems oscillator (MEMS), a ceramic resonator, a surface acoustic wave (SAW) oscillator, an inductor / capacitor (LC) oscillator, or another type of asynchronous clocking implementation.
[0020] Regardless of implementation, the local asynchronous media clock 203 can have a frequency and phase (offset) that are asynchronous (independent) from any other clock used by the audio device and / or audio system 100. There are many potential advantages to implementing the local asynchronous media clock 203 using a crystal-based oscillator (or MEMS, ceramic, SAW, or other oscillator). When compared to other high-precision clock generation devices such as chip atomic clocks or clocking chips synchronized to an external high-precision clock source, such oscillators can result in clocks with lower accuracy and / or lower precision. Such oscillators generally have a stable nominal frequency, but using such clocking technology can cause the frequency variation of the local asynchronous media clock 203 to be at least 1 ppm, or at least 10 ppm, or at least 100 ppm. These relative variations in clock frequency and phase can be a factor when designing the size of the receiver buffer and / or transmitter buffer to reduce buffer overflow. However, in certain audio applications such as, but not limited to, conference calls, video conferences, and public address systems, there can be a valuable trade-off where very high audio quality and very high audio synchronization are not required. Compared to such high-precision clock generation devices, these types of lower-accuracy oscillators are relatively inexpensive, require less printed circuit board footprint, have less complexity in both design and manufacturing, and consume less power. All of these factors can be advantageous over any audio device, but some can be particularly advantageous over audio devices that are portable and / or battery-powered and thus have limited space and / or power. Furthermore, as discussed below, using an oscillator designed to be of low accuracy or expected to be low can provide simplicity in that the exact frequency of the local asynchronous media clock 203 may not be crucial to achieving the intended purpose of the audio device and / or audio system.This simplicity and flexibility in the implementation of local media clocks that are asynchronous in each audio device 101 and / or 102 allows an audio device to not include, for example, a phase-locked loop (PLL), which might otherwise be used to synchronize local media clocks with each other. In a system where multiple devices (such as a first device 101 that transmits audio data and a second device 102 that receives the transmitted audio data) are communicating, the first and second devices can have their own local asynchronous media clocks, which are not necessarily synchronized with each other per se. Nevertheless, regardless of the type of clock used to implement the local asynchronous media clock, the first and second devices can efficiently transmit and receive audio data using the techniques described herein while managing the receiving buffer to potentially avoid buffer underruns and overruns.
[0021] The digital audio signal generated by the ADC 202 may be received by a transmitter buffer 204 that can temporarily store the audio data of the digital audio signal. The transmitter buffer 204 can be any type of buffer, such as a first-in-first-out (FIFO) buffer. The transmitter buffer 204 packages each part of the stored audio data into packets such as IP datagrams in multiple parts, and transmits each packet as packetized digital audio to the network stack and controller 205, thereby outputting the stored audio data. The transmitter buffer can compress, encrypt, and / or packetize the buffered digital audio data, and transmit the compressed, encrypted, and / or packetized digital audio data to the network stack controller 205 at a speed managed by either the local asynchronous media clock 203 or another clock herein called the network clock 206. Any compression scheme and / or encryption scheme can be used to compress and / or encrypt the digital audio data. For example, the digital audio data can be encrypted using AES-128 counter mode encryption, AES-256 counter mode encryption. The local asynchronous media clock 203 can be completely independent (asynchronous) of the network clock 206. In other words, the frequency and phase of the local asynchronous media clock 203 are completely independent of the frequency and phase of the network clock 206, and operate at its own frequency and phase regardless of the frequency and phase of the network clock 206. Furthermore, since the local asynchronous media clock 203 of each audio device 101 or 102 has relatively low accuracy, it can be expected that the local asynchronous media clock 203 of the audio device 101 may have a different nominal frequency from the local asynchronous media clock 203 of the audio device 102. The difference in these nominal frequencies can be on the order of at least 1 ppm, or at least 10 ppm, or at least 100 ppm.Furthermore, even if the corresponding nominal frequencies are the same, the actual frequencies over time are expected to be different from each other, where the instantaneous or average frequencies at various given times can be expected to differ from each other by at least 1 ppm, or at least 10 ppm, or at least 100 ppm.
[0022] The network stack and controller 205 can function as an interface between the audio device and the network 103. Thus, audio data packets (e.g., IP datagrams) received by the network stack and controller 205 from the transmission buffer 204 can be reformatted for the network 103 and transmitted through the network 103. For example, the network stack and controller 205 can reform the audio data IP datagram by at least encapsulating the IP datagram in an Ethernet frame. The network stack and controller 205 can transmit any packet received from the transmission buffer 204 to the network 103. The network stack and controller 205 can further label the packets (e.g., IP datagrams and / or Ethernet frames) using a sequence number and / or timestamp based on the network clock 206, where the network clock 206 can be generated based on a master clock associated with the network 103 (e.g., generated in synchronization therewith). For example, the network stack and controller 205 can generate the network clock 206 based on the master network clock according to the Precision Time Protocol (PTP) specified in IEEE 1588-2008 (PTP version 2) or IEEE 1588-2019. For example, according to PTP, the master clock of the network 103 can function as a synchronization reference (e.g., a grandmaster clock), and the network clock 206 can be generated to have a frequency and phase based on (e.g., equal to) the frequency and phase of the master clock. Thus, the network clock 206 can be synchronized with the master clock or, alternatively, generated based on the master clock.
[0023] In accordance with the above discussion, the local asynchronous media clock 203 for any given audio device in the audio system 100 can be completely independent of (asynchronous with) the master clock of the network 103. In other words, the frequency and phase of the local asynchronous media clock 203 are completely independent of the frequency and phase of the master clock and operate at its own frequency and phase regardless of the frequency and phase of the network clock 206. Further, the network clock 206 can be generated by an audio device based on the master clock of the network 103 (e.g., can be synchronized with both the frequency and phase of the master clock of the network 103).
[0024] As described above, packets that can be transmitted to network 103 can be labeled, for example, with a timestamp and / or a sequence number by the network stack and controller 205. The timestamp can be, for example, a PTP timestamp. The timestamp can be generated based on the network clock 206 or the master clock of network 103, and can also be generated in accordance with RFC 7273. For example, each timestamp can have a value based on the value of the network clock 206 associated with the packet, such as the value of the network clock 206 when the packet is generated or transmitted. The network clock 206 does not need to have an accurate timestamping capability such as hardware, and can be implemented in software, for example, to save complexity and external components. The sequence number is incremented for each packet transmitted (e.g., by one value). On the receiving side (e.g., a packet received by audio device 102 generated by audio device 101), various received packets can be organized using the timestamp and / or the sequence number, even if the local asynchronous media clock 203 of each audio device remains freely running in an asynchronous format. While using synchronous clocking for specific interactions between audio devices 101 and 102, in such a hybrid configuration where asynchronous clocking is used for specific internal processing within each of audio devices 101 and 102, complexity is reduced and it is possible to simplify the hardware and software of the audio devices. When a packet is received by another audio device 101 or 102 through network 103, the timestamp and / or the sequence number within each received packet is read and can be used to determine the order in which the packets are to be received, buffered, and / or processed, even if the latency of network 103 varies.
[0025] In addition to, or instead of, transmitting audio data to network 103, audio device 101 or 102 can receive audio data from network 103. In the example shown in FIG. 2, network stack and controller 205 can receive a network packet containing audio data (e.g., an Ethernet frame encapsulating an audio data packet such as an IP datagram), and can transmit the audio data packet to receive buffer 207, thereby depacketizing (e.g., extracting the audio data from the data packet), restoring if necessary, decrypting if necessary, extracting, restoring, and / or multiplexing the audio data and temporarily storing it. Receive buffer 207 can be any type of buffer, such as a FIFO buffer, can be combined with transmit buffer 204, or can be implemented as a completely separate buffer. Receive buffer 207 can depacketize an audio data packet (e.g., an IP datagram) received from network stack and controller 205 by extracting the audio data from the audio data packet and storing the audio data in receive buffer 207. Depacketization can be performed (e.g., managed) according to network clock 206 or by local asynchronous media clock 203. As discussed above, audio data packets can include a timestamp and / or a sequence number, which can be used by the receiving audio device to measure the correct order in which the audio packets can be processed. For example, the audio data extracted from a given audio data packet can be stored at a position within receive buffer 207 corresponding to (e.g., indexed by, or otherwise associated with) the timestamp and / or sequence number of the audio data packet, and / or in a certain order within receive buffer 207.By doing so, even when a later-transmitted packet is received earlier than a previously-transmitted packet due to the latency of network 103, the reception unit buffer 207 can ensure that it processes (for example, converts to an analog audio signal) the contents of various packets in the correct order. In a further embodiment, the reception unit buffer 207 can store audio data in the order in which it is received, such as in a first-in first-out (FIFO) configuration.
[0026] The receiving section buffer 207 can transmit a part of the stored audio data to a digital-to-analog converter (DAC) 208 to convert the digital audio data into an analog audio signal. The DAC 208 can operate according to the "local asynchronous media clock 203 (for example, its operation can be managed by the local asynchronous media clock 203). For example, the DAC 208 can receive (for example, extract or retrieve) the digital data stored in the receiving section buffer 207 at a speed based on the frequency of the local asynchronous media clock 203. For example, if the local asynchronous media clock 203 is a clock having a frequency of F Hz, the DAC 208 can receive the digital audio data from the receiving section buffer 207 and convert the digital audio data into an analog audio signal having a frequency of the local asynchronous media clock 203 by converting F bytes of data per second into an analog signal. The local asynchronous media clock is shown in FIG. 2 as providing the local asynchronous media clock 203 to the ADC 202 and the DAC 208, but the local asynchronous media clock signal may be provided to any one or more elements of the audio device 101 or 102 as desired. Further, the local asynchronous media clock 203 is shown in FIG. 2 as being connected to other elements in a specific manner and is shown as an element unique to the figure, but the local asynchronous media clock 203 can be located anywhere inside or outside the transmitting section chain (including at least elements 201, 202, 204, and 205), anywhere inside or outside the receiving section chain (including at least elements 205, 207, 208, and 209), and / or as part of any of the other elements in FIG. 2.
[0027] The DAC 208 can transmit the generated analog audio signal to the media receiver 209. The media receiver 209, which may be inside or outside the housing of the audio device, can be any type of media receiver such as a speaker, an audio storage device, a speakerphone, a telephone, or any other device capable of receiving and / or processing an audio signal such as the analog audio signal generated by the DAC 208. The media receiver 209 may be a device separate from the media source 201, or the two devices may be integrated as a single same device. For example, a speakerphone can include both a media source (e.g., its microphone and associated circuitry) and a media receiver (e.g., its speaker and associated circuitry). In some embodiments, the media source 201 and the media receiver 209 can be packaged together simultaneously as the same device as the rest of the circuitry of the audio device 101 or 102. For example, a single housing can include or at least partially include any or all of the components 201 - 209 shown in FIG. 2. In other embodiments, the media source 201 and / or the media receiver 209 can be physically separated while being communicatively connected to a device containing any of the remaining components 202 - 208. In such embodiments, the analog audio signal from the media source 201 to the ADC 202 and / or the analog audio signal from the DAC 208 to the media receiver 209 can be communicated via an external port and / or a cable. Further, the audio device 101 or 102 can include only a subset of the components shown in FIG. 2. For example, the audio device 101 or 102 can be configured to transmit audio to the network 103 and not receive audio from the network 103, or the audio device 101 or 102 can be configured to receive audio from the network 103 and not transmit audio to the network 103.In these embodiments, the audio device 101 or 102 may include at least components 202 to 206 (and optionally 201), and may not include components 207 to 209, or the audio device 101 or 102 may include at least components 203, and 205 to 208 (and optionally 209), and may not include components 201, 202, and 204.
[0028] As already mentioned, when packets are received by an audio device through network 103, the timestamps and / or sequence numbers within each received packet are read and can be used to determine in what order each packet should be read, buffered, and / or processed relative to other packets received, even if the network 103 latency varies. The timestamps and / or sequence numbers can also be used to detect dropped packets. For example, assume that two packets are to be sent from audio device 101 to audio device 102. Here, the first packet has a first timestamp and / or a first sequence number and is sent by audio device 101 before sending a second packet having a second timestamp and / or a second sequence number. Even if the first packet is received by audio device 102 after the second packet has been received (e.g., due to variable latency in network 103), audio device 102 can appropriately reorder the first and second packets entering receive buffer 207 based on the corresponding timestamps and / or sequence numbers so that audio device 102 can buffer the first packet before the second packet (e.g., within receive buffer 207), such that the first packet is processed (e.g., converted from digital to analog using DAC 208) before the second packet is processed. To accomplish this, receiving audio device 102 can store the audio data for each of the packets within receive buffer 207 in an order and / or storage location based on the corresponding timestamps and / or sequence numbers. Further, the buffered audio data can be read from receive buffer 207 and sent to DAC 208 in an order based on the corresponding storage locations within receive buffer 207.For example, when the receiving audio device 102 measures that the second sequence number is a sequence number more than two away from the first sequence number (e.g., one or more skipped sequence number values between the value of the first sequence number and the value of the second sequence number), the receiving audio device 102 can measure that one or more packets have been dropped (e.g., lost on the way to the receiving audio device 102 or, due to some other problem, not received by the receiving audio device 102). The receiving audio device 102 can measure the amount of dropped packets based on the number of sequence numbers missing from the received packets. For example, if the received packets contain sequence numbers [1, 2, 6, 7, 8, 9, …], the receiving audio device 102 can measure that the packets are missing sequence numbers 3, 4, and 5 and thus that three packets have been dropped. The receiving audio device 102 can measure that one or more packets have been dropped based on the period during which one or more packets containing sequence numbers are not received as expected. For example, if there is a value of a third sequence number between the value of the first sequence number and the value of the second sequence number and a packet containing the value of the third sequence number is not received after a threshold period from the packet containing the first sequence number or the second sequence number, the receiving audio device 102 can measure that the packet containing the expected third sequence number has been dropped. When the audio device 102 measures that a packet has been dropped, the audio device 102 can fill the receiving buffer with a series of created data instead of where the data from the dropped packet was stored, or perform some other operations such as generating a signal indicating the dropped packet. In this case, the signal can be used to indicate, for example, the state of the dropped packet to the user of the audio device 102.
[0029] As another example, the receiving audio device 102 can measure whether one or more packets are dropped using timestamps. The receiving audio device 102 can be configured with an expected time between packets, or an expected packet transmission speed that can be derived from the expected time between packets (based on the reciprocal of the expected packet transmission speed). The expected packet transmission speed between packets can be predetermined, or can be dynamically measured by the receiving audio device 102, such as by measuring the packet speed and / or time between packets and averaging these values over a sliding window of time. However, once the expected time between packets is determined, the receiving audio device 102 can save the value of the expected time between packets, referred to herein as T EP in this specification. When the receiving audio device 102 receives packets that are separated by a time approximately equal to T EP as indicated by the timestamps, the receiving audio device 102 can measure that no packets are dropped. However, if two timestamps are separated by a time T EP greater than T and there are no received packets with timestamps between these two timestamps, as measured by the receiving audio device 102, the receiving audio device 102 can measure that at least one packet is dropped. Further, the receiving audio device 102 can measure the number of one or more of these dropped packets as a multiple of T EP . In other words, the number of dropped packets between these two received packets can be measured as equal to T EP / T, which can be rounded as necessary.
[0030] To enhance the reliability of dropped packet measurements, the receiving audio device 102 can measure one or more dropped packets using both the sequence number and timestamp of the packets discussed above. For example, if using the packet sequence number indicates the measurement of one or more dropped audio packets, and using the timestamp indicates the measurement of one or more dropped packets, based on both of these indicating one or more dropped packets, the receiving audio device 102 can measure that one or more packets are dropped. In another example, if using the packet sequence number indicates the measurement of a specific quantity of one or more dropped audio packets (e.g., 3 packets), and using the timestamp indicates the measurement of the same specific quantity of one or more dropped packets (e.g., 3 packets), based on both of these indicating the same quantity of one or more dropped packets, the receiving audio device 102 can measure that a specific quantity of one or more packets (e.g., 3 packets) are dropped. If only one of the two measurements indicates one or more dropped packets (e.g., indicated by using the sequence number but not by using the timestamp, or indicated by using the timestamp but not by using the sequence number), the receiving audio device 102 cannot measure that one or more packets are dropped. If the two measurements (using the sequence number and using the timestamp) indicate that packets are dropped, but the measured number of dropped packets indicates using both methods, the receiving audio device 102 can, if desired, measure the number of dropped packets based on one of these methods, such as using a smaller or larger number of dropped packets.For example, for a given time frame between two received packets, if the receiving audio device 102 measures, using the sequence number, that two packets are missing between them, and if the receiving audio device 102 also measures, using the timestamps, that three packets are missing between them, the receiving audio device 102 can, if desired, measure that the fewer amount (two packets) or the greater amount (three packets) is missing.
[0031] In addition to measuring the order in which data within the received packet is buffered and / or converted to an analog signal, the receiving audio device 102 can further measure what the latency of the network 103 is for each received packet using the received timestamp. This is due to the method of generating the timestamp. For example, each timestamp can be generated based on the value of the network clock 206, which is known to the network stack and controller 205 of the transmitting audio device 101, where the network clock 206 of the transmitting audio device 101 can be synchronized with the master clock of the network 103. The receiving audio device 102 can have its own network clock 206 that is also synchronized with the master clock of the network 103. Thus, the receiving audio device 102 can measure what the latency of the received packet is (e.g., the time elapsed since the transmitting audio device 102 sent the packet) based on the timestamp of the received packet and the network clock 206 generated by itself. For example, the receiving audio device 102 can compare the timestamp of the received packet with the value of the network clock 206 at the time the packet was received (or some value derived from the value of the network clock 206). The receiving audio device 102 can perform several operations based on the measured latency.For example, when the latency becomes very large (e.g., when the latency is measured to exceed a known threshold), the receiving audio device 102 can take a first action such as adding filter data (e.g., zeros or interpolated audio data) to the audio data stored in the receiving buffer 207, or the receiving audio device 102 can send a signal to the audio device 101 that sends the signal, and can indicate to the sending audio device 101 that the receiving audio device 102 is experiencing a large latency, or the audio device 102 can show a message to the user (e.g., through a display or other user interface) to indicate that the audio device 102 is experiencing a large latency. Similarly, when the latency is measured to be below the threshold, the receiving audio device 102 can take a second action such as modifying the data in the receiving buffer 207 by removing (e.g., deleting, ignoring, or overwriting) a subset of the data from the buffer 207. The receiving audio device 102 can also measure the latency of the audio of the entire system using timestamps and can achieve the target latency by adjusting the amount of audio stored in the buffer 207 (and / or using sample rate conversion). For example, the receiving audio device 102 can be configured by a set point for how the buffer 207 should be filled. The receiving audio device 102 can measure the latency from the timestamp in the received packet by comparing the received timestamp with its own clock or the network clock. The receiving audio device 10 can respond to the measured latency by dropping data from the buffer 207 or other time-corrected data in the buffer and / or by storing the generated, time-corrected, and / or interpolated data in the buffer 207, and can maintain the buffer 207 near or at the set point.By maintaining the buffer 207 in a filled state at or near the set point, the receiving audio device 102 can provide audio with approximately the same latency regardless of the audio transmission source and regardless of the audio network path. For example, if the receiving audio device 102 receives two transmitting audio devices and the audio from one of the transmitting audio devices has a greater latency (is subject to latency) than the other transmitting audio device, one or more portions of the transmitting audio device with low latency (little latency) may be saved to the buffer 207 or, in another case, dropped from time correction, and / or one or more portions of the audio data may be created, time corrected, and / or interpolated by the receiving audio device 102 and, in another case, fill the latency gap that occurs in the audio data saved to the buffer 207. The receiving audio device 102 can further select a target latency, for example, a target latency based on a known or expected network latency, for example, a target latency slightly greater than the known or expected network latency.
[0032] The receiving audio device 102 can further measure the rate at which data (e.g., sample rate or packet rate) is received by the audio device 102, and the receiving audio device 102 can take one or more actions based on the measured data rate. For example, the receiving audio device 102 can measure, over time, the rate at which packets are received at the network stack and controller 205, and / or the receiving audio device 102 can measure, over time, the rate of audio samples received by or stored in the receive buffer 207. The receiving audio device 102 can compare the measured rate to a threshold rate and take one or more actions based on the comparison. For example, the receiving audio device 102 can compare the measured packet rate to a threshold packet rate, or the measured sample rate to a threshold sample rate, or any other measure of the incoming data rate to a threshold data rate. The threshold sample rate can be any rate, such as 96 samples per 2 milliseconds (or 48 samples per millisecond), which may correspond to an expected audio rate of 48 kHz. If the audio rate is expected to be another rate, the threshold may be a different value. For example, if the expected audio rate is 32 kHz, the threshold sample rate may be 32 kHz (e.g., 32 samples / millisecond, or 64 samples / 2 milliseconds). Still further, in an example, the threshold data rate may be equal to the nominal clock rate of the local asynchronous media clock 203. When the measured data rate (e.g., sample rate or packet rate) is below the threshold, the receiving audio device 102 can fill the receive buffer 207, such as by adding data that approximately creates lost expected data and adds to the receive buffer 207 enough to approximately achieve the expected data rate. The added data may be a predetermined value such as all zeros, or other data such as audio data interpolated from the actually received audio data.When the measured data rate (e.g., sample rate or packet rate) exceeds a threshold, the receiving audio device 102 can prevent the data from being stored in the receiving buffer 207, which is sufficient to bring about an expected data rate. If the threshold data rate is equal to (or, in another case, based on) the nominal clock rate of the local asynchronous media clock 203, the DAC 208 (which can operate according to the local asynchronous media clock 203) can continuously extract data from the receiving buffer 207 at a rate controlled by the local asynchronous media clock 203. Ideally, this process can generally keep the receiving buffer 207 partially full at all times while potentially avoiding underflow or overflow conditions in the receiving buffer 207. Further, when the measured data rate (e.g., sample rate or packet rate) exceeds or falls below the threshold, the audio device 102 can perform sample rate conversion with reference to the sample rate converter 301 and, as described later with reference to FIG. 3 and the like.
[0033] Therefore, each audio device within the audio system 100 can operate using a combination of an asynchronous clock and a synchronous clock. Specifically, for example, each audio device within the audio system 100 can generate a network clock 206 based on the master clock of the network 103 by synchronizing its own network clock 206 with the master clock, for example, in accordance with IEEE 1588-2008 or IEEE 1588-2019. The network clock 206 can be used for one or more aspects of communication between the audio devices 101 and 102 via the network 103 (for example, a time stamp for a packet can be generated using the network clock 206). Further, each audio device within the audio system 100 can generate its own local asynchronous media clock 203 that is asynchronous to both the network clock 206 and the master clock of the network 103. The local asynchronous media clock 203 can be used to manage communication and / or processing within the audio device and / or with respect to the media source 102 and / or the media receiver 209. Specifically, for example, the local asynchronous media clock 203 can be used to control the rate at which an analog audio signal from the media source 201 is converted by the ADC 202 into a digital audio signal stored in the transmitter buffer 204. Further, or alternatively, the local asynchronous media clock 203 can be used to control the rate at which a digital audio signal from the receiver buffer 207 is converted by the DAC 208 into an analog audio signal received by the media receiver 209.
[0034] Figure 3 is a block diagram showing another exemplary detail of an audio device that can be part of an audio system such as the audio system of FIG. 1. In this embodiment, the audio device 101 or 102 can include any of the components as described above with respect to FIG. 2 and can also include a sample rate converter (SRC) 301. The SRC 301 can assist in reducing the amount of overrun or underrun that might otherwise be experienced by the receive buffer 207. The SRC 301 can receive digital audio data stored in the receive buffer 207 and convert the digital audio data to a different sample rate. For example, if the digital audio data represents audio sampled at a first rate R1, the SRC 301 can convert the digital audio data to represent audio sampled at a second rate R2. R2 can be a rate faster than R1 or R2 can be a rate slower than R1. When the SRC 301 is converting from the slow rate R1 to the fast rate R2, the SRC 301 can perform an upsampling process, such as by inserting additional digital audio data between the existing R1 rate samples. The digital audio data inserted can be, for example, a predetermined one or more values (e.g., all zeros, sometimes referred to as "zero stuffing"), or can be interpolated values calculated based on the original digital audio data values. When the SRC 301 is converting from the fast rate R1 to the slow rate R2, the SRC 301 can perform a downsampling process, such as by removing a selected subset of the original digital audio data values / samples. However, other upsampling or downsampling processes can be used. The SRC 301 can be used to efficiently convert digital audio data that can be sampled at a rate that matches one clocking domain (e.g., frequency and / or phase) to be sampled at a rate that matches another different clocking domain (e.g., another different frequency and / or phase).For example, the speed R1 may be the speed at which packets are received by the receiving buffer 207 from the network stack and the controller 205 (in this case, R1 can be based on the frequency of the local asynchronous media clock 203 of the transmitting audio device), and the speed R2 may be the nominal frequency of the local asynchronous media clock 203 of the receiving audio device. Therefore, the SRC 301 can be used to at least partially compensate for the mismatch in the two local asynchronous media clocks 203 of the transmitting and receiving audio devices. Although the SRC 301 is shown after the receiving buffer 207 (for example, between the receiving buffer 207 and the DAC 208), the SRC 301 can alternatively be located before the receiving buffer 207 (for example, between the network stack and the controller 205 and the receiving buffer 207). In addition to, or instead of, providing a receiving-side SRC that functions in conjunction with the receiving buffer, a transmitting-side SRC can be added between the ADC 202 and the transmitting buffer 204, or between the transmitting buffer 204 and the network stack and the controller 205, etc., and made to function in conjunction with the transmitting buffer 204. In this case, the transmitting-side SRC can convert the sample rate of the digital audio received from the ADC 202 into audio data with a different sample rate that can be stored in the transmitting buffer 204 and / or transmitted to the network stack and the controller 205.
[0035] FIG. 4 is a block diagram of an exemplary audio system, such as the audio system of FIG. 1, including exemplary details of two audio devices in the audio system. In the illustrated embodiment, audio device 101 may be communicatively coupled to audio device 102 via network 103. Each of audio devices 101 and 102 may operate in accordance with the description of this specification with respect to FIGS. 2 and / or 3. Thus, for example, audio device 101 may transmit audio data packets to audio device 102 via network 103. Further, audio device 102 may transmit audio data packets to audio device 101 via network 103. In other examples, audio device 101 may be able to transmit audio data packets to audio device 102, but may not be able to receive any audio data packets from audio device 102, or audio device 101 may be able to receive audio data packets from audio device 102, but may not be able to transmit audio data packets to audio device 102. In such examples, audio device 101 may, for example, not include components 207-209, and audio device 102 may, for example, not include components 201, 202, and 204.
[0036] An example of the operation method of the system of FIG. 4 is described below. The audio device 101 may include its media source 201 or, in another case, may be connected to its media source 201. The media source 201 of the audio device 201 may be, for example, a microphone and associated circuitry for operating the microphone. The microphone generates an analog audio signal, which can be received by the ADC 202 of the audio device 101 and converted into a digital audio signal. The analog-to-digital conversion by the ADC 202 can be performed at a speed (frequency) and / or phase based on (e.g., synchronized with) the speed (frequency) and / or phase of the local asynchronous media clock 203 of the audio device 101. The digital audio signal is received by the transmitter buffer 203 of the audio device 101, and digital audio data can be stored based on the digital audio signal. The transmitter buffer 204 can packetize the stored digital audio into packets (e.g., IP datagrams) and transmit these packets to the network stack and controller 205 of the audio device 101. The packetization of the stored digital audio data and / or the transmission of the packets can be performed at a speed (frequency) and / or phase based on (e.g., synchronized with) the speed (frequency) and / or phase of the local asynchronous media clock 203 or the network clock 206. The network stack and controller 205 of the audio device can further packetize the packets, for example, by encapsulating the IP datagram in an Ethernet frame, and finally transmit the processed packets to the audio device 102 via the network 103. The transmission of the packets via the network 103 can be performed at a speed (frequency) and / or phase based on (e.g., synchronized with) the speed (frequency) and / or phase of the master clock.
[0037] After crossing network 103, the packet is received by the network stack and controller 102 of the audio device 102, and the received packet is at least partially depacketized (for example, by extracting the IP datagram from the encapsulated Ethernet frame), and the resulting audio data packet (for example, the IP datagram) can be transmitted to the receive buffer 207 of the audio device 102. The reception of the packet from network 103 can be carried out at a speed and / or phase based on (for example, synchronized with) the speed (frequency) and / or phase of the master clock. The receive buffer 207 of the audio device 102 can further depacketize the audio packet received from the network stack and controller 205, extract the audio data, and store the audio data in the audio packet. For example, the receive buffer 207 can extract the audio data stored in the IP datagram received from the network stack and controller 205. The depacketization of the audio packet and the storage of the digital audio data by the receive buffer 207 of the audio device 102 can be carried out at a speed and / or phase based on (for example, synchronized with) the speed (frequency) and / or phase of the network clock 206 of the audio device 102, or the local asynchronous media clock 203 of the audio device 102. The DAC 208 of the audio device 102 can receive the stored digital audio data of the receive buffer 207 and convert the received digital audio data into an analog audio signal that can be transmitted to the media receiver 209 of the audio device 102. The analog-to-digital conversion can be carried out at a speed and / or phase based on (for example, synchronized with) the speed (frequency) and / or phase of the local asynchronous media clock 203 of the audio device 102. Subsequently, the media receiver 209 of the audio device 102 can process the received analog audio signal.For example, when the media receiving unit 209 of the audio device 102 is a speaker, the media receiving unit 209 can generate sound based on an analog audio signal.
[0038] The audio flow can also move from the audio device 102 to the audio device 101, and the operation is the same as that described above with respect to FIG. 4, except that the references to the audio device 101 and the audio device 102 can be reversed. Further, the audio device 101 can receive audio from the audio device 102 and at the same time transmit audio to the audio device 102, and can receive audio from the audio device 102 at the same time as transmitting audio to the audio device 102. Although only two audio devices 101 and 102 are shown in FIG. 4, the audio system 100 can include three or more audio devices, such as three audio devices, four audio devices, or more, interconnected via the network 103. If there are three or more audio devices in the audio system 100, any given audio device can transmit audio to two or more other audio devices (simultaneously or at different times), and any audio device can receive audio from two or more other audio devices (simultaneously or at other times). For example, an audio device can transmit packets via the network 103 addressed to one or more other audio devices. Audio packets between any one or more audio devices and another one or more audio devices can be transmitted via one or more streams, such as one or more IP streams.
[0039] The local asynchronous media clocks 203 of each of the audio devices 101 and 102 may be asynchronous with each other and with any other clock within the system 100. Thus, for example, the local asynchronous media clock 203 of the audio device 101 may include a first oscillator, and the local asynchronous media clock 203 of the audio device 102 may include a second oscillator independent of the first oscillator. Each of the two local asynchronous media clocks 203 can be implemented using techniques that generally result in less accurate clocks, such as using, by way of non-limiting example, a crystal oscillator, a ceramic resonator, a MEMS oscillator, a SAW oscillator, or an LC oscillator. The local asynchronous media clocks 203 of the plurality of audio devices 101 and 102 can have the same nominal frequency or different nominal frequencies. For example, the local asynchronous media clock 203 of the audio device 101 can have a nominal frequency of 32 kHz, and the local asynchronous media clock 203 of the audio device 102 can have a nominal frequency of 48 kHz. Alternatively, the local asynchronous media clocks 203 of the audio devices 101 and 102 can both have a nominal frequency of 32 kHz or both have a nominal frequency of 48 kHz. The specific frequency values described herein are merely examples, and any one or more of the nominal frequencies of the local asynchronous media clock 203 can be used.
[0040] Multiple audio devices (e.g., audio devices 101 and 102) within the audio system 100 can transmit non-audio data in addition to audio data via the network 103. Examples of such non-audio data can include configuration settings, status indicators, capability indicators, or handshake protocol signaling. For example, audio devices can communicate with each other to indicate the nominal frequency of the local asynchronous media clock 203, or one or more configured, or preferred, audio compression settings (e.g., one or more compression ratios that are configured, available, or preferred), or an audio compression method (e.g., one or more types of coders / decoders (CODECs) that are configured, available, or preferred and can be used). Any kind of indicator can be used. For example, to indicate a clock speed of 48 kHz, a data packet may be transmitted by an audio device that includes a number such as "48" or "48,000". Alternatively, an audio device can transmit a data packet that indicates a specific shorthand value known to other audio devices within the audio system 100. For example, 32 kHz can be represented by a specific bit set to zero, and 48 kHz can be represented by a specific bit set to one. Such non-audio data can be transmitted in data packets dedicated to non-audio data (e.g., datagrams), in which case the non-audio data packet can be distinguished from an audio data packet by including first information in the packet header to indicate that it is a non-audio data packet, and including different second information in the packet header to indicate an audio data packet. Alternatively, both audio data and non-audio data can be combined within the same data packet. In either case, the audio data and non-audio data can be included in one or more payload portions of one or more packets.One potential advantage of such audio devices communicating such information with each other is that the audio devices can use the communicated information to configure themselves in a particular way, configure the other of the audio devices within the audio system 100 in a particular way, or generally negotiate one or more particular configurations such that the audio devices within the audio system 100 operate and communicate with each other in a compatible way. For example, two audio devices within the audio system 100 can have two different local asynchronous media clock speeds and can negotiate a particular audio compression ratio based on one or both of the corresponding local asynchronous media clock speeds, using devised clock speed information or other configuration information. Such negotiation can be performed automatically between the audio devices within the audio system 100. This can provide a convenience to the user of the audio system 100 in that the user may not need to worry about the local asynchronous media clock speeds of the various audio devices within the audio system 100, thereby potentially imparting flexibility in the selection of audio devices to operate in conjunction within the audio system 100.
[0041] FIG. 5 is a block diagram showing exemplary details of an audio device that can be part of an audio system such as the audio system of FIG. 1. For example, the audio device can be audio device 101 or audio device 102. The audio device can be implemented, for example, as an arithmetic device that executes stored instructions and / or as a wired-connected circuit, or, in another case, for example, these can be cited, and / or one or more processors can execute stored computer-readable instructions. In the illustrated embodiment, the arithmetic device can include or be connected to any one of the following: one or more processors 501, a storage device 502 (which can include one or more computer-readable media such as a memory), an external interface such as a network interface 502 (configured to communicate with network 103), a user interface 504, one or more microphones and / or associated circuits 505 configured to detect sound and convert the detected sound into an audio signal such as an analog audio signal or a digital audio signal, one or more digital signal processors 506 configured to implement one or more digital signal processing features of the audio device, one or more speakers and / or associated circuits 507 configured to generate sound in response to a received audio signal such as an analog audio signal or a digital audio signal, and / or a local oscillator 508. The one or more processors 501 can be communicably connected to any of the other components 502-508 via one or more data buses and / or one or more other types of connections.
[0042] In the example of FIG. 5, media source 201 is shown as one or more microphones of component 505, and media receiver 209 is shown as one or more speakers of component 507. However, media source 201 and media receiver 209 can be any other type of media source and media receiver discussed above. ADC 202 and / or transmitter buffer 204 can be implemented by the circuitry of component 505 and / or one or more processors 501, and DAC 208 and / or receiver buffer 207 can be implemented by the circuitry of component 507 and / or one or more processors 501. The circuitry of components 505 and 507 may, if desired, be separate circuits or may be examples of a combined single circuit. Network stack and controller 205 and / or network clock 206 can be implemented by network interface 503 and / or one or more processors 501. Local asynchronous media clock 203 can be implemented by local oscillator 508. In the illustrated embodiment, local oscillator 508 can provide a local asynchronous media clock signal to one or more processors 501 (e.g., for controlling the operation of ADC 202 and / or transmitter buffer 204), the circuitry of component 505, and the circuitry of component 507 (e.g., for controlling the operation of DAC 208 and / or receiver buffer 207). However, the local asynchronous media clock can be provided to any of the components of FIG. 5, if desired. In one embodiment, one or more processors 501 can receive a signal from local oscillator 508, and one or more processors 501 can generate an asynchronous local media clock based on the signal from local oscillator 508. For example, one or more processors 501 can include a phase-locked loop (PLL) circuit, and the signal from local oscillator 508 can be an input to the PLL circuit (e.g., for driving the PLL circuit).
[0043] One or more processors 501 can be configured to execute instructions stored in a storage device 502. When executed by the one or more processors 501, the instructions can cause a computer device (and thus, an audio device) to perform any of the functionality described herein that is to be performed by an audio device (such as audio device 101 or audio device 102). For example, the one or more processors 501 can control the operation of any of the other components 502-508 of the audio device, and / or can command various signals (such as audio signals and / or clock signals) between the various components 502-508 of the audio device.
[0044] Optionally, power can be supplied to the audio device and / or to any of the components of the audio device (such as any of components 501-508). Although not explicitly stated, the audio device may include an internal power source and / or an external power connection.
[0045] The following highlights various features of a series of numbered sections or paragraphs. These features should not be construed as limiting the invention or the inventive concept, but should be provided only to highlight some of the features described herein, and do not recommend a particular order of importance or relevance of these features.
[0046] Clause 1 Receiving digital audio data based on a network clock synchronized with a master clock of the network via the network, Generating a local synchronous media clock, Comparing the speed of the received digital audio data with a threshold data speed, wherein the threshold data speed is based on a nominal speed of a local asynchronous media clock. Storing at least a part of the received digital audio data in a buffer, wherein the at least a part of the received digital audio data is based on the comparison, the storing, and Generating an analog audio signal based on the at least a part of the digital audio data stored in the buffer using the local asynchronous media clock; and Generating sound based on the analog audio signal by using a speaker. A method including these steps.
[0047] Clause 2. The generating of the local asynchronous media clock includes generating the local asynchronous media clock by using a low-precision clocking technique that uses at least one of a crystal oscillator, a MEMS oscillator, a ceramic resonator, a SAW oscillator, or an LC oscillator. The method according to Clause 1.
[0048] Clause 3. The receiving of the digital audio data based on the network clock includes receiving the digital audio data based on the network clock and / or based on a plurality of timestamps generated based on a plurality of sequence numbers included in the digital audio data. The method according to Clause 1 or 2.
[0049] Clause 4. Further including buffering the digital audio data at a plurality of buffer positions based on the plurality of timestamps and / or based on the plurality of sequence numbers. The method according to any one of Clauses 1 to 3.
[0050] Clause 5. Further including, based on the comparison, preventing at least a part of the received digital audio data from being stored in the buffer. The method according to any one of Clauses 1 to 4.
[0051] Clause 6. The method according to any one of Clauses 1 to 4, further comprising stuffing additional data into the buffer based on the comparison.
[0052] Clause 7. Receiving the digital audio data includes receiving a plurality of data packets including the digital audio data, and the method further includes extracting the digital audio data from the plurality of data packets. The method according to any one of Clauses 1 to 6.
[0053] Clause 8. The method according to Clause 7, wherein the plurality of data packets includes one or both of a plurality of Internet Protocol data grams and a plurality of Ethernet frames.
[0054] Clause 9. The method according to Clause 7 or 8, wherein the extracting is managed by the local asynchronous media clock.
[0055] Clause 10. The receiving is performed by a first audio device, and the method further includes: receiving a second analog audio signal based on the detected sound; and generating second digital audio data based on the second analog audio signal using the local asynchronous media clock; and transmitting the second digital audio data via the network. The method according to any one of Clauses 1 to 9.
[0056] Clause 11. The method according to any one of Clauses 1 to 10, further comprising restoring at least a portion of the digital audio data, and generating the analog audio signal based on at least a portion of the digital audio signal includes generating the analog audio signal based on the restored digital audio data.
[0057] Clause 12. The method according to any one of Clauses 1 to 11, further comprising decrypting at least a part of the digital audio data, and generating the analog audio signal based on at least a part of the digital audio signal includes generating the analog audio signal based on the decrypted digital audio data.
[0058] Clause 13. The method according to any one of Clauses 1 to 12, wherein the storing is managed by the local asynchronous media clock.
[0059] Clause 14. The method according to any one of Clauses 1 to 13, wherein generating the analog audio signal based on the digital audio data includes converting at least a part of the digital audio data stored in the buffer into the analog audio signal using a digital-to-analog converter.
[0060] Clause 15. The method according to Clause 14, further comprising performing sample rate conversion on at least a part of the digital audio data stored in the buffer.
[0061] Clause 16. The method according to any one of Clauses 1 to 15, further comprising measuring one or more lost packets based on the received packet sequence number and / or the received packet timestamp, and adjusting data in the receive buffer based on the measurement of the one or more lost packets.
[0062] Clause 17. A first audio device, one or more processors, one or more computer-readable media storing instructions that, when executed by the one or more processors, cause the first audio device to perform the method according to any one of Clauses 1 to 16.
[0063] A non-transitory computer-readable medium that stores instructions to cause a first audio device to perform the method described in any one of clauses 1 to 16 when executed.
[0064] Clause 19: Receiving an analog audio signal based on a detected sound, Generating a local synchronous media clock, Generating the digital audio data based on the analog audio signal using the local asynchronous media clock, Generating a network clock based on a master clock of a network, Transmitting the digital audio data based on the network clock via the network, a method comprising.
[0065] Clause 20: The generating the local asynchronous media clock includes generating the local asynchronous media clock using a low-precision clocking technique that uses at least one of a crystal oscillator, a MEMS oscillator, a ceramic resonator, a SAW oscillator, or an LC oscillator, the method described in clause 19.
[0066] Clause 21: The method described in clause 19 or 20 further includes packetizing the digital audio data into a plurality of data packets, and the transmitting includes transmitting the plurality of data packets.
[0067] Clause 22: The plurality of data packets include one or both of a plurality of Internet protocol data grams or a plurality of Ethernet frames, the method described in any one of clauses 19 to 21.
[0068] Clause 23: The packetizing is the method described in any one of clauses 19 to 22, managed by the local asynchronous media clock.
[0069] Clause 24. The method according to any one of Clauses 19 to 23, wherein transmitting the digital audio data based on the network clock includes transmitting the digital audio data in a plurality of packets each including a time stamp based on the network clock and / or each including a sequence number.
[0070] Clause 25. The method according to any one of Clauses 19 to 24, wherein transmitting the digital audio data includes transmitting the digital audio data in a plurality of packets at a speed based on the frequency of the local asynchronous media clock.
[0071] Clause 26. The transmitting is performed by a first audio device, and the method further includes: receiving, by the first audio device, second digital audio data via the network; generating, using the local asynchronous media clock, a second analog audio signal based on the digital audio data; and generating sound based on the second analog audio signal by using a speaker associated with the first audio device. The method is according to any one of Clauses 19 to 25.
[0072] Clause 27. The method according to any one of Clauses 19 to 26, further including compressing the digital audio data, and the transmitting includes transmitting the compressed data packets.
[0073] Clause 28. The method according to any one of Clauses 19 to 27, further including decrypting the digital audio data, and the transmitting includes transmitting the decrypted data packets.
[0074] Clause 29 further includes storing the digital audio data in a buffer, and the storing is performed by the method according to any one of Clauses 19 to 28, which is managed by the local asynchronous media clock.
[0075] Clause 30, wherein the transmitting is performed by a first audio device, and the method further includes: generating a second local asynchronous media clock using a second crystal oscillator; receiving the digital audio data via the network by a second audio device; generating a second analog audio signal based on the received digital audio data using the second local asynchronous media clock, the method according to any one of Clauses 19 to 29.
[0076] Clause 31, the method according to Clause 30, further including buffering the received digital audio data at a buffer position based on a time stamp and / or a sequence number associated with the received digital audio data.
[0077] Clause 32, the method according to Clause 31, further including generating a second network clock based on the master clock of the network, wherein the buffer position is based on both the time stamp and the second network clock.
[0078] Clause 33, the method according to any one of Clauses 30 to 32, further including generating sound based on the second analog audio signal by using a speaker associated with the second audio device.
[0079] Clause 34 The local asynchronous media clock has a first nominal clock frequency, the second local asynchronous media clock has a second nominal clock frequency, and the first nominal clock frequency is different from the second nominal clock frequency, according to the method described in any one of Clauses 30 to 33.
[0080] Clause 35 The local asynchronous media clock has a first nominal clock frequency of one of 32 kHz or 48 kHz, and the second local asynchronous media clock has a different second nominal clock frequency of the other of 32 kHz or 48 kHz, according to the method described in any one of Clauses 30 to 33.
[0081] Clause 36 A first audio device, One or more processors, One or more computer-readable media storing instructions that, when executed by the one or more processors, cause the first audio device to perform the method described in any one of Clauses 19 to 29, the first audio device comprising the same.
[0082] Clause 37 A system, A first audio device, One or more processors, One or more computer-readable media storing instructions that, when executed by the one or more processors of the first audio device, cause the first audio device to perform the method described in any one of Clauses 19 to 29, the first audio device comprising the same, and A second audio device, One or more processors, One or more computer-readable media storing instructions that, when executed by the one or more processors of the second audio device, cause the second audio device to perform steps further described in any one of Clauses 30 to 35, the system comprising the second audio device comprising the same.
[0083] A non-transitory computer-readable medium that stores, when clause 38 is executed, an instruction to cause a first audio device to perform the method described in any one of clauses 19 to 29.
[0084] Although the embodiments have been described above, the features and / or steps of these embodiments can be combined, divided, omitted, reconfigured, modified, and / or enhanced in any desired manner. Various changes, modifications, and improvements will be readily envisioned by those skilled in the art. Such changes, modifications, and improvements are not explicitly described herein but are intended to be part of this description and are intended to be included within the spirit and scope of this disclosure. Accordingly, the foregoing description is merely illustrative and not restrictive.
Claims
1. It is an audio device, A local asynchronous media clock having a frequency variation of at least 1 ppm, Buffer and One or more processors, One or more computer-readable media for storing instructions, wherein the instructions, when executed by the one or more processors, Digital audio data is received via the network, based on a network clock synchronized with the master clock of the said network. The speed of the received digital audio data is compared with a threshold data speed based on the nominal speed of the local asynchronous media clock. At least a portion of the received digital audio data is stored in the buffer, where the storage of at least a portion of the received digital audio data is based on comparing the rate of the received digital audio data with the threshold data rate. Using the local asynchronous media clock, an audio signal is generated based on at least a portion of the digital audio data stored in the buffer. One or more computer-readable media that constitute the audio device to transmit the aforementioned audio signal for sound generation by a speaker, Audio devices, including those mentioned above.
2. The audio device according to claim 1, wherein the local asynchronous media clock includes one or more of the following: a crystal oscillator, a micro-electromechanical (MEMS) oscillator, a ceramic resonator, a surface acoustic wave (SAW) oscillator, or an inductor / capacitor (LC) oscillator.
3. The audio device according to claim 1, wherein, when the instruction is executed by one or more processors, the audio device is configured such that, based on the comparison, at least a portion of the received digital audio data is not stored in the buffer.
4. The digital audio data includes one or both of a plurality of timestamps or a plurality of packet sequence numbers, and the instruction, when executed by one or more processors, is based on one or both of the plurality of packet sequence numbers or the plurality of timestamps. To prevent at least a portion of the received digital audio data from being stored in the buffer, The locally generated audio data is saved to the buffer, Interpolating the data into the buffer, or The audio device according to claim 1, wherein the audio device is configured to perform at least one of the following: time correction of the data in the buffer.
5. When the instruction is executed by one or more processors, at least, Receiving multiple data packets including the aforementioned digital audio data, The audio device according to claim 1, wherein the audio device is configured to receive the digital audio data by extracting the digital audio data from the plurality of data packets.
6. The audio device according to claim 1, wherein the instruction, when executed by one or more processors, is configured to extract the digital audio data from a plurality of data packets at a timing managed by the local asynchronous media clock.
7. The audio device according to claim 1, wherein when the instruction is executed by one or more processors, the audio device is configured to store at least a portion of the received digital audio data in the buffer using timing managed by the local asynchronous media clock.
8. The audio device according to claim 1, wherein the instruction, when executed by one or more processors, causes at least a portion of the digital audio data to be restored, and generates at least the audio signal based on the restored at least portion of the digital audio data, thereby configuring the audio device to generate the audio signal based on at least a portion of the digital audio data.
9. The audio device according to claim 1, wherein the instruction, when executed by one or more processors, causes the processor to decode at least a portion of the digital audio data, and generates at least the audio signal based on the decoded at least portion of the digital audio data, thereby configuring the audio device to generate the audio signal based on at least a portion of the digital audio data.
10. The audio device according to claim 1, wherein the instruction, when executed by one or more processors, is configured to perform sample rate conversion on at least a portion of the digital audio data stored in the buffer.
11. It is an audio device, A local asynchronous media clock having a frequency variation of at least 1 ppm, One or more processors, One or more computer-readable media for storing instructions, wherein the instructions, when executed by the one or more processors, Based on the detected sound, an audio signal is received. Using the aforementioned local asynchronous media clock, digital audio data is generated based on the audio signal. Based on the network's master clock, a network clock is generated. One or more computer-readable media configured to transmit the digital audio data via the network based on the network clock, Audio devices, including those mentioned above.
12. The audio device according to claim 11, wherein the local asynchronous media clock includes one or more of the following: a crystal oscillator, a micro-electromechanical (MEMS) oscillator, a ceramic resonator, a surface acoustic wave (SAW) oscillator, or an inductor / capacitor (LC) oscillator.
13. The audio device according to claim 11, wherein, when the instruction is executed by one or more processors, the audio device is configured to packetize the digital audio data into multiple data packets using timing managed by the local asynchronous media clock.
14. The audio device according to claim 11, wherein when the instruction is executed by one or more processors, the audio device is configured to transmit the digital audio data based on the network clock by transmitting the digital audio data in at least a plurality of packets, each including a timestamp based on the network clock.
15. The audio device according to claim 11, wherein the instruction, when executed by one or more processors, is configured to transmit the digital audio data by transmitting the digital audio data in a plurality of packets at a rate based on the frequency of the local asynchronous media clock.
16. The audio device according to claim 11, wherein the audio device is configured to transmit the digital audio data by compressing the digital audio data and transmitting at least the compressed digital audio data when the instruction is executed by one or more processors.
17. The audio device according to claim 11, wherein the audio device is configured to transmit the digital audio data by encrypting the digital audio data and transmitting at least the encrypted digital audio data when the instruction is executed by one or more processors.
18. The audio device according to claim 11, wherein the instruction is configured to store the digital audio data in a buffer using timing managed by the local asynchronous media clock when executed by one or more processors.
19. Receiving digital audio data from a transmitting audio device via the network, based on a network clock synchronized with the network's master clock, The comparison involves comparing the speed of the received digital audio data with a threshold data speed, wherein the threshold data speed is based on the nominal speed of a local asynchronous media clock, and the local asynchronous media clock has a frequency different from that of the local asynchronous media clock of the transmitting audio device. The method of storing at least a portion of the received digital audio data in a buffer, wherein the storage of at least a portion of the received digital audio data is based on comparing the speed of the received digital audio data with the threshold data speed. Using the local asynchronous media clock, an audio signal is generated based on at least a portion of the digital audio data stored in the buffer. The speaker generates sound based on the aforementioned audio signal, A method that includes this.
20. The method according to claim 19, further comprising extracting the digital audio data from a plurality of data packets at a timing managed by the local asynchronous media clock.
21. The method according to claim 19, further comprising storing at least a portion of the received digital audio data in a buffer using timing managed by the local asynchronous media clock.