A method, a first audio device, and a computer program product, all arranged for synchronized media output in a multi-device system
By separating media playback data into audio and timing components and designating a master device, the system addresses latency issues in multi-device audio systems, ensuring synchronized audio playback and flexible hardware configurations.
Patent Information
- Application Number
- PCT/NL2025/050366
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-07-30
- Filing Date
- 2025-07-28
- Publication Date
- 2026-02-05
AI Technical Summary
Existing multi-device audio systems face challenges in synchronizing audio playback across devices with different latencies, leading to asynchronous sounds, and are limited by the need for bundled hardware solutions that restrict flexibility and interoperability.
A method and system that separates media playback data into audio data and playback timing information, communicated over different interfaces, allowing devices with varying latencies to synchronize playback by adjusting based on individual audio latencies and designating a master device to manage timing.
Ensures synchronized audio output across devices of different builds by distributing bandwidth and load, enabling flexible hardware configurations and secure, plug-and-play connectivity.
Smart Images

Figure NL2025050366_05022026_PF_FP_ABST
Abstract
Description
[0001] TITLE
[0002] A method, a first audio device, and a computer program product, all arranged for synchronized media output in a multi-device system.
[0003] TECHNICAL FIELD
[0004] The present disclosure relates to a method, a first audio device, and a computer program product, all arranged for synchronized media output in a multi-device system, such that audible sounds coming from multiple audio devices are perceived to arrive simultaneously.
[0005] BACKGROUND OF THE DISCLOSURE
[0006] In multi-device audio systems, wireless media data is output across multiple, coordinated audio devices. This allows for an immersive experience, where different audio channels of a media stream can be played from different physical locations, be it in one small or large space, or for playing audio at multiple locations.
[0007] To send the media data wirelessly across multiple devices, solutions today rely on using standardized technologies like Bluetooth or Wi-Fi, using proprietary solutions where general processing of the audio stream is bundled together with the audio computer hardware, running proprietary software. Those devices that have the same proprietary software are able to participate in the audio system, since their latencies are predefined by the manufacturer. Only then, the system allows the multiple devices to output audio synchronized enough to be indiscernible to human hearing.
[0008] Bluetooth has a number of usability challenges, because it relies on a processing device to be nearby. That processing device manages connectivity, remote service authentication, and decoding of the media stream. The media stream is received over the audio device’s wireless hardware, then decoded and output over the device’s audio hardware. This solution, therefore, requires a processing device with multiple wireless connections between the original audio data source and the audio hardware, resulting in reliability challenges, as well as connectivity issues related to range and number of devices that can connect to it. These challenges and limitations for Bluetooth lead to the development of bundled solutions, where the hardware in the processing device was merged with the audio hardware, creating one single device. This removes one of the wireless connections in the communication and improves the connectivity issues related to range and number of devices one can connect to. Moreso, it limits the number of playback timing variables to consider, because the time it takes for audio data to travel from the general processor to be human audible is fixed.
[0009] However, bundling all the computer hardware into one device does introduce other limitations. Once all the computer hardware is bundled in one device, it becomes impossible to mix and match audio / processing solutions to meet different price points. For instance, a simple and cheap speaker would still require the expensive processing circuitry with too beefy amps than needed, in case the audio and processing solutions are bundled. Additionally, it becomes difficult to have one company produce the bundled device, and another produce the audio device enclosures and speakers, since the company producing the bundled device has full control over a key component; the user experience of the overall audio device.
[0010] Both Bluetooth and proprietary Wi-Fi solutions also have limitations on the wireless technologies that can be used because the audio device must have enough bandwidth for both blocks of data of the media stream.
[0011] Therefore, it is the goal of the invention of the disclosure to provide a solution for synchronized media output in a multi-device system, which can be used by audio devices having completely different audio hardware resulting in different latencies, while still outputting audible sounds coming from the multiple audio devices that are emitted simultaneously.
[0012] SUMMARY OF THE DISCLOSURE
[0013] A first aspect of the disclosure pertains to a method for synchronized media output in a multi-device system. The multi-device system comprises multiple audio devices, and the method comprises the steps of: determining, by a first of said multiple audio devices, a latency related to processing of audio data by a processing unit, comprised by the first audio device, and an audible sound related to said audio data; providing, by said first audio device, to a second of said multiple audio devices, said determined latency; receiving, by said first audio device, from said second audio device, media playback data over a first interface, wherein said media playback data comprises audio data to be played by said first audio device, and playback timing information over a second interface, wherein the playback timing information determines when the first audio device is to play said audio data from said media playback data.
[0014] A multi-device system can be understood as an arrangement of multiple audio devices in communication with each other, which are intended to produce an audible sound in a synchronized manner, such that a user of the system cannot perceive any audible delays. The audio devices are ’’smart” devices such that they not only comprise a speaker (system), but also processing power to process communication signal and audio signals. Prior art multi-device systems typically only function when audio devices of identical built are connected, since then the latency related to processing of audio data by a processing unit until sound output is identical. However, the multi-device system of the disclosure is intended to work also with audio devices of a different built. Since these audio devices may have different latencies related to processing of audio data, additional information about the internals of the audio devices has to be communicated and the playback of the audible sound on all devices has to be adjusted accordingly. In an example, it may be adjusted to the “weakest link” (the audio device with the longest audio latency).
[0015] In order to achieve effective communication between the plurality of audio devices, one “master” device (second audio device) is used which will decide when a “slave” audio device (first audio device) is to play an audible sound. Some examples of audio devices are smart speakers (systems), computers, user equipment, such as a mobile phone, a smart hub, or the like. In essence, an audio device is defined as a device being able to produce an audible sound and at least having a processing unit able to receive and process media playback data.
[0016] Furthermore to communicate effectively, the media playback data is separated into audio data and playback timing information, which are communicated over a first interface and a second interface, respectively. Audio data can be understood as the information of an audible sound wave, such as mp3, wav, FLAG, AAC, or the like. The playback timing information on the other hand can be understood as a rhythm, a cadence, a heartbeat, or a centralized clock time to which the audible sound wave should be played against. Alternatively, the playback timing information could also be understood as a simple instruction to delay audio playback. Some examples of the first and the second interface may be one of the following communication methods like Wi-Fi, Bluetooth, LTE / 3G / 4G / 5G, Ethernet, Thread, Zigbee or the like.
[0017] Separating the media playback data has the advantage of distributing the required bandwidth and spreading the load on the network communication. This way, for instance, large data that would typically transfer slowly can be transmitted on an interface having adequate resources. For instance, it could be beneficial to transmit the audio data on an interface having a large bandwidth and transmitting the playback timing information on an interface having only a small bandwidth, such that the playback timing information can be transmitted with higher reliability. This way, unstable connections for audio devices which are moving, which are further removed the master device, or which are being obstructed will not experience unstable connections. These causes would otherwise lead to the emission of asynchronous audible sounds of the audio devices.
[0018] Note, however, that the first interface and the second interface are not required to be different communication method. They may also use the same communication method as long as they operate independently. The exact choice on which communication method and whether to use the same method depends on the situation and the implementation of the multi-device system and the environment it is in. Only in a first example of the method, the first interface actually differs from said second interface, such that two different communication methods are to be used.
[0019] The inventors have found that it may be beneficial to divide the media playback data into audio data and playback timing information. This insight is based on the concept that these two aspects may have different requirements: The audio data may need sufficient bandwidth to ensure that the data can be transmitted. The playback timing information may require accurate timing parameters to ensure that the audio devices can time synchronize with one another. Separating the audio data and the playback timing information thus allows for utilizing different interfaces, such that the interfaces can be matched to those specific requirements.
[0020] Furthermore, determining a latency, by the first of the multiple audio devices, related to processing of audio data by a processing unit, is to be understood as an audio latency, which would be the delay between sending the signal of the processing unit to output an audible sound until said audible sound is actually produced. This latency or delay could be measured, estimated, looked up, calculated, simulated, or traced. It is of importance to determine this audio latency and to communicate it with the second audio device, such that it can be assured that this audio latency is considered by the second audio device when addressing to the first audio device on when to produce the audible sound.
[0021] For instance, the audio latency may be determined during the manufacturing of the first audio device and stored on a local storage of the device, which latency information is then transferred to the second audio device during initialization of the audio devices in the multi-device system. Alternatively, the first audio device may comprise a microphone and the audio latency is simply measured by playing an audible sound, wherein the measurement is to be understood as the time tracked between sending the command to the speaker (system) by the processing unit of playing an audible sound until the moment the sound is observed at the microphone.
[0022] The processing unit may thus be understood to comprise a microprocessor, a processor, a FPGA, an ASIC, or the like. It may also further comprise a storage memory, network modules, and / or sensors. The storage memory may be used to store history data of previous latencies of audio devices (itself and other devices) or audio playback data. The Network modules may be comprised in the processing unit as well to transfer any of these data to any other audio device. Sensors may comprise the aforementioned microphone or any other sensor.
[0023] Lastly, the audible sound can be understood in the classical sense of such multidevice speakers as music or movie audio. However, it is not limited thereto. For instance, Artificial Intelligence-based sounds are becoming more prominent. These sounds range from adaptive soundscapes, which immerse the user of the system into a different world, to personalized and interactive audio experiences, wherein the user might interact with Al- assistants, speakbots / chatbots, or the like.
[0024] In another example of the method, said method further comprises the steps of discovering, by said first audio device, other audio devices of said multiple audio devices in a proximity of said first audio device.
[0025] Discovery allows the devices in the network to identify and locate the other devices. This helps with monitoring the latencies and administrating the connectivity. Furthermore, continuous discovery allows the network to act as self-healing, such that routing of media playback data can be adjusted in case of blockage or failure. Additionally, this feature allows for plug-and-play functionality aiding towards user-friendliness, wherein new devices are automatically recognized and assigned, thereby simplifying setup and reducing intervention of the user of the multi-device system.
[0026] In a further specified example of the method, said method comprises the step of authenticating, by said first audio device, to verify other audio devices using public key cryptography.
[0027] Public key cryptography allows for simplified key management, since only the devices’ own private keys need to be managed. Furthermore, public key cryptography allows for secure communication with any audio device in the multi-device system, even if no previous connection has been made and / or key has been shared. This functionality also aids the plug-and-play functionality of the multi-device system, thereby improving the userexperience of the multi-device system.
[0028] In yet another example of the method, wherein the first device is arranged to communicate using multiple interfaces, the method comprises the step of selecting, by said first audio device, one of said multiple interfaces as said first and / or second interface based on latency information measured over said multiple interfaces.
[0029] The multi-device system may establish multiple different communication interfaces all having different communication characteristics. Based on the actual implementation of the system and its environment, it may be beneficial to select the first and / or the second interface based on latency information of the different interfaces. This way, an optimal interface may be chosen for communication in that particular moment for that particular system. Moreover, this selection may also occur during use of the system, such that a stable connection between the plurality of audio devices can be assured in the system.
[0030] In a further example of the method, said method further comprises the step of receiving, by said first audio device, from said second audio device, latency information related to processing of audio data by a processing unit, comprised by the second audio device, and an audible sound related to said audio data, by said second audio device; and selecting, by said first audio device, said second audio device as a master audio device, such that said second audio device is to distribute at least said playback timing information.
[0031] Exchanging audio latency information between the audio devices allows the system to output synchronous audible sounds, since all the individual audio latencies can be taken into account by the audio devices when playing an audible sound. Only this way it can be assured that synchronized audio is outputted for audio devices even if they are of a different built, thereby having different audio latencies. The synchronized playback is also assured due to the separation of the playback media data, this way the audio data can be loaded into a buffer memory of the audio device and the playback timing information can dictate the synchronized playback. For instance, now the first audio device may output an audible sound based on its own audio latency, the audio latency of the second audio device, and the playback timing information. In one particular example, the longest (audio + network) latency may be used to dictate the delay times for playing the audible sound to all the other audio devices.
[0032] Additionally, selecting the second audio device as a master ensures that only one device in the plurality of audio devices can distribute at least the playback timing information. This master may be chosen based on the reliability and / or the network latency of the network connection of the second audio device. For instance, it may also be considered to check whether the audio device is plugged into the wall for power, what its network latency is, what its history of latency and connectivity is etc. Additionally, it may be the case that the master is selected again over time, based on updated reliability variables, such that the optimum performance of the multi-device system is guaranteed. Having only one stable audio device act as the master will assure that the playback of the audio sound is managed centrally. Otherwise, multiple rhythms, cadences or heartbeats would clash and none of the audio devices would know when to output the audible sound.
[0033] All in all, with the above-described features the method for synchronized media output in a multi-device system according to the disclosure is able to output synchronized audible sounds, especially in a system wherein any of the plurality of devices are of a different built, thereby having different audio latencies. By splitting the communication scheme into two parts (audio data and playback timing information), transmitted over two interfaces the system is able to ensure synchronized audio playback.
[0034] The disclosure pertains, in another aspect, to a first audio device arranged for synchronized media output in a multi-device system, said multi-device system comprising multiple audio devices, said first audio device arranged to execute an instance, wherein said instance is arranged to: determine a latency related to processing of audio data by a processing unit, comprised by the first audio device, and an audible sound related to said audio data; provide, to a second of said multiple audio devices, said determined latency; receive, from said second audio device, media playback data over a first interface, wherein said media playback data comprises audio data to be played by said first audio device, and playback timing information over a second interface, wherein the playback timing information determines when the first audio device is to play said audio data from said media playback data.
[0035] The first step of determining a latency related to processing of audio data can be understood as an audio latency. This delay could be measured, estimated, looked up, calculated, simulated, or traced. For instance, the audio latency may be determined during the manufacturing of the first audio device and stored on a local storage of the device. Alternatively, the first audio device may comprise a microphone and the audio latency is measured by playing an audible sound, wherein the measurement is to be understood as the time tracked between detecting the sound at the microphone and the sending the signal to play an audible sound by the processing unit.
[0036] In the second step of providing said determined latency, the determined audio latency is communicated with the second audio device such that the audio latency can be utilized for indicating playback adjustment by the second audio device to the first audio device when addressing to when the audible sound should be produced.
[0037] Lastly, in the third step of receiving media playback data over a first interface and playback timing information over a second interface, an effective way of communication is established between the first and the second audio devices. Herein, the second audio device (master) will decide when the first audio device (slave) is to play an audible sound. Additionally, by separating the media playback data and the playback timing information, the required bandwidth can be distributed and the load on the network communication can be spread. This way, for instance, large data that would typically transfer slowly can be transmitted on an interface having adequate resources, whereas the playback timing information can be transmitted on a faster channel needing less resources. Examples of such interfaces may be communication methods like Wi-Fi, Bluetooth, LTE / 3G / 4G / 5G, Ethernet, Thread or the like.
[0038] Note, however, that the first interface and the second interface are not required to be different communication method. They may also be the same communication method as long as they operate independently. The exact choice of communication method and whether to use the same method depends on the situation and the implementation of the multi-device system and the environment it is in. Only in a first example of the first audio device, the first interface actually differs from said second interface, such that two different communication methods are used.
[0039] An audible sound can be understood in the classical sense of such multi-device speakers as music or movie audio. However, it is not limited thereto. For instance, Artificial Intelligence-based sounds are becoming more prominent. These sounds range from adaptive soundscapes, which immerse the user of the system into a different world, to personalized and interactive audio experiences, wherein the user might interact with Al- assistants, speakbots / chatbots, or the like.
[0040] In a further example of the first audio device, said first audio device comprises audio hardware comprising said processing unit, and wherein said first audio device comprising a processor arranged for executing said instance. The processing unit may thus comprise a microprocessor, a processor, a FPGA, an ASIC, or the like for processing and may further comprise a storage memory, network modules, and / or sensors for storing, communicating, and measuring, respectively.
[0041] In an example of the first audio device, the audio hardware and the processor are arranged to communicate via a physical interface. Such physical interfaces may be included but are not limited to I2S, LIART, USB, and Ethernet, such that an abstraction layer of the audio is created, which allows the audio hardware to be exchangeable and modifiable, while the processing unit ensuring connectivity of the audio devices may remain.
[0042] In yet another example of the first audio device, the instance is further arranged to perform: discovering other audio devices of said multiple audio devices in a proximity of said first audio device.
[0043] The discovering of other audio devices is of importance in order to establish a communication between the devices. Furthermore, this discovering may be performed at a regular time intervals, such that changes in the connectivity between the multiple audio devices can be picked up and new connections can be established. Additionally, discovering and connecting to further audio devices of the multiple of audio devices may also allow for the generation of a network, wherein a communication path may run over a plurality of audio devices. For instance, the second audio device may not be in direct communication with the first audio devices, but through a third audio device the multimedia playback data the media playback data and / or the playback timing information could still be communicated. In a further example of the first audio device, the instance is further arranged to perform: authenticating to discovered other audio devices using public key cryptography.
[0044] Public key cryptography allows for simplified key management, since only the devices’ own private keys need to be managed. Furthermore, public key cryptography allows for secure communication with any audio device in the multi-device system, even if no previous connection has been made and / or key has been shared. This aids the plug- and-play functionality of the multi-device system.
[0045] In another example of the first audio device, the first device is arranged to communicate using multiple interfaces, wherein the instance is further arranged to perform: selecting one of said multiple interfaces as said first and / or second interface based on latency information measured over said multiple interfaces.
[0046] Selecting the first and / or second interface based on network latency information may be understood as selecting which type of interface is more adequate for the specific situation of wherein the multi-device system is present. For instance, connections may be made over three interfaces such as Wi-Fi, Bluetooth, and 5G. Then the instance selects which one or two of these three interfaces are the best to use based on the individual network latency information of each of those interfaces.
[0047] Note that this selection process may be occurring continuously such that any of said first and / or second interface could be switched based on occurrences in the connectivity between the multiple audio devices.
[0048] The instance may further be arranged to perform, according to another example of the first audio device of the disclosure: receiving, from said second audio device, latency related to processing of audio data by a processing unit, comprised by the second audio device, and an audible sound related to said audio data, by said second audio device; selecting, said second audio device as a master audio device, such that said second audio device is to distribute at least said playback timing information.
[0049] Receiving information of the audio latency of other audio devices is of importance to achieve synchronized media output. Namely, only with the information of the audio latencies of other devices or media playback data inherently comprising said latency information synchronous playback can be assured. For instance, the second audio device may communicate an individual delay time to every audio device, such that each device upon receiving of playback timing information knows to output the audible sound after that individual delay time has lapsed. Alternatively, the second audio device may broadcast audio playback data specifically for the first audio device, which has been shifted in time based on the audio latency. Then upon receival of the playback timing information, the first audio device simply can playback the audio playback data at the timing of the playback timing information.
[0050] The second step of selecting the second audio device as a master is of importance, since one of the audio devices must emit the playback timing information, such as a heartbeat, cadence, or centralized clock-information. This audio device may be selected based on its position in the network of audio devices related to all the individual latencies (network and / or audio). Note that, this device does not necessarily have to have the shortest latencies, but could also be selected based on how stable the connections with the other audio devices are. Furthermore, the selection of the second audio device as a master may occur continuously, such that variations / changes in the connectivity between the audio devices can be accommodated.
[0051] A last aspect of the disclosure pertains to a computer program product comprising a computer readable medium having instructions stored thereon which, when executed by an instance of a first audio device, cause said instance to perform any of the examples of the method as described above.
[0052] All in all, with the above-described features the first audio device and the computer program output synchronized media in relation to a multi-device system. This system in particular is tailored towards comprising a plurality of devices which are of a different built therefore having different audio latencies. Said system is able to ensure synchronized audio playback by splitting the communication scheme into two different parts, transmitted over two interfaces.
[0053] SHORT DESCRIPTION OF THE FIGURES
[0054] Fig. 1 shows the mains steps of the method according to the disclosure.
[0055] Fig. 2 shows an example of a multi-device system comprising three audio devices and their respective communication in the discovery phase. Fig. 3 shows an example of a multi-device system comprising three audio devices and their respective communication in the operation phase, where they exchange audio playback data and playback timing information.
[0056] Fig. 4 shows another example of a multi-device system comprising three audio devices and their respective communication in the operation phase, where they exchange audio playback data and playback timing information over two different interfaces.
[0057] Fig. 5 shows yet another example of a multi-device system comprising three audio devices and their respective communication in the operation phase, where they exchange audio playback data and playback timing information over two different interfaces, wherein one of the audio devices loses connection.
[0058] Fig. 6 shows another example of a multi-device system comprising three audio devices and their respective communication in both the discovery and operation phase, as continuation on Fig. 5, wherein the lost audio device is reconnected through discovery.
[0059] FULL DESCRIPTION OF THE DISCLOSURE
[0060] An explanation of the invention according to the disclosure is given with reference to the figures. For better understanding, similar parts and features of the figures are denoted with the same reference number as used in the text. Note that, these figures are only given as particular examples of the invention such that the invention may be understood better. These figures are therefore not meant to limit the disclosure.
[0061] It may further be understood that throughout this disclosure, when latency is mentioned, audio latency, being the delay between sending a signal of a processing unit to produce an audible sound until the production of said sound, should be understood. If other latencies, such as network latency, is meant this would be explicitly mentioned.
[0062] In Fig. 1 , the main steps of the method for synchronized media output in a multidevice system are shown. The multi-device system comprises multiple audio devices, and the method comprising the steps of:
[0063] 1) determining, by a first of said multiple audio devices, a latency related to processing of audio data by a processing unit, comprised by the first audio device, and an audible sound related to said audio data;
[0064] 2) providing, by said first audio device, to a second of said multiple audio devices, said determined latency; and 3) receiving, by said first audio device, from said second audio device, media playback data over a first interface, wherein said media playback data comprises audio data to be played by said first audio device, and playback timing information over a second interface, wherein the playback timing information determines when the first audio device is to play said audio data from said media playback data.
[0065] First and foremost, this method, audio device, and computer software are especially tailored towards a multi-device system which comprises audio devices of a different built. This means that every audio device has its own unique latency between the processing unit and the output of an audible sound, which is not a-priori known by the other audio devices. Therefore, it is beneficial to set-up an efficient communication structure, using the above- mentioned steps, such that the audio devices can output synchronized audible sounds. The inventors have found that this can be achieved when selecting one “master” device (the second audio device) and have the rest (including the first audio device) function as “slave” devices in the multi-device system.
[0066] These multi-device systems are intended to produce an audible sound in a synchronized manner, such that a user of the system cannot perceive any audible delays. Prior art multi-device systems only contain devices of identical built. In that case, the task of synchronized playback is easy and trivial, because the delay from the processing unit to the output of an audible sound are equal among all audio devices. Then only the network latency has to be corrected for. Note that, 2 different version of a same brand speaker system are not considered being of a different built in the eyes of the inventors. Namely, in most cases the exact same or similar sound boards are used, causing these devices to still only rely on network latency to adjust the synchronization of an audible sound. Yet, the multi-device system of the disclosure is intended to work with audio devices of different built needing not only adjustment based on the network latency, but also on the audio latency of each individual speaker. The latter not a-priori known to any of the other audio devices in the multi-device system.
[0067] Further, the inventors have found that by separating the media playback data into audio data and playback timing information, and to communicate those over a first interface and a second interface, respectively, the communication would be more stable and more effective, since bulky audio data can be sent over a large bandwidth interface shortly before the playback timing information is sent to indicate when to output the audible sound based on the just received audio data. The inventors thus found that, separating the media playback data has the advantage of distributing the required bandwidth and spreading the load on the network communication. This way, the audio data may be (temporarily) stored in a memory buffer on the audio devices to await audible emission upon receival of the correct cue in the playback timing information.
[0068] Even though the first interface and the second interface are different interfaces; the proposed method for synchronized media output does not require them to be different communication methods. They merely must be separate instances, communicating over their own connection channels and could thus also be utilizing the same communication method. Some examples of these interfaces are communication methods like Wi-Fi, Bluetooth, LTE / 3G / 4G / 5G, Ethernet, Thread or the like. The actual implementation on which communication method to use heavily depends on the situation and the local environment of the multi-device system. However, in a preferable example, the communication method of the first interface actually differs from said second interface, such that two different communication methods are used.
[0069] Thus far, latency related to processing of audio data by a processing unit has been mentioned, but not specified any further. This latency can be understood as audio latency, which in other words means a delay that is intrinsic to the audio device, which is, for instance, the time needed for one audio device to playback an audible sound upon receival and processing of the audio data. This audio latency is thus different from network / communication latency, which will also be present based on the communication method / interface that is used between the two audio devices, the physical environment which the two audio devices are in, and their respective built of communication and / or processing units. This network latency may also be considered and utilized to adjust the playback timing of a particular audio device, but the disclosure pertains to adjusting the playback based on at least the audio latency, since that is needed to tackle the problem of asynchronous playback of audio devices of different built in a multi-device system.
[0070] The audio latency can be measured, estimated, looked up, calculated, simulated, or traced. It is important to determine this audio latency and to communicate it with the second audio device, such that it can be assured that this audio latency is considered by the second audio device when addressing to the first audio device on when to produce the audible sound, especially since the audio latency of the first audio device is not a-priori known to the second audio device. For instance, the audio latency may be determined during the manufacturing of the first audio device and stored on a local storage of the device. This audio latency information may then be transferred to the second audio device during the initialization phase. Alternatively, the first audio device may comprise a microphone and the audio latency is simply measured, during initialization, by playing an audible sound and measuring the delay time. Herein, the measurement of the audio latency is to be understood as the time tracked between measuring the sound at the microphone and the sending the command of playing the sound by the processing unit.
[0071] To meet different price points of all the audio devices, the audio device may comprise audio hardware comprising said processing unit, and wherein said processing unit comprising a processor arranged for executing said instance. Part of the audio hardware (not being the processing unit) may be governed with production of a sound from a digital and / or electrical sound wave signal, whereas the processing unit may be governed with the communication with and / or the processing signals from the other audio devices. The audio hardware and the processing unit are arranged to communicate via a physical interface, which allows the audio hardware to be exchangeable and modifiable, while the processing unit ensures that connectivity of the audio devices may remain.
[0072] Separating the processing computer hardware from the audio hardware and communicating between the components over physical interfaces, allows products to be created for different price points, thereby creating interoperability and reliability between the different audio devices. Namely, this allows speaker builders to focus on building the speaker and the electronics to drive it, whereas the communication, connectivity, and userexperience can be taken care of by the invention of this disclosure. It also allows for audio products of different companies to include these different components allowing interoperability between audio devices of a different built, thereby having different audio latencies.
[0073] The downfall of this approach is that separating the audio hardware into two parts does create additional audio latency, since the time it takes for audio data to go from the processing unit to an audible sound will be different for all the different audio devices.
[0074] To solve this problem, the inventors have found that separating the timing data from the audio stream data allows for ensuring synchronized multi-media playback. Furthermore, it allows the use of wireless technologies that would otherwise not be applicable, because they do not provide enough bandwidth for both blocks of data to be transmitted together. Audio devices in the multi-device system can then communicate over whatever wireless technology will produce the best results, as opposed to being limited to a predefined wireless technology with the required bandwidth.
[0075] The method according to the disclosure executes on top of a hardware abstraction layer, wherein the audio hardware is exchangeable and communicates with a second processing part via standard physical interfaces, including but not limited to I2S, UART, USB, and Ethernet. During development and manufacturing, the delay between processor execution and audible sound will be measured and stored in a local memory.
[0076] In Fig. 2 an example of a multi-device system is shown comprising three audio devices, which go through an initialization process. During the initialization process, the audio devices utilize access, given to them by the user during set up, to a wide area network, such as Wi-Fi, which will allow them to communicate with a set of remote services. Then as a first step using this wide area connection and the remote services, the audio devices will discover or be provided a list of other audio devices in close physical proximity. Many wireless technologies have the ability to send out broadcast messages on different frequencies, or on the local network. These messages can be seen and responded to by close by devices running the method of the application. Examples of the discovery process include dns-sd on technologies like Thread or Wi-Fi, or Discovery in Bluetooth.
[0077] In Fig. 3 the subsequent authentication between each of the audio devices using public key cryptography is shown. Here, each audio device communicates a public key with the other audio devices and keeps a private key to itself. To establish a data communication channel between two specific audio devices, first a with the private key encrypted message is sent to said specific audio device, which uses the public key to decrypt the message. Only upon success, the specific audio device will let the original sending audio device know that the data communication channel may be established.
[0078] After that, all other possible communication interfaces are established among the two audio devices. This may also be done using public key cryptography for each interface, but may be understood that it is not necessary, since the already established interface / data communication channel may be used to prove the identity of both audio devices to each other. For this process, we speak of the exchanging of information phase, since a lot of information (about the identity) of the audio devices is communicated.
[0079] Later in that same phase, the network latencies over all the different interfaces are determined and communicated among the audio devices to find which interface is the best to use in the current environment of the multi-device system, based on the connection speed, bandwidth, and stability of the interfaces between all audio devices in the multidevice system.
[0080] Later during the use phase, the established communication interfaces will be tested periodically using the same process to check whether the optimal interface is still being utilized. This way, changes in the environment and connectivity can be observed and mitigated by for instance switching to another interface.
[0081] Once a communication interface has been chosen, each audio device will exchange additional information, including but not limited to, the audio latency between the processor and an audible sound, power source information, computing hardware resource information, configured remote audio source, and network connection information. The audio devices will then elect a single device based on the exchanged information to act as a master to dictate a sense of time, a cadence, a centralized clock, or any other form of timing.
[0082] In Fig. 3 for instance, the elected master could be audio device 100i . After election, it calculates a timescale to be used, wherein the timescale is calculated based on the information the other audio devices have provided to said master, including but not limited to, audio latency between the processor and an audible sound, configured remote audio source, network latency between the elected master and each other audio device.
[0083] Then as shown in Fig. 4, the elected master 100i communicates with remote audio devices to control their audio playback by using two different interfaces. It distributes information about that playback, including but not limited to, a list of upcoming sources of audio, and playback commands over a first interface. The elected master also utilizes a second interface of all its active communication interfaces to broadcast the current time, in the timescale that was calculated and distributed.
[0084] All three audio devices 100i IOO2 IOO3 store and use the distributed information about upcoming sources of audio to ensure there is a local copy available, and that there is enough information to playback the content at the point in the timescale the master dictates. This process is done independently, and the exact size of the buffer of the information in the list of upcoming sources is determined by each individual audio device, based on the individual audio device’s available resources.
[0085] The audio devices IOO2 - IOO3 that are not elected as master use the interface on which the master’s time is being broadcast to adjust their internal sense of time accordingly, such that the output of an audible sound becomes synchronous with the other audio devices in the multi-device system. The adjustment of time may even be based on (network + audio) latencies among all audio devices. How often the audio device checks its internal playback time against the master is determined by the audio device, based on but not limited to, available resources and the magnitude of recent adjustments.
[0086] In Figs. 7 and 8, examples of more complex networks of the multi-device system are shown, now comprising four audio devices. The examples of Figs. 7 and 8 are similar to the example of Fig. 4 in that the same audio device 100i is the master device and therefore devices IOO2 - IOO4 are slave devices and that all audio devices are initialized and are thus in the use phase by actively communicating with each other over at least two interfaces.
[0087] In Fig. 7 the audio device IOO4 is in direct communication with the master device 100i . The audio device IOO4 is thus simply receiving the audio data and the playback timing information as the other two audio devices IOO2 - IOO3 are receiving. However, there is no direct communication between the audio device IOO4 and the other audio devices IOO2 - 10O3. This does not hinder the functioning of the audio device 10O4 directly, but would mean that if the master device 100i becomes unresponsive, the audio device IOO4 would also lose connection with the other audio devices in the multi-device system. This is because the audio device IOO4 does not have any direct communion connection with any of the other slave devices IOO2 - IOO3.
[0088] To clarify, a device may become unresponsive, in case its battery dies, its being unplugged from the wall, or when its communication path is being changed or blocked, for instance by moving one audio device away from the rest.
[0089] In Fig. 8, a similar multi-device system configuration has yet again a different behavior, since the audio device IOO4 is not in direct communication with the master audio device 10O1. In this case, the master audio device sends the audio data and playback timing data through one of the other two audio devices IOO2 -IOO3. It should be noted that, in that case the audio latency is only dictated by the audio device IOO4, but the network latency has become combined with which ever audio device IOO2 - IOO3 the master is communicating through.
[0090] Fig 5 shows the unfortunate event that the master device 100i has become unresponsive or unavailable. In that case, the other audio devices will perform the abovedescribed process again to elect a new master device. For instance audio device IOO3 may be elected as new master device as shown in Fig. 5. After completion, the audio devices follow the phases as described above, but this time with a new master device IOO3. Lastly in Fig. 6, the event is shown wherein a new (or previously disappeared) audio device appears in the proximity of the multi-device system. In that case, the two existing audio devices IOO2 - IOO3 will simply continue the communication between them, such that the process with which the audio devices IOO2 - IOO3 are busy is not interrupted. That way, the user does not experience that the discovery, authentication, and information exchange phases are initiated with the new audio device 100i by each of the two existing audio devices IOO2 - IOO3. Only, after the new audio device 100i is completely initialized, all interfaces are established, and all latency information is exchanged, then the new audio device will participate in the output of an audible sound. Subsequently, a new check to elect a new master device may be performed as well. In general, these processes occur such that the user does not experience a change in the performance of the audio device. In other words these processes are running in the background and ensure that the user has a pleasant and friendly experience.
[0091] All in all, the above-described method, audio device and computer software based on the examples of the figures ensures synchronized media output in a multi-device system. The method is able to output synchronized audible sounds in a plurality of audio devices even if they are of a different built, thereby having different audio latencies. Therefore, overcoming the problems of prior-art methods for multi-device systems, which only rely on adjusting the playback of an audible sound based on the network latencies between the audio devices. Since the audio latencies differ between audio devices of a different built, asynchronous audio would still be output. Typically, prior art multi-audio systems therefore only allow devices of similar built to connect thereto.
[0092] The method, audio device and computer software according to the disclosure overcome these problems by splitting the media playback data into audio data and playback timing information, which are communicated over two different interfaces. This allows the audio devices to communicate the different audio latency information among each other, such that the playback timing information can be adjusted accordingly. Moreover, one leader is elected, which dictates the timing of the audible sound, whereas the other audio devices adjust they playback timing based on transferred latency information of the other audio devices. Reference numbers
[0093] 1000 multi-device system
[0094] IOO1-IOO2-IOO3-IOO4 audio device one, two, three, four
[0095] 110 first audio device
[0096] 120 second audio device
[0097] 200 master
[0098] 301 first interface
[0099] 302 second interface
Claims
CLAIMS1. A method for synchronized media output in a multi-device system, said multidevice system comprising multiple audio devices, said method comprising the steps of: determining, by a first of said multiple audio devices, a latency related to processing of audio data by a processing unit, comprised by the first audio device, and an audible sound related to said audio data; providing, by said first audio device, to a second of said multiple audio devices, said determined latency; receiving, by said first audio device, from said second audio device, media playback data over a first interface, wherein said media playback data comprises audio data to be played by said first audio device, and playback timing information over a second interface, wherein the playback timing information determines when the first audio device is to play said audio data from said media playback data.
2. The method in accordance with claim 1 , wherein said first interface differs from said second interface.
3. The method in accordance with any of the previous claims, wherein said method further comprises the steps of: discovering, by said first audio device, other audio devices of said multiple audio devices in a proximity of said first audio device.
4. The method in accordance with claim 3, wherein said method comprises the step of: authenticating, by said first audio device, to verify other audio devices using public key cryptography.
5. The method in accordance with any of the previous claims, wherein said first audio device is arranged to communicate using multiple interfaces, wherein the method comprises the step of:selecting, by said first audio device, one of said multiple interfaces as said first and / or second interface based on latency information measured over said multiple interfaces.
6. The method in accordance with any of the previous claims, wherein said method further comprises the step of: receiving, by said first audio device, from said second audio device, latency information related to processing of audio data by a processing unit, comprised by the second audio device, and an audible sound related to said audio data, by said second audio device; selecting, by said first audio device, said second audio device as a master audio device, such that said second audio device is to distribute at least said playback timing information.
7. A first audio device arranged for synchronized media output in a multi-device system, said multi-device system comprising multiple audio devices, said first audio device arranged to execute an instance, wherein said instance is arranged to: determine a latency related to processing of audio data by a processing unit, comprised by the first audio device, and an audible sound related to said audio data; provide, to a second of said multiple audio devices, said determined latency; receive, from said second audio device, media playback data over a first interface, wherein said media playback data comprises audio data to be played by said first audio device, and playback timing information over a second interface, wherein the playback timing information determines when the first audio device is to play said audio data from said media playback data.
8. The first audio device in accordance with claim 7, wherein said first interface differs from said second interface.
9. The first audio device in accordance with any of the claims 7 - 8, wherein said first audio device comprises audio hardware comprising said processing unit, and wherein said processing unit comprising a processor arranged for executing said instance.
10. The first audio device in accordance with claim 9, wherein said audio hardware and said processor are arranged to communicate via a physical interface.11 . The first audio device in accordance with any of the previous claims 7 - 10, wherein said instance is further arranged to perform: discovering other audio devices of said multiple audio devices in a proximity of said first audio device.
12. The first audio device in accordance with claim 11 , wherein said instance is further arranged to perform: authenticating to verify other audio devices using public key cryptography.
13. The first audio device in accordance with any of the claims 7 - 12, wherein said first audio device is arranged to communicate using multiple interfaces, wherein the instance is further arranged to perform: selecting one of said multiple interfaces as said first and / or second interface based on latency information measured over said multiple interfaces.
14. The first audio device in accordance with any of the claims 7 - 13, wherein said instance is further arranged to perform: receiving, from said second audio device, latency information related to processing of audio data by a processing unit, comprised by the second audio device, and an audible sound related to said audio data, by said second audio device; selecting, said second audio device as a master audio device, such that said second audio device is to distribute at least said playback timing information.
15. A computer program product comprising a computer readable medium having instructions stored thereon which, when executed by an instance of a first and / or second audio device, cause said instance to perform any of the methods of claims 1 - 6.
Citation Information
Patent Citations
Methods and apparatus for using the unused tv spectrum by devices supporting several technologies
EP2685773A1
Security mechanism for wireless video area networks
US20090010438A1
Systems and methods for syncronizing multiple electronic devices
WO2014179155A2
Distributed synchronization
WO2020081554A1
Latency negotiation in a heterogeneous network of synchronized speakers
WO2020163476A1