Synchronization of audio stream exchanges during an audio conference
The method synchronizes audio data between client and server devices using timestamp information and resampling to address echo and intelligibility issues in audio conferences, enhancing audio quality by distributing processing load and ensuring synchronized playback across multiple devices.
Patent Information
- Authority / Receiving Office
- FR · FR
- Patent Type
- Patents
- Current Assignee / Owner
- ORANGE SA
- Filing Date
- 2023-10-18
- Publication Date
- 2026-04-10
AI Technical Summary
Current audio conference systems fail to synchronize microphone and speaker audio signals accurately, leading to echo and intelligibility issues for participants in the same room, especially when using a single device for advanced synchronization processing, which can overload equipment and reduce processing reliability.
A method for synchronizing audio data between client devices and a server device using timestamp information to distribute processing load, resample audio packets based on clock drift, and align audio streams for synchronized playback, reducing echo and improving intelligibility by using client devices' microphones and speakers as close as possible to participants.
The method ensures reliable synchronization of audio signals across multiple devices, reducing echo and improving audio intelligibility for participants in the same room by distributing processing load and using timestamp-based synchronization, resulting in clearer audio playback.
Smart Images

Figure 00000026_0000 
Figure 00000027_0000 
Figure 00000027_0001
Abstract
Description
Title of the invention: Synchronization of audio stream exchanges during an audio conference technical field
[0001] The present invention relates to the field of audio or video conferencing and more particularly to improving audio quality when several people participating in the conference share the same environment. Prior art
[0002] Conference systems (audio and video conferencing) allow for virtual meetings between remote individuals. Due to the Covid crisis and the associated lockdowns, this use has become widespread, and today regular teleworking has made this tool indispensable.
[0003] These systems can be presented in standalone form to equip so-called dedicated meeting rooms, or in software form on PC or on your smart mobile phone so-called smartphone via applications such as Teams®, Zoom® or Skype® to name just a few.
[0004] These applications are installed on communication devices comprising software modules called conference clients that control this type of application. The devices can connect to each other directly (in "peer-to-peer" mode), or use a conference bridge, which allows the signals from the participants' microphones to be aggregated and sent back to each participant an individualized audio mix (the sum of all the microphones except their own).
[0005] In the specific case of a so-called "hybrid" meeting, that is to say a meeting involving at least one meeting room and one or more remote participants (or even one or more other remote meeting rooms), part of the participants (the co-participants) are co-located in the same space (the meeting room).
[0006] Each participant may bring their PC and / or smartphone into the room to connect individually to the virtual meeting in order to follow the visual elements of the meeting (video, presentation, etc.). However, it is necessary that, with the exception of one participant, whom we will call the primary participant, all participants disable their microphone and speaker on their respective devices.
[0007] Indeed: - If one of the co-participants (whom we will call the secondary co-participant) leaves their microphone activated, the remote participants will receive audio signals from two microphones (primary and secondary co-participant) at different times because the systems Conferences do not guarantee synchronization of sound sources. This results in an intolerable echo for remote participants. It should be noted that if the microphone signals of conference clients were synchronized with an accuracy of less than approximately 50ms, the echo would become tolerable but still audible, and virtually inaudible below 20ms. However, the primary participant's speaker will transmit the sound picked up by the secondary participant's microphone, thus creating an intolerable echo in the meeting room (a participant in the meeting room will be heard live, then through the primary participant's speaker with a delay of several hundred milliseconds).
[0008] - if one of the secondary co-participants leaves their speaker activated, all the co Participants will hear audio signals from remote participants through the two speakers (primary and secondary co-participant) at different times. This also results in an intolerable echo for the co-participants.
[0009] Since current conference systems do not guarantee synchronization of microphone and speaker audio signals, the sound in the meeting room is generally only picked up by the microphone of the main participant, and the sound from remote participants is only reproduced by the speaker of the main participant. As a result, participants located far from the main participant:
[0010] - will be difficult for remote participants to understand, due to the attenuation acoustics related to their distance from the microphone, but also due to the reverberation of the room which reduces intelligibility indicators
[0011] - will have difficulty understanding remote participants, due to the attenuation acoustics related to their distance from the loudspeaker, but also due to the reverberation of the room which reduces intelligibility indices.
[0012] US patent document 11425258 attempts to solve the echo problem by dedicating one of the devices in the same space, or the conference bridge, to advanced synchronization processing of the data it receives from the other devices in the common space before transmitting a mixed stream to other remote participants.
[0013] This type of system focuses the processing load on a single device, which can be problematic for the equipment in question, which may not be designed to absorb such a processing load, especially if the number of participants is large.
[0014] On the other hand, the synchronization implemented in this document is performed on data stored in the memory of a common device; it is therefore based on a correlation analysis of the received and stored audio signals in order to obtain the time differences between the signals, independent of latencies due to acoustic propagation distances. This type of analysis is costly in terms of processing load and lack of reliability since it relies on audio data which can be noisy or unstable.
[0015] There is therefore a need to improve state-of-the-art methods in this type of situation. Description of the invention
[0016] The invention improves upon the state of the art.
[0017] To this end, the invention relates to a method for synchronizing audio data during an audio conference between a plurality of participants via client devices, the method being implemented by a first client device of the conference sharing the same common space as a second client device and in which the following steps are executed: - exchange of timestamp information between the client device and a server device connected to the client devices in the common space; - calculation of a clock drift of the client device relative to the server device from the information exchanged; - when capturing audio data via a microphone on the client device, processing of captured audio packets by resampling of the packets according to the calculated clock drift value and transmission of the resampled packets to the server device with packet timestamp information.
[0018] The invention thus makes it possible to overcome the intelligibility problems of participants in an audio conference that can arise when several conference participants are in the same room and use only one microphone. The synchronization method as described allows the client devices' microphones to be used as close as possible to each participant. This method distributes the processing load across the client equipment rather than concentrating it on a single device and provides reliable synchronization of the audio signals, since it is based on timestamp information, which is reliable data.
[0019] In a particular embodiment, the process further comprises the following preliminary steps: - activation of a conference application; - Configuring the audio settings of the conference application to disconnect an internal sound card driving at least the microphone of the client device.
[0020] These preliminary steps ensure that the audio data captured by the client device's microphone is not transmitted directly to the remote conference participants. This audio data is synchronized using the process described above before being sent to the server device, which, after processing, will send it to the remote participants either directly or via a conference bridge.
[0021] In a complementary embodiment, when receiving audio data from the server device, the method implements a step of converting the received data by resampling according to the value of the calculated clock drift and a step of transmitting the resampled packets to at least one speaker of the client device.
[0022] Thus, the audio data is reproduced using the speakers of the client devices, in a synchronized manner and as close as possible to the participants using these devices, and is therefore more understandable.
[0023] In a particular embodiment, the exchange of timestamp information is carried out by regular sending steps of a data packet and collection of sending and receiving time data of the data packet, over a time window.
[0024] These exchanges of timestamping information make it possible to measure the network propagation time between a client device and a server device and thus obtain the information allowing the transformation of the time of one device on another, these exchanges being done regularly, they allow to have good accuracy of the synchronization data over time to adapt the synchronization of the audio data to a possible change in network conditions.
[0025] In a particular embodiment, the calculation of the clock drift between the client device and the server device is performed from the time data collected over a sliding time window and according to a linear regression function applied to the timestamp data collected during the returns of the data packet from the server device to the client device.
[0026] This calculation method makes it possible to obtain the clock drift between the equipment accurately, and thus makes it possible to adjust the clock offset values less often.
[0027] In a particular embodiment, a linear regression function is also applied to the timestamp data collected during the round trips of the data packet from the client device to the server device.
[0028] This allows verification that the calculation performed on the return data is consistent with the calculation performed on the forward data. If not, filtering of the problematic measurement or correction of the drift value can be applied.
[0029] In one embodiment, the linear regression function is applied randomly to the timestamp data collected during the return or outward journeys of the data packet.
[0030] This variant makes it possible to overcome differences that might exist between the outgoing and return data, for example network conditions that would not be identical or the sending and receiving conditions would be different.
[0031] In one embodiment, the timestamp information of the resampled audio packets is adapted to the clock of the server device before being sent to the server device.
[0032] These timestamp data then allow the server device to create a mixed stream with the different resampled data received from the different client devices, without having to perform additional synchronization.
[0033] The invention also relates to a method for processing audio data during an audio conference between a plurality of participants via client devices, the method being implemented by a server device connected with at least one client device sharing the same common space and in which the following steps are performed: - exchange of timestamp information between a client device in the common space and the server device; - when receiving audio data packets with associated timestamp information from client devices in the shared space:
[0034] - storing packets in temporary memory; - temporal alignment of stored packets, based on timestamp information; - mixing of aligned packets to obtain a mixed audio stream; - Sending the mixed audio stream to remote participants of the audio conference.
[0035] The method implemented by a server device receiving data synchronized with its own clock allows the different data to be mixed to create a stream to be transmitted to remote participants directly or via a conference bridge. This stream, once played by the various speakers of the remote participants, does not cause the echo phenomena that can occur when the received data is not synchronized.
[0036] According to one embodiment, the process further comprises the following preliminary steps: - activation of a conference application; - Configuring the audio settings of the conference application so that the application's input audio data comes from a processing module producing the mixed audio stream.
[0037] Thus, the audio signals transmitted by the conference application to remote participants or to a conference bridge do not originate directly from the audio data captured by the microphones of the devices in the common area, but come from a processing module implementing the audio data processing method received from client devices in the shared space. The captured audio signals are thus routed to this processing module in order to obtain a synchronized audio stream for transmission.
[0038] The invention relates to a client device sharing the same common space as a second client device during an audio conference between a plurality of participants, the client device comprising at least one microphone controlled by an internal sound card, and at least one client audio module controlled by a microprocessor and capable of implementing an audio data synchronization method, the method comprising the following steps: - exchange of timestamp information between the client device and a server device connected to the client devices in the common space; - calculation of a clock drift of the client device relative to the server device from the information exchanged; - when capturing audio data via at least one microphone of the client device, processing of captured audio packets by resampling of the packets according to the value of the calculated clock drift and transmission of the resampled packets to the server device with packet timestamp information.
[0039] In one embodiment, the device further includes a conference client module capable of activating a conference application, the internal sound card being configured as disconnected from the conference client module for the conference application.
[0040] The client device has the same advantages as the synchronization process it implements.
[0041] The invention also relates to a server device connected with at least one client device sharing the same common space during an audio conference between a plurality of participants, the server device comprising at least one audio server module controlled by a microprocessor and capable of implementing an audio data processing method, the method comprising the following steps: - exchange of timestamp information between a client device in the common space and the server device; - when receiving audio data packets with associated timestamp information from client devices in the shared space: - storing packets in temporary memory; - temporal alignment of stored packets, based on timestamp information; - mixing of aligned packets to obtain a mixed audio stream; - Sending the mixed audio stream to remote participants of the audio conference.
[0042] According to one embodiment, the server device includes a conference client module capable of activating a conference application and at least one virtual sound card configured as the audio data input for the conference client module of the audio conference, the virtual sound card receiving the mixed audio stream from the server audio module.
[0043] The server device has the same advantages as the processing method it implements.
[0044] The invention relates to a terminal comprising a client device and / or a server device as described.
[0045] The invention relates to a computer program comprising instructions for implementing the synchronization or data processing methods as described above, when executed by a processor.
[0046] Finally, the invention relates to a storage medium, readable by a processor, storing a computer program containing instructions for the execution of the synchronization process and / or the data processing process described above. Brief description of the drawings
[0047] Other features and advantages of the invention will become more apparent upon reading the following description of particular embodiments, given by way of simple illustrative and non-limiting examples, and the accompanying drawings, among which:
[0048] [Fig. la] illustrates a conference system comprising a conference bridge and a plurality of client devices, at least two of them being grouped in a common space and implementing a synchronization method according to an embodiment of the invention, one of the client devices also playing the role of a server device implementing a data processing method according to an embodiment of the invention;
[0049] [Fig. 1b] illustrates a conference system comprising a conference bridge and a plurality of client devices, at least two of them being grouped in a common space and implementing a synchronization method according to an embodiment of the invention, a server device also being grouped in the common space and implementing a data processing method according to an embodiment of the invention;
[0050] [Fig.2] illustrates in flowchart form the main steps of a synchronization process implemented in a customer device and a processing process of data implemented in a server device, according to an embodiment of the invention;
[0051] [Fig.3] illustrates the time-stamping information exchange step between a client device and a server device implemented in the synchronization process and the data processing process according to the invention;
[0052] [Fig.4] illustrates a representation of the data recorded in a temporary memory of the server device during a data processing method according to the invention; and
[0053] [Fig.5] illustrates examples of structural realization of a client device or a server device according to an embodiment of the invention. Description of the implementation methods
[0054] Figure 1a describes a conference system according to an embodiment of the invention. This conference system comprises a conference bridge (PC) and a plurality of client devices. The remote client devices DD1 and DD2 are represented as being alone in their respective spaces E1 and E2. The client devices DC1 and DC2 are represented as belonging to the same common space (CS), such as a meeting room. These devices DC1 and DC2 belong to two participants represented as U1 and U2, who are therefore present in this common space.
[0055] Obviously, such a conference can be planned between participants from several common areas, each grouping several participants. In this case, in each common area, the methods according to the invention described below are implemented.
[0056] Each of the client devices (DD1, DD2, DC1, and DC2) includes at least one microphone (Mi) and at least one speaker (HP) controlled by an internal sound card (ISC) which captures sounds from the shared space via the microphone and reproduces sounds from remote users via the speaker. The devices may also include a conference client software module (CC or PCC) controlling a conference application such as, for example, a web-based videoconferencing application.
[0057] The remote client devices DD1 and DD2 operate in a conventional manner, meaning that the internal sound card driving the microphone and speaker is configured as the audio input and output parameter in the conference application controlled by the conference client (CC) software module. This conference client retrieves the sounds captured by the ISC sound card and transmits them to the PC conference bridge. Conversely, the conference client retrieves the audio data from the conference bridge and transmits it to the sound card for playback on the speaker. In a variant In terms of implementation, the conference system does not include a conference bridge; the conference client devices then communicate directly with each other.
[0058] The data exchange arrows shown in Figures 1a and 1b represent, for those illustrated in solid lines, the exchange of audio signals through a communication network, in mixed lines, the exchange of data other than audio through a communication network, for example video and those illustrated in dashed lines, the transfers of acoustic signals through the common space.
[0059] In the common space SR illustrated in [Fig. 1a], two client devices are shown. This common space can obviously include more than two client devices. In [Fig. 1a], a client device DC1 is considered the primary client device because it hosts the conference client responsible for audio communication with the conference bridge or with remote devices for remote participants in the common space. This device also hosts the device, which will be referred to hereafter as the server device, designed to exchange audio data with the conference client.
[0060] This device has a particular function in relation to another client device in this common space. Indeed, this server device includes an audio module, also called an audio server module (AS), which implements a processing method according to an embodiment of the invention and as described later in [Fig.2].
[0061] For this server device, a virtual sound card (VSC) is configured as the audio input and output settings for the conferencing application driven by the main conference client module (PCC). This virtual sound card is a software module that allows audio signals to be routed from one application to another within the same device. In this case, the virtual sound card receives the processed audio data from the server audio module and transmits it to the main conference client module (PCC) for dispatch to the conference bridge (PC).
[0062] According to another embodiment, the server device can make the mixed audio stream available to the conference client via the use of shared memory space, such as, for example, a Linux pipe, or an IP connection. It should be noted that such an embodiment requires the conference client to provide a specific programming interface adapted to the server device.
[0063] The audio server module (AS) receives audio data from its own audio client module (AC) which retrieves the data captured by the microphone via the internal sound card (ISC).
[0064] This server audio module of the server device DC1 also retrieves the synchronized audio data from the various client audio modules of the client devices in the common space, in the illustration of [Fig.1a], the client device DC2.
[0065] The audio data was obtained using a synchronization method described with reference to [Fig. 2] and implemented by the various client audio modules of the client devices in the shared space. This synchronized audio data is in the form of data packets resampled to the clock frequency of the server device, with packet timestamp information, the packet timestamp information being adapted to the clock of the server device.
[0066] Upon receiving these audio data packets, the server audio module of the server device stores these packets in temporary memory along with the associated timestamp information. It then performs a time alignment step on the received packets according to the associated timestamp information to obtain a mixed audio stream, which it sends to the main conference client module via the virtual sound card. The processed data is then transmitted to the conference bridge PC, which forwards the audio data to the other remote client devices DD1 and DD2.
[0067] In this embodiment, this server device is a device such as a communication terminal or a participant's computer. In this case, the server device is also a client device.
[0068] In one embodiment, the DC2 client device illustrated in [Fig. 1a] includes, in addition to its internal sound card (ISC), and a conference client module (CC) as described above, a virtual sound card (VSC) which is configured as audio input and / or output parameters of the conference application driven by the conference client module (CC).
[0069] Indeed, when configuring the audio input and output settings of the client device's conference application, several options are available, as indicated by the illustrated connector. The choice here is between the internal sound card (ISC) and the virtual sound card (VSC). Unlike the server device, this virtual sound card does not receive audio data from an audio module, so the conference client module does not transmit audio data to the conference bridge. Similarly, if the conference client module receives audio data from the conference bridge, it transmits it to the virtual sound card configured as the audio output setting. However, since this virtual sound card is not connected to the client audio module (AC), the data will not be played through the speaker.
[0070] This embodiment allows compatibility with any type of commercially available conference application.
[0071] In another embodiment, the conference application may provide a configuration mode in which the audio outputs and inputs are disconnected.
[0072] In another embodiment, the conference application may provide a configuration mode in which the audio output is connected to the internal sound card to output directly to the speaker, and the audio input is connected to the sound card virtual to be powered by the audio server module, which is itself powered by the client devices.
[0073] In another embodiment, the conference application may provide a configuration mode in which the audio input is connected to the internal sound card to directly acquire the signal from the internal microphone, and the audio output is connected to the virtual sound card to feed the server audio, which will itself feed the client devices.
[0074] In another embodiment, the DC2 client device may not include a conference client module or a virtual sound card. In this case, the user of this client device is not connected to a conference application but participates in the conference via the microphone of their device and an audio client (AC) module.
[0075] The DC2 client device includes an audio client (AC) software module which retrieves the data captured by the microphone via the internal sound card (ISC) and in return sends the data received from the audio server module to the speaker, via the internal sound card.
[0076] This audio client (AC) module implements a synchronization process as described with reference to [Fig. 2]. The resulting data is then transmitted to the audio server (AS) module of the server device (DC1) via a communication network, for example, a local area network. This local area network can be, for example, a wireless network such as Wi-Fi or Bluetooth, or a wired network such as Ethernet.
[0077] This audio server (AS) module, as described above, implements a data processing method as described with reference to [Fig.2].
[0078] When the server device receives an audio stream from the conference PC bridge via its main conference client module (PCC), it transmits this stream to the client devices in its shared space via its virtual sound card and its server audio module. The client audio modules of the client devices implement the synchronization process before sending the resulting audio data to their respective internal sound cards so that this data can be played on their respective speakers. Thus, the data is played synchronously on the client devices in the shared space, as close as possible to the device user.
[0079] Thus, the microphones and speakers of the client devices in the shared space are used to capture and reproduce sound as close as possible to the device's user. The captured sound is therefore more intelligible than if it had been captured by a device further away from the user. However, this captured sound is not transmitted directly by the client device but is transmitted, after being time-stamped according to the server clock by the client device, to the server device, which is responsible for temporally aligning and mixing the audio data it receives, thus synchronizing them.
[0080] The processing load on the server device is reduced since it does not have to perform specific synchronization processing, such as correlation analysis between audio signals. Similarly, in the embodiment described here, the audio server device also does not have to perform the resampling processing that ensures fine synchronization over time, as this processing is handled by the client device. Remote participants, using devices DD1 and DD2, will therefore hear participants U1 and U2 more distinctly while avoiding any echo effect.
[0081] It should be noted that all the audio streams mentioned can be mono, but also stereo or multichannel.
[0082] Figure 1b represents an alternative embodiment in which a device (DD), independent of the participants' client devices, is dedicated to implementing the data processing method as described with reference to Figure 2. This data processing method is implemented by an audio server module (AS) of the dedicated processing device (DD). This audio server module, as described above, retrieves the synchronized audio data from the various audio client modules of the client devices in the shared space. In one embodiment, this device can be, for example, a terminal, a computer, or a network cabling system located in the meeting room.
[0083] This dedicated device may also include a sound card (ISC) driving at least one microphone and one speaker and an audio client (AC) module which retrieves the data captured by the microphone via its internal sound card (ISC) and in return, sends the received data to the speaker via its internal sound card.
[0084] This dedicated device acts as a server device, in the sense that it receives synchronized audio data from other client devices in the common space, in the form of data packets resampled at the clock frequency of this server device, with timestamp information for the data packets, the timestamp information being adapted to the clock of the server device.
[0085] Upon receiving these audio data packets, the audio server (AS) module of this dedicated device stores these packets in temporary memory along with the associated timestamp information. It then performs a time alignment step on the received packets according to the associated timestamp information to obtain a mixed audio stream, which it sends to the main conference client module located on a main client device, represented here as DC1. This transmission occurs via a communication bus, such as an external sound card (ESC) of this client device. This is accomplished, for example, by a wired connection between the two devices, such as a USB connection.
[0086] For this main client device, the configuration of the input and output audio data is done by choosing this external sound card instead of choosing its internal sound card (ISC).
[0087] In one embodiment, the server device (DD) can serve as a wireless network access point (e.g. WiFi) so that client devices in the common area can communicate with each other.
[0088] [Fig.2] illustrates the main steps of a synchronization process implemented in a client device (DC) as represented in DC2 in [Fig. aa] and of a data processing process implemented in a server device (DS) as represented in DC1 in [Fig. aa], according to an embodiment of the invention.
[0089] Optionally, client devices can activate, in a prior step E21, a conference application, to participate in a conference between several participants, such as a video conference or an audio conference.
[0090] The server device (DS) implements a preliminary step of activating a conference application in E27.
[0091] To do this, in the usual way, the software application installed on the device is opened. The audio inputs and outputs to be used for this conference are then configured in E22, if the conference application has been activated for the client device, and in E28 for the server device. This configuration is performed on the conference application's interface, for example, by the audio conference user.
[0092] For the client device, configuring the audio settings of the conference application involves disconnecting an internal sound card that drives at least the client device's microphone. In one embodiment, this disconnection is achieved by configuring a virtual sound card in place of the internal sound card; this virtual sound card does not receive any audio data for the client device.
[0093] In another embodiment where the conference application offers a "none" mode as input / output configuration, i.e. without input and output audio data, then this mode can be selected.
[0094] It is also possible to configure the audio input and output of the application differently and to apply, for example, only the configuration described above to the input data.
[0095] For the server device, the configuration of the audio parameters of the conference application is carried out so that the input audio data of the application comes from a processing module producing the mixed audio stream.
[0096] In one embodiment, this configuration can be achieved by connecting a virtual sound card to the conference module that controls the conference application, instead of the server device's internal sound card. This virtual sound card receives, for this server device, the audio data processed in E30 detailed below, before sending to the cong bridge or to remote devices directly, in E34 via the conference client module.
[0097] Similarly, it is possible to configure the application's audio input and output differently and, for example, to apply the previously described configuration only to the input data. In one embodiment, the server device can make the mixed audio stream available to the conference client via shared memory. In this case, the conference client provides a specific programming interface, adapted to the server device, to enable such a configuration of the input and output data.
[0098] Steps E23 and E29 represent the exchange of timestamp information between a client device and a server device in the common space.
[0099] This exchange of information allows the client device to know both the clock offset and a clock drift that may exist between the client device and the server device.
[0100] Figure 3 illustrates this step of exchanging timestamp information between the client devices and the server device of a common space. These exchanges can take place, for example, according to the NTP protocol (for "Network Time Protocol" in English) as described, for example, in the technical specification RFC-5905 "Network Time Protocol Version 4" of June 2010.
[0101] Other network synchronization protocols can of course be used here such as, for example, the PTP protocol (for "Precision Time Protocol").
[0102] Initially at time Tcs, the client device emits a message M, in the form of a data packet, which can be called synchronization data, to query the server device on the current time, it sends this message in which it fills a first field with the current time Tcs indicated by its local clock;
[0103] At Tmr, when the server device receives the message M, it immediately completes a second field of the message M with the current time TMR indicated by its local clock;
[0104] At Tms, when the server device sends its reply message, it completes a third field with the current TMS time indicated by its local clock;
[0105] At Tcr, when the client device receives the reply message, it immediately notes the Tcr time of receipt indicated by its local clock.
[0106] The client device can then calculate at step E24 illustrated in [Fig.2], the round-trip delays "delta" of these two messages which allows the evaluation of the network latency as well as the difference "theta" between its local clock and that of the server device.
[0107] Here, it is assumed that the forward delay and the return delay (delta) are the same, that is to say that the network conditions remain unchanged during this message exchange.
[0108] The calculation performed also takes into account a possible drift "alpha" of the clock of the client device relative to that of the server device.
[0109] Indeed, two independently operating devices each have their own clock. A clock is defined as a monotonic function equal to a time that increases at the rate determined by the clock frequency. It can originate at the instant the device is switched on, or at an arbitrarily defined instant such as, for example, January 1, 1970 for NTP.
[0110] The clocks of two devices are necessarily different and three parameters are defined:
[0111] - the clock offset (theta): difference in time at the origin between two clocks;
[0112] - clock drift (alpha): frequency ratio between two clocks;
[0113] - clock deviation: variation of the drift over time, or derivative second of the clock in relation to time.
[0114] A classical model of a clock neglects clock drift, mainly caused by the construction of the component called quartz, but also by temperature changes. Thus, in a server / client network context, as here between a server device and a client device, the clock of the client device Te is expressed as a function of the clock of the server device TM according to an equation such as Tc = f(TM) = alpha * TM + theta where alpha represents the clock drift of the client with respect to that of the server device, and theta represents the offset of the client's clock.
[0115] Thus, in the present case, at step E24 illustrated in [Fig.2], the following equations are used to estimate the delta delay:
[0116] Tcs + delta = alpha* TMR + theta
[0117] Tcr - delta = alpha* TMS + theta
[0118] Which gives the value of delta according to the equation:
[0119] delta = ((alpha* TMR - Tcs) + (TCR - alpha* TMS)) / 2
[0120] This value representing the offset between the clocks of the server device and that of the client device is a time that can be expressed in seconds or milliseconds.
[0121] This exchange of timestamp information (E23, E29) is carried out by regularly sending a data packet (message M) and collecting timestamp data for the sending and receiving of this packet. In one embodiment, a data packet is sent every second. A different sending frequency can, of course, be provided.
[0122] The information is collected over a sliding time window which can be, for example, one hour. In order to obtain the value of the clock drift "alpha", a linear regression (for example a least squares function) is performed taking into account the TMS, TCR timestamp values collected in this time window, i.e. the timestamp values collected during the return of the data packet from the server device to the client device.
[0123] This gives us a value of alpha representing the clock drift, that is to say the frequency ratio between the clock of the server device and the clock of the client device.
[0124] When a client audio application implementing the synchronization process starts, the time window is empty or has not reached the predefined size. In this case, the drift value is set to 1 and a minimum delay is observed before calculating a new drift value using the data collected during this minimum time interval. This interval could be, for example, 7 seconds.
[0125] Gradually, the size of the data history window reaches a predetermined value, for example 30 minutes or 1 hour. Only the data within this sliding window is then retained; previous data is erased from the device's memory. Indeed, a change in the network configuration could, for example, lead to a change in the delay or drift values.
[0126] In a particular embodiment, a linear regression function is also applied to the data going from the data packet of the client device to the server device. Thus, if a large difference in the alpha value is observed with this new calculation, the problematic measurement can either be filtered out or a correction can be applied to the alpha value. In the latter case, an average between the two values could, for example, be calculated.
[0127] In one embodiment, the linear regression function is applied randomly to either the return or forward data of the data packet. This embodiment eliminates potential differences between the forward and return data, such as network conditions that are not identical or different sending and receiving conditions.
[0128] The value of the drift obtained is a dimensionless value equal to the ratio of the clock frequencies of the server device and the client device.
[0129] This value is then used to adapt, in E25, the data captured by the microphone via the ISC sound card of the client device before transferring them in E26 to the server device or to adapt the data received from the server device before broadcasting to the speaker of the client device via its ISC sound card.
[0130] For this purpose, in E25, the data captured by the microphone of the client device are processed by resampling according to the value of the calculated clock drift.
[0131] Knowing the value of the clock frequency of the FM server device, for each audio packet, a resampling module is implemented to adapt the sampling frequency (Fc) of the client device to that of the server device in the following way: Fc = FM * alpha.
[0132] The instant Tccapture of the capture of an audio data packet by the client device is converted into timestamp data adapted to the clock of the server device using the inverse function of f defined above, in the following way: TMcapture=f'(Te capture), that is here TMcapture =( Tccapture - theta) / alpha.
[0133] This results in data packets resampled at the clock frequency of the server device Fc with capture times adapted to the clock of the server device.
[0134] These resampled data packets are transmitted to the server device with timestamp information for this packet according to the clock of the server device.
[0135] Similarly, when receiving audio data by the server device which itself received it from the conference bridge, the data packets are resampled in E25 according to the value of the calculated clock drift and transmitted to the speaker of the client device via its ISC sound card, so that the audio data can be played.
[0136] Thus, the distribution of audio streams to the various speakers of client devices in the same meeting room does not generate echo phenomena since there is no longer any delay or drift in the time of the sounds played. Furthermore, the playback is carried out as close as possible to the users of these client devices for better understanding of the audio stream.
[0137] If the clock drift information was not shared, it would inevitably result, at the time when the clock offset should be updated due to the time drift, in either a time overlap or a silence, between 2 successive audio packets from the same client at the time of temporally aligning the received packets.
[0138] If resampling was not used to give the audio packets the appropriate duration in the server clock, the result would inevitably be either too much data accumulation in the server if the server clock frequency is lower than that of the client, or a starvation of data in the server if the server clock frequency is higher than that of the client, resulting in both cases in audible audio artifacts ("clicks").
[0139] On the server device side, an audio data processing method is put into artwork.
[0140] In the same way as described for the client device, an exchange of timestamp information between a client device in the common space and the server device is carried out by steps E23 and E29 illustrated in [Fig.2].
[0141] As described previously, the client device transmits in E26 data packets resampled at the clock frequency of the server device, with a timestamp information, the timestamp information being adapted to the clock of the server device.
[0142] When receiving these audio data packets in E30, the server device stores these packets in E31 in a temporary FIFO (First In First Out) type memory with the associated timestamp information.
[0143] Then in E32, a time alignment step of the received packets is performed according to the associated timestamp information. A representation of these data packets aligned in the FIFO-type temporary memory is illustrated in [Fig.4].
[0144] This [Fig. 4] shows that the data packets (illustrated by consecutive AP blocks) received from the different client devices (DCi, DC2, ..., DCk) are stored in memory with the associated timestamp information. Thus, the data packets received from DCi were stored in a FIFO memory dedicated to DCi with the associated timestamp information Ti>M1+n,..., TijM1+3,TijM1+2, TijM1+1 to Tim1.
[0145] The same applies to data packets received from DC2 to DCk client devices with corresponding timestamp information.
[0146] A time alignment according to step E32 of [Fig.2] allows the data blocks to be shifted so that at a time T, the timestamp information of the blocks recorded in the different FIFOs correspond.
[0147] This offset is illustrated in [Fig.4]. The time adjustment is performed by aligning all the FIFO data so that a vertical axis corresponds to the same instant T.
[0148] At step E33 of [Fig. 2], the server device mixes the various data packets thus aligned at a given time T to create a mixed audio stream (Stream-T) to be transmitted to the conference bridge at E34. The mixing step can be a simple summation of all the aligned signals, or a summation with an arbitrary weighting or one resulting from an analysis. An analysis of the volume variations between different speakers, for example, allows the definition of a compensation gain to be applied as a weighting to the low-volume signals.
[0149] In the case where there are no data packets at this time T for one of the client devices, then a silence data can be inserted.
[0150] When receiving audio streams from the conference bridge (PC) in E30, the server device sends the audio packets to the client devices in space common in E26 so that they process them as explained previously in step E25.
[0151] Figure [Fig. 5] illustrates a structural realization of a client device or a server device (100) according to an embodiment of the invention.
[0152] Such devices can be included in a communication terminal such as a smartphone, electronic tablet or a personal computer, or even a dedicated device such as an audio conferencing octopus.
[0153] The client device includes a processing circuit typically comprising: - a MEM memory for storing instruction data of a computer program as defined in the invention as well as various applications, an audio or video conferencing application being, for example, one of these applications; - an INT interface capable of receiving user instructions such as, for example, audio parameter configuration information and of displaying various data, such as, for example, video images of a video conference; - an ISC sound card capable of retrieving audio data captured by a microphone Mi and sending audio data to a speaker HP for playback;
[0154] - an AC client audio software module capable of retrieving data captured by the microphone via the internal ISC sound card. This audio client module can also send the received data to the speaker via the internal sound card. This AC audio client module implements a synchronization process as described with reference to [Fig.2];
[0155] - a COM communication interface capable of sending or receiving data, including audio data, to or from another device, such as a server device via an R network;
[0156] - a PROC processor for executing computer program instructions which the MEM memory stores, in particular the instructions for the implementation of the synchronization process as described with reference to [Fig.2]; the processor being able to control the different modules described.
[0157] The client device may optionally include a CC conference client software module capable of controlling a conference application such as, for example, a web-based video conferencing application.
[0158] In this case, it may include a virtual sound card (VSC) configured as the audio input and / or output module for the conference application driven by the conference client module (CC). This virtual sound card is a software module that allows audio signals to be routed from one application to another within the device. In the case of a client device, the virtual sound card does not receive any audio data. It can be used to disconnect the ISC sound card from the conference client module.
[0159] Similarly, the server device, in one embodiment, comprises a processing circuit typically including: - a MEM memory for storing instruction data of a computer program as defined in the invention as well as various applications, an audio or video conferencing application being, for example, part of these applications; - an INT interface capable of receiving user instructions such as audio parameter configuration information and displaying various data, for example video images from a video conference; - an ISC sound card capable of retrieving audio data captured by a Mi microphone and sending it to an AC client audio module and conversely of receiving audio data from the client audio module and sending it to an HP speaker for playback.
[0160] - a CC conference client software module, which is also called in the case of the server device, main conference client, capable of driving a conference application such as, for example, a web-based video conferencing application and sending audio streams to other remote devices or to a conference bridge via the COM communication interface.
[0161] - an AC client audio software module capable of retrieving data captured by the microphone via the internal ISC sound card and send the received data to the speaker via the internal sound card.
[0162] - an AS server audio module that implements a processing method according to a embodiment of the invention and as described with reference to [Fig.2].
[0163] The AS server audio module receives audio data from its own AC client audio module and also retrieves synchronized audio data from the various client audio modules of client devices in a common space, via the COM communication interface to implement the processing method.
[0164] - a VSC virtual sound card configured as the audio input and output module of the conferencing application driven by the CC conference client module. This virtual sound card is a software module that allows audio signals to be routed from one application to another within the system. In the case of a server system, the virtual sound card receives processed audio data from the server audio module and transmits it to the CC conference client module; the virtual sound card can also send audio data from the CC conference client module to the server audio module.
[0165] - a COM communication interface capable of sending or receiving data, including audio data, to or from other devices, such as client devices or a conference bridge via an R network; - a PROC processor to execute the computer program instructions stored in the MEM memory, including instructions for implementing the data processing method as described with reference to [Fig.2]; the processor being capable of controlling the various modules described above.
[0166] In another embodiment, the server device may not include a CC conference client software module or a VSC virtual card. This server device may, however, have a USB connection port for communicating audio data to another client device. This type of server device is, for example, a device dedicated to audio conferencing, such as an audio conferencing hub.
[0167] Of course, this [Fig. 5] illustrates an example of a structural embodiment of a device (client or server) within the meaning of the invention. Figures 1 to 4, discussed above, describe in detail functional embodiments of these devices.
Claims
Demands
1. A method for synchronizing audio data during an audio conference between a plurality of participants via client devices, the method being implemented by a first client device of the conference sharing the same common space as a second client device and in which the following steps are executed: - exchange of timestamp information (E23, E29) between the client device and a server device connected with the client devices of the common space; - calculation (E24) of a clock drift of the client device with respect to the server device from the information exchanged; - when capturing audio data via a microphone of the client device, processing (E25) of the captured audio packets by resampling the packets according to the value of the calculated clock drift and then transmission (E26) of the resampled packets to the server device with timestamp information of the packets.
2. A method according to claim 1, further comprising the following preliminary steps: - activation of a conference application (E21); - configuration (E22) of the audio parameters of the conference application to disconnect an internal sound card driving at least the microphone of the client device.
3. A method according to any one of claims 1 or 2, wherein upon receiving audio data from the server device, the method implements a step of converting the received data by resampling according to the value of the calculated clock drift and a step of transmitting the resampled packets to at least one speaker of the client device.
4. A method according to any one of claims 1 or 2, wherein the exchange of timestamp information is carried out by regular sending steps of a data packet and collection of sending and receiving time data of the data packet, over a time window.
5. A method according to claim 4, wherein the calculation of the clock drift between the client device and the server device is performed from time data collected over a sliding time window and according to an applied linear regression function to timestamp data collected during the return of the data packet from the server device to the client device.
6. A method according to claim 5, wherein a linear regression function is also applied to the timestamp data collected during the round trips of the data packet from the client device to the server device.
7. A method according to claim 5, wherein the linear regression function is applied randomly either to the timestamp data collected during the return or the outbound movements of the data packet.
8. A method according to any one of the preceding claims, wherein the timestamp information of the resampled audio packets is adapted to the clock of the server device before being sent to the server device.
9. A method for processing audio data during an audio conference between a plurality of participants via client devices, the method being implemented by a server device connected with at least one client device sharing the same common space and in which the following steps are performed: - exchange (E23, E29) of timestamp information between a client device in the common space and the server device; - upon receipt (E30) of audio data packets sampled at the clock frequency of the server device, with associated timestamp information, from the client devices in the common space: - storage (E31) of the packets in temporary memory; - temporal alignment (E32) of the stored packets, according to the timestamp information; - mixing (E33) of the aligned packets to obtain a mixed audio stream;- sending (E34) the mixed audio stream to remote participants of the audio conference.;
10. A method according to claim 9, further comprising the following preliminary steps: - activation of a conference application (E27); - configuration (E28) of the audio parameters of the conference application so that the audio input data of the application comes from a processing module producing the mixed audio stream.
11. A client device sharing the same common space as a second client device during an audio conference between a plurality of participants, the client device comprising at least one microphone (Mi) driven by an internal sound card, and at least one client audio module (AC) driven by a microprocessor and capable of implementing an audio data synchronization process, the process comprising the following steps: - exchange of timestamp information between the client device and a server device connected with the client devices in the common space; - calculation of a clock drift of the client device relative to the server device from the information exchanged;- when capturing audio data via at least one microphone of the client device, processing of the captured audio packets by resampling the packets according to the calculated clock drift value and then transmission of the resampled packets to the server device with packet timestamp information.
12. Client device according to claim 11, further comprising a conference client (CC) module capable of activating a conference application, the internal sound card being configured as being disconnected from the conference client module for the conference application.
13. A server device connected with at least one client device sharing the same common space during an audio conference between a plurality of participants, the server device comprising at least one microprocessor-driven audio server (AS) module capable of implementing an audio data processing method, the method comprising the following steps: - exchange of timestamp information between a client device in the common space and the server device; - upon receiving audio data packets sampled at the clock frequency of the server device, with associated timestamp information, from the client devices in the common space: - storage of the packets in temporary memory; - temporal alignment of the stored packets, according to the timestamp information; - mixing of the aligned packets to obtain a mixed audio stream; - Sending the mixed audio stream to remote participants of the audio conference.
14. Server device according to claim 13, comprising a conference client module (CC) capable of activating a conference application and at least one virtual sound card configured as the audio input for the conference client module of the audio conference, the virtual sound card receiving the mixed audio stream from the audio server module.
15. Communication terminal comprising a client device according to one of claims 11 to 12 and / or a server device according to one of claims 13 to 14.
16. A processor-readable storage medium storing a computer program containing instructions for executing the synchronization method according to any one of claims 1 to 8 and / or the data processing method according to any one of claims 9 to 10.