Method and receiver for jitter compensation during reception of audio content over an IP-based network, and method and device for transmitting and receiving audio content with jitter compensation - Patents.com

Jitter compensation in VoIP systems synchronizes receiver clocks with sender times and dynamically adjusts buffers to address jitter-related challenges, ensuring consistent audio processing and reducing interruptions across varying transmission channels.

JP7721687B2Active Publication Date: 2025-08-12DFS GERMAN AIR TRAFFIC CONTROL GMBH
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
JP2023574861
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2021-07-09
Filing Date
2022-07-08
Publication Date
2025-08-12
Estimated Expiration
2042-07-08

AI Technical Summary

Technical Problem

Existing voice over IP (VoIP) systems face challenges in accurately predicting the total time duration for audio packet transmission due to random fluctuations in jitter, leading to potential voice packet loss and interruptions in continuous audio processing, especially in applications requiring real-time communication and multiple parallel transmission channels.

Method used

Implementing jitter compensation by synchronizing the receiver clock with the sender time using time information from voice packets, adjusting the receiver time to minimize relative packet transmission duration, and dynamically determining the buffer length based on actual jitter to ensure continuous audio processing.

Benefits of technology

Ensures consistent time duration from audio packet transmission to processing, minimizing interruptions and voice packet loss, even with varying transmission channels, thereby maintaining continuous audio output and avoiding echoes in applications like operational ground-to-air communications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007721687000001
    Figure 0007721687000001
  • Figure 0007721687000002
    Figure 0007721687000002
  • Figure 0007721687000003
    Figure 0007721687000003
Patent Text Reader

Abstract

A jitter compensation during reception over an IP-based network of audio content (1, 2, 3, 4) in an audio packet (110, 120; 210) having a header (H) and a payload (PL) is described, in which one time information of a transmitter time of a transmitter is included in the audio packet (110, 120; 210), indicating a transmission time (ts) of said audio packet (110, 120; 210). The following is provided: a receiver initializes a receiver clock by using the transmitter time according to the time information in an initial packet, the receiver determines a minimum relative packet transmission duration (delta) after initialization of the receiver clock, adjusts the receiver time according to the minimum relative packet transmission duration (delta), during reception of a first audio packet (111) having audio content, determines an actual reception time of this audio packet, and determines a buffer (DJB) according to the actual reception time.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to a method for jitter compensation of random time variations (called jitter) between receiver times of audio packets when receiving audio content over an IP-based network, i.e. in particular for use in methods for transmitting audio content (digitized as audio data in the case of analog audio content) in digital packets (such as in telephone or wireless connections), also called Voice over IP or VoIP); and to a receiver set up to carry out the method. Additionally, the present invention relates to methods and devices for transmitting and receiving audio content, in which the jitter compensation described is used. [Background technology]

[0002] In this method, as is usual in the case of VoIP, a series of digital voice packets, with or without voice content, is received by a receiver for processing the voice content contained in the voice packets (in particular a receiver for transmission for output, over a wireless link or other use of the voice data). The digital voice packets with voice content have a section called a header, which contains data for communication control (e.g. corresponding to a communication protocol or standard), and a section called a payload. The payload of a voice packet consists of a part of the (total) voice content in the form of digital voice data. The transmission of the (total) voice content is therefore divided into a series of voice packets with continuous digital voice data (payload PL): i.e., the digital voice data contained (directly or indirectly) in the continuous voice packets, when combined with each other, represent the (total) voice content.

[0003] An audio packet with no audio content has a header with no payload and is therefore usually correspondingly shorter. However, it is also possible that an audio packet with no audio data also has a payload with zero data and is exactly the same length as an audio packet with data content (audio data). Zero data is data with arbitrary or random content: i.e., uncorrelated with any useful audio content.

[0004] Since the payload in this case does not contain any information, such voice packets are also called payload-less voice packets in the context of this application. However, in principle, it would be advantageous if voice packets without voice content were shorter and the payload section could simply be omitted instead of being filled with zero data. However, depending on the communication standard, it may also be required that all voice packets, regardless of whether they contain voice data, must be the same length. Voice packets without voice content are used, for example, to maintain existing communication connections between a transmitter and a receiver in the long term so that they can be monitored. One example of such a use of voice over IP could be the transmission of operational ground-to-air communications (in some sections via a radio connection) in air traffic, which represents a particularly preferred application of the present invention.

[0005] Each voice packet (especially in the voice packet header) includes at least one time information of the transmitter time (indicating the time of transmission of the voice packet), which applies to voice packets with and without voice content.

[0006] After receiving the first voice packet with audio content, the receiver waits to process the audio content for a waiting time called a buffer (jitter buffer), meaning the buffering time. This allows the receiver to receive subsequent voice packets with audio content during the buffering (i.e., waiting time) to compensate for jitter between the reception of subsequent voice packets and to achieve continuous processing of the audio content without interrupting the audio output due to the delay in reception caused by jitter. This processing can be a direct reproduction of the audio content as audible speech (i.e., analog output). However, the receiver can also broadcast the received audio content via a wireless transmitter. In this case, the reproduction of the audio content occurs only in the wireless receiver. In this case, the receiver only indirectly reproduces the digital audio packets (in the sense that the received audio packets are processed and the contained audio content is sent as a radio message by the wireless transmitter). In this case, for example, the generated audio content (radio message) of air traffic control (as a transmitter) is transmitted as a voice-over-IP communication via an IP-based network to a ground station's radio transmitter (as a receiver). The radio transmitter then sends out a radio message (for communication from air traffic control, e.g., to an aircraft) which is output by the radio receiver. For communication in the reverse direction (e.g., from an aircraft to air traffic control), the radio transmitter is in the aircraft and the radio receiver is in the ground station. The radio receiver is then the transmitter of the voice-over-IP communication. The terms "transmitter" and "receiver" in this document refer to the transmission of audio content over an IP-based network.

[0007] The physical transmission process of voice packets is briefly described below to provide background on voice over IP communications. A more detailed description will occur in the context of the description of exemplary embodiments.

[0008] The total delay (packet transmission duration) of a voice packet begins with the time of its sending by the transmitter and is derived from the sum of the delay (also called simply delay) of the voice packet through the transmission channel (to which the length of the voice packet itself belongs) and jitter (meaning random fluctuations in receiver time). Furthermore, latencies such as buffers (jitter buffers) arise before the voice content contained in the voice packet is processed (e.g., output processing or other processing).

[0009] Therefore, the time from sending out an audio packet to processing the audio content is given by the actual packet transmission duration (delay and jitter) plus the buffer (and, if appropriate, any additional latency after receiving the audio packet and before outputting the audio content).

[0010] The longer the waiting time, the more subsequent voice packets can be received and temporarily stored (buffered), so that after processing the voice content of the preceding voice packets, the voice content of the buffered voice packets can be processed additionally early on, so that even if the jitter of the subsequent voice packets is particularly long, the voice content of the subsequent voice packets is already available for further processing, and continuous processing (e.g., voice playback) is guaranteed. A long buffer in terms of time leads to correspondingly long delays, which are undesirable for Voice-over-IP applications with high requirements for real-time communication. Therefore, a short buffer in terms of time leads to a higher communication speed, but also to a higher risk of voice packet loss, which may hinder continuous processing of the voice content.

[0011] Another drawback of this known solution with buffers is that the total time consisting of the packet transmission duration and, if necessary, additional waiting time (such as for buffering) still depends on the random duration of jitter during the transmission of the voice packets and therefore cannot be accurately predicted. This can lead to drawbacks in the case of some communication applications, such as wireless applications with multiple parallel transmission channels (e.g., sending wireless messages in parallel by several wireless transmitters at various ground stations), where jitter of different lengths on the various transmission channels leads to echoes. Summary of the Invention [Means for solving the problem]

[0012] The object of the present invention is therefore to propose jitter compensation and an audio transmission using such jitter compensation, which allows for the definition of the total time duration in the case of transmission of an audio packet (at least within a predetermined period) from the output of the audio packet by the transmitter to the processing of the audio information by the transmitter, while ensuring continuous processing (e.g. radio transmission or output) of the audio content.

[0013] This object is achieved according to the invention by the features of claims 1, 9, 11 and 15.

[0014] In the context of the jitter compensation proposed according to the invention, the following is provided in particular: during a connection period in a transmission channel (a connection period set up to a voice packet sender or a connection period existing between a receiver and a sender), after receiving a first voice packet (with or without voice content or voice data), called an initial packet, the receiver initializes its receiver clock with the sender time according to the time information in the initial packet so that the receiver time correlates with the sender time when sending the voice packet. As a result, the receiver clock displays a time that depends on the time of sending the first voice packet from the connection period, the delay of the voice packet through the transmission channel, and random jitter in time. According to an embodiment of the present invention, it is easy to synchronize the receiver time with the time information from the voice packet. Here, the time information is corrected, if applicable, by the sample time of the packet's payload so that the receiver time corresponds to the sender time of a voice packet without data content and the random jitter of the currently transmitted voice packet. Therefore, there is a time offset (relative to absolute time, which is constant throughout the entire system) between transmitter time and receiver time, and this time offset corresponds exactly to the packet transmission time of the first voice packet (which has no voice content) within a connection period.

[0015] A connection period (also called a session) begins with the setup of a transmission channel between a sender and a receiver and ends with the interruption of the transmission channel (i.e. corresponds to the actual connection duration). However, it is also possible to choose a shorter connection period within the actual connection duration, for example by defining a fixed period and / or depending on events at the receiver and / or at the sender. A connection period therefore corresponds to a defined or definable period within the physical connection duration within the transmission channel (i.e. a portion of the total connection duration (but at most the total connection duration)).

[0016] For subsequent voice packets received at the receiver with exactly the same jitter (assuming constant delay), the receiver time will be correlated with the transmitter time of the transmitter sending the packet. Thus, in the type of correlation described above, the transmitter time and receiver time are identical if the jitter is the same. Deviations in the correlation are caused by various jitters.

[0017] This allows the receiver to determine the minimum relative packet transmission duration after the receiver clock has been initialized in that: during reception of subsequent voice packets, in either case: The receiver time is compared with the time information contained in the header of the subsequent voice packet at the transmitter time when the voice packet was sent, and the relative packet transmission duration is determined relative to the receiver time. This can be done, for example, by the difference between the transmitter time and the receiver time with respect to the previous difference. If the receiver time and the transmitter time are synchronized as described in the previous embodiment, a difference "receiver time - transmitter time" of zero means that the subsequently transmitted voice packet was exactly as fast as the voice packet (during initialization or subsequent adjustment) from which the receiver time was determined (as further explained below). If the difference is less than zero, the receiver time is earlier than the transmitter time; i.e., the voice packet was faster (relatively shorter packet transmission duration) than the voice packet from which the receiver time was determined. In contrast, if the difference is greater than zero, the receiver time is later than the transmitter time; i.e., the voice packet was slower (relatively longer packet transmission duration) than the voice packet from which the receiver time was determined. The resulting minimum relative packet transmission duration or interval within the connection period is temporarily stored. Next, within the connection period (or the interval within the connection period used as the basis for the comparison), the relative packet transmission duration is newly stored as the minimum relative packet transmission duration if the currently determined relative packet transmission duration is shorter than the previous minimum (and temporarily stored) relative packet transmission duration. The receiver time, which depends on the minimum relative packet transmission duration temporarily stored at a defined evaluation time (e.g., at the end of an interval within the connection period), is adjusted in such a way that the receiver time is correlated with (in particular synchronized with) the transmitter time when sending out an audio packet having the minimum relative packet transmission duration. In one embodiment, this can occur in the sense that the value of the minimum relative packet transmission duration stored at the evaluation time is subtracted from the current receiver time to determine the new receiver time for the remainder of the method. As a result, the time difference between the receiver clock and the transmitter clock is also automatically and continuously corrected (e.g., thanks to the clock difference of the timers used).

[0018] The actual embodiments defined in the Summary of the Invention above will be used for illustration purposes, and the present invention is not necessarily limited to these embodiments.

[0019] Whenever a voice packet with voice content (or a sequence of voice packets with voice content) is received after at least one received voice packet without voice content, the receiver assumes that new voice content (e.g., a new radio message within a connection) is to be transmitted. The start and end of new voice content may also (additionally or alternatively) be identified by encoding in the header.

[0020] Thus, when receiving a first voice packet having voice content after at least one voice packet having no voice content, the receiver determines the actual reception time of this voice packet by comparing the time information in the received voice packet with the receiver time at the time of reception, such that the actual reception time is indicated, in particular, relative to the receiver time. Thus, this actual reception time indicates the actual jitter when the first voice packet having voice content is received. The receiver then determines the buffer (i.e., buffer time or waiting time) until the voice content is processed according to / depending on the actual reception time (i.e., according to the actual jitter when the first voice packet having voice content is received). In principle, the buffer can be formed by subtracting the actual jitter from a predetermined maximum buffer that would be used when a voice packet with no or minimum jitter is received as the first voice packet having data content. If the actual jitter corresponds to minimum or no jitter, the predetermined maximum buffer is used as the buffer.

[0021] The preferred embodiment of the present invention for determining the buffer consists in forming the difference between the time information in the voice packet and the actual reception time. This difference then indicates the jitter of the current voice packet, which is the first voice packet with voice content to be determined in accordance with the present invention. If this difference (time information in the voice packet - actual reception time) is less than zero, the (positive) value of the difference indicates the actual jitter. This actual jitter is then subtracted from the predetermined maximum buffer. Otherwise, if the difference (time information in the voice packet - actual reception time) is greater than or equal to zero, the maximum buffer is used as the buffer. This simple embodiment applies if the receiver time is equal to the time information (send time) in a voice packet with no voice content and the shortest packet transmission time within the connection period. Therefore, for a voice packet with voice content and the shortest packet transmission time within the connection period, the duration of the payload is used as a correction to ensure the same standardization. If the sender time and the receiver time are correlated according to a different dependency (i.e., they are not synchronized with the voice packet with the shortest packet transmission time), this correlation dependency is used accordingly.

[0022] Therefore, according to the present invention, a dynamic buffer (jitter buffer) is determined by jitter compensation, the time length of which is determined according to the length of the actual jitter of a voice packet. This is determined after the reception of a voice packet. This is the first voice packet with voice content after at least one voice packet without voice content. The buffer is preferably determined according to the present invention in such a way that the maximum buffer is used for a voice packet (i.e., a voice packet with the earliest possible reception time) having a packet transmission time with no actual jitter or the smallest actual jitter, and the actual jitter is subtracted from this maximum buffer. In other words, the buffer is determined according to the present invention so that the sum of the actual jitter (of the voice packet used to determine the dynamic buffer) and the dynamic buffer is constant. The length is preferably selected so that this constant value of the sum ensures continuous processing (e.g., audio playback and / or radio transmission) even when consecutive voice packets with maximally different jitter are received. Therefore, this constant value of the sum of the actual jitter and the dynamic buffer should be at least greater than the maximum difference that occurs when receiving a voice packet. Smaller transient interferences can also be caught with an additional safety margin (eg, twice the expected jitter).

[0023] As a result, not only is the shortest possible time between sending out an audio packet with audio content and processing the audio content (e.g., consecutive output of a series of audio packets with adjacent audio content) achieved, but the time between sending out an audio packet with audio content and processing it is also always the same length, even if various transmission channels are used in parallel or absolute transmission times have to be known for other reasons.

[0024] A further preferred embodiment of the present invention may provide the following: during initialization of the receiver clock and during determination of the relative packet transmission duration, the difference in length between voice packets with and without voice content is taken into account (in particular in that the reception time is normalized relative to the transmission of voice packets without voice content), whereby the additional and known length of the payload of the data telegram can be subtracted, for example, when processing the reception time. As a result, the effective transmission duration (packet transmission time) can be calculated in both cases relative to the shorter voice packet without data content. According to the present invention in a particularly preferred embodiment, the receiver clock can be synchronized with the time information in the voice packet during initialization and normalized relative to the voice packet without voice content. Therefore, for a voice packet with the shortest packet transmission time (i.e., no jitter occurs or the shortest jitter occurs), this actually corresponds to the delay of the voice packet on the transmission channel. Therefore, the deviation (e.g., difference) between the time information in the voice packet and the receiver time indicates the current jitter when receiving the voice packet currently under consideration. Handling is thereby facilitated and the method can be used (e.g. for initializing and / or adjusting receiver time) without distinction between voice packets with voice data and voice packets without voice data.

[0025] According to another aspect of the present invention, the determination of the minimum packet transmission duration (which, according to the present invention, results in an adjustment of the receiver time to the temporarily stored minimum packet transmission duration) can occur for the duration of the connection or at multiple intervals during the connection (the intervals preferably directly follow each other). As a result, the receiver time is continuously adjusted for possible changes in the transmission channel, even during the duration of the connection. This preferably occurs in each case at the end of the interval. The continued output of analog audio content corresponding to successively received audio packets having audio content is preferably unaffected by the adjustment of the receiver time and continues. The adjustment of the receiver time only affects the determination of the buffer (which, according to the present invention, is dynamic) if an audio packet having a first audio content is received again after at least one audio packet without audio content has been received (after the receiver time has been adjusted).

[0026] The length of the interval may be defined by a length of time (duration) or by a defined number of received voice packets with no voice content and / or with voice content.

[0027] According to the invention, it can also be provided that the connection duration of the receiver to the sender (or more generally between the sender and the receiver) is limited to a maximum connection duration. The sender and / or the receiver can be set up to prevent a connection after the maximum connection time has been reached, at least if the audio content is to be transmitted again, with the consequence that the connection must be set up anew. This leads to a reinitialization of the receiver clock and a repetition of the above method. Thus, changes in transmission conditions in the transmission channel can be automatically taken into account.

[0028] In another practical embodiment of the present invention, the adjustment of the receiver time according to the temporarily stored minimum relative packet transmission duration may occur in such a way that "the receiver time is advanced by the minimum relative packet transmission duration" upon indication that an audio packet having a shorter packet transmission duration (relative to the current receiver time) has been received. By advancing the receiver clock, the receiver clock is synchronized with the transmission time of the audio packet having the shortest jitter or shortest packet transmission duration.

[0029] Furthermore, the present invention may provide the following: the adjustment of the receiver time according to the temporarily stored minimum relative packet transmission duration may occur in such a way that if the temporarily stored minimum relative packet transmission duration indicates that only voice packets with longer packet transmission times relative to the current receiver time have been received, the receiver time is returned (the receiver time is returned by a specified or definable duration). Returning the receiver time takes into account a longer packet transmission duration for all voice packets and corrects the use of an unrealistically short jitter if, in the case of no reception during the period under consideration (particularly within an interval within the connection period), the previously shortest packet transmission duration (i.e., a small or shortest value of jitter) during previous reception is reached again. This may indicate that transmission conditions are changing. According to a particularly preferred embodiment of the present invention, the definable duration may be realized by the minimum relative packet transmission duration relative to the receiver time, in particular as a time deviation (weighted by a weighting factor) from the receiver time when receiving voice packets. A judicious weighting factor may be, for example, 1 / 3, without this embodiment of the present invention being limited thereto. As a result, time differences between the receiver clock and the transmitter clock (due, for example, to clock differences in the timers used) are also automatically corrected continuously.

[0030] As indicated previously, in a particularly preferred embodiment of the present invention, the determination of the minimum relative packet transmission duration for voice packets without voice content and for voice packets with voice content can occur in the same way (i.e., independently of) even during the output or processing of the voice content. In this case, the above-mentioned standardization for voice packets without voice content can be undertaken. Alternatively, standardization for voice packets with voice content would also be possible. This allows the minimum relative packet transmission duration to be determined continuously during the connection period.

[0031] According to a preferred embodiment of the present invention, the buffer may start with a maximum buffer that adjusts the maximum jitter to be considered based on the delay to reception (i.e., based on the actual jitter) resulting from comparing the time information in the first audio packet with audio content with the receiver time when this first audio packet is received. The maximum jitter to be considered should be understood as the maximum delay to reception caused by random fluctuations in the reception time of audio packets in the receiver. It can therefore be provided that the adjustment of the maximum buffer occurs by shortening this maximum buffer by the actual jitter. If a negative delay to reception occurs, i.e., if the first audio packet received with audio content is faster than the audio packets (including the initial packet) previously used to adjust the receiver time, it is also possible according to the present invention to dispense with adjusting the maximum buffer.

[0032] According to a preferred embodiment of the present invention, audio packets having the same audio content may be received by a receiver via several different transmission channels, and the receiver may apply delay compensation for the audio packets having the same audio content via the different transmission channels. This may occur, for example, so that the delays of the audio packets are compared with the delay compensations of the different transmission channels, and the delay compensation for each transmission channel is formed by the difference in delay from the transmission channel with the longest transmission time. This delay compensation is also called dynamic delay compensation (DDC). The delay compensation may be already known for each transmission channel and may be derived, for example, from the physical conditions of the transmission channel. Those skilled in the art are aware of the basic procedure for such delay compensation (DDC), as well as the physical laws involved.

[0033] In addition to the buffer (jitter buffer), this type of delay compensation also represents an additional latency between receiving a voice packet and processing the voice content from the voice packet. If such a delay compensation is known or can already be determined by the transmitter, it can also be transmitted in the voice packet (e.g., in the voice packet's header). As a result, particularly with regard to the jitter compensation proposed according to the present invention, the following is achieved: the time from sending voice packets with the same voice content on various transmission channels to processing of the voice content by the receiver (radio sending and / or outputting) is the same even in the case of various delays of these voice packets with the same voice content in various transmission channels (via an IP-based network). This avoids echoes in radio communications, for example, in the context of the particularly preferred use of voice-over-IP transmission for operational ground-to-air communications in air traffic, where several receivers (for voice-over-IP communications) simultaneously send voice information as radio transmitters at various ground stations (located in a spatially distributed manner). This means that large airspaces can be covered even with limited radio coverage.

[0034] According to the present invention, such delay corrections can also be determined by the receiver according to the present invention by applying a method of jitter compensation, in which the receiver clock is initialized in the above-described manner for each transmission channel separately during the connection, and after initialization of the receiver clock, the minimum relative packet transmission duration is determined in each transmission channel in the above-described manner, completed by the above-described adjustment of the receiver clock. Since audio packets with the same audio content are transmitted on the various transmission channels, and the receiver clocks are respectively correlated with the transmitter time when sending audio packets with the minimum relative packet transmission duration on the respective transmission channels, delay differences on the various transmission channels can be derived for the various transmission channels from a comparison of the receiver times, in each case associated with the packet transmission duration with the minimum jitter, and a corresponding correction of the delay (dynamic delay compensation DDC) can be derived for each transmission channel.

[0035] According to a preferred embodiment, this can be done, for example, by each receiver of one of several transmission channels, which sends its receiver time back to the sender (round trip delay) after receiving an audio packet (having the identifier of the audio packet and the identifier of the receiver). The sender can then determine the minimum packet transmission duration between the sender and each receiver (in a manner similar to that previously described). From this, the sender can determine a respective correction for each receiver's delay (Dynamic Delay Compensation DDC) and, if necessary, transmit this correction to the receiver, or else include this correction in the time information of the transmitter time contained in the header of each audio packet. This allows automatic Dynamic Delay Compensation DDC to be performed.

[0036] The present invention also relates to a receiver for receiving digital audio packets via an IP-based network and for processing analog audio content contained in the audio packets as claimed in claim 9. The receiver comprises a receiving unit connectable to (adapted to be connected to) an IP-based network, the receiving unit being set up (adapted) to receive audio packets transmitted via the IP-based network, and an arithmetic and logic unit having a receiver clock, the arithmetic and logic unit being set up to process the received audio packets, e.g., payloads carrying digital audio data transmitted in the audio packets being extracted from the audio packets and further processed. The further processing consists, for example, in converting the audio data extracted from the sequence of audio packets into audio content, transmitting this audio data wirelessly (e.g., by digital HF modulation by a modulator), and / or directly outputting this audio data by an audio output unit (in particular a loudspeaker). According to the present invention, the arithmetic and logic unit is set up (adapted) to perform the above-mentioned method for jitter compensation or part thereof according to any one of claims 1 to 8 when receiving and outputting the audio content.

[0037] According to a preferred embodiment, the receiver can be set up to receive audio packets having the same audio content via several different transmission channels and to apply the method of claim 8, where the arithmetic logic unit is set up to process the audio content from the received audio packets having the same audio content from several different transmission channels. In particular, the same audio content can be output (e.g., transmitted wirelessly) from each of several audio packets (or even all audio packets) of the audio packets transmitted via the various transmission channels. By the method according to the invention, the output times of the audio packets having the same audio content are synchronized across the various transmission channels with respect to both compensation for delays of the audio packets via the various transmission channels (dynamic delay compensation (DDC)) and compensation for random time variations (jitter) between the reception times. For the preferred use case of ground-to-air communications, the receiver particularly includes several radio transmitters in various spatially separated ground stations, which transmit the same audio information wirelessly in a time-synchronized manner according to the above method.

[0038] The present invention also relates to a method for transmitting and receiving audio content according to the features of claim 11, in which a transmitter converts audio data into a sequence of digital audio packets with or without audio content, where the audio packets with audio content have a section (called a header) carrying data for communication control and a section (called a payload) carrying audio data digitized from part of the audio content. Audio packets without audio content have a header without a payload. Each audio packet (especially the audio packet header) contains at least one time information related to the transmitter time, which indicates the transmission time of the audio packet. The transmitter sends the sequence of audio packets over at least one transmission channel of an IP-based network. In this method, a receiver receives the sequence of digital audio packets with or without audio content and processes the audio content, and the aforementioned method for jitter compensation (especially as described in any one of claims 1 to 8) is applied in the receiver.

[0039] According to a preferred embodiment of the method, a permanent communication connection can be set up in an IP-based network during a connection period between a sender and a receiver via at least one communication channel, during which time, when no audio content should be transmitted, audio packets without audio content are exchanged (preferably at a specified normal transmission clock rate) to maintain the relationship between the sender and the receiver, and during which time audio content should be transmitted, a series of digital packets or digital audio packets with audio content are transmitted from the sender to the receiver.

[0040] As a result, the jitter compensation method proposed according to the present invention can be performed over a long period of time (continuously) so that when audio content is transmitted, a dynamically adjusted buffer is available at any time according to the current conditions in the transmission channel.

[0041] According to the present invention, in the case of a two-way audio connection in which both communication partners function as a transmitter and a receiver, a combined transmitter / receiver device can also be used, in which the transmitter and receiver functions are both realized in one device. In this sense, functions described here separately for the sake of clarity with respect to a transmitter and a receiver also apply to the combined transmitter / receiver device, respectively. In these devices, where appropriate, some functions are not uniquely assigned to a transmitter or a receiver. The present invention also relates to embodiments in which all of the described features are assigned to the combined transmitter / receiver device, without explicitly assigning functions and features (some of which also overlap) to a transmitter or a receiver. This applies, for example, to the setup of a permanent communication connection, which can in principle be initiated by both a transmitter and a receiver. In the context of a communication setup, as provided for in common communication technologies, especially for IP-based networks, a (pure) receiver of audio content can also have a sending function for the communication setup, and a (pure) transmitter of audio content can also have a receiving function for the communication setup. In the case of a combined transmitter / receiver device, both communication participants may act as a receiver and as a transmitter, for example, to maintain a communication connection, and may exchange voice packets without voice content in both communication directions. This may be done in a common transmission channel for both communication directions, or in one transmission channel for each of the two communication directions, respectively. The same applies if there are several different transmission channels in one communication direction.

[0042] Insofar as they are used to implement the present invention, communication technologies known to those skilled in the art also belong to the subject matter of the present invention, since this expertise can be assumed even if not fully explained herein.

[0043] According to one embodiment of the present invention, it may be provided that a transmitter sends out voice packets via several different transmission channels and a receiver receives the voice packets sent via the various transmission channels. Preferably, it may be provided that the receiver receives the voice channels transmitted via the various transmission channels and processes them simultaneously, so that for example echoes in the output can be avoided thanks to the proposed invention. The receiver may include several spatially distributed wireless transmitters that simultaneously send voice messages by wireless.

[0044] The method can therefore be particularly advantageously used in accordance with the invention in radio communications, in particular for the transmission of operational ground-to-air communications in CLIMAX operations, where several different radio channels (as or in the sense of transmission channels) are used to transmit radio messages (as audio content).

[0045] To this end, the present invention also relates to a device for transmitting and receiving digital audio packets or digital audio packets containing audio content over an IP-based network, the device comprising a transmitter and a receiver, the transmitter comprising an audio recording unit (in particular a microphone) for recording analog audio content and an arithmetic logic unit having a transmitter clock, the arithmetic logic unit being set up (adapted) to process the recorded audio content according to the above-mentioned method or part thereof (in particular as described in any one of claims 11 to 14), and the receiver being constructed according to the features as described above (in particular as described in claim 9 or claim 10) "so that the transmitter and receiver are set up (adapted) as a whole to apply the features of the method (in particular as described in any one of claims 11 to 14) for transmitting and receiving audio content or part thereof".

[0046] Further advantages, features and possibilities for the use of the present invention also arise from the following description and drawings of exemplary embodiments, all described and / or illustrated features belong to the subject matter of the present invention, together or in any sensible combination for a person skilled in the art, even if unrelated to their summary in the described and shown exemplary embodiments or in the claims. [Brief explanation of the drawings]

[0047] [Figure 1] 1 illustrates a schematic time sequence for the transmission of analog audio content over an IP-based network (Voice over IP: VoIP). [Figure 2] 1 shows a schematic time sequence for sending audio packets with the same audio content on two different transmission channels without delay compensation. [Figure 3] 3 shows a schematic time sequence for sending audio packets with the same audio content on two different transmission channels corresponding to FIG. 2 extended by a delay correction. [Figure 4] 4 shows a schematic time sequence for sending audio packets with the same audio content on two different transmission channels corresponding to FIG. 3 with dynamic jitter compensation according to the invention; [Figure 5] 5 shows a schematic time sequence for sending audio packets with the same audio content on two different transmission channels corresponding to FIG. 4 with dynamic jitter compensation according to the present invention; here, several different audio contents are transmitted consecutively in different audio packet sequences. [Figure 6] 2 shows a schematic representation of a connection period between a transmitter and a receiver with voice packets with and without voice information being received by the receiver; [Figure 7] 4 shows a flow chart for the inventive initialization of a receiver clock and for the inventive determination of a minimum relative packet transmission duration; [Figure 8] 9 shows an extract of a log file for the execution of the method according to FIG. 8. [Figure 9] 10 shows a flowchart for delaying the first voice packet of a wireless message by dynamic buffer adjustment. [Figure 10] 10 shows an excerpt of a log file for the execution of the method according to FIG. DETAILED DESCRIPTION OF THE INVENTION

[0048] Before describing practical embodiments of the present invention, the known principles of voice transmission over IP-based networks (Voice over IP or VoIP applications) and the terminology used in this text should be explained.

[0049] 1 shows a schematic time sequence of a method for transmitting and receiving analog and / or digital audio content, as is already known in principle in the prior art and also as is applicable in principle in the context of the method according to the invention. It should be noted that the time scale is only intended to be understood schematically and therefore does not make any statements regarding real time ratios. In particular, the actual transmission times of audio packets are substantially shorter than those shown schematically in the sketch.

[0050] When transmitting audio (e.g., analog audio content 1) over an IP-based network (Voice over IP / VoIP), for example, recorded speech is digitized in known manner (e.g., by pulse code modulation (PCM)) so that audio content 1 exists in its digitized form as digital audio data, also referred to as audio content 1.

[0051] The digital audio data is combined into data packets (audio packets 110), where only a portion of the entire audio content 1 is contained in one audio packet 110. This is made clear by the vertical lines in audio content 1. Each digital audio packet 110 is also referred to as an audio sample.

[0052] Therefore, the audio content 1 is usually divided into a sequence 119 consisting of several consecutive audio packets 110. A sequence 119 with four audio packets is shown by way of example in FIG.

[0053] Each audio packet 110 contains digital audio data, which includes a section of the audio content 1 (i.e., a portion of the audio content 1) of, for example, a few milliseconds. Such a section of (digital or digitized analog) audio data (contained in the audio packet 110) is also referred to as a payload PL. The audio packet 110 also includes, as an important component, a header H containing data for communication control. For example, the sampling time of the first audio sample is included in the header H as a timestamp (in the sense of time information specifying the transmission time ts). This is adopted as the transmission time ts of the audio packet 110. Sampling of the audio content 1 is understood to mean sampling the analog audio content 1 at a sampling frequency (e.g., 8 kHz), during which digital values are respectively assigned to the audio content 1 (i.e., the audio content is digitized as audio data and divided into packets (payload PL)).

[0054] An audio packet 110 is generated by a transmitter (also referred to in technical terms as a source (of the audio packet)) and then transmitted to a receiver (also referred to in technical terms as a data sink (of the audio packet)). The transmitter inserts a transmission time ts into the header and transmits the audio packet 110. If an audio packet 120 (not shown in FIG. 1 but described further below) without audio content 1 (in which the payload PL is omitted) is transmitted, the transmission time ts corresponds to the transmission time ta at which the audio packet 120 was transmitted. If an audio packet 110 with audio content 1 (shown in FIG. 1) is transmitted, the transmission time is delayed by the sampling of the audio packet 110. This takes a period TPS, and as a result, the actual transmission time (actual transmission time ta) of the audio packet 110 is shifted by the period TPS (i.e., ta = ts + TPS). This period TPS required to sample the payload PL is also described as the payload length and can be considered constant (a correction for this is described further below). The sending of the audio packets 110 is shown with reference to timeline 50 and occurs by a sender (meaning the source of the audio packets).

[0055] As shown in Figure 1, there is a certain time after this transmission until the voice packet 110 is received at reception time te. In principle, this could start with the output of the payload PL. This is shown in Figure 1 in an IP-based network transmission channel 10 of four voice packets 110, where each voice packet 110 is shown on a row below each other with respect to the timeline 51.

[0056] The delay DEL of the voice packet 110 and the jitter JIT of the network represent two important influencing variables during the transmission of the voice packet 110 through the transmission channel 10 .

[0057] Delay DEL or latency essentially refers to the total (physical) signal delay on the transmission channel. Voice transmission over IP networks involves additional delays due to temporary storage and, if necessary, data reduction, data compression and decompression. This is not further considered here, as these delays can be considered constant over short time ranges (as can be assumed for the application of the present invention). These constant delays should also be included within delay DEL in the context of this text.

[0058] The random time variation (technically unavoidable) between the reception of two voice packets 110 is called jitter (JIT). This jitter (JIT) has the consequence that voice packets 110 of the same length are transmitted at exactly the same time interval relative to each other but arrive at the receiver at different time intervals. This is illustrated in Figure 1 by the different width boxes (JIT) of each voice packet 110 (of which only the payload PL, which is essential for the present invention, is shown), and the headers (H) have been omitted for clarity.

[0059] It should be noted again that the time lengths of the individual components (H, PL, DEL, JIT, SBJ) in the diagram do not reproduce the actual durations of the components relative to each other, nor do they represent a timeline with uniform periods over its length. The diagram is used only to understand the sequence of the method described in accordance with the present invention in the context of the present invention, without using a uniform (absolute or relative) time scale.

[0060] These various time intervals between the reception of consecutive audio packets 110 may have the effect that during further processing (e.g., playback or use) of the transmitted audio information 1, it may not be possible to connect the payload PL of the individual audio packets 110 to a continuous output stream, as outlined in FIG. 1 for the first variant of audio output 90 (as an example of processing received audio data). Here, output of audio information begins immediately after receiving the payload PL of the first audio packet 110 at reception time te1. The payload PL of the second audio packet 110 is received at time te2, before processing (e.g., playback) of the audio information from the payload PL of the first audio packet is completed. Thus, the payload PL of the second audio packet 110 is temporarily stored in the receiver until processing of the payload PL of the first audio packet 110 is completed. Therefore, processing of the payload PL of the second audio packet 110 may directly follow processing of the payload PL of the first audio packet 110. The same applies to the payload PL of the third audio packet 110, so that, for example, a continuous output (in the sense of possible processing of the payload PL) of audio information is possible up to this point.

[0061] However, due to the particularly long jitter JIT, the payload PL of the fourth voice packet 110 is only received at the reception time te4, when processing of the payload PL of the third voice packet 110 has already been completed.

[0062] Therefore, any possible desired playback (or other processing) must wait until reception time te4, which leads to a pause or interruption 92 in the audio output (processing the payload). Continuous processing is therefore not possible.

[0063] What are known as "buffers" (jitter buffers) are used in the prior art to compensate for these various time intervals (jitter JIT) in receiving successive voice packets 110 and to enable continuous voice output. Thus, for voice transmission, a defined time (latency) is awaited after the payload PL of the first voice packet 110 from a sequence 119 of voice packets 110 is received at a reception time te1 before further processing or voice output. This is shown in the second voice output variant 91 of FIG. 1, where there is a wait for a static, fixed latency period (called a (static) buffer SJB (static jitter buffer)) before outputting the voice information of the payload PL of the first voice packet 110 at a time tw after its reception at time te1. Thus, after the reception time te1 of the first voice packet 110, there is a wait for a fixed buffer SJB (meaning a period) before voicemail playback / processing can begin. During the buffer SJB, subsequent voice packets 110 continue to be received and temporarily stored. This means that after playback of the audio information (payload PL) from the first audio packet 110, the payload (PL) of the subsequent audio packet 110 has already been received at reception time te2 and is temporarily stored so that this subsequent audio packet can be played back directly following the payload PL from the first audio packet 110. The same applies to further subsequent audio packets 110, thereby enabling continuous playback by the receiver of the entire transmitted audio content 1. The term "playback" should be understood as synonymous with "processing" in the sense of the present invention and should also include other processing of the received payload PL from the audio packet 110, such as audible playback. An example of this is described below.

[0064] The (static) buffer SJB therefore provides an additional intentional delay in processing (audio output in the described example) in order to subsequently output the audio data concurrently. Audio packets 110 arriving later than the delay of the static buffer SJB can no longer be incorporated into the output data stream. The size of the buffer SJB is added to the delay DEL. The size of the buffer SJB therefore allows for the choice between more delay (and a lower packet loss rate) or less delay (and a higher packet loss rate). If a subsequent audio packet 110 is lost in the sense that its payload PL is not yet present for processing / output after the processing / output of the payload PL of the preceding audio packet, there is an interruption in the continuous processing / audio output.

[0065] Thus, the following methods are known from the prior art:

[0066] Audio packets 110 having a defined portion of the entire audio content 1 (digital or digital audio content) as payload PL are generated in a transmitter (data source) (in practice, with a payload length of, for example, 10 ms) and transmitted at equal time intervals (i.e., at the same transmission frequency). This is shown in the timeline 50 by equidistant transmission times ts1 to ts4. The audio packets 110 arrive late at the receiver (data sink) due to delay DEL and jitter (JIT) on the transmission channel 10. The delay DEL on the transmission channel 10, applied in the context of the present invention for successive audio packets 110 (e.g., radio messages in a digital wireless connection), should be considered constant over the short period under consideration, while the jitter JIT per audio packet 110 may differ and usually also lie within a constant jitter bandwidth. This leads to audio packets 110 no longer arriving at the receiver with constant time intervals relative to each other. This is shown in the timeline 51 by non-equidistant reception times te.

[0067] To eliminate these time differences in the further processing of the packets, the audio output (audio output 90) does not occur immediately upon arrival of the first audio packet 110 at reception time te1. The audio output 90 is further additionally delayed by a buffer SJB (jitter buffer time) relative to time tw.

[0068] If this were not done and instead further processing started immediately upon arrival of the first voice packet, this would lead to an audio gap (interruption in voice output 92) between the consecutive third and fourth voice packets 110 in the above example according to FIG. 1, since the fourth voice packet 110 would not have arrived yet after the voice output of the payload PL of the third voice packet 110.

[0069] It can be clearly seen from Fig. 1 that the start of the audio output tw depends on when the first audio packet 110 arrives at the receiver (time te1) and how large the buffer SJB is. Therefore, the start of the audio output depends on the delay DEL, the jitter JIT of the first audio packet of a sequence of audio packets 110, and the length of the buffer SJB (jitter buffer). The delay DEL and the buffer SJB can be considered constant in the context of normal audio transmission (e.g., as a radio message). The jitter JIT remains a variable influence variable.

[0070] For example, with regard to the use of voice over IP for the transmission of operational ground-to-air communications between aircraft and air traffic control, there is an additional factor (echo) that directly contributes to the jitter JIT and delay DEL of the voice packets.

[0071] Regardless of the particular example, this applies whenever a radio message (e.g., from air traffic control) is transmitted, for example, from air traffic control monitoring the appropriate airspace to aircraft moving within the airspace. The actual affected airspace is not always covered by a single radio transmission route from a single radio transmitter. Rather, in many cases, several radio transmitters transmitting radio messages from various locations to aircraft in the airspace are provided in a spatially distributed manner (as ground stations) so that radio receivers onboard the aircraft can receive the radio messages within the entire airspace. Thus, safety-related communications can be realized as needed (e.g., in the case of ground-to-air communications in radio operations).

[0072] Radio transmission by radio transmitters on the ground to radio receivers in aircraft occurs essentially at the speed of light. For the radio links in question herein, delay differences during transmission can be ignored. At this time, radio messages (formed by the transmitted audio content 1) received by the radio receiver are typically output directly by the radio receiver. This is conventional and therefore not the subject of the present invention. Typically, this radio link between the radio transmitter and the radio receiver is also not implemented as a voice-over-IP transmission. Regarding the particularly preferred embodiment of ground-to-air communications described herein, the present invention relates to jitter compensation over an IP-based network for transmitting voice mail from a central source (referred to herein as a "transmitter" in the sense of a transmitter in a voice-over-IP transmission) to a radio transmitter (referred to herein as a "receiver" in the sense of a receiver in a voice-over-IP transmission). Then, several radio transmitters wirelessly transmit voice mail into the airspace for reception by radio receivers onboard aircraft.

[0073] Thus, the wireless transmitter receives digital voice packets 110 via a voice-over-IP connection and combines the payload PL to form voice mail 1, which is then sent out over the air in a modulated manner using conventional wireless technology. This is what is meant by "processing" the voice content contained within the voice packets and also ultimately by outputting the voice content.

[0074] Furthermore, the delay of a wireless message through an IP-based network via Voice over IP (VoIP) from a transmitter (the source of the wireless message), i.e., a wireless transmitter for sending a wireless message via radio, to a receiver (the sink of the wireless message in a Voice over IP transmission) is different, unless the associated delay variations of various transmissions of the same wireless message (leading to echoes during playback of the same wireless message through various wireless transmitters (CLIMAX operation)) typically occur during the actual wireless transmission between the wireless transmitter and the wireless receiver. This leads to different sending times of the same wireless message through various wireless transmitters and generates echoes during playback of the wireless message at the receiver. The entire transmission path of a wireless message from the source of the wireless message (e.g., air traffic control) to playback after reception of the wireless transmission by the wireless receiver is called a wireless channel, where delay differences that typically cause echoes occur during Voice over IP transmission through an IP-based network and cause echoes. In this regard, a wireless channel in the sense of the present invention also always means a transmission channel in the general sense. The reception (and transmission) of continuous audio content 1 is also referred to as a wireless message in the described example. Therefore, a radio message also always means audio content in the general sense.

[0075] If a voice packet 110 having the same voice content 1 is successfully transmitted by the transmitter over two or more wireless channels, the voice packet 110 is transmitted by the transmitter as voice packet 110 over the first channel 10 and as voice packet 210 over the second wireless channel 20. Thus, the receiver also receives the first voice packet 110 and the second voice packet 210 having the same voice content 1 over the respective wireless channels 10 and 20 (FIG. 2). The structure and time sequence of the voice packets 110, 210 correspond to the previous description with reference to FIG. 1 with the voice packet 110. The payloads of the voice packets 110, 210 are identical.

[0076] Such transmission of voice packets 110 and 210 (also referred to as voice packets 110 and 210 to distinguish between the wireless channels 10 and 20) having the same voice content 1 on the wireless channel 10 and the wireless channel 20 is also referred to as CLIMAX operation, which is not limited to the two wireless channels 10 and 20, but rather applies generally to several wireless channels.

[0077] CLIMAX operation (German: "Ueberdeckung" / English: "overlap") is understood to mean the parallel and essentially simultaneous transmission (broadcast) of voice packets 110 and 210 by transmitters from several transmitting locations (radio transmitters) on the same frequency. Thus, a transmitter (as a source of radio messages) uses several different radio transmitters (sending the same voice packets 110 and 210 as simultaneously as possible on the same radio frequency) at various locations, so that the voice contents overlap at the radio receiver insofar as the radio messages sent by the various radio transmitters are received at the (same) receiver. This use case is a particularly preferred application of the present invention.

[0078] Thus, this type of CLIMAX operation is used, for example, in ground-to-air communications to cover very large airspace or areas with spatially difficult coverage (e.g., caused by mountains) by using one transmission frequency. Upon receiving radio message 1 (i.e., of the transmitted voice content), the pilot may be in an area where he receives voice packets 110 and 210 with the same voice content 1 from two or more radio transmitter locations.

[0079] Therefore, it must be ensured by suitable technical means that the transmission of radio messages occurs as simultaneously as possible (in accordance with Regulation ED-137 within 10 ms) at all radio transmitters to prevent the occurrence of echoes of the pilot's radio messages.

[0080] 2 illustrates a situation in which a wireless message originating from a central transmitter (as the source of the wireless message) is sent by sending four voice packets 110 on a first wireless channel 10 and four voice packets 210 on a second wireless channel 20, in both cases via a voice-over-IP connection. Here, the four voice packets 110 and the four voice packets 210 have the same voice content 1. The voice packets should be processed simultaneously within the receivers. In a particularly preferred use case of CLIMAX operation as described herein, this means that several receivers are provided, each with one wireless transmitter (which broadcasts the received voice packets 110 and 210 wirelessly via a voice-over-IP connection for reception by one wireless receiver). To prevent any echoes from occurring at the wireless receivers, the voice packets 110 and 120 received via the voice-over-IP connection should be sent simultaneously by each wireless transmitter (which in both cases corresponds to one receiver of the voice-over-IP connection) at various wireless transmitter locations.

[0081] It can be seen in this example that the delay DEL from the transmitter to the receiver on the radio channel 20 is longer than the delay to the receiver on the radio channel 10. The difference in delay occurs due to various interferences or characteristics within the radio channels 10 and 20 (or more generally within the transmission channels 10 and 20). These (different) delays DEL within the respective radio channels 10 and 20 should be understood to be constant for a given short period of time under consideration of the radio message, and therefore these differences in delay may be referred to as static delay differences. In principle, these could also be determined, for example, from the current positions of the transmitter and receiver on the respective radio channels 10, 20 (delay differences due to various distances) or by determining interference (e.g. in the context of a connection setup on the radio channels 10 and 20).

[0082] Therefore, due solely to the different delays DEL of the audio packets 110 and 210 on the different wireless channels 10, 20, there are different output times tw of the same audio content 1 at the receiver due to the transmission of the same audio packets 110 and 210 over the different wireless channels 10, 20 (i.e., output time tw(10) on wireless channel 10 and output time tw(20) on wireless channel 20).

[0083] In the preferred use case, the output time ta at the receiver refers to the sending of the audio content by the wireless transmitter. The audio content 1 is then simply output as analog speech at the wireless receiver. Sending the audio content over the radio then leads to undesirable echoes in the (analog) playback of the audio information by the wireless receiver. Therefore, the entire system for processing audio packets during and after reception over an IP-based network connection (Voice over IP) can be understood as a receiver in the sense of the present invention. In the exemplary case described, this is therefore a system consisting of a wireless transmitter (as a receiver of audio packets over a Voice over IP connection) and its output (in the sense of further processing) by sending audio information over the radio and receiving it by a wireless receiver, which then outputs the audio content in an audible manner. The wireless link downstream of the receiver (in the sense of further processing of the received audio packets for output at the end of the wireless link) effectively no longer introduces (new or additional) delay differences on the various wireless channels. In the following, only the behavior during Voice over IP transmission over an IP-based network is considered, regardless of whether the direct (analog) audio output occurs immediately after the Voice over IP transmission or only after a subsequent wireless transmission between the wireless transmitter and the wireless receiver.

[0084] A function called "Dynamic Delay Compensation (DDC)" is described in standard ED-137 to compensate for static delay differences from the transmitter to the receiver over the various transmission or radio channels 10, 20. Here, the various values of the delay DEL of the various channels 10 and 20 due to the additional delay compensation DDC are adjusted in such a way that the transmission on the "fast" radio channel 10 is delayed by the delay compensation DDC until it corresponds to the delay DEL on the "slow" radio channel 20. This delay caused by the delay compensation DDC is called dynamic delay compensation. This delay compensation can be achieved by a delayed transmission at the transmitter's location (i.e., within the transmitter's area) of the "faster" radio channel 10, or by a waiting time (similar to a buffer) during the processing of the voice information by the receiver. The delay difference can be known and / or calculable and applied accordingly at the transmitter or receiver. The delay difference can also be determined and applied during the connection by the receiver in the manner previously described.

[0085] The result of applying the delay compensation is shown in Figure 3 for the case according to Figure 2. Consequently, the total actual delay and the delay compensation DDC (if any) of the voice packets 110 and 210 is the same for each radio channel 10, 20. The occurrence of echoes then no longer depends on the different delays DEL.

[0086] 3 even makes it clear that the jitter JIT (i.e., the arrival time of the audio packets 110, 210) varies for each of the wireless channels 10, 20 (transmission channels). It is not difficult to understand that during CLIMAX operation of the wireless channels 10 and 20, due to the delay difference, there may also be problems with echoes that are solely responsible for the different jitter JIT. This also leads to a difference between the output time tw(10) via the wireless channel 10 and the output time tw(20) via the wireless channel 20 of the same audio content 1.

[0087] As mentioned above, the reception time te1 (see FIG. 1) of the first voice packet 110 or 210 from the sequence 119 determines the start of the voice output tw. If the first voice packet 210 arrives earlier relative to the voice packet 110 (i.e., if the current jitter of this voice packet 210 is small), voice playout will start earlier. If the first voice packet 110 arrives later relative to the voice packet 110 (i.e., if the current jitter of the voice packet 110 is large), voice playout will start later, and here "voice playout" or "voice output" also refers to retransmission over the air.

[0088] This issue is not taken into account in the ED-137 standard due to delay compensation when using some radio or transmission channels.

[0089] However, this problem is overcome by the proposed method for dynamic jitter compensation by means of a dynamic buffer: it is realized that the voice output start time tw of a radio message (or a sequence 119 of voice packets 110 and 210) is no longer dependent on the jitter of the first voice packet 110 and 210 of the sequence 119 of voice packets.

[0090] This is achieved in accordance with the present invention in that the actual receipt time (synonymous with actual jitter) is determined for the first voice packet 110, 210 of the sequence 119 of voice packets 110 and 210 representing the voice mail (i.e., voice information 1), and the buffer is fixed (in the sense of being determined or defined) as a dynamic buffer DJB depending on the actual receipt time (synonymous with actual jitter). The length of the dynamic buffer is fixed to depend on the actual jitter in such a way that the period made up of jitter JIT and the dynamic buffer DJB is constant.

[0091] This is shown in Figure 4 with respect to the time sequences described in Figures 2 and 3. Instead of a fixed buffer SJB according to Figure 2 or 3, according to the invention there is a dynamic buffer DJB which dynamically adapts to the jitter (jitter value, actual reception time) of the first voice packets 110 and 210 of a radio message or audio content 1. If the first voice packets 110 and 210 arrive relatively early (with a low jitter value), this leads to a buffer DJB which is longer in time. Therefore, the sum of the actual jitter JIT (jitter value in the sense of no jitter or random reception time for the shortest jitter occurring) and the dynamic buffer DJB is constant according to the invention. Reception here means reception in Voice over IP transmission.

[0092] Considering now several radio messages 1, 2, 3, 4 on various radio channels 10, 20, it can be seen in FIG. 5 that the actual jitter JIT of each radio message 1, 2, 3, 4 and each radio channel 10, 20 is different.

[0093] Nevertheless, the start time tw of the audio output is the same for all radio messages 1, 2, 3, 4 within radio channels 10, 20 and for all radio messages 1, 2, 3, 4 on all radio channels 10 and 20. Therefore, further processing of the audio packets occurs simultaneously for all radio messages and for all of the transmission channels.

[0094] When the first voice packet 111, 211 of the sequence 119 of voice packets 110, 210 (i.e., one of radio messages 1, 2, 3, 4, for example) arrives, the challenge in implementing the present invention lies in detecting how high the actual jitter of this first voice packet 111, 211 actually is, or, expressed differently, how large the relative packet transmission duration of this first voice packet 111, 211 is relative to the total occurring random time fluctuations (jitter) between the reception times of voice packets during reception of voice content over an IP-based network. Therefore, it must be quantitatively determined whether the first voice packet 111, 211, which fixes the length of the buffer DJB, arrived relatively early or late.

[0095] This feature is unknown in "normal" Voice-over-IP applications and is also not described in the ED-137 standard. It is possible to determine the actual jitter of the first voice packets 111 and 211 of the sequence 119 of voice packets 110 and 210 whenever a Voice-over-IP connection (in the sense of the transmission of analog voice content 1, 2, 3, 4 over an IP-based network) between a source (sender) and a sink (receiver) is not set up only when a sequence 119 of voice packets 110 and 210 with voice content representing digital voice data (payload PL) is to be broadcast, but rather the connection exists for a certain connection period during which several (distinct) voice contents (radio messages) 1, 2, 3, 4 are to be transmitted. This is possible, for example, in the case of a telephone call with a Voice-over-IP connection if the connection is set up (permanently) for the duration of the telephone call, but during pauses in the conversation, not voice packets 110, 210 with voice content are transmitted, but rather voice packets 120 without voice content that are used to maintain the connection. As soon as a new audio content 1, 2, 3, 4 is sent, a new dynamic buffer DJB in the sense of the present invention is established.

[0096] A further use case is, for example, a radio connection in ground-to-air communication, where it is important that the connection exists permanently because messages need to be transmitted quickly. Setting up a new connection for each radio message 1, 2, 3, 4 would take too long and involve too many risks. This type of radio connection can be developed, for example, according to the ED-137 standard. This type of application is described in detail below. However, the invention is not limited thereto, but rather can be used in all applications where a voice communication connection exists for a relatively long time (permanently).

[0097] In the context of use according to the ED-137 standard, a permanent voice-over-IP connection exists within an IP-based network, during which voice packets 110, 120, 210 are continuously transmitted in RTP format over each radio channel 10, 20 used for both voice transmission and connection maintenance. This is illustrated in FIG. 6 for connection duration 30 and radio channel 10, where three radio messages 1, 2, 3 (containing voice content (i.e., voice data) in digitized form) are shown as an example, with three dots indicating that not all radio messages 1, 2, 3, 4 and voice packets 110, 111, 120 are shown. Radio messages 1, 2, 3 in each case include a voice packet 110 containing voice content (i.e., having a header H and a payload PL) and form a sequence of voice packets 119. The same applies analogously to the other radio channels 20.

[0098] If no radio messages 1, 2, 3 are currently being transmitted, then an "empty" voice packet 120 with no voice content (without any voice data (payload PL) and only a header H) is transmitted. The sending of voice packets 110 and 120 preferably occurs at a predetermined transmit clock rate.

[0099] The first voice packet 111 of a sequence of voice packets 110 forming a radio message 1, 2, 3 is in each case an identified voice packet 111, which is identified in that it is the first voice packet 110 having voice content, which follows one or more voice packets 120 having no voice content and thus indicates the start of a new radio message 1, 2, 3. For this first voice packet 111 having voice content of the sequence of voice packets, the actual time of receipt (i.e., actual jitter) of this voice packet 111 is determined by comparing the time information in the received voice packet 111 with the receiver time at the time of receipt. The buffer DJB (Dynamic Buffer / Jitter Buffer) is determined in accordance with the present invention depending on the actual time of receipt.

[0100] With the support of this permanent connection during the connection period 30 and the voice packets 110 and 120 exchanged during the connection period 30, it is possible to determine in a simple way the shortest packet delay (considered relatively) and thus to calculate the actual jitter in the receiver. When the first voice packet 111 to be played out of the sequence 119 of voice packets 110 arrives, the actual jitter (jitter value) or reception time of this first voice packet 111 is known and the delay time to the start of the playout of the voice packet (i.e. the buffer DJB according to the invention) can be dynamically adjusted accordingly.

[0101] The procedure according to the invention therefore consists of two functional parts: Part A: Determining minimum packet delay between a sender (data source) and a receiver (data sink) during transmission over an IP-based network (Voice over IP) Part B: delaying the first voice packet 111 of radio messages 1, 2, 3 (sequence 119 of voice packets) by dynamically adjusting (in terms of the size of the buffer) the buffer DJB;

[0102] These two parts operate independently of each other and both run continuously. This procedure is presented in detail below.

[0103] Part A The audio packets 110 and 120 are generated by a sender (data source) according to the ED-137 standard and are labeled as RTP packets. The sampling time of the first audio sample is the timestamp (T RTP ) into the RTP header H. This is adopted as the transmission time ts of the audio packet 110 (also the RTP packet). Sampling of the audio content 1 is understood to mean sampling the analog audio content 1 at a sampling frequency (e.g., 8 kHz), during which digital values are assigned to the audio content 1, i.e., the audio content is digitized as audio data and divided into packets (payload PL).

[0104] After sampling, the voice packet 110 is transmitted via Voice over IP over an IP-based network to a data sink, where the payload PL is also attached to a header H. Thus, the relative or actual time of departure of the voice packet 110 at the sender is given by the timestamp T RTP The payload PL is given by the sum of the transmission time ts and the length or duration TPS of the payload PL (i.e., the time length or size of the included audio data). PL and in practice may be, for example, 10 ms.

[0105] For audio packets 120 (also called R2S packets) that do not have audio content, the sampling period is omitted: the R2S packets have the length or duration TPS (RTP of the RTP packet) of the payload PL. PL Therefore, for a voice packet 120 that does not have voice content, the timestamp T RTPcorresponds directly to the transmission time (transmission time ta). This difference between the two types of voice packets 110 and 120 is in fact that the receiver time is normalized or correlated to the voice packet 120 which has no voice content, and the expected reception time is accordingly scaled with the duration of the payload RTP PL This is taken into consideration in that it is corrected by

[0106] A timer (receiver clock T) with the same nominal frequency as the transmitter timer (data source transmitter clock). RECEIVER , below T SINK ) runs in the receiver (data sink). In other words, the clocks (timers) of the transmitter and receiver run at the same rate for a period considered in the context of the present invention, which is on the order of the size of the radio message. Over longer periods of time, design-induced deviations of the sink and receiver timers occur but are automatically corrected in accordance with preferred embodiments of the present invention. This has already been explained and is also shown once again in the algorithms according to Figures 7 and 9 described below.

[0107] After the session is set up (meaning a connection between the sender and receiver that is maintained for a connection period of 30), the timer T SINK is the time value T of the first arriving voice packet 110 with voice content (RTP packet) or the first arriving voice packet 120 without voice content (R2S packet). RTP +RTP PL where RTP PL = 0 applies to audio packets 120 that have no audio content. Therefore, the time required for sampling an audio packet 110 of payload PL (i.e., duration RTP PL ) which in effect corresponds to the normalization of the receiver clock to the transmitter clock for delay of voice packets 120 that have no voice content.

[0108] As a result, after receiving the first voice packet (called the initial packet), for the duration of the receiver's connection on the voice packet transmission channel, the receiver clock is initialized by using the transmitter time according to the time information in the initial packet so that the reception time correlates with the transmission time when the voice packet is sent out.

[0109] Since this time, the time (timer value) T for arriving RTP and R2S packets RTP +RTP PL : That is, the timestamps T of the voice packets 110 and 120 RTP (If appropriate, the duration of the payload of the voice packet 110 having the voice content is RTP PL The time (corrected by ) is compared with the receiver time (the timer value of the sink's timer) at the time of receipt of the voice packet to determine the minimum relative packet transmission duration.

[0110] Algorithmically, this can be achieved as follows:

[0111] The algorithm shown in FIG. 7 begins 300; then the session setup is established 301; the first voice packets 110 and 120 are received 302; then at receiver time T SINK =T RTP +RTP PL Therefore, the receiver time T SINK is the transmitter time T at the transmission time ts of the voice packets 110 and 120. SOURCE and put into the header H of the voice packet 110, 120: SOURCE is T RTP The payload RTP of the audio packet 110 having the audio content 1 is PL The possible duration of the

[0112] Adjustable interval T during connection duration 30 i (e.g., i maxis equal to 100 or 200 voice packets), the following processing occurs after every reception 304 of a voice packet 110 and 120 (RTP or R2S packet):

[0113] Each interval T i At the start of MIN is the value MAX JITTER (e.g. 200), where MAX JITTER defines the maximum delay of the reception time that must be taken into account. This corresponds to the maximum jitter that can occur. This time is defined in this case as a value without a unit of measurement. For example, the unit of measurement could be 125 μs (1 / 8 kHz): i.e. the sampling time. However, the invention also works with any other desired sampling time.

[0114] First step 305: Interval T i Within the time, the deviation of the actual reception time at the receiver (sink) is then calculated for each voice packet 110 and 120 (RTP packet, R2S packet). This occurs in principle by forming a delta difference of the actual transmission time (for a voice packet 120 with no voice content, R2S packet: timestamp T RTP , RTP PL = 0; for voice packets 110 that have no voice content, RTP packets: timestamp T RTP +RTP PL ) and receiver time (timer T SINK ): Delta = T RTP +RTP PL -T SINK where RTP PL =0 applies to voice packets 120 that have no voice content (R2S packets).

[0115] This is based on the following idea: after the described initialization of the receiver time, the difference delta is equal to zero if the jitter at the time of reception of the voice packets 110 and 120 for initializing the receiver time is equal to the jitter at the time of reception of the voice packets 110 and 120 currently considered, since the correlation of the receiver time corresponds exactly to the sender time plus delay and jitter.

[0116] If delta is greater than zero, the current voice packets 110 and 120 are faster than the fastest voice packet ever that formed the basis for initializing or post-adjusting the receiver time. On the other hand, if delta is less than zero, the current voice packets 110 and 120 are slower than the voice packet that formed the basis for the receiver time.

[0117] Therefore, in this regard, the receiver time (here: T SINK ) is the transmitter time (here: T RTP +RTP if appropriate PL ) and the relative packet transmission duration is determined relative to the receiver time corresponding to the value delta.

[0118] Second step 306: Therefore, the fastest relative voice packets 110 and 120 during the connection period 30 can in principle be determined as a result.

[0119] According to a practical embodiment, it is additionally proposed that: initially, each interval T i , the voice packets 110 and 120 with the shortest relative delay (i.e., with the lowest relative jitter) are assigned to the variable T MIN To this end, the highest relative jitter value (T) that can be absolutely conceived or practically unattained is found by comparing the delta. MIN =-MAX JITTER ) to each interval T i The variable T, which is set in principle at the start of MINhas a shorter relative delay (i.e., T MIN <Delta is applied) voice packets 110 and 120 are in interval T i If found in , it is set to the value of delta.

[0120] Therefore, the measurement interval T i When finished, T MIN is the measurement interval T i This defines the minimum relative delay of voice packets 110 and 120 occurring within the

[0121] The resulting minimum relative packet transmission duration is the interval T i T until the end MIN is temporarily stored as

[0122] Third step 308: Each interval T i After the completion of (307), the receiver time T SINK is the temporarily stored minimum relative packet transmission duration T MIN (when sending a voice packet having a minimum relative packet transmission duration, the receiver time T SINK is the transmitter time T SOURCE (which correlates with

[0123] This can be achieved algorithmically by using the following steps.

[0124] T MIN If is greater than zero (requirement 309), this is the "final interval T i This means that at least one voice packet 110 and 120 has arrived within the time period with a shorter jitter JIT (i.e., shorter packet transmission time, delay) than before. In this case, the receiver time (the timer of the data sink, T SINK ) is corrected as follows (Receiver Time Correction 312): SINK =T SINK +T MIN .

[0125] Therefore, the resulting receiver time is the value T MIN The time correction or timer correction described herein is the time difference between the receiver clock (sink timer) and the interval T i Or it is used to synchronize the fastest voice packets 110 and 120 within the entire connection period 30: i.e., the transmitter time and receiver time are equal with respect to the sending time and receiving time of the voice packet with the shortest occurring jitter JIT (i.e., the shortest occurring packet transmission time until reception).

[0126] Therefore, in the presented example, the adjustment of the receiver time according to the temporarily stored minimum relative packet transmission duration occurs in such a way that if the receiver time indicates that an audio packet with a shorter packet transmission time has been received, the receiver time is set back by the minimum relative packet transmission duration. This occurs in a simple and reliable way using the embodiments described herein. However, other implementations are also conceivable, for example, in which at least a certain number of faster packets must occur within the measurement interval before a correction of the receiver time is attempted. Random measurement errors can be compensated for as a result.

[0127] T MIN If is equal to 0, this is the "final measurement interval T i This means that at least one voice packet 110 and 120 with the same (shortest) jitter JIT (i.e., shortest packet transmission time, delay) as before has arrived within the time limit. In this case, the receiver time (the timer of the data sink, T SINK ) is not corrected but rather remains unchanged.

[0128] T MIN If is less than 0 (request 310), this is the "final measurement interval T i This means that no voice packet with a jitter (packet transmission time, delay) shorter than or equal to the previous one has arrived during the connection period 30. In this case, the receiver time (data sink timer, T SINK ) is corrected as follows:SINK =T SINK +(T MIN / G), where G is a weighting factor.

[0129] T MIN Due to the weighting (311) of the temporarily stored minimum relative packet transmission duration T, the receiver time is returned if the temporarily stored minimum relative packet transmission duration indicates that only voice packets with a longer packet transmission duration have been received. MIN where the receiver time is set back by a specified or definable duration. The receiver time or timer correction described herein is used so that the receiver clock is not corrected exclusively in the direction of faster voice packets 110 and 102 (which have shorter jitter JIT). The correction is also sensitive in the opposite direction (only slower voice packets 110 and 102 are received over a relatively long period of time). Thus, for example, the fact that the clocks at the source and sink are not running at exactly the same speed (especially in the case of a relatively long connection duration 30) can be taken into account.

[0130] It is also automatically taken into account that due to changes in the transmission route (i.e., changes in the transmission channel 10 in particular), all voice packets 110 and 120 arrive late, and therefore the already realized values of the earliest voice packets 110 and 120 (with the shortest jitter JIT) are no longer valid.

[0131] The correction method described herein takes into account that if "later" packets occur exclusively within the measurement interval, the receiver time is corrected by a weighting factor (e.g., a weighting factor of 3 is set). Other implementation variations are conceivable and possible, such as correction by a fixed value or a different weighting factor (preferably in the range 1 to 10, but also beyond this range if necessary).

[0132] This determination of minimum relative packet transmission duration occurs for voice packets with no voice content and for voice packets with voice content, where the different lengths of different voice packets are preferably compensated for in accordance with the present invention.

[0133] After these preceding steps are performed, the measurement interval T i and T MIN =-MAX JITTER The counter "i" is the interval T i is returned in the reinitialization 313 of the

[0134] By using the above method and algorithm, the receiver time T SINK It is possible to determine the relative minimum packet delay for voice packets by: This can be done entirely in the receiver.

[0135] The various previously described steps can also be read based on the flowchart of the receiver clock initialization and the determination of the minimum relative packet transmission duration according to Figure 7. The individual steps of the above method steps can be understood based on an exemplary log file, an excerpt of which is shown in Figure 8. For clarity, the interval T i has 10 voice packets, but in practice it may be significantly larger (e.g., 100 or 200 voice packets) and may be fixed by a person skilled in the art in a suitable manner. The weighting factor applied as an example in the second part of the log file was 3. The "measurement interval" according to FIG. 8 is the previously explained interval T i is synonymous with.

[0136] Part B After initializing and adjusting the receiver time by determining the minimum packet delay between the sender (data source) and the receiver (data sink) after the connection between the sender and the receiver is set up by evaluating the received voice packets 110 and 120 with and without voice content according to FIG. 7 (preferably occurring continuously during the connection period 30), the voice packets 110 and 120 are permanently exchanged according to the ED-137 standard until the connection between the sender and the receiver is terminated.

[0137] According to the method illustrated in FIG. 9, whose initiation 400 also occurs with the setup of a session 401 and which runs in principle parallel to the method illustrated in FIG. 7 according to the invention, the buffer DJB proposed according to the invention is considered as a dynamic jitter buffer depending on the actual reception times of the voice packets.

[0138] Typically, during connection 30, voice packets 110 with voice content are transmitted whenever radio messages 1, 2, 3, and 4 are to be sent. Voice packets 120 without voice content are used for line monitoring of radio channels 10 and 20 during times when no active radio transmissions occur during connection 30 (see FIG. 6). A field PT (Payload Type) indicating the type of voice packet 110 or 120 is present in the header H of voice packets 110 with voice content (RTP packets) and voice packets 120 without voice content (R2S packets), as defined as a header extension in the ED-137 standard. A separate PTT Type field determines whether voice content is to be broadcast via the aviation radio transmitter. If this is set to PTT≠0, the receiver is informed that voice content will be transmitted with this voice packet 110. In contrast, if the PTT Type field is set to PTT=0, the receiver is informed that this voice packet 120 will be transmitted without voice content. Therefore, during normal operation, voice packets 120 that do not have voice content are PTT off (PTT=0) and the voice packet 110 with voice content is sent with PTTon (PTT ≠ 0) is transmitted. Therefore, PTT off From PTT on indicates the start of radio messages 1, 2, 3, and 4. Thus, PTT on From PTT off indicates the end of radio messages 1, 2, 3, and 4.

[0139] To detect such a change, according to FIG. 9, the PTT type field of two consecutive voice packets 110 and 120 is evaluated in each case, where the variable PTT old indicates which PTT type the previous voice packets 110 and 120 had, and the variable PTT indicates which PTT (in other words, the voice packet that follows the previous voice packet) the current voice packet 110 and 120 has.

[0140] Therefore, PTT off A voice packet 120 having on is received at the same time, this voice packet is the first voice packet 111 of the new radio message 1, 2, 3, 4. This voice packet 111 is then delayed by a dynamic buffer DJB depending on the relative time of reception according to the receiver time at the time of reception relative to the time information in the header H of the voice packet 111 for sending out the voice packet 111, until the receiver starts to play the voice content 1, 2, 3, 4 (or radio message).

[0141] After starting 400 and setting up a session 401, first set the variable PTT old Value of PTT old =PTT off In accordance with the present invention, the receiver initializes the value of the variable PTT to PTT upon receipt of each voice packet 110 and 120. off or PTT on The variable PTT is determined as follows (403). oldand querying (404) the PTT to determine whether the voice packet is the first voice packet 111 with voice content after at least one voice packet 120 with no voice content, e.g., query PTT=PPT on && PTT old =PPT off Otherwise, the variable PTT old The value of is set to the value of the variable PTT (step 410), and new voice packets 110 and 120 are expected to be received.

[0142] If this is indeed the first voice packet 111 with voice content after the preceding voice packet 120 without voice content, the reception time of this first voice packet 111 is determined by comparing the time information in the received voice packet 111 with the receiver time at the time of reception. Then the (dynamic) buffer DJB according to the invention is determined according to the actual reception time.

[0143] According to a preferred embodiment, this works, for example, by using the procedure described below.

[0144] For voice packet 111, receiver time T SINK Actual reception time T from RTP +RTP PL The difference deltaJB is determined (405), which is the expected time of receipt of the voice packet 110 having the voice content with the shortest jitter encountered so far: deltaJB=T RTP +RTP PL -T SINK Shows.

[0145] For voice packets 120 with no voice content, which should never occur in this method step, the duration of the payload is RTP PL =0.

[0146] This is in accordance with determining the relative packet transmission duration 305 by the difference delta according to FIG.

[0147] Then, a query 406 of the difference deltaJB occurs. If deltaJB is greater than zero (deltaJB>0; query 406 "j"), this means that the audio packet 111 has a shorter relative packet transmission time (i.e., lower jitter) than previously found during the adjustment of the receiver time according to Part A. Then, the playback of audio contents 1, 2, 3, 4 (radio messages) should be delayed by a predetermined maximum buffer length (Defaultjitterbuffersize; for example, 160) fixed in advance in a manner suitable for a person skilled in the art for the audio packet with the shortest possible jitter JIT (setting buffer DJB to maximum size MAX=Defaultjitterbuffersize 407): DJB=MAX The DJB in this case describes the buffer (ie the size or duration of the buffer).

[0148] If deltaJB is less than 0 (i.e., deltaJB<0; query 406 "n"), this means that the audio packet 111 has a longer relative packet transmission time (i.e., higher jitter) than previously found during the receiver time adjustment according to part A. Then, the playback of audio contents 1, 2, 3, 4 (radio messages) should not be delayed by the full length of the predetermined maximum buffer (Defaultjitterbuffersize=MAX). The dynamic buffer DJB according to the present invention is reduced by the deviation from the minimum delay (i.e., deltaJB: set buffer DJB to reduced size RED 408): DJB=Defaultjitterbuffersize-|deltaJB|, |deltaJB| is used herein to describe the (positive) value of the essentially negative variable deltaJB.

[0149] If the size of the buffer DJB is less than zero, then JB SIZE =0 is set (reset buffer DJB to zero 409).

[0150] Also, after this, the variable PTT oldThe value of =PPT is set to the value of the variable PTT (step 410) and new voice packets 110 and 120 are expected to be received.

[0151] The various previously described steps may also be read based on the flowchart for delaying the first voice packet 111 (sequence of voice packets 119) of radio messages 1, 2, 3, 4 by dynamically adjusting the buffer DJB according to Figure 9. The individual steps of the above method steps may be understood from an exemplary log file, an excerpt of which is shown in Figure 10. The third voice packet shown contains voice content (PTT on = 1). This packet is 7 time units too late (relatively speaking (deltaJB = delta = -7)), so the default maximum buffer (DefaultJitterBufferSize = 160) is reduced to 153 according to the calculation 160 - 7 = 153. [Explanation of symbols]

[0152] 1. Audio content, radio messages 2. Audio content, radio messages 3. Audio content, radio messages 4. Audio content, radio messages 10. The first transmission channel of an IP-based network 20 Secondary Transmission Channel of IP-Based Networks 30 connection periods 50 Transmitting Timeline 51 Timeline on Transmission Channel 90 Second Variant of Audio Output 91 First Variant of Audio Output 92 Interruption during audio output 110 voice packets with voice content on the first transmission channel 111 audio packet with first audio content 119 A sequence of voice packets 120 Audio packets with no audio content 210 voice packets with voice content on the second transmission channel 211 audio packet having first audio content 230 Packet Transmission Time or Duration 300 Start of algorithm to determine minimum packet delay 301 Session Setup 302 First voice packet received after session setup 303 Setting the receiver time 304 Receive voice packets 305 First Step: Determining Relative Packet Transmission Durations 306 Second step: Determining the voice packet with the minimum relative packet transmission duration 307 Checking the end of the interval 308 Third step: Adjusting the receiver time 309T MIN Inquire whether is greater than zero 310T MIN Inquire whether is less than zero 311 T MIN Weighting of 312 Receiver time correction 313 Reinitialize Interval 400 Start of algorithm for fixing dynamic buffers 401 Session Setup 402 Variable PTT old To initialize 403 Received voice packet 404 Variable PTT old and PTT inquiry 407 Set buffer DJB to maximum size 408 Set buffer DJB to a reduced size 409 Zeroing the buffer DJB H Voice packet (communication control) header PL payload (audio content as digital audio data) DEL Delay (delay) of voice packets on the transmission channel JIT jitter (random time fluctuation between the receipt times of voice packets) SJB Static Buffer (Jitter Buffer) (Conventional Technology) DJB Dynamic Buffer (Jitter Buffer) (present invention) TPS payload length (duration of sending payload) DDC additional delay compensation for various transmission channels TWD The period from sending the first voice packet to the start of processing PPT PTT field of the current voice packet PowerPoint OLD PPT of preceding audio packet T SINK Receiver time T SOURCE Transmitter Time Delta: The difference between the actual reception time and the receiver time. deltaJB The difference between the actual reception time and the receiver time ts The time when the voice packet was sent ta Actual sending time (sent time) te The time when the voice packet was received tw Start of audio output (output time or processing time)

Claims

1. 1. A method for compensating for jitter (called jitter (JIT)) of random time variations between the reception times (te) of audio packets (110, 120; 210) when receiving audio content (1, 2, 3, 4) over an IP-based network, the method comprising: a sequence of digital audio packets (110, 120; 210) with or without audio content (1, 2, 3, 4) received by a receiver for processing the audio content (1, 2, 3, 4) contained in said audio packets (110, 120; 210), the method comprising: A voice packet (110; 210) having a voice content has a section (called a header (H)) having data for communication control and a section (called a payload (PL)) having digital voice data from a part of the voice content (1, 2, 3, 4), An audio packet (120) without audio content has the header (H) without the payload; At least one time information of a transmitter time of a transmitter is included in the voice packet (110, 120; 210), the time information indicating a transmission time (ts) of the voice packet (110, 120; 210), and the sampling time of a first sample when sampling the voice content (1, 2, 3, 4) is adopted as the transmission time of the voice packet (110, 120, 210); The method, wherein the receiver waits to process the audio content (1, 2, 3, 4) for a waiting time (called a buffer (DJB)) after receiving a first audio packet (111) having the audio content (1, 2, 3, 4), After receiving a first voice packet (110, 120; 210) (called an initial packet) during a connection period (30) on a transmission channel (10, 20) of said voice packets (110, 120; 210), said receiver initializes said receiver clock by using a transmitter time according to said time information in said initial packet so that the receiver time of said receiver clock correlates with the transmitter time at the time of sending said voice packet (110, 120; 210); After the receiver clock is initialized, for each subsequent reception of a voice packet (110, 120; 210), the receiver: - comparing the receiver time to the transmitter time when the voice packet was sent with the time information (contained in the header (H) of the subsequent voice packet (110, 120; 210)) and determining a relative packet transmission duration (delta) with respect to the receiver time; temporarily storing a minimum relative packet transmission duration (delta) resulting from the minimum of said relative packet transmission durations (delta); determining a minimum relative packet transmission duration (delta) by adjusting the receiver time dependent on the temporarily stored minimum relative packet transmission duration (delta) to correlate with the transmitter time when sending the voice packet (110, 120; 210) having the minimum relative packet transmission duration; When receiving a first voice packet (111) having voice content after at least one voice packet (120) having no voice content, the receiver determines the actual reception time of this voice packet by comparing the time information in the received voice packet (111) with the receiver time at the time of reception (deltaJB) and determines the buffer (DJB) according to the actual reception time.

2. 2. The method of claim 1, wherein during the initialization of the receiver clock and during the determination of the relative packet transmission durations (delta), the difference in length between voice packets (110; 210) with voice content and voice packets (120) without voice content is taken into account.

3. 3. The method of claim 1 or 2, wherein the determination of the minimum relative packet transmission duration (delta) occurs at multiple intervals during the duration of the connection (30).

4. 4. The method of claim 1, wherein the adjustment of the receiver time in response to the temporarily stored minimum relative packet transmission duration (delta) occurs by advancing the receiver time by the minimum relative packet transmission duration (delta) to an earlier time if an indication is received that the voice packet (110, 120; 210) has a shorter packet transmission time.

5. 5. The method of claim 1, wherein the adjustment of the receiver time in response to the temporarily stored minimum relative packet transmission duration (delta) occurs by moving the receiver time back to a later time if the temporarily stored minimum relative packet transmission duration (delta) indicates that "only voice packets having a longer packet transmission duration have been received," wherein the receiver time is moved back by a specified or definable duration.

6. 6. The method of claim 1, wherein the determination of the minimum relative packet transmission duration (delta) occurs for voice packets (120) that do not have voice content and for voice packets (110; 210) that do have voice content.

7. 7. The method according to claim 1, wherein the buffer (DJB), starting from a maximum buffer, adjusts the maximum jitter to be considered based on a delay (deltaJB) for reception resulting from a comparison of the time information in the first audio packet (111) having audio content with the receiver time when the first audio packet (111) is received.

8. 8. The method according to any one of claims 1 to 7, characterized in that the audio packets (110; 210) having the same audio content are received by the receiver via a plurality of different transmission channels (10, 20), and the delay (DDC) correction of the audio packets (110, 210) having the same audio content via the different transmission channels is applied by the receiver.

9. 1. A receiver for receiving digital audio packets (110, 120; 210) over an IP-based network and for processing audio content (1, 2, 3, 4) contained within said audio packets (110, 120; 210), comprising: The voice packets (110, 120; 210) that are voice packets (110; 210) that have voice content have a section (called a header (H)) that has data for communication control and a section (called a payload (PL)) that has digital voice data from a part of the voice content (1, 2, 3, 4), or the voice packets (120) that do not have voice content have the header (H) that does not have the payload, and at least one time information of a transmitter time of a transmitter is included in the voice packets (110, 120; 210), and the time information indicates a transmission time (ts) of the voice packets (110, 120; 210), and the sampling time of a first sample when sampling the voice content (1, 2, 3, 4) is adopted as the transmission time of the voice packets (110, 120, 210), the receiver comprises a receiving unit connectable to the IP-based network, the receiving unit being set up to receive voice packets (110, 120; 210) transmitted over the IP-based network; and 1. A receiver having an arithmetic logic unit with a receiver clock, the arithmetic logic unit being set up to process said received voice packets (110, 120; 210), The receiver, wherein the arithmetic logic unit is set up to perform a method of jitter compensation when receiving and processing the audio content (1, 2, 3, 4) as follows: waiting for a waiting time (called a buffer (DJB)) after receiving a first audio packet (111) having audio content (1, 2, 3, 4) to process said audio content (1, 2, 3, 4); After receiving a first voice packet (110, 120; 210) (called an initial packet) during a connection period (30) on the transmission channel (10, 20) of said voice packets (110, 120; 210), initializing said receiver clock by using a transmitter time according to said time information in said initial packet so that the receiver time of said receiver clock correlates with the transmitter time at the time of sending said voice packet (110, 120; 210); After the receiver clock has been initialized, for each reception of a subsequent voice packet (110, 120; 210): - comparing the receiver time to the transmitter time when the voice packet was sent with the time information (contained in the header (H) of the subsequent voice packet (110, 120; 210)) and determining a relative packet transmission duration (delta) with respect to the receiver time; temporarily storing a minimum relative packet transmission duration (delta) resulting from the minimum of said relative packet transmission durations (delta); determining a minimum relative packet transmission duration (delta) by adjusting the receiver time dependent on the temporarily stored minimum relative packet transmission duration (delta) to correlate with the transmitter time when sending the voice packet (110, 120; 210) having the minimum relative packet transmission duration; When receiving a first voice packet (111) having voice content after at least one voice packet (120) having no voice content, the actual receiving time of this voice packet is determined by comparing the time information in the received voice packet (111) with the receiver time at the time of reception (deltaJB) and determining the buffer (DJB) according to the actual receiving time.

10. Receiver according to claim 9, characterized in that the receiver is set up to receive the voice packets (110; 210) having the same voice content via a plurality of different transmission channels (10, 20) and to apply the method of claim 8 in the receiving processing, wherein the arithmetic logic unit is set up to process the voice content (1, 2, 3, 4) from the received voice packets (110, 120; 210) having the same voice content from a plurality of the different transmission channels (10, 20).

11. A method for transmitting and receiving audio content, The transmitter converts the audio contents (1, 2, 3, 4) into a series of digital audio packets (110, 120; 210) with or without audio contents, and the audio packets (110; 210) with audio contents have a section (called a header (H)) with data for communication control and a section (called a payload (PL)) with audio data digitized from a part of the audio contents (1, 2, 3, 4), and the audio packets (120) without audio contents have the header (H) without the payload, and each audio packet (110, 120) 20; 210) (in particular the header (H) of the voice packet (110, 120; 210)) contains at least one time information of a transmitter time (indicating the transmission time (ts) of the voice packet (110, 120; 210)), the sampling time of a first sample when sampling the voice content (1, 2, 3, 4) being taken as the transmission time of the voice packet (110, 120, 210), and the transmitter sends out the sequence of the voice packets (110, 120; 210) via at least one transmission channel (10, 20) of an IP-based network, A method, wherein a receiver receives the sequence of digital audio packets (110, 120; 210) with or without audio content and processes the audio content (1, 2, 3, 4), The following method is applied in the receiver: waiting for a waiting time (called a buffer (DJB)) after receiving a first audio packet (111) having audio content (1, 2, 3, 4) to process said audio content (1, 2, 3, 4); After receiving a first voice packet (110, 120; 210) (called an initial packet) during a connection period (30) on the transmission channel (10, 20) of said voice packets (110, 120; 210), initializing said receiver clock by using a transmitter time according to said time information in said initial packet so that the receiver time of said receiver clock correlates with the transmitter time at the time of sending said voice packet (110, 120; 210); After the receiver clock has been initialized, for each reception of a subsequent voice packet (110, 120; 210): - comparing the receiver time to the transmitter time when the voice packet was sent with the time information (contained in the header (H) of the subsequent voice packet (110, 120; 210)) and determining a relative packet transmission duration (delta) with respect to the receiver time; temporarily storing a minimum relative packet transmission duration (delta) resulting from the minimum of said relative packet transmission durations (delta); determining a minimum relative packet transmission duration (delta) by adjusting the receiver time dependent on the temporarily stored minimum relative packet transmission duration (delta) to correlate with the transmitter time when sending the voice packet (110, 120; 210) having the minimum relative packet transmission duration; When receiving a first voice packet (111) having voice content after at least one voice packet (120) having no voice content, the actual receiving time of this voice packet is determined by comparing the time information in the received voice packet (111) with the receiver time at the time of reception (deltaJB) and determining the buffer (DJB) according to the actual receiving time.

12. 12. The method according to claim 11, characterized in that a permanent communication connection is set up in the IP-based network during the connection period (30) between the sender and the receiver via at least one communication channel (10, 20), During the connection period (30), when no voice content (1, 2, 3, 4) should be transmitted, voice packets (120) having no voice content are exchanged to maintain the connection between the sender and the receiver; At times during the connection period (30) when the audio content (1, 2, 3, 4) is to be transmitted, the sequence of digital audio packets (110; 210) comprising the audio content is transmitted from the transmitter to the receiver.

13. 13. The method according to claim 11 or 12, characterized in that the transmitter sends out voice packets (110, 120; 210) via a plurality of different transmission channels (10, 20), and the receiver receives the voice packets (110, 120; 210) sent via the different transmission channels (10, 20).

14. 14. The method according to any one of claims 11 to 13, characterized in that it is used in radio communications, in which a plurality of different transmission channels (10, 20) are used to transmit radio messages as audio content (1, 2, 3, 4), in particular for the transmission of operational ground-to-air communications in CLIMAX operations.

15. 1. A device for transmitting and receiving digital voice packets (110, 120; 210) over an IP-based network, the voice packets (110; 210) having voice content (1, 2, 3, 4) contained therein, the device comprising a transmitter and a receiver, the transmitter comprises an audio recording unit for recording the audio content (1, 2, 3, 4) and an arithmetic logic unit having a transmitter clock, the arithmetic logic unit being set up to process the audio content (1, 2, 3, 4) according to the method of any one of claims 11 to 14, A device, wherein the receiver is constructed according to the features of claim 9 or 10.

Citation Information

Patent Citations

  • Satellite broadcast transmitting / receiving equipment and digital voice transmission system by same transmitting equipment

    JP1991006150A

  • Device and method for absorbing delay jitter caused in data transmission

    JP2001352316A

  • Estimation of playback delay

    JP2011505743A