Processing device, receiving device, transmitting device, transmitting / receiving device, and program related to smpte st2110
The processing device leverages RTP timestamp values and NTP synchronization to derive PTP seconds, addressing the cost barrier of SMPTE ST2110-compatible devices by enabling their use with non-PTP compatible equipment.
Patent Information
- Application Number
- JP2024073266
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-04-27
- Publication Date
- 2025-11-07
- Estimated Expiration
- 2044-04-27
AI Technical Summary
Existing SMPTE ST2110-compatible devices require expensive PTP-compatible products and precise time management, creating a barrier to their development and operation.
A processing device that uses the RTP timestamp value as a reference, combined with NTP synchronization, to derive PTP seconds without requiring PTP synchronization, allowing the use of non-PTP compatible devices and reducing system costs.
Enables cost-effective SMPTE ST2110-compliant transmission/reception without the need for PTP synchronization, facilitating the use of non-PTP compatible switches and devices.
Smart Images

Figure 2025168122000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to a processing device, a receiving device, a transmitting device, a transmitting / receiving device, and a program in accordance with the SMPTE ST2110 standard. [Background technology]
[0002] SMPTE ST2110 transmits video, audio, and auxiliary data in separate streams over an IP network. The video, audio, and ancillary data are all in separate RTP (Real-time Transport Protocol) streams, and each can be handled independently. Because packet transmission over the network is asynchronous, each node must lock to highly accurate time information called PTP (IEEE-1588 Precision Time Protocol) to synchronize the separately transmitted RTP streams.
[0003] The transmitting device stamps the time information obtained from an internal clock synchronized with PTP into the timestamp field of each RTP packet.
[0004] In the RTP standard, timestamp values are 32 bits, so they overflow and wrap around after a certain period of time. As a result, a receiving device cannot normally determine the time at which a packet was sent just by looking at the timestamp value stamped on an RTP packet. For this reason, the receiving device also synchronizes with PTP, deriving the true time stamped on the RTP packet by comparing it with its own internal clock synchronized with PTP, and then using this time information to synchronize the video, audio, and auxiliary data. [Prior art documents] [Patent documents]
[0005] [Patent Document 1] Patent Publication No. 2018-42020 [Patent Document 2] Patent Publication No. 2021-077940 [Non-patent literature]
[0006] [Non-Patent Document 1] SMPTE ST2110-10 [Non-patent document 2] SMPTE ST2110-20 [Non-patent document 3] SMPTE ST2110-30 [Non-patent document 4] SMPTE ST2110-31 [Non-Patent Document 5] SMPTE ST2110-40 [Non-patent document 6] SMPTE ST2059-1 [Non-Patent Document 7] SMPTE ST2059-1 Summary of the Invention [Problem to be solved by the invention]
[0007] While the commonly used NTP (Network Time Protocol) has an accuracy of milliseconds, PTP can synchronize the time between devices with an accuracy of less than one microsecond. Therefore, to use PTP, the PTP generator, known as the PTP GM (Grandmaster), as well as all related network switches, devices, and computers must be PTP-compatible devices. However, PTP-compatible products are generally expensive and require ultra-precise time management, which creates a barrier to the development and operation of SMPTE ST2110-compatible devices.
[0008] Devices and computers using the present invention can receive and transmit packets of the SMPTE ST2110 standard without being synchronized to PTP, and can use switches, devices, and computers that do not support PTP. Therefore, it is possible to reduce the cost of the entire system, and the purpose is to be able to more easily provide SMPTE ST2110-compliant transmission / reception devices and transmission / reception servers.
Means for Solving the Problems
[0009] To solve the above problems, the processing device according to claim 1 includes a receiving unit for RTP packets in a packet format defined by the ST2110 standard, and is characterized by using the RTP timestamp value included in the received RTP packet as a reference.
[0010] Further, in addition to the features of the processing device according to claim 1, the processing device according to claim 2 receives video or auxiliary data packets and audio packets in the RTP packet receiving unit, using the combination of the RTP timestamp values included in each received RTP packet and the property that each RTP timestamp value increases at regular intervals, it is characterized by deriving the elapsed seconds p seconds (0≦p<N) within a period with a period of N seconds as an element of the PTP seconds marked on the received RTP packet, The N seconds is equal to the period in which all phases of the period in which each RTP timestamp value increases at regular intervals and the period in which each RTP timestamp value wraps around are aligned.
[0011] ; Further, in addition to the features of the processing device according to claim 2, the processing device according to claim 3 assumes that the value a is an arbitrary non-negative integer, and is characterized by deriving the PTP seconds marked on the received RTP packet as a×N + p.
[0012] In addition to the features of the processing device of claim 2, the processing device of claim 4 is characterized in that it derives the value a from the PTP seconds stamped in the received RTP packet using a value in the range up to N seconds before, and derives the PTP seconds stamped in the received RTP packet as a×N+p.
[0013] Furthermore, the processing device of claim 5 is characterized in that, in addition to the features of the processing device of claim 1, it is equipped with a time synchronization unit using NTP (Network Time Protocol) in an environment where the clock of the PTP GM (PTP Grandmaster) is operating in real time, and derives the number of PTP seconds stamped in the received RTP packet and the cumulative leap second value DTAI using the RTP timestamp value included in the received RTP packet and the current time synchronized by the NTP.
[0014] In addition, the processing device of claim 6 is characterized in that, in addition to the characteristics of the processing device of any one of claims 3 to 5, it outputs the value obtained by adding an adjustment value α to the PTP second number stamped in the last received RTP packet as the PTP second number to subsequent processing.
[0015] In addition to the features of the processing device of claim 6, the receiving device of claim 7 has a separate receiving unit for RTP packets in a packet format defined by the ST2110 standard, and processes the RTP packets received by the separate receiving unit.
[0016] In addition, the transmitting device of claim 8 is a transmitting device that has the characteristics of the processing device of claim 6 and is equipped with a transmitting unit for RTP packets in a packet format defined in the ST2110 standard, and transmits the RTP packets from the transmitting unit.
[0017] Furthermore, the transmitting / receiving device of claim 9 combines the features of the receiving device of claim 7 and the features of the transmitting device of claim 8, and receives at least one of video, audio, and auxiliary data, processes it, and transmits it as a separate RTP packet.
[0018] Furthermore, the transmission / reception device of claim 10 is a transmission / reception device that, in addition to the features of the processing device of claim 1, is equipped with a packet rewriting unit and a transmission unit for RTP packets in a packet format defined by the ST2110 standard, and transmits another RTP packet from the transmission unit by reusing information including the RTP timestamp value of the RTP packet received by the reception unit.
[0019] Furthermore, a general-purpose computer may be configured to have the functions of the above-mentioned device. [Effects of the Invention]
[0020] According to the present invention, it is possible to receive and transmit packets conforming to the SMPTE ST2110 standard without synchronizing with PTP, and it is possible to use non-PTP compatible switches and non-PTP compatible devices and computers, which not only reduces the cost of the entire system but also makes it easier to provide SMPTE ST2110 compliant transmitting and receiving devices and transmitting and receiving servers. [Brief explanation of the drawings]
[0021] [Figure 1] 1 is a diagram illustrating an overall configuration of a system to which some embodiments of the present invention can be applied. [Figure 2] 1 is a block diagram showing the functional configuration of a PTP seconds number derivation device according to a first embodiment of the present invention. [Figure 3] 10 is a flowchart showing the processing of the number of seconds in a period deriving unit according to the first embodiment. [Figure 4] 4 is a flowchart showing details of the video timestamp normalization process of FIG. 3. [Figure 5] 4 is a flowchart showing the details of the timestamp combination calculation process of FIG. 3. [Figure 6] FIG. 10 is a block diagram showing the functional configuration of a PTP seconds number derivation device according to a second embodiment of the present invention. [Figure 7] FIG. 10 is a block diagram showing a functional configuration of a receiving device according to a third embodiment of the present invention. [Figure 8]FIG. 10 is a block diagram showing a functional configuration of a transmission device according to a fourth embodiment of the present invention. [Figure 9] FIG. 11 is a block diagram showing a functional configuration of a transmission device according to a fifth embodiment of the present invention. [Figure 10] 10 is an example of a packet structure for video RTP packets in a packet format defined by the SMPTE ST2110 standard. [Figure 11] This is the packet structure for audio in RTP packets in the packet format defined by the SMPTE ST2110 standard. [Figure 12] 10 is an example of a packet structure for ancillary data of an RTP packet in a packet format defined in the SMPTE ST2110 standard. DETAILED DESCRIPTION OF THE INVENTION
[0022] Hereinafter, embodiments of the present invention will be described in detail with reference to the drawings. Note that the same elements are given the same reference numerals, and duplicate explanations will be omitted. Furthermore, the following embodiments are merely examples for explaining the present invention, and are not intended to limit the present invention to only these embodiments. Furthermore, the present invention can be modified in various ways without departing from the gist of the invention.
[0023] The standards that are the basis for understanding the description of this specification are briefly described below.
[0024] The SMPTE ST2059 standard specifies that the SMPTE Epoch begins at 00:00:00 TAI (International Electronic Time) on January 1, 1970, and that the SMPTE Epoch is the same as the PTP Epoch. In other words, the number of PTP seconds at the time of the SMPTE Epoch is 0. Also, an alignment point is defined that serves as the starting time of each frame, starting from the SMPTE Epoch.
[0025] In the RFC3551 standard that defines the RTP standard, the RTP timestamp value represents the time when the oldest data contained in a packet was generated. Based on this, the RTP timestamp value in the SMPTE ST2110 standard represents the time of each frame at the alignment point for video and ancillary data; in the case of interlaced data, the first field represents the frame time, and the second field represents the time exactly halfway between the frame and the next frame. For audio, it represents the time of the first sample. RTP packets for video and ancillary data have the same timestamp value within a frame in the case of progressive data, and within a field in the case of interlaced data.
[0026] Similarly, the SMPTE ST2110 standard specifies that the random offset in the RTP timestamp value specified in the RFC3550 standard is 0. This means that multiple SMPTE ST2110 streams locked to the same PTP have the same timestamp value at the same time, allowing them to be synchronized with each other. However, because the clock frequencies for video and ancillary data are different from those for audio, the RTP timestamp values will be different even if they are the same time.
[0027] The RTP timestamp value is 32 bits by standard, and since it overflows and wraps around after a certain period of time, it is usually not possible to synchronize clocks with different frequencies just by looking at the RTP timestamp value. Therefore, each node that sends and receives SMPTE ST2110 streams must lock to the common clock, PTP, with high precision.
[0028] Although the SMPTE ST2110 standard does not describe the logic for achieving synchronization, the following is one example of a commonly conceivable method.
[0029] The sender obtains the current PTP timestamp, calculates the PTP timestamp for the alignment point specified in SMPTE ST2059-1 in the immediate future, synchronizes the video frame and audio to that alignment point, and in the case of video, stamps the lowest 32 bits of the value (rounded down to the nearest integer) obtained by multiplying the alignment point timestamp by the clock frequency. In the case of audio, assuming that the RTP timestamp value has been increasing regularly since the SMPTE Epoch in accordance with the number of samples contained in the packet, stamps the lowest 32 bits of the value (rounded down to the nearest integer) obtained by multiplying the alignment point timestamp by the clock frequency in accordance with the number of samples contained in the packet, at a timing that is a multiple of the number of samples contained in the packet.
[0030] The receiving side calculates the PTP seconds of the alignment point immediately before the current PTP seconds, and compares the lower 32 bits of the value obtained by multiplying that value by the clock frequency (truncating any decimal points) with the received RTP timestamp value to calculate the elapsed time and align synchronization.
[0031] The SMPTE ST2110 standard specifies a clock frequency of 90 kHz for the RTP timestamps of video and ancillary data. Therefore, for a progressive frame rate of 30 / 1.001 fps (frames per second) (hereafter referred to as "29.97 fps"), the RTP timestamp values for video and ancillary data increase in increments of 90,000 / (30 / 1.001) = 3,003 (0 → 3,003 → 6,006 → 9,009 → ...). For an interlaced frame rate of 29.97 fps, the timestamp for the second field is not an integer, so the fractional part is rounded down (0 → 1,501 → 3,003 → 4,504 → 6,006 → ...).
[0032] Similarly, the SMPTE ST2110 standard specifies that the clock frequency of the audio RTP timestamp must match the audio sampling frequency. The standard audio sampling frequency is 48 kHz, but 96 kHz and 44.1 kHz can also be used. Audio RTP packets always transmit the same number of samples at regular intervals, and the RTP timestamp value also increases simply by the number of samples. For a sampling frequency of 48 kHz, the increase is usually in increments of 6 at 0.125 ms intervals, or 48 at 1 ms intervals.
[0033] The period for the RTP timestamp value to wrap around is 2^32 / 90,000 = approximately 13.3 hours for video and auxiliary data, and 2^32 / 48,000 = approximately 24.9 hours for audio at a sampling frequency of 48 kHz.
[0034] In this specification, the period in which the RTP timestamp value wraps around and the number of times it wraps around are referred to as "cycles," with the video or auxiliary data cycle being referred to as the "video cycle" and the audio cycle being referred to as the "audio cycle." Furthermore, in the calculation formula, the video cycle is referred to as "Cv" and the audio cycle as "Ca." Cv and Ca are equal to the quotient (integer) obtained by multiplying the number of PTP seconds by the clock frequency (decimals truncated), divided by 2^32.
[0035] Other basic variables in the calculation formula are defined as follows: f: sampling frequency Rv: A number that periodically increases the video timestamp. Ra: Number of audio timestamps that increase periodically at regular intervals
[0036] The SMPTE ST2110 standard specifies information such as the frame rate and sampling frequency of the transmitting side and the time on which PTP operates in the Session Description Protocol (SDP) text, but does not specify how this information is transmitted to the receiving side. Common methods include using the Advanced Media Workflow Association (AMWA) Networked Media Open Specification (NMOS) IS-05 protocol, operating with fixed values, or manually entering information. An apparatus according to one embodiment of the present invention performs processing under the assumption that it knows the SDP information of the RTP stream it receives. Furthermore, a transmitting apparatus according to one embodiment of the present invention is assumed to have a means for generating SDP and transmitting it to the receiving side as a common means.
[0037] FIG. 1 is a diagram showing the overall configuration of a system to which some embodiments of the present invention can be applied. A PTP GM (Grandmaster) 2, a reference ST2110 transmitter 3, and a PTP-incompatible SW 4 are connected via a PTP (Precision Time Protocol)-compatible SW 1. In addition, if required in some embodiments of the present invention, an ST2110 transmitter 6 or an ST2110 receiver 7 is connected to the PTP-compatible SW 1. The PTP-enabled SW 1 is connected to any one of the devices 10, 20, 30, 40, or 50 of an embodiment of the present invention via the non-PTP-enabled SW 4. If required in some embodiments of the present invention, an NTP server 5 is connected to the device via another network.
[0038] The reference ST2110 transmitter 3 may be a general SMPTE ST2110 compatible IP gateway or generator. The device of the embodiment of the present invention may be directly connected to the PTP-compatible SW1 without using the PTP-incompatible SW4. The NTP server 5 may be connected to the PTP-compatible SW1 or the PTP-incompatible SW4, rather than via a separate network. Each device according to the embodiment of the present invention may be configured as a dedicated device, or a general-purpose computer may be configured to have the functions of the device according to the embodiment of the present invention.
[0039] 2 is a block diagram showing the functional configuration of a PTP seconds derivation device 10 according to a first embodiment of the present invention. The PTP seconds derivation device 10 includes an RTP packet receiving unit 11, an in-period seconds derivation unit 12, a PTP seconds derivation unit 13, and a PTP seconds adjustment unit 15.
[0040] The RTP packet receiving unit 11 receives an RTP packet in a packet format defined by the SMPTE ST2110 standard, transmitted from the reference transmitting device 3, and performs processing to extract the RTP timestamp value contained in the packet.
[0041] Video RTP packets and auxiliary data RTP packets in the same frame or field have the same timestamp value, so it doesn't matter which one you receive. Because video data is very large, receiving auxiliary data reduces processing time.
[0042] Hereinafter, in this specification, the RTP timestamp value of video or auxiliary data will be referred to as "video timestamp" and "Tv", and the RTP timestamp value of audio will be referred to as "audio timestamp" and "Ta".
[0043] The RTP packet receiving unit 11 outputs the extracted video timestamp or audio timestamp to subsequent processing.
[0044] In the case of the PTP second number derivation device 10, the RTP packet receiving unit 11 receives video or auxiliary data, and audio.
[0045] FIG. 3 is a flowchart showing the processing of the in-period second number deriving unit 12. The RTP timestamp value received from the RTP packet receiving unit 11 is input in step S150. In step S150, the type of timestamp is confirmed, and if it is a video timestamp, the process proceeds to step S200, and if it is an audio timestamp, the process proceeds to both step S300 and step S400.
[0046] FIG. 4 is a flowchart showing the process of step S200. In step S200, the video timestamp is normalized. In the first embodiment, the RTP timestamp value increases monotonically, but depending on the frame rate, the value of PTP seconds x 90 kHz at the alignment point, which is the basis of the video timestamp for each frame, may not be an integer. Therefore, only video timestamps for which this value is an integer are used, and the others are ignored during processing.
[0047] Specific processing steps will be described.
[0048] In step S200, the video timestamp Tv from step S150 is input (step S201). It is determined whether the input video timestamp Tv has changed from the previous value Told (step S202). A change in value indicates a boundary between video frames or video fields. If the value has not changed (step S202: No), the process returns to step S201. If the value has changed (step S202: Yes), the process proceeds to step S203.
[0049] Next, in step S203, the array D that records the difference from the previous video timestamp value is shifted (D(n)←D(n-1)), and the video timestamp difference, Tv-Told, is substituted into D(0). This calculation is handled as int32 so that the correct difference value is calculated even if Tv or Told wraps around.
[0050] In step S204, Tv is substituted into Told and is held as the previous video timestamp.
[0051] Table 1 shows the increment pattern of the video timestamp for each frame rate and the value Rv at which the video timestamp increases regularly at regular intervals. As mentioned above, the RTP timestamp value in ST2110 is a standard that truncates if the value is not an integer, so if the increment pattern for each frame rate matches, it can be determined that the timestamp is an integer.
[0052] In step S205, it is determined whether array D matches the video timestamp increment pattern in Table 1. The determination is made by checking whether the values of D(0), D(1), ... match for all the video stamp increment patterns in Table 1.
[0053] If the sequences do not match (step S205: No), the process returns to step S201. If the sequences match (step S205: Yes), the process proceeds to step S206, where Tv is output as the normalized video timestamp Tvn to both step S250 and step S400.
[0054] [Table 1]
[0055] The above is the processing of S200.
[0056] Table 2 is a table summarizing the congruence equations for deriving the remainder Cv(mod V) when Cv is divided by V for each frame rate. Hereinafter, in this specification, the remainder when X is divided by Y will be expressed as "X(mod Y)".
[0057] [Table 2]
[0058] In step S250, based on Tvn input from step S200, it is determined which cycle of the V video cycle the video cycle Cv is, using Table 2. w=Cv(mod V) is set and output to step S500 together with Tvn.
[0059] The calculation theory of step S250 will be explained using specific values as an example for 29.97 fps.
[0060] At 29.97 fps, the video timestamp increases by 3,003 for each frame. Because 2^32 is not divisible by 3,003, when the video timestamp wraps around, it does not become exactly 0, but rather a fractional part is generated. By utilizing this property, since the greatest common divisor of 2^32 and 3,003 is 1, it is possible to determine which of the 3,003 video cycles it is, i.e., Cv(mod 3,003), from the video timestamp. This means that the video timestamp starts from 0 and returns to 0 only after the 3,003 cycle.
[0061] This can be calculated using the remainder theorem as follows:
[0062] Starting from PTP seconds = 0, the next video timestamp after one video cycle has elapsed is calculated as 3,003-(2^32 mod 3,003)=1,382.
[0063] From this, starting from PTP seconds = 0, the next video timestamp after the Cv video cycle has elapsed is (1,382 × Cv) mod 3,003.
[0064] This can be expressed using the congruence formula as follows:
[0065] [Number 1] 1,382×Cv ≡ Tvn(mod 3,003)
[0066] When you divide a multiple of 3,003 by 3,003, the remainder will always be 0, so if you use an integer m,
[0067] [Number 2] 3,003×m×Cv ≡ 0(mod 3,003)
[0068] Next, find the combination of integers j and k such that 1,382×j-1=3,003×k. The above formula holds when j=2,714 and k=1,249 (1,382x2,714-1=3,003x1,249=3,750,747). Therefore, multiplying both sides of equation 1 by 2,714 gives
[0069] [Number 3] 2,714×1,382×Cv ≡ 3,750,748×Cv ≡ 2,714×Tvn (mod 3,003)
[0070] Substituting 1,249 for m in formula 2, we get
[0071] [Number 4] 3003×1,249×Cv ≡ 3,750,747×Cv ≡ 0(mod 3,003)
[0072] Subtracting formula 4 from formula 3 gives
[0073] [Number 5] Cv ≡ 2,714×Tvn(mod 3,003)
[0074] The SMPTE ST2110-20 standard does not specify the frame rate to be used, but Table 2 above summarizes the formulas for calculating the congruence of Cv for the frame rates used in SDI (Serial Digital Interface) signals specified in the SMPTE ST352 and SMPTE ST2082-10 standards.
[0075] Even if other frame rates are used in SMPTE ST2110 in the future, Cv can be expressed by a congruence formula using this theory.
[0076] Table 3 is a table summarizing congruence formulas for deriving Ca(mod A) when the audio timestamp has a value Ra that increases at regular intervals.
[0077] [Table 3]
[0078] In step S300, based on Ta input from step S150, it is determined which cycle of the A audio cycles the audio cycle Ca is, using Table 3. y=Ca(mod A) is set and output together with Ta to step S500.
[0079] The calculation theory of step S300 will be explained using specific values as an example where Ra=6 (the number of voice samples per packet is 6).
[0080] Because 2^32 is not divisible by 6, when it wraps around it does not return to 0, but instead produces a fraction. By utilizing this property, since the greatest common divisor of 2^32 and 6 is 2, it is possible to determine which of the three audio cycles it is, i.e., Ca(mod 3), from the audio timestamp. This means that the audio timestamp starts from 0 and returns to 0 only in the third cycle.
[0081] This can be calculated using the remainder theorem as follows:
[0082] Starting from PTP seconds = 0, the next audio timestamp after one audio cycle has elapsed is calculated as 6-(2^32 mod 6)=2.
[0083] Starting from PTP seconds = 0, the next audio timestamp after the Ca audio cycle has elapsed is (2 × Ca) mod 6.
[0084] From the above, it can be expressed as follows using the congruence formula:
[0085] [Number 6] 2×Ca ≡ Ta(mod 6)
[0086] The modulus of the equation 6, 6, can be factorized into 2 and 3, so it can be expressed as follows:
[0087] [Number 7] 2 × Ca ≡ Ta(mod 3)
[0088] On the other hand, the remainder of any multiple of 3 is 0, so equation 8 holds true.
[0089] [Number 8] 3×Ca ≡ 0(mod 3)
[0090] Subtracting formula 7 from formula 8 gives
[0091] [Number 9] Ca ≡ -Ta(mod 3)
[0092] Table 3 above summarizes the congruence formulas for Ca for audio timestamp increments (number of samples per packet) defined in the SMPTE ST2110-30 and SMPTE ST2110-31 standards.
[0093] Even if in the future SMPTE ST2110 uses a number of audio samples per packet that is not included in the current standard, Ca can be expressed by a congruence formula using the same theory.
[0094] If the number of audio samples in one packet is a power of 2, such as 4 or 8, it will always be exactly 0 when wrapped around, leaving no remainder, so A = 1 and it is not possible to derive the period of the audio cycle. However, step S300 is merely one element for deriving the number of PTP seconds, and does not affect the overall theory of the present invention.
[0095] FIG. 5 is a flowchart showing the process of step S400. In this step, since the video timestamp and audio timestamp have different RTP packet reception timings, the most recently input Tvn and Ta are used for processing.
[0096] In step S401, the audio timestamp value Ta' at the start of the current video cycle is calculated based on Tvn and Ta using the following formula: Ta'=Tvn+Ta' / Ta'.
[0097] [Number 10] Ta' = int(Ta-Tvn / 90,000×f)mod 2^32
[0098] Cv(mod Mv) is derived (set as r) by adopting the video cycle value Ta' that is closest to the "audio timestamp at the start of the video cycle" in Table 4 (Table 4a, Table 4b, Table 4c), and output to step S403. "Close" here means that the value within the range of ±ε×f (for example, ε = 250 milliseconds, where ε is a deviation range that is well within any frame rate) matches Ta' with respect to the audio timestamp at the start of the video cycle in Table 4. In this case, calculations are made assuming that 0 and 2^32-1 are consecutive.
[0099] [Table 4a]
[0100] [Table 4b]
[0101] [Table 4c]
[0102] The calculation theory of step S401 will be specifically explained for the case of a sampling frequency of 48 kHz.
[0103] The video timestamp has a clock frequency of 90 kHz, and the audio timestamp has a clock frequency of 48 kHz, the same as the sampling frequency, so the relationship is 90,000:48,000=15:8. In other words, starting from PTP seconds = 0, the video timestamp and audio timestamp wrap around at the same time every 15 video cycles and 8 audio cycles, and the phase within that cycle is also in a constant relationship. Using this property, the value of the audio timestamp at the start of the video cycle (mod 15) is int(video cycles x 2^32 / 90,000 x 48,000) mod 2^32 This can be calculated and the above-mentioned Table 4a can be created.
[0104] Table 4 is a table of sampling frequencies currently specified in SMPTE ST2110, but the same theory can be used to derive other sampling frequencies in the future.
[0105] In step S402, the calculation is performed inversely to that in step S401. Based on Tvn and Ta, the video timestamp value Tv' at the start of the current audio cycle is calculated using the following formula.
[0106] [Number 11] Tv' = int(Tvn-Ta / f×90,000)mod 2^32
[0107] Ca (mod Ma) is derived (set as z) by adopting the audio cycle value Tv' that is closest to the "video timestamp at the start of the audio cycle" in Table 5 (Table 5a, Table 5b, Table 5c), and output to step S403. "Close" here means that a value within the range of ±ε × 90,000 (for example, ε = 250 milliseconds, which is a deviation range that is well within any frame rate) matches Tv' with respect to the video timestamp at the start of the audio cycle in Table 5. In this case, 0 and 2^32-1 are assumed to be consecutive in the calculation.
[0108] [Table 5a]
[0109] [Table 5b]
[0110] [Table 5c]
[0111] The calculation theory of step S402 is the same as step S401, except that the relationship between video and audio is reversed.
[0112] For a sampling frequency of 48 kHz, the video timestamp value at the start of the audio cycle (mod 8) is int(audiocycles x 2^32 / 48,000 x 90,000) mod 2^32 This can be calculated and the above-mentioned Table 5a can be created.
[0113] In step S403, a cycle difference S between r=Cv(mod Mv) derived in step S401 and z=Ca(mod Ma) derived in step S402 is calculated.
[0114] From the relationship between the clock frequencies of the video timestamp and the audio timestamp, 90,000:f = Mv:Ma, it is possible to calculate the "audio cycle (mod Ma)" by inverse calculation from the values of r and Tv. However, since there is a deviation in the timing of receiving RTP packets, the "audio cycle (mod Ma)" obtained by inverse calculation may deviate from Ca (mod Ma) derived in step S402. For subsequent calculations, the difference S is calculated. S is usually 0, but it is a value that becomes 1 when Ca is ahead and -1 when Ca is behind with respect to the audio cycle obtained by inverse calculation from Cv and Tv.
[0115] S can be obtained by the following formula.
[0116] [Equation 12] S = (z - int((r × 2^32 + Tv) / 90,000 × f / 2^32)) mod Ma However, when S = Ma - 1, let S = -1.
[0117] Output the derived z = Ca (mod Ma) and S to step S500.
[0118] In step S500, Tvn and w (= Cv mod V) input from step S250, Ta and y (= Ca mod A) input from step S300, z (= Ca mod Ma) and S input from step S400, Based on these, as an element of the PTP seconds number marked on each RTP packet, the elapsed seconds p seconds (0 ≦ p < N) within a period with a period of N seconds are derived and output to the PTP seconds number derivation unit 13.
[0119] In this step, since the receiving timings of the video timestamp and the audio timestamp in the RTP packet are different, w, y, z, and S derived based on them also have different generation timings. Therefore, the processing is performed using the most recently input Tvn, Ta, w, y, z, and S to step S500 respectively.
[0120] Since it is difficult to express the calculation method of step S500 as a general solution, we will specifically explain the case where the frame rate is 29.97 fps, f=48,000 (sampling frequency 48 kHz), and Ra=6 (number of audio samples per packet is 6).
[0121] Under the above conditions, V=3,003, A=3, and Ma=8.
[0122] According to the Chinese Remainder Theorem, 3 and 8 are mutually prime, so Ca(mod 24) modulo the least common multiple 24 can be found using the following formula:
[0123] [Number 13] Ca ≡ 9×z-8×y(mod 24)
[0124] Using Ca(mod 24) (= g) obtained from Equation 13, Cv(mod 45) is calculated using the following equation based on the relationship between the clock frequencies of the video timestamp and audio timestamp (90 kHz:48 kHz). The cycle difference S is also taken into account in the calculation.
[0125] [Number 14] Cv ≡ int((g×2^32+Ta) / 48,000×90,000 / 2^32)-S(mod 45)
[0126] Using the Chinese Remainder Theorem and its related theorems, we can use Cv(mod 45) (=h) obtained from Equation 14 to find Cv(mod 45,045) modulo 45,045, the least common multiple of 45 and 3,003, using the following equation.
[0127] [Number 15] Cv ≡ 4,005×w-4,004×h(mod 45,045)
[0128] From Cv(mod 45,045) (=i) obtained by Equation 15, Ca(mod 24,024) can be calculated using the following equation. [Number 16] Ca ≡ int((i × 2^32+Tvn) / 90,000×48,000 / 2^32)+S(mod 24,024)
[0129] 45,045 video cycles and 24,024 audio cycles are the same length of time, and are the time when the video timestamp values and audio timestamp values are perfectly phase-aligned. This is also synonymous with the time period N when the periods in which the video timestamps and audio timestamps increase at regular intervals and the periods in which the video timestamps and audio timestamps wrap around are all phase-aligned.
[0130] As described above, the period N is uniquely determined by the combination of frame rate, sampling frequency, and the number of samples included in the audio RTP packet. Let Nv be the number of video cycles corresponding to period N, and Na be the number of audio cycles corresponding to period N. Table 6 shows the formulas for calculating g, h, and i for common combinations of frame rate, sampling frequency, and audio RTP packet transmission interval, as well as a table of Nv, Na, and N. Note that the 44.1 kHz sampling frequency case is not widely expected, so it is omitted and only the case with a frame rate of 29.97 fps and Ra = 6 is listed. The case with Ra = 4 is also not widely expected, so it is omitted and only the case with frame rates of 29.97 fps and 25 fps is listed. Patterns not listed in this table can also be calculated in the same way using the above calculation theory.
[0131] [Table 6]
[0132] Note that period N is the period in which the phases of the periods in which the video timestamps and audio timestamps increase at regular intervals and the periods in which the video timestamps and audio timestamps wrap around all align exactly, so it can be expressed as a general solution using the times of each period. The period in which the phases of multiple periods align exactly is obtained by expressing each period value as a fraction and then dividing the least common multiple of the numerators by the greatest common divisor of the denominators. Using this calculation method, period N can be expressed by the following formula, where lcd is the least common multiple and gcd is the greatest common divisor.
[0133]
number
[0134] Using the values derived using the above method, the number of elapsed seconds Pv within a period of N seconds, which is the PTP number of seconds stamped into the RTP packets of the received video or auxiliary data, can be calculated using the following formula.
[0135] [Number 18] Pv = ((Cv mod Nv)×2^32+Tv) / 90,000)
[0136] The number of seconds Pa elapsed within a period of N seconds, which is the number of PTP seconds stamped on the received audio RTP packet, is calculated using the following formula.
[0137] [Number 19] Pa = ((Ca mod Na)×2^32+Ta) / f)
[0138] Preferably, the derived Pa and N are output to the PTP seconds derivation unit 13. The reason for outputting Pa rather than Pv is that Pv is the PTP seconds in units of frames or fields, and there is a gap due to blanking time, and the granularity is coarse, whereas audio RTP packets arrive regularly at fine granularity (125 microsecond intervals, 1 millisecond intervals, etc.).
[0139] The PTP seconds derivation unit 13 uses the input Pa and N, sets the value a to a non-negative integer, calculates the PTP seconds t stamped on each RTP packet using the following formula, and outputs the result to the PTP seconds adjustment unit 15. The method for deriving the value a will be described later.
[0140] [Number 20] t = a×N+Pa
[0141] In one embodiment of the present invention, for the purpose of synchronizing received RTP packets, the period N described above is the time when the video timestamp and audio timestamp are perfectly synchronized, and therefore a in Equation 20 can be set to any value, even if it is not an exact PTP number of seconds. For example, it is possible to assume a=0 and use Pa as the number of PTP seconds for subsequent processing.
[0142] In one embodiment of the present invention, the value a in equation 20 can be derived by providing a value τ (tN<τ≦t) ranging from the PTP second number t stamped in the received RTP packet to N seconds ago. Specifically, if the quotient of τ divided by N is c and the remainder is q, then if q > Pa, then a = c + 1, and if q <= Pa, then a = c. This gives the value a, allowing us to derive the exact number of PTP seconds stamped on the received RTP packet.
[0143] According to Table 6, for the 29.97 fps and 59.94 fps used in Japan, the period N is approximately 68 years, and since the PTP Epoch was January 1, 1970, as of 2024, the period N has not yet gone through one cycle. In other words, in an environment where the PTP GM clock within the system operates on real time or in an environment where a local time source is used with real time as the initial value (an environment where timeSource=F1h described in Chapter 5.5.4 of SMPTE ST2059-2), for example, by setting a fixed value for April 2024, the year the present invention was invented, as τ, this calculation will hold up until 2092, and it can be said that the actual number of PTP seconds can be accurately derived. Even in an environment where the PTP GM clock is operating on a local time source (an environment where timeSource=F0h as described in Chapter 5.5.4 of SMPTE ST2059-2), if τ is set to 0, this calculation will be valid for 68 years from the time the PTP GM has been in operation, so it can be said that the actual number of PTP seconds can be accurately derived.
[0144] According to Table 6, regardless of the combination of frame rate and sampling frequency, the period N is at least approximately 124 days, and in an environment where the PTP GM clock is running in real time, for example, when running this embodiment, the exact number of PTP seconds can be derived by entering the year and month of the day.
[0145] After t is found, by substituting t for τ, it is possible to continue to derive the number of PTP seconds stamped on RTP packets received continuously thereafter.
[0146] The PTP seconds adjustment unit 15 records the number of seconds t' obtained by adding the adjustment value α to the input PTP seconds t, and outputs it when requested by an external process.
[0147] The adjustment value α takes into account the time it takes for the reference RTP packet to reach the RTP packet receiver 11, based on the PTP seconds stamped on the packet. When using an audio timestamp as a reference, the adjustment value α is greater than the packet transmission interval of an audio RTP packet. For example, if the sampling frequency is 48 kHz and the packet contains six samples, it takes 6 / 48,000 seconds (125 microseconds) to sample the number of samples. In this case, α is set to 125 microseconds. In practice, α can be adjusted to suit the environment due to the time lag between when the reference ST2110 transmitter 3 outputs a packet and the time it takes for the packet to propagate through the network. However, IP networks are based on best-effort processing, and any time errors are absorbed by the receiver's buffer, so there is no problem using α as the audio packet transmission interval.
[0148] As a guide to the buffer size of a typical receiving device, many receiving devices have a redundancy mechanism that receives the same RTP packet from multiple different networks (usually two) according to SMPTE ST2022-7, and has a mechanism that adopts the packet that arrives first. If packets do not arrive from both networks for a certain period of time, it is considered a packet loss. The waiting time for that packet is defined as a class, with Class A: 10 milliseconds, Class B: 50 milliseconds, Class C: 150 milliseconds, and Class D: 150 microseconds. Therefore, errors within this time are considered to be within an absorbable range.
[0149] 6 is a block diagram showing the functional configuration of a PTP seconds derivation device 20 according to a second embodiment of the present invention. The PTP seconds derivation device 20 includes an RTP packet receiving unit 11, an NTP (Network Time Protocol) time synchronization unit 21, an internal device clock 22, a DTAI / PTP seconds calculation unit 23, and a PTP seconds adjustment unit 15.
[0150] The NTP time synchronization unit 21 connects to and synchronizes with the NTP server 5, and adjusts the time of the device clock 22. This is a typical NTP operation, and the time accuracy is on the order of milliseconds to several tens of milliseconds.
[0151] Preferably, in the case of the PTP seconds number derivation device 20, the RTP packet receiving unit 11 receives the audio RTP packets transmitted from the ST2110 transmission device 3 that serves as a reference.
[0152] The DTAI / PTP seconds calculation unit 23 uses the current time unitxtime (the number of seconds originating from 00:00:00 UTC on January 1, 1970) u obtained from the internal clock 22 and the audio timestamp Ta obtained from the RTP packet reception unit 11 to derive the PTP seconds t stamped in the RTP packets received by the RTP packet reception unit 11 and the accumulated leap seconds DTAI. The derived PTP seconds t is output to the PTP seconds adjustment unit 15. The derived DTAI is output as needed to subsequent processing that requires DTAI.
[0153] The operation of the DTAI / PTP seconds calculation unit 23 will now be described in detail.
[0154] NTP is based on UTC, which is adjusted for leap seconds, while PTP is based on TAI, which is not adjusted for leap seconds. TAI is advanced by an accumulated leap second, called DTAI. In other words, DTAI = TAI - UTC, and since leap seconds are inserted or deleted in seconds, DTAI is an integer. As of April 2024, DTAI is 37 seconds, and leap seconds have been inserted 27 times in the 50 years since their introduction in 1972.
[0155] Using this property of DTAI, if the lower 32 bits of the value obtained by multiplying u by the clock frequency (rounded down) are Tu, and the value obtained by multiplying u by the clock frequency (rounded down) and shifting it 32 bits to the right is Cu, then the value obtained by subtracting Tu from Ta (DiffT) will be close to a multiple of the clock frequency when viewed as a 32-bit integer (signed integer). Therefore, DTAI can be derived by dividing this value by the clock frequency and rounding off the decimal point.
[0156] In this case, the NTP error and the time it takes for the RTP packet to reach the RTP packet receiver 11 range from a few microseconds at most to a few hundred milliseconds (even considering an extreme case where the reference ST2110 transmitter 3 is on the other side of the world), so this error can be absorbed by rounding off. This calculation method does not work if the DTAI exceeds half the cycle period, but the cycle period is at least a little over 12 hours, and DTAI does not change suddenly in the first place, so it is thought that there is no problem even if such extreme cases are not taken into account.
[0157] The audio cycle Ca can be calculated as follows:
[0158] If DiffT is positive and Tu>Ta, then Ca = Cu+1, If DiffT is negative and Ta > Tu, then Ca = Cu-1. (However, this case does not usually occur, since it is difficult to imagine that DTAI will be negative.) Otherwise, Ca = Cu.
[0159] From the above, the PTP seconds t stamped on the received RTP packet can be derived using the following formula:
[0160] [Number 21] t = (Ca×2^32+Ta) / f
[0161] Once the PTP seconds and DTAI are determined, UTC = PTP + DTAI, which is equivalent to being able to derive the UTC stamped in the RTP packet. The SMPTE ST2059-1 standard specifies a formula for calculating the time code using the PTP seconds and DTAI value, and by using this embodiment, it is also possible to derive the time code of a received video frame without directly locking to PTP.
[0162] FIG. 7 is a diagram illustrating a receiving device 30 according to a third embodiment of the present invention. Receiving device 30 is configured to include the functions of PTP second number derivation device 10 or PTP second number derivation device 20, an RTP packet receiving unit 31, a synchronization processing unit 32, and a decapsulating unit 33. The NTP server 5, which is necessary when using the functions of PTP second number derivation device 20, is omitted from the drawing.
[0163] The RTP packet receiving section 31 receives at least one packet of video, audio or auxiliary data transmitted from the ST2110 transmitting device 6, and outputs it to subsequent processing.
[0164] The received RTP packets may be the same as the RTP packets transmitted from the reference ST2110 transmission device 3. Depending on the purpose of reception, for example, reception of multiple types of video or multiple types of audio may be performed.
[0165] The frame rate of the reference RTP packet and the RTP packet to be received may be different. For example, the reference RTP packet may be 29.97 fps RTP packet, and the RTP packet to be received may be 25 fps RTP packet.
[0166] The synchronization processing unit 32 extracts the RTP timestamp value from the RTP packet received from the RTP packet receiving unit 31, derives the PTP seconds that form the basis of the RTP timestamp value based on the PTP seconds output from the function of the PTP seconds derivation device 10 or 20, and performs synchronization processing. This is the same process as the general ST2110 receiving device described above, but the original PTP seconds are not from an internal clock locked to PTP, but are PTP seconds obtained using the first or second embodiment of the present invention. The synchronization processing unit 32 outputs the derived PTP seconds stamped in the RTP packet and the RTP packet to the decapsulation unit 33.
[0167] The decapsulation unit 33 converts the RTP packets into frame images, ancillary data, and audio data in accordance with the ST2110 standard, performs phase adjustment using the PTP seconds derived by the synchronization processing unit 32, and outputs the data to subsequent processing.
[0168] The receiving device 30 can perform various processes based on the information decapsulated by the above process. For example, it can act as a recording server to convert video and audio into files, monitor the video and audio for inappropriate content, receive and control inter-station control signal packets included in the auxiliary data, and so on.
[0169] FIG. 8 is a diagram illustrating a transmitting device 40 according to a fourth embodiment of the present invention. Transmitting device 40 is configured to include the functions of PTP second number derivation device 10 or PTP second number derivation device 20, an encapsulating unit 41, a timestamp assigning unit 42, and a transmitting unit 43. The NTP server 5, which is necessary when using the functions of PTP second number derivation device 20, is omitted from the drawing.
[0170] The encapsulation unit 41 receives at least one of the video frame data, ancillary data, and audio data to be transmitted, encapsulates it into an RTP packet in a packet format defined by the ST2110 standard, and outputs it to the timestamp assignment unit 42. At this time, the RTP timestamp value is set to null.
[0171] The timestamp assigning unit 42 assigns a timestamp value to the RTP packet based on the PTP seconds output from the function of the PTP seconds derivation device 10 or the PTP seconds derivation device 20 , and outputs the packet to the transmitting unit 43 .
[0172] This is similar to the processing performed by the general ST2110 transmission device described above, but the original PTP seconds are not those of an internal clock locked to PTP, but rather those obtained using the first or second embodiment of the present invention.
[0173] The transmitter 43 transmits the RTP packets received from the previous processing to the network. The ST2110 receiver 7 receives the RTP packets as needed.
[0174] Since there is no difference between the RTP packets output by the present invention and the RTP packets output from a general ST2110 transmitter (locked to PTP), they can be received by the ST2110 receiver 7.
[0175] Through the above processing, the transmitting device 40 can output various data as RTP packets conforming to the ST2110 standard. Examples include a generator that outputs color bars over IP, a transmission server that decodes and plays video files and outputs them over IP, and a generator that uses the functions of the PTP second number derivation device 20 to obtain DTAI and output time codes over IP.
[0176] As one embodiment of the present invention, the receiving device 30 and the transmitting device 40 can be combined to form a transmitting / receiving device that receives at least one of video, audio, and auxiliary data, rewrites a portion of the data, and transmits it. As an example, it is possible to transmit video with a caption or the like superimposed on the received video.
[0177] FIG. 9 is a diagram illustrating a transmitting / receiving device 50 according to a fifth embodiment of the present invention. The transmitting / receiving device 50 includes an RTP packet receiving unit 31, a packet rewriting unit 51, and a transmitting unit 43. Packet rewriting unit 51 rewrites part of the RTP packets input from RTP packet receiving unit 31, using necessary information including the timestamp, and outputs the rewritten RTP packets to transmitting unit 43. The RTP packets output to the network by transmitting unit 43 are received by ST2110 receiving device 7 as needed.
[0178] The operation of the packet rewriting unit 51 will now be described in detail.
[0179] First, the case of rewriting a video will be described.
[0180] Figure 10 shows an example of the packet structure for video RTP packets in a packet format defined by the SMPTE ST2110 standard. A video RTP packet consists of an RTP header 101, an RTP payload header 102, and an RTP payload 103. While we will not go into detailed explanations of the standards for each item in each header, the header contains, in addition to a timestamp, an SSRC 200 that is an RTP session identifier, information such as the vertical and horizontal position of the image from which the data originates, and how many bytes it contains.
[0181] Packet rewriting unit 51 reuses the header information other than SSRC 200 of the video RTP packet input from RTP packet receiving unit 31 as is, replaces RTP payload 103 with pixel data of the same vertical and horizontal positions from the data to be transmitted input to packet rewriting unit 51, and inputs it to transmitting unit 43. SSRC 200 is an identifier for the RTP session, and may be determined randomly, but must remain the same value while this process is operating continuously.
[0182] Next, the case of rewriting audio will be described.
[0183] 11 shows the packet structure for audio in an RTP packet in the packet format defined by the ST2110 standard. It is simpler than video, and does not have an RTP payload header 102; it consists only of an RTP header 101 and an RTP payload 103.
[0184] Packet rewriting unit 51 reuses the header information other than SSRC 200 of the audio RTP packet input from RTP packet receiving unit 31 as is, replaces RTP payload 103 with the audio data to be transmitted that was input to packet rewriting unit 51, and inputs the packet to transmitting unit 43. SSRC 200 is rewritten in the same way as in the case of video.
[0185] However, as described above, in the case of audio, the RTP timestamp value is the time of the first sample in the RTP packet, so at least the time equivalent to the packet transmission interval has passed by the time this packet arrives at packet rewriting unit 51. Therefore, it is preferable to rewrite the RTP timestamp value to a value obtained by adding the number of samples included in the RTP packet.
[0186] Next, a case where auxiliary data is rewritten will be described.
[0187] 12 shows an example of the packet structure for ancillary data in an RTP packet in the packet format defined in the ST2110 standard. It consists of an RTP header 101, an RTP payload header 104, and an RTP payload 105. There may be multiple RTP payloads 105, depending on the number specified in the ANC_Count field of the RTP payload header 104.
[0188] Packet rewriting unit 51 reuses the header information other than SSRC 200 of the auxiliary data RTP packet input from RTP packet receiving unit 31 as is, replaces RTP payload 105 with the auxiliary data to be transmitted that was input to packet rewriting unit 51, and inputs the packet to transmitting unit 43. SSRC 200 is rewritten in the same way as in the case of video.
[0189] Since packet rewriting unit 51 is driven by the input RTP packet as a trigger, the RTP packet output by ST2110 transmitting device 3 that serves as the reference and the data to be transmitted that is input to transmitting device 50 must match in resolution and frame rate in the case of video, and in sampling frequency in the case of audio. In the case of ancillary data, the reference ancillary data RTP packet must be an ancillary data RTP packet other than a keep-alive packet.
[0190] Also, a part of the RTP payload may be reused and only the necessary parts may be rewritten.
[0191] As described above, the transmitting / receiving device 50 can more easily transmit RTP packets in the packet format defined by the ST2110 standard when precise synchronization between video frames and audio is not required.
[0192] The present invention is not limited to the above-described embodiment, and can be implemented in various other forms without departing from the spirit of the present invention. Therefore, the above-described embodiment is merely illustrative in all respects and should not be interpreted as limiting. For example, although there is only one PTP-non-compatible SW4 in FIG. 1, multiple PTP-non-compatible SWs may be used. Similarly, in FIG. 1, the PTP-compatible SW1 and PTP-compatible SW4 may be connected via a wide area network. [Explanation of symbols]
[0193] 1 PTP compatible SW 2 PTP GM (Grandmaster) 3 Reference ST2110 transmitter 4. PTP-incompatible SW 5 NTP Server 6 ST2110 transmitter 7 ST2110 receiver 10 PTP seconds derivation device 11 RTP packet receiver 12 Seconds in period derivation part 13 PTP seconds derivation part 15 PTP seconds adjustment unit 20 PTP seconds derivation device 21 NTP time synchronization section 22 Internal clock 23 DTAI / PTP seconds derivation part 30 Receiving device 31 RTP packet receiver 32 Synchronization processing section 33 Decapsulation Unit 40 Transmitting device 41 Encapsulation Department 42 Time stamp assignment section 43 Transmitter 50 Transmitting and Receiving Device 51 Packet rewriting unit 101 RTP Header 102,104 RTP payload header 103,105 RTP payload a number representing how many N cycles have passed (any non-negative number) c quotient of τ divided by N f sampling frequency g is the variable obtained from Equation 13 h is the variable whose value is obtained in Equation 14 i is the variable whose value is obtained in Equation 15 j is a temporary variable used to derive Equation 3 k is a temporary variable used to derive Equation 3 m Temporary variable used in formula 2 n is a non-negative integer number used for explanation when shifting array D p is the number of seconds elapsed within a period of N seconds, the number of PTP seconds stamped in the received RTP packet Remainder when qt is divided by N r Ca mod Mv t is the number of PTP seconds stamped in the received RTP packet t' t plus α u Current unixtime w Variable representing Cv mod V Variable representing y Ca mod A z Ca mod Ma A is the video cycle period, which can be derived from the remainder when the timestamp is divided by Ra. Ca Audio Cycle Cu u multiplied by the clock frequency (decimal parts truncated) and shifted 32 bits to the right Cv Video Cycle D Array that records video timestamp increments DiffT The value obtained by subtracting Tu from Ta, viewed as a 32-bit integer Ma The number of audio cycles in which the video timestamp and audio timestamp periods are in phase. Mv The number of video cycles in which the video and audio timestamp periods are in phase. N The period in which the video timestamp and audio timestamp increase at regular intervals and the period in which they wrap around are all phase-aligned. Na Number of audio cycles corresponding to period N Nv Number of video cycles corresponding to period N Pa is the number of seconds elapsed within a period of N seconds, the number of PTP seconds stamped in the received audio RTP packet Pv The number of seconds elapsed within a period of N seconds, the number of PTP seconds stamped in the received RTP packets of video or auxiliary data Ra is the number of audio timestamps that increase at regular intervals. Rv is a regularly increasing number of video timestamps S cycle difference Ta Audio Timestamp Ta' Audio timestamp when video timestamp is counted backward to 0 Told Previous value of video timestamp Tu: The lower 32 bits of the value (decimal parts truncated) obtained by multiplying u by the clock frequency TV video timestamp Tv' Video timestamp when audio timestamp is counted backward to 0 Tvn normalized video timestamp V is the period of the video cycle, which can be derived from the remainder when the timestamp is divided by Rv. α Adjustment to add to the PTP seconds stamped in the received RTP packet ε The amount of time allowed for the difference between video and audio timestamps The number of seconds given in the range tN < τ ≦ t to calculate τ a
Claims
1. A processing device comprising a receiving unit for an RTP packet in a packet format defined by the SMPTE ST2110 standard, the processing device using an RTP timestamp value included in the received RTP packet as a reference.
2. The RTP packet receiving unit receives video or auxiliary data packets and audio packets; Using a combination of RTP timestamp values included in each received RTP packet and the property that each RTP timestamp value increases at regular intervals, The method is characterized in that, as an element of the PTP seconds stamped on the received RTP packet, the number of elapsed seconds p seconds (0≦p<N) within a period of N seconds is derived. The N seconds is equal to a period in which the phases of the period in which the RTP timestamp values increase at regular intervals and the period in which the RTP timestamp values wrap around are all aligned. The processing device of claim 1 .
3. 3. The processing device according to claim 2, wherein the value a is assumed to be an arbitrary non-negative integer, and the number of PTP seconds stamped on the received RTP packet is calculated as a×N+p.
4. The processing device according to claim 2, characterized in that the value a is derived from the PTP seconds stamped in the received RTP packet using a value in the range up to the N seconds before, and the PTP seconds stamped in the received RTP packet is derived as a × N + p.
5. 2. The processing device according to claim 1, further comprising a time synchronization unit using NTP (Network Time Protocol) in an environment where a PTP GM (PTP Grandmaster) clock operates in real time, and deriving the PTP seconds stamped in the received RTP packet and the accumulated leap second value DTAI using an RTP timestamp value included in the received RTP packet and a current time synchronized by the NTP.
6. 6. A processing device according to claim 3, wherein the processing device outputs a value obtained by adding an adjustment value α to the PTP number of seconds stamped in the last received RTP packet as the PTP number of seconds to a subsequent processing stage.
7. 7. A receiving device comprising the processing device according to claim 6, further comprising a separate receiving section for RTP packets in a packet format defined by the ST2110 standard, and for processing the RTP packets received by said separate receiving section.
8. 7. A transmitting device comprising the processing device according to claim 6 and a transmitting section for RTP packets in a packet format defined by the ST2110 standard, the transmitting section transmitting the RTP packets.
9. A transmitting / receiving device that integrates the receiving device of claim 7 and the transmitting device of claim 8, and receives at least one of video, audio and auxiliary data, processes the data and transmits it as another RTP packet.
10. A transmitting / receiving device comprising a packet rewriting unit and a transmitting unit for RTP packets in a packet format defined by the ST2110 standard, in addition to the processing device described in claim 1, wherein the transmitting unit transmits another RTP packet by reusing information including an RTP timestamp value of the RTP packet received by the receiving unit.
11. A program that causes a computer to function as any one of the processing devices according to claims 1 to 5.
12. A program that causes a computer to function as the processing device according to claim 6.
13. A program that causes a computer to function as the receiving device according to claim 7.
14. A program that causes a computer to function as the transmitting / receiving device according to claim 8.
15. A program that causes a computer to function as the transmitting / receiving device according to claim 9.
16. A program that causes a computer to function as the transmitting / receiving device according to claim 10.
Citation Information
Patent Citations
Method and apparatus for transmitting or receiving broadcast content over one or more networks
JP2017509199A
Broadcast signal processing system and broadcast signal processing method
JP2020162078A
Broadcasting system, encoder, multiplexing device, multiplexing method, system switching device, and synchronization control device
JP2023038249A
Video transmission system
JP2018042020A
Receiver receiving video image signal
JP2021077940A