FAULT-TOLERANT CLOCK NETWORK

A redundant large master clock system with primary and backup configurations addresses synchronization challenges in time-sensitive networks, ensuring continuous operation and reliability by dynamically switching between clocks using the BMCA algorithm.

DE102014204752B4Active Publication Date: 2025-08-14AVAGO TECHNOLOGIES INTERNATIONAL SALES PTE LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
DE102014204752
Authority / Receiving Office
DE · DE
Patent Type
Patents
Current Assignee / Owner
Priority Date
2014-03-05
Filing Date
2014-03-14
Publication Date
2025-08-14
Estimated Expiration
2034-03-14

AI Technical Summary

Technical Problem

Existing communication networks face challenges in maintaining reliable clock synchronization, particularly in time-sensitive networks, where failures in primary synchronization sources can disrupt the seamless operation of devices and applications.

Method used

Implementing a redundant large master clock system with a primary and backup clock configuration, utilizing the Best Master Clock Algorithm (BMCA) to select and dynamically switch between primary and backup clocks, ensuring seamless transition and synchronization through active or passive modes, and maintaining synchronization within predetermined tolerances.

Benefits of technology

Ensures continuous and reliable clock synchronization in time-sensitive networks by providing fault-tolerant operation, allowing devices to maintain synchronization even during primary clock failures, thus enhancing network stability and performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

System (100) with: a primary master clock device (110); and a reserve main clock device (120); wherein the primary grandmaster clock device (110) and the backup grandmaster clock device (120) are selected from an alternative grandmaster device list comprising a list of multiple grandmaster clock devices (110, 120) in a network (190) by means of a selection process based on an order of priority, a clock quality, a traceability of the clock source and / or a position of the clock device (110, 120) with respect to the network (190); an interface (112) configured to communicate on the network (190) and to receive a primary synchronization message (170), the primary synchronization message (170) being received from the primary master clock device (110) via the network (190), the primary synchronization message (170) comprising a clock signal from the primary master clock device (110); a processor configured to keep the clock of the backup master clock device (120) substantially synchronous with the clock signal received in the primary synchronization message (170); and wherein the processor is further configured to generate a backup synchronization message (180) based on the clock of the backup master clock device (120), and wherein the interface (112) is further configured to transmit the generated backup synchronization message (180) to other devices via the network (190); wherein the processor is further configured to: extract a global identifier from the received primary synchronization message (170); and inserting the global identifier into the generated reserve synchronization message (180) before transmitting the reserve synchronization message (180).
Need to check novelty before this filing date? Find Prior Art

Description

Technical area

[0001] This disclosure relates to synchronizing clocks of different nodes arranged in a distributed network, for example by using redundant grandmaster clocks to protect against faults. background

[0002] Communication protocols are widely used in networks such as local area networks (LANs) and metropolitan area networks (MANs). Examples include Ethernet, Token Ring, wireless LANs, bridged local area networks, and virtually bridged local area networks, as defined in the Institute of Electrical and Electronics Engineers (IEEE) 802 standard. The IEEE 802 standards refer to networks that carry variable-size packets. The services and protocols specified in IEEE 802 are mapped to the bottom two layers (data link and physical) of the seven-layer network reference model for open systems communications (OSI). The OSI data link layer is divided into two sublayers called Logical Link Control (LLC) and Media Access Control (MAC).

[0003] The clocks in devices in a computer network can be synchronized so that the devices operate cooperatively. The granularity to which the clocks, or simply the devices, can be synchronized depends on the purpose of the network. Therefore, processing and motion or other control-oriented network applications in mission-critical networks and networks called time-sensitive networks (TSNs), such as those that may be used in an automotive control system such as a powertrain, a traction control system, and in a manufacturing environment such as high-speed motion control, in power grid control systems, in a financial transaction network, in security networks, and other such networks that support time-sensitive applications, depend on a reliable clock source that keeps the devices or end stations in the network synchronized.Furthermore, with the advancement in mobile networks such as 3G, 4G, 4G LTE, WiFi, and various other such networks, synchronization of network-connected devices has become even more important.

[0004] US 2008 / 0168182 A1 discloses a system comprising a plurality of communication devices. Each communication device has a clock whose value is passed on to neighboring communication devices.

[0005] From the document US 4,736,393 A a timing control arrangement is known which dynamically controls the distribution of timing information in a digital communication system.

[0006] The document DE 693 29 294 T4 describes a hierarchical synchronization method for a telecommunications system with a plurality of nodes, wherein the nodes exchange signals containing synchronization messages with information about the priority of the respective signal in the synchronization hierarchy.

[0007] According to the present invention, a method, a system, and an end station device in a network having the features of the independent claims are provided.

[0008] Advantageous further developments of the invention are specified in the subclaims. Short description of the drawings

[0009] The innovation can be better understood by reference to the following drawings and description. In the figures, like reference numerals designate corresponding parts throughout the different views. Fig. 1 is a block diagram of an exemplary system using a primary grand base clock (primary grand base clock) and a backup grand base clock. Fig. 2 is a block diagram of another exemplary system using a primary grand base clock and a backup grand base clock. Fig. 3A and Fig. 3B are block diagrams of example configurations of a primary grand base clock and a backup grand base clock operating with a common or separate clock source. Fig. 4 is a flowchart illustrating exemplary steps performed by an exemplary reserve master clock. Fig. 5 is a block diagram of an exemplary system for healing a primary grandmaster primary synchronization message, where the primary grandmaster clock and the backup grandmaster clock share a common primary clock source. Fig. 6 is a block diagram of an exemplary system for healing the primary grandmaster clock primary synchronization message, wherein the primary grandmaster clock and the backup grandmaster clock do not share a primary clock source. Fig. 7 is a flowchart illustrating exemplary steps performed during healing of the primary grandmaster clock primary synchronization message, where the primary grandmaster clock and the backup grandmaster clock do not share a primary clock source. Detailed description

[0010] The following discussion refers to synchronizing clocks (clock generators) of different nodes arranged in a distributed network.

[0011] In this context, a clock can be any device with a network connection, and can be either the source of a synchronization reference (master) or a destination for the synchronization reference (slave). A synchronization master can be selected for each network segment in the distributed network. Further, the source timing reference can be referred to as a grand-base clock. Therefore, a grand-base clock can be the clock that serves as the primary source of time to which all devices in the network are ultimately synchronized. Two or more clocks are generally said to be "synchronized" within a specified uncertainty if they have the same period, and measurements of each time segment by both clocks differ from each other by no more than the specified uncertainty.Therefore, timestamps generated by two synchronized clocks for the same event cannot differ from each other by more than the specified uncertainty. The specified uncertainty provides an engineering tolerance that can vary based on the mission criticality of the network. For example, in a mission-critical setting, such as an industrial manufacturing line, the engineering tolerance can be a very small time period on the order of milliseconds, microseconds, or even smaller. Conversely, in a relatively lax setting, such as an audio-video transmission, the engineering tolerance of the time period can be one second or longer. The described system and method are not limited by any particular engineering tolerance value.

[0012] Redundant master clocks can be used to protect against synchronization clock errors in the network. The primary master clock (pGM) and the backup master clock (bGM) can be preconfigured or dynamically determined. The devices selected as the pGM and bGM can be selected and configured through a selection process based on clock quality, priority (preference), and other parameters using selection procedures such as the Best Master Clock Algorithm (BMCA) as defined in IEEE 1588-2008 and IEEE 802.1AS, or by using other selection techniques. For example, clock precedence can be determined based on a clock source, such as a GPS, or a clock stratum level, such as Stratum 1 or Stratum 2. Traceability of the clock source can also be a factor in determining clock precedence.For example, if a large-scale clock device is directly traceable to a clock source such as a GPS, the large-scale clock device may have higher precedence than another large-scale clock device that derives a clock from another device that also uses a GPS clock source. Precedence may also be based on whether a large-scale clock device is located at a relatively central location relative to the network, whether the large-scale clock device has a robust power backup, and other such network-specific details. The relatively central location of the large-scale clock device relative to the network may be determined based on empirical latency values ​​for transmitting messages from the large-scale clock device to the devices in the network, or on any other performance-related condition.

[0013] In one example, the BMCA may pre-select the pGM and the bGM, and the corresponding devices may be identified as the pGM and the bGM with respect to their respective selection when the devices are installed in the network. In another example, the devices may be selected as the pGM and the bGM after installation in the network, wherein the BCMA process may be performed dynamically. The BCMA or any other selection process used may be configured to select more than one bGM device. Therefore, an alternate grandmaster device list may be generated by a selection process such that the alternate grandmaster device list contains one or more bGM devices or potential bGM devices. The alternate grandmaster device list may list the potential bGMs in order of priority, clock quality, or any other parameter.

[0014] A bGM device provides a seamless transition for a network device in the TSN in the event of a failure to receive a primary synchronization signal from the pGM. The failure (error) can be a failure of the pGM or a failure of a link in the TSN. For example, if a pGM fails, a bGM can become the new grandmaster. Therefore, a redundant grandmaster clock system and method is described. It may also be desirable to transition the frequency and phase from the pGM to the bGM under failure conditions. For example, switching from the pGM to the bGM (and from the bGM to the pGM) can involve controlled phase and frequency deviation. The bGM can be provided as active (sending sync) or passive (not sending sync). Redundant grandmaster clocks that are synchronized with each other can also be supported.There can be multiple instances of bGM, either preconfigured or dynamically determined and configured using an alternative best base clock selection as part of the BCMA as defined in IEEE 1588-2008 and IEEE 802.1AS, or using other equivalent dynamic selection mechanisms. When multiple bGMs are used, the list of bGMs is in priority order (or precedence order) in both provisioning and dynamic selection cases.

[0015] Fig. 1 is a block diagram of an exemplary system 100 employing a primary clock source including a pGM 110 and a bGM 120. The pGM 110 and the bGM 120 may be selected together using a primary clock selection algorithm such as BMCA or any other algorithm. The bGM 120 may be selected while the pGM 110 is operating and before any failure is detected. The pGM 110 and / or the bGM 120 may be provided, for example, with a clock source such as a global positioning system (GPS), a world time server, or any other such clock source. The pGM 110 and the bGM 120 may each include an interface 112 for transmitting and / or receiving messages over the network. In some cases, the pGM 110 and the bGM 120 may have separate interfaces for transmitting and receiving messages.The messages may include a pGM clock synchronization message 170 (pSync 170) and a bGM clock synchronization message 180 (bSync 180). The pSync 170 and bSync 180 messages may both be used to derive a clock for end stations such as Ethernet stations 130. The end stations may also be referred to as terminals, network devices, network nodes, or simply nodes. The end stations may also have interfaces 112 for transmitting and receiving messages over the network. In the system of FIG. Fig. 1, the pSync 170 and bSync 180 messages can be sent to and received from the Ethernet stations 130 via Ethernet bridges 140, 150, and Ethernet or other IEEE 802-compliant time-sensitive networks (TSN) 190. The pSync message can also be sent to the bGM 120, and the bSync message can be sent to the pGM 110, as well as via the Ethernet bridges 140, 150.

[0016] The pGM 110 may be a grand base clock device selected to be the grand base clock of the system 100. The selection may be based on an algorithm such as the Best Base Clock (BMC) algorithm, or any other process. The selection may be based on several factors such as network speed, uptime, variance, assigned priority, or other such factors. The bGM 120 may be a grand base clock device selected to be a backup grand base clock of the system 100. The bGM selection may be based on the same algorithms and factors as those of the pGM. Alternatively, the bGM 120 may be selected based on a different algorithm and / or other factors. Additionally, to provide a seamless backup clock signal, the bGM 120 may be selected while the pGM is functional. Alternatively, the bGM may be selected after a pGM failure is detected. Although Fig. 1 illustrates only one bGM, multiple reserve master clocks may be selected. Therefore, at any given time, one pGM and at least one bGM may be functionally active or operating in system 100.

[0017] The pSync 170 and bSync messages may be configured according to a time protocol such as a Network Time Protocol (NTP), a Precision Time Protocol (PTP), or other such protocols. Furthermore, the messages may conform to protocol standards such as IEEE 1588-2002, IEEE 1588-2008, or any other standard. Messages may be transported over the network 190 using multicast, unicast, or any other communication mechanism or protocol. Additionally or alternatively, the messages may be transported using Internet Protocol (IP) packets such as IPv4 or IPv6 packets. Alternatively or additionally, the messages may be encapsulated using DeviceNet, ControlNet, IEEE 802.3 Ethernet, PTP, or other such protocols.

[0018] For the purpose of explanation, the terminal stations in Fig. 1 as Ethernet stations 130, however, the end stations may also include other types of nodes in the network, for example, end stations for Token Ring, wireless local area networks, bridging, and virtually bridged local area network-type networks. The end stations may be nodes connected to the network, such as network bridges, routers, modems, workstations, mobile phones, laptop computers, desktop computers, servers, tablet devices, smartphones, or any other device that may be connected in the network 190. The end stations may also be machines such as industrial robots, conveyor belts, or any other such industrial machine. The end stations may also be vehicles such as cars, trucks, airplanes, spaceships, or other devices that may be synchronized. Although end stations 130 in Fig. 1 as a single block, it is understood that end stations may comprise multiple network nodes distributed throughout the network. End stations may be intermediate nodes in the network, although they are referred to as "end" stations. The end stations may also be boundary clocks. A boundary clock may typically be used to transfer synchronization from one network segment with a single time domain, such as an Internet Protocol (IP) subnet, to another, typically by means of a router that blocks all other synchronization messages. The end stations may comprise one or more processors and one or more non-volatile memory devices. The processors may be responsible for performing the various functions at an end station.The end stations may also have a local clock that may be synchronized with one or more master clock devices using synchronization messages from the master clock devices. The end stations may be part of a distributed network system, and the end station operations may be coordinated based on the local clock signals at each respective end station. Therefore, maintaining synchronization of the local clock signals throughout the end stations may enable the distributed network system, such as system 100, to operate at scheduled timing intervals and / or events.

[0019] As indicated above, the bGM 120 can operate in an active or a passive mode. In the active mode, the bGM 120 can continuously send the bSync message 180 over the network, even when the pGM is operating and transmitting the pSync message 170. Therefore, in the active mode, the end stations 130 can receive both the pSync 170 and the bSync messages 180. The pSync 170 and bSync messages 180 can be received substantially simultaneously or within a certain time interval of each other. The time interval within which the messages are received can be approximately half the time interval between consecutive pSync 170 (or bSync 180) messages. In such a case, if the end stations 130 receive the pSync 170 (or bSync 180) at a frequency F, the end stations 130 can receive the combination of the two messages at twice the frequency, 2F.In passive mode, the bGM 120 can capture the pSync message, and upon failure to capture or receive the pSync message 170 within an expected time period, the bGM 120 can send the bSync message 180. The bGM timeout period is often shorter than the required hold-over time, where the hold-over time is the time for which the end stations could continue to operate within a specified clock tolerance. The hold-over time is application-specific and can be of different values ​​(e.g., lowest common denominator) for each time-sensitive network. Therefore, the Ethernet station 130 can only receive one of the pSync message 170 or the bSync message 180 in passive mode. However, the bGM 120 itself can be operational in passive mode and generate the bSync message 180.In one example, the bGM 180 may transmit the generated bSync message 180 to a temporary buffer such as a silent drop instead of to the network. Once the bGM 120 detects a non-receipt of the pSync message 170, also referred to as a pSync failure, the bGM 120 may change the destination of the generated bSync messages 180 so that the messages are transmitted to the end stations rather than to the temporary buffer. In cases where there are multiple backup grandmaster devices, the Ethernet station 130 may receive multiple bSync messages in the event of an interruption in reception of the pSync message 170. If multiple bGMs exist, their relative precedence may be known, such as based on the order of the bGMs in the alternate grandmaster device list. Accordingly, the Ethernet station 130 may use the bSync message from the bGM with the highest precedence according to the list.Alternatively, the Ethernet station 130 may use all of the received bSync messages. Alternatively or additionally, in the case of multiple bGMs, the bGM 120 may detect bSync messages from the other bGMs in the list. The bGM 120 may compare its own assigned priority and the priority of the bGM from which another bSync was received. If the bGM 120 detects that the other bGM has a higher priority, the bGM 120 may stop transmitting the bSync. Alternatively or additionally, the bGM 120 may continue transmitting the bSync until the pSync message is detected from the original pGM 110.

[0020] In both active and passive modes, the Ethernet stations 130 can use all of the received synchronization messages to derive their clocks, regardless of the source clock of the received synchronization or clock signals. Alternatively, the Ethernet stations 130 can use clock signals received from a specific source clock to derive their clocks. A synchronization message can include an identifier indicating the source of the message. The identifier can be a global identifier included in a synchronization message regardless of the source. In such cases, the end stations 130 cannot distinguish between the received synchronization messages. The bGM 120 can extract the global identifier from the pSync message 170 and embed or include it in the bSync message 180 generated by the bGM 120.The bGM 120 may store the global identifier for embedding. Alternatively or additionally, a synchronization message may include a unique identifier from the grandmaster clock device that generated the message. Therefore, in addition to the global identifier, the bGM 120 may embed an identifier in the bSync message 120 that represents the bGM 120. In another example, the bGM 120 may add only the unique identifier to the bSync message 180, and not the global identifier. The unique identifier may enable the end stations 130 to identify the source of a received synchronization message and perform analysis regarding the reliability of the source grandmaster clock device, such as the reliability of the source device. The end station behavior may remain the same, although the synchronization rate may change.For example, in the active mode, the Ethernet station 130 may receive multiple synchronization messages, such as the pSync 170 and bSync 180 messages. The synchronization messages from multiple sources may be received within a predetermined time period of each other. Alternatively, the synchronization messages may be received substantially simultaneously. Alternatively or additionally, the synchronization messages may be received at a particular rate. The Ethernet station 130 may derive the local clock based on all of the received synchronization messages. In the passive mode, the Ethernet station 130 may receive only one pSync 170 or bSync 180 message at a given time. In such a case, the Ethernet station 130 may derive the clock based only on the received pSync 170 or bSync 180 message.Therefore, regardless of the frequency at which the synchronization messages are received, the Ethernet station 130 can continue to act synchronously in each of these cases.

[0021] Fig. 2 is a block diagram of an exemplary system 200 in which the pGM 110 has failed (failed) to send accurate pSync messages 170. The bGM 120, operating in an active mode, may continue to send bSync messages 180 to the Ethernet station 130. As described elsewhere, the Ethernet station 130 may continue to operate seamlessly based on the continuously received bSync messages 180 because the Ethernet station 130 may derive the clock based on the bSync message 180.

[0022] Alternatively, when the bGM 120 operates in a passive mode, the bSync message 180 is not transmitted continuously. In this case, as described elsewhere, the bGM 120 may wait for a predetermined timeout or hold-over time before transmitting the bSync message 180. A hold-over time is the period of time used to keep a device synchronization-stabilized when a device's synchronization source is interrupted or temporarily unavailable. The hold-over time after which the bGM 120 sends the bSync message 180 may be shorter than a hold-over time used by the Ethernet station 130. For example, the bGM 120 hold-over time may be a number of milliseconds, while the Ethernet station 130 hold-over time may be one second. Therefore, the Ethernet station 130 can receive the bSync message 180 within the hold time of the Ethernet station 130.Consequently, Ethernet station 130 can continue to derive its clock seamlessly from bSync message 180 (instead of pSync message 170) and can continue synchronized operation. Therefore, Ethernet stations 130 can continue to seamlessly synchronize their clocks, even during a pGM error or other condition where pSync message 170 is not received. Alternatively, during passive mode, Ethernet stations 130 can enter a hold until they are synchronized with bGM 120 using bSync message 180.

[0023] The pGM 110 and the bGM 120 can derive their respective clock signals to derive the pSync 170 and bSync 180 messages from a traceable clock source. For example, as shown in Fig. 3A, pGM 110 and bGM 120 use a common clock source 310. Alternatively, as shown in Fig. 3B, pGM 110 and bGM 120 each include independent clock sources 350 and 360. Each of clock sources 310, 350, and 360 may be a GPS, a common clock, or any other clock signal providing device. A clock source may be external to the master clock device, and the master clock device may derive a local clock relative to the clock source. Alternatively or additionally, a master clock device may include an internal clock source. The clock sources may be used by pGM 110 and / or bGM 120 to establish a Coordinated Universal Time (UTC) time base. Clock source 350 used by pGM 110 may be a preferred clock reference, while clock source 360 ​​used by bGM 120 may be a common clock or a non-common clock source.

[0024] In the initial case where pGM 110 and bGM 120 have the traceable common clock source 310, the pSync 170 and bSync 180 messages can provide clock signals within a desired and / or predetermined tolerance. The tolerance may also be referred to as engineering tolerance and is a permissible limit on a variation of the clock signal. Tolerances are typically set to allow reasonable margin for defects and inherent variability without compromising performance and without significantly impacting the functioning of the entire system and / or individual devices. The tolerance may be based on a jitter-wander tolerance, such as according to a maximum time interval error (MTIE) mask for the system.Therefore, the clocks are derived based on the pSync 170 and / or bSync 180 messages so that the clocks at the end stations act substantially synchronously within the predetermined tolerance.

[0025] In the case where pGM 110 uses a first clock source such as clock source 350, and bGM 120 uses a second clock source such as clock source 360, extra steps may be taken to maintain synchronization if the two clock sources are separate, autonomous clock sources. For example, one of the master clock devices may derive a clock based on the other master clock device. Fig. 4 illustrates at least some exemplary steps that may be taken in this regard. In this example, in step 410, the bGM 120 may derive a local clock signal based on the pSync messages 170. This may involve aligning a local clock at the bGM 120 with the clock signal provided by the pSync messages 170. As long as a pSync failure is not detected in step 420, the bGM 120 may generate the bSync message 180 based on the derived local clock signal. Further, if the bGM 120 is operating in an active mode, which may be determined in step 450, the generated bSync message 180 may be transmitted over the network 190 in step 460. Alternatively, if the bGM 120 is not operating in the active mode in step 450, operation may return to block 410.Alternatively, if a pSync failure is detected in step 420, such as the pSync message 170 not being received for more than the hold time, the bGM 120 may generate the bSync message 180 based on the clock source 360 ​​(instead of the derived local clock). However, shifting the clock source from the pGM 110 to the clock source 360 ​​may cause an abrupt change that is above the predetermined tolerance discussed above. Therefore, the bGM 120 may transition the local clock from the pGM 110 to the clock source 360 ​​in small steps. Each step of transitioning the local clock may be performed within the allowed predetermined tolerance until the local clock is substantially synchronous with the clock source 360. Since a pSync failure has been detected, in both the active mode and the passive mode, the bGM 120 may transmit the generated bSync message 180 over the network as in step 460.

[0026] Fig. 5 is a block diagram of an exemplary system 500 for recovering or healing the pGM pSync messages 170, where the pGM 110 and the bGM 120 share a common clock source, such as a GPS clock source. As described above, in the event of a failure to receive the pGM pSync messages 170, the bGM 120 may assume the role of the system's primary grandmaster clock device and may be responsible for synchronization messages to the end stations 130 in the system. Once a failure regarding the pSync messages 170 is resolved, the system, such as system 500, may transition back to the previous state with the pGM 110 as the primary grandmaster clock device and the bGM 120 as the backup grandmaster clock device.The failure may be caused by a failure in the pGM 110, a failure in the communication channel used to convey the pGM pSync messages 170, or another such failure in any component in the system. Upon recovery from such a failure, the pGM 110 may not immediately begin sending the pSync messages 170. Instead, the pGM 110 may check for synchronization with the bGM 120 before beginning to send the pSync messages 170 again. This may be implemented because the clock signal of the bGM 120 may have drifted away from the clock signal of the pGM 110, for example, if the bGM 120 has switched to a different clock source as described above. In such a case, the bGM 120 may send the bSync messages 180 to the pGM 110. The pGM 110 can check for synchronization to the bGM 120 based on the bSync messages 180.The pGM 110 may align the local clock signal at the pGM 110 with the clock information contained in the bSync messages 180. Once the pGM 110 has derived a stable clock based on the bSync messages 180, the pGM 110 may generate the pSync messages 170 based on the local clock and transmit the pSync messages 170. The bGM 120 in the active mode may continue to transmit the bSync messages 180 as before. The bGM 120 in the passive mode may detect that the pGM 110 has resumed transmitting pSync messages 170 and continue transmitting bSync messages 180 for the holdover period. After the holdover period, the bGM 120 may stop transmitting the bSync messages 180.

[0027] Alternatively, the bGM 120 may not have drifted away from the pGM 110 during the time the pSync messages 170 are interrupted by a failure. This may be the case if the clock sources of pGM 110 and bGM 120 are traceable to the same clock source 310, and / or if the pSync failure occurred for a negligibly short period of time. In such cases, the pGM 110 may not coordinate with the bGM 120 during recovery. Recovery from the failure may also be referred to as healing or heal-back. Furthermore, the bGM 120 may not coordinate with the pGM 110 once the pGM is restored or healed back and the sending of pSync messages 170 begins. The Ethernet stations 130 can use both the pSync 170 and the bSync 180 to derive their clocks.

[0028] Fig. 6 is a block diagram of an exemplary system 600 illustrating recovery from a pSync-related failure, and wherein the pGM 110 and the bGM 120 of the system 600 do not share a primary clock source or have respective separate clock sources 350 and 360. In this context, Fig.7 illustrates exemplary steps that may be performed during and after recovery from a failure in such a case. If the clock sources of pGM 110 and bGM 120 are not traceable to the same clock source as in system 600, bGM 120 may drift in frequency and / or phase relative to pGM 110 once bGM 120 stops receiving pSync messages 170. Therefore, in step 710, pGM 110 may synchronize with bGM 120 toward recovery. Since end stations 130 may use both pSync 170 and bSync 180 for local synchronization, the respective clock signals of pSync 170 and bSync 180 may be maintained within the predetermined tolerance. To maintain tolerance, the pGM 110 may adjust the local clock at the pGM 110 using the clock information in the bSync messages 180 from the bGM 120.Once the tuning is complete, timing corrections may be made using a timing procedure such as MTIE to maintain the desired engineering tolerance. Upon reaching stable lock, the pGM 110 may generate and transmit pSync messages 170 based on the tuned local clock in step 720. Further, in step 730, the pGM 110 may move the tuned local clock to the primary reference clock source 350 by a permitted jitter wander tolerance for the timing procedure such as MTIE. Upon detecting pSync, the bGM 120 may, in turn, tune its local clock to the pGM 110 clock using the pSync messages 170 in step 740. Typically, the tuning may be performed by transitioning the clock within a jitter wander tolerance of the pGM 110, as described elsewhere in this document.When the bGM 120 operates in the active mode, the bGM 120 may continuously transmit the bSync message 180. In the passive mode, the bGM 120 may transmit the bSync message 180 for a holdover period after detecting a pSync message 170 after recovery, and may stop transmitting the bSync message 180 after the holdover period ends.

[0029] Therefore, generally during simultaneous operation of redundant master clocks, if the pGM 110 and the bGM 120 have the same primary clock source 310, the respective clocks at the pGM 110 and the bGM 120 may be synchronized, and extra steps may not be taken to synchronize the clocks. Alternatively, if the pGM 110 and the bGM 120 have different respective primary clock sources 350 and 360, the bGM 120 may initially be aligned with the pGM 110. For example, the local clock of the bGM 120 may be adjusted to operate at the same frequency and phase as the local clock of the pGM 110. This may involve processing the clock information contained in the pSync messages 170. The bGM-aligned clock may be used if it is stabilized as needed within the clock tolerances for MTIE.

[0030] Failure detection by bGM 120 can be the same regardless of whether the pGM 110 and bGM 120 use a common clock source or different clock sources. End stations such as Ethernet stations 130 can receive pSync 170 and bSync 180 messages from both the pGM 110 and bGM 120 at nominally twice the rate as with a single grandmaster. These messages may appear identical from a synchronization perspective even if a clock ID inserted into the messages identifies one of two different message sources.

[0031] The bGM 120 can operate in either an active mode or a passive mode. In the active mode, the bGM 120 can generate and transmit a bSync message regardless of whether the pSync message 170 is transmitted or not. Therefore, during a pSync failure, such as a pGM failure, an active bGM 120 can continue to send bSync messages 180 to all end stations or nodes in the network. The end stations can receive the bSync messages 180 and derive and / or adjust their respective clocks according to the bSync messages 180. During operation, if the pSync message 170 is operational, end stations 130 can receive pSync 170 and bSync 180, respectively, from the pGM 110 and the bGM 120 at twice the rate as with a single GM. These messages may appear identical from a synchronization perspective, although the clock IDs may be different in different received messages.The end stations can process the messages as if they were from the same GM because they are synchronized. Therefore, in the case of the bGM 120 in active mode, the end stations can continue to operate seamlessly with or without an operational pSync message 170, just as if the pGM 110 may be down.

[0032] In passive mode, the bGM 120 cannot send bSync messages 180 if the pGM 110 is operating. During pGM failure operation using a passive bGM 120, the bGM 120 may begin sending bSync messages 180 after a pGM 110 timeout. The timeout period for which the bGM 120 waits for a pSync message 170 may be shorter than a holdover period configured at the end stations. After the timeout period, the bGM 120 may begin sending bSync messages 180 to the end stations. The end stations may operate in a hold until bGM synchronization is achieved using the bSync messages 180. Therefore, in the case of the bGM 120 in passive mode, the end stations continue to operate seamlessly with or without an operating pGM 110.

[0033] Upon detection of a pSync failure, such as a failure of the pGM 110 or a network connection, the bGM 120 assumes the role of the current pGM 110. Furthermore, a new backup grandmaster device can be selected by triggering a bGM selection such as a BMCA in response to detecting the failure and transitioning to the bGM 120. The new bGM can be added to the alternate grandmaster device list. If the list is ordered according to the relative precedence of the bGMs, the newly selected bGM can either be appended to the list or inserted at an earlier position in the list. The selection can be performed based on a grandmaster selection algorithm such as the BMCA. Other algorithms can also be used for the selection.The clock selection algorithm, such as the BMCA, may be unchanged if the potential new backup master clock devices have clock sources within the predetermined tolerance relative to the bGM 120. Furthermore, the clock selection may not be performed if a secondary backup master clock device has already been preselected, or if there are multiple master clock devices actively transmitting a respective synchronization message over the network.

[0034] The pGM 110 and the bGM 120 may operate in a network supporting two or more time-sensitive applications. In such cases, the network may support more than one independent time domain that is independent of each other. For example, a first application in the network may depend on a precision clock timing clock, which may be derived from a GPS clock, where leap-second (or leap-microsecond) corrections are desirable or essential for the first application to function. At the same time, a second application operating in the network may depend on precision repeat cycles, where leap-second corrections may be non-essential or even undesirable. Therefore, the two applications may be part of independent time domains and may derive the corresponding clocks from independent clock sources.Each independent time domain can have respective pGM and bGM devices as described throughout this document. Therefore, for each independent time domain support in the network, respective pGM and bGM devices can be duplicated to achieve fault tolerance. Dual time domains could be extended to cover overlapping pSync and bSync timing paths.

[0035] For simultaneous, redundant GM recovery, the failure can be corrected and the original pGM 110 restored, or a new pGM 110 can be inserted into the network. Due to the failure of the pSync messages 170 as well as the failure of the pGM 110, the bGM 120 can take over responsibility from the pGM and therefore be considered the current pGM. The current pGM (original bGM 120) may have drifted away from the clock reference of the original pGM 110. During a pSync 170 restoration, to maintain synchronization between the GM clock sources, the pGM 110 can first be synchronized with the current pGM (original bGM 120). Once synchronized, the original pGM 110 can be reinstated as the pGM and can begin sending pSync 170 messages. The pGM 110 can stabilize the clock after synchronization before sending the pSync messages 170.The original bGM 120 (current pGM) may, upon detecting the recovered pSync messages 170, reinstate itself as the standby GM and may synchronize its clock with the pGM 110 using the pSync messages 170. The bGM 120 may optionally stop sending the bSync messages 180 while the switchover occurs. In this optional case, the bGM 120 may enter an initiation state with respect to the newly discovered and / or recovered pGM. In passive mode, the bGM 120 may stop sending the bSync 180 after a pGM detection and a predetermined timeout. The pGM 110 may move the synchronized clock to a reference clock source, as discussed elsewhere in this disclosure. The pGM 110 can transition the clock to the primary reference by gradually applying the time difference to avoid non-continuous steps in the synchronized time.The transition can therefore be performed in small steps within the predetermined tolerance. The predetermined tolerance can be based on the system's jitter migration tolerance. The predetermined tolerance can be limited by the system's MTIE.

[0036] Alternatively, upon recovery from the pSync failure, the current pGM (original bGM 120) may continue to act as the pGM. The original pGM 110 may take over as the current bGM upon recovery from the pSync failure. The current bGM (original pGM 110) may coordinate with the current pGM (original bGM 120) upon recovery as discussed above. The current bGM (original pGM 110) may, in the active mode, send the bSync message 180 along with the pSync message 170 from the current pGM (original bGM 120). Alternatively, in the passive mode, the current bGM (original pGM 110) may send the bSync message 180 in the case where the pSync message 170 from the current pGM (original bGM 120) is not received for the hold time.The change of roles between the pGM and the bGM may not affect the synchronization of the respective local clocks at the end stations 130. The end stations 130 may continue to derive their respective local clocks based on the synchronization messages received from the current primary and backup master clock devices.

[0037] Operation of an end station, such as an Ethernet station or other IEEE 802.1AS and IEEE 1588-capable network node, may include operation during entry into a holdover period when no synchronization message is received. The holdover period may accommodate clock tolerance values ​​such as the predetermined tolerance during a free run, such as when no synchronization message is received within an expected period. An end station may not distinguish pSync from bSync, and one or both pSync and bSync may be used for synchronization and / or tuning to derive a local clock. For reliability and protection, three or more synchronization messages may be received at the end stations from three or more large base clock sources.A weighted selection such as a simple majority, a weighted majority, or other calculation may be performed to qualify a synchronization message before use. For reliability and protection, each of the synchronization messages may be validated. Validation of the synchronization message, and thereby a clock signal from a particular large base clock source, may be based on a timing difference between successive messages of the clock signal from the particular large base clock source. The timing difference between a current synchronization message and a previous synchronization message from the particular source may be determined. The end station may then ensure that the timing difference lies within an expected predetermined range.Such a validation check may be performed before the synchronization messages from the particular source are used. If an out-of-range time value step is detected over a specific and / or configurable time or sequence, the particular master clock may be considered an unreliable or invalid clock source, and this status may be signaled to a network management entity, such as a network administrator. The network management entity may be an automated system or a network administrator responsible for maintaining the network in an operational state. Alternatively or additionally, an unreliable status of the particular master clock may be reported to a common repository that can receive status updates regarding the network components.Such status updates can be represented in a visual representation of the system. The end station can use the clock ID contained in the messages to determine the source of the message and thereby the clock source of the received clock signal.

[0038] The operations described herein can also be used to migrate the synchronized network from one time domain under one pGM to another time domain under a different pGM. A time domain can be a logical grouping of clocks that synchronize with each other, generally using a protocol such as one of the PTP protocols mentioned above. The clocks in one time domain may not necessarily be synchronized with clocks in another time domain. Time domains provide a way to implement separate sets of clocks that share a common network but maintain independent synchronization within each set. Using the operations described, a time domain can be migrated from one pGM to another pGM. The previous time domain pGM can then enter a sleep state (retire).For example, synchronization messages from pGM' including a time domain different from the time domain of pGM' can be sent to the end stations. Synchronization of a clock of the end station can then be switched from the pGM to the synchronization message from pGM' of a different time domain. The switchover can be performed by gradually applying the time difference between pGM and pGM' to avoid discontinuous steps in the synchronized time. As the switchover process progresses closer to the time domain of pGM', pGM can be put into a sleep state (retired).

[0039] End stations can receive and process multiple synchronization messages from multiple different time domains. The multiple time domains can overlap within a single network, and synchronization messages (pSync, bSync, and other messages) can be carried over any fault-tolerant network path. The network paths can include Ethernet, wireless local area network, coaxial, or power lines. End stations can also perform grandmaster validation functions based on synchronization messages received from a specific grandmaster clock device. The time value difference between consecutive messages or messages over a specific time period can be used for this purpose.The end station may employ algorithms such as N-1 out of N agreement, weighted difference, or any other such algorithms to determine reliability of the timing information received from the particular base clock device.

[0040] The methods and devices described above are applicable to all time-sensitive systems and networks such as Ethernet, Coordinated Shared Networks (CSNs), such as WLAN Coax, Powerline, and other such time-sensitive networks. The master clock devices may be devices specifically configured to provide the clock synchronization messages described throughout the disclosure. The described master clock devices may include one or more processors and one or more non-volatile memory devices. The processors may be responsible for performing the various functions described throughout the disclosure.The large base clock devices may also include network interfaces and corresponding logic and circuitry for transmitting and receiving messages over various communication networks such as those described throughout the disclosure. The large base clock devices may further include a local clock that may be synchronized with other clocks. Additionally, or alternatively, the large base clock may receive a reference clock signal from devices such as GPS.

[0041] The methods, devices, and logic described above may be implemented in many different ways, in many different combinations of hardware, software, or both hardware and software. For example, all parts of the system may include circuitry in a controller, a microprocessor, or an application-specific integrated circuit (ASIC), or may be implemented with discrete logic or components or a combination of other types of analog or digital circuitry, combined on a single integrated circuit, or distributed across multiple integrated circuits.All or part of the logic described above may be implemented as instructions for execution by a processor, controller, or other processing device, and may be stored in a tangible or non-transitory machine-readable or computer-readable medium such as flash memory, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), or other machine-readable medium such as compact disk read-only memory (CDROM), or a magnetic or optical disk. Thus, a product such as a computer program product may include a storage medium and computer-readable instructions stored on the medium that, when executed at an endpoint, computer system, or other device, cause the device to perform operations consistent with all of the foregoing description.

[0042] The processing capability of the system may be distributed across multiple system components such as multiple processors and memories, optionally including multiple distributed processing systems. Parameters, databases, and other data structures may be stored and managed separately, may be integrated into a single memory or database, may be logically and physically organized in many different ways, and may be implemented in many different ways, including data structures such as linked lists, hash tables, or implicit storage mechanisms. Programs may be parts (e.g., subroutines) of a single program, may be distributed across different memories and processors, or may be implemented in many different ways, such as in a library such as a shared library (e.g., a dynamic link library, DLL).For example, the DLL may store code that performs any (all) of the system processing described above.

[0043] The processing capability of the system may be distributed across multiple system components such as multiple processors and memories, optionally including multiple distributed processing systems. Parameters, databases, and other data structures may be stored and managed separately, may be integrated into a single memory or database, may be logically and physically organized in many different ways, and may be implemented in many different ways, including data structures such as linked lists, hash tables, or implicit storage mechanisms. Programs may be parts (e.g., subroutines) of a single program, may be distributed across different memories and processors, or may be implemented in many different ways, such as in a library such as a shared library (e.g., a dynamic link library, DLL).For example, the DLL may store code that performs any (all) of the system processing described above.

[0044] Various implementations have been described in detail. However, many other implementations are also possible.

[0045] A fault-tolerant and redundant grand base clock scheme can reduce or eliminate precision time transitions caused by network link or equipment failure. A primary synchronization message can be sent by a primary grand base clock, and one or more backup synchronization messages can be sent by corresponding backup grand base clocks. The primary and backup grand base clocks can operate concurrently. The primary and backup synchronization messages can be sent to an end station over a network. The end station can derive a local clock based on one, some, or all of the received messages. The end station can distinguish between the messages or not based on the clock source. The end station can validate messages received from a specific clock source.

Claims

[1] System (100) with: a primary master clock device (110); and a reserve main clock device (120); wherein the primary grandmaster clock device (110) and the backup grandmaster clock device (120) are selected from an alternative grandmaster device list comprising a list of multiple grandmaster clock devices (110, 120) in a network (190) by means of a selection process based on an order of priority, a clock quality, a traceability of the clock source and / or a position of the clock device (110, 120) with respect to the network (190); an interface (112) configured to communicate on the network (190) and to receive a primary synchronization message (170), the primary synchronization message (170) being received from the primary master clock device (110) via the network (190), the primary synchronization message (170) comprising a clock signal from the primary master clock device (110); a processor configured to keep the clock of the backup master clock device (120) substantially synchronous with the clock signal received in the primary synchronization message (170); and wherein the processor is further configured to generate a backup synchronization message (180) based on the clock of the backup master clock device (120), and wherein the interface (112) is further configured to transmit the generated backup synchronization message (180) to other devices via the network (190); wherein the processor is further configured to: extract a global identifier from the received primary synchronization message (170); and inserting the global identifier into the generated reserve synchronization message (180) before transmitting the reserve synchronization message (180). [2] The system (100) of claim 1, wherein the backup synchronization message (180) is transmitted to a temporary buffer when the primary synchronization message (170) is received. [3] The system (100) of claim 2, wherein the processor is configured to: detect an absence of receipt of the primary synchronization message (170) from the primary master clock device (110) for a predetermined hold time, and in response, triggering a transmission of the backup synchronization message (180) for receipt by a network device (130). [4] The system (100) of claim 3, wherein the predetermined hold time after which a transmission of the backup synchronization message (180) to the network device (130) is triggered is shorter than a predetermined hold time of the network device (130). [5] The system (100) of claim 3, wherein the processor is configured to: detecting receipt of the primary synchronization message (170) from the primary master clock device (110) after the predetermined hold time, and in response, interrupt the transmission of the backup synchronization message (180) for receipt by the network device (130). [6] The system (100) of claim 3, wherein in response to the absence of receipt of the primary synchronization message (170) from the primary grand base clock device (110), the processor is configured to initiate the backup grand base clock device (120) as a new primary grand base clock device (110) of the network (190) and the primary grand base clock device (110) as a new backup grand base clock device (120) of the network (190). [7] The system (100) of claim 1, wherein the backup synchronization message (180) is transmitted for receipt by a network device (130) independent of receipt of the primary synchronization message (170). [8] The system (100) of claim 1, wherein the processor is further configured to insert an identifier representative of the primary master clock device (110) into the generated backup synchronization message (180) prior to transmission of the backup synchronization message (180). [9] An end station device in a network (190), the end station device (130) comprising: a clock; an interface (112) configured to receive a first clock signal over the network (190) from a first master clock device (110), the interface (112) further configured to receive a second clock signal over the network (190) from a second master clock device (120); wherein the first grandmaster clock device (110) and the second grandmaster clock device (120) are selected from an alternative grandmaster device list comprising a list of multiple grandmaster clock devices (110, 120) in the network (190) by means of a selection process based on an order of priority, a clock quality, a traceability of the clock source, and / or a position of the clock device (110, 120) with respect to the network (190); and a processor configured to adjust the clock of the second master clock device (120) based on the received first clock signal and the received second clock signal; wherein the processor is further configured to generate a backup synchronization message (180) based on the clock of the second master clock device (120), and wherein the interface (112) is further configured to transmit the generated backup synchronization message (180) to other devices via the network (190); wherein the processor is further configured to: extract a global identifier from a received primary synchronization message (170); and inserting the global identifier into the generated reserve synchronization message (180) before transmitting the reserve synchronization message (180). [10] The end station device of claim 9, wherein the processor is configured to adjust the clock at a first frequency, the first frequency being a rate at which the first clock signal and the second clock signal are received. [11] The end station device of claim 10, wherein the processor is configured to adjust the clock at a second frequency, the second frequency being a rate at which the second clock signal is received. [12] The end station device of claim 9, wherein the interface (112) is configured to receive a third clock signal from a third master clock device, and the processor is configured to adjust the clock based on the first, second, and third received clock signals independent of a source of the received clock signals. [13] Terminal device according to claim 9, wherein: the processor is configured to identify a source of a received clock signal based on an identifier in the clock signal representing an identity of the source; and the processor is further configured to validate the clock signal from the source based on a time value difference between consecutive messages of the clock signal from the source. [14] The end station device of claim 13, wherein the processor is configured to indicate to a network management device the source as an unreliable source of a clock signal based on the time value difference being outside a predetermined range. [15] Method comprising the steps: Receiving, at a network device (130), a primary synchronization message (170) from a primary master clock device (110) over a network (190); Receiving, at the network device (130), a backup synchronization message (180) from a backup master clock device (120) over the network (190); and Configuring a local clock at the network device (130) based on the primary synchronization message (170) and the backup synchronization message (180); wherein the primary grandmaster clock device (110) and the backup grandmaster clock device (120) are selected from an alternative grandmaster device list comprising a list of multiple grandmaster clock devices (110, 120) in the network (190) by means of a selection process based on an order of priority, a clock quality, a traceability of the clock source, and / or a position of the clock device (110, 120) relative to the network (190); wherein the backup synchronization message (180) comprises a global identifier extracted from the primary synchronization message (170). [16] The method of claim 15, wherein the local clock is configured at the network device (130) in response to receipt of each of the primary synchronization message (170) and the backup synchronization message (180). [17] The method of claim 15, wherein: the local clock at the network device (130) is configured to be substantially synchronous with a primary reference clock source (350, 360) of the primary master clock device (110) based on the primary synchronization message (170) and the backup synchronization message (180); wherein the primary synchronization message (170) is generated at the primary master clock device (110) according to the primary reference clock source (350, 360); and the backup synchronization message (180) is generated at the backup grandmaster clock device (120) according to a local clock of the backup grandmaster clock device (120), wherein the local clock of the backup grandmaster clock device (120) is aligned with the primary reference clock source (350, 360) based on the primary synchronization message (170). [18] The method of claim 17, wherein in case of a failure to receive the primary synchronization message (170) at the network device (130), the method further comprises: overriding the local clock at the network device (130) to be substantially synchronous with a secondary reference clock source of the backup master clock device (120) based on the backup synchronization message (180); wherein the reserve synchronization message (180) is generated at the reserve master clock device (120) according to the local clock of the reserve master clock device (120); wherein the local clock of the backup master clock device (120) transitions to be synchronous with the secondary reference clock by changing the local clock by a predetermined tolerance value. [19] The method of claim 15, further comprising: identifying, at the network device (130), the backup master clock device (120) as the source of the backup synchronization message (180) based on a clock identifier in the backup synchronization message (180); calculating, at the network device (130), a difference between clock signals in consecutive reserve synchronization messages (180); validating, at the network device (130), the reserve master clock device (120) based on the difference being within a predetermined range; and reporting, by the network device (130), the backup master clock device (120) as an invalid clock source in response to the difference being outside the predetermined range.

Citation Information

Patent Citations

  • hierarchical SYNCHRONIZATION METHOD AND TELECOMMUNICATIONS SYSTEM WITH MESSAGE-BASED SYNCHRONIZATION

    DE69329294T2

  • System and Method of Synchronizing Real Time Clock Values in Arbitrary Distributed Systems

    US20080168182A1

  • Distributed timing control for a distributed digital communication system

    US4736393A