Collaborative testing of devices or sub-segments within a network
By using conventionally deployed devices in time-sensitive networks in industrial plants to generate and transmit test data flows, detect performance metrics and locate problems, the problem of difficulty in effectively testing and locate network problems in the prior art is solved, and the reliability and stability of the network is improved.
Patent Information
- Application Number
- CN202211196758.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2021-10-08
- Filing Date
- 2022-09-29
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2042-09-29
AI Technical Summary
The prior art is difficult to effectively test and locate equipment or sub-segment problems in time-sensitive network TSNs in industrial plants, making it difficult to detect and solve network problems in a timely manner.
By using existing conventionally deployed sending and receiving devices in the network to generate and transmit test data streams, detect performance metrics of the measured device or sub-segment, determine whether it complies with predetermined criteria, and then locate and resolve network problems.
This method can detect multiple network problems, including hardware failures, configuration errors and installation issues, without changing the network hardware and software configuration, improving the reliability and stability of the network during routine operations.
Smart Images

Figure CN115967648B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to testing of devices or sub - segments in a communication network such as a Time - Sensitive Network (TSN) for use in a Distributed Control System (DCS). Background Art
[0002] Monitoring and active control of an industrial plant are based on various information collected within the plant. For example, sensor devices output measurement values, and actuator devices report their operating states. This information can be used by a Distributed Control System (DCS). For example, sensor data can be used as feedback in a control loop that controls the plant so that certain state variables such as temperature or pressure remain at desired set - point values.
[0003] This communication has time urgency. It requires ensuring that any change in the measurement value is conveyed in a timely manner so that the DCS can react to the change. Additionally, any action taken by the DCS needs to be conveyed to the actuators in the plant in a timely manner. A Time - Sensitive Network (TSN) nominally guarantees end - to - end determinism in communication. However, the network may deviate from this nominal state for various reasons, including configuration errors, hardware or wiring faults, or hardware and / or software components that promise a certain performance but fail to deliver.
[0004] WO 2020 / 143 900 A1 discloses a method for handling TSN communication link failures in a TSN network.
[0005] Object of the Invention
[0006] The object of the present invention is to facilitate testing of devices or sub - segments within a network in order to allow localization and correction of the root cause of network problems.
[0007] This object is achieved by a test method according to the independent claims. Further advantageous embodiments are detailed in the dependent claims. Summary of the Invention
[0008] The present invention provides a method for testing a device under test or a sub - segment under test within a network. Specifically, the sub - segment may include multiple devices interconnected by links. The sub - segment may form a continuous domain within the network, which is connected to the rest of the network via certain links. For example, the devices may include end - devices such as computers, sensors, actuators, or controllers in a Distributed Control System, as well as network infrastructure devices such as switches, routers, bridges, or wireless access points.
[0009] During the process of the method, at least one test data stream is generated by at least one sending device. The sending device is a regular participant in the network and is contiguous to the device under test or the sub-segment under test. That is, the sending device exists in the network according to the normal deployment of the network for its intended operation. It is not inserted into the network specifically for testing and / or troubleshooting. The sending device is connected to the device under test or the sub-segment under test via a link of the network. Preferably, the link is a direct link without passing through intermediate devices.
[0010] The sending device sends the test data stream to at least one receiving device different from the sending device. Just like the sending device, the receiving device is a regular participant in the network and is contiguous to the device under test or the sub-segment under test. The test data stream is sent on the path through the device under test or the sub-segment under test. The receiving device determines at least one performance metric of the received test data stream. If there are problems with the network communication between the device under test or the sub-segment under test, it can be expected that this will be visible in at least one performance metric. Examples of performance metrics include throughput, latency, reliability, and jitter. These are among the quantities for which guarantees are given in the context of a TSN network.
[0011] Based on at least one performance metric, it is then determined whether the device under test or the sub-segment under test and / or the link between the device under test or the sub-segment under test and the sending device or the receiving device performs according to at least one predetermined criterion. For example, if one or more performance metrics are all nominal performance metrics, this indicates that everything is working properly everywhere in the path between the sending device and the receiving device. However, if one or more performance metrics are below the nominal value, this indicates that there is a problem somewhere in the path between the sending device and the receiving device.
[0012] The inventors have found that in this way, it is possible to test a device or a sub-segment using only the hardware and links that are already part of the normal deployment of the network for its intended use. Without inserting specific hardware into the network, it can remain completely in the deployment state intended for normal operation. Specifically, the network can be tested in exactly the same hardware, software, and installation configuration and in exactly the same physical environment intended for the normal operation of the network. In this way, more types of problems can be detected. Once these problems are solved, the likelihood of more problems occurring during the normal operation of the network is greatly reduced.
[0013] For example, if the transceiver module is defective and the physical layer just barely outputs the voltage level required for a logical "1" to the physical link, it depends on the exact physical configuration of the link and the device at the other end of the link whether the logical "1" is correctly detected or misdetected as a logical "0". Thus, special test and / or troubleshooting equipment with a high-quality transceiver may still correctly receive all the bits sent by a defective transceiver module. However, the transceivers in the devices as part of the actual deployment may be of poor quality and / or have degraded due to wear, so it may exhibit bit errors. Thus, although the test and / or troubleshooting equipment finds no problems, problems may occur during normal operation. Using regular participants in the network as the sending and receiving devices helps to detect such problems in a timely manner.
[0014] In another example, if there are loose contacts or a defective cable that causes sporadic errors, inserting the test and / or troubleshooting equipment requires physically manipulating the cable, which may cause the problem to disappear temporarily during testing. Thus, the problem may be overlooked and reappear during the normal operation of the network. Using regular participants in the network as the sending and receiving devices effectively avoids physically touching any device or cable of the network.
[0015] In addition, problems may occur during the installation of other execution devices. For example, the wrong type of cable may be selected (e.g., Ethernet Cat-5 instead of Cat-6e), or the required cable grounding may be omitted. The bridge may be misconfigured, and / or the cable may be inserted into the wrong port. All these installation problems can be detected by testing completely in the deployment state to be used for normal operation.
[0016] Moreover, the software configuration of the network is the same as the software configuration for normal operation, so problems caused by a specific software configuration can also be detected. For example, if the devices operate with different software versions, the mutual communication between these devices may be hindered.
[0017] In the main use where the network is configured to carry the network traffic of the distributed control system (DCS) in an industrial plant, it is particularly important to be able to detect a wide variety of problems. If any unexpected problems occur during the normal operation of the network, they may immediately affect the operation of the plant before there is time to correct these problems. Thus, network problems may directly affect the productivity of the plant and even the safety or integrity of the plant.
[0018] In addition, in industrial plants, there are more specific error causes that can only be discovered when tests are performed in the actual deployment state for routine operations. For example, industrial equipment may send electromagnetic crosstalk onto cables, so that transmissions with already marginal physical quality may become garbled. In another example, a network device being close to a heating industrial device may cause the temperature inside the network device to rise too high, so that the processor of the network device gradually reduces to a slower clock speed and the performance of the network device degrades.
[0019] Specifically, a test data stream can be created and sent before the industrial plant starts up and / or during the maintenance of the industrial plant. In this way, if any problems occur during the test, they will not affect the routine operation of the plant. However, a test data stream can also be created and then sent when the equipment under test and / or the equipment of the industrial plant controlled by the sub-segment is not in use and / or is disabled. In this way, more opportunities to perform tests can be created, allowing for periodic retesting to detect any emerging problems (such as the degradation of electronic components).
[0020] In a particularly advantageous embodiment, a plurality of transmitting devices and a plurality of receiving devices are selected that are contiguous to the device under test or the sub-segment under test, such that each device that is a regular participant in the network and contiguous to the device under test or the sub-segment under test assumes the role of a transmitting device and / or a receiving device with respect to at least one test data stream. In this way, multiple paths through the device under test or the sub-segment under test can be investigated, so that all components in the device under test or the sub-segment under test that may potentially cause problems are in at least one investigation path. If the performance metrics on all investigation paths are nominal performance metrics, the likelihood of an unexpected problem occurring during the routine operation of the network is very low. Figuratively speaking, the investigation of one path from the transmitting device to the receiving device corresponds to an X-ray image of the device under test or the sub-segment under test taken from one angle, and images taken from multiple angles are needed to fully evaluate the state of the device under test or the sub-segment under test. Through the organized efforts of the neighbors of the device under test or the sub-segment under test, all of this can be automatically performed in a distributed and collaborative manner without any manual intervention.
[0021] Therefore, in another particularly advantageous embodiment, the performance metrics from the plurality of receiving devices are aggregated. This aggregation can be performed on one or more receiving devices and / or on a centralized network management or monitoring application entity. Based on the aggregated performance metrics, it is then determined whether the device, the substation under test, or the link is operating according to at least one predetermined criterion.
[0022] In a particularly advantageous embodiment, the network is selected to be deterministic with respect to resource allocation and / or traffic scheduling. For example, the network can be a TSN network. Then, at least one performance metric and / or at least one predefined criterion can be determined at least in part based on the resource allocation and / or traffic scheduling of the network. For example, in a deployment state ready for expected use of the network, throughput and / or transfer time can be configured for different types of network traffic. For each type of simulated traffic used as a test data stream, it can then be reasonably expected that the performance metric corresponds to the content that has been allocated to that specific type. For example, information regarding resource allocation and / or traffic scheduling can be obtained from a centralized network configuration CNC entity and / or a centralized user configuration CUC entity of the TSN network.
[0023] In another advantageous embodiment, at least one performance metric and / or at least one predefined criterion is determined at least in part based on the nominal performance specification of the device or sub-segment under test and / or the nominal performance specification of the link between the device or sub-segment under test and the sending or receiving device. This allows for the detection of underperforming devices or links. For example, as discussed above, a device may be prevented from outputting its full performance because its internal temperature is too high and its processor needs to gradually reduce the clock speed. However, it is also possible that the device manufacturer deliberately deceives and labels the device with a higher performance than it can actually output, hoping that the customer will not attempt to fully utilize the advertised performance and thus will not notice that he has been deceived.
[0024] In another particularly advantageous embodiment, at least one test data stream is determined at least in part based on one or more of the following:
[0025] · The traffic patterns expected during normal operation of the network;
[0026] · The network configuration; and
[0027] · The traffic scheduling of the network.
[0028] Specifically, mimicking the traffic patterns expected during normal operation of the network causes the scenario being tested to more realistically correspond to the expected normal operation. Specifically, some problems may only occur during certain phases of high load. For example, a high traffic volume may cause buffers, queues, address tables, or other system resources of network infrastructure devices to become exhausted.
[0029] In addition, the network configuration and / or traffic scheduling may have largely given away where in the device or sub-segment under test what kind of how much traffic is expected. Using this information to generate test data streams can make the test as close as possible to the scenario intended for normal operation.
[0030] Devices or sub - segments in a network can be tested one by one so that the testing gradually covers more and more of the network until all relevant entities have been tested. However, the method can also be used to gradually narrow down the scope of testing until the root cause of the problem is found.
[0031] Thus, in another advantageous embodiment, in response to determining that a tested sub - segment does not perform according to at least one predetermined criterion, at least one device belonging to the tested sub - segment and / or a subset of the tested sub - segment is determined as a new tested device or a new tested sub - segment. Then, the method is repeated for the new tested device or the new tested sub - segment. In this way, the root cause of the poor performance of the original tested sub - segment can be narrowed down.
[0032] The method can be implemented in whole or in part by a computer. Accordingly, the present invention also relates to one or more computer programs having machine - readable instructions that, when executed on one or more computers and / or computing instances, cause one or more computers to perform the method. In this context, virtualization platforms, hardware controllers, network infrastructure devices (such as switches, bridges, routers, or wireless access points) capable of executing machine - readable instructions, and end - devices in the network (such as sensors, actuators, or other industrial field devices) are also to be regarded as computers. Specifically, the network can be embodied as a software - defined network running on general - purpose hardware.
[0033] Accordingly, the present invention also relates to a non - transitory storage medium and / or a download product having one or more computer programs. A download product is a product that can be sold in an online store for immediate implementation by downloading. The present invention also provides one or more computers and / or computing instances having one or more computer programs and / or one or more non - transitory machine - readable storage media and / or download products. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] Hereinafter, the present invention is illustrated with reference to the drawings, which are not intended to limit the scope of the present invention. The drawings show:
[0035] Figure 1 is an exemplary embodiment of a test method 100 for a tested device or a tested sub - segment 4 within a network 1;
[0036] Figure 2 is an exemplary application of the method 100 to a tested device in a TSN network. DETAILED DESCRIPTION
[0037] Figure 1 is a schematic flowchart of an exemplary embodiment of a test method 100 for a tested device or a tested sub - segment 4 within a network 1. As Figure 2 is further detailed, the network 1 includes TSN bridges 2a to 2e and TSN end - stations 3a to 3d.
[0038] In step 105, network 1 can be selected to be deterministic with respect to resource allocation and / or traffic scheduling.
[0039] In step 106, network 1 can be the network 1 configured to carry the network traffic of the distributed control system (DCS) in the industrial plant.
[0040] In step 107, network 1 can be selected to be in a deployment state intended for regular operation.
[0041] In step 110, at least one sending device 5 generates at least one test data stream 6. The sending device 5 is a regular participant of network 1 and is adjacent to the device under test or the sub-segment under test 4.
[0042] In step 120, the sending device sends the test data stream 6 across the device under test or the sub-segment under test 4 to at least one receiving device 7. The receiving device 7 is another regular participant of network 1 different from the sending device 5. It is also adjacent to the device under test or the sub-segment under test 4.
[0043] In step 130, the receiving device 7 determines at least one performance metric 8 of the received test data stream 6.
[0044] In step 140, based on at least one performance metric 8, it is determined whether the device under test or the sub-segment under test 4 and / or the link between the device under test or the sub-segment under test 4 and the sending device 5 or the receiving device 6 complies with at least one predetermined criterion 9. This determination can include, for example, a numerical score or a binary determination. If the determination is a binary determination, the result may be "OK" or "not OK" (NOK).
[0045] In step 150, it is checked whether the determination is "OK". If this is not the case (truth value 0), then in step 160, at least one of the devices 2a to 2e, 3a to 3d belonging to the sub-segment under test 4 and / or the subset 4' of the sub-segment under test 4 is determined to be a new device under test or a new sub-segment under test 4. In step 170, then, for the new device under test or the new sub-segment under test 4, method 100 is repeated to narrow down the root cause of the poor performance in the original sub-segment under test 4.
[0046] According to block 111, multiple sending devices 5 can be selected; and according to block 121, multiple receiving devices 7 adjacent to the device under test or the sub-segment under test 4 are selected. Options are made such that each TSN bridge 2a to 2e and each TSN end station 3a to 3d that is a regular participant of network 1 and is adjacent to the device under test or the sub-segment under test 4 assumes the role of the sending device 5 and / or the receiving device 7 with respect to at least one test data stream 6.
[0047] According to block 112, at least one test data stream can be determined based at least in part on one or more of the following.
[0048] · The traffic patterns expected during normal operation of Network 1;
[0049] · The configuration of Network 1; and
[0050] · The traffic scheduling of Network 1.
[0051] According to block 113, if the network carries control traffic in an industrial plant, test data stream 6 can be generated (and / or sent)
[0052] · Before the industrial plant is started; and / or
[0053] · During maintenance of the industrial plant; and / or
[0054] · When the equipment of the industrial plant controlled by the device under test and / or sub-segment (4) is not in use and / or is disabled.
[0055] According to block 131, at least one performance metric 8 can include one or more of the following: throughput, latency, reliability, and jitter.
[0056] According to block 141, if Network 1 is deterministic with respect to resource allocation and / or traffic scheduling, it can be determined based at least in part on the resource allocation and / or traffic scheduling of Network 1.
[0057] According to block 142, at least one performance metric 8 and at least one predetermined criterion 9 can be determined based at least in part on the nominal performance specifications of the device under test or the tested sub-segment 4 and / or the nominal performance specifications of the link between the device under test or the tested sub-segment 4 and the sending device 5 or the receiving device 7.
[0058] According to block 143, performance metrics 8 can be aggregated from multiple receiving devices 7. According to block 144, it can then be determined whether the device, the tested sub-segment 4, or the link is operating according to at least one predetermined criterion 9 based on the aggregated performance metrics 8.
[0059] Figure 2 An exemplary TSN network 1 is shown, which includes five TSN bridges 2a to 2c and four TSN end stations 3a to 3d. In Figure 2 the example shown, TSN bridge 2c is the device under test 4. TSN bridges 2a, 2b, and 2d and TSN end station 3b are adjacent to the device under test 4. That is, they are neighbors of the device under test 4, and these neighbors are connected to the device under test 4 by direct links. The adjacent devices 2a, 2b, 2d, and 3b take turns acting as the sending device 5 and the receiving device 7 to transmit five different test data streams 6, 6a to 6e over the path through the device under test 4.
[0060] The first test data stream 6a is sent by the TSN bridge 2a acting as the sending device 5 to the TSN bridge 2d acting as the receiving device 7. The second test data stream 6b is sent by the TSN bridge 2b acting as the sending device 5 to the same TSN bridge 2d as the receiving device 7. The TSN bridge 2d, in its role as the receiving device 7, determines that both the test data streams 6a and 6b have been received with acceptable performance (OK).
[0061] Compared with the first test data stream 6a, the third test data stream exchanges the roles of the TSN bridges 2a and 2d. That is to say, the TSN bridge 2d now acts as the sending device 5, while the TSN bridge 2a now acts as the receiving device 7. Therefore, the TSN bridge 2a determines that the test data stream 6c has been received with acceptable performance (OK).
[0062] The fourth test data stream 6d is sent from the TSN end station 3b acting as the sending device 5 to the TSN bridge 2b acting as the receiving device 7. The TSN bridge 2b determines that this test data stream 6d is not received with acceptable performance (NOK).
[0063] In contrast, if the roles are exchanged again for the fifth test data stream 6e (the TSN bridge 2b is the sending device 5 and the TSN end station 3b is the receiving device 7), then the TSN end station 3b determines that this fifth test data stream 6e is received with acceptable performance.
[0064] The final result of the test is that there is a problem with the device under test 4, but it only affects one communication direction between the TSN end station 3b and the TSN bridge 2b. A possible root cause of this problem is a configuration error on the TSN bridge 2c acting as the device under test 4, which causes this TSN bridge 2c to reject the TSN end station 3b from sending data streams to the network 1, such as incorrect or missing entries in the schedule or access control whitelist.
[0065] List of Reference Signs
[0066] 1 Network
[0067] 2a to 2e TSN bridges as devices in Network 1
[0068] 3a to 3d TSN end stations as devices in Network 1
[0069] 4 Device under test or sub - segment under test
[0070] 4' Sub - set of the device under test or sub - segment under test 4
[0071] 5 Sending device
[0072] 6, 6a to 6e Test data streams
[0073] 7 Receiving device
[0074] 8 Performance metric
[0075] 9 Predetermined criterion for performance metric
[0076] 100 Test method 4 for the device under test or the sub - segment under test
[0077] 110 Generate test data stream 6
[0078] 111 Select multiple transmitting devices 5 for cooperation
[0079] 112 Create test data stream 6 based on specific information
[0080] 113 Generate data stream 6 at a specific time point
[0081] 120 Send test data stream 6 to receiving device 7
[0082] 121 Select multiple receiving devices 7 for cooperation
[0083] 130 Determine performance metric 8
[0084] 131 Examples of performance metric
[0085] 140 Check performance metric 8 using predetermined criterion 9
[0086] 141 Determine criterion 9 based on specific information
[0087] 142 Use the nominal performance specifications of device 4
[0088] 143 Aggregate performance metric 8
[0089] 144 Check the aggregated metric 8 using predetermined criterion 9
[0090] 150 Determine whether the performance is acceptable
[0091] 160 Determine a new device under test or a new sub - segment under test 4
[0092] 170 Repeat the test 4 for the new device under test or the new sub - segment under test
Claims
1. A method (100) for testing a device under test or a sub - segment under test (4) within a network (1), comprising the following steps: · Generating (110) at least one test data stream (6) by at least one transmitting device (5), the at least one transmitting device (5) being a regular participant of the network (1) and being adjacent to the device under test or the sub - segment under test (4); · Transmitting (120) the test data stream (6) to at least one receiving device (7) by the transmitting device (5) on a path passing through the device under test or the sub - segment under test (4), the at least one receiving device (7) being a regular participant of the network (1) different from the transmitting device (5) and also being adjacent to the device under test or the sub - segment under test (4); · Determining (130) at least one performance metric (8) of the received test data stream (6) by the receiving device (7); · Determining (140) whether the device under test or the sub - segment under test (4) and / or the link between the device under test, or the sub - segment under test (4) and the transmitting device (5) or receiving device (7) is operating according to at least one predetermined criterion (9) based on the at least one performance metric (8); wherein a plurality of transmitting devices (5) and a plurality of receiving devices (7) adjacent to the device under test or the sub - segment under test (4) are selected (111, 121) such that each device (2a to 2e, 3a to 3d) that is a regular participant of the network (1) and is adjacent to the device under test or the sub - segment under test (4) assumes the role of the transmitting device (5) and / or the receiving device (7) with respect to at least one test data stream (6).
2. The method (100) according to claim 1, wherein the network (1) is selected (105) to be deterministic with respect to resource allocation and / or traffic scheduling.
3. The method (100) according to claim 2, wherein the at least one performance metric (8) and / or the at least one predetermined criterion (9) is determined (141) at least in part based on the resource allocation and / or traffic scheduling of the network (1).
4. The method (100) according to any one of claims 1 to 3, wherein the at least one performance metric (8) and / or the at least one predetermined criterion (9) is determined (142) at least in part based on the nominal performance specification of the device under test or the sub - segment under test (4) and / or the nominal performance specification of the link between the device under test or the sub - segment under test (4) and the transmitting device (5) or receiving device (7).
5. The method (100) according to any one of claims 1 to 3, wherein at least one test data stream (6) is determined (112) at least in part based on one or more of the following: · The traffic pattern expected during the normal operation of the network (1); · The configuration of the network (1); and · The traffic scheduling of the network (1).
6. The method (100) according to any one of claims 1 to 3, wherein the network (1) is selected (106) as a network (1) in an industrial plant configured to carry network traffic of a distributed control system DCS.
7. The method (100) according to claim 6, wherein the test data stream (6) is generated (113) and / or sent at the following times: · Before the start-up of the industrial plant; and / or · During the maintenance of the industrial plant; and / or · When the equipment in the industrial plant controlled by the device under test and / or the sub-segment (4) is not in use and / or is out of service.
8. The method (100) according to any one of claims 1 to 3, wherein the network (1) is selected (107) to be in a deployment state intended for normal operation.
9. The method (100) according to any one of claims 1 to 3, wherein the at least one performance metric (8) comprises (131) one or more of the following: throughput, latency, reliability, and jitter.
10. The method (100) according to any one of claims 1 to 3, further comprises: · Aggregating (143) performance metrics (8) from a plurality of receiving devices (7); and · Determining (144) whether the device under test, the sub-segment under test (4), or the link is performing according to the at least one predetermined criterion (9) based on the aggregated performance metrics (8).
11. The method (100) according to any one of claims 1 to 3, further comprises: In response to determining (150) that the sub-segment under test (4) is not performing according to the at least one predetermined criterion (9), · Determining (160) at least one device (2a to 2e, 3a to 3d) belonging to the sub-segment under test (4) and / or a subset (4') of the sub-segment under test (4) as a new device under test or a new sub-segment under test (4); and · Repeating (170) the method (100) for the new device under test or the new sub-segment under test (4) in order to narrow down the root cause of poor performance in the original sub-segment under test (4).
12. A computer program product comprising machine-readable instructions which, when executed on one or more computers, cause the one or more computers to perform the method (100) according to any one of claims 1 to 11.
13. A non-transitory storage medium having the computer program product according to claim 12.
14. A download product having the computer program product according to claim 12.
15. A computer having the computer program product according to claim 12 and / or having the non-transitory storage medium according to claim 13 and / or having the download product according to claim 14.
Citation Information
Patent Citations
Failure handling of a TSN communication link
WO2020143900A1
Data path evaluation system and method
US20020183952A1
Intent based network data path tracing and instant diagnostics
US20200162589A1