Method, system and computer-readable storage medium for performing network performance analysis

By grouping the nodes and switches of high-performance computing systems to measure the delay and comparing them with the expected value, identifying the part of the network that causes excessive delay, solving the problem of identifying network problems and improving system performance.

CN117955877BActive Publication Date: 2025-08-08HEWLETT PACKARD ENTERPRISE DEV LP
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311073115.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2022-10-28
Filing Date
2023-08-24
Publication Date
2025-08-08
Estimated Expiration
2043-08-24

AI Technical Summary

Technical Problem

In high-performance computing systems, it is difficult to identify the location of the problem in the network, especially whether there are problems that lead to low performance at the application level, library level or hardware level.

Method used

By grouping multiple nodes and switches of the computing system, the delays of different levels and communication routes are measured, the measurement results are compared with the expected delay, and the part that causes excessive delays are identified.

Benefits of technology

It realizes precise positioning of network performance issues and promotes performance improvement and troubleshooting of computing systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117955877B_ABST
    Figure CN117955877B_ABST
Patent Text Reader

Abstract

The present disclosure relates to performing network performance analysis. A method for performing network performance analysis, the method comprising measuring the delay of a plurality of data packets transmitted through a network, comprising determining delay representations for a plurality of layers of the network, for a plurality of communication routes, and / or for a plurality of communication types. The delay representations comprise delay measurement results, statistical representations of the delay measurement results, and / or delay metrics derived from the delay measurement results. The method comprises comparing the determined delay representations with expected delay representations, the expected delay representations comprising expected delays, expected statistical representations of the delays, and / or expected delay metrics. Based on the comparison, the method comprises determining a difference between an expected delay representation in the expected delay representations and a delay representation in the determined delay representations; and based on the determined difference, identifying a layer of the network, one of the communication routes, and / or one of the communication types as a portion of the network causing excessive delay.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to a method, medium, and system for performing network performance analysis. Background Art

[0002] With the recent surge in high-performance computing (HPC) systems (e.g., Exascale computing systems), extracting every bit of performance from every component of the system (e.g., compute, memory, network, and storage) is more important than ever to ensure that real-world problems, such as computing, memory, and storage, can be effectively solved. Maximizing system performance and co-designing software to extract every bit of performance from all system components is crucial to providing the full functionality of HPC systems. Meeting the performance promises of HPC applications is also crucial and is an important part of delivering and gaining acceptance for Exascale systems. A detailed understanding of the performance achieved by each component of an HPC system is the first step to maximizing system performance. Summary of the Invention

[0003] A method for performing network performance analysis on a computing system comprising a plurality of networked nodes, the method comprising: measuring, by a node among the nodes of the computing system, a latency of a plurality of data packets transmitted by the node through a network, wherein the plurality of nodes and switches are divided into a plurality of groups, each group comprising one or more nodes and one or more switches; determining latency representations for a plurality of layers of the network and three different types of communication routes, wherein the latency representations comprise latency measurement results, statistical representations of the latency measurement results, and latency metrics derived from the latency measurement results, wherein the plurality of layers of the network comprise an application layer, a library layer, and a hardware layer, wherein the three different types of communication routes comprise a first type of communication route between two nodes in the same group and coupled to the same switch, a second type of communication route between two nodes in the same group and coupled to different switches, and a third type of communication route between two nodes in different groups and coupled to different switches, and wherein determining the delay metric comprises determining a first average delay for the first type of communication route and a second average delay for the second type of communication route; comparing the determined delay representation with an expected delay representation, the expected delay representation comprising an expected delay, an expected statistical representation of the delay, and / or an expected delay metric; determining a difference between an expected delay representation in the expected delay representations and a determined delay representation in the determined delay representations based on the comparison; and identifying one of the three different types of communication routes as a portion of the network causing excessive delay based on the determined difference.

[0004] A non-transitory computer-readable storage medium having a set of computer-executable instructions stored thereon, the instructions being configured to cause one or more processors to perform the following operations: obtaining, from a node of a computing system comprising a plurality of networked nodes, a plurality of switches, latency measurements of a plurality of data packets transmitted by the node through the network, wherein the plurality of nodes and switches are divided into a plurality of groups, each group comprising one or more nodes and one or more switches; determining latency representations for a plurality of layers of the network and three different types of communication routes, wherein the latency representations comprise the latency measurements, statistical representations of the latency measurements, and latency metrics derived from the latency measurements, wherein the plurality of layers of the network comprise an application layer, a library layer, and a hardware layer, and wherein the three different types of communication routes are included in the same group and coupled a first type of communication route between two nodes to the same switch, a second type of communication route between two nodes in the same group and coupled to different switches, and a third type of communication route between two nodes in different groups and coupled to different switches, and wherein determining the delay metric comprises determining a first average delay for the first type of communication route and a second average delay for the second type of communication route; comparing the determined delay representation with an expected delay representation, the expected delay representation comprising an expected delay, an expected statistical representation of the delay, and / or an expected delay metric; determining a difference between an expected delay representation in the expected delay representations and a delay representation in the determined delay representations based on the comparison; and identifying one of the three different types of communication routes as a portion of the network causing excessive delay based on the determined difference.

[0005] A system for performing network performance analysis on a computing system including a plurality of networked nodes, the system comprising: one or more processors configured to: obtain latency measurement results of a plurality of data packets transmitted by a node of the computing system through a network, wherein the plurality of nodes and switches are divided into a plurality of groups, each group comprising one or more nodes and one or more switches; determine latency representations for a plurality of layers and three different types of communication routes of the network and / or for a plurality of communication types, wherein the latency representations comprise the latency measurement results, statistical representations of the latency measurement results, and latency metrics derived from the latency measurement results, wherein the plurality of layers of the network comprise an application layer, a library layer, and a hardware layer, and wherein the three different types of communication routes are included in the same group. and a first type of communication route between two nodes that are in the same group and coupled to the same switch, a second type of communication route between two nodes that are in the same group and coupled to different switches, and a third type of communication route between two nodes that are in different groups and coupled to different switches, and wherein determining the delay metric comprises determining a first average delay for the communication routes of the first type and a second average delay for the communication routes of the second type; comparing the determined delay representation with an expected delay representation, the expected delay representation comprising an expected delay, an expected statistical representation of the delay and / or an expected delay metric; determining a difference between an expected delay representation in the expected delay representations and a delay representation in the determined delay representations based on the comparison; and identifying one of the three different types of communication routes as a portion of the network that causes excessive delay based on the determined difference. BRIEF DESCRIPTION OF THE DRAWINGS

[0006] Various examples will be described below with reference to the following drawings.

[0007] Figure 1A is a block diagram of an example computing system according to some implementations of the present disclosure.

[0008] Figure 1B is a block diagram of an example computing system illustrating a first example communication, according to some implementations of the present disclosure.

[0009] Figure 1C is a block diagram of an example computing system illustrating a second example communication, according to some implementations of the present disclosure.

[0010] Figure 1D is a block diagram of an example computing system illustrating a third example communication, according to some implementations of the present disclosure.

[0011] Figure 2is a diagram illustrating an example histogram that may be used for network performance analysis according to an example embodiment of the present disclosure.

[0012] Figure 3 is a flow chart showing an illustrative method for conducting network performance analysis according to an example embodiment of the present disclosure.

[0013] Figure 4 is a block diagram illustrating an example computing system according to some implementations of the present disclosure.

[0014] Throughout the drawings, the same reference numerals designate similar, but not necessarily identical, elements. In addition, the drawings provide examples and / or implementations consistent with the description; however, the description is not limited to the examples and / or implementations provided in the drawings. DETAILED DESCRIPTION

[0015] Multi-node computing systems such as high-performance computing (HPC) systems include multiple computing nodes interconnected by a network. In such systems, the overall performance of the entire system depends in part on the performance of the network that interconnects the nodes. In addition, as each generation of computing nodes in such systems becomes faster and the amount of data shared by these nodes increases, it is more important that the network that interconnects the nodes has high performance to achieve maximum system performance and avoid the network from becoming a bottleneck in the system. However, the high-performance network of such an HPC system can be composed of many different elements (including hardware and software), which can interact in various ways. For example, each node in the system can include a network interface element (e.g., a network interface card (NIC)), an application that communicates through the network, and an interface library layer that facilitates the application to communicate through the network. These can be referred to as different "layers" of the network (e.g., application layer, library layer, and hardware layer). In addition, the network can include various networking devices (e.g., switches) outside the node, which connect the nodes in various topological structures. For example, in some HPC systems, nodes can be arranged in multiple groups (e.g., 32 nodes per group in some example systems), each group including multiple switches (e.g., four switches in some example systems) to interconnect the nodes of these groups. The switches of one group can also be connected to the switches of other groups to interconnect these groups. These external networking devices are also part of the network hardware layer described above. Therefore, because all of these elements participate in communication on the network at various different layers of the network, it is difficult to fully understand how each element contributes to the overall performance of the network. In particular, it is difficult for system or system administrators or users to identify where problems may exist in the network. For example, if the overall performance of the entire system is lower than expected, it is difficult to know whether this is due to problems in the network, and if so, it is difficult to know whether the problem is at the application level, the library level, or the hardware level.

[0016] Thus, the present disclosure provides for performing end-to-end network performance analysis that can identify problem areas of a network by decomposing performance latency measurements based on different hardware and software components and comparing (e.g., by a processor) the latency measurements, statistical representations of the latency measurements, and / or metrics derived from the latency measurements or statistical representations with expected latency / statistical representations / metrics. Based on this comparison, it can be determined where (e.g., at which layer or layers of the network, for which communication routes, and / or for which communication types) higher-than-expected latency is occurring. In some embodiments, latency measurements can be determined at multiple different network layers (e.g., application layer, library application layer, and hardware layer), for multiple different types of communication routes, and / or for multiple different types of communications. For example, examples of these different communication routes used in a dragonfly network include a route between two nodes in the same group and connected to the same switch (SGSS), a route between two nodes in the same group and connected to different switches (SGDS), and a route between two nodes in different groups and connected to different switches (DGDS), or other types of routes within the network. Examples of different types of communications include GET / read communications, PUT / write communications, or other types of communications. Latency measurements, statistical representations thereof (e.g., mean, mode, median, histogram, etc.), and / or metrics derived therefrom can be compared to expected values. These expected values can be predetermined, configurable, and / or dynamically determined (e.g., based on measured latency or metrics derived therefrom).

[0017] Based on the comparison, a difference between an expected value and a measured latency (or statistical representation or derived metric) can be determined for at least one of a network level, a communication route, and / or a communication type. For example, a difference can be determined when the difference between two latency measurements / statistical representations / metrics compared to each other exceeds a threshold or a given percentage, when the latency measurements / statistical representations / metrics are more than a certain number of standard deviations apart, or when there is any other predetermined or dynamically determined difference between them. In response to determining a difference in a network level, a communication route, and / or a communication type, the network level, route, and / or communication type can be identified (e.g., by a processor) as a problem that potentially causes excessive latency, and further troubleshooting can therefore be focused on that potentially problematic portion of the system. Thus, the examples disclosed herein can allow improved performance in computing systems such as HPC systems by facilitating the identification (and subsequent correction) of network problems that may impair system performance. Such identification and correction can be performed by the system's processor and / or can be output for user analysis.

[0018] For example, in one analysis mode, average delays can be determined for multiple combinations of network levels, communication routes, and communication types. Intermediate metrics can then be determined based on these average delays. For example, another average delay value for the same network level, communication route, and communication type but a different communication route can be subtracted from the average delay value for one network level, communication route, and communication type, and the resulting value can be a metric that identifies the delay difference due to the communication route difference. This same type of metric can be determined for multiple different communication types at the same network level, and these metrics can be compared with each other, with one derived metric value being used as an "expected" value relative to another derived metric, and any differences therebetween can be identified. Other metrics (in addition to or in place of the aforementioned examples) can be derived from the measured delays, and these metrics can be compared with each other, with some of these metrics being used as expected values compared to other values.

[0019] As another example, in the second analysis mode, a statistical distribution of delay measurements can be determined for a given network layer, communication route, and / or communication type, and the statistical distribution can be compared with an expected distribution to identify differences. For example, in some embodiments, the statistical distribution includes a histogram formed by the determined delay measurements grouped into a plurality of bars for a given network layer and communication type. The histogram of the determined delay measurements can be compared with the expected histogram. Based on the comparison, the difference between the expected distribution and the determined distribution can be determined. For example, the number and / or position of peaks in the distribution can be compared with the expected number and / or expected position of peaks in the expected distribution, and any differences can be identified. Based on the location of these differences (e.g., which network layer, communication route, and / or communication type these difference distributions correspond to), one of the network layer, communication route, and / or communication type can be identified as problematic (i.e., the portion of the network that causes excessive delay).

[0020] The following detailed description refers to the accompanying drawings. Whenever possible, the same reference numerals are used in the drawings and the following description to refer to the same or similar parts. However, it should be expressly understood that these drawings are for illustration and description purposes only. Although several examples are described in this document, modifications, adaptations, and other embodiments are possible. Therefore, the following description does not limit the disclosed examples. Instead, the proper scope of the disclosed examples may be defined by the appended claims.

[0021] Figures 1A to 1D FIG1 is a block diagram of a computing system 100 configured for performing end-to-end network performance analysis according to some embodiments of the present disclosure. The computing system 100 includes a plurality of nodes 110. In FIG1 , two nodes are illustrated, namely, node A 110 and node B 110. Figures 1B to 1D, twelve nodes 110 are illustrated. A person of ordinary skill in the art will understand that Figures 1A to 1D The number of nodes shown in is not intended to be limiting, and computing systems having any number of nodes are contemplated within the scope of embodiments of the present disclosure. By way of example and not limitation, each node 110 may include a computing device. Figure 1A As shown, each node 110 includes a network interface element (e.g., a network interface card (NIC)) 114, 116, respectively, which is configured to connect the corresponding node 110 to a network 118. Node A 110 is further illustrated as including an application layer 120, which is configured to contain applications that communicate over the network 118, and a library layer 122, which is configured to facilitate the communication of the applications over the network 118. (It should be understood that although not illustrated with respect to node B 110, it is within the scope of embodiments of the present disclosure that node B may also include an application layer and a library layer.)

[0022] Computing system 100 also includes at least one switch 124 (and in most systems, multiple switches 124) configured to connect node A 110 and node B 110. Network 118 includes switch(es) 124 and cables or other interconnects extending between the nodes and switches 124. For the purposes of analyzing network performance, network 118 may also be considered to include application layer 120, library layer 122, and NICs 114 and 116, as these layers all participate in the delivery of packets across network 118 and may have an impact on the performance of network 118. For example, in one embodiment, system 100 is an HPC system (e.g., an HPE Cray EX system, an HPE Cray supercomputer, or an HPC cluster), switch 124 comprises an HPC switch (e.g., an HPE Slingshot switch), NIC 114 comprises an HPC NIC (e.g., a Cassini NIC) configured for use with switch 124, library layer 122 comprises an interface (e.g., a Cassini Xascale Interface (CXI)) configured for use with NIC 114, and application layer 120 comprises one or more HPC applications (e.g., a Symmetric Hierarchical Memory (SHMEM) application).

[0023] In some examples, the nodes may be arranged in a plurality of groups 190, each group 190 including a plurality of switches 124, such as Figures 1B to 1DFor example, in one embodiment, each group 190 has 32 nodes 110 and four switches 124. Any number of groups 190 may be provided in the system 100, any number of switches 124 may be provided per group 190, and any number of nodes 110 may be provided per switch 124 (constrained only by the capabilities of the hardware involved and the desired topology and performance of the system 100).

[0024] Each node 110 is connected to at least one of the switches 124 of its group 190, and the switches 124 of one group 190 are connected to the switches 124 of another group 190. Thus, if node A 110 and node B 110 are part of the same group and connected to the same switch 124, the communication between these nodes may be classified herein as SGSS. For example, Figure 1B One such SGSS communication is illustrated between node A:1 100 and node A:2 100, both of which are connected to the same switch A:1 124 of the same group A 190. This communication does not require any switch hops (i.e., transfers from one switch 124 to another switch 124) and can therefore be expected to have relatively low latency.

[0025] If node A 110 and node B 110 are part of the same group and are connected to different switches 124 of the group, the communication between these nodes may be classified herein as SGDS. Figure 1C One such SGDS communication is illustrated between node A:1 100 and node A:6 100, which are respectively connected to switch A:1 124 and switch A:2 124 of the same group A 190. This communication requires one local switch hop (i.e., a transfer from one switch 124 to another switch 124 in the same group) and can therefore be expected to have slightly higher latency.

[0026] If node A 110 and node B 110 are part of different groups and are connected to different switches 124 of their respective groups, then the communication between these nodes may be classified herein as DGDS. Figure 1DOne such DGDS communication is illustrated between node A:1 100 and node A:6 100, which are connected to switch A:1 124 and switch B:1 124 of group A 190 and group B 190, respectively. This communication requires at least one global switch hop (i.e., a transfer from one switch 124 to another switch 124 in a different group 190) and can therefore be expected to have the highest latency of the three routes described above. In particular, in some embodiments, the various groups 190 can correspond to physically (and in some cases, logically) separate locations, such as different chassis, different racks and / or rack rows, or in some cases even different geographic locations (e.g., different data centers). As a result, in some embodiments, global switch hops can be expected to have slightly higher latency than intra-group switch hops.

[0027] The present disclosure provides a framework configured to measure performance through packet latency at different layers of a network, making it possible to examine the contribution of each network layer to the overall achieved latency. Figures 1A to 1D In computing system 100, application layer 120 of node A 110 represents application layer 126 of network 118, library layer 122 of node A 110 represents library layer 128 of network 118, and the combination of NIC 114 of node A 110, switch 124, and NIC 116 of node B 110 represents hardware layer 130 of network 118. Each layer 126, 128, and 130 of network 118 may include one or more tools or utilities associated therewith for measuring latency at that layer. For example, application layer 126 may have a tool for measuring latency. An example of such a tool is a microbenchmark. Library layer 128 may have a tool for measuring latency at that layer. An example is a Cassini eXascale Interface (CXI) utility used in the Cassini eXascale Interface library. The NIC may have hardware counters for measuring precise latency at hardware layer 130. It should be understood that while some tools can measure latency directly, other tools can make measurements that require the application of mathematical functions to convert them into latency measurements. Any and all such measurements and / or conversions, and any combination thereof, are contemplated as being within the scope of embodiments of the present disclosure.

[0028] In some examples, once the latency is measured, a statistical representation of the latency measurement can be determined, and / or other metrics can be derived from the latency measurement. The statistical representation can include a mean, variance, distribution (e.g., a histogram), or other ways of statistically aggregating or representing the measured delay values. Metrics derived from the latency measurements can include values derived from the latency measurements that quantitatively represent aspects of system performance. For example, a metric can represent the contribution of the application layer to the total latency and can be derived from the measured latency by subtracting the delay measured at the library layer from the delay measured at the application layer. A similar metric can be derived to represent the contribution of the library layer to the total latency by subtracting the measured hardware level delay from the measured library level delay. These are just two examples of metrics, and any number of other metrics can be derived from measured values, such as by adding or subtracting or otherwise performing mathematical operations on the measured values.

[0029] Once latency measurements are obtained and / or derived from tools at the various network layers 126, 128, and 130, the measurements, statistical representations of the measurements, and / or metrics derived from the measurements or statistical representations may be compared (e.g., by a processor of the computing system 100) to expected latency measurements / statistical representations / metrics. By way of example and not limitation, the expected latency measurements / statistical representations / metrics may include predetermined thresholds, ranges, or representations stored in a memory of the computing system 100 and may be based on specifications associated with the appropriate network layer. In some examples, in addition to or in lieu of predetermined expected values, expected values may be determined dynamically, such as by deriving expected values from the measurements themselves—e.g., some measured values or metrics derived therefrom may be used as expected values to compare to other measured values or metrics, as discussed in more detail below. Based on a comparison of the latency measurement results / statistical representations / metrics with the expected latency measurement results / representations / metrics, it can be determined (e.g., by a processor of the computing system 100) where (e.g., at which layer or layers of the network, for which communication routes, and / or for which communication types) the higher-than-expected latency occurs.

[0030] In one example, average latency (i.e., the average latency for multiple packets or a given number of packets routed through a particular network layer within a given time interval) can be determined for multiple combinations of network layer, communication route, and communication type. Table 1 illustrates examples of such average latency in the rows titled "Software Latency," "Library Latency," and "Hardware Latency."

[0031] Table 1

[0032]

[0033] In the illustrated example, looking at the Software Latency row, we find that the average latency at the application level for GET / read communications for the SGSS communication route is 2.79 μs, the average latency at the application level for GET / read communications for the SGDS communication route is 3.46 μs, the average latency at the application level for PUT / write communications for the SGSS communication route is 2.53 μs, and the average latency at the application level for PUT / write communications for the SGDS communication route is 3.73 μs. Now looking at the Repository Latency row, we find that the average latency at the repository level for GET / read communications for the SGSS communication route is 2.62 μs, the average latency at the repository level for GET / read communications for the SGDS communication route is 3.33 μs, the average latency at the repository level for GET / read communications for the SGDS communication route is 2.02 μs, and the average latency at the repository level for PUT / write communications for the SGSS communication route is 2.63 μs. And finally, looking at the hardware latency row, we found that the average latency at the hardware level for GET / read communication and for SGSS communication routing was 1.6μs, the average latency at the hardware level for GET / read communication and for SGDS communication routing was 2.25μs, the average latency at the hardware level for PUT / write communication and for SGSS communication routing was 1.18μs, and the average latency at the hardware level for PUT / write communication and for SGDS communication routing was 1.78μs.

[0034] Intermediate metrics can then be determined based on these average delays. Such intermediate metrics are represented in the remaining rows of Table 1. For example, for a given network layer and communication type, the average delay value for one communication route can be subtracted from the average delay values for different communication routes. The resulting value can be a metric that identifies the delay differences that are attributable to the communication route differences. This is illustrated in the shaded rows of Table 1. For example, the Variation in Software Latency w / r / t SGSS row indicates that the application-level delay difference for GET / read communications between the SGSS communication route and the SGDS communication route is 0.67 μs (i.e., 3.46 μs - 2.79 μs), and the application-level delay difference for PUT / write communications between the SGSS communication route and the SGDS communication route is 1.2 μs (i.e., 3.73 μs - 2.53 μs). In addition, the Variation of Repository Latency w / r / t SGSS row indicates that the difference in repository-level latency for GET / read communications between the SGSS communication route and the SGDS communication route is 0.71 μs (i.e., 3.33 μs - 2.62 μs), and the difference in repository-level latency for PUT / write communications between the SGSS communication route and the SGDS communication route is 0.61 μs (i.e., 2.63 μs - 2.02 μs). Finally, the Variation of Hardware Latency w / r / t SGSS row indicates that the difference in hardware-level latency for GET / read communications between the SGSS communication route and the SGDS communication route is 0.65 μs (i.e., 2.25 μs - 1.6 μs), and the difference in hardware-level latency for PUT / write communications between the SGSS communication route and the SGDS communication route is 0.60 μs (i.e., 1.78 μs - 1.18 μs). These metrics may be referred to herein simply as SGSS to SGDS metrics, with the associated level and / or communication type appended to identify the specific metric, such as a repository SGSS to SGDS GET metric.

[0035] These derived intermediate metrics can be compared to each other, with one derived metric value used as the "expected" value relative to the other derived metrics, and any discrepancies between them can be identified. In some embodiments, the derived metric value used as the "expected" value can be the value representing the lowest latency. For example, the SGSS-to-SGDS metrics associated with the differences observed between SGDS and SGSS configurations at different levels of measurement can be compared, and several observations can be made. First, for GET / read communications, the SGSS-to-SGDS metrics are roughly the same (i.e., for GET / read measurements at all three levels, there is roughly the same latency difference between the SGSS communication route and the SGDS communication route (approximately 0.65 μs)). Second, for PUT / write communications, the SGSS-to-SGDS metrics are not exactly the same: there is a discrepancy between the SGSS-to-SGDS PUT metric at the application layer and the SGSS-to-SGDS metrics at the other two layers (1.2 μs versus 0.6 μs and 0.61 μs). This suggests that the application layer may be the portion of the network causing excessive latency, and that PUT / write communications, in particular, may be a possible source of application-layer issues.

[0036] In some examples, system 100 may determine that a difference exists between two measurements or metrics when they are more than a predetermined distance apart, such as measured by actual value, percentage, number of standard deviations, and the like. Specifically, in some examples, a difference may be defined as any difference of 20% or more between a value and an expected value. In other examples, a difference may be defined as any difference of 40% or more between a value and an expected value. In other examples, a difference may be defined as any difference of 50% or more between a value and an expected value. In other examples, a difference may be defined as any difference exceeding one standard deviation. In other examples, a difference may be defined as any difference of 0.1 μs or more between a value and an expected value. In other examples, a difference may be defined as any difference of 0.2 μs or more between a value and an expected value. In some examples, the aforementioned definitions of difference may be combined, such that a difference is detected if any one of a set of definitions is met.

[0037] More specifically, in the example above, the SGSS-to-SGDS PUT metrics at the library and / or hardware level were used as "expected" values compared to the SGSS-to-SGDS PUT metrics at the application level, and when compared, a discrepancy was identified. Because the discrepancy occurred at the application level, this suggests a possible issue at the application level. In other words, changing from SGSS to SGDS is expected to increase latency somewhat due to the addition of another switch hop. This change is expected to result in a roughly even increase in latency across all layers (since adding a new switch hop should affect all layers equally), and indeed, this is what is seen for the GET communication and both PUT communications: an increase of approximately 0.6μs is observed for these metrics. However, for the application-level PUT communication, the changes to SGSS and SGDS resulted in nearly double the expected additional latency (1.2μs). This doubling does not occur at other layers, and this indicates a possible application-level-specific issue. Furthermore, the fact that the same problem is not seen with application-layer GET communications, but is seen with application-layer PUT communications, indicates that the problem is not occurring at the entire application layer but is specific to application-layer PUT communications (e.g., there may be a problem with how the application layer handles PUT communications).

[0038] In some examples, the system 100 is preconfigured to use certain metrics as "expected" values compared to other metrics. In other examples, the system 100 (or in some cases, the user) can dynamically determine which values should be used as "expected" values. For example, certain measurements or metrics can be grouped together based on the fact that they are expected to be similar to each other (i.e., within a predetermined threshold of each other). These groupings can be inferred based on knowledge of the system, its network topology, the components used in the system, and any other relevant data. For example, two communications expected to travel the same distance in the network may be expected to have similar delays, and therefore the delay measurements or metrics derived therefrom can be grouped together. The "expected" value for comparison with each measurement / metric in the group can then be determined to be one of the values (or range of values) in the group, such as the lowest value in the group or the most common or consistent value in the group of measurements / metrics (e.g., where, for this purpose, values within a predetermined threshold of each other are considered to be the same). For example, in the example scenario above, the SGSS to SGDS metric constitutes a set of metrics that are expected to have similar values, and the most consistent value of the SGSS to SGDS metric is approximately 0.6 μs, and therefore this can be determined as the "expected" value for comparison with all other SGSS to SGDS metrics. As another example, when two metrics are expected to be similar (e.g., less than a predetermined distance apart, as measured by actual values, percentages, number of standard deviations, etc.) but are not similar, the lower one of the metrics can be considered the "expected" value of the two metrics (e.g., because it represents the lowest latency, which is generally preferred). For example, in the example scenario, the SGSS to SGDS metrics for PUT communications at the application level and the library level can be expected to be similar to each other, and therefore the lower value (in this case, the library-level SGSS to SGDSPUT metric) can be used as the "expected" value for comparison. In other examples, the user can specify which values are to be used as "expected" values relative to other values.

[0039] It should be noted that the metrics and comparisons described with respect to the example scenario above are not the only comparisons that can be made in that scenario, and the same insights and / or additional insights can be gained by comparing other metrics. For example, in some examples, additional metrics can be derived and compared. For example, metrics that indicate the overhead attributable to different network layers can be determined by the system. For the example scenario, these metrics are shown in the last two rows of Table 1 above. As the network layers build upon each other, the measurements made at each layer indicate not only the latency of that layer, but also the latency of lower layers of the network. Thus, as Figure 1AAs shown, measurements made at the application layer include application-level latency in addition to library-level and hardware-level latency. To arrive at a latency measurement that is attributable only to a specific layer of the network, the latency measurement of the next lowest layer in the communication stack can be subtracted from it (because the next lowest layer includes not only the latency attributable to that layer, but also the latency of all other lower layers of the network). Therefore, the application-level overhead for GET / read communication and SGSS communication routing is 0.17μs (i.e., 2.79μs-2.62μs), the application-level overhead for GET read communication and SGDS communication routing is 0.13μs (i.e., 3.46μs-3.33μs), the application-level overhead for PUT / write communication and SGSS communication routing is 0.51μs (i.e., 2.53μs-2.02μs), and the application-level overhead for PUT / write communication and SGDS communication routing is 1.1μs (3.73μs-2.63μs). The library-level overhead for SGSS communication routing is 1.02 μs (i.e., 2.62 μs - 1.6 μs) for GET / read communication, 1.08 μs (i.e., 3.33 μs - 2.25 μs) for GET read communication, 0.84 μs (i.e., 2.02 μs - 1.18 μs) for PUT / write communication, and 0.85 μs (2.63 μs - 1.78 μs) for SGDS communication routing for PUT / write communication. The latency values measured at the hardware level also represent the hardware-level overhead because the hardware level is the lowest level of the network communication stack.

[0040] These overhead metrics can also be compared. In the example scenario above, comparing the overhead metrics confirmed the observation that the portion of the supporting network causing excessive latency is likely the application-level overhead for PUT / write communications, as the overhead at this level for the SGDS communication route and PUT / write communications (i.e., 1.1 μs) is significantly higher than the application-level overhead for SGSS (i.e., 0.51 μs) (i.e., almost twice as high, which may indicate unacceptable latency, as all packets are generally expected to follow the same path and therefore have approximately the same latency). Significantly higher here means higher by an amount that exceeds the thresholds described above for determining differences. More specifically, the application overhead is expected to be different between GET and PUT communications because they have different processing requirements, but the application overhead is expected to be constant when switching between SGSS and SGDS communication routes because the routes traversed by the communications should not affect the application's overhead. Therefore, the application overhead metric for SGSS communications can be used as an "expected" value compared to SGDS communications at the same level. When this was performed for GET communications, no differences were detected (approximately the same overhead of 0.13 μs and 0.17 μs was found). On the other hand, for PUT communications, the application overhead of SGDS (1.1 μs) is much higher than that of SGSS (0.51 μs), and this difference indicates a problem with how SGDS communications are handled. Furthermore, since this difference between SGDS and SGSS overhead occurs only at the application layer and only for PUT communications, this indicates that the problem may lie in how the application layer handles PUT communications routed through SGDS.

[0041] The examples above are described with respect to specific connection topologies, where communications can have zero switch hops (SGSS), 1 intra-group switch hop (SGDS), or 1 global (inter-group) switch hop. However, in other examples, the connection topology may be different. For example, another system topology is a so-called "fat tree" topology. In such a fat tree topology, communications can take 0, 2, 4, 6, or more switch hops, depending on the size of the system. In other examples, other connection topologies may be used. Regardless of the specific connection topology used, similar types of analysis as described above can be applied. For example, for any given connection topology, the type of communication route expected can be determined, including, for example, its relative distance, the number and / or type of switch hops involved, and other similar factors, and based on this understanding, the delay variation between routes and / or the overhead attributable to each network layer can be determined.

[0042] Thus, the above example illustrates one manner in which performing end-to-end network performance analysis can allow problem areas of a network to be identified by breaking down performance latency measurements based on different hardware and software components, comparing the latency measurements and / or metrics derived from the latency measurements to expected latency measurements / metrics, and determining where higher than expected latency is occurring based on the comparison.

[0043] As another example, a statistical distribution of latency measurements for a given network layer, communication route, and / or communication type can be determined, and the statistical distribution can be compared to an expected distribution to identify differences. For example, in some embodiments, the statistical distribution includes a histogram of latency measurements determined for a given network layer, communication type, and communication route type (e.g., SGSS, SGDS, or DGDS), where the latency is grouped into a plurality of bars. This is in Figure 2 2 shows a histogram 210 of SGSS communication routes, a histogram 212 of SGDS communication routes, and a histogram 214 of DGDS communication routes for GET / read traffic, using values from a hypothetical scenario.

[0044] The histogram of the determined latency measurements can be compared to an expected histogram. The expected histogram can be an actual statistical representation (not shown) or it can be an expectation based on what is known about the subject network. For example, in the illustrated embodiment, it is known that data packets are expected to have latencies within a specific range (e.g., approximately 1600 nanoseconds) at the hardware level and that these data packets are expected to take the shortest possible communication route, and therefore, it can be expected that these measurements will result in a histogram with a single peak at the expected latency measurement (i.e., indicating that most data packets follow the same, shortest route). The expected latency can be known based on the hardware specifications or can be derived from the measurements. Given a known expected latency and a known number of peaks, an expected histogram can be created with peaks at the expected values. It should be noted that for this analysis, what is compared is the number and / or location of the peaks and therefore the size of the peaks in the expected histogram can be set arbitrarily. A peak can be any bin in the histogram with a count greater than a threshold, which can be a fixed value or a percentage value (eg, 10% or more of the total counts observed), both of which can be configurable or predefined. Figure 2 Some example expected histograms 215 , 216 , and 217 are illustrated in dashed lines for comparison with histograms 210 , 212 , and 214 , respectively.

[0045] Based on the comparison, differences between the expected distribution and the determined distribution can be determined, and based on where the differences are located (e.g., which network layer, communication route, and / or communication type the different distributions correspond to), one of the network layers, communication routes, and / or communication types can be identified as problematic (i.e., the portion of the network that causes excessive latency). Figure 2 In the illustrated scenario, the expected unimodal behavior is shown to be occurring for histograms 210 and 212, and therefore these histograms 210 and 212 match their expected histograms 215 and 216 (in terms of the number and location of peaks). However, for histogram 214 (DGDS communication routing), there are three peaks compared to the one peak in the expected histogram 217, and therefore a discrepancy is detected. This discrepancy indicates that for the DGDS GET communication, the packets are taking different routes of different lengths. This is contrary to expected behavior and indicates that the hardware may be misconfigured or experiencing other issues. Therefore, the hardware layer can be identified as the portion of the network that is causing the excessive latency for the DGDS communication routing. Therefore, Figure 2 Another way in which performing network performance analysis may allow problem areas of a network to be identified is illustrated.

[0046] The examples above are described with respect to particular connection topologies where communications may have zero switch hops (SGSS), 1 intra-group switch hop (SGDS), or 1 global (inter-group) switch hop, but in other examples, the connection topology may be different. For example, in a fat tree topology, communications may take 0, 2, 4, 6, or more switch hops, depending on the size of the system. In other examples, other connection topologies may be used. Similar types of analysis as described above may be applied regardless of the particular connection topology used. For example, for any given connection topology, the type of communication route expected may be determined, including, for example, its relative distance, the number and / or type of switch hops involved, and other similar factors, and based on this understanding, the number and / or location of peaks in the histogram may be predicted.

[0047] Figure 3 FIG3 is a flow chart illustrating an example method 300 for performing network performance analysis according to an embodiment of the present disclosure. Method 300 may be performed, for example, by computing system 100 of FIG1 . At step 310 , the latency of a plurality of packets transmitted over a network may be measured. By way of example only, packet latency may be measured using a computing system application, library, or hardware device-specific tool and / or utility.

[0048] At step 312, latency representations may be determined for multiple layers of the network (e.g., application layer, library layer, and / or hardware layer), for multiple communication routes (e.g., SGSS, SGDS, and / or DGDS), and / or for multiple communication types (e.g., GET / read and / or PUT / write communications). The latency representations may include the latency measurements themselves, statistical representations of the latency measurements (e.g., a histogram), and / or latency metrics derived from the latency measurements (e.g., resulting in measurement differences in latency attributable to a particular network layer, communication route, and / or communication type).

[0049] At step 314, the determined delay representation can be compared with the expected delay representation. The expected delay representation can include the expected delay, the expected statistical representation of the delay, and / or the expected delay metric. These expected delay representations can be predetermined, configurable, and / or dynamically determined. Based on the comparison, at step 316, a difference between one of the expected delay representations and one of the determined delay representations can be determined.

[0050] At step 318, based on the determined discrepancy, one of the plurality of layers of the network, one of the plurality of communication routes, and / or one of the plurality of communication types may be determined as a portion of the network causing excessive latency. In some embodiments, remedial action may be taken to address and / or correct the discrepancy.

[0051] Figure 4 is a block diagram illustrating an example computing device 400 according to some embodiments of the present disclosure. Figure 4 In the illustrated example, the computing device 400 includes a processing resource 410 coupled to a non-transitory computer-readable storage medium 412 encoded with computer-executable instructions for performing a system setting memory reset. The processing resource 410 may include a microcontroller, a microprocessor, a central processing unit core(s), a graphics processing unit core(s), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), and / or another hardware device suitable for retrieving and / or executing instructions from the computer-readable storage medium 412 to perform functions associated with the various examples described herein. Additionally or alternatively, the processing resource 410 may include electronic circuitry for performing the functions of the instructions described herein.

[0052] Computer readable storage medium 412 may be any medium suitable for storing executable instructions. Non-limiting examples of computer readable storage medium 412 include RAM, ROM, EEPROM, flash memory, hard drive, optical disk, etc. Computer readable storage medium 412 may be provided within computing device 400, such as Figure 4As shown, in this case, the executable instructions can be considered to be "installed" or "embedded" on the computing device 400. Alternatively, the computer-readable storage medium 412 can be a portable (e.g., external) storage medium and can be part of an "installation package." The instructions stored on the computer-readable storage medium 412 can be used to implement at least the methods described herein (e.g., Figure 3 ).

[0053] In the context of this example, computer readable storage medium 412 is encoded with a set of executable instructions 414-424. It should be understood that some or all executable instructions and / or electronic circuits included in one block may be included in different blocks shown in the figures or in different blocks not shown in alternative implementations.

[0054] When executed, the instructions 414 may cause the processing resource 410 to measure (or otherwise obtain) the latency of a plurality of data packets transmitted over the network. By way of example only, the packet latency may be measured / obtained using a computing system application, library, or hardware device-specific tool and / or utility.

[0055] The instructions 416, when executed, may cause the processing resource 410 to determine (a plurality of) latency representations for multiple layers of a network, for multiple communication routes, and / or for multiple communication types. These latency representations may include latency measurements themselves, statistical representations of the latency measurements (e.g., a histogram), and / or latency metrics derived from the latency measurements (e.g., resulting in measurement differences in latency attributable to a particular network layer, communication route, and / or communication type).

[0056] The instructions 418, when executed, may cause the processing resource 410 to compare the determined latency representation(s) with the expected latency representation(s). These expected latency representations may include expected latency(s), expected statistical representation(s) of the latency(s), and / or expected latency metric(s).

[0057] The instructions 420, when executed, may cause the processing resource 410 to determine a difference between one of the expected latency representation(s) and one of the determined latency representations based on the comparison.

[0058] The instructions 422, when executed, may cause the processing resource 410 to identify a tier among multiple tiers of the network, a communication route among multiple communication routes, and / or a communication type among multiple communication types as a portion of the network causing excessive latency based on the determined difference.

[0059] When executed, the instructions 424 may cause the processing resource 410 to perform troubleshooting or take remedial measures focused on the identified portion causing excessive latency. For example, the processing resource 410 may cause a communication to be sent instructing a user to change one or more configuration settings of the identified portion of the network causing excessive latency.

[0060] This article refers to Figures 1 to Figure 4 The described processes may be implemented in the form of executable instructions stored on a machine-readable medium and executed by a processing resource (e.g., a microcontroller, a microprocessor, (multiple) central processing unit cores, an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), etc.) and / or in the form of other types of electronic circuitry. For example, the processes may be performed by one or more computing systems or nodes of various forms, as described above with reference to Figures 1A to 1D Described system.

[0061] The techniques described herein include various steps, examples of which are described above. As further described above, these steps can be performed by hardware components or can be embodied in machine-executable instructions that can be used to perform these steps on a processor programmed with the instructions. Alternatively, at least some of the steps can be performed by a combination of hardware, software, and / or firmware.

[0062] The technology described herein can be provided as a computer program product, which can include a tangible machine-readable storage medium having instructions embodied thereon, which can be used to program a computer (or other electronic device) to perform a process. The machine-readable medium can include, but is not limited to, a fixed (hard) drive, a magnetic tape, a floppy disk, an optical disk, a compact disk read-only memory (CD-ROM), a magneto-optical disk, a semiconductor memory such as a ROM, a PROM, a random access memory (RAM), a programmable read-only memory (PROM), an erasable PROM (EPROM), an electrically erasable PROM (EEPROM), a flash memory, a magnetic or optical card, or other types of media / machine-readable media suitable for storing electronic instructions (e.g., computer programming code, such as software or firmware).

[0063] In the technical description herein, numerous specific details are set forth to provide a thorough understanding of the exemplary embodiments. However, it will be apparent to those skilled in the art that the embodiments described herein can be practiced without these specific details. In other instances, well-known structures and devices are shown in block diagram form.

[0064] The terms used herein are for the purpose of describing example embodiments only and are not intended to be limiting. Unless the context clearly indicates otherwise, the singular forms "a", "an" and "the" as used herein are intended to also include the plural forms. As used herein, the term "multiple" is defined as two or more. The term "and / or" as used herein refers to and covers any and all possible combinations of one or more items in the associated enumerated items. As used herein, the term "includes" means including but not limited to, and the term "including" means including but not limited to. The term "based on" means at least partially based on. If the specification stipulates that a certain component or feature "may / can / could / might" be included or have a certain characteristic, then the specific component or feature does not need to be included or have that characteristic. As used in the specification herein, unless the context clearly indicates otherwise, the meaning of "in..." includes "in..." and "on...".

[0065] The various methods described herein may be performed by combining one or more machine-readable storage media containing code according to the example embodiments described herein with appropriate standard computer hardware to execute the code contained therein. Apparatus for practicing the various embodiments described herein may involve one or more computing elements or computers (or one or more processors within a single computer) and a storage system containing or having network access to computer programs encoded according to the various methods described herein, and the method steps of the various embodiments described herein may be implemented by modules, routines, subroutines, or sub-portions of a computer program product.

[0066] In the foregoing description, numerous details have been set forth to provide an understanding of the subject matter disclosed herein. However, embodiments may be practiced without some or all of these details. Other embodiments may include modifications and variations of the details discussed above. The appended claims are intended to cover such modifications and variations.

Claims

1. A method for performing network performance analysis on a computing system comprising a plurality of nodes networked together, the method comprising: measuring, by a node among the nodes of the computing system, a latency of a plurality of data packets transmitted by the node through a network, wherein the plurality of nodes and switches are divided into a plurality of groups, each group including one or more nodes and one or more switches; determining latency representations for a plurality of layers of the network and three different types of communication routes, wherein the latency representations include latency measurements, statistical representations of the latency measurements, and latency metrics derived from the latency measurements, wherein the plurality of layers of the network include an application layer, a library layer, and a hardware layer, wherein the three different types of communication routes include a first type of communication route between two nodes in the same group and coupled to a same switch, a second type of communication route between two nodes in the same group and coupled to different switches, and a third type of communication route between two nodes in different groups and coupled to different switches, and wherein determining the latency metrics includes determining a first average latency for the first type of communication routes and a second average latency for the second type of communication routes; comparing the determined representation of delay to a representation of expected delay, the representation of expected delay comprising an expected delay, an expected statistical representation of the delay, and / or an expected delay metric; determining a difference between one of the expected delay representations and one of the determined delay representations based on the comparison; and Based on the determined differences, one of the three different types of communication routes is identified as a portion of the network causing excessive latency.

2. The method according to claim 1, wherein Determining the latency metric further includes determining a difference between the first average latency and a second average latency and determining a latency attributable to a tier in the plurality of tiers of the network.

3. The method according to claim 1, wherein The delay representation includes a statistical distribution of delay values.

4. The method according to claim 3, wherein: The statistical distribution of delay values comprises a histogram.

5. The method according to claim 3, wherein: The statistical distribution corresponds to a corresponding type of communication route.

6. The method of claim 1, further comprising: Troubleshoot or take remedial action focused on the identified portion of the network causing the excessive latency.

7. A non-transitory computer-readable storage medium having stored thereon a set of computer-executable instructions for causing one or more processors to perform the following operations: Obtaining latency measurements of a plurality of data packets transmitted by a node of a computing system comprising a plurality of nodes networked together through a network via a plurality of switches, wherein: The plurality of nodes and switches are divided into a plurality of groups, each group including one or more nodes and one or more switches; determining latency representations for a plurality of layers of the network and three different types of communication routes, wherein the latency representations include the latency measurements, statistical representations of the latency measurements, and latency metrics derived from the latency measurements, wherein the plurality of layers of the network include an application layer, a library layer, and a hardware layer, wherein the three different types of communication routes include a first type of communication route between two nodes in the same group and coupled to a same switch, a second type of communication route between two nodes in the same group and coupled to different switches, and a third type of communication route between two nodes in different groups and coupled to different switches, and wherein determining the latency metrics includes determining a first average latency for the first type of communication routes and a second average latency for the second type of communication routes; comparing the determined representation of delay to a representation of expected delay, the representation of expected delay comprising an expected delay, an expected statistical representation of the delay, and / or an expected delay metric; determining a difference between one of the expected delay representations and one of the determined delay representations based on the comparison; and Based on the determined differences, one of the three different types of communication routes is identified as a portion of the network causing excessive latency.

8. The computer-readable storage medium of claim 7, in, Determining the latency metric further includes determining a difference between the first average latency and the second average latency and determining a latency attributable to a tier in the plurality of tiers of the network.

9. The computer-readable storage medium of claim 7, wherein: The delay representation comprises a statistical distribution of delay values, and wherein the statistical distribution of delay values comprises a histogram.

10. The computer-readable storage medium of claim 9, wherein: The statistical distribution corresponds to a corresponding type of communication route.

11. The computer-readable storage medium of claim 10, wherein: The set of computer executable instructions further cause the one or more processors to: Troubleshoot or take remedial action focused on the identified portion of the network causing the excessive latency.

12. A system for performing network performance analysis on a computing system comprising a plurality of nodes networked together, the system comprising: One or more processors configured to: obtaining, from a node among the nodes of the computing system, latency measurements of a plurality of data packets transmitted by the node over a network, wherein the plurality of nodes and switches are divided into a plurality of groups, each group including one or more nodes and one or more switches; determining latency representations for a plurality of layers of the network and three different types of communication routes, wherein the latency representations include the latency measurements, statistical representations of the latency measurements, and latency metrics derived from the latency measurements, wherein the plurality of layers of the network include an application layer, a library layer, and a hardware layer, wherein the three different types of communication routes include a first type of communication route between two nodes in the same group and coupled to a same switch, a second type of communication route between two nodes in the same group and coupled to different switches, and a third type of communication route between two nodes in different groups and coupled to different switches, and wherein determining the latency metrics includes determining a first average latency for the first type of communication routes and a second average latency for the second type of communication routes; comparing the determined representation of delay to a representation of expected delay, the representation of expected delay comprising an expected delay, an expected statistical representation of the delay, and / or an expected delay metric; determining a difference between one of the expected delay representations and one of the determined delay representations based on the comparison; and Based on the determined differences, one of the three different types of communication routes is identified as a portion of the network causing excessive latency.

13. The system of claim 12, in, Determining the latency metric further includes determining a difference between the first average latency and the second average latency and determining a latency attributable to a tier in the plurality of tiers of the network.

14. The system of claim 12, wherein: The delay representation comprises a statistical distribution of delay values, and wherein the statistical distribution of delay values comprises a histogram.

15. The system of claim 14, wherein: The statistical distribution corresponds to a corresponding type of communication route.

Citation Information

Patent Citations

  • Network state analysis method, device and equipment and machine readable storage medium

    CN113507396A

  • Determination and indication of network traffic congestion

    US20190052565A1