Service level target violation diagnosis method and device, equipment and medium
Through probability detection and path coding technology, the problems of low overhead and high accuracy in existing network violation diagnosis methods are solved, and the rapid and accurate location of network failures is achieved, and system overhead and information reporting are reduced.
Patent Information
- Application Number
- CN202510624919.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-15
- Publication Date
- 2025-07-08
AI Technical Summary
Existing network violation diagnosis methods cannot take into account low overhead and high accuracy, and cannot quickly and accurately locate the location and cause of network failures.
Probability detection method is adopted, and packets are compressed and random sampling is used to combine the media access control code for native path encoding. Diagnostic information based on location and details is collected probabilistically, and a threshold-triggered violation reporting mechanism is used for reporting.
It reduces system overhead, improves the accuracy and efficiency of diagnostic information, can quickly locate network violations, reduces unnecessary information reporting, and reduces bandwidth consumption.
Smart Images

Figure CN120281684A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of network measurement technologies, and in particular, to a method, apparatus, device, and medium for diagnosing violations of service level objectives. Background Art
[0002] Network applications have strict requirements for the availability of application programs. Service providers provide users with availability guarantees, namely the so-called Service Level Agreement (SLA). For example, a software vendor promises that the percentage of normal running time of 10% of its users per month is higher than 99%. These commitments are achieved after network topology design and traffic engineering, but the network does not develop as expected. A Service Level Objective (SLO) is a defined objective for performance. An SLO may be that the data stream used by a latency-sensitive application (such as a video game) must ensure that its latency is less than 50 milliseconds. The objective is usually more stringent than the commitment. If the SLO is violated due to certain reasons, this problem can be solved before the SLA is violated.
[0003] In practical applications, a highly available network should not frequently fail. Even if a failure occurs, it should be able to recover quickly. In SLO violations, network administrators hope to exactly know the specific location or details of the network violation and reduce the recovery time to meet their SLA commitments to users. The purpose of SLO violation diagnosis is precisely to detect violations and locate the root cause. The reasons for application anomalies include bandwidth degradation, long-tail latency, and packet loss, etc. Network administrators usually use two types of SLOs: bandwidth SLO and latency SLO. Existing methods cannot balance low overhead and high precision. The time granularity, the coverage of events, and whether it reflects the actual traffic of the application constitute the precision. Summary of the Invention
[0004] In view of this, the present invention provides a method, apparatus, device, and medium for diagnosing violations of service level objectives, which are used to at least partially solve problems such as the inability to balance low overhead and high precision in existing violation diagnosis.
[0005] The first aspect of the present invention provides a method for diagnosing violations of service level objectives, which implements the method for diagnosing violations of service level objectives in the data plane, including: compressing and grouping devices on a network path based on path information to probabilistically collect location-based diagnostic information, where the location-based diagnostic information includes status information of specific points on the network path; probabilistically detecting packet loss port information of all devices on the network path based on the compression and grouping results to obtain detail-based diagnostic information, where the detail-based diagnostic information refers to the specific data required for diagnosing suspicious areas; performing native path encoding according to the media access control codes of each device in the network to determine the target devices corresponding to the violation events included in the location-based diagnostic information and / or the detail-based diagnostic information.
[0006] According to an embodiment of the present invention, compressing and grouping all devices on a network path based on path information to probabilistically collect location-based diagnostic information includes: compressing and grouping the path information into three parts: group ID, group information, and intra-group information, where the group ID represents the sequence number of the sampled group, the intra-group information represents the information of the sampled group, and the group information records all violation groups on the network path; randomly sampling the devices within the group based on the group ID, group information, and intra-group information to probabilistically collect location-based diagnostic information.
[0007] According to an embodiment of the present invention, an approximate algorithm is used to randomly sample the devices within the group.
[0008] According to an embodiment of the present invention, the time period when the data packet enters the network and the path length of the current data packet are used as random numbers, and a hash function is used as the work function to randomly sample the devices within the group.
[0009] According to an embodiment of the present invention, probabilistically detecting packet loss port information of all devices on the network path based on the compression and grouping results to obtain detail-based diagnostic information includes: probabilistically collecting detail-based diagnostic information through variable-length 0 sequences and padding bits.
[0010] According to an embodiment of the present invention, probabilistically collecting detail-based diagnostic information through variable-length 0 sequences and padding bits includes: in the case of detecting packet loss, obtaining the group ID of the current device within the group through telemetry message information; obtaining the value i of the group ID, filling in i 0s to obtain a variable-length 0 sequence, sequentially filling in the port where the data packet enters the network and the port where the data packet leaves the network after the variable-length 0 sequence, and using 0s to fill in the number of bits.
[0011] According to an embodiment of the present invention, performing native path encoding based on the media access control codes of devices in the network to determine the target device corresponding to the violation event included in the location-based diagnosis information and / or the detail-based diagnosis information includes: calculating the end-to-end path ID of all data packets in the network according to the media access control codes of devices in the network; recording the path IDs of different real service paths at each egress node to determine the target device corresponding to the violation event.
[0012] According to an embodiment of the present invention, the method further includes: recovering the violation event included in the location-based diagnosis information and / or the detail-based diagnosis information.
[0013] According to an embodiment of the present invention, the service level objective violation diagnosis method further includes: reporting the violation event by using a threshold-triggered violation reporting mechanism.
[0014] According to an embodiment of the present invention, reporting the violation event by using a threshold-triggered violation reporting mechanism includes: setting a threshold for the number of reports; in response to the number of reports of the violation event corresponding to the current data stream being no greater than the threshold for the number of reports, normally reporting the violation event; in response to the number of reports of the violation event corresponding to the current data stream being greater than the threshold for the number of reports, stopping reporting the violation event and continuously configuring the current data stream as a violation state until it is detected that there is no violation event in the current data stream, and then resuming normal reporting of the violation event.
[0015] According to an embodiment of the present invention, setting the threshold for the number of reports according to the maximum hop count of the network path.
[0016] According to an embodiment of the present invention, reporting the status change information in the initial stage of reporting violations and the initial stage of recovery.
[0017] A second aspect of the present invention provides a service level objective violation diagnosis device, which implements the service level objective violation diagnosis device in the data plane, including: a path information probability compression module, configured to perform compression grouping on devices on the network path based on path information to probabilistically collect location-based diagnosis information, where the location-based diagnosis information includes the status information of specific points on the network path; a detail information probability compression module, configured to probabilistically detect the packet loss port information of all devices on the network path based on the compression grouping result to obtain detail-based diagnosis information, where the detail-based diagnosis information refers to the specific data required for suspicious area diagnosis; a native path encoding module, configured to perform native path encoding according to the media access control codes of devices in the network to determine the target device corresponding to the violation event included in the location-based diagnosis information and / or the detail-based diagnosis information.
[0018] A third aspect of the present invention provides an electronic device, including: one or more processors; a memory for storing one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors are caused to implement the above method.
[0019] A fourth aspect of the present invention provides a computer-readable storage medium having executable instructions stored thereon, and when the instructions are executed by a processor, the processor is caused to implement the method as claimed in the above claims.
[0020] The service level objective violation diagnosis method, device, equipment and medium provided according to the embodiments of the present invention at least include the following beneficial effects:
[0021] By using a probabilistic detection method to detect violation diagnosis information, the overhead cost of the system can be reduced. Classifying and grouping the violation diagnosis information for detection, and probabilistically collecting location-based diagnosis information and detail-based diagnosis information can reduce the overhead of the diagnosis information telemetry head on the basis of ensuring detection accuracy. Determining the target device corresponding to the violation event by performing native path encoding according to the media access control codes of each device in the network can further reduce the overhead of the system.
[0022] Further, during the process of diagnosing the violation location, the path information is compressed and grouped into group ID, group information and intra-group information, and diagnosis is performed in units of groups, so that when the system performs information recovery, it can be determined whether all data has been obtained, ensuring the completeness and knowability of the violation diagnosis.
[0023] Further, based on an approximation algorithm, the time period when the data packet enters the network and the path length of the current data packet are used as random numbers, and a hash function is used as the work function to randomly sample the devices within the group, so that the results of the sampled groups are the same within a time period, and further enabling the service level objective violation diagnosis method to be completed in the network data plane without relying on the control plane capabilities.
[0024] Further, by probabilistically collecting detail-based diagnosis information through variable-length 0 sequences and padding bits, combined with the grouping mechanism, it is possible to prepare complete detection of detail-based violation diagnosis information.
[0025] Furthermore, a threshold-triggered violation reporting mechanism is used to report violation events, reducing the number of reported violation events, and only reporting the state change information at the initial stage of the violation and the initial stage of the recovery, and no longer reporting all detected violations, further reducing the overhead of a specific bandwidth. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] Through the following description of the embodiments of the present invention with reference to the accompanying drawings, the above and other objects, features and advantages of the present invention will become clearer. In the drawings:
[0027] Figure 1 Schematically shows a flowchart of a service level objective violation diagnosis method according to an embodiment of the present invention.
[0028] Figure 2 Schematically shows a flowchart of a location-based diagnostic information collection method according to an embodiment of the present invention.
[0029] Figure 3 Schematically shows a schematic diagram of a location-based diagnostic information collection method according to an embodiment of the present invention.
[0030] Figure 4 Schematically shows a flowchart of a detail-based diagnostic information collection method according to an embodiment of the present invention.
[0031] Figure 5 Schematically shows a schematic diagram of a detail-based diagnostic information collection method according to an embodiment of the present invention.
[0032] Figure 6 Schematically shows a flowchart of a method for determining a violation target device according to an embodiment of the present invention.
[0033] Figure 7 Schematically shows a flowchart of a method for reporting a violation event according to an embodiment of the present invention.
[0034] Figure 8 Schematically shows a comparison experimental result graph of the service level objective violation diagnosis method according to an embodiment of the present invention and the existing method in terms of Static Random-Access Memory (SRAM) consumption.
[0035] Figure 9 Schematically shows a comparison experimental result graph of the service level objective violation diagnosis method according to an embodiment of the present invention and the existing method in terms of violation reporting rate.
[0036] Figure 10 Schematically shows a block diagram of a service level objective violation diagnosis device according to an embodiment of the present invention.
[0037] Figure 11 Schematically shows a detailed diagram of a service level objective violation diagnosis device according to an embodiment of the present invention.
[0038] Figure 12 Schematically shows a block diagram of an electronic device suitable for implementing the method described above according to an embodiment of the present invention. Detailed implementation manner
[0039] To make the objectives, technical solutions and advantages of the present invention more clearly understood, the following further describes the present invention in detail with reference to specific embodiments and the accompanying drawings. Apparently, the described embodiments are only a part rather than all of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the scope of protection of the present invention.
[0040] The terms used herein are only for describing specific embodiments and are not intended to limit the present invention. The terms "including", "comprising", etc. used herein indicate the presence of the described features, steps, operations and / or components, but do not exclude the presence or addition of one or more other features, steps, operations or components.
[0041] In the present invention, unless otherwise clearly defined and limited, terms such as "installed", "connected", "joined", "fixed", etc. shall be construed in a broad sense. For example, it may be a fixed connection, a detachable connection, or integrated; it may be a mechanical connection, an electrical connection or communication with each other; it may be a direct connection or an indirect connection through an intermediate medium, and may be the internal connection of two elements or the interaction relationship between two elements. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.
[0042] In the description of the present invention, it should be understood that the orientation or positional relationship indicated by terms such as "longitudinal", "length", "circumferential", "front", "rear", "left", "right", "top", "bottom", "inner", "outer", etc. is based on the orientation or positional relationship shown in the drawings, and is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the described subsystems or elements must have a specific orientation, be constructed and operated in a specific orientation, and thus should not be construed as a limitation of the present invention.
[0043] Throughout the drawings, the same elements are denoted by the same or similar reference numerals. When it may cause confusion in the understanding of the present invention, conventional structures or configurations will be omitted. Also, the shapes, sizes and positional relationships of the components in the drawings do not reflect the actual sizes, ratios and actual positional relationships. Additionally, in the present invention, any reference signs located between parentheses should not be construed as a limitation of the present invention.
[0044] Similarly, to streamline the present invention and assist in understanding one or more of the various disclosed aspects, in the above description of the exemplary embodiments of the present invention, the various features of the present invention are sometimes grouped together into a single embodiment, figure, or description thereof. Descriptions referring to terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples" mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described may be combined in any one or more embodiments or examples in a suitable manner.
[0045] In addition, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, features defined with "first" and "second" may explicitly or implicitly include one or more of such features. In the description of the present invention, the meaning of "a plurality" is at least two, such as two, three, etc., unless otherwise specifically defined.
[0046] Figure 1 A flowchart of a service level objective violation diagnosis method according to an embodiment of the present invention is schematically shown.
[0047] As Figure 1 shown, the service level objective violation diagnosis method is implemented in the data plane of the switch and may include, for example, operations S101 to S103.
[0048] In operation S101, devices on the network path are compressed for packets based on path information to probabilistically collect location-based diagnostic information.
[0049] In operation S102, based on the compressed packet result, the packet loss port information of all devices on the network path is probabilistically detected to obtain detail-based diagnostic information.
[0050] In operation S103, native path encoding is performed according to the media access control codes of the devices in the network to determine the target devices corresponding to the violation events included in the location-based diagnostic information and / or the detail-based diagnostic information.
[0051] In the embodiment of the present invention, the network SLO violation diagnosis information is classified into location-based diagnostic information and detail-based diagnostic information. The location-based diagnostic information may include the status information of specific points on the network path. For example, whether the queue depth exceeds a threshold. The detail-based diagnostic information refers to the specific data required for suspicious area diagnosis, such as relevant ports or precise measurement values.
[0052] The method for diagnosing service level objective violations according to an embodiment of the present invention uses a probabilistic detection method to detect violation diagnosis information, which can reduce the overhead cost of the system. Classifying and grouping the violation diagnosis information for detection, and probabilistically collecting location-based diagnosis information and detail-based diagnosis information can reduce the overhead of the diagnosis information telemetry head on the basis of ensuring detection accuracy. Determining the target device corresponding to the violation event through native path encoding according to the media access control codes of each device in the network can further reduce the overhead of the system.
[0053] The following will further illustrate Figures 2 - 9 the Figure 1 service level objective violation diagnosis method shown in the figure.
[0054] Figure 2 The figure schematically shows a flowchart of a method for collecting location-based diagnosis information according to an embodiment of the present invention.
[0055] As Figure 2 shown, the method for collecting location-based diagnosis information may include operations S201 to S202, for example.
[0056] In operation S201, the path information is compressed and grouped into three parts: group ID, group information, and intra-group information.
[0057] In operation S202, random sampling is performed on the devices within the group based on the group ID, group information, and intra-group information to probabilistically collect location-based diagnosis information.
[0058] In an embodiment of the present invention, the group ID represents the sequence number of the sampled group, the intra-group information represents the information of the sampled group, and the group information records all the violation groups on the network path.
[0059] Figure 3 The figure schematically shows a schematic diagram of a method for collecting location-based diagnosis information according to an embodiment of the present invention.
[0060] Exemplarily, assume that the path passed by an n-hop traffic f is represented as r = (d1, d2,..., d n ), where d is a single device. The packet size g is obtained by prediction, that is, each group contains no more than g devices, and the grouping is automatically grouped. For example, for r x = (d1, d2, d3), g = 2, (d1, d2) is the first group, (d3) is the second group, and the number of groups is represented by l.
[0061] As Figure 3 shown, assume that the maximum number of hops h of the network path maxEqual to 48, let the number of groups l be 6, the group size g be 8, and the path length occupies 6 bits in this network. Among them, 3 bits are for the group length, and 3 bits are for the in-group length. At this time, the information in the original packet indicates that it carries the information of the fourth group (there is an abnormality in the 28th hop device). At the same time, the original packet also declares that there is a suspicious location in the third group. Therefore, when the system performs information recovery, it can determine whether all the data has been obtained.
[0062] Furthermore, if reservoir sampling is used in the process of collecting location-based violation information, since reservoir sampling does not require prior knowledge of the problem scale, when traversing to the m-th (m>k) element, the probability that any one of the first m elements is retained is k / m. When there are suspicious situations within the group (such as an overly long exchange queue, packet loss, etc.), sample the group. At this time, count how many suspicious groups there are currently and perform sampling according to reservoir sampling. Reservoir sampling requires the use of division and random numbers, and these operations cannot be performed on the data plane.
[0063] To solve this problem, the present invention proposes an approximation algorithm, and uses the approximation algorithm to randomly sample the devices within the group. Random sampling can include, for example: using the time period when the data packet enters the network (Ingress epoch) and the path length of the current data packet as random numbers, and using the hash function as the work function to randomly sample the devices within the group. The sampling process can be as follows:
[0064] Input: Mask function F, Hash function H1 and H2,
[0065] Correction mask M c , i-th sample, Ingress epoch,
[0066] Output: Ture or False
[0067] mask := F(i) & M c
[0068] a := H1(Ingress epoch)
[0069] b := H2(Path length)
[0070] c := mask & (a⊕b)
[0071] if c = 0 then
[0072] return Ture
[0073] else
[0074] return False
[0075] end if
[0076] Based on the above hash functions H1 and H2, it is ensured that the sampling results of the groups within one epoch are the same.
[0077] In addition, the left shift operation can be used for approximate calculation to overcome the limitation of not being able to perform division. h max is the maximum number of hops of a predictable network path (such as 48 hops). F(x) returns the mask M as all bits with the maximum number of bits of x. For example, when h max = 8. Among them:
[0078]
[0079] For more accurate diagnosis, the correction mask Mc can be further applied. Mc is pre-calculated according to h max The correction mask Mc helps improve the accuracy of the sampling probability approximation value. All operations can be completed in O(1) time on the data plane.
[0080] In the embodiments of the present invention, the detail-based diagnostic information may include: probabilistically collecting detail-based diagnostic information through variable-length 0 sequences and padding bits.
[0081] Figure 4 Schematically shows a flowchart of a method for collecting detail-based diagnostic information according to an embodiment of the present invention.
[0082] Such as Figure 4 shown, the method for collecting detail-based diagnostic information may include operations S401 to S402 for example.
[0083] In operation S401, in the case of detecting packet loss, obtain the group ID of the current device within the packet through telemetry message information.
[0084] In operation S402, obtain the value i of the group ID, fill in i zeros to obtain a variable-length 0 sequence, sequentially fill in the port where the data packet enters the network and the port where the data packet leaves the network after the variable-length 0 sequence, and use 0 to make up the number of bits.
[0085] Figure 5 Schematically shows a schematic diagram of a method for collecting detail-based diagnostic information according to an embodiment of the present invention.
[0086] Such as Figure 5As shown, exemplarily, the packet loss port information is detailed information that the controller needs to know when a violation occurs, and the switch does not need to process it. Encoding uses the switch data plane method, but decoding is completed by the controller that can perform complex algorithms. A piece of packet loss port information usually includes the in / out switch (48 * 2 = 96 bits) and the in / out port (9 * 2 = 18 bits). The system can load n pieces of packet loss port information in 26 bits using the previous PathID and the packet. When the switch detects packet loss, it obtains the current order i of the device in the packet through telemetry message information. Then, i zeros are filled in as a variable-length zero sequence. After that, the in / out ports are filled in sequentially, and the number of bits is filled with zeros. If different groups trigger the packet mechanism, they are completely replaced. When there are multiple pieces of packet loss port information in the same group, XOR operations are probabilistically performed. Because of the property of XOR, the complete information can be restored after multiple reports.
[0087] Figure 6 Schematically shows a flowchart of a method for determining a target device for violation according to an embodiment of the present invention.
[0088] As Figure 6 shown, the method for determining the target device for violation may include, for example, operation S601 to operation S602.
[0089] In operation S601, calculate the end-to-end path ID of all data packets in the network according to the media access control codes of each device in the network.
[0090] In operation S602, record the path IDs of different actual service paths at each egress node and determine the target device corresponding to the violation event.
[0091] Exemplarily, each network device has a unique 48-bit MAC code (media access control code), and the manufacturer assigns the last 3 bytes. Since multiple devices are obtained from the same manufacturer in the network environment, the last μ bits of the MAC are taken and θ bits of anti-collision bits are set to form the device ID. The anti-collision bits are defaulted to 0. Define the flow f<s,t> corresponding to the data packet, and its path set R f =(r1,r2,…,r x ). Since d represents a device, r x can be composed of (d1,d2,…,d y ). Among them, the device ID is c y . Then the path ID is:
[0092]
[0093] There may be hash collisions in PathID. Verify whether different paths use the same PathID at each egress node. If a collision occurs, the controller calculates the native MAC code of the device in the network and issues a device ID correction bit to modify the anti-collision bit of one of the devices in the conflicting path, and returns for re-verification until there are no more conflicts.
[0094] It should be noted that Figure 1 The shown operation sequence is not used to limit the present invention. Operation S103 can be executed in parallel with operations S101 and S102. By collecting location-based diagnostic information and detail-based diagnostic information through operations S101 and S102, the specific hop count where a violation occurs in the network can be determined, but it is not clear which specific device the hop count belongs to. Therefore, flow detection can be achieved by transmitting path information through the native MAC code of the devices in the network to determine the target device where the violation occurs.
[0095] Based on the technology of the above embodiment, the service level objective violation diagnosis method further includes: recovering the violation events included in the location-based diagnostic information and / or the detail-based diagnostic information. By quickly recovering, the high availability of the network is ensured.
[0096] Based on the technology of the above embodiment, the service level objective violation diagnosis method further includes: reporting the violation events by using a threshold-triggered violation reporting mechanism.
[0097] Figure 7 Schematically shows a flowchart of a method for reporting violation events according to an embodiment of the present invention.
[0098] As Figure 7 shown, the method for reporting violation events may include, for example, operations S701 to S702.
[0099] In operation S701, set a reporting times threshold.
[0100] In operation S702, in response to the reporting times of the violation event corresponding to the current data flow not being greater than the reporting times threshold, normally report the violation event.
[0101] In operation S703, in response to the reporting times of the violation event corresponding to the current data flow being greater than the reporting times threshold, stop reporting the violation event, and continuously configure the current data flow as a violation state until it is detected that there is no violation event in the current data flow, and then resume normal reporting of the violation event.
[0102] Exemplarily, when the system detects a violation of the flow corresponding to a data packet, it will first report it normally. Then, when the number of reports exceeds the reporting times threshold ζ, the system will stop reporting. After that, this traffic will be continuously regarded as the SLO violation state. Until a normal event of this traffic is detected, it starts to report normally and subtracts the count value. When the count value is zero, this flow is considered to be in a normal state again.
[0103] The reporting times threshold ζ can be set according to the maximum number of hops of the network path, and the recommended ζ depends on h max :
[0104]
[0105] It should be noted that the above-mentioned various data can be integrated into in-band telemetry meta-information. The in-band telemetry meta-information is shown in Table 1.
[0106] Table 1
[0107]
[0108] Path ID is the actual service path that the flow passes through. Ingress epoch and Egress epoch are the last ingress and egress epochs of the network respectively. RxPks is the total number of packets received by the network in the epoch currently. TxPks and TxBits are the total number of packets and bytes forwarded by the network in the epoch currently respectively. Path information represents the diagnostic information in the path. Port information represents the specific suspicious port information in the path.
[0109] To further verify the effectiveness of the above service level objective violation diagnosis method, a specific example is given below for illustration.
[0110] The experiment was carried out on REPITITA. REPITITA is a framework for repeatable traffic engineering algorithm experiments on a large-scale dataset of a real network. This experiment verified that this example meets the SLO violation detection requirements and has lower overhead than other similar methods.
[0111] Figure 8 Schematically shows the experimental result graph of the comparison of the service level objective violation diagnosis method according to the embodiment of the present invention with the existing method in terms of the consumption of Static Random-Access Memory (SRAM).
[0112] Figure 9 Schematically shows the experimental result graph of the comparison of the service level objective violation diagnosis method according to the embodiment of the present invention with the existing method in terms of the violation reporting rate.
[0113] As Figure 8 and Figure 9 shown, benefiting from the probability compression of telemetry information, this example achieves the best SRAM usage and the lowest event reporting rate without loss of accuracy compared with the comparative method. In fact, there is still enough space in the switch for additional services, with at least 71.9% of the Ternary Content Addressable Memory (TCAM) and 63.3% of the SRAM still remaining).
[0114] Figure 10 Schematically shows a block diagram of a service level objective violation diagnosis device according to an embodiment of the present invention.
[0115] As Figure 10 shown, the service level objective violation diagnosis device 1000 may include, for example, a path information probability compression module 1001, a detailed information probability compression module 1002, and a native path encoding module 1003.
[0116] The path information probability compression module 1001 is configured to compress and group devices on a network path based on path information to probabilistically collect location-based diagnostic information, where the location-based diagnostic information includes status information of specific points on the network path.
[0117] The detailed information probability compression module 1002 is configured to probabilistically detect packet loss port information of all devices on the network path based on the compression grouping result to obtain detailed-based diagnostic information, where the detailed-based diagnostic information refers to specific data required for suspicious area diagnosis.
[0118] The native path encoding module 1003 is configured to perform native path encoding according to the media access control codes of devices in the network to determine the target devices corresponding to the violation events included in the location-based diagnostic information and / or the detailed-based diagnostic information.
[0119] Figure 11 Schematically shows a detailed diagram of a service level objective violation diagnosis device according to an embodiment of the present invention.
[0120] As Figure 11 shown, in the data plane, the collection of location-based diagnostic information, the collection of detailed-based diagnostic information, and native path encoding can be completed.
[0121] It should be noted that the part of the service level objective violation diagnosis device in the embodiment of the present invention corresponds to the part of the service level objective violation diagnosis method in the embodiment of the present invention, and their specific implementation details and the technical effects brought are also the same, which will not be elaborated here.
[0122] Any of a plurality of modules, sub-modules, units, and sub-units according to embodiments of the present invention, or at least part of the functions of any of them, can be implemented in one module. Any one or more of the modules, sub-modules, units, and sub-units according to embodiments of the present invention can be split into multiple modules for implementation. Any one or more of the modules, sub-modules, units, and sub-units according to embodiments of the present invention can be at least partially implemented as a hardware circuit, such as a field programmable gate array (FPGA), a programmable logic array (PLA), a system on chip, a system on a substrate, a system on a package, an application specific integrated circuit (ASIC), or can be implemented by any other reasonable way of integrating or packaging circuits, or in any one of the three implementation manners of software, hardware, and firmware, or in an appropriate combination of any of them. Alternatively, one or more of the modules, sub-modules, units, and sub-units according to embodiments of the present invention can be at least partially implemented as a computer program module, and when the computer program module is run, the corresponding functions can be executed.
[0123] For example, any of the path information probability compression module 1001, the detail information probability compression module 1002, and the native path encoding module 1003 can be combined and implemented in one module / unit / sub-unit, or any one of the modules / units / sub-units can be split into multiple modules / units / sub-units. Alternatively, at least part of the functions of one or more of these modules / units / sub-units can be combined with at least part of the functions of other modules / units / sub-units and implemented in one module / unit / sub-unit. According to embodiments of the present invention, at least one of the path information probability compression module 1001, the detail information probability compression module 1002, and the native path encoding module 1003 can be at least partially implemented as a hardware circuit, such as a field programmable gate array (FPGA), a programmable logic array (PLA), a system on chip, a system on a substrate, a system on a package, an application specific integrated circuit (ASIC), or can be implemented by any other reasonable way of integrating or packaging circuits, or in any one of the three implementation manners of software, hardware, and firmware, or in an appropriate combination of any of them. Alternatively, at least one of the path information probability compression module 1001, the detail information probability compression module 1002, and the native path encoding module 1003 can be at least partially implemented as a computer program module, and when the computer program module is run, the corresponding functions can be executed.
[0124] Figure 12 A block diagram of an electronic device suitable for implementing the method described above according to embodiments of the present invention is schematically shown. Figure 12 The electronic device shown is only an example and should not impose any limitation on the functions and usage scope of embodiments of the present invention.
[0125] As shown Figure 12 in FIG. 1, an electronic device 1200 according to an embodiment of the present invention includes a processor 1201, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 1202 or a program loaded from a storage section 1208 into a random access memory (RAM) 1203. The processor 1201 may include, for example, a general-purpose microprocessor (such as a CPU), an instruction set processor, and / or a related chipset, and / or a dedicated microprocessor (such as an application-specific integrated circuit (ASIC)), and so on. The processor 1201 may also include on-board memory for caching purposes. The processor 1201 may include a single processing unit or multiple processing units for performing different actions of the method flow according to an embodiment of the present invention.
[0126] In the RAM 1203, various programs and data required for the operation of the electronic device 1200 are stored. The processor 1201, the ROM 1202, and the RAM 1203 are connected to each other via a bus 1204. The processor 1201 performs various operations of the method flow according to an embodiment of the present invention by executing programs in the ROM 1202 and / or the RAM 1203. It should be noted that the program may also be stored in one or more memories other than the ROM 1202 and the RAM 1203. The processor 1201 may also perform various operations of the method flow according to an embodiment of the present invention by executing programs stored in the one or more memories.
[0127] According to an embodiment of the present invention, the electronic device 1200 may further include an input / output (I / O) interface 1205, and the input / output (I / O) interface 1205 is also connected to the bus 1204. The electronic device 1200 may further include one or more of the following components connected to the I / O interface 1205: an input portion 1206 including a keyboard, a mouse, etc.; an output portion 1207 including a cathode ray tube (CRT), a liquid crystal display (LCD), etc., and a speaker, etc.; a storage portion 1208 including a hard disk, etc.; and a communication portion 1209 including a network interface card such as a LAN card, a modem, etc. The communication portion 1209 performs communication processing via a network such as the Internet. A drive 1210 is also connected to the I / O interface 1205 as needed. A removable medium 1211, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 1210 as needed so that a computer program read from it can be installed into the storage portion 1208 as needed.
[0128] According to an embodiment of the present invention, the method flow according to the embodiment of the present invention can be implemented as a computer software program. For example, an embodiment of the present invention includes a computer program product, which includes a computer program carried on a computer-readable storage medium, and the computer program contains program codes for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network through the communication part 1209, and / or installed from the removable medium 1211. When the computer program is executed by the processor 1201, the above functions defined in the system of the embodiment of the present invention are executed. According to an embodiment of the present invention, the above-described system, device, apparatus, module, unit, etc. can be implemented by computer program modules.
[0129] The present invention also provides a computer-readable storage medium, which can be included in the device / device / system described in the above embodiment; or can exist alone without being assembled into the device / device / system. The above computer-readable storage medium carries one or more programs, and when the above one or more programs are executed, the method according to the embodiment of the present invention is implemented.
[0130] According to an embodiment of the present invention, the computer-readable storage medium can be a non-volatile computer-readable storage medium. For example, it can include but is not limited to: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the above. In the present invention, the computer-readable storage medium can be any tangible medium that contains or stores a program, and the program can be used by or in combination with an instruction execution system, device, or device.
[0131] For example, according to an embodiment of the present invention, the computer-readable storage medium can include one or more memories other than the above-described ROM 1202 and / or RAM 1203 and / or ROM 1202 and RAM 1203.
[0132] In the technical solution of the present invention, the information and data involved (including but not limited to data for analysis, stored data, displayed data, etc.) are all authorized information and data, and the collection, storage, use, processing, transmission, provision, disclosure, and application of relevant data and other processes all comply with relevant laws, regulations, and standards, take necessary confidentiality measures, do not violate public order and good customs, and provide corresponding operation entrances for selecting authorization or refusal.
[0133] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a portion of code, which contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions noted in the blocks may occur in a different order than that noted in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram or flowchart, and combinations of blocks in the block diagram or flowchart, can be implemented by a dedicated hardware-based system that performs the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions. Those skilled in the art will understand that the features described in the various embodiments of the present invention can be combined and / or combined in various ways, even if such combinations or combinations are not explicitly described in the present invention. In particular, without departing from the spirit and teachings of the present invention, the features described in the various embodiments of the present invention can be combined and / or combined in various ways. All such combinations and / or combinations fall within the scope of the present invention.
Claims
1. A method for diagnosing violations of service level objectives, characterized in that, Implement the service level objective violation diagnosis method in the data plane, including: Compressively group the devices on the network path based on the path information to probabilistically collect location-based diagnostic information, where the location-based diagnostic information includes the status information of specific points on the network path; Based on the compressive grouping result, probabilistically detect the packet loss port information of all devices on the network path to obtain detail-based diagnostic information, where the detail-based diagnostic information refers to the specific data required for suspicious area diagnosis; Perform native path encoding according to the media access control codes of each device in the network to determine the target device corresponding to the violation event included in the location-based diagnostic information and / or the detail-based diagnostic information.
2. The method for diagnosing service level objective violations according to claim 1, wherein The step of compressing and grouping all devices on the network path based on the path information to probabilistically collect location-based diagnostic information includes: Compressively group the path information into three parts: group ID, group information, and in-group information, where the group ID represents the sequence number of the sampled group, the in-group information represents the information of the sampled group, and the group information records all violation groups on the network path; Randomly sample the devices within the group based on the group ID, group information, and in-group information to probabilistically collect location-based diagnostic information.
3. The service level objective violation diagnosis method according to claim 2, characterized in that, Use an approximation algorithm to randomly sample the devices within the group.
4. The method for diagnosing service level objective violations according to claim 2 or 3, characterized in that, Use the time period when the data packet enters the network and the path length of the current data packet as random numbers, and use the hash function as the work function to randomly sample the devices within the group.
5. The service level objective violation diagnosis method according to claim 2, characterized in that The step of probabilistically detecting the packet loss port information of all devices on the network path based on the compressive grouping result to obtain detail-based diagnostic information includes: Probabilistically collect detail-based diagnostic information through variable-length 0 sequences and padding bits.
6. The method for diagnosing service level objective violations according to claim 5, characterized in that, The step of probabilistically collecting detail-based diagnostic information through variable-length 0 sequences and padding bits includes: In the case of detecting packet loss, obtain the group ID of the current device within the group through telemetry message information; Obtain the value i of the group ID, fill in i 0s to obtain a variable-length 0 sequence, sequentially fill in the port where the data packet enters the network and the port where the data packet leaves the network after the variable-length 0 sequence, and use 0 to fill in the number of bits.
7. The method for diagnosing service level objective violations according to claim 2, wherein The step of performing native path encoding according to the media access control codes of each device in the network to determine the target device corresponding to the violation event included in the location-based diagnostic information and / or the detail-based diagnostic information includes: Calculate the end-to-end path ID of all data packets in the network according to the media access control codes of each device in the network; Record the path IDs of different real service paths at each egress node to determine the target device corresponding to the violation event.
8. The method for diagnosing service level objective violations according to claim 1, wherein The service level objective violation diagnosis method further includes: Recover the violation events included in the location-based diagnostic information and / or the detail-based diagnostic information.
9. The method for diagnosing service level objective violations according to claim 8, wherein The service level objective violation diagnosis method further includes: Report the violation events using a threshold-triggered violation reporting mechanism.
10. The method for diagnosing violation of service level objective according to claim 9, wherein The step of reporting the violation events using a threshold-triggered violation reporting mechanism includes: Set the threshold for the number of reports; In response to the number of reported violations corresponding to the current data stream being no greater than the reporting threshold, report the violation event normally; In response to the number of reported violations corresponding to the current data stream being greater than the reporting threshold, stop reporting the violation event, and continuously configure the current data stream as a violation state until no violation event is detected in the current data stream, and resume normal reporting of violation events.
11. The method for diagnosing service level objective violations according to claim 10, characterized in that, Set the reporting threshold according to the maximum number of hops of the network path.
12. The method for diagnosing service level objective violations according to any one of claims 9-11, characterized in that, Report the status change information in the initial stage of reporting violations and the initial stage of recovery.
13. A service level objective violation diagnosis device, characterized in that Implement the service level objective violation diagnosis device in the data plane, including: A path information probability compression module, configured to compress and group devices on the network path based on path information to probabilistically collect location-based diagnosis information, where the location-based diagnosis information includes the status information of specific points on the network path; A detailed information probability compression module, configured to probabilistically detect the packet loss port information of all devices on the network path based on the compression grouping result to obtain detailed-based diagnosis information, where the detailed-based diagnosis information refers to the specific data required for suspicious area diagnosis; A native path encoding module, configured to perform native path encoding according to the media access control codes of devices in the network to determine the target device corresponding to the violation event included in the location-based diagnosis information and / or the detailed-based diagnosis information.
14. An electronic device, characterized in that, Including: One or more processors; A memory, configured to store one or more programs, wherein, when the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1 to 12.
15. A computer-readable storage medium, characterized in that, Stored thereon are executable instructions, which when executed by a processor cause the processor to implement the method according to any one of claims 1 to 12.