Fault detection method and computing device

By calculating the difference ratio of traffic rate pairs and the chi-square test statistic to identify abnormal traffic, the problem of silent chip faults being difficult to track and the high false alarm rate is solved, achieving efficient and accurate fault detection and ensuring data communication security.

CN116633824BActive Publication Date: 2025-09-12SHENZHEN HUAWEI CLOUD COMPUTING TECHNOLOGIES CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310489230.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-04-28
Publication Date
2025-09-12
Estimated Expiration
2043-04-28

AI Technical Summary

Technical Problem

Silent chip failures caused by processor calculation errors cannot be tracked at the hardware level, resulting in packet loss during communication and the inability to issue timely alarms, affecting data security. Existing abnormal traffic detection methods have a high false alarm rate and are difficult to be consistent with actual link traffic fluctuations.

Method used

By obtaining the input and output difference ratio sequence and the input and output difference ratio sequence of the traffic rate pair, calculating the test statistic and failure probability value, and using the packet conservation algorithm and chi-square test statistic to identify abnormal traffic, the false alarm rate is reduced and the detection accuracy is improved.

Benefits of technology

Effectively identify abnormal traffic, reduce false alarm rates, improve detection accuracy, promptly discover silent chip failures, prevent communication interruptions, and ensure data security.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116633824B_ABST
    Figure CN116633824B_ABST
Patent Text Reader

Abstract

A fault detection method includes: after obtaining a set of test statistics, obtaining the traffic rate pairs of the target node in multiple detection time periods, determining the input and output difference ratio sequence and the input and output difference ratio sequence of the detection time period based on the traffic rate pairs of the detection time period, and then determining the test statistics of the detection time period based on the above two sequences, and then calculating the fault probability value of the detection time period based on the test statistics of the detection time period and the test statistics set, and judging whether the target node has abnormal traffic based on the fault probability values ​​of multiple detection time periods. The test statistics can reflect the difference between the input and output difference ratio distribution and the input and output difference ratio distribution, and the above difference is related to the node failure. Therefore, based on the above test statistics, it is possible to judge whether the node traffic is abnormal. The present application also provides a computing device that can implement the above fault detection method.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of data security, and in particular to a fault detection method and a computing device. Background Art

[0002] Processors have a certain chance of making calculation errors. The resulting erroneous data cannot be tracked at the hardware level. This error is called a silent chip failure. Silent chip failures can cause packet loss during communication. This type of packet loss typically does not generate alarms or logs, significantly impacting data security.

[0003] Currently, there is a method for detecting abnormal traffic that is roughly as follows: obtain link traffic, and if the minimum value of the link traffic within a detection period is greater than the traffic lower limit of the detection period, it is determined that abnormal traffic has occurred.

[0004] During normal communication, link traffic will fluctuate. The dynamic baseline of the detection period is difficult to be consistent with the actual link traffic fluctuation, which is prone to false alarms. Summary of the Invention

[0005] In view of this, the present application provides a fault detection method that can obtain a test statistic related to the difference in traffic entering and leaving a node, and judge traffic faults based on this test statistic, thereby reducing false positives and improving the accuracy of identifying abnormal traffic. The present application also provides a computing device capable of implementing the above fault detection method.

[0006] A first aspect provides a fault detection method, which includes: after obtaining a set of test statistics, obtaining traffic rate pairs of a target node in multiple detection time periods; determining the input-output difference ratio sequence and the input-output difference ratio sequence of the detection time period based on the multiple traffic rate pairs of each detection time period, and then determining the test statistic of the detection time period based on the input-output difference ratio sequence and the input-output difference ratio sequence of the detection time period; and then calculating the fault probability value of the detection time period based on the test statistic of the detection time period and the set of test statistics. When, among the fault probability values ​​of multiple detection time periods, the node failure probability values ​​of N consecutive detection time periods are less than or equal to a preset significance level and N is greater than or equal to a first threshold, it is determined that there is abnormal traffic at the target node.

[0007] Each traffic rate pair includes the inbound traffic rate and the outbound traffic rate collected at a measurement moment. The inbound and outbound difference ratios in the inbound and outbound difference ratio sequence correspond one-to-one to the traffic rate pair, and the inbound and outbound difference ratios in the inbound and outbound difference ratio sequence correspond one-to-one to the traffic rate pair. The first threshold can be any integer between 3 and 10. The specific value can be set based on actual conditions and is not limited in this application.

[0008] According to this implementation, the test statistics included in the test statistics set are calculated based on the normal inbound traffic rate and the normal outbound traffic rate, so the test statistics set can reflect the normal distribution of the test statistics. After calculating the test statistics of the detection period according to the method of the first aspect, the test statistics of the detection period are related to the distribution of the proportion of inbound and outbound differences and the distribution of the proportion of inbound and outbound differences in the detection period, and can reflect the maximum difference between the above two distributions. According to the test statistics of the detection period and the normal test statistics, the failure probability value of the detection period can be calculated, thereby determining whether the traffic of the node is abnormal. Compared with the minimum value of the existing link traffic rate, identifying abnormal traffic based on the test statistics can reduce the false alarm rate and increase the accuracy of detecting abnormal traffic.

[0009] In one possible implementation, the traffic rate pair is a unicast traffic rate pair. Test data shows that, compared with multicast traffic rate pairs or broadcast traffic rate pairs, the test statistic calculated based on the unicast traffic rate pair can more accurately identify abnormal traffic.

[0010] In another possible implementation, determining the test statistic of the detection period based on the input-output difference ratio sequence and the input-output difference ratio sequence of the detection period includes: determining the first distribution function of the detection period based on the input-output difference ratio sequence of the detection period, and determining the second distribution function of the detection period based on the input-output difference ratio sequence of the detection period, and then determining the target function of the detection period based on the first distribution function of the detection period and the second distribution function of the detection period, and determining that the test statistic of the detection period is equal to the upper bound of the target function of the detection period. Wherein, the first distribution function Second distribution function And the objective function F′(z) satisfies the formula: z is the independent variable of the objective function of the detection period. The independent variable interval of the objective function may include the independent variable interval of the first distribution function and the independent variable interval of the second distribution function. Optionally, the minimum value of the independent variable interval of the first distribution function is the minimum value of the input and output difference ratio sequence, and the maximum value of the independent variable interval of the first distribution function is the maximum value of the input and output difference ratio sequence. The minimum value of the independent variable interval of the second distribution function is the minimum value of the input and output difference ratio sequence, and the maximum value of the independent variable interval of the second distribution function is the maximum value of the input and output difference ratio sequence. The test statistic value of the detection period is related to the input and output difference ratio distribution of the detection period and the input and output difference ratio distribution of the detection period, and can reflect the maximum difference between the above two distributions. This provides a specific feasible solution for calculating the test statistic value of the test period.

[0011] In another possible implementation, calculating a failure probability value for the detection period based on a test statistic value for the detection period and a test statistic value set includes: when M test statistic values ​​in the test statistic value set are greater than or equal to the test statistic value for the detection period, determining the failure probability value for the detection period as M divided by the total number of test statistic values ​​in the test statistic value set. When the test statistic value for the detection period exceeds a normal test statistic value, a failure may have occurred, thereby providing a method for calculating a failure probability value for the detection period.

[0012] In another possible implementation, obtaining a test statistic value set includes steps A, B, C, D, E, and F, wherein step A includes extracting a set of flow rate pairs from the flow rate pairs in the statistical period; step B includes determining, based on the flow rate pair set, a sequence of input and output difference ratios and a sequence of input and output difference ratios corresponding to the flow rate pair set; step C includes determining a first distribution function of the flow rate pair set based on the sequence of input and output difference ratios corresponding to the flow rate pair set; step D includes determining a second distribution function of the flow rate pair set based on the sequence of input and output difference ratios corresponding to the flow rate pair set; step E includes determining an objective function of the flow rate pair set based on the first distribution function of the flow rate pair set and the second distribution function of the flow rate pair set; step F includes determining that a test statistic value of the flow rate pair set is equal to the upper bound of the objective function of the flow rate pair set; and steps A to F are repeated until the number of test statistic values ​​reaches a preset number of the test statistic value set. The preset number is the total number of test statistic values ​​of the test statistic value set. The specific value can be set according to actual conditions and is not limited in this application.

[0013] In another possible implementation, the fault detection method of the present application further includes: determining the output flow rate proportion of each interface at multiple measurement moments in the interface group corresponding to the ECMP group; determining the chi-square test statistic at the measurement moment based on the output flow rate proportion of the interface at the measurement moment, the expected value of the output flow rate proportion of the interface, the sample variance of the output flow rate proportion of the interface, and the total number of flows passing through the interface group; calculating the fault probability value at the measurement moment based on the chi-square test statistic at the measurement moment; when the probability values ​​of L consecutive measurement moments among the fault probability values ​​at the multiple measurement moments are less than or equal to a preset significance level and L is greater than or equal to a second threshold, determining that the flow of the ECMP group is abnormal. L is a positive integer, and the second threshold can be any integer from 3 to 10. The specific value can be set according to actual conditions and is not limited by this application.

[0014] In this implementation, a chi-square test statistic can be determined based on the interface traffic rate ratio. This chi-square test statistic reflects the degree of deviation between the output traffic rate ratio of all interfaces in the ECMP group and the expected output traffic rate ratio. The greater the deviation, the greater the probability of a failure. Therefore, the chi-square test statistic can be used to calculate the failure probability at each measurement time. The failure probability values ​​at multiple measurement times can then be used to determine whether the ECMP group's traffic is abnormal.

[0015] In another possible implementation, the output traffic rate ratio of the interface at the measurement time, the expected value of the output traffic rate ratio of the interface, the sample variance of the output traffic rate ratio of the interface, the total number of flows passing through the interface group, and the chi-square test statistic at the measurement time satisfy the following formula:

[0016]

[0017] T is the chi-square test statistic at the measurement moment, is the output traffic rate ratio of the i-th interface, is the expected value of the output traffic rate ratio of the i-th interface, is the sample variance of the interface traffic rate ratio of the i-th interface, d is the total number of interfaces in the interface group, and n is the total number of flows passing through the interface group.

[0018] In another possible implementation, the chi-square test statistic value and the failure probability value at the measurement time satisfy the following formula:

[0019]

[0020] p′ is the failure probability value at the measurement time, is the chi-square test statistic at the measurement moment.

[0021] In another possible implementation, the fault detection method of the present application further includes: determining that traffic on the target interface is abnormal when the chi-square test statistic of the target interface is greater than or equal to a third threshold. This enables fault location of the interface with abnormal traffic.

[0022] The second aspect provides a computing device, which includes an acquisition module, a measurement module, a computing module and a judgment module; the acquisition module is used to acquire a test statistic value set; the measurement module is used to acquire the traffic rate pairs of the target node in multiple detection time periods; the computing module is used to determine the input and output difference ratio sequence and the input and output difference ratio sequence of the detection period according to the multiple traffic rate pairs of the detection period for each detection period; determine the test statistic value of the detection period according to the input and output difference ratio sequence of the detection period and the input and output difference ratio sequence of the detection period; calculate the fault probability value of the detection period according to the test statistic value of the detection period and the test statistic value set; the judgment module is used to determine that there is abnormal traffic at the target node when the failure probability values ​​of N consecutive detection periods among the failure probability values ​​of multiple detection periods are less than or equal to a preset significance level and N is greater than or equal to a first threshold.

[0023] In one possible implementation, the calculation module is specifically used to determine the first distribution function of the detection period based on the sequence of the proportion of the input and output differences of the detection period; determine the second distribution function of the detection period based on the sequence of the proportion of the input and output differences of the detection period; determine the objective function of the detection period based on the first distribution function of the detection period and the second distribution function of the detection period; determine that the test statistic value of the detection period is equal to the upper bound of the objective function of the detection period.

[0024] In another possible implementation, the calculation module is specifically configured to determine, when M test statistics in the test statistic set are greater than or equal to the test statistic of the detection period, the failure probability value of the detection period as M divided by the total number of test statistics in the test statistic set.

[0025] In another possible implementation, the acquisition module is specifically configured to repeatedly execute steps A to F until the number of test statistic values ​​reaches a preset number of the test statistic value set.

[0026] In another possible implementation, the measurement module is further used to determine the output traffic rate proportion of each interface at multiple measurement moments in the interface group corresponding to the ECMP group; the calculation module is further used to determine, for each measurement moment, the chi-square test statistic of the interface at the measurement moment, based on the output traffic rate proportion of the interface at the measurement moment, the expected value of the output traffic rate proportion of the interface, the sample variance of the output traffic rate proportion of the interface, and the total number of flows passing through the interface group; calculate the failure probability value at the measurement moment based on the chi-square test statistic at the measurement moment; and the judgment module is further used to determine that the traffic of the ECMP group is abnormal when the failure probability values ​​of L consecutive measurement moments among the failure probability values ​​at multiple measurement moments are less than or equal to a preset significance level and L is greater than or equal to a second threshold.

[0027] In another possible implementation, the judgment module is further configured to determine that the traffic of the target interface is abnormal when the chi-square test statistic value of the target interface is greater than or equal to a third threshold.

[0028] For the explanation of terms in the second aspect, the steps performed by each module and the beneficial effects, please refer to the corresponding description of the first aspect.

[0029] The third aspect provides a computing device cluster, which includes at least one computing device, each computing device includes a processor and a memory, and the processor of at least one computing device is used to execute instructions stored in the memory of at least one computing device, so that the computing device cluster performs the method of the first aspect.

[0030] A fourth aspect provides a computer-readable storage medium having computer program instructions stored therein. When the computer program is executed by a computing device, the computing device executes the method of the first aspect.

[0031] A fifth aspect provides a computer program product comprising instructions, which, when executed by a computing device, causes the computing device to perform the method of the first aspect. BRIEF DESCRIPTION OF THE DRAWINGS

[0032] Figure 1 A schematic diagram of a fault detection scenario in an embodiment of the present application;

[0033] Figure 2 This is another schematic diagram of a fault detection scenario in an embodiment of the present application;

[0034] Figure 3A This is another schematic diagram of the fault detection method in an embodiment of the present application;

[0035] Figure 3B This is another schematic diagram of the fault detection method in an embodiment of the present application;

[0036] Figure 4 A flowchart of a fault detection method in an embodiment of the present application;

[0037] Figure 5A A schematic diagram of the flow detection results in an embodiment of the present application;

[0038] Figure 5B This is another schematic diagram of the flow detection results in the embodiment of the present application;

[0039] Figure 5C This is another schematic diagram of the flow detection results in the embodiment of the present application;

[0040] Figure 6 A schematic diagram of a scenario in which traffic data is forwarded through an ECMP group in an embodiment of the present application;

[0041] Figure 7 A schematic diagram of a method for detecting ECMP group failure in an embodiment of the present application;

[0042] Figure 8 This is another flow chart of the fault detection method in an embodiment of the present application;

[0043] Figure 9 This is another schematic diagram of the flow detection results in the embodiment of the present application;

[0044] Figure 10 A structural diagram of a computing device in an embodiment of the present application;

[0045] Figure 11 This is another structural diagram of a computing device in an embodiment of the present application;

[0046] Figure 12 A schematic diagram of a computing device cluster in an embodiment of the present application;

[0047] Figure 13 This is another schematic diagram of a computing device cluster in an embodiment of the present application. DETAILED DESCRIPTION

[0048] The fault detection method of the present application can be applied to traffic detection scenarios, which can be, but are not limited to, chip silence scenarios. The network device used to perform the fault detection method in the traffic detection scenario can be, but is not limited to, a switch, a router, a switch board, or a router board.

[0049] See Figure 1 In one embodiment, after receiving unicast traffic data, multicast traffic data, and broadcast traffic data, a network device may transmit the unicast traffic data, multicast traffic data, and broadcast traffic data. It should be understood that the traffic data received and transmitted by the network device includes one or more of unicast traffic data, multicast traffic data, and broadcast traffic data. Unicast is a communication method between a sender and a single receiver. Multicast is a communication method between a sender and multiple receivers. Broadcast is a communication method between a sender and all receivers on the network.

[0050] For traffic data received during any time period, a network device can measure the inbound traffic rate. For traffic data sent during any time period, a network device can measure the outbound traffic rate. When traffic rate is expressed as a bit rate, the unit can be bits per second (bps).

[0051] See Figure 2In another example, a network device is configured with a computing delivery unit, which includes computing delivery points (PODs) 1 to 3, four processing cores, storage delivery point 1, and storage delivery point 2. The four processing cores are processing core 1 to processing core 4, and each processing core includes four slots.

[0052] When the computing delivery point receives traffic data (such as unicast traffic data, multicast traffic data or broadcast traffic data), the computing delivery point sends the traffic data to the slots of each processing core respectively. After the processing core processes the traffic data through the slots, it sends the processing results to the storage delivery point or output.

[0053] When a silent fault occurs in slot 4 of processor core 4, it has a significant impact on the network because it does not trigger alarms or fault logs. In an example, the impact range and fault recovery time of a silent fault are shown in Table 1:

[0054]

[0055] Table 1

[0056] See Figure 3A ,For each network device, after obtaining the inbound traffic rate and outbound traffic rate, ,the packet conservation algorithm is used to process the inbound traffic rate and outbound traffic rate to ,obtain the node status.,The node status includes normal traffic state and abnormal traffic state.

[0057] For any type of traffic data, a network device can also measure its corresponding inbound traffic rate and outbound traffic rate separately. For example, for unicast traffic data, a network device can measure the inbound unicast traffic rate and the outbound unicast traffic rate. The unit of unicast traffic rate can be packets per second (pps).

[0058] See Figure 3B ,The network device obtains the inbound unicast traffic rate and the outbound unicast traffic rate, and uses the packet conservation algorithm to process the inbound unicast traffic rate and the outbound unicast traffic rate to obtain the node ,status.

[0059] The applicant found that since the inbound traffic rate and the outbound traffic rate correspond to traffic data at different times, there is a traffic difference. Under normal circumstances, the fluctuation conditions of the inbound traffic rate and the outbound traffic rate are similar, and they can maintain a balance state, which can be expressed by y obs,t =x obs,t ±O(1) means, where y obs,t is the inbound traffic rate at time t, x obs,tis the outbound traffic rate at time t, and O(1) can be used to represent the difference between the inbound unicast traffic rate and the outbound unicast traffic rate. In abnormal situations, the fluctuation states of the inbound traffic rate and the outbound traffic rate are greatly different.

[0060] It should be noted that there is an error between the collected data and the actual data. The actual inbound traffic rate at time t is denoted as y t , the actual outbound traffic rate at time t is recorded as x t , then the inbound rate error ε Y , the actual inbound traffic rate y t and the collected inbound traffic rate y obs,t Satisfies the following formula:

[0061] y obs,t =y t +y obs,t ε Y .

[0062] The outbound rate error is ε X , actual inbound traffic rate x t and the collected inbound traffic rate x obs,t Satisfies the following formula:

[0063] x obs,t =x t +x obs,t ε X .

[0064] Since the acquisition method is independent of time, the inbound rate error ε can be considered Y and the outbound rate error ε X It will not change over time. The above error distribution is independent and identically distributed and symmetrical about the vertical axis.

[0065] In order to describe the above equilibrium state, this application defines two variables Z 1,t and Z 2,t , which are related to the inbound traffic rate variable Y obs,t , outbound traffic rate variable X obs,t , outbound rate error variable E X and the inbound rate error variable E Y The relationship satisfies the following formula:

[0066]

[0067]

[0068] In the absence of any abnormality, and is close to 1, so it can be considered that Z 1,t and Z2,t There is no time dependency.

[0069] Z 1,t The distribution function of Z 2,t The distribution function of according to and It can be determined The range of the independent variable z includes Z 1,t The value range and Z 2,t The value range of .

[0070] The normal situation corresponds to the null hypothesis, which is The abnormal situation corresponds to the alternative hypothesis, which is Test statistic This test statistic is similar to the test statistic for the Kolmogorov-Smirnov test.

[0071] The following is a detailed introduction to the fault detection method based on the packet conservation algorithm. Figure 4 In one embodiment, the fault detection method of the present application includes:

[0072] Step 401: Obtain a set of test statistics.

[0073] In this embodiment, the test statistic value set includes multiple test statistic values. The test statistic value set can be pre-configured or calculated based on the traffic rate of the statistical period.

[0074] The process of calculating the test statistic value set is introduced below. Optionally, step 401 includes:

[0075] Step A: extracting a set of flow rate pairs from the flow rate pairs in the statistical period;

[0076] Step B: determining, according to the flow rate pair set, a sequence of input and output difference ratios corresponding to the flow rate pair set and a sequence of input and output difference ratios corresponding to the flow rate pair set;

[0077] Step C: determining a first distribution function of the flow rate pair set according to a sequence of proportions of in-out differences corresponding to the flow rate pair set;

[0078] Step D: determining a second distribution function of the traffic rate pair set according to a sequence of input and output difference ratios corresponding to the traffic rate pair set;

[0079] Step E: determining an objective function of the flow rate pair set according to the first distribution function of the flow rate pair set and the second distribution function of the flow rate pair set;

[0080] Step F: Determine that the test statistic of the set of flow rate pairs is equal to the supremum of the objective function of the set of flow rate pairs;

[0081] Steps A to F are repeatedly performed until the number of test statistic values ​​reaches a preset number of test statistic value sets.

[0082] Specifically, the flow rate pair set includes n flow rate pairs. According to each flow rate pair, the input and output difference ratio and the input and output difference ratio are calculated to obtain n input and output difference ratios and n input and output difference ratios. The i-th flow rate pair is recorded as (y′ obs,i , x′ obs,i ), the proportion of the i-th difference is recorded as z′ 1,i The proportion of the difference between the i-th input and output is recorded as z′ 2,i , which satisfy the following formula:

[0083]

[0084]

[0085] Among the n difference ratios, (-∞,z′ 1,i ] has s differences in the proportion of the difference, then z′ 1,i The corresponding first distribution function value s and n satisfy the following formula:

[0086] Among the n input and output difference ratios, (-∞,z′ 2,i ] has s′ as the percentage of the difference between input and output, then z′ 2,i The corresponding second distribution function value s′ and n satisfy the following formula:

[0087] After determining the first distribution function of the flow rate pair set and the second distribution function of the flow rate pair set according to the above formula, the target function of the flow rate pair set is determined according to the first distribution function of the flow rate pair set and the second distribution function of the flow rate pair set. The second distribution function And the objective function F″(z) satisfies the following formula:

[0088] The value interval of z includes the independent variable interval of the first distribution function of the flow rate pair set and the independent variable interval of the second distribution function of the flow rate pair set. The lower limit of the independent variable interval of the first distribution function of the flow rate pair set is less than or equal to the minimum value of the n input and output difference ratios in step 401, and the upper limit of the independent variable interval of the first distribution function of the flow rate pair set is greater than or equal to the maximum value of the n input and output difference ratios in step 401. The lower limit of the independent variable interval of the second distribution function of the flow rate pair set is less than or equal to the minimum value of the n input and output difference ratios in step 401, and the upper limit of the independent variable interval of the second distribution function of the flow rate pair set is greater than or equal to the maximum value of the n input and output difference ratios in step 401.

[0089] The supremum of the objective function is used as the test statistic of the flow rate set. Repeat the process nboot times to obtain nboot test statistics (i.e., a set of test statistics). The supremum in this application can be replaced by the maximum value.

[0090] Step 402: Obtain traffic rate pairs of the target node in multiple detection periods.

[0091] Specifically, each detection period can obtain multiple traffic rate pairs of the target node, each traffic rate pair includes the inbound traffic rate and the outbound traffic rate collected at a measurement time, and the target node can be any network node or computing device. Optionally, the starting time of the multiple detection periods is the same, the ending time of the multiple detection periods increases in sequence, and the range of the subsequent detection period is larger than that of the previous detection period. Optionally, the starting time of the multiple detection periods increases in sequence, the ending time of the multiple detection periods increases in sequence, and the duration of each detection period in the multiple detection periods is the same.

[0092] It should be noted that the traffic rate can be calculated based on all traffic data or based on a type of traffic data (e.g., unicast traffic data, multicast traffic data, or broadcast traffic data). All traffic data includes one or more of unicast traffic data, multicast traffic data, or broadcast traffic data. Optionally, the traffic rate pair is a unicast traffic rate pair. Compared to multicast traffic rates or broadcast traffic rates, the test statistic calculated based on the unicast traffic rate pair can more accurately identify abnormal traffic.

[0093] Step 403: For each detection period, determine the input / output difference ratio sequence and the input / output difference ratio sequence of the detection period according to multiple traffic rate pairs of the detection period.

[0094] During the detection period, obtain n traffic rate pairs. Calculate the input / output difference ratio and the input / output difference ratio for each traffic rate pair, resulting in n input / output difference ratios and n input / output difference ratios. Sorting the n input / output difference ratios by time yields the input / output difference ratio sequence. Sorting the n input / output difference ratios by time yields the input / output difference ratio sequence for the detection period.

[0095] Step 404: Determine the test statistic value of the detection period according to the sequence of proportions of the input and output differences of the detection period and the sequence of proportions of the input and output differences of the detection period.

[0096] Optionally, step 404 includes: determining a first distribution function of the detection period based on a sequence of input and output difference ratios of the detection period; determining a second distribution function of the detection period based on a sequence of input and output difference ratios of the detection period; determining an objective function of the detection period based on the first distribution function of the detection period and the second distribution function of the detection period; and determining that a test statistic value of the detection period is equal to the upper bound of the objective function of the detection period.

[0097] Specifically, the i-th flow rate pair is recorded as (y obs,i , x obs,i ), the proportion of the i-th difference is recorded as z 1,i The proportion of the difference between the i-th input and output is recorded as z 2,i , which satisfy the following formula:

[0098]

[0099]

[0100] Among the n difference ratios, (-∞,z 1,i ] has s percentages of discrepancies, then z 1,i The corresponding first distribution function value s and n satisfy the following formula:

[0101] Among the n input and output difference ratios, (-∞,z 2,i ] has s′ as the proportion of the input and output difference, then z 2,i The corresponding second distribution function value s′ and n satisfy the following formula:

[0102] Where i is a variable, and its value is an integer in [1, n].

[0103] After determining the first distribution function of the flow rate pair set and the second distribution function of the flow rate pair set according to the above formula, the first distribution function of the detection period The second distribution function of the detection period The objective function F′(z) of the detection period satisfies the following formula: z is the independent variable of the objective function of the detection period. The value range of z includes the independent variable range of the first distribution function of the detection period and the independent variable range of the second distribution function of the detection period. The lower limit of the independent variable range of the first distribution function of the detection period is less than or equal to the minimum value of the proportion of the n input and output differences in step 404, and the upper limit of the independent variable range of the first distribution function of the detection period is greater than or equal to the maximum value of the proportion of the n input and output differences in step 404. The lower limit of the independent variable range of the second distribution function of the detection period is less than or equal to the minimum value of the proportion of the n input and output differences in step 404, and the upper limit of the independent variable range of the second distribution function of the detection period is greater than or equal to the maximum value of the proportion of the n input and output differences in step 404.

[0104] After obtaining the objective function, traverse the values ​​of the independent variable interval to determine the upper bound of the objective function Take the supremum of the objective function as the test statistic for the test period

[0105] Step 405: Calculate the failure probability value of the detection period according to the test statistic value of the detection period and the test statistic value set.

[0106] In this application, the failure probability value of a detection period refers to the probability value of a target node failing during the detection period, and each detection period has a failure probability value of a detection period. The failure probability value of a detection period can be, but is not limited to, a P value. The P value is the probability of a result more extreme than the sample observation result occurring when the null hypothesis is true. If the P value is very small, it means that the probability of the null hypothesis occurring is very small. The smaller the P value, the more compelling the reason for rejecting the null hypothesis, indicating that the result of the hypothesis being invalid is more significant.

[0107] Specifically, when M test statistics in the test statistic set are greater than or equal to the test statistic for the detection period, the failure probability for the detection period is determined as M divided by the total number of test statistics in the test statistic set. The test statistic set represents the normal distribution of the test statistics. A larger M value results in a larger P value, indicating a lower failure probability. A smaller M value results in a smaller P value, indicating a higher failure probability. M is a positive integer and can be set based on actual conditions.

[0108] The p-value, M, and the total number of test statistics nboot included in the test statistic set satisfy the following formula:

[0109]

[0110] M can be expressed by the following formula:

[0111]

[0112] T i is the i-th test statistic in the test statistic set, is the test statistic value for the detection period.

[0113] Step 406: When the failure probability values ​​of N consecutive detection periods in the multiple detection period periods are less than or equal to the preset significance level and N is greater than or equal to the first threshold, it is determined that abnormal traffic exists at the target node.

[0114] When the failure probability values ​​for multiple detection periods are less than or equal to the preset significance level, it indicates a high probability of abnormal traffic. The first threshold can be any value between 3 and 10, and the first threshold can be set based on actual circumstances and is not limited in this application. The preset significance level can be, but is not limited to, 5%, and can be set based on actual circumstances.

[0115] In this embodiment, the test statistic can reflect the difference in traffic entering and leaving the node. The test statistics included in the test statistic set are calculated based on the normal inbound traffic rate and the normal outbound traffic rate. Therefore, the test statistic set can reflect the normal distribution of the test statistics. After calculating the test statistics for the detection period, the failure probability value for the detection period can be calculated based on the test statistics for the detection period and the normal test statistics, thereby determining whether the node traffic is abnormal. Compared with the existing difference in inbound traffic rate, the accuracy of identifying abnormal traffic based on this test statistic is higher.

[0116] The following combination Figure 5A 、 Figure 5B and Figure 5C The fault detection results of this application are introduced. In one example, in period 1, according to Figure 4 The fault detection method shown in the figure detects node 1. The p value of node 1 is as follows: Figure 5A As shown. Figure 5A It can be seen that there are consecutive p-values ​​below the significance level of 5%, so node 1 has traffic anomalies in period 1. Figure 4 The fault detection method shown in the figure detects node 2. The p value of node 2 is as follows: Figure 5B As shown. Figure 5B It can be seen that the p-values ​​are all above the 5% significance level, so node 2 is normal in period 1.

[0117] After node 1 is repaired, node 1 is tested in period 2. The p value of node 1 is as follows Figure 5C As shown. Figure 5C It can be seen that the p-values ​​are all above the 5% significance level, so node 1 is normal in period 2.

[0118] For a sample of inbound and outbound traffic rates with 720 timestamps, when calculating a test statistic set, fault detection took approximately 2.5367 seconds on a computer with an Intel i7-10700 2.90GHz CPU. With a preconfigured test statistic set, fault detection took approximately 0.25 seconds. This demonstrates the high fault detection efficiency of this application, enabling timely fault reporting and preventing silent chip failures from severely impacting services.

[0119] Equal cost multipath routing (ECMP) is a technology that implements equal cost multipath load balancing and link backup. ECMP is used in network environments where multiple links reach the same destination. Using multiple links simultaneously in such a network increases transmission bandwidth and enables data backup on failed links without delay or packet loss. When an ECMP group is configured on a network device, traffic from the ECMP group can be distributed across multiple interfaces for transmission.

[0120] In an example, the source IP addresses of four data flows are 11.23.128.16, 11.23.128.17, 11.23.130.16, and 11.23.130.17, and their destination IP address is 11.23.134.60. If their costs are the same, they can be sent through the interface of the ECMP group.

[0121] The applicant discovered that, when no interface anomalies occur, while the outbound traffic rate value on each equivalent path interface changes over time, the output traffic rate percentage of each interface remains virtually constant. When the output traffic rate percentage of a particular interface fluctuates significantly, the traffic equivalence algorithm of this application can be used to identify the fault.

[0122] The traffic rate flowing through interface j at time t is recorded as bps j,t , the traffic rate ratio of interface j at time t is recorded as q j,t , bps j,t ,q j,t And the total number of interfaces d satisfies the following formula:

[0123]

[0124] The applicant believes that when there is no abnormality in the ECMP group, the traffic rate of interface i accounts for and the expected value of the traffic rate ratio of interface i Satisfies the following formula: This formula can be considered as the null hypothesis H0. In the case of abnormality in the ECMP group, This formula can be considered as the alternative hypothesis H A .

[0125] See Figure 7 This application can obtain the output traffic rate ratios of multiple ECMP group interfaces and then use a traffic equivalence algorithm to determine the status of the ECMP group. Specifically, based on the above assumptions, this application constructs a chi-square test statistic, calculates the failure probability value at each measurement time based on the chi-square test statistic, and then determines whether the ECMP group traffic is abnormal based on the failure probability value at the measurement time.

[0126] The following describes this method in detail. Figure 8 In an optional embodiment, the fault detection method of the present application further includes:

[0127] Step 801: Determine the output traffic rate proportion of each interface at multiple measurement times in the interface group corresponding to the ECMP group.

[0128] Step 802: Determine the chi-square test statistic at the measurement time based on the output flow rate ratio of the interface at the measurement time, the expected value of the output flow rate ratio of the interface, the sample variance of the output flow rate ratio of the interface, and the total number of flows passing through the interface group.

[0129] Optionally, the output traffic rate ratio of the interface at the measurement time, the expected value of the output traffic rate ratio of the interface, the sample variance of the output traffic rate ratio of the interface, the total number of flows passing through the interface group, and the chi-square test statistic at the measurement time satisfy the following formula:

[0130]

[0131] T is the chi-square test statistic of the target interface at the measurement time, is the output traffic rate ratio of the i-th interface, is the expected value of the output traffic rate ratio of the i-th interface, is the sample variance of the interface traffic rate ratio of the i-th interface, d is the total number of interfaces in the interface group, and n is the total number of flows passing through the interface group.

[0132] For the i-th interface, the output traffic rate ratio of the interface, the expected value of the output traffic rate ratio of the interface, and the sample variance of the interface traffic rate ratio satisfy the following formula:

[0133]

[0134] k is the number of measurement moments, j is a variable and the value of j is a positive integer in [1, k].

[0135] The expected value of the output traffic rate ratio of each interface can be determined based on the output traffic rate ratio of k measurement moments during a normal period. The specific formula is as follows:

[0136]

[0137] Optionally, the time interval between adjacent measurement moments in the k measurement moments during the normal period is greater than a preset time interval (e.g., 10 minutes). The output flow rate ratio collected at the preset time interval meets the requirements of a stationary time series and has no time correlation. It should be understood that the above time intervals are illustrative and can be adjusted according to actual conditions.

[0138] Step 803: Calculate the fault probability value at the measurement time according to the chi-square test statistic value at the measurement time.

[0139] The chi-square test statistic is correlated with the output traffic rate ratio of all interfaces and can indicate whether the output traffic rate ratio of the ECMP group deviates from the expected value. A smaller chi-square test statistic indicates a smaller deviation between the output traffic rate ratio of the ECMP group and the expected value, and a lower probability of failure. A larger chi-square test statistic indicates a greater deviation between the output traffic rate ratio of the ECMP group and the expected value, and a higher probability of failure.

[0140] Optionally, the chi-square test statistic value and the failure probability value at the measurement time satisfy the following formula: p′ is the failure probability value at the measurement time, is the chi-square test statistic at the measurement moment.

[0141] Step 804: If the failure probability values ​​at L consecutive measurement times among the multiple measurement times are less than or equal to the preset significance level, and L is greater than or equal to a second threshold, traffic abnormality of the ECMP group is determined. The second threshold can be any integer between 3 and 10. The value of L can be set based on actual conditions and is not limited in this application.

[0142] This embodiment can obtain a chi-square test statistic related to the output traffic rate ratio of the interface. Since the chi-square test statistic can reflect whether the output traffic rate ratio of the ECMP group deviates from the expected output traffic rate ratio, the failure probability value at each measurement time can be calculated based on the chi-square test statistic. Then, based on the failure probability values ​​at multiple measurement times, it can be determined whether the traffic of the ECMP group is abnormal.

[0143] In an optional embodiment, the fault detection method of the present application further includes: when the chi-square test statistic value of the target interface is greater than or equal to a third threshold, determining that the traffic of the target interface is abnormal.

[0144] In this embodiment, the target interface is any one of the interface groups. The third threshold can be pre-configured or can be configured based on the The formula is calculated, Represents the inverse function of the chi-square distribution with 1 degree of freedom. The value of α can be, but is not limited to, 0.05 and can be set based on actual conditions.

[0145] The following describes the fault detection results of the ECMP group. In one embodiment, the ECMP group includes four interfaces, and the traffic rates of interface 1, interface 2, and interface 3 account for Figure 8 As shown. Figure 8 Not shown, the traffic rate ratio of interface 4 is almost the same as the traffic rate ratio of interface 1 .

[0146] See Figure 8 Before interface 3 failed, the traffic rate of interface 1 and interface 4 were both between 0.3% and 0.35%, respectively. The traffic rate of interfaces 2 and 3 were both between 0.15% and 0.2%.

[0147] After the failure of interface 3 occurs, the traffic rate proportions of interface 1 and interface 4 are 0.35-0.4, the traffic rate proportions of interface 2 are 0.2-0.25, and the traffic rate proportions of interface 3 are 0-0.05.

[0148] As can be seen, after the failure of interface 3, a continuous failure probability indicator appears, indicating that the p-value is below the significance level. Before the failure of interface 3, the failure probability indicator was discontinuous, which is usually caused by normal fluctuations in the interface traffic rate ratio.

[0149] Taking the Intel i7-10700 processor with a main frequency of 2.9 GHz as an example, for the interface of the ECMP group, after obtaining the output traffic rate samples of 571 timestamps, the processor executes Figure 8 The method of the embodiment shown takes 0.8035 seconds. It can be seen that the fault detection efficiency of the present application is very high, and faults can be reported in a timely manner to prevent silent chip faults from causing serious impact on services.

[0150] See Figure 10 In one embodiment, the computing device 1000 of the present application includes an acquisition module 1001 , a measurement module 1002 , a calculation module 1003 and a judgment module 1004 .

[0151] The acquisition module 1001 is used to obtain a set of test statistics values;

[0152] The measurement module 1002 is configured to obtain a traffic rate pair of a target node during a plurality of detection periods, each traffic rate pair including an inbound traffic rate and an outbound traffic rate collected at a measurement moment;

[0153] The calculation module 1003 is configured to determine, for each detection period, a sequence of input-output difference ratios and a sequence of input-output difference ratios for the detection period based on multiple traffic rate pairs during the detection period; determine a test statistic for the detection period based on the sequence of input-output difference ratios and the sequence of input-output difference ratios for the detection period; and calculate a failure probability value for the detection period based on the test statistic and the set of test statistics for the detection period.

[0154] The judgment module 1004 is configured to determine that abnormal traffic exists at the target node when N consecutive failure probability values ​​in the failure probability values ​​of the multiple detection periods are less than or equal to a preset significance level and N is greater than or equal to a first threshold.

[0155] Acquisition module 1001, measurement module 1002, calculation module 1003, and judgment module 1004 can all be implemented in software or hardware. For example, the following describes the implementation of acquisition module 1001, taking acquisition module 1001 as an example. Similarly, the implementation of measurement module 1002, calculation module 1003, and judgment module 1004 can refer to the implementation of acquisition module 1001.

[0156] As an example of a software functional unit, the acquisition module 1001 may include code running on a computing instance. The computing instance may include at least one of a physical host (computing device), a virtual machine, and a container. Furthermore, the computing instance may be one or more. For example, the acquisition module 1001 may include code running on multiple hosts / virtual machines / containers. It should be noted that the multiple hosts / virtual machines / containers used to run the code may be distributed in the same region or in different regions. Furthermore, the multiple hosts / virtual machines / containers used to run the code may be distributed in the same availability zone (AZ) or in different AZs, each AZ including one data center or multiple geographically close data centers. Typically, a region may include multiple AZs.

[0157] Similarly, multiple hosts / virtual machines / containers running the code can be distributed within the same virtual private cloud (VPC) or across multiple VPCs. Typically, a VPC is set up within a region. Cross-region communication between two VPCs within the same region, or between VPCs in different regions, requires a communication gateway within each VPC to interconnect the VPCs.

[0158] As an example of a hardware functional unit, the acquisition module 1001 may include at least one computing device, such as a server. Alternatively, the acquisition module 1001 may be implemented using an application-specific integrated circuit (ASIC) or a programmable logic device (PLD). The PLD may be a complex programmable logical device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL), or any combination thereof.

[0159] The multiple computing devices included in acquisition module 1001 can be distributed in the same region or in different regions. The multiple computing devices included in acquisition module 1001 can be distributed in the same AZ or in different AZs. Similarly, the multiple computing devices included in acquisition module 1001 can be distributed in the same VPC or in multiple VPCs. The multiple computing devices can be any combination of servers, ASICs, PLDs, CPLDs, FPGAs, GALs, and other computing devices.

[0160] It should be noted that, in other embodiments, the acquisition module 1001 may be used to execute Figure 4 The embodiment shown or Figure 8 In any step of the fault detection method in the embodiment shown, the measurement module 1002 can be used to perform Figure 4 The embodiment shown or Figure 8 In any step of the fault detection method in the embodiment shown, the calculation module 1003 can be used to perform Figure 4 The embodiment shown or Figure 8 In any step of the fault detection method in the embodiment shown, the judgment module 1004 can be used to perform Figure 4 The embodiment shown or Figure 8For any step in the fault detection method in the illustrated embodiment, the steps that each module is responsible for implementing can be specified as needed, and all functions of the computing device can be realized by each module implementing different steps in the fault detection method.

[0161] The present application also provides a computing device 1100, such as Figure 11 As shown, computing device 1100 includes a bus 1102, a processor 1104, a memory 1106, and a communication interface 1108. Processor 1104, memory 1106, and communication interface 1108 communicate with each other via bus 1102. Computing device 1100 can be a server or a terminal device. It should be understood that this application does not limit the number of processors and memories in computing device 1100.

[0162] The bus 1102 may be a peripheral component interconnect (PCI) bus or an extended industry standard architecture (EISA) bus. The bus may be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 11 The bus 1104 may include a path for transmitting information between various components of the computing device 1100 (eg, memory 1106, processor 1104, communication interface 1108).

[0163] The processor 1104 may include any one or more processors such as a central processing unit (CPU), a graphics processing unit (GPU), a microprocessor (MP), or a digital signal processor (DSP).

[0164] The memory 1106 may include volatile memory, such as random access memory (RAM). The processor 1104 may also include non-volatile memory, such as read-only memory (ROM), flash memory, a hard disk drive (HDD), or a solid state drive (SSD).

[0165] Memory 1106 stores executable program code, which processor 1104 executes to implement the functions of acquisition module 1001, measurement module 1002, calculation module 1003, and judgment module 1004, thereby implementing the fault detection method. Specifically, memory 1106 stores instructions for executing the fault detection method.

[0166] The communication interface 1103 uses a transceiver module such as, but not limited to, a network interface card or a transceiver to implement communication between the computing device 1100 and other devices or a communication network.

[0167] Embodiments of the present application also provide a computing device cluster. The computing device cluster includes at least one computing device. The computing device can be a server, such as a central server, an edge server, or a local server in a local data center. In some embodiments, the computing device can also be a terminal device such as a desktop computer, a laptop computer, or a smartphone.

[0168] like Figure 12 As shown, the computing device cluster includes at least one computing device 1100. The memory 1106 in one or more computing devices 1100 in the computing device cluster may store the same instructions for executing the fault detection method.

[0169] In some possible implementations, the memory 1106 of one or more computing devices 1100 in the computing device cluster may also store some instructions for executing the fault detection method. In other words, the combination of one or more computing devices 1100 can jointly execute the instructions for executing the fault detection method.

[0170] It should be noted that the memory 1106 in different computing devices 1100 in the computing device cluster can store different instructions, each for performing a portion of the functions of the computing device. In other words, the instructions stored in the memory 1106 in different computing devices 1100 can implement the functions of one or more of the acquisition module 1001, measurement module 1002, calculation module 1003, and judgment module 1004.

[0171] In some possible implementations, one or more computing devices in a computing device cluster may be connected via a network, which may be a wide area network or a local area network. Figure 13 A possible implementation is shown. Figure 13As shown, two computing devices 1100A and 1100B are connected via a network. Specifically, the connection to the network is achieved through a communication interface within each computing device. In this possible implementation, the memory 1106 within computing device 1100A stores instructions for executing the functions of the acquisition module. Simultaneously, the memory 1106 within computing device 1100B stores instructions for executing the functions of the measurement module, the calculation module, and the judgment module.

[0172] Embodiments of the present application also provide a computer program product including instructions. The computer program product may be software or a program product including instructions that can be run on a computing device or stored on any available medium. When the computer program product is run on at least one computing device, the at least one computing device executes a fault detection method.

[0173] The present application also provides a computer-readable storage medium. The computer-readable storage medium can be any available medium that can be stored by a computer or a data storage device such as a data center that contains one or more available media. The available medium can be a magnetic medium (e.g., a floppy disk, a hard disk, a magnetic tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state drive). The computer-readable storage medium includes instructions that instruct the computer to execute the fault detection method.

[0174] The present application also provides a chip system, which includes a processor and a memory coupled to each other. The memory is used to store computer programs or instructions, and the processor is used to execute the computer programs or instructions stored in the memory, so that the computer executes the steps performed by the acquisition module, measurement module, calculation module or judgment module in the above embodiment. Optionally, the memory is a memory within the chip, such as a register, cache, etc. The memory can also be a memory located outside the chip within the site, such as a read-only memory or other types of static storage devices that can store static information and instructions, random access memory, etc. The processor mentioned in any of the above places can be a general-purpose central processing unit, a microprocessor, a dedicated integrated circuit or one or more integrated circuits for implementing the above-mentioned fault detection method.

[0175] The above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present application.

Claims

1. A fault detection method, characterized in that: include: Get the test statistic value set; Obtaining traffic rate pairs of the target node during a plurality of detection periods, each of the traffic rate pairs comprising an inbound traffic rate and an outbound traffic rate collected at a measurement moment; For each detection period, determining, according to a plurality of flow rate pairs of the detection period, a sequence of proportions of inbound and outbound differences and a sequence of proportions of inbound and outbound differences in the detection period; Determining a test statistic value for the detection period according to a sequence of proportions of input and output differences in the detection period and a sequence of proportions of input and output differences in the detection period; Calculating a failure probability value for the detection period based on the test statistic value for the detection period and the test statistic value set; When the failure probability values ​​of N consecutive detection periods in the failure probability values ​​of the multiple detection periods are less than or equal to the preset significance level and N is greater than or equal to the first threshold, it is determined that abnormal traffic exists in the target node.

2. The method according to claim 1, characterized in that The step of determining the test statistic value of the detection period according to the sequence of proportions of the input and output difference values ​​of the detection period and the sequence of proportions of the input and output difference values ​​of the detection period includes: Determine a first distribution function of the detection period according to a sequence of proportions of the difference between the detection period and the input and output values; Determining a second distribution function of the detection period according to a sequence of proportions of input and output differences in the detection period; The target function of the detection period is determined according to the first distribution function of the detection period and the second distribution function of the detection period, wherein the first distribution function of the detection period The second distribution function of the detection period The objective function F′(z) of the detection period satisfies the following formula: z is the independent variable of the objective function during the detection period; A test statistic value of the detection period is determined to be equal to the supremum of the objective function of the detection period.

3. The method according to claim 1, characterized in that Calculating the failure probability value of the detection period based on the test statistic value of the detection period and the test statistic value set includes: When M test statistics in the test statistic value set are greater than or equal to the test statistic value of the detection period, the failure probability value of the detection period is determined to be M divided by the total number of test statistics in the test statistic value set.

4. The method according to claim 1, characterized in that The obtaining of the test statistic value set includes: Step A: extracting a set of flow rate pairs from the flow rate pairs in the statistical period; Step B: determining, according to the flow rate pair set, a sequence of input and output difference ratios corresponding to the flow rate pair set and a sequence of input and output difference ratios corresponding to the flow rate pair set; Step C: determining a first distribution function of the flow rate pair set according to a sequence of proportions of in-out differences corresponding to the flow rate pair set; Step D: determining a second distribution function of the traffic rate pair set according to a sequence of input and output difference ratios corresponding to the traffic rate pair set; Step E: determining the objective function of the flow rate pair set according to the first distribution function of the flow rate pair set and the second distribution function of the flow rate pair set; Step F: Determine that the test statistic of the flow rate pair set is equal to the supremum of the objective function of the flow rate pair set; Steps A to F are repeatedly performed until the number of test statistic values ​​reaches a preset number of the test statistic value set.

5. The method according to any one of claims 1 to 4, characterized in that The method further comprises: Determine the output traffic rate proportion of each interface at multiple measurement times in the interface group corresponding to the equal-cost multi-path routing ECMP group; Determine a chi-square test statistic at the measurement time based on the output traffic rate proportion of the interface at the measurement time, the expected value of the output traffic rate proportion of the interface, the sample variance of the output traffic rate proportion of the interface, and the total number of flows passing through the interface group; Calculating a failure probability value at the measurement moment according to a chi-square test statistic value at the measurement moment; When the failure probability values ​​at L consecutive measurement moments among the multiple failure probability values ​​at the measurement moments are less than or equal to a preset significance level and L is greater than or equal to a second threshold, it is determined that the traffic of the ECMP group is abnormal.

6. The method according to claim 5, characterized in that The output traffic rate proportion of the interface at the measurement time, the expected value of the output traffic rate proportion of the interface, the sample variance of the output traffic rate proportion of the interface, the total number of flows passing through the interface group, and the chi-square test statistic at the measurement time satisfy the following formula: T is the chi-square test statistic at the measurement moment, is the output traffic rate ratio of the i-th interface, is the expected value of the output traffic rate ratio of the i-th interface, is the sample variance of the interface traffic rate ratio of the i-th interface, d is the total number of interfaces in the interface group, and n is the total number of flows passing through the interface group.

7. The method according to claim 6, characterized in that The chi-square test statistic value at the measurement time and the failure probability value at the measurement time satisfy the following formula: p′ is the failure probability value at the measurement moment, is the chi-square test statistic at the measurement moment.

8. The method according to claim 7, characterized in that The method further comprises: When the chi-square test statistic value of the target interface is greater than or equal to a third threshold, it is determined that traffic of the target interface is abnormal, and the target interface is any one of the interface group.

9. A computing device, characterized in that: include: An acquisition module is used to obtain a set of test statistic values; a measurement module, configured to obtain a flow rate pair of a target node during a plurality of detection periods, each flow rate pair comprising an inbound flow rate and an outbound flow rate collected at a measurement moment; A calculation module, configured to determine, for each detection period, a sequence of input-output difference ratios and a sequence of input-output difference ratios for the detection period according to a plurality of flow rate pairs in the detection period; Determine the test statistic value of the detection period according to the sequence of proportions of the input and output differences of the detection period and the sequence of proportions of the input and output differences of the detection period; Calculating a failure probability value for the detection period based on the test statistic value for the detection period and the test statistic value set; The judgment module is used to determine that abnormal traffic exists in the target node when the failure probability values ​​of N consecutive detection periods are less than or equal to a preset significance level among the failure probability values ​​of multiple detection periods and N is greater than or equal to a first threshold.

10. The computing device according to claim 9, wherein: The calculation module is specifically configured to determine a first distribution function of the detection period according to a sequence of proportions of input and output differences of the detection period; and determine a second distribution function of the detection period according to a sequence of proportions of input and output differences of the detection period; The target function of the detection period is determined according to the first distribution function of the detection period and the second distribution function of the detection period, wherein the first distribution function The second distribution function The objective function F′(z) of the detection period satisfies the following formula: z is the independent variable of the objective function during the detection period; A test statistic value of the detection period is determined to be equal to the supremum of the objective function of the detection period.

11. The computing device according to claim 9, wherein: The calculation module is specifically configured to determine, when M test statistics in the test statistics set are greater than or equal to the test statistics of the detection period, a failure probability value of the detection period as M divided by the total number of test statistics in the test statistics set.

12. The computing device according to claim 9, wherein: The acquisition module is specifically configured to perform the following steps: Step A: extracting a set of flow rate pairs from the flow rate pairs in the statistical period; Step B: determining, according to the flow rate pair set, a sequence of input and output difference ratios corresponding to the flow rate pair set and a sequence of input and output difference ratios corresponding to the flow rate pair set; Step C: determining a first distribution function of the flow rate pair set according to a sequence of proportions of in-out differences corresponding to the flow rate pair set; Step D: determining a second distribution function of the traffic rate pair set according to a sequence of input and output difference ratios corresponding to the traffic rate pair set; Step E: determining the objective function of the flow rate pair set according to the first distribution function of the flow rate pair set and the second distribution function of the flow rate pair set; Step F: Determine that the test statistic of the flow rate pair set is equal to the supremum of the objective function of the flow rate pair set; Steps A to F are repeatedly performed until the number of test statistic values ​​reaches a preset number of the test statistic value set.

13. The computing device according to any one of claims 9 to 12, characterized in that The measurement module is further configured to determine the output flow rate proportion of each interface at multiple measurement moments in the interface group corresponding to the equal cost multipath routing ECMP group; The calculation module is further configured to determine, for each measurement moment, a chi-square test statistic at the measurement moment based on the output flow rate ratio of the interface at the measurement moment, the expected value of the output flow rate ratio of the interface, the sample variance of the output flow rate ratio of the interface, and the total number of flows passing through the interface group; and calculate a failure probability value at the measurement moment based on the chi-square test statistic at the measurement moment; The judgment module is further configured to determine that the flow of the ECMP group is abnormal when the failure probability values ​​at L consecutive measurement moments among the failure probability values ​​at the multiple measurement moments are less than or equal to a preset significance level and L is greater than or equal to a second threshold.

14. The computing device according to claim 13, wherein: The output traffic rate proportion of the interface at the measurement time, the expected value of the output traffic rate proportion of the interface, the sample variance of the output traffic rate proportion of the interface, the total number of flows passing through the interface group, and the chi-square test statistic at the measurement time satisfy the following formula: T is the chi-square test statistic at the measurement moment, is the output traffic rate ratio of the i-th interface, is the expected value of the output traffic rate ratio of the i-th interface, is the sample variance of the interface traffic rate ratio of the i-th interface, d is the total number of interfaces in the interface group, and n is the total number of flows passing through the interface group.

15. The computing device according to claim 14, wherein: The chi-square test statistic value at the measurement time and the failure probability value at the measurement time satisfy the following formula: p′ is the failure probability value at the measurement moment, is the chi-square test statistic at the measurement moment.

16. The computing device according to claim 13, wherein: The judgment module is further configured to determine that traffic of the target interface is abnormal when a chi-square test statistic value of the target interface is greater than or equal to a third threshold, and the target interface is any one of the interface group.

17. A computing device cluster, characterized in that: comprising at least one computing device, each of said computing devices comprising a processor and a memory; The processor of the at least one computing device is configured to execute instructions stored in a memory of the at least one computing device, so that the computing device cluster executes the method according to any one of claims 1 to 8.

18. A computer-readable storage medium, characterized in that The method comprises computer program instructions, and when the computer program instructions are executed by a computing device, the computing device performs the method according to any one of claims 1 to 8.

19. A computer program product comprising instructions, characterized in that When the instructions are executed by a computing device, the computing device is caused to perform the method according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Data detection method, device and system

    CN107888455A

  • Industrial safety protection system based on flow analysis and control

    CN115174211A