Network Fault Location Method and Device
By collecting and analyzing the data traffic of the data center edge access switch, the problem of network failure location in the data center is solved, and the effect of accurately determining the cause of boundary failure and reducing the amount of data acquisition is achieved.
Patent Information
- Application Number
- CN202210507852.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-05-11
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2042-05-11
AI Technical Summary
Due to the lack of storage service network modeling in the data center, the correlation between network performance and storage service quality is unclear. The existing monitoring methods have fine granularity in measurement but low correlation with services, resulting in difficulty in fault location.
By collecting data traffic forwarded by the edge access switch, packet analysis and feature analysis are carried out to determine the possibility of congestion in data traffic, and determining the location of network failure through correlation matching and life cycle identification.
It realizes accurate determination of the cause of bound network failure, reduces the total amount of collected data, avoids data acquisition bottlenecks, and accurately distinguishes and delimits the causes of abnormalities based on the decomposition of the total delay of data traffic.
Smart Images

Figure CN115604089B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of data processing and can also be used in the financial field. Specifically, it relates to a network fault location method and device. Background Art
[0002] In data center storage services, there is a common risk of network congestion leading to a large number of packet losses across the entire link and a decrease in business transaction volume. Currently, there is a lack of network modeling for storage services in data centers, resulting in an unclear correlation between network performance and the quality of storage services. Network metrics such as packet loss and throughput can only explain some basic network problems. The current mainstream basic network monitoring methods have a fine measurement granularity but a very low correlation with storage services, making it difficult to delimit specific problems in storage services, and it is very difficult to explain business performance problems, unable to meet the needs of the business. Summary of the Invention
[0003] In view of the problems in the prior art, this application provides a network fault location method and device that can accurately delimit the cause of network faults.
[0004] To solve at least one of the above problems, this application provides the following technical solutions:
[0005] In a first aspect, this application provides a network fault location method, including:
[0006] Collect the data traffic forwarded by the edge access switch and perform packet analysis to determine the data traffic that may be congested;
[0007] Determine whether the data traffic that may be congested is a microburst data stream. If not, perform correlation matching and life cycle identification on the round-trip characteristic packets of the data traffic;
[0008] After the correlation matching and life cycle identification are completed, perform delay analysis on the data traffic to determine the network fault location.
[0009] Further, the step of collecting the data traffic forwarded by the edge access switch and performing packet analysis to determine the data traffic that may be congested includes:
[0010] Real-time monitor the data traffic forwarded by the edge access switch of the storage service and collect the packet header information of the data traffic;
[0011] Perform feature analysis on the packet header information of the data traffic to determine the data traffic that may be congested.
[0012] Further, the step of determining whether the data traffic that may be congested is a microburst data stream includes:
[0013] Judging the microburst data flow of the data traffic with possible congestion according to the chip cache queue depth information read within the set time period.
[0014] Further, the correlation matching and life cycle identification of the round-trip characteristic packets of the data traffic include:
[0015] Performing correlation matching on the round-trip characteristic packets of the data traffic on the edge access switch processors respectively accessed on the client side and the storage array side;
[0016] Identifying the end-side data access delay, data preparation delay, data transmission delay, and data confirmation delay in the life cycle of the data traffic.
[0017] Further, the delay analysis of the data traffic after the correlation matching and life cycle identification includes:
[0018] Distinguishing the large and small flow types of the data traffic according to the correlation matching result of the round-trip characteristic packets;
[0019] If the data traffic belongs to the small flow type, it is sent to the set collection server for delay analysis by means of data mirroring;
[0020] If the data traffic belongs to the large flow type, it is preprocessed first and then sent to the set collection server for delay analysis.
[0021] Further, the delay analysis of the data traffic after the correlation matching and life cycle identification to determine the network fault location includes:
[0022] If the delay ratio of the data transmission delay exceeds the threshold, it is determined that a fault occurs on the network side, and the network switch bandwidth utilization rate and packet loss rate are sent to the network operation and maintenance end for emergency handling;
[0023] If the delay ratio of the data preparation delay and / or data confirmation delay exceeds the threshold, it is determined that a fault occurs on the end side, and the system operation and maintenance end is notified to check the client and the storage array.
[0024] In a second aspect, the present application provides a network fault location device, including:
[0025] An information collection module, configured to collect the data traffic forwarded by the edge access switch and perform packet analysis to determine the data traffic with possible congestion;
[0026] A data volume correlation module, configured to judge whether the data traffic with possible congestion is a microburst data flow. If not, perform correlation matching and life cycle identification on the round-trip characteristic packets of the data traffic;
[0027] A delay fault analysis module, which is used to perform delay analysis on the data traffic after the correlation matching and life cycle identification are completed, so as to determine the network fault location.
[0028] Further, the information collection module includes:
[0029] A packet header information collection unit, which is used to monitor the data traffic forwarded by the storage service edge access switch in real time and collect the packet header information of the data traffic;
[0030] A packet header feature analysis unit, which is used to perform feature analysis on the packet header information of the data traffic to determine the data traffic with congestion possibility.
[0031] Further, the data volume correlation module includes:
[0032] A microburst judgment unit, which is used to judge the microburst data stream of the data traffic with congestion possibility according to the chip cache queue depth information read within a set time period.
[0033] Further, the data volume correlation module includes:
[0034] A correlation matching unit, which is used to perform correlation matching on the round-trip feature packets of the data traffic on the edge access switch processors accessed on the client side and the storage array side respectively;
[0035] A life cycle identification unit, which is used to identify the end-side data access delay, data preparation delay, data transmission delay, and data confirmation delay in the data traffic life cycle.
[0036] Further, the delay fault analysis module includes:
[0037] A large / small flow differentiation unit, which is used to differentiate the data traffic into large / small flow types according to the correlation matching result of the round-trip feature packets;
[0038] A small flow uploading unit, which is used to upload the data traffic to a set collection server for delay analysis in the form of data mirroring if the data traffic belongs to the small flow type;
[0039] A large flow uploading unit, which is used to upload the data traffic to a set collection server for delay analysis after data preprocessing if the data traffic belongs to the large flow type.
[0040] Further, the delay fault analysis module includes:
[0041] The network - side fault location unit is used to determine that a network - side fault occurs if the delay ratio of the data transmission delay exceeds a threshold, and send the network switch bandwidth utilization rate and packet loss rate to the network operation and maintenance end for emergency handling;
[0042] The terminal - side fault location unit is used to determine that a terminal - side fault occurs if the delay ratio of the data preparation delay and / or data confirmation delay exceeds a threshold, and notify the system operation and maintenance end to check the client and the storage array.
[0043] In a third aspect, the present application provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the steps of the network fault location method are implemented.
[0044] In a fourth aspect, the present application provides a computer - readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the network fault location method are implemented.
[0045] In a fifth aspect, the present application provides a computer program product, including a computer program / instructions. When the computer program / instructions are executed by a processor, the steps of the network fault location method are implemented.
[0046] As can be seen from the above technical solutions, the present application provides a network fault location method and device. By collecting the data traffic forwarded by the edge access switch and performing packet analysis, the micro - burst data stream is accurately partitioned, the total amount of collected data is greatly reduced, the data collection bottleneck is avoided, and the abnormal causes are accurately distinguished and delimited according to the decomposition of the total data traffic delay. BRIEF DESCRIPTION OF THE DRAWINGS
[0047] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0048] Figure 1 It is one of the flow diagrams of the network fault location method in the embodiments of the present application;
[0049] Figure 2 It is another flow diagram of the network fault location method in the embodiments of the present application;
[0050] Figure 3 It is yet another flow diagram of the network fault location method in the embodiments of the present application;
[0051] Figure 4The fourth flowchart of the network fault location method in the embodiments of the present application;
[0052] Figure 5 The fifth flowchart of the network fault location method in the embodiments of the present application;
[0053] Figure 6 The first structure diagram of the network fault location device in the embodiments of the present application;
[0054] Figure 7 The second structure diagram of the network fault location device in the embodiments of the present application;
[0055] Figure 8 The third structure diagram of the network fault location device in the embodiments of the present application;
[0056] Figure 9 The fourth structure diagram of the network fault location device in the embodiments of the present application;
[0057] Figure 10 The fifth structure diagram of the network fault location device in the embodiments of the present application;
[0058] Figure 11 The sixth structure diagram of the network fault location device in the embodiments of the present application;
[0059] Figure 12 The flowchart of the network fault location method in a specific embodiment of the present application;
[0060] Figure 13 The structure diagram of the electronic device in the embodiments of the present application. Detailed implementation manners
[0061] To make the objectives, technical solutions and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are some but not all of the embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the scope of protection of the present application.
[0062] In the technical solutions of the present application, the acquisition, storage, use, processing, etc. of data all comply with the relevant regulations of national laws and regulations.
[0063] In view of the problem of inaccurate network fault location in the prior art, the present application provides a network fault location method and device. By collecting the data traffic forwarded by the edge access switch and performing packet analysis, the microburst data stream is accurately partitioned, the total amount of collected data is greatly reduced, the data collection bottleneck is avoided, and the abnormal causes are accurately distinguished and delimited according to the decomposition of the total delay of the data traffic.
[0064] In order to accurately delimit the cause of network faults, the present application provides an embodiment of a network fault location method. Refer to Figure 1 , the network fault location method specifically includes the following content:
[0065] Step S101: Collect the data traffic forwarded by the edge access switch and perform packet analysis to determine the data traffic that may have congestion.
[0066] Optionally, the present application can monitor in real time the data traffic forwarded by the storage service edge access switch and collect the packet header information of the data traffic, perform feature analysis on the packet header information of the data traffic, and determine the data traffic that may have congestion.
[0067] For example, the present application can analyze the qos statistical information of the storage service data traffic. If congestion control identification information such as ECN, CNP, and PFC appears, it indicates that the business data IO may be congested.
[0068] Step S102: Determine whether the data traffic that may have congestion is a microburst data stream. If not, perform correlation matching and life cycle identification on the round-trip characteristic packets of the data traffic.
[0069] Optionally, the present application can determine whether the data traffic that may have congestion is a microburst data stream according to the chip cache queue depth information read within a set time period.
[0070] For example, if individual peaks appear in the queue depth information read in a certain period and no queue depth peaks appear in adjacent periods, it indicates that there is a microburst IO.
[0071] Optionally, the present application can also perform correlation matching on the round-trip characteristic packets of the data traffic on the edge access switch processors connected to the client side and the storage array side respectively. Identify the end-side data access delay, data preparation delay, data transmission delay, and data confirmation delay in the life cycle of the data traffic.
[0072] Specifically, performing correlation matching on the round-trip characteristic packets of the data traffic means: inserting a unique identification field into the packet header based on the storage network protocol capabilities, and the round-trip characteristic packets in an IO operation are correlated and matched according to the packet header identification field information.
[0073] Step S103: After the correlation matching and life cycle identification are completed, perform a delay analysis on the data traffic to determine the network fault location.
[0074] In an embodiment of the present application, the present application can distinguish the size flow type of the data traffic according to the correlation matching result of the round-trip characteristic message; if the data traffic belongs to the small traffic type, it is sent to the set collection server for delay analysis through the data mirroring method; if the data traffic belongs to the large traffic type, it is pre-processed first and then sent to the set collection server for delay analysis.
[0075] Optionally, the delay analysis is, for example, the respective proportions of each identifier (access delay of end-side data, data preparation delay, data transmission delay, data confirmation delay).
[0076] Specifically, if the delay proportion of the data transmission delay exceeds the threshold, it is determined that a fault has occurred on the network side, and the network switch bandwidth utilization rate and packet loss rate are sent to the network operation and maintenance side for emergency handling. If the delay proportion of the data preparation delay and / or the data confirmation delay exceeds the threshold, it is determined that a fault has occurred on the end side, and the system operation and maintenance side is notified to check the client and the storage array.
[0077] As can be seen from the above description, the network fault location method provided by the embodiment of the present application can accurately partition the microburst data stream by collecting the data traffic forwarded by the edge access switch and performing message analysis, greatly reducing the total amount of collected data, avoiding the data collection bottleneck, and accurately distinguishing and delimiting the abnormal cause according to the decomposition of the total delay of the data traffic.
[0078] In order to be able to determine the data traffic with possible congestion, in an embodiment of the network fault location method of the present application, refer to Figure 2 , the above step S101 may further specifically include the following contents:
[0079] Step S201: Real-time monitor the data traffic forwarded by the storage service edge access switch and collect the message header information of the data traffic.
[0080] Step S202: Perform feature analysis on the message header information of the data traffic to determine the data traffic with possible congestion.
[0081] In order to be able to screen the microburst data stream, in an embodiment of the network fault location method of the present application, the above step S102 may further specifically include the following contents:
[0082] Judge the microburst data stream for the data traffic with possible congestion according to the chip cache queue depth information read within the set time period.
[0083] In order to accurately perform correlation matching and life cycle identification, in an embodiment of the network fault location method of the present application, refer to Figure 3 , the above step S102 may specifically include the following contents:
[0084] Step S301: Perform correlation matching on the round-trip characteristic packets of the data traffic on the edge access switch processors accessed on the client side and the storage array side respectively.
[0085] Step S302: Identify the end-side data access delay, data preparation delay, data transmission delay, and data confirmation delay in the life cycle of the data traffic.
[0086] In order to accurately perform delay analysis, in an embodiment of the network fault location method of the present application, refer to Figure 4 , the above step S103 may specifically include the following contents:
[0087] Step S401: Distinguish the size flow types of the data traffic according to the correlation matching results of the round-trip characteristic packets.
[0088] Step S402: If the data traffic belongs to the small flow type, upload it to the set acquisition server for delay analysis through the data mirroring method.
[0089] Step S403: If the data traffic belongs to the large flow type, perform data preprocessing first and then upload it to the set acquisition server for delay analysis.
[0090] In order to accurately determine the network fault location, in an embodiment of the network fault location method of the present application, refer to Figure 5 , the above step S103 may specifically include the following contents:
[0091] Step S501: If the delay ratio of the data transmission delay exceeds the threshold, it is determined that a fault occurs on the network side, and the network switch bandwidth utilization rate and packet loss rate are sent to the network operation and maintenance end for emergency processing.
[0092] Step S502: If the delay ratio of the data preparation delay and / or the data confirmation delay exceeds the threshold, it is determined that a fault occurs on the end side, and the system operation and maintenance end is notified to check the client and the storage array.
[0093] In order to accurately delimit the cause of the network fault, the present application provides an embodiment of a network fault location device for implementing all or part of the content of the network fault location method, refer to Figure 6 , the network fault location device specifically includes the following contents:
[0094] The information collection module 10 is used to collect the data traffic forwarded by the edge access switch and perform packet analysis to determine the data traffic that may have congestion.
[0095] The data volume association module 20 is used to determine whether the data traffic that may have congestion is a microburst data stream. If not, it performs correlation matching and life cycle identification on the round-trip characteristic packets of the data traffic.
[0096] The delay fault analysis module 30 is used to perform delay analysis on the data traffic after the correlation matching and life cycle identification are completed to determine the network fault location.
[0097] As can be seen from the above description, the network fault location device provided by the embodiment of the present application can collect the data traffic forwarded by the edge access switch and perform packet analysis, accurately partition the microburst data stream, greatly reduce the total amount of collected data, avoid encountering data collection bottlenecks, and accurately distinguish and delimit the abnormal reasons according to the decomposition of the total delay of the data traffic.
[0098] In order to be able to determine the data traffic that may have congestion, in an embodiment of the network fault location device of the present application, refer to Figure 7 , the information collection module 10 includes:
[0099] The packet header information collection unit 11 is used to monitor the data traffic forwarded by the storage service edge access switch in real time and collect the packet header information of the data traffic.
[0100] The packet header feature analysis unit 12 is used to perform feature analysis on the packet header information of the data traffic to determine the data traffic that may have congestion.
[0101] In order to be able to screen the microburst data stream, in an embodiment of the network fault location device of the present application, refer to Figure 8 , the data volume association module 20 includes:
[0102] The microburst judgment unit 21 is used to judge whether the data traffic that may have congestion is a microburst data stream according to the depth information of the chip cache queue read within a set time period.
[0103] In order to be able to accurately perform correlation matching and life cycle identification, in an embodiment of the network fault location device of the present application, refer to Figure 9 , the data volume association module 20 includes:
[0104] The correlation matching unit 22 is used to perform correlation matching on the round-trip characteristic packets of the data traffic on the edge access switch processors accessed on the client side and the storage array side respectively.
[0105] A lifecycle identification unit 23 is used to identify the end - side data access delay, data preparation delay, data transmission delay, and data confirmation delay in the data traffic lifecycle.
[0106] In order to accurately perform delay analysis, in an embodiment of the network fault location device of the present application, refer to Figure 10 , the delay fault analysis module 30 includes:
[0107] A large - small flow discrimination unit 31 is used to distinguish the large - small flow types of the data traffic according to the correlation matching result of the round - trip characteristic packets.
[0108] A small - flow uploading unit 32 is used to, if the data traffic belongs to the small - flow type, upload it to a set acquisition server through data mirroring for delay analysis.
[0109] A large - flow uploading unit 33 is used to, if the data traffic belongs to the large - flow type, first perform data pre - processing and then upload it to a set acquisition server for delay analysis.
[0110] In order to accurately determine the network fault location, in an embodiment of the network fault location device of the present application, refer to Figure 11 , the delay fault analysis module 30 includes:
[0111] A network - side fault location unit 34 is used to, if the delay ratio of the data transmission delay exceeds the threshold, determine that a fault occurs on the network side, and send the network switch bandwidth utilization rate and packet loss rate to the network operation and maintenance end for emergency processing.
[0112] An end - side fault location unit 35 is used to, if the delay ratio of the data preparation delay and / or data confirmation delay exceeds the threshold, determine that a fault occurs on the end side, and notify the system operation and maintenance end to check the client and storage array.
[0113] To further illustrate the present solution, the present application also provides a specific application example of a network fault location method implemented by using the above - mentioned network fault location device. Refer to Figure 12 , which specifically includes the following content:
[0114] S1: The monitoring and analysis platform monitors the IO traffic forwarded by the storage service edge access switch in real - time and obtains the device log to observe the IO traffic index information.
[0115] S2: The monitoring and analysis platform cooperates with the switch to only collect the information of the IO traffic packet header part and upload it to the index, and analyze the characteristic packet information of the packet header, greatly reducing the total amount of analysis and acquisition data.
[0116] S3: After initial screening through business modeling analysis to identify IOs that may be congested, read the chip cache queue depth information within the cycle to analyze whether the IO is a microburst IO traffic. Microburst is a normal phenomenon; if it is not a microburst, further judgment is required.
[0117] S3a: Perform correlation matching on the round-trip characteristic packets of the IO on the edge access switch processors connected to the client side and the storage array side respectively, and identify each life cycle of the IO (end-side data access delay, data preparation delay, data transmission delay, data confirmation delay).
[0118] S3b: Distinguish large and small flows according to the IO association situation. Small IOs are sent to the acquisition server for analysis and processing using mirroring or erpsan methods; large IOs are pre-processed through feature extraction, aggregation, etc. in the device class and then sent to the acquisition server for analysis and processing.
[0119] S3c: Complete the decomposition of the total IO delay on the analysis server.
[0120] S4: Among them, the data transmission delay is related to bandwidth, network delay, etc., and the data preparation delay and data confirmation delay are related to the client and the storage array. Determine whether the delay anomaly occurs on the network side or the end side.
[0121] S5: According to the result of S4, if the data transmission delay is abnormal, that is, the network side is abnormal, the network operation and maintenance personnel check indicators such as the bandwidth utilization rate and packet loss rate of the network switch and perform emergency handling.
[0122] S6: According to the result of S4, if the data preparation delay and data confirmation delay are abnormal, the system operation and maintenance personnel check the client and the storage array and perform targeted processing.
[0123] S7: Complete the delimitation of the problems causing storage service anomalies.
[0124] As can be seen from the above content, the present application can at least further achieve the following technical effects:
[0125] 1. Build a business network model to accurately distinguish data flow types of microbursts and large and small IOs.
[0126] 2. Greatly reduce the total amount of collected data by analyzing the collection of IO packet header information, and avoid encountering data collection bottlenecks.
[0127] 3. Distinguish and delimit the causes of storage service anomalies according to the decomposition of the total IO data flow delay.
[0128] From the hardware level, in order to accurately delimit the cause of network faults, the present application provides an embodiment of an electronic device for implementing all or part of the content in the network fault location method. The electronic device specifically includes the following content:
[0129] A processor, a memory, a communications interface, and a bus; wherein, the processor, the memory, and the communications interface complete communication with each other through the bus; the communications interface is used to implement information transmission between the network fault location device and related devices such as a core business system, a user terminal, and a related database, etc.; this logic controller can be a desktop computer, a tablet computer, a mobile terminal, etc., and this embodiment is not limited thereto. In this embodiment, this logic controller can be implemented with reference to the embodiments of the network fault location method in the embodiments, as well as the embodiments of the network fault location device, the content of which is incorporated herein, and the repeated parts will not be elaborated again.
[0130] It can be understood that the user terminal can include a smart phone, a tablet electronic device, a network set-top box, a portable computer, a desktop computer, a personal digital assistant (PDA), a vehicle-mounted device, a smart wearable device, etc. Among them, the smart wearable device can include smart glasses, a smart watch, a smart bracelet, etc.
[0131] In practical applications, part of the network fault location method can be executed on the side of the electronic device as described above, or all operations can be completed in the client device. Specifically, it can be selected according to the processing capacity of the client device, as well as the limitations of the user usage scenario, etc. This application does not make any limitations in this regard. If all operations are completed in the client device, the client device may further include a processor.
[0132] The above-mentioned client device can have a communication module (i.e., a communication unit), and can be communicatively connected to a remote server to achieve data transmission with the server. The server can include a server on the side of the task scheduling center, and in other implementation scenarios, it can also include a server of an intermediate platform, such as a server of a third-party server platform having a communication link with the task scheduling center server. The server can include a single computer device, or can include a server cluster composed of multiple servers, or a server structure of a distributed device.
[0133] Figure 13 This is a schematic block diagram of the system composition of the electronic device 9600 according to the embodiment of the present application. As Figure 13 shown, the electronic device 9600 can include a central processing unit 9100 and a memory 9140; the memory 9140 is coupled to the central processing unit 9100. It should be noted that this Figure 13 is exemplary; other types of structures can also be used to supplement or replace this structure to implement telecommunications functions or other functions.
[0134] In one embodiment, the function of the network fault location method can be integrated into the central processing unit 9100. Among them, the central processing unit 9100 can be configured to perform the following controls:
[0135] Step S101: Collect the data traffic forwarded by the edge access switch and perform packet analysis to determine the data traffic with possible congestion.
[0136] Step S102: Determine whether the data traffic with possible congestion is a microburst data stream. If not, perform correlation matching and life cycle identification on the round-trip characteristic packets of the data traffic.
[0137] Step S103: Perform delay analysis on the data traffic after the correlation matching and life cycle identification are completed to determine the network fault location.
[0138] As can be seen from the above description, the electronic device provided by the embodiment of the present application accurately partitions the microburst data stream by collecting the data traffic forwarded by the edge access switch and performing packet analysis, greatly reducing the total amount of collected data, avoiding the data collection bottleneck, and accurately distinguishing and delimiting the abnormal reasons according to the decomposition of the total delay of the data traffic.
[0139] In another embodiment, the network fault location device can be separately configured from the central processing unit 9100. For example, the network fault location device can be configured as a chip connected to the central processing unit 9100, and the function of the network fault location method can be realized through the control of the central processing unit.
[0140] As Figure 13 shown, the electronic device 9600 may further include: a communication module 9110, an input unit 9120, an audio processor 9130, a display 9160, and a power supply 9170. It should be noted that the electronic device 9600 does not necessarily have to include all the components shown in Figure 13 ; in addition, the electronic device 9600 may further include components not shown in Figure 13 , and reference can be made to the prior art.
[0141] As Figure 13 shown, the central processing unit 9100 is sometimes also referred to as a controller or an operation control, and may include a microprocessor or other processor devices and / or logic devices. The central processing unit 9100 receives inputs and controls the operations of the various components of the electronic device 9600.
[0142] Among them, the memory 9140 can be, for example, one or more of a buffer, a flash memory, a hard drive, a removable medium, a volatile memory, a non-volatile memory, or other suitable devices. The above-mentioned information related to failures can be stored, and in addition, a program for executing relevant information can also be stored. And the central processing unit 9100 can execute the program stored in the memory 9140 to implement information storage or processing, etc.
[0143] The input unit 9120 provides an input to the central processing unit 9100. The input unit 9120 is, for example, a key or a touch input device. The power supply 9170 is used to supply power to the electronic device 9600. The display 9160 is used to display display objects such as images and texts. The display can be, for example, an LCD display, but is not limited thereto.
[0144] The memory 9140 can be a solid-state memory. For example, it can be a read-only memory (ROM), a random access memory (RAM), a SIM card, etc. It can also be a memory that stores information even when power is off, can be selectively erased and has more data. An example of this memory is sometimes called an EPROM, etc. The memory 9140 can also be some other type of device. The memory 9140 includes a buffer memory 9141 (sometimes called a buffer). The memory 9140 can include an application / function storage unit 9142, and the application / function storage unit 9142 is used to store application programs and function programs or the processes for operating the electronic device 9600 through the central processing unit 9100.
[0145] The memory 9140 can also include a data storage unit 9143, and the data storage unit 9143 is used to store data, such as contacts, digital data, pictures, sounds, and / or any other data used by the electronic device. The driver storage unit 9144 of the memory 9140 can include various drivers of the electronic device for communication functions and / or for executing other functions of the electronic device (such as a messaging application, an address book application, etc.).
[0146] The communication module 9110 is a transmitter / receiver 9110 that transmits and receives signals via the antenna 9111. The communication module (transmitter / receiver) 9110 is coupled to the central processing unit 9100 to provide an input signal and receive an output signal, which can be the same as the case of a conventional mobile communication terminal.
[0147] Based on different communication technologies, in the same electronic device, multiple communication modules 9110 can be provided, such as a cellular network module, a Bluetooth module, and / or a wireless local area network module, etc. The communication module (transmitter / receiver) 9110 is also coupled to a speaker 9131 and a microphone 9132 via an audio processor 9130 to provide an audio output via the speaker 9131 and receive an audio input from the microphone 9132, so as to implement normal telecommunication functions. The audio processor 9130 can include any suitable buffer, decoder, amplifier, etc. In addition, the audio processor 9130 is also coupled to a central processor 9100, so that recording can be performed on the local machine through the microphone 9132, and the sound stored on the local machine can be played through the speaker 9131.
[0148] An embodiment of the present application also provides a computer-readable storage medium capable of implementing all steps in the network fault location method where the execution subject in the above embodiment is a server or a client. A computer program is stored on the computer-readable storage medium, and when the computer program is executed by a processor, all steps in the network fault location method where the execution subject in the above embodiment is a server or a client are implemented. For example, when the processor executes the computer program, the following steps are implemented:
[0149] Step S101: Collect the data traffic forwarded by the edge access switch and perform packet analysis to determine the data traffic that may be congested.
[0150] Step S102: Determine whether the data traffic that may be congested is a microburst data stream. If not, perform correlation matching and life cycle identification on the round-trip characteristic packets of the data traffic.
[0151] Step S103: After the correlation matching and life cycle identification are completed, perform delay analysis on the data traffic to determine the network fault location.
[0152] As can be seen from the above description, the computer-readable storage medium provided by the embodiment of the present application accurately partitions the microburst data stream by collecting the data traffic forwarded by the edge access switch and performing packet analysis, greatly reducing the total amount of collected data, avoiding encountering a data collection bottleneck, and accurately distinguishing and delimiting the abnormal cause according to the decomposition of the total delay of the data traffic.
[0153] An embodiment of the present application also provides a computer program product capable of implementing all steps in the network fault location method where the execution subject in the above embodiment is a server or a client. When the computer program / instructions are executed by a processor, the steps of the network fault location method are implemented. For example, the computer program / instructions implement the following steps:
[0154] Step S101: Collect the data traffic forwarded by the edge access switch and perform packet analysis to determine the data traffic that may be congested.
[0155] Step S102: Determine whether the data traffic that may be congested is a microburst data stream. If not, perform correlation matching and life cycle identification on the round-trip characteristic packets of the data traffic.
[0156] Step S103: Perform delay analysis on the data traffic after the correlation matching and life cycle identification are completed to determine the network fault location.
[0157] As can be seen from the above description, the computer program product provided by the embodiments of the present application collects the data traffic forwarded by the edge access switch and performs packet analysis, accurately partitions the microburst data stream, greatly reduces the total amount of collected data, avoids encountering data collection bottlenecks, and accurately distinguishes and delimits the abnormal causes according to the decomposition of the total delay of the data traffic.
[0158] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, apparatus, or computer program product. Therefore, the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0159] The present invention is described with reference to the flowcharts and / or block diagrams of methods, devices (apparatus), and computer program products according to embodiments of the present invention. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and the combination of flows and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the functions specified in Figure 1 one or more flows and / or blocks Figure 1 one or more blocks.
[0160] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device that implements the functions specified in Figure 1 one or more flows and / or blocks Figure 1 one or more blocks.
[0161] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus, so that a series of operation steps are executed on the computer or other programmable apparatus to produce a computer-implemented process, thereby providing instructions for implementing the steps of the function specified in one process or a plurality of processes and / or blocks Figure 1 in one block or a plurality of blocks Figure 1 in the method.
[0162] Specific embodiments are used in the present invention to illustrate the principles and implementation manners of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention; at the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation to the present invention.
Claims
1. A network fault location method, characterized in that, The method includes: Collecting the data traffic forwarded by the edge access switch and performing packet analysis to determine the data traffic that may be congested; Judging whether the data traffic that may be congested is a microburst data stream. If not, performing correlation matching and life cycle identification on the round-trip characteristic packets of the data traffic, including: performing correlation matching on the round-trip characteristic packets of the data traffic on the edge access switch processors accessed on the client side and the storage array side respectively; identifying the end-side data access delay, data preparation delay, data transmission delay, and data confirmation delay in the life cycle of the data traffic; After the correlation matching and life cycle identification are completed, performing delay analysis on the data traffic to determine the network fault location, including: if the delay ratio of the data transmission delay exceeds the threshold, determining that a fault occurs on the network side and sending the network switch bandwidth utilization rate and packet loss rate to the network operation and maintenance end for emergency handling; if the delay ratio of the data preparation delay and / or data confirmation delay exceeds the threshold, determining that a fault occurs on the end side and notifying the system operation and maintenance end to check the client and the storage array.
2. The network fault location method according to claim 1, characterized in that, The collecting the data traffic forwarded by the edge access switch and performing packet analysis to determine the data traffic that may be congested includes: Real-time monitoring the data traffic forwarded by the storage service edge access switch and collecting the packet header information of the data traffic; Performing feature analysis on the packet header information of the data traffic to determine the data traffic that may be congested.
3. The network fault location method according to claim 1, characterized in that, The judging whether the data traffic that may be congested is a microburst data stream includes: Judging the microburst data stream of the data traffic that may be congested according to the chip cache queue depth information read within a set time period.
4. The network fault location method according to claim 1, characterized in that, The performing delay analysis on the data traffic after the correlation matching and life cycle identification are completed includes: Distinguishing the large and small flow types of the data traffic according to the correlation matching result of the round-trip characteristic packets; If the data traffic belongs to the small flow type, uploading it to the set collection server for delay analysis through the data mirroring method; If the data traffic belongs to the large flow type, first performing data preprocessing and then uploading it to the set collection server for delay analysis.
5. A network fault location device, characterized in that, Including: An information collection module, which is used to collect the data traffic forwarded by the edge access switch and perform packet analysis to determine the data traffic that may be congested; A data volume correlation module, which is used to judge whether the data traffic that may be congested is a microburst data stream. If not, performing correlation matching and life cycle identification on the round-trip characteristic packets of the data traffic, including: performing correlation matching on the round-trip characteristic packets of the data traffic on the edge access switch processors accessed on the client side and the storage array side respectively; identifying the end-side data access delay, data preparation delay, data transmission delay, and data confirmation delay in the life cycle of the data traffic; The delay fault analysis module is used to perform delay analysis on the data traffic after the correlation matching and life cycle identification are completed, and determine the network fault location, including: if the delay ratio of the data transmission delay exceeds the threshold, it is determined that a fault occurs on the network side, and the network switch bandwidth utilization rate and packet loss rate are sent to the network operation and maintenance end for emergency processing; if the delay ratio of the data preparation delay and / or data confirmation delay exceeds the threshold, it is determined that a fault occurs on the terminal side, and the system operation and maintenance end is notified to check the client and storage array.
6. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the steps of the network fault location method according to any one of claims 1 to 4.
7. A computer-readable storage medium, on which a computer program is stored, characterized in that, When the computer program is executed by the processor, it implements the steps of the network fault location method according to any one of claims 1 to 4.
8. A computer program product, comprising a computer program / instructions, characterized in that, When the computer program / instructions are executed by the processor, it implements the steps of the network fault location method according to any one of claims 1 to 4.
Citation Information
Patent Citations
Network fault early warning method and device
CN110601900A
Knowledge base radio and core network prescriptive root cause analysis
US20160241429A1