Network fault automatic processing method, device and equipment and storage medium

CN117097601BActive Publication Date: 2026-08-07CCB FINTECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
CCB FINTECH CO LTD
Filing Date
2023-09-05
Publication Date
2026-08-07

AI Technical Summary

Technical Problem

网络作为基础服务,可以造成网络故障的原因多种多样,导致的网络故障问题、现象可能各不相同,根据人工经验排查网络故障,难度较大

Benefits of technology

[0043] According to embodiments of this disclosure, by identifying abnormal servers and corresponding first keywords based on monitoring thresholds and detected fault information, M data packets from the abnormal servers are automatically captured within a preset time period. Based on these M data packets, second keywords are determined, thus abstracting network fault phenomena into standard error types using first and second keywords. Then, based on the first and second keywords and an artificial intelligence library, the cause of the target network fault is obtained. After inputting the first and second keywords into the artificial intelligence library, the target network fault cause is automatically analyzed and located using artificial intelligence algorithms, improving troubleshooting efficiency. Then, the configuration information corresponding to the abnormal server is adjusted according to the cause of the target network fault, restoring the network status of the abnormal server to normal. This achieves automatic repair of the network status of abnormal servers without any manual intervention, reducing manual maintenance costs and improving the efficiency of automatic network fault handling.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117097601B_ABST
    Figure CN117097601B_ABST
Patent Text Reader

Abstract

The present disclosure provides a network fault automatic processing method and device, equipment and storage medium, which can be applied to the fields of computer technology, intelligent operation and maintenance technology, financial technology and big data technology. The method comprises the following steps: determining an abnormal server and a first keyword corresponding to the abnormal server according to a monitoring threshold and monitored fault information; automatically capturing M data packets fed back by the abnormal server in a preset period; determining a second keyword according to the M data packets; obtaining a target network fault cause according to the first keyword, the second keyword and an artificial intelligence library; and adjusting configuration information corresponding to the abnormal server according to the target network fault cause, so as to restore the network state of the abnormal server to normal.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the fields of computer technology, intelligent operation and maintenance technology, financial technology technology and big data technology, and more specifically to a method, apparatus, device, medium and program product for automatic network fault handling. Background Technology

[0002] In the process of troubleshooting network faults, fault location mainly relies on manual analysis. As a fundamental service, networks are susceptible to a variety of causes, resulting in diverse problems and symptoms. Troubleshooting network faults based solely on human experience is extremely difficult.

[0003] In realizing the concept disclosed herein, the inventors discovered at least the following problems in the related technologies: manual troubleshooting of network faults is slow and has low troubleshooting efficiency. Summary of the Invention

[0004] In view of the above problems, this disclosure provides a method, apparatus, device, medium and program product for automatic network fault handling.

[0005] The first aspect of this disclosure provides an automatic network fault handling method, comprising:

[0006] Based on the monitoring thresholds and the detected fault information, identify the abnormal servers and the first keyword corresponding to the aforementioned abnormal servers;

[0007] During a preset time period, automatically capture M data packets returned by the aforementioned abnormal server, where M is an integer greater than or equal to 1;

[0008] Based on the above M data packets, determine the second keyword;

[0009] Based on the first keyword and the second keyword mentioned above, as well as the artificial intelligence library, the cause of the target network failure is obtained.

[0010] Based on the cause of the aforementioned target network failure, the configuration information corresponding to the abnormal server is adjusted to restore the network status of the abnormal server to normal.

[0011] According to embodiments of this disclosure, determining the second keyword based on the aforementioned M data packets includes:

[0012] Based on the above M data packets, N network connection status identifiers and the number of identifiers corresponding to each of the N network connection status identifiers are obtained, where N is an integer greater than or equal to 1 and less than or equal to M;

[0013] Based on the above N network connection status identifiers and the above number of identifiers, the above second keyword is determined.

[0014] According to embodiments of this disclosure, the aforementioned fault information includes multiple indicator information, and the determination of the abnormal server and the first keyword corresponding to the abnormal server based on the monitoring threshold and the monitored fault information includes:

[0015] For each of the multiple metrics, if the value of the metric exceeds the monitoring threshold corresponding to the metric, the monitored server is identified as the abnormal server.

[0016] Based on the above indicator information, determine the format of the first keyword corresponding to the indicator from the first keyword library;

[0017] Based on the above indicator information and the format of the first keyword corresponding to the indicator, the above first keyword is obtained.

[0018] According to embodiments of this disclosure, the aforementioned fault information includes fault type information reported by the client corresponding to the aforementioned abnormal server. The determination of the abnormal server and the first keyword corresponding to the aforementioned abnormal server based on the monitoring threshold and the detected fault information further includes:

[0019] Based on the above fault type information, determine the format of the first keyword corresponding to the above fault type information from the first keyword library;

[0020] Based on the above fault type information and the format of the first keyword corresponding to the above fault type information, the above first keyword is obtained.

[0021] According to embodiments of this disclosure, based on the aforementioned first keyword, the aforementioned second keyword, and the artificial intelligence library, the causes of target network failures include:

[0022] Based on the first keyword and the second keyword mentioned above, the keyword group is obtained;

[0023] The above keyword groups are matched with the index terms of the above artificial intelligence library to obtain the target index terms;

[0024] Based on the target index terms and the artificial intelligence library mentioned above, the initial network failure cause is obtained;

[0025] Based on the initial network failure cause mentioned above, the configuration information corresponding to the abnormal server is checked to determine the cause of the target network failure.

[0026] According to embodiments of this disclosure, the above-mentioned detection of configuration information corresponding to the abnormal server based on the above-mentioned initial network failure cause, and determination of the above-mentioned target network failure cause, includes:

[0027] Based on the initial network failure cause described above, determine the abnormal configuration corresponding to the abnormal server.

[0028] Retrieve the configuration information corresponding to the above-mentioned abnormal configuration;

[0029] If the initial network failure cause matches the configuration information, the initial network failure cause will be identified as the target network failure cause.

[0030] According to embodiments of this disclosure, the first keyword mentioned above includes at least one of the following:

[0031] Slow access, error code return, network interruption, inability to access, incomplete data return, host retransmission rate reaching the first detection value, domain name dialing success rate lower than the second detection value, service latency exceeding the third detection value, and bandwidth continuously surging to the fourth detection value.

[0032] According to embodiments of this disclosure, the second keyword mentioned above includes at least one of the following:

[0033] Retransmission, port reuse, reset, network window, timestamp, increased host retransmission rate, multiple retransmission packets received, multiple connection reset packets sent, timestamp value field is 0.

[0034] A second aspect of this disclosure provides an automatic network fault handling device, comprising:

[0035] The first determination module is used to determine the abnormal server and the first keyword corresponding to the abnormal server based on the monitoring threshold and the detected fault information.

[0036] The capture module is used to automatically capture M data packets returned by the abnormal server within a preset time period, where M is an integer greater than or equal to 1;

[0037] The second determining module is used to determine the second keyword based on the above M data packets;

[0038] The module is used to obtain the cause of the target network failure based on the first keyword, the second keyword and the artificial intelligence library mentioned above;

[0039] The adjustment module is used to adjust the configuration information corresponding to the abnormal server according to the cause of the above-mentioned target network failure, so as to restore the network status of the abnormal server to normal.

[0040] A third aspect of this disclosure provides an electronic device comprising: one or more processors; and a storage device for storing one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors perform the network fault automatic handling method.

[0041] A fourth aspect of this disclosure also provides a computer-readable storage medium having executable instructions stored thereon, which, when executed by a processor, cause the processor to perform the above-described automatic network fault handling method.

[0042] The fifth aspect of this disclosure also provides a computer program product, including a computer program that, when executed by a processor, implements the above-described automatic network fault handling method.

[0043] According to embodiments of this disclosure, by identifying abnormal servers and corresponding first keywords based on monitoring thresholds and detected fault information, M data packets from the abnormal servers are automatically captured within a preset time period. Based on these M data packets, second keywords are determined, thus abstracting network fault phenomena into standard error types using first and second keywords. Then, based on the first and second keywords and an artificial intelligence library, the cause of the target network fault is obtained. After inputting the first and second keywords into the artificial intelligence library, the target network fault cause is automatically analyzed and located using artificial intelligence algorithms, improving troubleshooting efficiency. Then, the configuration information corresponding to the abnormal server is adjusted according to the cause of the target network fault, restoring the network status of the abnormal server to normal. This achieves automatic repair of the network status of abnormal servers without any manual intervention, reducing manual maintenance costs and improving the efficiency of automatic network fault handling. Attached Figure Description

[0044] The foregoing contents, as well as other objects, features, and advantages of this disclosure, will become clearer from the following description of embodiments with reference to the accompanying drawings, in which:

[0045] Figure 1 This diagram illustrates an application scenario of the automatic network fault handling method according to an embodiment of the present disclosure.

[0046] Figure 2 A flowchart illustrating an automatic network fault handling method according to an embodiment of the present disclosure is shown schematically.

[0047] Figure 3 This illustration schematically depicts communication between servers based on TCP / IP according to embodiments of the present disclosure;

[0048] Figure 4 A schematic diagram illustrating fault analysis according to an embodiment of the present disclosure is shown.

[0049] Figure 5 A flowchart illustrating automatic network fault handling according to another embodiment of the present disclosure is shown schematically;

[0050] Figure 6A schematic block diagram of a network fault automatic handling apparatus according to an embodiment of the present disclosure is shown; and

[0051] Figure 7 A block diagram schematically illustrates an electronic device suitable for implementing an automatic network fault handling method according to an embodiment of the present disclosure. Detailed Implementation

[0052] The embodiments of the present disclosure will now be described with reference to the accompanying drawings. However, it should be understood that these descriptions are exemplary only and are not intended to limit the scope of the disclosure. In the following detailed description, numerous specific details are set forth to provide a thorough understanding of the embodiments of the present disclosure for ease of explanation. However, it will be apparent that one or more embodiments may be practiced without these specific details. Furthermore, descriptions of well-known structures and techniques are omitted in the following description to avoid unnecessarily obscuring the concepts of the present disclosure.

[0053] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit this disclosure. The terms “comprising,” “including,” etc., as used herein indicate the presence of the stated features, steps, operations, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, or components.

[0054] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those skilled in the art, unless otherwise defined. It should be noted that the terms used herein are to be interpreted in a manner consistent with the context of this specification, and not in an idealized or overly rigid way.

[0055] When using expressions such as "at least one of A, B, and C", they should generally be interpreted in accordance with the meaning that is commonly understood by a person skilled in the art (e.g., "a system having at least one of A, B, and C" should include, but is not limited to, a system having A alone, a system having B alone, a system having C alone, a system having A and B, a system having A and C, a system having B and C, and / or a system having A, B, and C, etc.).

[0056] In the technical solution of this invention, the user information (including but not limited to user personal information, user image information, user device information, such as location information) and data (including but not limited to data used for analysis, stored data, and displayed data) involved are all information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, storage, use, processing, transmission, provision, disclosure, and application of related data all comply with the relevant laws, regulations, and standards of the relevant countries and regions, take necessary confidentiality measures, do not violate public order and good morals, and provide corresponding operation entry points for users to choose to authorize or refuse.

[0057] Current network troubleshooting relies heavily on individual experience, leading to a lack of standardized skills among operations and maintenance personnel and significant individual differences in troubleshooting abilities. As a fundamental service, networks experience a wide range of causes and manifestations of faults. Faults originating from applications, systems, and other sources can all present as network faults, sometimes exhibiting identical symptoms. The broad scope and numerous causes of network faults make troubleshooting difficult, lacking clear direction and hindering analysis and resolution. For operations and maintenance personnel with limited troubleshooting experience, analyzing network faults presents considerable challenges.

[0058] Currently, network fault emergency response requires manual intervention to confirm relevant information before any action can be taken. Related technologies involve identifying the abnormal server host through alarms when a network fault occurs. Operations personnel then log into the abnormal server host to query relevant application logs, filter abnormal logs, and confirm the fault location before taking action. This requires operations personnel to be familiar with the application and the overall system architecture, placing high demands on their skills. Furthermore, network fault emergency response is slow, and troubleshooting efficiency is low.

[0059] As internet applications develop, applications and systems are constantly iterating, leading to more and more problems. Individuals have limited learning abilities and high learning costs, making it even more difficult to rely on human experience for network troubleshooting.

[0060] In order to at least partially solve the technical problems existing in the related technologies, the embodiments of this disclosure provide an automatic network fault handling method, apparatus, device and storage medium, which can be applied to the fields of computer technology, intelligent operation and maintenance technology, financial technology technology and big data technology.

[0061] Embodiments of this disclosure provide an automatic network fault handling method, comprising: determining an abnormal server and a first keyword corresponding to the abnormal server based on a monitoring threshold and monitored fault information; automatically capturing M data packets fed back by the abnormal server within a preset time period, wherein M is an integer greater than or equal to 1; determining a second keyword based on the M data packets; obtaining the target network fault cause based on the first keyword, the second keyword, and an artificial intelligence library; and adjusting the configuration information corresponding to the abnormal server according to the target network fault cause to restore the network status of the abnormal server to normal.

[0062] Figure 1 The diagram illustrates an application scenario of the automatic network fault handling method according to an embodiment of the present disclosure.

[0063] like Figure 1As shown, application scenario 100 according to this embodiment may include a first terminal device 101, a second terminal device 102, a third terminal device 103, a network 104, and a server 105. The network 104 serves as a medium for providing a communication link between the first terminal device 101, the second terminal device 102, the third terminal device 103, and the server 105. The network 104 may include various connection types, such as wired or wireless communication links, or fiber optic cables, etc.

[0064] Users can use the first terminal device 101, the second terminal device 102, and the third terminal device 103 to interact with the server 105 via the network 104 to receive or send messages, etc. Various communication client applications can be installed on the first terminal device 101, the second terminal device 102, and the third terminal device 103, such as shopping applications, web browser applications, search applications, instant messaging tools, email clients, social media platform software, etc. (for example only).

[0065] The first terminal device 101, the second terminal device 102, and the third terminal device 103 can be various electronic devices with displays and support web browsing, including but not limited to smartphones, tablets, laptops, and desktop computers.

[0066] Server 105 can be a server that provides various services, such as a backend management server that supports websites browsed by users using the first terminal device 101, the second terminal device 102, and the third terminal device 103 (this is just an example). The backend management server can analyze and process data such as received user requests, and feed back the processing results (such as web pages, information, or data obtained or generated according to user requests) to the terminal devices.

[0067] It should be noted that the automatic network fault handling method provided in this embodiment can generally be executed by server 105. Correspondingly, the automatic network fault handling device provided in this embodiment can generally be located in server 105. The automatic network fault handling method provided in this embodiment can also be executed by a server or server cluster that is different from server 105 and capable of communicating with the first terminal device 101, the second terminal device 102, the third terminal device 103, and / or server 105. Correspondingly, the automatic network fault handling device provided in this embodiment can also be located in a server or server cluster that is different from server 105 and capable of communicating with the first terminal device 101, the second terminal device 102, the third terminal device 103, and / or server 105.

[0068] It should be understood that Figure 1The number of terminal devices, networks, and servers shown is merely illustrative. Depending on implementation needs, any number of terminal devices, networks, and servers can be included.

[0069] The following will be based on Figure 1 The described scene, through Figures 2-5 The automatic network fault handling method of the disclosed embodiments is described in detail.

[0070] Figure 2 A flowchart illustrating an automatic network fault handling method according to an embodiment of the present disclosure is shown schematically.

[0071] like Figure 2 As shown, the automatic network fault handling method of this embodiment includes operations S210 to S250.

[0072] In operation S210, based on monitoring thresholds and detected fault information, abnormal servers and the first keyword corresponding to the abnormal servers are identified.

[0073] According to embodiments of this disclosure, fault information characterizes network anomalies in the server.

[0074] According to embodiments of this disclosure, the first key information characterizes the standard error type of the network fault extracted from the fault information.

[0075] According to embodiments of this disclosure, before determining the abnormal server and the first keyword corresponding to the abnormal server based on the monitoring threshold and the detected fault information, the method further includes: obtaining fault information.

[0076] According to embodiments of this disclosure, a monitoring platform can be used to monitor the server implementing network services in real time and collect network anomaly information reported by the client corresponding to the server in real time, thereby obtaining fault information corresponding to the server.

[0077] According to embodiments of this disclosure, when servers implementing network services communicate with each other, the commonly used network communication protocol is TCP / IP (Transmission Control Protocol / Internet Protocol). TCP / IP refers to a suite of protocols capable of transmitting information between multiple different networks. The TCP / IP protocol does not simply refer to the two protocols TCP and IP, but rather to a protocol suite composed of protocols such as FTP, SMTP, TCP, UDP, and IP. Because TCP and IP are the most representative protocols within the TCP / IP protocol suite, it is called the TCP / IP protocol suite.

[0078] According to embodiments of this disclosure, servers implementing network services typically communicate using TCP / IP. Communication pairs are established for calls between different services, i.e., common TCP communication pairs are established. By using a monitoring platform to collect all communication pairs corresponding to a service host, all associated applications of the service host can be statistically identified. By monitoring the connection status of the communication pairs corresponding to the service host in real time and collecting network anomaly information reported by clients corresponding to the service host in real time, network fault information corresponding to the service host can be obtained. Here, the service host represents the server implementing the network service.

[0079] According to embodiments of this disclosure, for each server monitored by the monitoring platform, after obtaining the fault information corresponding to the server, each fault information included in the fault information can be compared with the monitoring threshold corresponding to each fault information to determine the abnormal server.

[0080] When operating S220, within a preset time period, M data packets are automatically captured from the abnormal server, where M is an integer greater than or equal to 1.

[0081] According to embodiments of this disclosure, when automatically capturing M data packets returned by an abnormal server, capture can be performed using default capture parameters, or capture parameters can be defined as needed. For example, capture size, capture address, save path, etc., can be defined as required.

[0082] According to embodiments of this disclosure, the preset time period can be selected based on actual business conditions, and is not limited thereto.

[0083] For example, the preset time period can be 5 minutes, 10 minutes, or 20 minutes, etc.

[0084] According to embodiments of this disclosure, M data packets fed back by the abnormal server are automatically captured within a preset time period, which facilitates the quantitative statistical analysis of the fault information included in the data packets based on the number of data packets and the size of the preset time period.

[0085] In operation S230, the second keyword is determined based on M data packets.

[0086] According to embodiments of this disclosure, the second key information characterizes the standard error type that reflects network failures extracted from data packets.

[0087] According to embodiments of this disclosure, an abnormal server communicating based on the TCP / IP protocol sends data packets containing a TCP header. This TCP header includes flags, which are network connection status identifiers; different flags indicate different network connection states. By analyzing the network status identifiers included in the data, the second keyword can be determined.

[0088] According to the embodiments of this disclosure, since the data packets are obtained by real-time monitoring of the network connection status of the abnormal server when the network of the abnormal server is detected to be abnormal, the second keyword can be determined based on M data packets, and the second keyword can accurately reflect the network connection status and network fault.

[0089] In operation S240, the cause of the target network failure is obtained based on the first keyword, the second keyword, and the artificial intelligence library.

[0090] According to embodiments of this disclosure, the artificial intelligence library stores solutions for resolving network failures and the causes of network failures.

[0091] According to embodiments of this disclosure, by inputting a first keyword and a second keyword into an artificial intelligence library, the first keyword and the second keyword can be used as partial matching conditions based on the artificial intelligence algorithms included in the artificial intelligence library, and the target network fault cause corresponding to the first keyword and the second keyword can be automatically obtained from the artificial intelligence library.

[0092] According to embodiments of this disclosure, the artificial intelligence library can continuously learn from network fault handling cases. After inputting the first keyword and the second keyword into the artificial intelligence library, the library can automatically obtain a more accurate cause of the target network fault through artificial intelligence algorithm matching, thereby improving troubleshooting efficiency.

[0093] When operating the S250, the configuration information corresponding to the abnormal server is adjusted according to the cause of the target network failure, so that the network status of the abnormal server is restored to normal.

[0094] For example, a target network failure could be caused by proxies like HAProxy performing TCP health checks on application services, leading to packet loss and retransmission. Adjusting the configuration information corresponding to the abnormal server based on the cause of the target network failure could involve changing the TCP health check to an HTTP health check.

[0095] According to embodiments of this disclosure, by identifying abnormal servers and corresponding first keywords based on monitoring thresholds and detected fault information, M data packets from the abnormal servers are automatically captured within a preset time period. Based on these M data packets, second keywords are determined, thus abstracting network fault phenomena into standard error types using first and second keywords. Then, based on the first and second keywords and an artificial intelligence library, the cause of the target network fault is obtained. After inputting the first and second keywords into the artificial intelligence library, the target network fault cause is automatically analyzed and located using artificial intelligence algorithms, improving troubleshooting efficiency. Then, the configuration information corresponding to the abnormal server is adjusted according to the cause of the target network fault, restoring the network status of the abnormal server to normal. This achieves automatic repair of the network status of abnormal servers without any manual intervention, reducing manual maintenance costs and improving the efficiency of automatic network fault handling.

[0096] According to the embodiments of this disclosure, the automatic network fault handling method provided by this disclosure can obtain a more accurate target network fault cause by combining the first keyword and the second keyword with an artificial intelligence library. Then, based on the more accurate target network fault cause, the configuration information corresponding to the abnormal server is automatically adjusted to restore the network status of the abnormal server to normal. This completes one-stop automatic operation and maintenance from fault discovery, fault analysis to fault handling, which greatly improves operation and maintenance efficiency and troubleshooting efficiency. Moreover, production emergencies have high requirements for timeliness. Timely discovery and resolution of production fault points (hosts, applications, etc.) can greatly ensure the availability of the production system.

[0097] According to embodiments of this disclosure, for example, Figure 2 Operation S230, as shown, determines the second keyword based on M data packets, and may include the following operations:

[0098] Based on M data packets, obtain N network connection status identifiers and the number of identifiers corresponding to each of the N network connection status identifiers, where N is an integer greater than or equal to 1 and less than or equal to M;

[0099] The second keyword is determined based on N network connection status identifiers and the number of identifiers.

[0100] According to embodiments of this disclosure, network failures will display relevant network failure characteristics on M data packets, such as slow network access, inability to access data, connection timeout, incomplete data display, intermittent disconnection, etc. Relevant identifiers and parameters will be displayed on the data packets, such as retransmission, reset, port reuse, Windows window congestion, timestamp of 0, etc.

[0101] According to embodiments of this disclosure, an abnormal server communicating based on the TCP / IP protocol sends data packets containing a TCP header. This TCP header includes flags, which are network connection status identifiers; different flags indicate different network connection states. By analyzing the network status identifiers included in the data, the second keyword can be determined.

[0102] Figure 3 This illustration schematically depicts communication between servers based on TCP / IP according to embodiments of the present disclosure.

[0103] like Figure 3 As shown in (a), when communication is initiated, client 301 sends a data packet with the SYN flag to server 302, that is, sends a request to server 302 to establish a connection.

[0104] like Figure 3 As shown in (b), the client 303 sends a data packet with the SYN flag to the server 304. After the server 304 receives the connection establishment request sent by the client 303, in response to the connection establishment request, the server 304 sends back a data packet with the RST flag, that is, a data packet that needs to be reset, indicating that the server 304 service has not started.

[0105] According to embodiments of this disclosure, it is possible to... Figure 3 (b) shows the communication process, which automatically captures M data packets fed back by the abnormal server. At this time, the client 303 is equivalent to the end where the server of this embodiment is located, and the server 304 is equivalent to the end where the abnormal server is located.

[0106] According to embodiments of this disclosure, the N network status identifiers may include at least one of the following: connection establishment identifier, response identifier, connection closure identifier, response identifier after connection establishment, connection orderly termination identifier, connection reset identifier, DATA data transmission identifier, urgent pointer valid identifier, retransmission identifier, port reuse identifier, and Windows window congestion identifier.

[0107] According to embodiments of this disclosure, the second keyword includes at least one of the following:

[0108] Retransmission, port reuse, reset, network window, timestamp, increased host retransmission rate, multiple retransmission packets received, multiple connection reset packets sent, timestamp value field is 0.

[0109] According to embodiments of this disclosure, SYN can represent establishing a connection, FIN can represent closing a connection, ACK can represent a response, PSH can represent data transmission, RST can represent a connection reset, and TSval=0 can represent a timestamp value field of 0.

[0110] According to embodiments of this disclosure, ACK can be used simultaneously with SYN, FIN, etc. For example, SYN / ACK used simultaneously, and both being 1, indicates a response after a connection is established; if it is only a single SYN, it simply indicates that a connection has been established. FIN / ACK used simultaneously indicates that the connection is terminated in an orderly manner.

[0111] According to embodiments of this disclosure, a second keyword corresponding to N network connection status identifiers can be determined by judging whether the number of identifiers is within the identifier number threshold range.

[0112] For example, N network connection status identifiers can include 2 RST packets. The threshold for the number of identifiers corresponding to the RST identifier is 1. If 2 > 1, it is confirmed that the second key corresponding to the RST identifier is multiple resets.

[0113] For example, N network connection status identifiers can include 1 RST packet. The threshold for the number of identifiers corresponding to the RST identifier is 1. When 1 = 1, the second key corresponding to the RST identifier is confirmed to be reset.

[0114] According to embodiments of this disclosure, if only RST packets are received multiple times after a connection is established and a SYN packet is sent, it may be that the server service is not running. In this case, the artificial intelligence library can recommend checking whether the server service is running based on the second keyword corresponding to this situation, such as multiple resets.

[0115] According to embodiments of this disclosure, when the network status connection identifier is only SYN, it indicates that the network connection has been established but no ACK packet has been received, and several retransmissions are required, indicating that the network is not working. In this case, the artificial intelligence library can recommend checking security group access, routing, firewall settings, etc., based on the second keyword corresponding to this situation, such as multiple retransmissions.

[0116] According to embodiments of this disclosure, during the chain establishment process, if multiple retransmission packets are received and multiple RST packets are sent, and the Timestamp TSval is 0, etc., M data packets are characterized, including multiple data packets identified as retransmissions, multiple data packets identified as RST, and multiple data packets identified as TSval=0. At this time, the artificial intelligence library can recommend checking whether the host TCP retransmission rate has reached the threshold alarm based on the second keyword corresponding to this situation, such as multiple retransmissions, multiple resets, and timestamp 0.

[0117] According to embodiments of this disclosure, based on M data packets, N network connection status identifiers and the number of identifiers corresponding to each of the N network connection status identifiers are obtained. This enables the analysis of the data stream state formed by data frames during network connection establishment and closure. This allows for the acquisition of N network connection status identifiers and the number of identifiers that accurately reflect the network connection status of the abnormal database in real time. Then, based on the N network connection status identifiers and the number of identifiers, a second keyword is determined. This enables the automatic abstraction of network fault phenomena into a standard error type second keyword, while simultaneously obtaining a more accurate second keyword that comprehensively reflects the network fault characteristics of the abnormal server.

[0118] According to embodiments of this disclosure, determining the second keyword based on M data packets may further include the following operations: obtaining N network connection status identifiers and identifier parameters corresponding to the N network connection status identifiers based on the M data packets, wherein N is an integer greater than or equal to 1 and less than or equal to M; determining the second keyword based on the N network connection status identifiers and the identifier parameters.

[0119] For example, the network connection status identifier is TSval, and the identifier parameter corresponding to TSval is 0, representing a timestamp of 0. By determining whether the identifier TSval is 0, the second keyword corresponding to TSval can be determined. If the identifier TSval is 0, the second keyword corresponding to TSval can be determined as a timestamp of 0; if the identifier TSval is not 0, no keyword corresponding to TSval can be set.

[0120] According to embodiments of this disclosure, the fault information includes multiple indicator information, such as... Figure 2 The operation S210 shown, which determines the abnormal server and the first keyword corresponding to the abnormal server based on the monitoring threshold and the detected fault information, may include the following operations:

[0121] For each of the multiple metrics, if the metric exceeds the monitoring threshold corresponding to the metric, the monitored server is identified as an abnormal server.

[0122] Based on the indicator information, determine the format of the first keyword corresponding to the indicator from the first keyword library;

[0123] The first keyword is obtained based on the indicator information and the format of the first keyword corresponding to the indicator.

[0124] According to embodiments of this disclosure, multiple metrics may include at least one of the following: retransmission rate, packet error rate, packet loss rate, and response time.

[0125] According to embodiments of this disclosure, the indicator information can be any one of the following: information corresponding to retransmission rate, information corresponding to packet error rate, information corresponding to packet loss rate, and information corresponding to response time.

[0126] According to embodiments of this disclosure, the first keyword includes at least one of the following:

[0127] Slow access, error code return, network interruption, inability to access, incomplete data return, host retransmission rate reaching the first detection value, domain name dialing success rate lower than the second detection value, service latency exceeding the third detection value, and bandwidth continuously surging to the fourth detection value.

[0128] For example, the monitoring threshold corresponding to response time could be 200ms. If the monitored response time is 500ms, it is confirmed that the response time is too long. At this point, based on the response time metric, multiple first keywords corresponding to the response time are matched from the first keyword library. Based on the condition that the monitored response time is greater than the monitoring threshold corresponding to the response time, the final first keyword corresponding to the metric response time is matched from these multiple first keywords. Thus, the format corresponding to the final first keyword corresponding to the metric response time can be determined as the first keyword format corresponding to the metric response time, where the first keyword format corresponding to the metric response time is "service latency exceeds the third detection value".

[0129] Given that the first keyword corresponding to the indicator response time is "service latency exceeds the third detection value", the monitored response time of 500ms is replaced with the third detection value, and finally the first keyword corresponding to the indicator response time is "service latency exceeds 500ms".

[0130] According to embodiments of this disclosure, for each of multiple indicator information, if the indicator information is greater than the monitoring threshold corresponding to the indicator, the monitored server is identified as an abnormal server. This enables real-time monitoring of servers based on indicator information and the monitoring threshold corresponding to the indicator, automatically and accurately identifying abnormal servers. Then, based on the indicator information, a first keyword format corresponding to the indicator is determined from a first keyword library. Based on the indicator information and the first keyword format corresponding to the indicator, a first keyword is obtained. This enables the automatic abstraction of network fault phenomena into a first keyword of standard error type, while obtaining a more accurate first keyword that can comprehensively reflect the network fault characteristics of abnormal servers.

[0131] According to embodiments of this disclosure, the fault information includes fault type information reported by the client corresponding to the faulty server, for example... Figure 2The operation S210 shown, which determines the abnormal server and the first keyword corresponding to the abnormal server based on the monitoring threshold and the detected fault information, may also include the following operations:

[0132] Based on the fault type information, determine the format of the first keyword corresponding to the fault type information from the first keyword library;

[0133] The first keyword is obtained based on the fault type information and the format of the first keyword corresponding to the fault type information.

[0134] According to embodiments of this disclosure, the monitoring platform obtains fault type information by collecting network anomaly information reported by clients corresponding to the server in real time. The fault type information collected by the monitoring platform can be manually reported fault information from clients, or fault information directly displayed on communication packets automatically captured by the clients.

[0135] According to embodiments of this disclosure, the fault type information reported by the client corresponding to the faulty server may be, for example, slow network access, inability to obtain data, network interruption, connection timeout, return of exception code, incomplete data display, intermittent access, and data access error.

[0136] For example, the fault type information could be "slow network access." Based on this, multiple primary keywords corresponding to "slow network access" can be matched from the primary keyword database. Examples include "slow access," "slow network access," "waiting time 1 minute," and "slow network access, network speed 10kb / s." Since "slow network access" provides no numerical feedback, the final primary keyword corresponding to "slow network access" is obtained from these multiple primary keywords. Therefore, the format of this final primary keyword corresponding to "slow network access" can be determined as the primary keyword format for "slow network access." Because "slow network access" provides no numerical feedback, it can be determined as the final primary keyword.

[0137] According to an embodiment of this disclosure, when the fault type information includes a feedback value and the matched first keyword format also includes a value, the feedback value in the fault type information replaces the value in the first keyword format to obtain the final first keyword.

[0138] According to embodiments of this disclosure, based on fault type information, a first keyword format corresponding to the fault type information is determined from a first keyword library. Based on the fault type information and the first keyword format corresponding to the fault type information, a first keyword is obtained, thereby automatically abstracting network fault phenomena into standard error types of first keywords, and obtaining more accurate first keywords that can comprehensively reflect the network fault characteristics of abnormal servers.

[0139] According to embodiments of this disclosure, for example, Figure 2 Operation S240, as shown, obtains the cause of the target network failure based on the first keyword, the second keyword, and the artificial intelligence library. This may include the following operations:

[0140] Based on the first and second keywords, we obtain the keyword group;

[0141] The keyword groups are matched with the index terms of the artificial intelligence library to obtain the target index terms;

[0142] Based on the target index terms and the artificial intelligence library, the initial network failure cause is determined;

[0143] Based on the initial cause of the network failure, the configuration information corresponding to the abnormal server is checked to determine the cause of the target network failure.

[0144] According to embodiments of this disclosure, the configuration information corresponding to the abnormal server is detected based on the initial network failure cause to determine the target network failure cause, including:

[0145] Based on the cause of the initial network failure, determine the abnormal configuration corresponding to the abnormal server;

[0146] Retrieve the configuration information corresponding to the abnormal configuration;

[0147] If the initial network failure cause matches the configuration information, the initial network failure cause will be identified as the target network failure cause.

[0148] According to embodiments of this disclosure, duplicate keywords in the first keyword and the second keyword can be removed to obtain a keyword group.

[0149] According to embodiments of this disclosure, keyword groups can be input into an artificial intelligence (AI) library. The AI ​​algorithms included in the library are used to match the keyword groups with the indexes in the AI ​​library to obtain target index terms. Then, using AI algorithms, based on the target index terms, the initial network failure cause is retrieved from the AI ​​library.

[0150] For example, keyword phrases could include: increased host retransmission rate, multiple receipt of retransmission packets, multiple sending of connection reset packets, and a timestamp value of 0. The initial network failure cause output by the AI ​​library could be: a bug in the Linux 3.10 kernel, or that proxies like HAProxy causing packet loss and retransmission when performing TCP health checks on application services.

[0151] Then, based on the initial network failure cause, the system can automatically obtain the kernel version information corresponding to the abnormal server and the health check method information of the proxy applications deployed on the abnormal server. It can automatically detect the kernel version corresponding to the abnormal server to determine if it is 3.10, and detect the health check method information of the deployed proxy applications to determine if they are using TCP health checks. If the initial network failure cause recommended by the AI ​​library is confirmed to be correct, it is determined as the target network failure cause. After determining the target network failure cause, the configuration information corresponding to the abnormal server can be adjusted and repaired based on the target network failure cause, such as automatically upgrading the system kernel or changing TCP health checks to HTTP health checks.

[0152] According to embodiments of this disclosure, a technical solution is implemented by obtaining a keyword group based on a first keyword and a second keyword, matching the keyword group with an index term group in an artificial intelligence library to obtain a target index term, obtaining an initial network failure cause based on the target index term and the artificial intelligence library, and detecting the configuration information corresponding to the abnormal server based on the initial network failure cause to determine the target network failure cause. This solution can automatically analyze and locate the network failure cause based on the first keyword, the second keyword, and the artificial intelligence library, improving the efficiency of network failure detection, and automatically verifying the network failure cause to obtain a more accurate target network failure cause.

[0153] According to the embodiments of this disclosure, by utilizing the technical solution of determining the abnormal configuration corresponding to the abnormal server based on the initial network failure cause, obtaining the configuration information corresponding to the abnormal configuration, and determining the initial network failure cause as the target network failure cause when the initial network failure cause matches the configuration information, the technical solution can automatically verify the initial network failure cause and obtain a more accurate target network failure cause.

[0154] Figure 4 A schematic diagram illustrating fault analysis according to an embodiment of the present disclosure is shown.

[0155] In step 400, firstly, based on the indicator information, the format of the first keyword corresponding to the indicator is determined from the first keyword library 401. Based on the indicator information and the first keyword format corresponding to the indicator, the first keyword 402 is obtained. Similarly, based on the fault type information, the format of the first keyword corresponding to the fault type information is determined from the first keyword library 401. Based on the fault type information and the first keyword format corresponding to the fault type information, the first keyword 403 is obtained.

[0156] like Figure 4As shown, the first keyword database 401 may include: slow access, returned error code, network interruption, inaccessibility, incomplete returned data, host retransmission rate reaching the first detection value, domain name dialing success rate lower than the second detection value, service latency exceeding the third detection value, bandwidth continuously surging to the fourth detection value, etc.

[0157] Then, based on M data packets, N network connection status identifiers and the corresponding identifier counts are obtained. Based on the N network connection status identifiers and their counts, the second keyword 405 is determined. The format of the second keyword corresponding to each network connection status identifier can be determined from the second keyword library 404 based on the N network connection status identifiers and their counts. Finally, the second keyword 405 is obtained based on the network connection status identifiers and their corresponding second keyword formats.

[0158] like Figure 4 As shown, the second keyword library 404 may include: retransmission, port reuse, reset, network window, checksum, timestamp value field is 0, etc.

[0159] Based on the first keyword 402, the first keyword 403, and the second keyword 405, the keyword group 406 is obtained. The keyword group 406 may include the host retransmission rate reaching 10%, reset, and the timestamp value field being 0.

[0160] After obtaining keyword group 406, keyword group 406 is input into artificial intelligence library 407. The artificial intelligence algorithm included in artificial intelligence library 407 is used to analyze keyword group 406 and output the cause of target network failure.

[0161] like Figure 4 As shown, by using the first keyword and the second keyword, a keyword group is obtained. The keyword group is then input into an artificial intelligence library for analysis to determine the cause of the target network failure. This enables automatic analysis and location of network failure causes based on the first keyword, the second keyword, and the artificial intelligence library, thereby improving the efficiency of network failure detection.

[0162] Figure 5 A flowchart illustrating automatic network fault handling according to another embodiment of the present disclosure is shown.

[0163] like Figure 5 As shown, in step S510, the monitored fault information is obtained. Step S510 includes sub-steps S511 and S512. In step S511, the indicator information monitored by the monitoring platform is obtained. In step S512, the fault type information reported by the client corresponding to the abnormal server is obtained.

[0164] In step S520, based on the monitoring threshold and the detected fault information, the abnormal server and the first keyword corresponding to the abnormal server are determined.

[0165] In step S530, M data packets reported by the abnormal server are automatically captured within a preset time period.

[0166] In step S540, according to... Figure 4 The fault analysis method shown analyzes indicator information, fault type information, and M data packets to obtain keyword groups.

[0167] In step S550, the keyword group is input into the artificial intelligence library for analysis to obtain the initial network fault cause.

[0168] In step S560, the configuration information corresponding to the abnormal server is detected according to the initial network failure cause to determine the target network failure cause, and the configuration information corresponding to the abnormal server is adjusted according to the target network failure cause.

[0169] In step S570, the initial cause of the network fault is pushed to the operations and maintenance personnel using the network fault handling interface. In step S580, after the operations and maintenance personnel authorize automatic emergency handling, step S560 is executed.

[0170] According to the embodiments of this disclosure, after step S550 is executed, steps S570, S580 and S560 can be executed according to the system settings, or step S560 can be executed directly.

[0171] Executing steps S570, S580, and S560 after completing step S550 can improve the efficiency of network fault handling while increasing the control of maintenance personnel over the network fault repair process. Directly executing step S560 can speed up the network fault handling process.

[0172] like Figure 5 The network fault automatic handling method shown can obtain a relatively accurate target network fault cause by combining the first and second keywords with an artificial intelligence library. Then, based on the relatively accurate target network fault cause, it automatically adjusts the configuration information of the abnormal server to restore the network status of the abnormal server to normal. It completes one-stop automatic operation and maintenance from fault discovery, fault analysis to fault handling, which greatly improves operation and maintenance efficiency and troubleshooting efficiency. Moreover, production emergencies have high requirements for timeliness. Timely discovery and resolution of production fault points (hosts, applications, etc.) can greatly ensure the availability of production systems.

[0173] Based on the above-described automatic network fault handling method, this disclosure also provides an automatic network fault handling device. The following will be combined with... Figure 6The device is described in detail.

[0174] Figure 6 A schematic block diagram of a network fault automatic handling apparatus according to an embodiment of the present disclosure is shown.

[0175] like Figure 6 As shown, the network fault automatic processing device 600 of this embodiment includes a first determination module 610, a capture module 620, a second determination module 630, an acquisition module 640, and an adjustment module 650.

[0176] The first determining module 610 is used to determine the abnormal server and the first keyword corresponding to the abnormal server based on the monitoring threshold and the detected fault information. In one embodiment, the first determining module 610 can be used to perform the operation S210 described above, which will not be repeated here.

[0177] The capture module 620 is used to automatically capture M data packets returned by the abnormal server within a preset time period, where M is an integer greater than or equal to 1. In one embodiment, the capture module 620 can be used to perform the operation S220 described above, which will not be repeated here.

[0178] The second determining module 630 is used to determine the second keyword based on M data packets. In one embodiment, the second determining module 630 can be used to perform the operation S230 described above, which will not be repeated here.

[0179] The obtaining module 640 is used to obtain the cause of the target network failure based on the first keyword, the second keyword, and the artificial intelligence library. In one embodiment, the obtaining module 640 can be used to perform the operation S240 described above, which will not be repeated here.

[0180] The adjustment module 650 is used to adjust the configuration information corresponding to the abnormal server according to the cause of the target network failure, so as to restore the network status of the abnormal server to normal. In one embodiment, the adjustment module 650 can be used to perform the operation S250 described above, which will not be repeated here.

[0181] According to embodiments of this disclosure, the second determining module includes a first obtaining submodule and a first determining submodule.

[0182] The first submodule is used to obtain N network connection status identifiers and the number of identifiers corresponding to each of the N network connection status identifiers based on M data packets, where N is an integer greater than or equal to 1 and less than or equal to M.

[0183] The first determination submodule is used to determine the second keyword based on N network connection status identifiers and the number of identifiers.

[0184] According to embodiments of this disclosure, the fault information includes multiple indicator information, and the first determining module includes a second determining submodule, a third determining submodule, and a second obtaining submodule.

[0185] The second determination submodule is used to determine the monitored server as an abnormal server if the value of the indicator exceeds the monitoring threshold corresponding to the indicator.

[0186] The third determination submodule is used to determine the format of the first keyword corresponding to the indicator from the first keyword library based on the indicator information.

[0187] The second submodule is used to obtain the first keyword based on the indicator information and the format of the first keyword corresponding to the indicator.

[0188] According to embodiments of this disclosure, the fault information includes fault type information fed back by the client corresponding to the abnormal server, and the first determining module further includes a fourth determining submodule and a third obtaining submodule.

[0189] The fourth determination submodule is used to determine the format of the first keyword corresponding to the fault type information from the first keyword library based on the fault type information.

[0190] The third submodule is used to obtain the first keyword based on the fault type information and the first keyword format corresponding to the fault type information.

[0191] According to embodiments of this disclosure, the obtaining module includes a fourth obtaining submodule, a fifth obtaining submodule, a sixth obtaining submodule, and a fifth determining submodule.

[0192] The fourth submodule is used to obtain keyword groups based on the first and second keywords.

[0193] The fifth submodule is used to match keyword groups with indexed word groups in the artificial intelligence library to obtain target indexed words.

[0194] The sixth submodule is used to obtain the initial network fault cause based on the target index terms and the artificial intelligence library.

[0195] The fifth determination submodule is used to detect the configuration information corresponding to the abnormal server based on the initial network failure cause, and determine the cause of the target network failure.

[0196] According to embodiments of this disclosure, the fifth determining submodule includes a first determining unit, an acquiring unit, and a second determining unit.

[0197] The first determining unit is used to determine the abnormal configuration corresponding to the abnormal server based on the cause of the initial network failure.

[0198] The acquisition unit is used to acquire configuration information corresponding to the abnormal configuration.

[0199] The second determining unit is used to determine the initial network failure cause as the target network failure cause when the initial network failure cause matches the configuration information.

[0200] According to embodiments of this disclosure, the first keyword includes at least one of the following:

[0201] Slow access, error code return, network interruption, inability to access, incomplete data return, host retransmission rate reaching the first detection value, domain name dialing success rate lower than the second detection value, service latency exceeding the third detection value, and bandwidth continuously surging to the fourth detection value.

[0202] According to embodiments of this disclosure, the second keyword includes at least one of the following:

[0203] Retransmission, port reuse, reset, network window, timestamp, increased host retransmission rate, multiple retransmission packets received, multiple connection reset packets sent, timestamp value field is 0.

[0204] According to embodiments of this disclosure, any plurality of modules among the first determining module 610, grasping module 620, second determining module 630, obtaining module 640, and adjusting module 650 may be combined into one module, or any one of these modules may be split into multiple modules. Alternatively, at least part of the functionality of one or more of these modules may be combined with at least part of the functionality of other modules and implemented in one module. According to embodiments of this disclosure, at least one of the first determining module 610, grasping module 620, second determining module 630, obtaining module 640, and adjusting module 650 may be at least partially implemented as hardware circuitry, such as a field-programmable gate array (FPGA), a programmable logic array (PLA), a system-on-a-chip, a system-on-a-substrate, a system-on-package, an application-specific integrated circuit (ASIC), or implemented in hardware or firmware by any other reasonable means of integrating or packaging the circuitry, or implemented in software, hardware, or firmware, or in any suitable combination of any of these three implementation methods. Alternatively, at least one of the first determining module 610, the grasping module 620, the second determining module 630, the obtaining module 640, and the adjusting module 650 can be at least partially implemented as a computer program module, which can perform corresponding functions when the computer program module is run.

[0205] Figure 7 A block diagram schematically illustrates an electronic device suitable for implementing an automatic network fault handling method according to an embodiment of the present disclosure.

[0206] like Figure 7As shown, an electronic device 700 according to an embodiment of the present disclosure includes a processor 701, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 702 or a program loaded from a storage portion 708 into a random access memory (RAM) 703. The processor 701 may include, for example, a general-purpose microprocessor (e.g., a CPU), an instruction set processor and / or an associated chipset and / or a special-purpose microprocessor (e.g., an application-specific integrated circuit (ASIC)), etc. The processor 701 may also include onboard memory for caching purposes. The processor 701 may include a single processing unit or multiple processing units for performing different actions of the method flow according to an embodiment of the present disclosure.

[0207] RAM 703 stores various programs and data required for the operation of electronic device 700. Processor 701, ROM 702, and RAM 703 are interconnected via bus 704. Processor 701 performs various operations of the method flow according to embodiments of the present disclosure by executing programs in ROM 702 and / or RAM 703. It should be noted that the programs may also be stored in one or more memories other than ROM 702 and RAM 703. Processor 701 may also perform various operations of the method flow according to embodiments of the present disclosure by executing programs stored in said one or more memories.

[0208] According to embodiments of this disclosure, the electronic device 700 may further include an input / output (I / O) interface 705, which is also connected to a bus 704. The electronic device 700 may also include one or more of the following components connected to the I / O interface 705: an input section 706 including a keyboard, mouse, etc.; an output section 707 including a cathode ray tube (CRT), liquid crystal display (LCD), etc., and a speaker, etc.; a storage section 708 including a hard disk, etc.; and a communication section 709 including a network interface card such as a LAN card, modem, etc. The communication section 709 performs communication processing via a network such as the Internet. A drive 710 is also connected to the I / O interface 705 as needed. A removable medium 711, such as a disk, optical disk, magneto-optical disk, semiconductor memory, etc., is installed on the drive 710 as needed so that computer programs read from it can be installed into the storage section 708 as needed.

[0209] This disclosure also provides a computer-readable storage medium, which may be included in the device / apparatus / system described in the above embodiments; or it may exist independently and not assembled into the device / apparatus / system. The computer-readable storage medium carries one or more programs that, when executed, implement the method according to the embodiments of this disclosure.

[0210] According to embodiments of this disclosure, the computer-readable storage medium may be a non-volatile computer-readable storage medium, such as, but not limited to: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this disclosure, the computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. For example, according to embodiments of this disclosure, the computer-readable storage medium may include ROM 702 and / or RAM 703 and / or one or more memories other than ROM 702 and RAM 703 described above.

[0211] Embodiments of this disclosure also include a computer program product comprising a computer program containing program code for performing the methods shown in the flowchart. When the computer program product is run on a computer system, the program code enables the computer system to implement the automatic network fault handling method provided in the embodiments of this disclosure.

[0212] When the computer program is executed by the processor 701, it performs the functions defined in the system / apparatus of this disclosure embodiments. According to embodiments of this disclosure, the systems, apparatuses, modules, units, etc., described above can be implemented by computer program modules.

[0213] In one embodiment, the computer program may rely on a tangible storage medium such as an optical storage device or a magnetic storage device. In another embodiment, the computer program may also be transmitted and distributed in the form of signals over a network medium, and may be downloaded and installed via the communication section 709, and / or installed from a removable medium 711. The program code contained in the computer program can be transmitted using any suitable network medium, including but not limited to: wireless, wired, etc., or any suitable combination thereof.

[0214] In such an embodiment, the computer program can be downloaded and installed from a network via the communication section 709, and / or installed from the removable medium 711. When the computer program is executed by the processor 701, it performs the functions defined in the system of this disclosure embodiment. According to embodiments of this disclosure, the systems, devices, apparatuses, modules, units, etc., described above can be implemented by computer program modules.

[0215] According to embodiments of this disclosure, program code for executing the computer programs provided in embodiments of this disclosure can be written in any combination of one or more programming languages. Specifically, these computational programs can be implemented using high-level procedural and / or object-oriented programming languages, and / or assembly / machine languages. Programming languages ​​include, but are not limited to, languages ​​such as Java, C++, Python, "C", or similar programming languages. The program code can execute entirely on the user's computing device, partially on the user's device, partially on a remote computing device, or entirely on a remote computing device or server. In cases involving remote computing devices, the remote computing device can be connected to the user's computing device via any type of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computing device (e.g., via the Internet using an Internet service provider).

[0216] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in a block diagram or flowchart, and combinations of blocks in a block diagram or flowchart, may be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0217] Those skilled in the art will understand that the features described in the various embodiments and / or claims of this disclosure can be combined or combined in various ways, even if such combinations or combinations are not explicitly described in this disclosure. In particular, the features described in the various embodiments and / or claims of this disclosure can be combined or combined in various ways without departing from the spirit and teachings of this disclosure. All such combinations and / or combinations fall within the scope of this disclosure.

[0218] The embodiments of this disclosure have been described above. However, these embodiments are for illustrative purposes only and are not intended to limit the scope of this disclosure. Although various embodiments have been described above, this does not mean that the measures in the various embodiments cannot be used advantageously in combination. The scope of this disclosure is defined by the appended claims and their equivalents. Various substitutions and modifications can be made by those skilled in the art without departing from the scope of this disclosure, and all such substitutions and modifications should fall within the scope of this disclosure.

Claims

1. An automatic network fault handling method, comprising: Based on the monitoring threshold and the detected fault information, an abnormal server and a first keyword corresponding to the abnormal server are determined, wherein the first keyword represents the standard error type reflecting the network fault extracted from the fault information. During a preset time period, M data packets fed back by the abnormal server are automatically captured, where M is an integer greater than or equal to 1; Based on the M data packets, a second keyword is determined, wherein the second keyword represents the standard error type reflecting network failure extracted from the data packets; Based on the first keyword and the second keyword, a keyword group is obtained; the keyword group is matched with the index word group of the artificial intelligence library to obtain the target index word; based on the target index word and the artificial intelligence library, the initial network failure cause is obtained; based on the initial network failure cause, the abnormal configuration corresponding to the abnormal server is determined; the configuration information corresponding to the abnormal configuration is obtained; if the initial network failure cause matches the configuration information, the initial network failure cause is determined as the target network failure cause. Based on the cause of the target network failure, the configuration information corresponding to the abnormal server is automatically adjusted to restore the network status of the abnormal server to normal.

2. The method according to claim 1, wherein, The step of determining the second keyword based on the M data packets includes: Based on the M data packets, N network connection status identifiers and the number of identifiers corresponding to each of the N network connection status identifiers are obtained, where N is an integer greater than or equal to 1 and less than or equal to M; The second keyword is determined based on the N network connection status identifiers and the number of identifiers.

3. The method according to claim 1 or 2, wherein, The fault information includes multiple indicator information. The process of determining the abnormal server and the first keyword corresponding to the abnormal server based on the monitoring threshold and the detected fault information includes: For each of the multiple indicator information, if the indicator information is greater than the monitoring threshold corresponding to the indicator, the monitored server is identified as the abnormal server. Based on the indicator information, determine the format of the first keyword corresponding to the indicator from the first keyword library; The first keyword is obtained based on the indicator information and the format of the first keyword corresponding to the indicator.

4. The method according to claim 3, wherein, The fault information includes fault type information reported by the client corresponding to the abnormal server. The step of determining the abnormal server and the first keyword corresponding to the abnormal server based on the monitoring threshold and the detected fault information further includes: Based on the fault type information, determine the format of the first keyword corresponding to the fault type information from the first keyword library; The first keyword is obtained based on the fault type information and the first keyword format corresponding to the fault type information.

5. The method according to claim 1, wherein, The first keyword includes at least one of the following: Slow access, error code return, network interruption, inability to access, incomplete data return, host retransmission rate reaching the first detection value, domain name dialing success rate lower than the second detection value, service latency exceeding the third detection value, and bandwidth continuously surging to the fourth detection value.

6. The method according to claim 1, wherein, The second keyword includes at least one of the following: Retransmission, port reuse, reset, network window, timestamp, increased host retransmission rate, multiple retransmission packets received, multiple connection reset packets sent, timestamp value field is 0.

7. An automatic network fault handling device, comprising: The first determining module is used to determine the abnormal server and the first keyword corresponding to the abnormal server based on the monitoring threshold and the monitored fault information, wherein the first keyword represents the standard error type reflecting the network fault extracted from the fault information. The capture module is used to automatically capture M data packets returned by the abnormal server within a preset time period, where M is an integer greater than or equal to 1; The second determining module is used to determine a second keyword based on the M data packets, wherein the second keyword represents a standard error type reflecting a network fault extracted from the data packets; The module is used to obtain the cause of the target network failure based on the first keyword, the second keyword, and the artificial intelligence library; The obtaining module includes a fourth obtaining submodule, a fifth obtaining submodule, a sixth obtaining submodule, and a fifth determining submodule; The fourth submodule is used to obtain keyword groups based on the first and second keywords; The fifth submodule is used to match keyword groups with index word groups in the artificial intelligence library to obtain target index words; The sixth submodule is used to obtain the initial network fault cause based on the target index terms and the artificial intelligence library; The fifth determination submodule is used to detect the configuration information corresponding to the abnormal server based on the initial network failure cause to determine the cause of the target network failure. The fifth determining submodule includes a first determining unit, an acquisition unit, and a second determining unit; The first determining unit is used to determine the abnormal configuration corresponding to the abnormal server based on the cause of the initial network failure. The acquisition unit is used to acquire configuration information corresponding to the abnormal configuration. The second determining unit is used to determine the initial network failure cause as the target network failure cause when the initial network failure cause matches the configuration information. The adjustment module is used to automatically adjust the configuration information corresponding to the abnormal server according to the cause of the target network failure, so as to restore the network status of the abnormal server to normal.

8. An electronic device, comprising: One or more processors; Storage device for storing one or more programs. Wherein, when the one or more programs are executed by the one or more processors, the one or more processors perform the method according to any one of claims 1 to 6.

9. A computer-readable storage medium having stored thereon executable instructions that, when executed by a processor, cause the processor to perform the method according to any one of claims 1 to 6.

10. A computer program product comprising a computer program that, when executed by a processor, implements the method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Network transmission abnormity processing method and device, electronic equipment and storage medium

    CN114285727A

  • Network communication fault troubleshooting method and device, equipment and storage medium

    CN115580527A