Network fault loss limiting method and device, electronic equipment and readable storage medium
By combining multi-data source processing of network probe points and client-side embedded data, the problems of high cost and poor timeliness in public network link fault detection are solved, and rapid and accurate fault mitigation is achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- DUXIAOMAN TECH (BEIJING) CO LTD
- Filing Date
- 2023-07-31
- Publication Date
- 2026-07-21
AI Technical Summary
In existing technologies, public network link fault detection is costly and has poor timeliness, and fault mitigation solutions are outdated, leading to extended fault duration.
By combining network detection points and client-side embedded data, and through data processing and verification from multiple data sources, the accuracy and timeliness of fault detection are achieved, the number of detection points is reduced, and detection point data is used as basic alarms, while embedded data is used for verification to execute fault handling plans.
It improves the accuracy and timeliness of fault detection, reduces costs, and shortens the time for fault prevention.
Smart Images

Figure CN116866206B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of computer network communication technology, and in particular to a method, apparatus, electronic device, and readable storage medium for preventing network failures. Background Technology
[0002] Currently, various clients request information from servers via public network links from telecom operators. However, if the public network link and its access points fail, user traffic cannot be processed by the backend service, leading to serious stability issues and impacting user experience. Therefore, the stability of the public network link is crucial. Commonly used methods for fault detection using public network probes require deploying a large number of probes, resulting in high costs. Client-side fault detection methods involve large amounts of data and lack timeliness. Both of these methods rely solely on manual intervention after fault detection, extending the downtime. Summary of the Invention
[0003] In view of this, embodiments of the present invention provide a network fault loss mitigation method to solve the problems of high network fault perception cost, poor timeliness, and outdated fault loss mitigation solutions.
[0004] According to one aspect of the present invention, a method for mitigating network failures is provided, comprising:
[0005] Receive a first input, which is first data obtained by the network probe point. The first data includes, but is not limited to, HTTP status data, HTTP request time data, and domain name resolution data.
[0006] In response to the first input, first abnormal information is obtained based on the first data according to the location and the operator. The first abnormal information includes abnormal information on domain name availability, abnormal information on average response time, and abnormal information on hijacking.
[0007] Receive a second input, which includes second data uploaded by the client's embedded points;
[0008] In response to the second input, if the first abnormal information meets the preset abnormal conditions, the second data is used to verify the first abnormal information that meets the preset abnormal conditions. If the verification determines that the network is in a fault state, the network fault handling plan is executed.
[0009] Optionally, the step of responding to the first input and obtaining the first abnormal information based on the first data according to the home location and home operator includes:
[0010] HTTP status data, HTTP request time data, and domain name resolution data are aggregated according to the location and the operator, respectively, to obtain domain name availability information, average response time information, and hijacking information for different locations and operators.
[0011] The number of abnormal information in the domain name availability information, average response time information, and hijacking information for different origins and different operators is counted separately to obtain abnormal information in domain name availability, average response time, and hijacking situation for different origins and different operators.
[0012] Optionally, in response to the second input, if the first abnormal information meets preset abnormal conditions, the second data is used to verify the first abnormal information that meets the preset abnormal conditions. When the verification determines that the network is in a fault state, a network fault handling plan is executed, including:
[0013] Set separate thresholds for abnormal information regarding domain name availability, average response time, and hijacking for different domain name locations and different ISPs.
[0014] If the abnormal information of domain name availability, average response time, or hijacking exceeds the corresponding threshold for abnormal information of domain name availability, average response time, and hijacking, the first abnormal information is determined to meet the preset abnormal conditions.
[0015] According to the first abnormal information that meets the preset abnormal conditions, obtain the target second data of the same location, same operator and same sequence, and use the target second data to verify the first abnormal information;
[0016] If the verification result meets the preset verification conditions, the network is determined to be in a fault state, and the network fault handling plan is executed.
[0017] Optionally, the implementation of the network fault handling plan includes:
[0018] Cut off the IP traffic corresponding to the first abnormal information.
[0019] According to a second aspect of the present invention, a network failure mitigation device is provided, comprising:
[0020] The first receiving module is used to receive a first input, which is first data obtained by the network probe point. The first data includes, but is not limited to, HTTP status data, HTTP request time data, and domain name resolution data.
[0021] The first acquisition module, in response to the first input, acquires first abnormal information based on the first data according to the location and the operator, the first abnormal information including abnormal information on domain name availability, abnormal information on average response time, and abnormal information on hijacking.
[0022] The second receiving module is used to receive a second input, which includes second data uploaded by the client's embedded points;
[0023] The fault handling module, in response to the second input, if the first abnormal information meets the preset abnormal conditions, uses the second data to verify the first abnormal information that meets the preset abnormal conditions. When the verification determines that the network is in a fault state, it executes the network fault handling plan.
[0024] Optionally, the first acquisition module includes:
[0025] The aggregation module is used to aggregate HTTP status data, HTTP request time data, and domain name resolution data according to the location and carrier, respectively, to obtain domain name availability information, average response time information, and hijacking information for different locations and carriers.
[0026] The statistics module is used to count the number of abnormal information in domain name availability information, average response time information, and hijacking information for different home locations and different home operators, and to obtain abnormal information in domain name availability, average response time, and hijacking situations for different home locations and different home operators.
[0027] Optionally, the fault handling module includes:
[0028] The threshold setting module is used to set the abnormal information thresholds for domain name availability, average response time, and hijacking for different regions and different operators.
[0029] The determination module is used to determine that the first abnormal information meets the preset abnormal conditions when the abnormal information of domain name availability, the abnormal information of average response time, or the abnormal information of hijacking exceeds the corresponding abnormal information thresholds of domain name availability, average response time, and hijacking.
[0030] The verification module is used to obtain target second data of the same location, same operator and same sequence according to the first abnormal information that meets the preset abnormal conditions, and to verify the first abnormal information using the target second data.
[0031] The execution module is used to determine that the network is in a fault state when the verification result meets the preset verification conditions, and to execute the network fault handling plan.
[0032] Optionally, the fault handling module includes:
[0033] The traffic control module is used to cut off the IP traffic corresponding to the first abnormal information.
[0034] According to a third aspect of the present invention, an electronic device is provided, comprising:
[0035] Processor; and
[0036] Stored program memory,
[0037] The program includes instructions that, when executed by the processor, cause the processor to perform the method according to any one of the first aspects of the invention.
[0038] According to a fourth aspect of the present invention, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are configured to cause a computer to perform the method according to any one of the first aspects of the present invention.
[0039] The one or more technical solutions provided in this application embodiment utilize network probe points for fault detection and client-side embedded points for verifying the fault detection results. This makes the fault detection results more accurate, and the existence of verification data reduces the number of network probe points deployed, thereby lowering the cost of network fault detection. The two types of data complement each other, solving the problems of accuracy, timeliness, and cost in network detection. Furthermore, the contingency plan for traffic diversion is executed in fault conditions, greatly improving the speed of fault mitigation. Attached Figure Description
[0040] Further details, features, and advantages of the invention are disclosed in the following description of exemplary embodiments in conjunction with the accompanying drawings, in which:
[0041] Figure 1 A schematic diagram of an example system in which the various methods described herein may be implemented according to exemplary embodiments of the present invention is shown;
[0042] Figure 2 A flowchart of a network failure mitigation method according to an exemplary embodiment of the present invention is shown;
[0043] Figure 3 A schematic block diagram of a network failure mitigation device according to an exemplary embodiment of the present invention is shown;
[0044] Figure 4 A structural block diagram of an exemplary electronic device that can be used to implement embodiments of the present invention is shown. Detailed Implementation
[0045] Embodiments of the present invention will now be described in more detail with reference to the accompanying drawings. While some embodiments of the invention are shown in the drawings, it should be understood that the invention can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of the invention. It should be understood that the accompanying drawings and embodiments are for illustrative purposes only and are not intended to limit the scope of protection of the invention.
[0046] It should be understood that the various steps described in the method embodiments of the present invention may be performed in different orders and / or in parallel. Furthermore, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present invention is not limited in this respect.
[0047] The term "comprising" and its variations as used herein are open-ended, meaning "including but not limited to". The term "based on" means "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". Definitions of other terms will be given in the following description. It should be noted that the concepts of "first", "second", etc., mentioned in this invention are used only to distinguish different devices, modules, or units, and are not intended to limit the order of functions performed by these devices, modules, or units or their interdependencies.
[0048] It should be noted that the terms "a" and "a plurality of" used in this invention are illustrative rather than restrictive. Those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".
[0049] The names of the messages or information exchanged between the multiple devices in the embodiments of the present invention are for illustrative purposes only and are not intended to limit the scope of these messages or information.
[0050] The present invention will now be described with reference to the accompanying drawings. The technical solutions provided by the embodiments of this application will be explained in detail through specific examples and application scenarios.
[0051] With the full development of the mobile internet era, the technological barriers to customer acquisition among internet companies have gradually disappeared. Each client can easily acquire a large number of new users, making user retention the key issue. Only by continuously iterating and improving the user experience can user retention be maximized. In other words, current internet companies are essentially selling users a better service experience.
[0052] If a good user experience is the foundation for user growth, then good operability and stability are the foundation for building a good user experience. Operability largely relies on designers providing attractive interfaces and smooth interactions. Stability, on the other hand, requires a reliable and stable information system jointly built by clients (such as apps and H5 pages) and servers. Various clients request information from the server through the operator's public network link to complete the interaction. However, if the public network link and its access point fail, user traffic cannot be processed by the backend service, causing serious stability issues and impacting the user experience. Therefore, the stability of the public network link is extremely important.
[0053] Common public network link failure scenarios include public network access point failures in data centers and domain name hijacking by a certain operator in a certain province. In these failure scenarios, in addition to manual discovery or customer complaints, the following two methods are often used:
[0054] I. Method of using public network probe points for detection
[0055] Simulate user requests to actively execute probe tasks, report the probe results to the monitoring system, and use the network status of the probe point to characterize the network status of the network operator in the province and city where the probe point is located.
[0056] The drawback of this method is that the overall accuracy of monitoring depends on the accuracy of the detection points. Therefore, it is necessary to deploy a large number of detection points in various provinces, cities and operator data centers, or purchase detection data from testing vendors. While pursuing high accuracy, there are certain economic costs.
[0057] II. Methods of detecting data points on the client side
[0058] This method reports information related to the client's public network requests to the monitoring system. The drawback of this method is that it requires aggregating network data from a large number of users in the monitoring system. If all users' network quality reports are submitted, the data volume will be too large, and there will be a time delay after calculation and aggregation, resulting in poor timeliness.
[0059] Therefore, common public network fault detection methods have certain problems in terms of accuracy, timeliness, and cost. On the other hand, the loss mitigation actions after a fault occurs and is detected mainly rely on manual labor, resulting in excessively long fault durations.
[0060] like Figure 1As shown, different methods are used to acquire probe point data, which is stored in a time-series database and anomaly identification is performed. After anomaly identification, the acquired embedded data is used for anomaly verification, and the acquired embedded data is also stored in the time-series database. After verification, a fault contingency plan is executed. The contingency plan implements DNS management, endpoint configuration management, and management of the code deployment platform and the reach platform. Fault monitoring using data from multiple data sources makes the fault perception results more accurate. Both types of data can also reduce the number of network probe points deployed, reducing costs. This application embodiment uses both probe point data and embedded data in the fault discovery phase, complementing each other to solve the problems of accuracy, timeliness, and cost. Probe point data is used as the basic fault alarm, and embedded data is used as its calibration.
[0061] like Figure 2 As shown, Figure 2 This is a flowchart illustrating a network failure mitigation method provided in an embodiment of this application. The method may include the following steps S201 to S204:
[0062] S201, Receive a first input, the first input being first data obtained by the network probe point, the first data including but not limited to HTTP status data, HTTP request time data and domain name resolution data.
[0063] In this embodiment, the first data sources include self-built public network probe points and third-party testing platforms. The self-built public network probe points are established in the data centers of the three major telecommunications operators in key regions served by the internet company, reducing the required scale and lowering the cost of self-built probe points. The number and distribution of probe points are based on the actual situation of the internet company. This embodiment also requires obtaining probe sources from the three major telecommunications operators covering 34 provincial-level administrative regions nationwide from third-party testing platforms, totaling 102 probe points. Using these two types of probe data sources as the basis for public network quality perception, the multi-data source-based automatic public network fault mitigation system has the capability to perceive network quality nationwide and ensures the accuracy of probes in key areas.
[0064] The probe points report the raw data of the probe requests every 30 seconds. After the raw data is cleaned, processed, and aggregated, it is stored in the time-series database. The HTTP status data consists of HTTP status codes.
[0065] S202, in response to the first input, obtain first abnormal information according to the first data based on the home location and home operator. The first abnormal information includes abnormal information on domain name availability, abnormal information on average response time, and abnormal information on hijacking.
[0066] In this embodiment, the acquired first data is processed. HTTP status data, HTTP request time data, and domain name resolution data are categorized according to their origin and carrier. Domain name availability is obtained based on the HTTP status data, average response time is obtained based on the HTTP request time data, and hijacking information is obtained based on the domain name resolution data. Abnormal events and their numbers for each origin and carrier are statistically analyzed to obtain abnormal information regarding domain name availability, average response time, and hijacking.
[0067] In one optional embodiment, the anomaly identification method for different first anomaly information is as follows:
[0068] According to the domain name availability anomaly identification rules based on location and operator: an anomaly data point is defined as any domain name availability less than 0.8 from any province or operator. When 8 anomalies occur within 10 statistical periods, an anomaly event is triggered, generating the first anomaly information.
[0069] According to the anomaly identification rules based on the average response time of the originating location and the originating operator: an average response time greater than 3 seconds from any province and any operator to a certain domain name is defined as an anomaly data point. When 8 anomalies occur within 10 statistical periods, an anomaly event is triggered, and the first anomaly information is generated.
[0070] According to the hijacking anomaly identification rules based on location and carrier: a hijacking data point greater than 0 from any province and any carrier to a certain domain is defined as an anomaly. When 8 anomalies occur within 10 statistical periods, an anomaly event is triggered, generating the first anomaly information.
[0071] S203, Receive the second input, the second input including the second data uploaded by the client's embedded points.
[0072] In this embodiment, the client reports the event tracking information. To ensure the timeliness of data generation, the data is processed with a sampling rate of 10%. That is, 10% of the client-reported information is randomly retained in each log sampling period, while the remaining 90% is discarded without processing. Furthermore, processing the data with a 10% sampling rate also ensures that the data storage volume will not be excessive. The sampled and filtered data is stored as the second data in the time-series database. The data items of the second data are shown below:
[0073] Table 1. Data items and descriptions of the second data reported by the data terminal.
[0074]
[0075]
[0076] S204, in response to the second input, if the first abnormal information meets the preset abnormal conditions, the second data is used to verify the first abnormal information that meets the preset abnormal conditions. If the verification determines that the network is in a fault state, the network fault handling plan is executed.
[0077] In this embodiment, preset abnormal conditions are set. When the first abnormal information meets the preset abnormal conditions, the target second data corresponding to the first abnormal information is obtained from the time series database. The target second data is used to verify the first abnormal information. Based on the verification result, it is determined whether the network is in a fault state. In the event of a fault, the fault handling plan is executed quickly and efficiently to reduce the losses caused by the network fault.
[0078] In an optional embodiment, the step of obtaining first abnormal information based on the first data according to the location and carrier in response to the first input includes:
[0079] S2021, HTTP status data, HTTP request time data and domain name resolution data are aggregated according to the location and the operator, respectively, to obtain domain name availability information, average response time information and hijacking information for different locations and operators;
[0080] S2022, count the number of abnormal information in the domain name availability information, average response time information and hijacking information for different home locations and different home operators, and obtain abnormal information in domain name availability, average response time and hijacking situation for different home locations and different home operators.
[0081] In this embodiment, as shown in Table 2, the processing rules for the first data are as follows: the first data is aggregated according to the following rules, including HTTP status data, HTTP request time data, and domain name resolution data, to obtain domain name availability information, average response time information, and hijacking information for each home location and home operator. Specifically, for the HTTP status data, the number of requests with the same home location, the same operator, and the same return code is summed, and the total number of requests with the same home location and the same operator is summed.
[0082] Table 2. Processing rules for acquiring the first data from the detection points.
[0083]
[0084]
[0085] In an optional embodiment, in response to the second input, when the first abnormal information meets preset abnormal conditions, the second data is used to verify the first abnormal information that meets the preset abnormal conditions. When the verification determines that the network is in a fault state, a network fault handling plan is executed, including:
[0086] S2041, set abnormal information thresholds for domain name availability, average response time, and hijacking for different locations and different operators.
[0087] S2042, if the abnormal information of domain name availability, the abnormal information of average response time, or the abnormal information of hijacking in different locations and different operators is greater than the corresponding abnormal information threshold of domain name availability, the abnormal information threshold of average response time, and the abnormal information threshold of hijacking, it is determined that the first abnormal information satisfies the preset abnormal conditions.
[0088] S2043, obtain target second data of the same location, same operator and same sequence according to the first abnormal information that meets the preset abnormal conditions, and use the target second data to verify the first abnormal information;
[0089] S2044: When the verification result meets the preset verification conditions, the verification determines that the network is in a fault state and executes the network fault handling plan.
[0090] In one optional embodiment, a trigger-verification model is used to determine whether the first abnormal information meets the preset abnormal conditions. First, the first data in the time series database is input, and the first abnormal information is determined according to the abnormality rules. Then, the first abnormal information and the second data in the time series database are input for verification.
[0091] The verification method is as follows for different first-order anomaly information:
[0092] When a domain experiences a decrease in availability with a specific ISP in a particular province, the system retrieves data from the time-series database showing the number of successful requests for that domain to that ISP within the province during the abnormal period. If the ratio of successful requests to total requests is less than 0.8, then a decrease in availability is considered to have occurred. Otherwise, it is considered a false alarm and no action is taken.
[0093] When a domain experiences an anomaly in average response time with a specific ISP in a particular province, the system retrieves data from the time-series database showing the total request time for that domain within that province and ISP during the anomaly period. If the average response time (i.e., the sum of all request times / total number of requests) during that period is greater than 3 seconds, then an anomaly is considered to have indeed occurred. Otherwise, it is considered a false alarm and no action is taken.
[0094] When a domain experiences a DNS hijacking incident with a specific ISP in a particular province, the system retrieves a list of requesting server IPs for that domain from client-side reports within the abnormal time period from the time-series database. The entire list of requested IPs is deduplicated and merged. If any IPs not present in the authoritative DNS resolution results are found, then DNS hijacking is confirmed. Otherwise, it is considered a false alarm and no action is taken.
[0095] In an optional embodiment, the execution of the network fault handling plan includes:
[0096] S2045, cut off the IP traffic corresponding to the first abnormal information.
[0097] In this embodiment, when the first abnormal information is an availability reduction type fault, the IPs that failed to make requests during the fault period are aggregated. If all the aggregated IPs that failed to make requests are of type IPv4 or IPv6, then it is considered that there is an IPv4 or IPv6 fault, and the corresponding IPv4 or IPv6 traffic switching plan is executed. If all the aggregated IPs that failed to make requests are from the same public network access room, then it is considered that there is a fault in that access room, and the corresponding external network traffic switching plan for that room is executed. If it is impossible to distinguish the IP type and access room after aggregation, then it is considered that there is a fault in the network of that operator in that province, and the alarm plan is executed to contact the on-duty personnel.
[0098] When the first abnormal information is a response time increase type of fault, IPs with a total request time greater than 3 seconds during the fault period are aggregated. If all IPs with a total time greater than 3 seconds after aggregation are of IPv4 or IPv6 type, then an IPv4 or IPv6 fault is considered, and the corresponding IPv4 or IPv6 traffic switching plan is executed. If all IPs with a total time greater than 3 seconds after aggregation are from the same public network access data center, then a fault is considered to have occurred in that access data center, and the corresponding data center's external network traffic switching plan is executed. If the IP type and access data center cannot be distinguished after aggregation, then a fault is considered to have occurred in the network of that operator in that province, and an alarm plan is executed to contact the on-duty personnel.
[0099] When the first abnormal information is a DNS hijacking-type failure, the backup domain name switching plan is invoked, and the alarm plan is executed to reach the on-duty personnel.
[0100] In one optional embodiment, cutting off the IP traffic corresponding to the first abnormal information includes:
[0101] IPv4 traffic diversion contingency plan: Call the DNS management interface to delete all IPv4 A records for a specific domain.
[0102] IPv6 traffic diversion contingency plan: Call the DNS management interface to delete all IPv6 AAAA records for a specific domain.
[0103] Contingency plan for diverting external network traffic from the data center: Call the DNS management interface to delete the DNS resolution results for a specific domain name in a specific data center.
[0104] Backup domain name switching contingency plan: Invoke the client configuration management platform and code deployment platform to switch the domain name requested by the client to the backup domain name.
[0105] Alarm contingency plan: Invoke the contact platform to send a text message and alarm call to the on-duty personnel's mobile phone number.
[0106] The network fault mitigation method provided in this application utilizes network probe points for fault detection and client-side data points for verification of the fault detection results. This makes the fault detection results more accurate, and the presence of verification data reduces the number of network probe points required, lowering the cost of network fault detection. The two types of data complement each other, solving the problems of accuracy, timeliness, and cost in network detection. Furthermore, it executes a traffic diversion plan in fault conditions, significantly improving the speed of fault mitigation.
[0107] Corresponding to the above embodiments, see [link to relevant documentation]. Figure 3 This application also provides a network failure mitigation device 300, comprising:
[0108] The first receiving module 301 is used to receive a first input, the first input being first data obtained by the network probe point, the first data including but not limited to HTTP status data, HTTP request time data and domain name resolution data;
[0109] The first acquisition module 302, in response to the first input, acquires first abnormal information based on the first data according to the location and the operator, the first abnormal information including abnormal information on domain name availability, abnormal information on average response time and abnormal information on hijacking.
[0110] The second receiving module 303 is used to receive a second input, the second input including second data uploaded by the client's embedded points;
[0111] The fault handling module 304, in response to the second input, when the first abnormal information meets the preset abnormal conditions, uses the second data to verify the first abnormal information that meets the preset abnormal conditions. When the verification determines that the network is in a fault state, it executes the network fault handling plan.
[0112] Optionally, the first acquisition module 302 includes:
[0113] The aggregation module 3021 is used to aggregate HTTP status data, HTTP request time data and domain name resolution data according to the location and the operator, respectively, to obtain domain name availability information, average response time information and hijacking information for different locations and operators.
[0114] The statistics module 3022 is used to count the number of abnormal information in the domain name availability information, average response time information, and hijacking information for different home locations and different home operators, and to obtain abnormal information in domain name availability, average response time, and hijacking situations for different home locations and different home operators.
[0115] Optionally, the fault handling module 304 includes:
[0116] The threshold setting module 3041 is used to set the abnormal information thresholds for domain name availability, average response time, and hijacking situations for different home locations and different home operators, respectively.
[0117] The determination module 3042 is used to determine that the first abnormal information satisfies the preset abnormal conditions when the abnormal information of domain name availability, the abnormal information of average response time, or the abnormal information of hijacking exceeds the corresponding abnormal information thresholds of domain name availability, average response time, and hijacking.
[0118] Verification module 3043 is used to obtain target second data of the same location, same operator and same sequence according to the first abnormal information that meets the preset abnormal conditions, and to verify the first abnormal information using the target second data;
[0119] The execution module 3044 is used to determine that the network is in a fault state when the verification result meets the preset verification conditions, and to execute the network fault handling plan.
[0120] Optionally, the fault handling module 304 includes:
[0121] The traffic control module 3045 is used to cut off the IP traffic corresponding to the first abnormal information.
[0122] The network fault mitigation method provided in this application utilizes network probe points for fault detection and client-side data points for verification of the fault detection results. This makes the fault detection results more accurate, and the presence of verification data reduces the number of network probe points required, lowering the cost of network fault detection. The two types of data complement each other, solving the problems of accuracy, timeliness, and cost in network detection. Furthermore, it executes a traffic diversion plan in fault conditions, significantly improving the speed of fault mitigation.
[0123] An exemplary embodiment of the present invention also provides an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor. The memory stores a computer program executable by the at least one processor, the computer program being executed by the at least one processor to cause the electronic device to perform a method according to an embodiment of the present invention.
[0124] An exemplary embodiment of the present invention also provides a non-transitory computer-readable storage medium storing a computer program, wherein the computer program, when executed by a computer's processor, is used to cause the computer to perform a method according to an embodiment of the present invention.
[0125] An exemplary embodiment of the present invention also provides a computer program product, including a computer program, wherein, when executed by a computer's processor, the computer program is used to cause the computer to perform a method according to an embodiment of the present invention.
[0126] refer to Figure 4 The present invention will now be described in the form of a structural block diagram of an electronic device 400 that can serve as a server or client of the present invention, which is an example of a hardware device that can be applied to various aspects of the present invention. The electronic device is intended to represent various forms of digital electronic computer devices, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein.
[0127] like Figure 4 As shown, the electronic device 400 includes a computing unit 401, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 402 or a computer program loaded from a storage unit 408 into a random access memory (RAM) 403. The RAM 403 may also store various programs and data required for the operation of the device 400. The computing unit 401, ROM 402, and RAM 403 are interconnected via a bus 404. An input / output (I / O) interface 405 is also connected to the bus 404.
[0128] Multiple components in electronic device 400 are connected to I / O interface 405, including: input unit 406, output unit 404, storage unit 408, and communication unit 409. Input unit 406 can be any type of device capable of inputting information to electronic device 400. Input unit 406 can receive input digital or character information and generate key signal inputs related to user settings and / or function control of electronic device. Output unit 407 can be any type of device capable of presenting information and may include, but is not limited to, a display, speaker, video / audio output terminal, vibrator, and / or printer. Storage unit 408 may include, but is not limited to, disks and optical discs. Communication unit 409 allows electronic device 400 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks, and may include, but is not limited to, modems, network cards, infrared communication devices, wireless communication transceivers, and / or chipsets, such as Bluetooth™ devices, WiFi devices, WiMax devices, cellular communication devices, and / or the like.
[0129] The computing unit 401 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 401 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 401 performs the various methods and processes described above. For example, in some embodiments, the auditing method may be implemented as a computer software program tangibly contained in a machine-readable medium, such as storage unit 408. In some embodiments, part or all of the computer program may be loaded and / or installed on the electronic device 400 via ROM 402 and / or communication unit 409. In some embodiments, the computing unit 401 may be configured to perform the auditing method by any other suitable means (e.g., by means of firmware).
[0130] The program code used to implement the methods of the present invention can be written in any combination of one or more programming languages. This program code can be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing device, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code can be executed entirely on the machine, partially on the machine, as a standalone software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0131] In the context of this invention, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. Machine-readable media can include, but are not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0132] As used herein, the terms "machine-readable medium" and "computer-readable medium" refer to any computer program product, device, and / or apparatus (e.g., disk, optical disk, memory, programmable logic device (PLD)) for providing machine instructions and / or data to a programmable processor, including machine-readable media that receive machine instructions as machine-readable signals. The term "machine-readable signal" refers to any signal for providing machine instructions and / or data to a programmable processor.
[0133] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0134] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with embodiments of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.
[0135] Computer systems can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. Client-server relationships are created by computer programs running on the respective computers and having a client-server relationship with each other.
Claims
1. A method for mitigating network failure losses, characterized in that, include: Receive a first input, which is first data obtained by the network probe point. The first data includes, but is not limited to, HTTP status data, HTTP request time data, and domain name resolution data. In response to the first input, first abnormal information is obtained based on the first data according to the location and the operator. The first abnormal information includes abnormal information on domain name availability, abnormal information on average response time, and abnormal information on hijacking. Receive a second input, which includes second data uploaded by the client's embedded points; In response to the second input, if the first abnormal information meets preset abnormal conditions, the second data is used to verify the first abnormal information that meets the preset abnormal conditions. If the verification determines that the network is in a fault state, a network fault handling plan is executed, including: Set separate thresholds for abnormal information regarding domain name availability, average response time, and hijacking for different domain name locations and different ISPs. If the abnormal information of domain name availability, average response time, or hijacking exceeds the corresponding threshold for abnormal information of domain name availability, average response time, and hijacking, the first abnormal information is determined to meet the preset abnormal conditions. According to the first abnormal information that meets the preset abnormal conditions, obtain the target second data of the same location, same operator and same sequence, and use the target second data to verify the first abnormal information; If the verification result meets the preset verification conditions, the network is determined to be in a fault state, and the network fault handling plan is executed.
2. The network failure loss mitigation method according to claim 1, characterized in that, The response to the first input, obtaining first abnormal information based on the first data according to the location and carrier, includes: HTTP status data, HTTP request time data, and domain name resolution data are aggregated according to the location and the operator, respectively, to obtain domain name availability information, average response time information, and hijacking information for different locations and operators. The number of abnormal information in the domain name availability information, average response time information, and hijacking information for different origins and different operators is counted separately to obtain abnormal information in domain name availability, average response time, and hijacking situation for different origins and different operators.
3. The network failure loss mitigation method according to claim 1, characterized in that, The network failure handling plan includes: Cut off the IP traffic corresponding to the first abnormal information.
4. A network fault mitigation device, characterized in that, include: The first receiving module is used to receive a first input, which is first data obtained by the network probe point. The first data includes, but is not limited to, HTTP status data, HTTP request time data, and domain name resolution data. The first acquisition module, in response to the first input, acquires first abnormal information based on the first data according to the location and the operator, the first abnormal information including abnormal information on domain name availability, abnormal information on average response time, and abnormal information on hijacking. The second receiving module is used to receive the second input, which includes the second data uploaded by the client's embedded points; The fault handling module, in response to the second input, when the first abnormal information meets the preset abnormal conditions, uses the second data to verify the first abnormal information that meets the preset abnormal conditions. When the verification determines that the network is in a fault state, it executes the network fault handling plan. The fault handling module includes: The threshold setting module is used to set the abnormal information thresholds for domain name availability, average response time, and hijacking for different regions and different operators. The determination module is used to determine that the first abnormal information meets the preset abnormal conditions when the abnormal information of domain name availability, the abnormal information of average response time, or the abnormal information of hijacking exceeds the corresponding abnormal information thresholds of domain name availability, average response time, and hijacking. The verification module is used to obtain target second data of the same location, same operator and same sequence according to the first abnormal information that meets the preset abnormal conditions, and to verify the first abnormal information using the target second data. The execution module is used to determine that the network is in a fault state when the verification result meets the preset verification conditions, and to execute the network fault handling plan.
5. The network fault prevention device according to claim 4, characterized in that, The first acquisition module includes: The aggregation module is used to aggregate HTTP status data, HTTP request time data, and domain name resolution data according to the location and carrier, respectively, to obtain domain name availability information, average response time information, and hijacking information for different locations and carriers. The statistics module is used to count the number of abnormal information in domain name availability information, average response time information, and hijacking information for different home locations and different home operators, and to obtain abnormal information in domain name availability, average response time, and hijacking situations for different home locations and different home operators.
6. The network fault loss mitigation device according to claim 4, characterized in that, The fault handling module includes: The traffic control module is used to cut off the IP traffic corresponding to the first abnormal information.
7. An electronic device, comprising: processor; as well as Stored program memory, The program includes instructions that, when executed by the processor, cause the processor to perform the method according to any one of claims 1-3.
8. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to perform the method according to any one of claims 1-3.