Network operation and maintenance methods, devices, and software products based on end-to-end detection
By deploying monitoring nodes in various regions and using time-series algorithms and Markov Monte Carlo algorithms, network faults are proactively identified and handled. This solves the problem of low operation and maintenance efficiency caused by relying on customer complaints in existing technologies, and enables rapid and accurate fault location and handling, thereby improving the stability of network services and user experience.
Patent Information
- Application Number
- CN202510079366.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-17
- Publication Date
- 2025-11-14
- Estimated Expiration
- 2045-01-17
AI Technical Summary
Current network operations and maintenance rely on customer complaints, resulting in low efficiency, inability to detect and handle network faults in a timely manner, and potential risks.
By deploying monitoring nodes in various regions, network information is collected and time-series algorithms are used for anomaly detection. By combining Markov Monte Carlo algorithms and tomography algorithms, abnormal nodes and faulty network segments are identified, enabling proactive fault location and handling.
It improved fault response speed, reduced fault handling time, enhanced the efficiency and accuracy of network operation and maintenance, and ensured network stability and quality.
Smart Images

Figure CN119906629B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of distributed systems, and more specifically, to a network operation and maintenance method, apparatus, and program product based on end-to-end detection. Background Technology
[0002] End-to-end monitoring is a commonly used technique and method in network operations and maintenance. It is primarily used to comprehensively understand and evaluate the performance, availability, and reliability of a network. End-to-end monitoring means performing comprehensive monitoring and measurement from one end of the network (e.g., the client) to the other end (e.g., the server), including but not limited to monitoring network devices, ISPs, CDNs, DNS, etc., in between.
[0003] Existing technologies mainly rely on passive monitoring, primarily depending on customer complaints. Operational troubleshooting and emergency handling are then carried out based on the limited information provided by customers. This passive response model is extremely inefficient. Customer complaints are usually only obtained after a large number of users have experienced problems. Moreover, the issues reported by customers are generally unrelated to operation and maintenance work, and are only related to user experience. This makes it difficult for operation and maintenance personnel to locate problems, prolongs emergency handling time, and causes unnecessary risks to the network.
[0004] There is currently no effective solution to the problem that network operation and maintenance in related technologies relies on customer-reported complaints, and then conducts operation and maintenance investigations and emergency handling based on customer feedback, which results in low operation and maintenance efficiency and potential network risks. Summary of the Invention
[0005] The main purpose of this application is to provide a network operation and maintenance method, device and program product based on end-to-end detection, so as to solve the problem that network operation and maintenance relies on customer-reported complaints and then conducts operation and maintenance investigation and emergency handling based on customer feedback, which results in low operation and maintenance efficiency and potential network risks.
[0006] To achieve the above objectives, according to one aspect of this application, a network operation and maintenance method based on end-to-end detection is provided. The method includes: collecting network information of N nodes in various regions, wherein the network information of the N nodes includes at least: network information of a first node and network information of a second node, the first node representing a node sending a request, the second node representing a node receiving a request, and N being a positive integer; performing anomaly detection on the network information of the N nodes using a preset timing algorithm to obtain abnormal nodes among the N nodes; acquiring network segment information of the N nodes and determining the network segment information of the abnormal nodes based on the network segment information of the N nodes; determining fault information of the N nodes based on the network information and the network segment information of the N nodes, and processing the abnormal nodes based on the fault information of the N nodes.
[0007] Further, determining the fault information of the N nodes based on the network information and network segment information of the N nodes includes: calculating the number of successful requests and the number of failed requests for each request path based on the network information of the N nodes, wherein the request path represents the routing path between the first node and the second node; calculating the fault probability of each of the N nodes by using a tomographic scanning algorithm to calculate the network information of the N nodes; calculating the forwarding success rate of each network segment based on the network segment information of the N nodes, the number of successful requests and the number of failed requests for each request path using a Markov Monte Carlo algorithm; and determining the fault information of the N nodes based on the number of successful requests, the number of failed requests, the forwarding success rate of each network segment, and the fault probability of each node.
[0008] Further, the abnormal node is processed based on the fault information of the N nodes, including: locating the faulty network segment based on the forwarding success rate of each network segment, wherein the faulty network segment includes the faulty node; determining the network segment type of the faulty network segment based on the fault information of the N nodes; and sending the forwarding success rate of each network segment, the faulty network segment, the network segment type of the faulty network segment, and the network information of the abnormal node to the target device, wherein the abnormal node and the faulty node are processed by the target device.
[0009] Further, determining the network segment information of the abnormal node based on the network segment information of the N nodes includes: obtaining the routing information of the N nodes based on the network information of the N nodes using a preset tool, wherein the above includes at least: routing information between the first node and the second node; determining the network segment information of each request path based on the routing information and the network segment information of the N nodes; and locating the network segment information of the abnormal node based on the network segment information of each request path and the network information corresponding to the abnormal node.
[0010] Further, collecting network information of N nodes in each region includes: sending a network request from the first node among the N nodes in each region to the second node among the N nodes, and receiving response data returned by the second node; determining the original data of the N nodes based on the response data returned by the second node, wherein the network information includes at least: request path, IP information of the first node, IP information of the second node, and indicator data; and aggregating the original data of the N nodes to obtain the network information of the N nodes.
[0011] Further, the raw data of the N nodes are aggregated to obtain the network information of the N nodes, including: classifying the raw data of the N nodes according to the request time of the network request to obtain a first classification result; classifying the raw data of the N nodes according to the region to which the node sending the network request belongs to obtain a second classification result; classifying the raw data of the N nodes according to the indicator information of the network request to obtain a third classification result; and constructing the network information of the N nodes based on the first classification result, the second classification result, and the third classification result.
[0012] Further, obtaining the network segment information of the N nodes includes: parsing the return code in the response data returned by the second node; extracting normal data and abnormal data from the response data returned by the second node according to the return code, and storing the normal data and the abnormal data; performing data preprocessing on the normal data and the abnormal data respectively to obtain processed data; and using a preset query algorithm to query the network segment information of the first node and the network segment information of the second node in the processed data respectively to obtain the network segment information of the N nodes.
[0013] To achieve the above objectives, according to another aspect of this application, a network operation and maintenance device based on end-to-end detection is provided. The device includes: a collection unit for collecting network information of N nodes in various regions, wherein the network information of the N nodes includes at least: network information of a first node and network information of a second node, the first node representing a node sending a request, the second node representing a node receiving a request, and N being a positive integer; a detection unit for performing anomaly detection on the network information of the N nodes using a preset timing algorithm to obtain abnormal nodes among the N nodes; an acquisition unit for acquiring network segment information of the N nodes and determining the network segment information of the abnormal nodes based on the network segment information of the N nodes; and a processing unit for determining fault information of the N nodes based on the network information and network segment information of the N nodes, and processing the abnormal nodes based on the fault information of the N nodes.
[0014] Further, the processing unit includes: a first calculation subunit, configured to calculate the number of successful requests and the number of failed requests for each request path based on the network information of the N nodes, wherein the request path represents the routing path between the first node and the second node; a second calculation subunit, configured to calculate the network information of the N nodes using a tomographic scanning algorithm to determine the failure probability of each of the N nodes; a third calculation subunit, configured to calculate the forwarding success rate of each network segment based on the network segment information of the N nodes, the number of successful requests for each request path, and the number of failed requests for each request path using a Markov Monte Carlo algorithm; and a first determination subunit, configured to determine the failure information of the N nodes based on the number of successful requests for each request path, the number of failed requests for each request path, the forwarding success rate of each network segment, and the failure probability of each node.
[0015] Further, the processing unit includes: a first positioning subunit, configured to locate faulty network segments based on the forwarding success rate of each network segment, wherein the faulty network segment includes faulty nodes; a second determining subunit, configured to determine the network segment type of the faulty network segment based on the fault information of the N nodes; and a sending subunit, configured to send the forwarding success rate of each network segment, the faulty network segment, the network segment type of the faulty network segment, and the network information of the abnormal node to a target device, wherein the abnormal node and the faulty node are processed by the target device.
[0016] Further, the acquisition unit includes: an acquisition subunit, configured to acquire routing information of the N nodes based on the network information of the N nodes using a preset tool, wherein the acquisition information includes at least: routing information between the first node and the second node; a third determination subunit, configured to determine the network segment information of each request path based on the routing information of the N nodes and the network segment information of the N nodes; and a second positioning subunit, configured to locate the network segment information of the abnormal node based on the network segment information of each request path and the network information corresponding to the abnormal node.
[0017] Further, the acquisition unit includes: a receiving subunit, configured to send a network request from the first node among the N nodes in each region to the second node among the N nodes, and receive response data returned by the second node; a fourth determining subunit, configured to determine the original data of the N nodes based on the response data returned by the second node, wherein the network information includes at least: request path, IP information of the first node, IP information of the second node, and indicator data; and an aggregation subunit, configured to aggregate the original data of the N nodes to obtain the network information of the N nodes.
[0018] Further, the aggregation subunit includes: a first classification module, used to classify the raw data of the N nodes according to the request time of the network request, to obtain a first classification result; a second classification module, used to classify the raw data of the N nodes according to the region to which the node sending the network request belongs, to obtain a second classification result; a third classification module, used to classify the raw data of the N nodes according to the indicator information of the network request, to obtain a third classification result; and a construction module, used to construct the network information of the N nodes based on the first classification result, the second classification result, and the third classification result.
[0019] Further, the acquisition unit includes: a parsing subunit, used to parse the return code in the response data returned by the second node; an extraction subunit, used to extract normal data and abnormal data from the response data returned by the second node according to the return code, and store the normal data and the abnormal data; a processing subunit, used to perform data preprocessing on the normal data and the abnormal data respectively to obtain processed data; and a query subunit, used to query the network segment information of the first node and the network segment information of the second node respectively in the processed data using a preset query algorithm to obtain the network segment information of the N nodes.
[0020] To achieve the above objectives, according to one aspect of this application, a computer program product is provided, comprising a computer program that, when executed by a processor, implements the network operation and maintenance method based on end-to-end detection as described above, and when executed by a processor, implements the steps of the network operation and maintenance method based on end-to-end detection as described in various embodiments of this application.
[0021] To achieve the above objectives, according to one aspect of this application, a computer-readable storage medium is provided, the computer-readable storage medium including stored computer instructions, wherein, when the computer instructions are executed by a processor, the network operation and maintenance method based on end-to-end detection described above is implemented.
[0022] To achieve the above objectives, according to one aspect of this application, an electronic device is provided, including one or more processors and a memory, the memory being used to store one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors cause the one or more processors to implement the network operation and maintenance method based on end-to-end detection as described above.
[0023] In this embodiment, network information of N nodes in various regions is collected, including at least the network information of a first node and the network information of a second node. The first node represents the node sending the request, and the second node represents the node receiving the request, where N is a positive integer. A preset timing algorithm is used to detect anomalies in the network information of the N nodes, identifying abnormal nodes. The network segment information of the N nodes is obtained, and the network segment information of the abnormal nodes is determined based on this information. Fault information of the N nodes is determined based on their network information and network segment information, and the abnormal nodes are processed accordingly. This solves the technical problem of network maintenance relying on customer complaints and subsequent troubleshooting and emergency handling, which results in low maintenance efficiency and potential network risks.
[0024] By deploying monitoring nodes (first nodes) in prefecture-level cities across the country and proactively sending network requests to target nodes (second nodes), and employing pre-defined time-series algorithms for anomaly detection, abnormal nodes can be quickly identified when a fault first occurs or performance begins to deteriorate. Compared to relying on customer complaints or post-event analysis, this significantly improves fault response speed, effectively shortens fault handling time, and enhances service quality. Furthermore, by combining anomaly detection results to determine the network segment information of abnormal nodes, the specific location of network faults can be more accurately pinpointed, whether in access networks, backbone networks, or data center internal networks. This significantly improves the efficiency and accuracy of fault location, reduces the workload of maintenance personnel, and avoids potential blind spots and inefficiencies in fault investigation, thereby improving the efficiency and accuracy of network operations and maintenance, and further enhancing the stability and quality of network services. Attached Figure Description
[0025] The accompanying drawings, which form part of this application, are used to provide a further understanding of this application. The illustrative embodiments and descriptions of this application are used to explain this application and do not constitute an undue limitation of this application. In the drawings:
[0026] Figure 1 This is a hardware structure block diagram of a computer terminal (or mobile device) for implementing a network operation and maintenance method based on end-to-end detection, according to Embodiment 1 of this application.
[0027] Figure 2 This is a flowchart of an optional network operation and maintenance method based on end-to-end detection provided according to Embodiment 1 of this application;
[0028] Figure 3 This is a schematic diagram of an optional network operation and maintenance method based on end-to-end detection provided in Embodiment 1 of this application. Figure 1 ;
[0029] Figure 4 This is a schematic diagram of an optional network operation and maintenance method based on end-to-end detection provided in Embodiment 1 of this application. Figure 1 ;
[0030] Figure 5 This is a schematic diagram of an optional network operation and maintenance method based on end-to-end detection provided in Embodiment 1 of this application. Figure 1 ;
[0031] Figure 6 This is a schematic diagram of an optional network operation and maintenance method based on end-to-end detection provided in Embodiment 1 of this application. Figure 1 ;
[0032] Figure 7 This is a schematic diagram of a network operation and maintenance device based on end-to-end detection provided in Embodiment 2 of this application;
[0033] Figure 8 This is a schematic diagram of a network operation and maintenance electronic device based on end-to-end detection, provided according to Embodiment 5 of this application. Detailed Implementation
[0034] It should be noted that, unless otherwise specified, the embodiments and features described in this application can be combined with each other. This application will now be described in detail with reference to the accompanying drawings and embodiments.
[0035] It should be noted that the processing method, apparatus, storage medium, and electronic device specified in this application can be used in the network operation and maintenance process in the financial technology field to improve operation and maintenance efficiency, and can also be used in any field other than the financial technology field. The application field of the processing method, apparatus, storage medium, and electronic device specified in this application is not limited.
[0036] It should be noted that the information collected in this application (including but not limited to user device information, user personal information, collected data, used data, generated data, processed data, etc.) and the data (including but not limited to data used for analysis, stored data, displayed data, collected information, used information, generated information, processed information, etc.) are all information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, storage, use, processing, transmission, provision, disclosure, and application of related data all comply with the relevant laws, regulations, and standards of the relevant countries and regions, and necessary confidentiality measures have been taken. These measures do not violate public order and good morals, and corresponding operation entry points are provided for users to choose to authorize or refuse. For example, this system has interfaces with relevant users or organizations, providing users with corresponding operation entry points for users to choose to agree to or refuse automated decision results; if the user chooses to refuse, the process proceeds to the expert decision-making stage.
[0037] Example 1
[0038] According to an embodiment of this application, a method embodiment for network operation and maintenance based on end-to-end detection is also provided. It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Although a logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than that shown here.
[0039] The method embodiment provided in Embodiment 1 of this application can be executed on a mobile terminal, computer terminal, or similar computing device. Figure 1 A hardware block diagram of a computer terminal (or mobile device) for implementing a network operation and maintenance method based on end-to-end detection is shown. Figure 1As shown, the computer terminal 10 (or mobile device) may include one or more processors 102 (shown as 102a, 102n in the figure, ..., 102n), including but not limited to microprocessors (MCUs) or programmable logic devices (FPGAs), a memory 104 for storing data, and a transmission device 106 for communication functions. In addition, it may include: a display, an input / output interface (I / O interface), a universal serial bus (US) port (which may be included as one port of the US bus), a network interface, a power supply, and / or a camera. Those skilled in the art will understand that... Figure 1 The structure shown is for illustrative purposes only and does not limit the structure of the aforementioned electronic device. For example, computer terminal 10 may also include... Figure 1 The more or fewer components shown, or having the same Figure 1 The different configurations shown.
[0040] It should be noted that the aforementioned one or more processors 102 and / or other data processing circuits are generally referred to herein as "data processing circuits". These data processing circuits may be embodied, in whole or in part, in software, hardware, firmware, or any other combination thereof. Furthermore, the data processing circuits may be a single, independent processing module, or may be integrated, in whole or in part, into any other element within the computer terminal 10 (or mobile device). As involved in the embodiments of this application, the data processing circuits serve as a processor control mechanism (e.g., selection of a variable resistor termination path connected to an interface).
[0041] The memory 104 can be used to store software programs and modules of application software, such as the program instruction / data storage device corresponding to the network operation and maintenance method based on end-to-end detection in this embodiment of the application. The processor 102 executes various functional applications and data processing by running the software programs and modules stored in the memory 104, thereby realizing the aforementioned network operation and maintenance method based on end-to-end detection. The memory 104 may include high-speed random access memory, and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 104 may further include memory remotely located relative to the processor 102, and these remote memories can be connected to the computer terminal 10 via a network. Examples of the aforementioned networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.
[0042] The transmission device 106 is used to receive or send data via a network. Specific examples of the network described above may include a wireless network provided by the communication provider of the computer terminal 10. In one example, the transmission device 106 includes a Network Interface Controller (NIC), which can connect to other network devices via a base station to communicate with the Internet. In another example, the transmission device 106 may be a Radio Frequency (RF) module, used for wireless communication with the Internet.
[0043] The display can be, for example, a touchscreen liquid crystal display (LCD) that allows the user to interact with the user interface of the computer terminal 10 (or mobile device).
[0044] Under the aforementioned operating environment, this application provides the following: Figure 2 The network operation and maintenance method based on end-to-end detection is shown. Figure 2 This is a flowchart of a network operation and maintenance method based on end-to-end detection according to Embodiment 1 of this application.
[0045] Step S101: Collect network information of N nodes in each region. The network information of the N nodes includes at least the network information of the first node and the network information of the second node. The first node represents the node that sends the request and the second node represents the node that receives the request. N is a positive integer.
[0046] In this first embodiment, multiple monitoring nodes can be deployed in various regions, and the network performance and status can be evaluated by collecting network information from N nodes in each region. Each region can include different areas such as cities at the prefecture level or above across the country, and each region has a monitoring node. N is a positive integer, and its value can be adjusted according to monitoring needs and resource availability. For example, if more detailed data is needed, the value of N can be increased, that is, the city can be divided into smaller regions, and monitoring nodes can be deployed in each region, thereby collecting more detailed data through the monitoring nodes.
[0047] The first node typically refers to the node that sends the request, i.e., the monitoring node; the second node refers to the node that receives the request, usually the target server or a specific destination in the network. The information collected includes, but is not limited to, performance metrics such as network latency, packet loss rate, and response time from the monitoring node (first node) to the target point (second node). This information is crucial for analyzing network quality and locating faults.
[0048] By deploying multiple monitoring nodes in various regions and collecting network information between these nodes and the target node, end-to-end network monitoring can be achieved. This enables timely detection of network faults, performance bottlenecks, and other issues, as well as fault location and cause analysis, thereby improving the efficiency and response speed of network operation and maintenance.
[0049] Step S102: Use a preset time-series algorithm to perform anomaly detection on the network information of N nodes to obtain the abnormal nodes among the N nodes.
[0050] In this first embodiment, when monitoring network performance, the network information collected from N nodes is typically in the form of a time series. For example, performance metrics such as the response time from the monitoring node to the target IP, DNS resolution time, and SSL handshake time change over time. The preset time series algorithm refers to statistical or machine learning algorithms suitable for time series data, such as the 3Sigma algorithm, moving average method, autoregressive moving average model (ARIMA), and long short-term memory network (LSTM). These algorithms can identify trends, periodicity, and abnormal fluctuations in the data.
[0051] By performing anomaly detection on the network information of N nodes using a time-series algorithm, monitoring nodes exhibiting abnormal characteristics can be identified, i.e., the aforementioned anomalous nodes. The network information of anomalous nodes may indicate a fault or performance degradation in a certain part of the network, and is the key object for subsequent fault localization and cause analysis.
[0052] By following the steps above, network problems can be identified in a timely and proactive manner, which greatly improves the efficiency and proactivity of network operation and maintenance compared to the passive response model that relies on user complaints.
[0053] Step S103: Obtain the network segment information of N nodes, and determine the network segment information of abnormal nodes based on the network segment information of N nodes.
[0054] In this first embodiment, after identifying the abnormal node, it is necessary to locate the network segment information of the abnormal node, that is, the network range and operator to which the abnormal node belongs. This allows the scope of fault location to be narrowed down based on the network segment information, helping maintenance personnel to quickly pinpoint the network environment in which the problem may occur and improving maintenance efficiency.
[0055] Network segment information refers to the address range or identifier of the network where each monitoring node is located. By obtaining the network segment information of the monitoring nodes, the network environment in which the nodes are located can be determined, such as ISP (Internet Service Provider), geographical location, and other network environment information.
[0056] By combining anomaly detection in network performance with node network segment information, fault location can be made more accurate and faster. For example, if multiple monitoring nodes from the same network segment report abnormally high latency or frequent packet loss, operations and maintenance personnel can determine that the problem may lie in the network equipment or operator services of this network segment, thereby avoiding ineffective troubleshooting of other normal network segments and saving valuable troubleshooting time and resources.
[0057] By acquiring the network segment information of the monitoring nodes and determining the network segment of the abnormal nodes based on this information, the cause of the fault can be quickly determined based on the abnormal nodes and their network segment information, thus significantly improving the efficiency and accuracy of network operation and maintenance.
[0058] Step S104: Determine the fault information of N nodes based on the network information and network segment information of N nodes, and process the abnormal nodes based on the fault information of N nodes.
[0059] In this first embodiment, based on the aforementioned network information and network segment information, specific fault information can be further analyzed and identified. This includes determining the nature of the fault (e.g., whether it is a packet loss problem or a latency problem), the severity of the fault, and the possible root causes of the fault (e.g., whether it is a problem with the monitoring node's network access, the server node's network, or the intermediate forwarding network). The determination of fault information is based on in-depth analysis of abnormal network information and comprehensive consideration of network segment information, making fault location more accurate.
[0060] After identifying the fault information, maintenance personnel can take appropriate measures. For example, if the problem lies with the monitoring node's network access, it may be necessary to coordinate with the local ISP; if it's a problem with the intermediate forwarding network, it may be necessary to adjust the packet routing path; if it's a problem with the target server's network, it may be necessary to troubleshoot and repair the fault on the server side. The formulation of these measures depends on the accurate understanding and location of the fault information.
[0061] By collecting and analyzing network and network segment information from monitoring nodes, network faults can be proactively detected and located without waiting for user complaints or alarms, thus significantly reducing fault response time and improving network operation and maintenance efficiency and user experience. Furthermore, processing fault information allows for more targeted problem-solving, avoiding blind operations and improving the success rate and speed of fault resolution.
[0062] Optionally, in the network operation and maintenance method based on end-to-end detection provided in Embodiment 1 of this application, the fault information of N nodes is determined based on the network information of N nodes and the network segment information of N nodes, including: calculating the number of successful requests and the number of failed requests for each request path based on the network information of N nodes, wherein the request path represents the routing path between the first node and the second node; calculating the fault probability of each of the N nodes by using a tomographic scanning algorithm to calculate the network information of the N nodes; calculating the forwarding success rate of each network segment based on the network segment information of the N nodes, the number of successful requests and the number of failed requests for each request path using a Markov Monte Carlo algorithm; and determining the fault information of the N nodes based on the number of successful requests, the number of failed requests, the forwarding success rate of each network segment, and the fault probability of each node.
[0063] In this first embodiment, network information can be monitored to calculate whether each node's request to the target node was successful, and the number of successes or failures. Each request path refers to the routing path from the monitoring node (first node) to the target server node (second node). For example, if monitoring node A sends 100 requests to server B, with 90 successes and 10 failures, then for this request path, the number of successes is 90 and the number of failures is 10.
[0064] Then, based on the collected network information, especially the number of successes and failures for each path, the failure probability of each monitored node is calculated using a tomographic scanning algorithm. The tomographic scanning algorithm can analyze the path and status of data packets transmitted in the network, identify nodes or paths with failure rates higher than a preset value, and thus determine the failure probability of the nodes.
[0065] Secondly, based on the Markov Monte Carlo algorithm, the forwarding success rate of each network segment is calculated by considering the network segment information of N nodes, the number of successful requests for each path, and the number of failed requests for each path. The Markov Chain Monte Carlo (MCMC) algorithm is a method used to simulate state transitions and estimate probability distributions in complex systems, and is used to estimate the forwarding success rate of each network segment. The forwarding success rate reflects the probability of a data packet being successfully transmitted when passing through a network segment, and is crucial for assessing the health status of the network segment and locating faults.
[0066] Finally, by comprehensively analyzing the above information, the fault information of each monitoring node is determined, including the specific type of fault (such as packet loss, latency, connectivity issues), the severity of the fault, and the possible causes of the fault. For example, if the number of failed requests from a certain monitoring node to multiple target servers far exceeds the number of successful requests, and the forwarding success rate of the network segment where the monitoring node is located is low, this may indicate a problem with the network access to the monitoring node, requiring further investigation.
[0067] By using statistical analysis of monitoring data, fault location using the tomography algorithm, and network segment forwarding success rate calculation using the Markov Monte Carlo algorithm, network faults can be proactively and promptly detected and located, reducing fault response time and improving the efficiency and accuracy of network operation and maintenance. This overcomes the limitations of traditional passive monitoring methods, improves the operation and maintenance efficiency of financial institutions, and ultimately enhances the business stability of financial institutions.
[0068] Optionally, in the network operation and maintenance method based on end-to-end detection provided in Embodiment 1 of this application, abnormal nodes are processed according to the fault information of N nodes, including: locating faulty network segments with faults based on the forwarding success rate of each network segment, wherein the faulty network segment includes faulty nodes with faults; determining the network segment type of the faulty network segment based on the fault information of N nodes; and sending the forwarding success rate of each network segment, the faulty network segment, the network segment type of the faulty network segment, and the network information of the abnormal nodes to the target device, wherein the abnormal nodes and faulty nodes are processed by the target device.
[0069] In this first embodiment, if the forwarding success rate of a certain network segment is significantly lower than the average value or a set threshold, then the network segment is identified as a "faulty network segment". A faulty network segment may contain one or more "faulty nodes", i.e., network devices or connection points, whose abnormalities cause a decrease in the forwarding success rate. Locating the faulty network segment can help maintenance personnel quickly narrow down the scope of troubleshooting and improve the efficiency of fault location.
[0070] By analyzing the fault information of N nodes, the "network segment type" of the faulty network segment can be determined, which is the specific network environment in which the fault occurred, such as an access network (usually a connection from a client to an ISP), a backbone network (a connection between ISPs), or a data center network (an internal network on the server side). Determining the network segment type helps operations and maintenance personnel understand the background of the fault and take more appropriate fault handling strategies.
[0071] Finally, the forwarding success rate of each network segment, the faulty network segment, the network type of the faulty network segment, and the network information of the abnormal nodes are sent to the target device. The target device can be understood here as a centralized monitoring platform or a fault handling system used by maintenance personnel. Once the target device receives this fault information, it can initiate the corresponding fault handling process. This may include direct repair of abnormal nodes (e.g., restarting network devices, adjusting configurations, etc.), rerouting traffic from the faulty network segment (avoiding the fault point), and communicating with network operators to resolve backbone network issues. The processing is usually automated, based on preset fault handling strategies and workflows to achieve fast and efficient fault recovery.
[0072] Through data analysis and algorithm application, the system enables fault detection at monitoring nodes, fault location in network segments, and fault type determination. Finally, this information is transmitted to the target device for automated or semi-automated fault handling, significantly improving the efficiency of network fault detection and handling. This, in turn, reduces network outage time and improves the service quality of financial institutions.
[0073] Optionally, in the network operation and maintenance method based on end-to-end detection provided in Embodiment 1 of this application, determining the network segment information of an abnormal node based on the network segment information of N nodes includes: obtaining the routing information of N nodes based on the network information of N nodes using a preset tool, wherein at least the routing information between the first node and the second node is included; determining the network segment information of each request path based on the routing information and the network segment information of N nodes; and locating the network segment information of the abnormal node based on the network segment information of each request path and the network information corresponding to the abnormal node.
[0074] In this first embodiment, "preset tool" typically refers to network diagnostic tools, such as route tracing. Route tracing helps us understand the specific path data packets take in the network. By using the preset tool, we can obtain the routing information between the monitoring node (first node) and the target server (second node) based on the network information of these monitoring nodes. The routing information records in detail every hop of the data packet from the monitoring node (first node) to the target server, including the IP address of each intermediate router it passes through.
[0075] After obtaining routing information from N monitoring nodes, the network segment information is combined with the network segment information to determine the network segment information for each request path. Network segment information refers to the network range or ISP (Internet Service Provider) to which the monitoring nodes and routers along the path belong. By matching the IP addresses of each router in each request path with the network segment information, the specific network segment traversed during data packet transmission can be determined.
[0076] By utilizing the network segment information of each request path and the network performance information of abnormal nodes, the network environment in which these abnormal nodes are located can be more accurately identified. For example, if multiple monitoring nodes report abnormally high packet loss rates when communicating with the same target server, and the request paths of these nodes pass through the same network segment, then it can be determined that the network segment information of the abnormal nodes is likely that network segment, thereby inferring that there may be a network failure or performance bottleneck in that network segment.
[0077] By collecting network information, acquiring routing information, and combining it with network segment information for analysis, we have achieved proactive monitoring and precise location of network faults. At the same time, by leveraging the distributed advantages of monitoring nodes and the powerful capabilities of network diagnostic tools, we can help maintenance personnel discover and handle network faults before customers perceive the problems, greatly improving the stability of network services and user experience.
[0078] Optionally, in the network operation and maintenance method based on end-to-end detection provided in Embodiment 1 of this application, the network information of N nodes in each region is collected, including: sending a network request to a second node among the N nodes through a first node in each region, and receiving response data returned by the second node; determining the original data of the N nodes based on the response data returned by the second node, wherein the network information includes at least: request path, IP information of the first node, IP information of the second node, and indicator data; and aggregating the original data of the N nodes to obtain the network information of the N nodes.
[0079] In this first embodiment, the "first node" refers to the monitoring nodes distributed across various regions, while the "second node" typically refers to a key node in the target server or network. The monitoring nodes periodically or as needed send network requests to the target node, such as HTTP requests, PINGs, and test data, to test network connectivity and performance. Simultaneously, the monitoring nodes receive response data from the target node, including response status codes and response times, for subsequent network performance analysis.
[0080] Raw data refers to the detailed information received by the monitoring nodes regarding network requests and responses, including but not limited to request paths, the IP information of the monitoring node (first node), the IP information of the target node (second node), and various network performance metrics such as response time, packet loss rate, and connectivity. This raw data forms the basis of network monitoring, providing direct evidence of the network status when each monitoring node communicates with the target node.
[0081] After the raw data is collected, data aggregation is required to extract key network information, resulting in the network information for the aforementioned N nodes. This typically includes statistical analysis of the raw data by time window (e.g., every 5 minutes) and spatial range (e.g., by province), calculating aggregation metrics such as average response time and average packet loss rate, and determining the success rate and number of failures in communication between the monitored nodes and the target points. The aggregated data can more intuitively reflect the overall network performance and trends, helping operations and maintenance personnel quickly identify potential problems or faults in the network.
[0082] Through the above process, the health status of the network can be proactively monitored and evaluated, network performance degradation or faults can be detected and located in a timely manner, and corresponding measures can be taken for repair and optimization, thereby improving the efficiency of fault detection and thus enhancing network stability and user experience.
[0083] Optionally, in the network operation and maintenance method based on end-to-end detection provided in Embodiment 1 of this application, the raw data of N nodes are aggregated to obtain network information of N nodes, including: classifying the raw data of N nodes according to the request time of the network request to obtain a first classification result; classifying the raw data of N nodes according to the region to which the node that sent the network request belongs to obtain a second classification result; classifying the raw data of N nodes according to the indicator information of the network request to obtain a third classification result; and constructing network information of N nodes based on the first classification result, the second classification result, and the third classification result.
[0084] In this first embodiment, the raw data can be classified according to time sequence, typically to observe the trend of network performance changes over time. For example, the data can be grouped according to time granularities of 5 minutes, 1 hour, or 1 day to determine the network's performance at different times, such as performance differences between peak and off-peak periods. The first classification result helps identify periodic problems or sudden failures in the network and provides a basis for subsequent time series analysis.
[0085] Then, the data is categorized by geographical location, grouping it according to the different regions (such as provinces or cities) where the monitoring nodes are located. This allows for analysis of network conditions within specific regions, such as identifying areas with higher network latency than others. The second classification helps in understanding the geographical distribution characteristics of the network and identifying regional network problems.
[0086] Secondly, the raw data can be categorized according to metrics, which refer to quantitative indicators of network performance, including response time, packet loss rate, and connectivity status. This categorization method provides a clearer understanding of the network's performance under different performance metrics; for example, identifying which nodes have packet loss rates that reach or exceed a certain threshold. Thirdly, the categorization results help pinpoint performance bottlenecks or fault points, providing a basis for network optimization.
[0087] Finally, these classification results (i.e., the first, second, and third classification results) are combined to construct comprehensive network information for each monitoring node. This includes, for example, performance trends over time, performance overviews by region, and statistical analysis results for specific performance indicators.
[0088] This multi-dimensional classification and aggregation method can transform massive amounts of raw monitoring data into structured and meaningful network information, providing strong data support for network health monitoring, fault location, and performance optimization, thereby improving data processing efficiency and information availability.
[0089] Optionally, in the network operation and maintenance method based on end-to-end detection provided in Embodiment 1 of this application, obtaining the network segment information of N nodes includes: parsing the return code in the response data returned by the second node; extracting normal data and abnormal data from the response data returned by the second node according to the return code, and storing the normal data and abnormal data; performing data preprocessing on the normal data and abnormal data respectively to obtain processed data; and using a preset query algorithm to query the network segment information of the first node and the second node respectively in the processed data to obtain the network segment information of N nodes.
[0090] In this first embodiment, to obtain the network segment information of N nodes, the return codes in the response data returned by the second node can be parsed. Return codes represent the response status of a network request, typically represented by HTTP status codes in the HTTP protocol, such as 200 indicating a successful request, 404 indicating a resource not found, and the 500 series indicating a server-side error. By parsing the return codes, it can be determined whether the network request was successfully executed, thereby identifying whether the network is operating normally or in an abnormal state. Based on the return codes, the response data can be divided into two parts: normal data and abnormal data. Normal data refers to data where the return code indicates a successful request or a normal response, while abnormal data refers to data where the return code indicates a failed request or an abnormal response.
[0091] Then, data preprocessing refers to cleaning, transforming, or formatting the raw data before data analysis to remove noise, fill in missing values, and standardize data formats. For both normal and outlier data, the purpose of preprocessing is to ensure that subsequent algorithms can accurately analyze and utilize this data. Processed data will be cleaner and more uniformly formatted, facilitating subsequent analysis.
[0092] Secondly, a preset query algorithm is used to query the network segment information of the first node and the second node in the processed data, respectively, to obtain the network segment information of N nodes. The preset query algorithm is usually based on a network query tool (in this embodiment, it refers to a tool used to obtain information such as domain names and IP addresses) to perform a database query. The database contains the allocation information of the IP addresses of the detection nodes in each region and the corresponding network segment affiliation data. By querying the IP addresses of the first node (monitoring node) and the second node (target node) in the processed data, the network segment information of each node can be obtained.
[0093] This process effectively processes and categorizes network request response data, extracts key network segment information, and enables proactive discovery and rapid location of network faults. This, in turn, improves fault response speed, reduces customer complaints, and enhances network operation and maintenance efficiency and network service quality.
[0094] Optionally, in this first embodiment, the end-to-end network fault monitoring and location process of this solution can be as follows: Figure 3 As shown, network data is actively collected through monitoring nodes distributed across prefecture-level cities nationwide to obtain comprehensive network status information. Network query tools are used to query IP addresses and determine their network segments, providing crucial information for constructing the network topology. The raw network data is aggregated and categorized by time, space, and network metrics for easier subsequent in-depth analysis. The network topology is constructed, and a network link diagram is drawn based on monitored router information and network segment affiliation. Temporal anomaly detection (such as the 3sigma algorithm) is used to identify outliers in network performance metrics, which may indicate network failures. Tomography-based scanning technology, combined with temporal anomaly information and network topology, is used to pinpoint the specific location of the network failure. Finally, based on fault location and network metric analysis, the root cause of the failure is inferred, which may be network equipment failure, link congestion, or software problems. Figure 3 The process integrates data collection, analysis, fault location, and cause inference, enabling proactive monitoring and rapid response to network faults, and significantly improving the efficiency and accuracy of network operation and maintenance.
[0095] Optionally, in this first embodiment, the initial data processing flow for network fault monitoring in this solution can be as follows: Figure 4As shown. First, in step S401, monitoring nodes distributed across prefecture-level cities nationwide actively initiate network requests based on a preset URL list, collecting comprehensive network data including monitoring node IPs, target server IPs, response times, etc. Next, in step S402, the response return codes are parsed, and normal and abnormal response data are distinguished and marked, stored separately in the database, providing a clear dataset for subsequent analysis. Finally, in step S403, caching acceleration technology is used to quickly obtain network segment information of monitoring nodes and routers along the transmission path. This step helps to construct the network topology and understand the network segment affiliation in the data transmission path, providing a network structure reference for fault location.
[0096] Optionally, in this first embodiment, the process of analyzing monitoring data and constructing network topology can be as follows: Figure 5 As shown in step S501, the collected network information is aggregated. This process not only integrates network data from multiple monitoring nodes but also categorizes the data according to predefined time and spatial granularities (such as 5-minute intervals and provinces), facilitating subsequent detailed analysis. In step S502, the average network metrics, such as latency, bandwidth, and success rate, are calculated under specific time and spatial categories. This helps identify normal and abnormal network performance patterns. In step S503, traceroute technology is used to obtain the IP addresses of each router on the path from the monitoring node to the target server, providing crucial information for understanding the specific path of data packet transmission. Finally, in stage S504, a network query tool is used to obtain the network segment affiliation information of the aforementioned router IP addresses. Combining this with the data from S501 and S503, a detailed network topology map is constructed. This topology map not only reveals the complete path from the monitoring node to the server but also marks the affiliation of each network segment along the path, serving as the foundation for network fault location and performance optimization. Figure 5 The process ensures the full integration and utilization of data, providing a clear network view and data support for subsequent fault diagnosis and network operation and maintenance.
[0097] Optionally, in this first embodiment, the network fault monitoring and location process of this solution can be as follows: Figure 6 As shown. Starting from S601, the 3sigma algorithm is used to detect anomalies in the aggregated time-series data. This algorithm can identify measurement values that exceed the normal fluctuation range; these anomalies usually indicate potential problems or faults in the network. In S602, the monitored end-to-end data is used as the basis for analysis, and the number of successful and failed measurements for each path is counted to provide a quantitative basis for assessing the network health status. Moving to S603, based on... Figure 5The system uses the network topology information constructed in step S604 to determine the specific network segments traversed by each path, providing a concrete network structure reference for fault location. In S604, the Markov Monte Carlo method is used to estimate the forwarding success rate of each network segment by combining the number of successful and failed path measurements. This is used to measure network quality; for example, network segments with low forwarding success rates are likely the source of the fault. In S605, based on the analysis of the network segment forwarding success rates, the location of the faulty network segment, i.e., the fault point, is precisely located, enabling maintenance personnel to quickly focus on the problem and conduct targeted troubleshooting and repair. Finally, in S606, all analysis results, including outliers, forwarding success rates, and fault point locations, are uploaded to the centralized monitoring platform and relevant personnel are notified via email, ensuring timely transmission and response to fault information and improving the efficiency and speed of fault handling.
[0098] It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although a logical order is shown in the flowchart, in some cases the steps shown or described may be executed in a different order than that shown here.
[0099] In summary, the network operation and maintenance method based on end-to-end detection provided in this application collects network information from N nodes in various regions. The network information of the N nodes includes at least: network information of a first node and network information of a second node. The first node represents the node sending the request, and the second node represents the node receiving the request, where N is a positive integer. A preset timing algorithm is used to detect anomalies in the network information of the N nodes, identifying abnormal nodes. The network segment information of the N nodes is obtained, and the network segment information of the abnormal nodes is determined based on this information. The fault information of the N nodes is determined based on the network information and network segment information, and the abnormal nodes are processed accordingly. This method solves the problem in related technologies where network operation and maintenance relies on customer complaints and then relies on customer feedback for troubleshooting and emergency handling, resulting in low operation and maintenance efficiency and potential network risks.
[0100] By deploying monitoring nodes (first nodes) in prefecture-level cities across the country and proactively sending network requests to target nodes (second nodes), and employing pre-defined time-series algorithms for anomaly detection, abnormal nodes can be quickly identified when a fault first occurs or performance begins to deteriorate. Compared to relying on customer complaints or post-event analysis, this significantly improves fault response speed, effectively shortens fault handling time, and enhances service quality. Furthermore, by combining anomaly detection results to determine the network segment information of abnormal nodes, the specific location of network faults can be more accurately pinpointed, whether in access networks, backbone networks, or data center internal networks. This significantly improves the efficiency and accuracy of fault location, reduces the workload of maintenance personnel, and avoids potential blind spots and inefficiencies in fault investigation, thereby improving the efficiency and accuracy of network operations and maintenance, and further enhancing the stability and quality of network services.
[0101] Example 2
[0102] This application also provides a network operation and maintenance device based on end-to-end detection. It should be noted that this end-to-end detection-based network operation and maintenance device can be used to execute the network operation and maintenance method based on end-to-end detection provided in this application. The following describes the end-to-end detection-based network operation and maintenance device provided in this application.
[0103] According to embodiments of this application, an apparatus for implementing the above-described network operation and maintenance method based on end-to-end detection is also provided, such as... Figure 7 As shown, the device includes:
[0104] Specifically, the acquisition unit 701 is used to acquire network information of N nodes in each region. The network information of the N nodes includes at least the network information of the first node and the network information of the second node. The first node represents the node that sends the request, and the second node represents the node that receives the request. N is a positive integer.
[0105] The detection unit 702 is used to perform anomaly detection on the network information of N nodes using a preset timing algorithm to obtain the abnormal nodes among the N nodes.
[0106] The acquisition unit 703 is used to acquire the network segment information of N nodes and determine the network segment information of abnormal nodes based on the network segment information of N nodes.
[0107] The processing unit 704 is used to determine the fault information of N nodes based on the network information and network segment information of N nodes, and to process the abnormal nodes based on the fault information of N nodes.
[0108] The network operation and maintenance device based on end-to-end detection provided in this application embodiment collects network information of N nodes in each area through a collection unit 701. The network information of the N nodes includes at least: network information of a first node and network information of a second node. The first node represents the node that sends the request, and the second node represents the node that receives the request. N is a positive integer. The detection unit 702 uses a preset timing algorithm to perform anomaly detection on the network information of the N nodes to obtain abnormal nodes among the N nodes. The acquisition unit 703 acquires the network segment information of the N nodes and determines the network segment information of the abnormal nodes based on the network segment information of the N nodes. The processing unit 704 determines the fault information of the N nodes based on the network information and network segment information of the N nodes, and processes the abnormal nodes based on the fault information of the N nodes. This solves the problem in related technologies where network operation and maintenance relies on customer-reported complaints and then conducts operation and maintenance investigations and emergency handling based on customer feedback, resulting in low operation and maintenance efficiency and potential network risks.
[0109] By deploying monitoring nodes (first nodes) in prefecture-level cities across the country and proactively sending network requests to target nodes (second nodes), and employing pre-defined time-series algorithms for anomaly detection, abnormal nodes can be quickly identified when a fault first occurs or performance begins to deteriorate. Compared to relying on customer complaints or post-event analysis, this significantly improves fault response speed, effectively shortens fault handling time, and enhances service quality. Furthermore, by combining anomaly detection results to determine the network segment information of abnormal nodes, the specific location of network faults can be more accurately pinpointed, whether in access networks, backbone networks, or data center internal networks. This significantly improves the efficiency and accuracy of fault location, reduces the workload of maintenance personnel, and avoids potential blind spots and inefficiencies in fault investigation, thereby improving the efficiency and accuracy of network operations and maintenance, and further enhancing the stability and quality of network services.
[0110] Optionally, in the network operation and maintenance device based on end-to-end detection provided in Embodiment 2 of this application, the processing unit 704 includes: a first calculation subunit, used to calculate the number of successful requests and the number of failed requests for each request path based on the network information of N nodes, wherein the request path represents the routing path between the first node and the second node; a second calculation subunit, used to calculate the network information of N nodes using a tomographic scanning algorithm to determine the failure probability of each of the N nodes; a third calculation subunit, used to calculate the forwarding success rate of each network segment based on the network segment information of the N nodes, the number of successful requests and the number of failed requests for each request path using a Markov Monte Carlo algorithm; and a first determination subunit, used to determine the failure information of the N nodes based on the number of successful requests, the number of failed requests, the forwarding success rate of each network segment, and the failure probability of each node.
[0111] Optionally, in the network operation and maintenance device based on end-to-end detection provided in Embodiment 2 of this application, the processing unit 704 includes: a first positioning subunit, used to locate faulty network segments with faults based on the forwarding success rate of each network segment, wherein the faulty network segment includes faulty nodes with faults; a second determining subunit, used to determine the network segment type of the faulty network segment based on the fault information of N nodes; and a sending subunit, used to send the forwarding success rate of each network segment, the faulty network segment, the network segment type of the faulty network segment, and the network information of abnormal nodes to the target device, wherein the abnormal nodes and faulty nodes are processed by the target device.
[0112] Optionally, in the network operation and maintenance device based on end-to-end detection provided in Embodiment 2 of this application, the acquisition unit 703 includes: an acquisition subunit, used to acquire routing information of N nodes based on the network information of N nodes using a preset tool, wherein at least: routing information between the first node and the second node; a third determination subunit, used to determine the network segment information of each request path based on the routing information of N nodes and the network segment information of N nodes; and a second positioning subunit, used to locate the network segment information of the abnormal node based on the network segment information of each request path and the network information corresponding to the abnormal node.
[0113] Optionally, in the network operation and maintenance device based on end-to-end detection provided in Embodiment 2 of this application, the above-mentioned acquisition unit 701 includes: a receiving subunit, used to send a network request to a second node among N nodes through a first node among N nodes in each region, and receive response data returned by the second node; a fourth determining subunit, used to determine the original data of N nodes based on the response data returned by the second node, wherein the network information includes at least: request path, IP information of the first node, IP information of the second node, and indicator data; and an aggregation subunit, used to aggregate the original data of N nodes to obtain the network information of N nodes.
[0114] Optionally, in the network operation and maintenance device based on end-to-end detection provided in Embodiment 2 of this application, the above-mentioned aggregation subunit includes: a first classification module, used to classify the original data of N nodes according to the request time of the network request to obtain a first classification result; a second classification module, used to classify the original data of N nodes according to the region to which the node sending the network request belongs to obtain a second classification result; a third classification module, used to classify the original data of N nodes according to the indicator information of the network request to obtain a third classification result; and a construction module, used to construct the network information of N nodes based on the first classification result, the second classification result, and the third classification result.
[0115] Optionally, in the network operation and maintenance device based on end-to-end detection provided in Embodiment 2 of this application, the acquisition unit 703 includes: a parsing subunit, used to parse the return code in the response data returned by the second node; an extraction subunit, used to extract normal data and abnormal data from the response data returned by the second node according to the return code, and store the normal data and abnormal data; a processing subunit, used to perform data preprocessing on the normal data and abnormal data respectively to obtain processed data; and a query subunit, used to query the network segment information of the first node and the second node respectively in the processed data using a preset query algorithm to obtain the network segment information of N nodes.
[0116] It should be noted that the acquisition unit 701, detection unit 702, acquisition unit 703, and processing unit 704 mentioned above correspond to steps S201 to S204 in Embodiment 1. The two modules and the corresponding steps implement the same instances and application scenarios, but are not limited to the content disclosed in Embodiment 1. It should be noted that the above modules or units can be hardware or software components stored in memory (e.g., memory 104) and processed by one or more processors (e.g., processors 102a, 102 based on end-to-end detection network operation and maintenance, ..., 102n). The above modules can also be part of the device and run in the computer terminal 10 provided in Embodiment 1.
[0117] Example 3
[0118] Embodiments of this application may provide an electronic device. Figure 8 This is a structural block diagram of an electronic device according to an embodiment of this application. Figure 8 As shown, the electronic device may include: one or more ( Figure 8 Only one of the components is shown: processor 802, memory 804, memory controller, and peripheral interface, wherein the peripheral interface is connected to the radio frequency module, audio module, and display.
[0119] The memory can be used to store software programs and modules, such as the program instructions / modules corresponding to the methods and apparatus in the embodiments of this application. The processor executes various functional applications and data processing by running the software programs and modules stored in the memory, thereby implementing the above-described methods. The memory may include high-speed random access memory, and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory may further include memory remotely located relative to the processor, and these remote memories can be connected to the terminal via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.
[0120] The processor can access information and applications stored in memory via a transmission device to execute the following steps: Collect network information from N nodes in each region, where the network information of the N nodes includes at least: network information of a first node and network information of a second node, where the first node represents the node sending the request and the second node represents the node receiving the request, and N is a positive integer; perform anomaly detection on the network information of the N nodes using a preset timing algorithm to identify abnormal nodes among the N nodes; obtain the network segment information of the N nodes and determine the network segment information of the abnormal nodes based on the network segment information of the N nodes; determine the fault information of the N nodes based on the network information and network segment information of the N nodes, and process the abnormal nodes based on the fault information of the N nodes.
[0121] The processor can access information and applications stored in memory via a transmission device to execute the following steps: determining the fault information of N nodes based on the network information and network segment information of N nodes, including: calculating the number of successful requests and the number of failed requests for each request path based on the network information of N nodes, where the request path represents the routing path between the first node and the second node; calculating the fault probability of each of the N nodes by using a tomographic scanning algorithm to analyze the network information of the N nodes; calculating the forwarding success rate of each network segment based on the network segment information of the N nodes, the number of successful requests and the number of failed requests for each request path using a Markov Monte Carlo algorithm; and determining the fault information of the N nodes based on the number of successful requests, the number of failed requests, the forwarding success rate of each network segment, and the fault probability of each node.
[0122] The processor can access the information and application programs stored in the memory via the transmission device to execute the following steps: processing abnormal nodes based on the fault information of N nodes, including: locating faulty network segments based on the forwarding success rate of each network segment, wherein the faulty network segment includes faulty nodes; determining the network segment type of the faulty network segment based on the fault information of N nodes; and sending the forwarding success rate of each network segment, the faulty network segment, the network segment type of the faulty network segment, and the network information of the abnormal nodes to the target device, wherein the target device processes the abnormal nodes and faulty nodes.
[0123] The processor can invoke information and applications stored in the memory through the transmission device to perform the following steps: determining the network segment information of the abnormal node based on the network segment information of N nodes, including: obtaining the routing information of N nodes based on the network information of N nodes using a preset tool, wherein at least the routing information between the first node and the second node is included; determining the network segment information of each request path based on the routing information and the network segment information of N nodes; and locating the network segment information of the abnormal node based on the network segment information of each request path and the network information corresponding to the abnormal node.
[0124] The processor can access information and applications stored in memory via a transmission device to perform the following steps: collecting network information from N nodes in each region, including: sending a network request from the first node among the N nodes in each region to the second node among the N nodes, and receiving the response data returned by the second node; determining the original data of the N nodes based on the response data returned by the second node, wherein the network information includes at least: request path, IP information of the first node, IP information of the second node, and indicator data; and aggregating the original data of the N nodes to obtain the network information of the N nodes.
[0125] The processor can access information and applications stored in memory via a transmission device to perform the following steps: aggregating the raw data of N nodes to obtain network information of N nodes, including: classifying the raw data of N nodes according to the request time of the network request to obtain a first classification result; classifying the raw data of N nodes according to the region to which the node sending the network request belongs to obtain a second classification result; classifying the raw data of N nodes according to the indicator information of the network request to obtain a third classification result; and constructing the network information of N nodes based on the first classification result, the second classification result, and the third classification result.
[0126] The processor can access the information and application programs stored in the memory via the transmission device to perform the following steps: obtaining network segment information of N nodes, including: parsing the return code in the response data returned by the second node; extracting normal data and abnormal data from the response data returned by the second node according to the return code, and storing the normal data and abnormal data; performing data preprocessing on the normal data and abnormal data respectively to obtain processed data; and using a preset query algorithm to query the network segment information of the first node and the second node respectively in the processed data to obtain the network segment information of N nodes.
[0127] This application provides a network operation and maintenance method based on end-to-end detection. It collects network information from N nodes in various regions, where the network information of the N nodes includes at least: network information of a first node and network information of a second node, where the first node represents the node sending the request and the second node represents the node receiving the request, and N is a positive integer. A preset timing algorithm is used to detect anomalies in the network information of the N nodes, identifying abnormal nodes. The network segment information of the N nodes is obtained, and the network segment information of the abnormal nodes is determined based on this information. Fault information of the N nodes is determined based on the network information and network segment information, and the abnormal nodes are processed accordingly. This method solves the technical problem of network operation and maintenance relying on customer complaints and subsequent troubleshooting and emergency handling, which results in low efficiency and potential network risks.
[0128] By deploying monitoring nodes (first nodes) in prefecture-level cities across the country and proactively sending network requests to target nodes (second nodes), and employing pre-defined time-series algorithms for anomaly detection, abnormal nodes can be quickly identified when a fault first occurs or performance begins to deteriorate. Compared to relying on customer complaints or post-event analysis, this significantly improves fault response speed, effectively shortens fault handling time, and enhances service quality. Furthermore, by combining anomaly detection results to determine the network segment information of abnormal nodes, the specific location of network faults can be more accurately pinpointed, whether in access networks, backbone networks, or data center internal networks. This significantly improves the efficiency and accuracy of fault location, reduces the workload of maintenance personnel, and avoids potential blind spots and inefficiencies in fault investigation, thereby improving the efficiency and accuracy of network operations and maintenance, and further enhancing the stability and quality of network services.
[0129] Those skilled in the art will understand that Figure 8 The structure shown is for illustrative purposes only. Electronic devices can also be smartphones (such as Android phones, iOS phones, etc.), tablets, PDAs, mobile Internet devices (MIDs), PADs, and other terminal devices. Figure 8 This does not limit the structure of the aforementioned electronic device. For example, electronic devices may also include components that are more... Figure 8 The more or fewer components shown (such as network interfaces, display devices, etc.), or having the same Figure 8 The different configurations shown.
[0130] Those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be implemented by a program instructing the hardware related to the terminal device. The program can be stored in a computer-readable storage medium, which may include: flash drive, read-only memory (ROM), random access memory (RAM), disk or optical disk, etc.
[0131] Example 4
[0132] Embodiments of this application also provide a storage medium. Optionally, in this embodiment, the storage medium can be used to store the program code executed by the network operation and maintenance method based on end-to-end detection provided in Embodiment 1.
[0133] Optionally, in this embodiment, the storage medium may be located in any computer terminal in a group of computer terminals in a computer network, or in any mobile terminal in a group of mobile terminals.
[0134] This application also provides a computer program product that, when executed on a data processing device, is suitable for performing network operation and maintenance method steps based on end-to-end detection.
[0135] The sequence numbers of the embodiments in this application are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.
[0136] In the above embodiments of this application, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments.
[0137] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. The device embodiments described above are merely illustrative; for example, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the displayed or discussed mutual coupling, direct coupling, or communication connection may be through some interfaces; the indirect coupling or communication connection between units or modules may be electrical or other forms.
[0138] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0139] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0140] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as a USB flash drive, read-only memory (ROM), random access memory (RAM), portable hard drive, magnetic disk, or optical disk.
[0141] The above description is only a preferred embodiment of this application. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of this application, and these improvements and modifications should also be considered within the scope of protection of this application.
Claims
1. A network operation and maintenance method based on end-to-end detection, characterized in that, include: Collect network information of N nodes in each region, wherein the network information of the N nodes includes at least: network information of a first node and network information of a second node, the first node represents the node that sends the request, the second node represents the node that receives the request, and N is a positive integer; An anomaly detection is performed on the network information of the N nodes using a preset time-series algorithm to identify the abnormal nodes among the N nodes; Obtain the network segment information of the N nodes, and determine the network segment information of the abnormal node based on the network segment information of the N nodes; Based on the network information and network segment information of the N nodes, the fault information of the N nodes is determined, and the abnormal nodes are processed based on the fault information of the N nodes. The fault information of the N nodes is determined based on the network information and network segment information of the N nodes, including: Based on the network information of the N nodes, calculate the number of successes and the number of failures for each request path, where the request path represents the routing path between the first node and the second node; The network information of the N nodes is calculated using a tomography algorithm to determine the failure probability of each of the N nodes. The forwarding success rate of each network segment is calculated based on the network segment information of the N nodes, the number of successful requests for each request path, and the number of failed requests for each request path, according to the Markov Monte Carlo algorithm. The fault information of the N nodes is determined based on the number of successful requests for each request path, the number of failed requests for each request path, the forwarding success rate of each network segment, and the failure probability of each node.
2. The method according to claim 1, characterized in that, The abnormal nodes are processed based on the fault information of the N nodes, including: Based on the forwarding success rate of each network segment, faulty network segments are located, wherein the faulty network segments include faulty nodes. The network segment type of the faulty network segment is determined based on the fault information of the N nodes; The forwarding success rate of each network segment, the faulty network segment, the network segment type of the faulty network segment, and the network information of the abnormal node are sent to the target device, wherein the abnormal node and the faulty node are processed by the target device.
3. The method according to claim 1, characterized in that, Determining the network segment information of the abnormal node based on the network segment information of the N nodes includes: The routing information of the N nodes is obtained by using a preset tool based on the network information of the N nodes, wherein the routing information includes at least the routing information between the first node and the second node; The network segment information of each requested path is determined based on the routing information and network segment information of the N nodes; The network segment information of the abnormal node is located based on the network segment information of each request path and the network information corresponding to the abnormal node.
4. The method according to claim 1, characterized in that, Collect network information from N nodes in each region, including: The network request is sent from the first node among the N nodes in each region to the second node among the N nodes, and the response data returned by the second node is received. The original data of the N nodes are determined based on the response data returned by the second node, wherein the network information includes at least: request path, IP information of the first node, IP information of the second node, and indicator data; The original data of the N nodes are aggregated to obtain the network information of the N nodes.
5. The method according to claim 4, characterized in that, The original data of the N nodes are aggregated to obtain the network information of the N nodes, including: The raw data of the N nodes are classified according to the request time of the network request to obtain the first classification result; The original data of the N nodes are classified according to the region to which the node that sent the network request belongs, to obtain a second classification result; The raw data of the N nodes are classified according to the indicator information of the network request to obtain a third classification result; The network information of the N nodes is constructed based on the first classification result, the second classification result, and the third classification result.
6. The method according to claim 1, characterized in that, Obtaining the network segment information of the N nodes includes: Parse the return code in the response data returned by the second node; Based on the return code, normal data and abnormal data are extracted from the response data returned by the second node, and the normal data and abnormal data are stored. The normal data and the abnormal data are preprocessed separately to obtain the processed data; A preset query algorithm is used to query the network segment information of the first node and the network segment information of the second node in the processed data to obtain the network segment information of the N nodes.
7. A network operation and maintenance device based on end-to-end detection, characterized in that, include: The acquisition unit is used to acquire network information of N nodes in each area, wherein the network information of the N nodes includes at least: network information of a first node and network information of a second node, the first node represents the node that sends the request, the second node represents the node that receives the request, and N is a positive integer; The detection unit is used to perform anomaly detection on the network information of the N nodes using a preset time-series algorithm to obtain the abnormal nodes among the N nodes; The acquisition unit is used to acquire the network segment information of the N nodes and determine the network segment information of the abnormal node based on the network segment information of the N nodes. The processing unit is used to determine the fault information of the N nodes based on the network information and network segment information of the N nodes, and to process the abnormal nodes based on the fault information of the N nodes. The processing unit includes: a first calculation subunit, used to calculate the number of successful requests and the number of failed requests for each request path based on the network information of the N nodes, wherein the request path represents the routing path between the first node and the second node; a second calculation subunit, used to calculate the network information of the N nodes using a tomographic scanning algorithm to determine the failure probability of each of the N nodes; a third calculation subunit, used to calculate the forwarding success rate of each network segment based on the network segment information of the N nodes, the number of successful requests for each request path, and the number of failed requests for each request path using a Markov Monte Carlo algorithm; and a first determination subunit, used to determine the failure information of the N nodes based on the number of successful requests for each request path, the number of failed requests for each request path, the forwarding success rate of each network segment, and the failure probability of each node.
8. An electronic device, characterized in that, include: Memory, which stores executable programs; A processor for running the program, wherein the program, when running, performs the method according to any one of claims 1 to 6.
9. A computer program product comprising computer instructions, characterized in that, When the computer instructions are executed by the processor, they implement the steps of the method described in any one of claims 1 to 6.
Citation Information
Patent Citations
Network state monitoring method and device, electronic equipment and readable storage medium
CN113079059A
Link anomaly positioning method and device, computer equipment and storage medium
CN115987771A