A method, system, device, and medium for handling network packet loss in an available zone on the cloud
By deploying monitoring nodes and using a weighted fusion algorithm to evaluate packet loss metrics, the method addresses the challenge of network packet loss in cloud environments, enhancing fault detection and handling efficiency and accuracy.
Patent Information
- Application Number
- CN202411305414.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-19
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2044-09-19
AI Technical Summary
The prior art is difficult to efficiently and accurately monitor and handle network packet loss problems in a cloud computing environment, resulting in service interruption and data loss. The traditional methods are inefficient and error-prone, making it difficult to meet the needs of high reliability and high availability.
Deploy detection nodes in the cloud Availability Zone network, process multi-dimensional network packet loss index data through weighted fusion algorithms, monitor and evaluate network status in real time, and use personalized notifications and fault briefing generation rules to achieve accurate fault location and processing.
It improves the accuracy and comprehensiveness of network packet loss monitoring in cloud environments, improves fault processing efficiency, ensures business stability and data security, and adapts to network environments of different scales and complexities.
Smart Images

Figure CN119211079B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of network communication technologies, and in particular, to a method, system, device, and medium for handling network packet loss in a cloud available zone. Background Art
[0002] With the rapid development of cloud computing technologies, more and more enterprises choose to migrate their services to the cloud to achieve flexible resource configuration and efficient utilization. However, while enjoying the convenience brought by cloud computing, the stability and reliability issues of the cloud network are becoming increasingly prominent, and the network packet loss problem is particularly outstanding. Network packet loss refers to the situation where, during network transmission, due to various reasons such as network congestion, transmission errors, and device failures, some data packets fail to be successfully transmitted to the destination, thereby affecting the integrity and accuracy of data transmission.
[0003] The network packet loss problem not only reduces the efficiency of data transmission but may also lead to service interruptions and data loss, causing immeasurable losses to enterprises. In a cloud environment, due to the complex network structure, numerous devices, and wide distribution, it becomes more difficult to discover and locate the network packet loss problem. Traditional network monitoring and fault troubleshooting methods often rely on manual intervention, with low efficiency and high error rates, and are difficult to meet the requirements of high reliability and high availability for cloud services.
[0004] Therefore, how to monitor the network status of cloud available zones in real time and accurately, and promptly discover and handle network packet loss problems has become a technical problem urgently to be solved in the current cloud computing field. Currently, although there are some network monitoring and fault troubleshooting tools on the market, most of them have problems such as single monitoring indicators, low determination accuracy, and insufficient automation, and are difficult to meet the actual needs of cloud network management. Summary of the Invention
[0005] The purpose of the present invention is to provide a method, system, device, and medium for handling network packet loss in a cloud available zone, which effectively improves the accuracy and comprehensiveness of packet loss monitoring for complex cloud networks, realizes accurate fault location and personalized notifications, and significantly improves the efficiency of network fault handling, so as to solve at least one of the above-mentioned existing technical problems.
[0006] In a first aspect, the present invention provides a method for handling network packet loss in a cloud available zone, and the method specifically includes:
[0007] Deploy a number of detection nodes in each cloud available zone network, and obtain multi-dimensional network packet loss index data for each cloud available zone network according to the detection nodes;
[0008] Using a weighted fusion algorithm, process the multi-dimensional network packet loss metric data of all the probe nodes in each available zone network on the cloud to obtain the comprehensive network packet loss score for each available zone network on the cloud;
[0009] Determine the network packet loss score threshold for each available zone network on the cloud according to the network area division, service type, and time period configuration parameters. By comparing the comprehensive network packet loss score and the network packet loss score threshold of each available zone network on the cloud, determine whether a network packet loss failure event has occurred in each available zone network on the cloud;
[0010] Determine the packet loss failure event data of the available zone network on the cloud where the network packet loss failure event has occurred, and determine the packet loss event level according to the packet loss failure event data;
[0011] Based on the packet loss event level and the attribute information of each operation and maintenance personnel, determine different network fault briefing generation rules and different network fault briefing push rules for each operation and maintenance personnel. Generate corresponding network fault briefings for each operation and maintenance personnel according to different network fault briefing generation rules and different network fault briefing push rules and push them.
[0012] In a second aspect, the present invention provides a system for processing network packet loss in an available zone network on the cloud. The system specifically includes:
[0013] A first processing module, used to deploy a number of probe nodes in each available zone network on the cloud, and obtain the multi-dimensional network packet loss metric data of each available zone network on the cloud according to the probe nodes;
[0014] A second processing module, used to process the multi-dimensional network packet loss metric data of all the probe nodes in each available zone network on the cloud by using a weighted fusion algorithm to obtain the comprehensive network packet loss score for each available zone network on the cloud;
[0015] A third processing module, used to determine the network packet loss score threshold for each available zone network on the cloud according to the network area division, service type, and time period configuration parameters. By comparing the comprehensive network packet loss score and the network packet loss score threshold of each available zone network on the cloud, determine whether a network packet loss failure event has occurred in each available zone network on the cloud;
[0016] A fourth processing module, used to determine the packet loss failure event data of the available zone network on the cloud where the network packet loss failure event has occurred, and determine the packet loss event level according to the packet loss failure event data;
[0017] A fifth processing module, configured to determine different network fault briefing generation rules and different network fault briefing push rules for each operation and maintenance personnel based on the packet loss event level and the attribute information of each operation and maintenance personnel, and generate and push corresponding network fault briefings for each operation and maintenance personnel according to the different network fault briefing generation rules and different network fault briefing push rules.
[0018] In a third aspect, the present invention provides a computer device, including: a memory, a processor, and a computer program stored on the memory. When the computer program is executed on the processor, it implements the method for processing network packet loss in an available zone on the cloud as described in any one of the above methods.
[0019] In a fourth aspect, the present invention provides a computer-readable storage medium, on which a computer program is stored. When the computer program is run by a processor, it implements the method for processing network packet loss in an available zone on the cloud as described in any one of the above methods.
[0020] Compared with the prior art, the present invention has at least one of the following technical effects:
[0021] 1. Through real-time monitoring, accurate determination, flexible configuration, and automated processing, the present invention effectively solves the problem of network packet loss in the available zone on the cloud, ensures service continuity and data security, and reduces the losses suffered by enterprises due to network packet loss.
[0022] 2. The present invention can intelligently detect network packet loss faults in a multi-cloud environment, and provide personalized network fault briefings for different operation and maintenance personnel according to the severity and impact range of the faults, improving the timeliness and effectiveness of network fault handling and ensuring the stable operation of services.
[0023] 3. The present invention effectively improves the accuracy and comprehensiveness of packet loss monitoring for complex networks in the cloud environment, realizes accurate fault location and personalized notifications, and significantly improves the efficiency of network fault handling.
[0024] 4. By deploying detection nodes in each available zone network on the cloud, the present invention can obtain multi-dimensional network packet loss index data in real time, realize real-time monitoring of the network status on the cloud, and timely discover packet loss problems.
[0025] 5. The present invention uses a weighted fusion algorithm to process the multi-dimensional network packet loss index data of detection nodes, obtains a comprehensive network packet loss score, and determines a packet loss score threshold in combination with configuration parameters such as network area division, service type, and time period, so as to accurately judge whether packet loss occurs and its severity.
[0026] 6. The present invention supports customizing network packet loss scenarios and setting packet loss thresholds, adapts to network environments of different scales and complexities, supports multiple cloud platforms, and improves the flexibility and applicability of the method.
[0027] 7. The present invention automatically generates and pushes personalized network fault briefings according to the packet loss event level and the attribute information of operation and maintenance personnel, realizes the automatic determination and processing of network packet loss problems, and improves the efficiency of network operation. BRIEF DESCRIPTION OF THE DRAWINGS
[0028] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for use in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0029] Figure 1 is a schematic flowchart of a method for processing network packet loss in an available zone on the cloud provided by the first embodiment of the present invention;
[0030] Figure 2 is a schematic flowchart of a method for processing network packet loss in an available zone on the cloud provided by the second embodiment of the present invention;
[0031] Figure 3 is a schematic structural diagram of a system for processing network packet loss in an available zone on the cloud provided by an embodiment of the present invention;
[0032] Figure 4 is a schematic structural diagram of a computer device provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0033] In the following description, specific details such as specific system structures and technologies are presented for the purpose of illustration rather than limitation, so as to thoroughly understand the embodiments of the present application. However, those skilled in the art should clearly understand that the present application can also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid unnecessary details from interfering with the description of the present application.
[0034] It should be understood that when used in the specification of the present application and the appended claims, the term "comprising" indicates the presence of the described features, wholes, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components, and / or their combinations.
[0035] It should also be understood that the term "and / or" as used in the specification of the present application and the appended claims refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations.
[0036] As used in the specification of this application and the appended claims, the term "if" may be construed as "when", "once", "in response to determining", or "in response to detecting" depending on the context. Similarly, the phrases "if determined" or "if [the described condition or event] is detected" may be construed as meaning "once determined", "in response to determining", "once [the described condition or event] is detected", or "in response to detecting [the described condition or event]" depending on the context.
[0037] In addition, in the description of the specification of this application and the appended claims, the terms "first", "second", "third", etc. are only used for distinguishing descriptions and should not be construed as indicating or implying relative importance.
[0038] The reference to "one embodiment" or "some embodiments" etc. described in the specification of this application means that a specific feature, structure, or characteristic described in connection with the embodiment is included in one or more embodiments of this application. Thus, the statements "in one embodiment", "in some embodiments", "in other some embodiments", "in still other embodiments", etc. that appear in different places in this specification do not necessarily all refer to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized in other ways. The terms "comprising", "including", "having", and their variants all mean "including but not limited to", unless otherwise specifically emphasized in other ways.
[0039] In the embodiments of this application, the execution subject of the process includes a terminal device. The terminal device includes but is not limited to: devices such as servers, computers, smart phones, and tablet computers that can execute the methods disclosed in this application. Figure 1 The flowchart showing a method for processing network packet loss in a cloud available zone network disclosed in the first embodiment of the present invention is described in detail as follows:
[0040] S101, deploy a number of detection nodes in each cloud available zone network respectively, and obtain multi-dimensional network packet loss index data of each cloud available zone network according to the detection nodes.
[0041] In this embodiment, in order to monitor the multi-dimensional network packet loss index data of each cloud available zone network, a network monitoring system based on detection nodes is designed and deployed. This system deploys a number of detection nodes in each available zone of the cloud, and collects and analyzes key performance indicators such as network packet loss, latency, and jitter in real time, providing data support for network optimization and fault troubleshooting.
[0042] Specifically, select appropriate servers or virtual machines as probe nodes according to the network architecture and scale of the cloud availability zones. These nodes should have stable network connections and sufficient computing resources. Deploy probe nodes near the key network nodes in each cloud availability zone to ensure comprehensive coverage of the network conditions in that zone. Install network monitoring software, such as tools like Ping, Traceroute, iperf, etc., and customized network monitoring scripts on the probe nodes to send and receive network data packets and record relevant metric data.
[0043] Configure the probe nodes to regularly send ICMP protocol packets, TCP / UDP data packets, etc. to the target servers and their intermediate routing nodes, and collect multi-dimensional network performance metric data including packet loss rate, data latency, data jitter, available bandwidth estimation, routing hop count, etc. Use cloud network analysis technology to perform real-time analysis on the collected network data to identify abnormal conditions such as network congestion, latency, and packet loss. Through visualization tools, display the analysis results in the form of charts, trend graphs, etc. to help operation and maintenance personnel intuitively understand the network conditions. According to the analysis results, quickly locate network fault points, such as network device failures, line failures, etc. Based on the network performance bottlenecks, propose optimization suggestions, such as adjusting the network topology, increasing bandwidth, optimizing application programs, etc.
[0044] In this embodiment, by real-time monitoring and analyzing key metrics such as network packet loss, network problems are discovered and solved in a timely manner, improving the stability and reliability of the network. According to the monitoring data, perform refined management and optimization on the network, such as adjusting network bandwidth, optimizing routing strategies, etc., to improve the overall performance of the network.
[0045] In some embodiments, before the step of obtaining the network performance data of each cloud availability zone network according to the probe nodes in the above step S101, it further includes:
[0046] Obtain the topology structure information and path dynamic change information of each cloud availability zone network, and determine the initial probe path and probe frequency parameters according to the topology structure information and the path dynamic change information;
[0047] Adopt the ARIMA time series analysis algorithm to perform modeling analysis on the historical probe data of each cloud availability zone network to obtain the packet loss probability distributions under different paths and different frequencies;
[0048] According to the initial probe path, the probe frequency parameters, and the packet loss probability distributions under different paths and different frequencies, with the goal of minimizing the overall packet loss monitoring blind area, construct an integer programming probe strategy optimization model, and the integer programming probe strategy optimization model is used to determine the optimal probe strategy for each cloud availability zone network.
[0049] In this embodiment, the network topology structure information of the available zone on the cloud is obtained, including network nodes, links, and their connection relationships, etc., to construct a network topology graph model. The dynamic change information of each path in the available zone network on the cloud is monitored, including the changes in performance indicators such as path delay, packet loss rate, and bandwidth, and the path dynamic change data is recorded. According to the network topology structure information and the path dynamic change information, graph theory algorithms and machine learning algorithms are used to automatically plan and generate initial network detection paths to cover key paths and nodes. By analyzing the statistical characteristics of the path dynamic change data, clustering algorithms are used to group the paths to obtain stable path groups and unstable path groups. For different path groups, the detection frequency parameters are adaptively set, with a low detection frequency for the stable path group and a high detection frequency for the unstable path group, to dynamically optimize the detection efficiency.
[0050] The network historical detection data of the available zone on the cloud is obtained, and the data under different network paths and detection frequencies is classified and sorted. The ARIMA time series analysis algorithm is used to model and analyze the classified historical detection data. According to the results of the ARIMA algorithm modeling and analysis, the packet loss probability distribution under different network paths and detection frequencies is obtained. If the packet loss probability under a certain network path or detection frequency exceeds the preset threshold, it is determined that there is a network anomaly risk for this path or frequency. For the path or frequency with network anomaly risk, the exponential smoothing method is used to predict the packet loss probability within a certain period in the future.
[0051] The initial detection path, detection frequency parameters, and the packet loss probability distribution data under different paths and frequencies are obtained as the input for constructing an integer programming detection strategy optimization model. According to the obtained data, an integer programming algorithm is used to construct a detection strategy optimization model, and the objective function is set to minimize the overall packet loss monitoring blind area. In the constructed integer programming detection strategy optimization model, the decision variables are set as the detection paths and detection frequencies of each available zone network on the cloud. By solving the integer programming detection strategy optimization model, the optimal detection path and detection frequency combination of each available zone network on the cloud, that is, the optimal detection strategy, is obtained. If the obtained optimal detection strategy can meet the goal of minimizing the overall packet loss monitoring blind area, this strategy is applied to the packet loss monitoring of each available zone network on the cloud. After applying the optimal detection strategy, the packet loss probability distribution data of each available zone network on the cloud is continuously obtained for dynamically updating the integer programming detection strategy optimization model. According to the updated integer programming detection strategy optimization model, the optimal detection strategy is periodically re-solved and applied to the available zone network on the cloud to adapt to the changes in the network conditions and continuously minimize the overall packet loss monitoring blind area.
[0052] Exemplarily, in order to obtain the topological structure information and path dynamic change information of each available zone network on the cloud, a network topology discovery tool such as Nmap can be used to scan the target network to obtain device information, link information, routing information, etc. in the network. At the same time, by continuously monitoring the network traffic changes, machine learning algorithms such as support vector machine (SVM) are used to learn and predict the path change rules, so as to grasp the network topology dynamics in real time. Based on the above information, an initial detection path can be constructed through the minimum spanning tree algorithm. With the principle of covering the whole network and the shortest path, the initial detection frequency is set to once every 5 minutes according to factors such as network bandwidth and device load. After obtaining sufficient historical detection data, the ARIMA time series analysis algorithm is used to model the packet loss rate of each path, and the packet loss probability distribution under different paths and frequencies is obtained through parameter estimation. Finally, with the goal of comprehensive path coverage and the lowest total packet loss rate, an integer programming model is constructed to solve how to select the optimal detection path set and the detection frequency of each path, so as to minimize the blind area of the whole network packet loss monitoring under the condition of limited monitoring resources. Through the above technical solution, it is possible to adaptively optimize the detection strategy according to the dynamic characteristics of the cloud network, reduce the overhead while ensuring the monitoring quality, and improve the reliability of the cloud network.
[0053] S102. Use a weighted fusion algorithm to process the multi-dimensional network packet loss index data of all detection nodes in each available zone network on the cloud, and obtain the network packet loss comprehensive score of each available zone network on the cloud.
[0054] In this embodiment, in order to more accurately evaluate the network quality of each available zone network on the cloud, especially the network packet loss situation, a network packet loss comprehensive scoring system based on a weighted fusion algorithm is designed and implemented. This system performs weighted fusion processing on the multi-dimensional network packet loss index data of all detection nodes in each available zone network on the cloud, and finally obtains a comprehensive score to quantify the network packet loss situation.
[0055] Specifically, first determine the weight assignment principle of the weighted fusion algorithm. The weights can be assigned according to factors such as the location, importance, and historical performance of the detection nodes. For example, the detection nodes located on the key path of the network can have higher weights. For each available zone network on the cloud, use the multi-dimensional network packet loss index data of all its internal detection nodes as the input. According to the weight assignment principle, perform weighted processing on the data of each detection node. Fuse the weighted data to calculate the network packet loss comprehensive score of this available zone network on the cloud. The score can be represented in a percentage system or other forms for easy understanding and comparison. Display the calculated network packet loss comprehensive score in a visual form to relevant personnel, such as operation and maintenance personnel, management personnel, etc. Provide detailed score reports and data analysis to help relevant personnel understand the network situation and formulate corresponding optimization measures.
[0056] In this embodiment, the multi-dimensional data of multiple detection nodes are comprehensively processed through a weighted fusion algorithm, which can more comprehensively reflect the network packet loss situation and improve the accuracy of evaluation. Considering that the data of different detection nodes may have differences or errors, the weighted fusion algorithm can flexibly process the data according to the weight assignment principle, reduce the influence of individual abnormal data on the overall evaluation result, and enhance the robustness of the system.
[0057] In some embodiments, in the above step S102, the weighted fusion algorithm is used to process the multi-dimensional network packet loss index data of all detection nodes in each available zone network on the cloud respectively, and obtain the comprehensive network packet loss score of each available zone network on the cloud, specifically including:
[0058] Use the analytic hierarchy process to determine the weight coefficient of the network packet loss index data of each detection node;
[0059] According to the preset weighted fusion algorithm model and the weight coefficient, perform weighted calculation on the network packet loss index data of each detection node respectively to obtain the comprehensive network packet loss score of each detection node;
[0060] Use the K-means clustering algorithm to perform clustering analysis on the comprehensive network packet loss scores of all detection nodes to determine the network packet loss level data to which each detection node belongs;
[0061] Obtain the network packet loss level data of all detection nodes in each available zone network on the cloud, and obtain the network packet loss level distribution of the detection nodes in each available zone network on the cloud by performing frequency statistics on the network packet loss level data;
[0062] According to the preset network packet loss level weight coefficient, perform weighted calculation on the network packet loss level distribution of the detection nodes to obtain the comprehensive network packet loss score of each available zone network on the cloud.
[0063] In this embodiment, the network packet loss index data collected by multiple detection nodes, including indicators such as packet loss rate, delay, and jitter, are used as the input of the analytic hierarchy process. Through the analytic hierarchy process, a hierarchical structure model of the network packet loss index is constructed, the importance weights between the indicators are determined, and the weight coefficient of the network packet loss index data of each detection node is calculated. According to the preset weighted fusion algorithm model, the weight coefficient is weighted with the network packet loss index data of the corresponding detection node to obtain the comprehensive network packet loss score of each detection node.
[0064] Using the K-means clustering algorithm, taking the network packet loss scores of all detection nodes as input, the nodes are divided into different clustering clusters through iterative optimization, and each clustering cluster represents a network packet loss level. During the clustering process, by calculating the distance between the node scores and the cluster centers, the nodes are assigned to the nearest clustering cluster, and the cluster centers are continuously updated until the clustering results converge. According to the clustering results, the network packet loss level to which each detection node belongs is determined, and the network quality classification label of the node is obtained.
[0065] Obtain all detection node information in each available zone on the cloud, including metadata such as node IDs and the available zones where the nodes are located, and store it in the detection node information table. For each detection node, periodically conduct network packet loss tests, obtain the network packet loss rate data over a period of time, and convert the packet loss rate data into the corresponding packet loss level according to the preset packet loss rate level threshold, and store it in the detection node packet loss level data table. From the detection node information table and the detection node packet loss level data table, perform data association according to the available zone dimension, and conduct a frequency statistics on the packet loss level data of all detection nodes in each available zone to obtain the node quantity distribution of each packet loss level.
[0066] Obtain the preset network packet loss level weight coefficients, and for each detection node, determine its network packet loss level. According to the network packet loss level distribution of the detection nodes, use the weighted calculation method to obtain the network packet loss weighted score of each detection node. Aggregate the network packet loss weighted scores of the detection nodes in the same available zone on the cloud to obtain the network packet loss comprehensive score of this available zone.
[0067] Exemplarily, in order to determine the weight coefficients of the network packet loss metric data of each detection node, the analytic hierarchy process can be adopted. First, by means of expert scoring, the importance degrees of indicators such as delay, jitter, and retransmission rate are compared pairwise to construct a judgment matrix. Then, the matrix eigenvalue method is used to calculate the weight vectors of each indicator, and consistency tests are carried out until they pass. For example, the finally calculated weight coefficients are: delay 5, jitter 3, and retransmission rate 2. On this basis, using the weighted average model, the network packet loss metric data of each detection node is weighted and fused to calculate a comprehensive score. Then, the K-means clustering algorithm is used to perform clustering analysis on the comprehensive scores of all detection nodes. By calculating the Euclidean distances from each score sample to the cluster centers and iteratively updating the cluster centers until the clustering results converge, the detection nodes can finally be divided into three packet loss levels: low, medium, and high. In each cloud availability zone, the frequency distribution of detection nodes with different packet loss levels is counted, and according to the preset level weight coefficients (such as low 2, medium 3, high 5), the frequency distribution is weighted and averaged to obtain the overall packet loss comprehensive score of the network in this cloud availability zone for subsequent availability zone selection decisions. Through the above method, multi-dimensional indicators are converted into one-dimensional scores, and the importance degree differences of different indicators and levels are considered, which can more comprehensively and reasonably evaluate the packet loss level of the network in the cloud availability zone.
[0068] S103, determine the network packet loss scoring threshold for the network of each cloud availability zone according to the network area division, service type, and time period configuration parameters, and determine whether a network packet loss fault event has occurred in the network of each cloud availability zone by comparing the respective network packet loss comprehensive score and the network packet loss scoring threshold of the network of each cloud availability zone.
[0069] In this embodiment, in order to more accurately monitor the network packet loss situation of the network in the cloud availability zone and timely detect network packet loss fault events, a network packet loss fault monitoring system based on dynamic thresholds is designed and implemented. The system configures specific network packet loss scoring thresholds for the network of each cloud availability zone according to parameters such as network area division, service type, and time period. By comparing the network packet loss comprehensive score with these thresholds in real time, the system can automatically judge and report network packet loss fault events.
[0070] Specifically, according to factors such as the physical location and topological structure of the network in the cloud availability zone, the network is divided into different regions, and each region may have different network characteristics and business requirements. Identify and classify different business types running on the cloud, such as Web services, database services, video streaming media, etc. Different business types have different sensitivities to network packet loss. Considering the periodic changes in network traffic, time is divided into different periods (such as peak periods, off-peak periods, night periods, etc.), and the network load and packet loss conditions also vary during different periods. Based on historical data and business requirements, set reasonable network packet loss scoring thresholds for each cloud availability zone network, each business type, and each time period. Deploy detection nodes within each cloud availability zone network to collect multi-dimensional network performance index data such as network packet loss, latency, and jitter in real time. Use a weighted fusion algorithm to process the collected data to obtain the comprehensive network packet loss score for each cloud availability zone network. The system compares the comprehensive network packet loss score of each cloud availability zone network with its corresponding threshold in real time. When the score exceeds the threshold, the system automatically determines it as a network packet loss fault event.
[0071] In this embodiment, by configuring specific network packet loss scoring thresholds for each cloud availability zone network and comparing the score with the threshold in real time, the system can more accurately detect network packet loss fault events and improve the fault detection rate. The dynamic threshold mechanism can automatically adjust the threshold according to changes in parameters such as network area division, business type, and time period, thereby reducing the false alarm and missed alarm rates caused by improper setting of fixed thresholds.
[0072] In some embodiments, in the above step S103, determining the network packet loss scoring threshold for each cloud availability zone network according to the configuration parameters of network area division, business type, and time period specifically includes:
[0073] In each cloud availability zone network, respectively divide corresponding network regions for each user, and determine the first scoring coefficient of each network region according to the user characteristic information of the users corresponding to each network region;
[0074] Determine the second scoring coefficient of each business type according to the sensitivity of each business type to network packet loss;
[0075] Determine the third scoring coefficient of each business access time period according to the user traffic situation during the business access time period;
[0076] Adopt a weighted average algorithm to determine the network packet loss scoring threshold for each cloud availability zone network through the first scoring coefficient, the second scoring coefficient, and the third scoring coefficient.
[0077] In this embodiment, user information in each cloud availability zone network is obtained, including data such as user characteristics and network usage requirements. According to the obtained user characteristic information, clustering algorithms are used to group users, and users with similar characteristics are classified into the same network area. The user characteristic information in each network area is analyzed to extract key characteristic parameters, such as bandwidth requirements and latency sensitivity. According to the key characteristic parameters of users in the network area and the available network resources, the first scoring coefficient of each network area is calculated.
[0078] Business type information and corresponding network packet loss sensitivity data are obtained. According to a preset sensitivity threshold, the sensitivity level of the business type to network packet loss is judged, and corresponding second scoring coefficients are determined for different sensitivity levels. Specifically, machine learning algorithms are used to train the historical data of business types and network packet loss sensitivity to obtain an association model between business types and sensitivity levels. Through this model, the packet loss sensitivity of new business types is dynamically predicted, and its second scoring coefficient is determined.
[0079] Business access log data is obtained, and information such as user access time and access traffic is extracted. The data is divided into different time periods according to the access time, such as every hour, every day, etc. For each time period, the total access traffic within that time period is statistically calculated, and the proportion of the traffic in each time period to the total daily traffic is calculated. The traffic proportion of each time period is used as the third scoring coefficient of that time period, and a mapping relationship between the time period and the scoring coefficient is established. During subsequent business scoring, the corresponding third scoring coefficient is obtained from the mapping relationship according to the time period of the current business access.
[0080] The first scoring coefficient, the second scoring coefficient, and the third scoring coefficient are used as input parameters, and a weighted average algorithm is used to calculate the network packet loss scoring threshold for each cloud availability zone network. If the network packet loss scoring threshold exceeds the preset warning value, it is judged that there is a risk in the network quality of this cloud availability zone network, and network optimization and adjustment are required. According to the level of the network packet loss scoring threshold, the quality of each cloud availability zone network is evaluated and ranked, and the availability zone network with a higher scoring threshold is used as the preferred network resource to carry service requests for critical services and high-priority users.
[0081] Exemplarily, in order to determine the first scoring coefficient of each network area, the fuzzy comprehensive evaluation method can be adopted. First, a fuzzy evaluation matrix is constructed according to the characteristics of the user's bandwidth requirements, latency sensitivity, service importance, etc. Then, the qualitative indicators are quantified into membership degree values between 0 and 1 by using the membership function. Next, the comprehensive membership degree is calculated by the weighted average method as the first scoring coefficient of this network area. For example, if the membership degree of the bandwidth requirement of a certain area's user is 8, the membership degree of latency sensitivity is 6, and the membership degree of service importance is 9, then its first scoring coefficient is 8×3 + 6×3 + 9×4 = 78.
[0082] When determining the second scoring coefficient, according to expert experience, the packet loss sensitivity of different service types can be scored, and the score is between 0 and 10. Then, the scores of each service type are normalized, and the obtained weight coefficient is the second scoring coefficient. For example, if the packet loss sensitivity scores of video service, game service, and download service are 10, 8, and 6 respectively, the second scoring coefficients after normalization are 42, 33, and 25.
[0083] For the third scoring coefficient, by statistically analyzing the user traffic distribution in different time periods and combining expert experience, a weight coefficient between 0 and 1 is assigned to each time period. For example, a day is divided into four time periods: early morning, morning, afternoon, and evening, and the corresponding weight coefficients are 1, 3, 3, and 3 respectively. Finally, using the weighted average model and combining the above three scoring coefficients, the packet loss scoring threshold of each available zone network on the cloud is calculated.
[0084] Let the weights of the first, second, and third scoring coefficients be 4, 35, and 25 respectively, then the packet loss scoring threshold = the first scoring coefficient × 4 + the second scoring coefficient × 35 + the third scoring coefficient × 25. When the actual comprehensive packet loss score of the available zone network exceeds this threshold, it indicates that the network quality is poor and optimization and adjustment are required.
[0085] S104. Determine the packet loss fault event data of the available zone network on the cloud where the network packet loss fault event has occurred, and determine the packet loss event level according to the packet loss fault event data.
[0086] In this embodiment, after a network packet loss fault event occurs in the available zone network on the cloud, in order to more effectively evaluate the influence range and severity of the fault, a packet loss fault event level determination system is designed and implemented. Based on the occurred packet loss fault event data, through a series of evaluation indicators and algorithms, the level of the packet loss event is automatically determined so as to take corresponding countermeasures.
[0087] Specifically, when the system detects a network packet loss fault event in the available zone network on the cloud, it immediately collects relevant fault data. This data may include the time and location of the fault (specific available zone), the types of affected services, packet loss rate, duration, number of concurrent faults (such as increased latency and jitter), etc. Among them, the packet loss rate is used to measure the proportion of data packets lost during network transmission, the duration represents the time length of the fault event from occurrence to end, the types of affected services are used to evaluate the different sensitivities of different service types to network packet loss, and the number of concurrent faults is used to evaluate whether other network faults occurred during the same time period and the correlation between these faults.
[0088] Based on the above evaluation metrics, design and implement an evaluation algorithm to determine the level of the packet loss fault event. The algorithm can use methods such as weighted summation, fuzzy evaluation, and machine learning to convert the values of each evaluation metric into a comprehensive score, and then divide the fault event into different levels (such as severe, general, minor, etc.) according to the score. According to the output result of the evaluation algorithm, the packet loss fault event is divided into different levels. Each level corresponds to different response measures and priorities. For example, a severe-level fault event may require immediately activating the emergency plan, notifying key users, and allocating resources for emergency repair; while a minor-level fault event may only need to be recorded and monitored for subsequent optimization.
[0089] In this embodiment, by defining clear evaluation metrics and adopting a scientific evaluation algorithm, the system can more accurately evaluate the level of the network packet loss fault event in the available zone network on the cloud, reducing the errors and subjectivity of human judgment. According to the different fault levels, the system can reasonably allocate resources for fault handling. For severe-level fault events, more resources and manpower can be allocated for emergency repair; while for minor-level fault events, a more flexible handling method can be adopted.
[0090] In some embodiments, in the above step S104, determining the packet loss severity of the available zone network on the cloud where the network packet loss fault event has occurred, and determining the packet loss event level according to the packet loss severity specifically includes:
[0091] Obtain the historical packet loss data of the available zone network on the cloud, obtain the packet loss occurrence event metrics according to the historical packet loss data, and generate the original packet loss event log through the packet loss occurrence event metrics;
[0092] According to a preset time window, perform statistical analysis on the original packet loss event log to obtain the average packet loss rate and the maximum packet loss rate within each time window;
[0093] Determine the severity of the network packet loss fault according to the average packet loss rate and the maximum packet loss rate within all time windows, and form a packet loss event training data set;
[0094] Input the packet loss event training data set into a decision tree model for training to obtain a packet loss event classification model;
[0095] Obtain the packet loss fault event data of the cloud available zone network where a network packet loss fault event has occurred, input the packet loss fault event data into the packet loss event classification model for identification, and determine the packet loss event level of the cloud available zone network where the network packet loss fault event has occurred.
[0096] In this embodiment, historical packet loss data of the cloud available zone network is obtained, statistical analysis is performed on the packet loss data for different time periods to obtain indicators such as the average packet loss rate and the frequency of packet loss occurrences in each time period. According to the statistical analysis results of the historical packet loss data, a judgment threshold for packet loss occurrence events is determined. If the packet loss rate or the frequency of packet loss occurrences in a certain time period exceeds the preset threshold, it is determined that a packet loss event has occurred in that time period. For the time period determined to have a packet loss event, further analyze the packet loss data characteristics in that time period, extract key indicators such as the start and end times of packet loss occurrence and the peak packet loss rate, and generate packet loss event indicator data. Classify the packet loss event indicator data through a clustering algorithm, and identify different types of packet loss events according to the clustering results, such as packet loss caused by network congestion and packet loss caused by link failures. For each type of packet loss event, based on its packet loss data characteristics and occurrence rules, construct a corresponding packet loss event detection model for real-time monitoring of network packet loss conditions and timely discovery of packet loss events. After detecting a packet loss event, an original packet loss event log in a standard format is automatically generated according to indicators such as the time of event occurrence and the packet loss rate, recording the detailed information of the event occurrence.
[0097] Obtain preset time window parameters, determine the start time and end time of the time window; according to the time window parameters, screen out the packet loss event records falling within each time window from the original packet loss event log; for the packet loss event records within each time window, count the total number of packet loss events and the total number of data packets; calculate the average packet loss rate within each time window, and the formula is: average packet loss rate = number of packet loss events within the window / total number of data packets within the window; calculate the maximum packet loss rate within each time window by finding the maximum value of the packet loss rate within the window. According to the preset packet loss fault severity threshold, judge whether the maximum packet loss rate exceeds the threshold. If it exceeds, it is considered that a severe packet loss fault has occurred. Use the support vector machine algorithm, with the average packet loss rate and the maximum packet loss rate as features, to train a packet loss fault severity classification model. Use the decision tree algorithm, with the average packet loss rate and the maximum packet loss rate as features, to train a packet loss event classification model to judge whether a packet loss event has occurred. Take the average packet loss rate, the maximum packet loss rate, the packet loss fault severity, and the packet loss event as sample attributes to construct a packet loss event training data set.
[0098] According to the preprocessed training data set, a decision tree algorithm is used to construct a packet loss event classification model. Through the cross-validation method, the hyperparameters of the decision tree model are optimized to improve the classification accuracy of the model. The trained decision tree model is used to predict and classify the newly occurred packet loss event data. If the confidence level of the prediction result is lower than the preset threshold, the model retraining is triggered to update the packet loss event classification model.
[0099] The real-time packet loss data and related event information of the network in the cloud availability zone are continuously obtained through the network monitoring system. When a packet loss fault is detected, the subsequent packet loss event classification process is triggered. For the obtained packet loss fault event data, key feature parameters such as the time of packet loss occurrence, duration, and packet loss rate are extracted and formatted into a standardized data form for input into the packet loss event classification model. The standardized packet loss event data is input into the packet loss event classification model to automatically identify and analyze the severity of the packet loss event and give the corresponding rating result.
[0100] Further, after determining the packet loss event level of the network in the cloud availability zone where the network packet loss fault event has occurred, it further includes:
[0101] Obtain the network topology structure of the network in the cloud availability zone where the packet loss event level is the severe level. According to the network topology structure, determine the connection relationship and network device information between each host node of the network in the cloud availability zone to form a network topology diagram;
[0102] Based on the network topology diagram, use a graph neural network model to determine the fault propagation characteristic pattern, and determine the optimal corresponding measures for the network packet loss fault according to the fault propagation characteristic pattern.
[0103] In this embodiment, obtain the network device information of the network in the cloud availability zone where the packet loss event level is severe. By analyzing the connection relationship between network devices, construct a network topology structure model. According to the network topology structure model, use graph theory algorithms to analyze the connectivity between host nodes to obtain the connection relationship matrix between host nodes. Analyze the configuration information and alarm logs of network devices through deep learning algorithms to determine whether there are abnormalities in the network devices. If there are abnormalities, mark the abnormal devices in the topology diagram.
[0104] Construct a graph neural network model based on the network topology diagram, mapping network nodes and links to the nodes and edges of the graph. The attributes of the nodes and edges include network attribute parameters such as network device types and link bandwidths. For historical network fault data, extract feature parameters such as fault propagation paths, propagation times, and affected scopes, and train through the graph neural network model to obtain fault propagation feature patterns. When a packet loss fault occurs in the network, obtain the current network topology structure and fault alarm information, input them into the trained graph neural network model, and predict the fault propagation path and affected scope. According to the predicted fault propagation feature pattern, retrieve and match in the fault knowledge base to obtain candidate fault corresponding measures. The candidate measures include link load balancing, dynamic routing, etc. Adopt a reinforcement learning algorithm to evaluate the candidate fault corresponding measures. The evaluation metrics include packet loss rate recovery time, network throughput, etc., and obtain the optimal measure with the highest evaluation benefit through multiple iterations. Send the optimal fault corresponding measure to the network controller, and dynamically adjust the network forwarding rules and link policies through software-defined network technology to achieve fast location and recovery of network packet loss faults. Continuously monitor the network status, obtain the network performance metrics after fault recovery, verify the effectiveness of the measures through methods such as A / B testing, and optimize the graph neural network model and fault knowledge base according to the feedback results to continuously improve the intelligent level of fault diagnosis and recovery.
[0105] S105, based on the packet loss event level and the attribute information of each operation and maintenance personnel, determine different network fault briefing generation rules and different network fault briefing push rules for each operation and maintenance personnel, and generate and push corresponding network fault briefings for each operation and maintenance personnel according to the different network fault briefing generation rules and different network fault briefing push rules.
[0106] In this embodiment, in order to improve the response efficiency and handling ability of the operation and maintenance team for network fault events, a personalized network fault briefing generation and push system is designed and implemented. The system is based on the determined packet loss event level and the attribute information of each operation and maintenance personnel (such as skill specialties, responsibility scopes, preference settings, etc.), and customizes different network fault briefing generation rules and push rules for each operation and maintenance personnel. By generating and pushing personalized network fault briefings, the system can ensure that operation and maintenance personnel can obtain fault information related to their responsibilities and interests in a timely manner, so as to make responses and decisions faster.
[0107] Specifically, collect the attribute information of operation and maintenance personnel, and this information includes but is not limited to: skill specialties, the knowledge and experience of operation and maintenance personnel in specific technical or business fields; responsibility scopes, the specific network areas, business types or system components that operation and maintenance personnel are responsible for; preference settings, the personal preferences of operation and maintenance personnel for briefing content, format, push method, etc.
[0108] Based on the packet loss event level and the attribute information of the operation and maintenance personnel, the system formulates personalized briefing generation rules for each operation and maintenance personnel. These rules include: content screening, screening out relevant fault information according to the scope of responsibilities and skill specialties of the operation and maintenance personnel; priority sorting, sorting multiple fault events according to the event level and the urgency of the responsibilities of the operation and maintenance personnel; formatting customization: customizing the format, charts, colors, etc. of the briefing according to the preferences of the operation and maintenance personnel.
[0109] The system also formulates personalized briefing push rules for each operation and maintenance personnel. These rules include: push time, selecting a suitable push time according to the working hours and preferences of the operation and maintenance personnel; push method, pushing the briefing through multiple methods such as email, SMS, instant messaging, mobile applications, etc.; push frequency, setting a suitable push frequency according to the urgency of the fault event and the needs of the operation and maintenance personnel.
[0110] According to the formulated personalized briefing generation rules and push rules, the system generates corresponding network fault briefings for each operation and maintenance personnel and pushes them through the selected methods.
[0111] In this embodiment, by providing the operation and maintenance personnel with fault information closely related to their responsibilities and interests, the system can significantly improve the response efficiency of the operation and maintenance team to network fault events. The operation and maintenance personnel can identify and handle relevant fault events faster, reducing unnecessary interference and delays. The personalized briefing not only contains the basic information of the fault event, but also can include analysis and suggestions for the skill specialties of the operation and maintenance personnel, helping the operation and maintenance personnel to understand the fault situation more comprehensively and make more accurate decisions.
[0112] In some embodiments, in the above step S105, based on the packet loss event level and the attribute information of each operation and maintenance personnel, determining different network fault briefing generation rules and different network fault briefing push rules for each operation and maintenance personnel specifically includes:
[0113] Obtain the attribute information of each operation and maintenance personnel in the pre-established operation and maintenance personnel information database, where the attribute information includes level, professional scope, and responsibility attributes;
[0114] Adopt the K-means algorithm to perform clustering analysis on the attribute information, divide the operation and maintenance personnel with similar attribute information into the same category, and set corresponding network fault briefing templates for each category;
[0115] Use the level, professional scope, and responsibility attributes of the operation and maintenance personnel and the historical packet loss event data as the input features of the decision tree, and use the personalized network fault briefing generation rules and network fault briefing push rules as the output of the decision tree. By training the decision tree model, a network fault briefing rule setting model is formed;
[0116] Set the network fault briefing rule setting model and the network fault briefing template, and determine different network fault briefing generation rules and different network fault briefing push rules for each operation and maintenance personnel based on the packet loss event level and the attribute information of each operation and maintenance personnel.
[0117] In this embodiment, obtain the pre-established operation and maintenance personnel information database, extract the level, professional scope, and responsibility attribute information of each operation and maintenance personnel from the database; convert the extracted attribute information into a numerical vector, and construct an operation and maintenance personnel attribute feature matrix; use the K-means algorithm to perform clustering analysis on the operation and maintenance personnel attribute feature matrix, and divide the operation and maintenance personnel with similar attributes into the same category by calculating the distance between each data point and the cluster center; for each cluster of operation and maintenance personnel categories obtained, analyze the common attribute characteristics of the operation and maintenance personnel in this category, and set the corresponding network fault briefing template according to the attribute characteristics; when a network fault occurs, match the corresponding operation and maintenance personnel category according to the fault type and the specialty to which the faulty device belongs; obtain the network fault briefing template corresponding to this category, automatically fill in the fault information, and generate a targeted fault briefing.
[0118] Obtain the corresponding weight coefficients according to the level, professional scope, and responsibility attributes of the operation and maintenance personnel; count the efficiency and accuracy of different operation and maintenance personnel in handling faults according to the historical packet loss event data; use the operation and maintenance personnel attributes and historical data as the input features of the decision tree, and the personalized briefing rule as the output; use the ID3 algorithm to construct a decision tree model, and select the optimal division attribute through information gain; at each decision node, divide the data set into multiple subsets according to the attribute value; recursively construct sub-decision trees until all samples belong to the same category or cannot be divided further; generate personalized network fault briefing generation rules and push rules according to the path of the decision tree.
[0119] Based on the network fault briefing rule setting model and the network fault briefing template, combine the current packet loss event level and the attribute information of the operation and maintenance personnel related to the event to dynamically generate personalized network fault briefing content. According to the push rule in the network fault briefing rule setting model and the operation and maintenance personnel attributes, determine which operation and maintenance personnel the briefing content should be pushed to, and form a personalized network fault briefing push list. Send the personalized network fault briefing content to the corresponding operation and maintenance personnel according to the push list, and dynamically optimize the network fault briefing rule setting model according to the feedback to continuously improve the accuracy and timeliness of the network fault briefing.
[0120] Exemplarily, in order to obtain the attribute information of operation and maintenance personnel, a database containing the basic information of all operation and maintenance personnel can be established in advance. The records of each operation and maintenance personnel in the database include: level (such as junior, intermediate, senior), professional scope (such as network, system, security, etc.) and responsibility attributes (such as on-duty, inspection, emergency, etc.). Then, the K-means clustering algorithm is used to analyze these attribute information. The algorithm first randomly selects k clustering centers, and then iteratively assigns each data point to the nearest clustering center and updates the position of the clustering center until the clustering center no longer changes. Suppose the operation and maintenance personnel are divided into 5 categories, and the attribute characteristics of each category of personnel are similar, such as the level is intermediate and the professional scope is network. For each category, different network fault briefing templates are set, and the content and format in the template are customized according to the characteristics of the personnel in this category. Then, taking the attributes of operation and maintenance personnel and the historical packet loss event data as features, and taking personalized briefing generation rules (such as for junior network operation and maintenance personnel, the briefing content should be detailed, and for senior security operation and maintenance personnel, the briefing focuses on highlighting the fault impact) and briefing push rules (such as pushing within 5 minutes for severe faults and within 1 hour for general faults) as the output of the decision tree, the decision tree model is trained through the C5 algorithm. When a new packet loss fault occurs, according to the severity level of the fault and the attributes of the operation and maintenance personnel, the decision tree model determines the corresponding briefing generation rules and push rules for each person, and then combines the preset briefing template to automatically generate a personalized network fault briefing and send it to the relevant operation and maintenance personnel in a timely manner according to the push rules, so as to quickly start fault handling.
[0121] As an application example of the above embodiment, as Figure 2 shown, for a network monitoring system (such as an open-source zabbix system, etc.), a monitoring agent is deployed in the available zone to collect network packet loss data between available zones in real time through a simple ping full-interconnection method; when the network packet loss in the cloud available zone reaches the set threshold, the network monitoring system will trigger an alarm; the network monitoring system calls back the push platform interface; the push platform pulls the monitoring data of the network monitoring system through the interface; by pulling the monitoring data, combined with the preset threshold and judgment conditions of the push platform, through calculation and comparison, a network fault conclusion is output and written into the database. Outputting the network fault conclusion will trigger the push module to generate and push a network fault briefing. Before pushing the network fault briefing, first pull the previous fault conclusion and push time from the database for comparison. For the same conclusion within 5 minutes, it will not be pushed repeatedly. Call the third-party webhook interface to complete pushing the network fault briefing and notify the on-duty personnel by phone.
[0122] Referring to Figure 3 , an embodiment of the present invention provides a processing system 3 for network packet loss in a cloud available zone. The system 3 specifically includes:
[0123] The first processing module 301 is used to deploy a number of detection nodes in each available zone network on the cloud, and obtain multi-dimensional network packet loss metric data for each available zone network on the cloud according to the detection nodes;
[0124] The second processing module 302 is used to process the multi-dimensional network packet loss metric data of all detection nodes in each available zone network on the cloud by using a weighted fusion algorithm, and obtain a comprehensive network packet loss score for each available zone network on the cloud;
[0125] The third processing module 303 is used to determine a network packet loss score threshold for each available zone network on the cloud according to network area division, service type, and time period configuration parameters, and determine whether a network packet loss failure event has occurred in each available zone network on the cloud by comparing the comprehensive network packet loss score and the network packet loss score threshold of each available zone network on the cloud;
[0126] The fourth processing module 304 is used to determine the packet loss failure event data of the available zone network on the cloud where a network packet loss failure event has occurred, and determine the packet loss event level according to the packet loss failure event data;
[0127] The fifth processing module 305 is used to determine different network fault briefing generation rules and different network fault briefing push rules for each operation and maintenance personnel based on the packet loss event level and the attribute information of each operation and maintenance personnel, and generate corresponding network fault briefings for each operation and maintenance personnel according to different network fault briefing generation rules and different network fault briefing push rules and push them.
[0128] It can be understood that the content in the embodiment of the method for processing network packet loss in the available zone network on the cloud as Figure 1 shown is applicable to the embodiment of the system for processing network packet loss in the available zone network on the cloud in this embodiment. The functions specifically implemented in the embodiment of the system for processing network packet loss in the available zone network on the cloud in this embodiment are the same as those in the embodiment of the method for processing network packet loss in the available zone network on the cloud as Figure 1 shown, and the beneficial effects achieved are also the same as those in the embodiment of the method for processing network packet loss in the available zone network on the cloud as Figure 1 shown.
[0129] It should be noted that the information interaction, execution process, etc. between the above systems, due to being based on the same concept as the embodiment of the method of the present invention, for their specific functions and the technical effects brought, reference can be specifically made to the method embodiment part, and details are not described here again.
[0130] Those skilled in the art can clearly understand that, for the convenience and conciseness of description, only the above-mentioned division of each functional unit and module is used as an example. In actual applications, the above functions can be allocated to different functional units and modules as needed, that is, the internal structure of the system can be divided into different functional units or modules to complete all or part of the functions described above. Each functional unit and module in the embodiment can be integrated into a processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above integrated unit can be implemented in the form of hardware or in the form of a software functional unit. In addition, the specific names of each functional unit and module are only for the convenience of mutual distinction and do not limit the protection scope of this application. The specific working process of the units and modules in the above system can refer to the corresponding process in the foregoing method embodiment and will not be elaborated here.
[0131] Referring to Figure 4 , an embodiment of the present invention further provides a computer device 4, including: a memory 402, a processor 401, and a computer program 403 stored on the memory 402. When the computer program 403 is executed on the processor 401, it implements the method for processing network packet loss in an available zone on the cloud as described in any one of the above methods.
[0132] The computer device 4 may be a computing device such as a desktop computer, a notebook, a palm computer, and a cloud server. The computer device 4 may include, but is not limited to, a processor 401 and a memory 402. Those skilled in the art can understand that Figure 4 this is only an example of the computer device 4 and does not constitute a limitation on the computer device 4. It may include more or fewer components than shown in the figure, or combine some components, or different components. For example, it may also include input / output devices, network access devices, etc.
[0133] The so-called processor 401 may be a central processing unit (CPU), and the processor 401 may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.
[0134] In some embodiments, the memory 402 may be an internal storage unit of the computer device 4, such as a hard disk or memory of the computer device 4. In some other embodiments, the memory 402 may also be an external storage device of the computer device 4, such as a plug-in hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc., equipped on the computer device 4. Further, the memory 402 may also include both the internal storage unit and the external storage device of the computer device 4. The memory 402 is used to store an operating system, application programs, a Boot Loader, data, and other programs, such as program codes of the computer program. The memory 402 may also be used to temporarily store data that has been output or is to be output.
[0135] An embodiment of the present invention further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is run by a processor, it implements the method for processing network packet loss in an available zone on the cloud as described in any one of the above methods.
[0136] In this embodiment, if the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, to implement all or part of the processes in the above embodiment methods of the present application, a computer program can be used to instruct relevant hardware to complete. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, it can implement the steps of the above various method embodiments. Among them, the computer program includes computer program codes, and the computer program codes can be in the form of source code, object code, executable files, or some intermediate forms, etc. The computer-readable medium may at least include: any entity or device that can carry the computer program codes to the photographing device / terminal device, a recording medium, a computer memory, a Read-Only Memory (ROM), a Random Access Memory (RAM), an electrical carrier signal, a telecommunication signal, and a software distribution medium. For example, a USB flash drive, a mobile hard disk, a magnetic disk, or an optical disc, etc. In some jurisdictions, according to legislation and patent practice, the computer-readable medium may not be an electrical carrier signal and a telecommunication signal.
[0137] In the above embodiments, the descriptions of the various embodiments have their own emphases. For parts not detailed or recorded in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0138] Those of ordinary skill in the art will realize that the units and algorithm steps of the examples described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods for each specific application to implement the described functions, but such implementation should not be considered to exceed the scope of this application.
[0139] In the embodiments disclosed in this application, it should be understood that the disclosed apparatus / terminal device and method can be implemented in other ways. For example, the apparatus / terminal device embodiments described above are merely illustrative. For example, the division of the modules or units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the couplings or direct couplings or communication connections shown or discussed with each other can be through some interfaces. The indirect couplings or communication connections of the devices or units can be electrical, mechanical or other forms.
[0140] The units described as separate components may or may not be physically separated. The components shown as units may or may not be physical units, that is, they may be located in one place, or they may be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
Claims
1. A method for handling network packet loss in an available zone on the cloud, characterized in that The method specifically includes: Deploy a number of detection nodes in each available zone network on the cloud, and obtain multi-dimensional network packet loss index data for each available zone network on the cloud according to the detection nodes; Among them, before obtaining the multi-dimensional network packet loss index data for each available zone network on the cloud according to the detection nodes, it further includes: Obtain the topology structure information and path dynamic change information of each available zone network on the cloud, and determine the initial detection path and detection frequency parameters according to the topology structure information and the path dynamic change information; Adopt the ARIMA time series analysis algorithm to perform modeling analysis on the historical detection data of each available zone network on the cloud, and obtain the packet loss probability distributions under different paths and different frequencies; According to the initial detection path, the detection frequency parameters, and the packet loss probability distributions under different paths and different frequencies, with the goal of minimizing the overall packet loss monitoring blind area, construct an integer programming detection strategy optimization model, and the integer programming detection strategy optimization model is used to determine the optimal detection strategy for each available zone network on the cloud; Adopt a weighted fusion algorithm to process the multi-dimensional network packet loss index data of all detection nodes in each available zone network on the cloud respectively, and obtain the network packet loss comprehensive score for each available zone network on the cloud; Determine the network packet loss score threshold for each available zone network on the cloud according to the network area division, service type, and time period configuration parameters. By comparing the network packet loss comprehensive score and the network packet loss score threshold of each available zone network on the cloud respectively, determine whether a network packet loss failure event has occurred in each available zone network on the cloud; Determine the packet loss failure event data of the available zone network on the cloud where the network packet loss failure event has occurred, and determine the packet loss event level according to the packet loss failure event data; Based on the packet loss event level and the attribute information of each operation and maintenance personnel, determine different network fault briefing generation rules and different network fault briefing push rules for each operation and maintenance personnel, and generate corresponding network fault briefings for each operation and maintenance personnel according to different network fault briefing generation rules and different network fault briefing push rules and push them; 2. The method according to claim 1, wherein The adopting a weighted fusion algorithm to process the multi-dimensional network packet loss index data of all detection nodes in each available zone network on the cloud respectively, and obtaining the network packet loss comprehensive score for each available zone network on the cloud specifically includes: Adopt the analytic hierarchy process to determine the weight coefficient of the network packet loss index data of each detection node; According to the preset weighted fusion algorithm model and the weight coefficient, perform weighted calculation on the network packet loss index data of each detection node respectively, and obtain the network packet loss comprehensive score of each detection node; Adopt the K-means clustering algorithm to perform clustering analysis on the network packet loss comprehensive scores of all detection nodes, and determine the network packet loss level data to which each detection node belongs; Obtain the network packet loss level data of all detection nodes in each available zone network on the cloud, and obtain the network packet loss level distribution of the detection nodes in each available zone network on the cloud through frequency statistics of the network packet loss level data; According to the preset network packet loss level weight coefficient, perform weighted calculation on the network packet loss level distribution of the detection node to obtain the comprehensive network packet loss score of each available zone network on the cloud.
3. The method according to claim 1, wherein Determine the network packet loss score threshold of each available zone network on the cloud according to the network area division, service type, and time period configuration parameters, specifically including: In each available zone network on the cloud, divide corresponding network areas for each user respectively, and determine the first score coefficient of each network area according to the user characteristic information of the users corresponding to each network area; Determine the second score coefficient of each service type according to the sensitivity of each service type to network packet loss; Determine the third score coefficient of each service access time period according to the user traffic situation during the service access time period; Adopt the weighted average algorithm to determine the network packet loss score threshold of each available zone network on the cloud through the first score coefficient, the second score coefficient, and the third score coefficient.
4. The method according to claim 1, wherein Determine the packet loss fault event data of the available zone network on the cloud where the network packet loss fault event has occurred, and determine the packet loss event level according to the packet loss fault event data, specifically including: Obtain the historical packet loss data of the available zone network on the cloud, obtain the packet loss occurrence event index according to the historical packet loss data, and generate the original packet loss event log through the packet loss occurrence event index; According to the preset time window, perform statistical analysis on the original packet loss event log to obtain the average packet loss rate and the maximum packet loss rate within each time window; Determine the severity of the network packet loss fault according to the average packet loss rate and the maximum packet loss rate within all time windows, and form a packet loss event training data set; Input the packet loss event training data set into the decision tree model for training to obtain a packet loss event classification model; Obtain the packet loss fault event data of the available zone network on the cloud where the network packet loss fault event has occurred, input the packet loss fault event data into the packet loss event classification model for identification, and determine the packet loss event level of the available zone network on the cloud where the network packet loss fault event has occurred.
5. The method according to claim 4, characterized in that, After determining the packet loss event level of the available zone network on the cloud where the network packet loss fault event has occurred, it further includes: Obtain the network topology structure of the available zone network on the cloud with the packet loss event level of the severe level, determine the connection relationship and network device information between each host node of the available zone network on the cloud according to the network topology structure, and form a network topology diagram; Based on the network topology diagram, adopt a graph neural network model to determine the fault propagation characteristic pattern, and determine the optimal corresponding measures for network packet loss faults according to the fault propagation characteristic pattern.
6. The method according to any one of claims 1 to 5, characterized in that Based on the packet loss event level and the attribute information of each operation and maintenance personnel, determine different network fault briefing generation rules and different network fault briefing push rules for each operation and maintenance personnel, specifically including: Obtain the attribute information of each operation and maintenance personnel in the pre-established operation and maintenance personnel information database, and the attribute information includes level, professional scope, and responsibility attribute; The K-means algorithm is used to perform clustering analysis on the attribute information, and the operation and maintenance personnel with similar attribute information are divided into the same category, and a corresponding network fault briefing template is set for each category; Taking the level, professional scope and responsibility attributes of the operation and maintenance personnel and the historical packet loss event data as the input features of the decision tree, and taking the personalized network fault briefing generation rule and the network fault briefing push rule as the output of the decision tree, a network fault briefing rule setting model is formed by training the decision tree model; According to the network fault briefing rule setting model and the network fault briefing template, different network fault briefing generation rules and different network fault briefing push rules are determined for each operation and maintenance personnel through the packet loss event level and the attribute information of each operation and maintenance personnel.
7. A processing system for network packet loss in an available zone on the cloud, characterized in that, The system specifically includes: The first processing module is used to deploy a number of detection nodes in each cloud available zone network respectively, and obtain multi-dimensional network packet loss index data of each cloud available zone network according to the detection nodes; Among them, before obtaining the multi-dimensional network packet loss index data of each cloud available zone network according to the detection nodes, it also includes: Obtaining the topology structure information and path dynamic change information of each cloud available zone network, and determining the initial detection path and detection frequency parameters according to the topology structure information and the path dynamic change information; Using the ARIMA time series analysis algorithm to perform modeling analysis on the historical detection data of each cloud available zone network, and obtaining the packet loss probability distribution under different paths and different frequencies; According to the initial detection path, the detection frequency parameters and the packet loss probability distribution under different paths and different frequencies, with the goal of minimizing the overall packet loss monitoring blind area, an integer programming detection strategy optimization model is constructed, and the integer programming detection strategy optimization model is used to determine the optimal detection strategy of each cloud available zone network; The second processing module is used to process the multi-dimensional network packet loss index data of all detection nodes in each cloud available zone network respectively by using a weighted fusion algorithm, and obtain the network packet loss comprehensive score of each cloud available zone network; The third processing module is used to determine the network packet loss score threshold of each cloud available zone network according to the network area division, service type and time period configuration parameters, and determine whether a network packet loss fault event has occurred in each cloud available zone network by comparing the network packet loss comprehensive score and the network packet loss score threshold of each cloud available zone network; The fourth processing module is used to determine the packet loss fault event data of the cloud available zone network where the network packet loss fault event has occurred, and determine the packet loss event level according to the packet loss fault event data; The fifth processing module is used to determine different network fault briefing generation rules and different network fault briefing push rules for each operation and maintenance personnel based on the packet loss event level and the attribute information of each operation and maintenance personnel, and generate corresponding network fault briefings for each operation and maintenance personnel according to the different network fault briefing generation rules and different network fault briefing push rules and push them.
8. A computer device, characterized in that, Including: A memory, a processor, and a computer program stored in the memory, which, when executed on the processor, implement the method for processing network packet loss in an available zone on the cloud as described in any one of claims 1 to 6.
9. A computer-readable storage medium, characterized in that, A computer program is stored thereon, which, when run by a processor, implements the method for processing network packet loss in an available zone on the cloud as described in any one of claims 1 to 6.
Citation Information
Patent Citations
Multi-cloud network node detection method and device, computer equipment and storage medium
CN117221193A