Cloud platform monitoring method and device based on flow analysis and automatic testing

By deploying active monitoring probes and passive traffic analysis at key nodes of the cloud platform network, a panoramic view of network quality is constructed, which solves the problem of limited monitoring scope in existing cloud platform monitoring methods, achieves real-time performance and accuracy, and improves fault location efficiency and tenant-level service assurance.

CN121333987APending Publication Date: 2026-01-13XINYANG BRANCH HENAN CO LTD OF CHINA MOBILE COMM CORP +1
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202511774572.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-28
Publication Date
2026-01-13

AI Technical Summary

Technical Problem

Existing cloud platform monitoring methods lack proactive testing tools, resulting in limited monitoring scope and an inability to achieve real-time and accurate monitoring, especially in terms of fault diagnosis and tenant-level service assurance.

Method used

Deploy active monitoring probes at key nodes of the cloud platform network, combine them with passive traffic analysis, and transmit test data back to the cloud business monitoring and assurance platform through a dedicated data backhaul channel to build a panoramic view of network quality. By analyzing traffic information, identify tenant traffic and service types, and establish a correlation model between network performance and service quality.

Benefits of technology

It enables panoramic real-time monitoring of the cloud platform, improves business quality and user experience, supports rapid fault location and differentiated service assurance, and meets high availability and SLA management requirements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121333987A_ABST
    Figure CN121333987A_ABST
Patent Text Reader

Abstract

The invention provides a cloud platform monitoring method and device based on flow analysis and automatic testing, and the method comprises the steps: deploying active monitoring probes at five key nodes in combination with cloud platform network topology; transmitting test data monitored by the monitoring probe back to the cloud service monitoring and guarantee platform to obtain performance indexes of each segment in a cloud platform network, and constructing a network quality panoramic view; traffic collection and traffic analysis are carried out in the virtual environment, and service key performance indexes are identified by analyzing quintuple information or application feature fields in the traffic; and associating the performance index of each segment with the service key performance indexes of different service types, constructing a correlation model of the network performance and the service quality of different service types, quantifying the influence of the network performance on service awareness based on the model to generate a service health degree, and displaying a service health degree view. According to the method, panoramic real-time monitoring of the cloud platform is realized, and the service quality and the user experience are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer data processing technology, and more specifically, to a cloud platform monitoring method and apparatus based on traffic analysis and automated testing. Background Technology

[0002] Cloud platform monitoring is a key technology for ensuring the stability of cloud computing services. Existing cloud platform monitoring methods typically rely on installing data collection plugins on each node to collect hardware or software metrics, and then providing a unified query interface through a metric tree to achieve cross-database metric management. For example, invention CN115827380A discloses a cloud platform monitoring system that stores metric trees and configuration information in a cached database, generates a unified query statement based on query requests, retrieves metric data from the target database, and returns it to the user device. This method reduces monitoring costs and improves efficiency, but it only focuses on unified metric querying and fails to comprehensively cover real-time monitoring of network traffic and business performance.

[0003] Existing technologies are primarily based on passive data collection and lack proactive testing methods, resulting in limited monitoring scope. For example, while the indicator tree method solves the problem of multi-database access, it does not address the deployment of key nodes in the network topology, traffic analysis, or business health assessment. This limits the real-time performance and accuracy of cloud platform monitoring, particularly in troubleshooting and tenant-level service assurance. Summary of the Invention

[0004] The purpose of this application is to provide a cloud platform monitoring method and apparatus based on traffic analysis and automated testing, which integrates active traffic analysis and passive automated testing to achieve panoramic real-time monitoring of the cloud platform, thereby improving business quality and user experience.

[0005] Firstly, a cloud platform monitoring method based on traffic analysis and automated testing is provided, which may include: Based on the cloud platform network topology, proactive monitoring probes are deployed at five key nodes: the user side, the access PE, the provincial network PE, the cloud platform entry point, and the cloud host. Through a dedicated data backhaul channel, the test data monitored by the monitoring probe is transmitted back to the cloud business monitoring and assurance platform, so that the cloud business monitoring and assurance platform can process the test data in segments to obtain the performance indicators of each segment in the cloud platform network, thereby constructing a panoramic view of network quality. Deploy traffic acquisition devices on host machines or virtual switches in the cloud platform network to collect and analyze traffic in the virtual environment, and identify tenant traffic, service type and corresponding key performance indicators (KPIs) by parsing the five-tuple information or application feature fields in the traffic. The performance metrics of each segment in the cloud platform network are correlated with the key performance metrics of different business types to construct a correlation model between network performance and the service quality of different business types. Based on this model, the impact of network performance on service perception is quantified to generate service health and a service health view is displayed.

[0006] In one possible implementation, deploying proactive monitoring probes includes: For networks outside the platform network, deploy portable probes on the user side and rack-mounted high-performance probes at the provincial network PE and cloud platform network entrances; Within the cloud platform network, soft probes are used in the virtualization environment, and these soft probes are deployed on the host machine and / or tenant virtual machines.

[0007] In one possible implementation, the portable probe supports Ping testing, Traceroute testing, TCP / UDP testing, and service simulation testing; The rack-mounted high-performance probe supports 10 Gigabit network and high-concurrency application stress testing, as well as application layer service performance analysis. The soft probe is deployed as a software package on the host machine or as a virtual network function (VNF) within a tenant virtual machine to collect operating system environment and performance parameters.

[0008] In one possible implementation, the cloud service monitoring and assurance platform segments the test data to obtain performance metrics for each segment of the cloud platform network, thereby constructing a comprehensive view of network quality, including: The cloud platform network is divided into five segments: user side, access PE, provincial network PE, cloud platform entrance, and cloud host. Collect test data for each segment boundary node; For each segment, the data from its ingress and egress nodes are analyzed comprehensively to calculate the performance metrics of that segment.

[0009] By integrating the performance metrics of each segment, a comprehensive view of network quality is constructed.

[0010] In one possible implementation, traffic analysis includes statistical analysis of at least one of the following key business performance indicators: TCP connection establishment latency, success rate, DNS resolution success rate, HTTP service response latency, and success rate, and supports drill-down to the IP or host dimension for fault location.

[0011] In one possible implementation, a correlation model between network performance and service quality is constructed, including: Perform data synchronization to ensure that network performance metrics and key performance indicators (KPIs) are consistent in time. Establish a mapping relationship between network performance indicators and key performance indicators (KPIs) of services; Through data analysis, the impact weight of network performance on service quality is quantified; Based on the above mapping relationship and influence weights, the comprehensive business health score for each business is obtained.

[0012] In one possible implementation, traffic acquisition devices are deployed on host machines or virtual switches within the cloud platform network to perform traffic acquisition and analysis, including: Collect CPU, memory, disk, and network card utilization data from the host machine and virtual machines to generate resource utilization and overall health information; A session traffic monitoring model is established, and the number of sessions and session packets are collected in real time through the session traffic monitoring model. The number of sessions and session packets collected in real time are compared with the corresponding preset deviation thresholds, and an alarm is generated based on the comparison results.

[0013] Secondly, a cloud platform monitoring device based on traffic analysis and automated testing is provided, which may include: The deployment unit is used to deploy active monitoring probes at five key nodes—the user side, the access PE, the provincial network PE, the cloud platform entry point, and the cloud host—in combination with the cloud platform network topology. The backhaul unit is used to transmit the test data monitored by the monitoring probe back to the cloud business monitoring and assurance platform through a dedicated data backhaul channel, so that the cloud business monitoring and assurance platform can process the test data in segments to obtain the performance indicators of each segment in the cloud platform network, and build a panoramic view of network quality. The processing unit is used to deploy traffic acquisition devices on host machines or virtual switches in the cloud platform network to perform traffic acquisition and traffic analysis in the virtual environment. The identification unit is used to identify tenant traffic, service type and corresponding key performance indicators (KPIs) by parsing the five-tuple information or application feature fields in the traffic. The building unit is used to associate the performance indicators of each segment in the cloud platform network with the key performance indicators of different business types, and to build a correlation model between network performance and the service quality of different business types. The generation unit is used to generate service health based on the model by quantifying the impact of network performance on service perception. The display unit is used to show the business health view.

[0014] Thirdly, an electronic device is provided, which includes a processor, a communication interface, a memory, and a communication bus, wherein the processor, the communication interface, and the memory communicate with each other through the communication bus; Memory, used to store computer programs; When a processor executes a program stored in memory, it implements any of the steps described in the first aspect above.

[0015] Fourthly, a computer-readable storage medium is provided, wherein a computer program is stored therein, and when executed by a processor, the computer program implements the steps of any of the methods described in the first aspect above.

[0016] This application provides a cloud platform monitoring method and apparatus based on traffic analysis and automated testing. The method, combined with the cloud platform network topology, deploys active monitoring probes at five key nodes: the user side, access PE, provincial network PE, cloud platform entry point, and cloud hosts. Through a dedicated data backhaul channel, the test data monitored by the probes is transmitted back to the cloud service monitoring and assurance platform. This platform then segments the test data to obtain performance indicators for each segment of the cloud platform network, constructing a panoramic view of network quality. Traffic acquisition devices are deployed on the host machines or virtual switches in the cloud platform network to collect and analyze traffic in a virtual environment. By parsing the five-tuple information or application feature fields in the traffic, tenant traffic, service types, and corresponding key performance indicators are identified. The performance indicators of each segment of the cloud platform network are correlated with the key performance indicators of different service types to construct a correlation model between network performance and service quality for different service types. Based on this model, the impact of network performance on service perception is quantified to generate service health, and a service health view is displayed. This method achieves panoramic real-time monitoring of the cloud platform, improving service quality and user experience. Attached Figure Description

[0017] To more clearly illustrate the technical solutions of the embodiments of this application, the accompanying drawings used in the embodiments of this application will be briefly introduced below. It should be understood that the following drawings only show some embodiments of this application and should not be regarded as a limitation of the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.

[0018] Figure 1 A flowchart illustrating a cloud platform monitoring method based on traffic analysis and automated testing, provided for an embodiment of this application; Figure 2 A schematic diagram of the structure of a cloud platform monitoring device based on traffic analysis and automated testing, provided for an embodiment of this application; Figure 3 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation

[0019] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of the embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application. Unless otherwise defined, the technical or scientific terms used in this application should have the ordinary meaning understood by those skilled in the art. The words "first," "second," and similar terms used in this application do not indicate any order, quantity, or importance, but are only used to distinguish different components. The words "comprising" or "including," etc., mean that the element or object preceding the word covers the element or object listed after the word and its equivalents, but do not exclude other elements or objects. The words "connected," "coupled," or "connected," etc., are not limited to physical or mechanical connections, but can include electrical connections, whether direct or indirect. "Up," "down," "left," "right," etc., are only used to indicate relative positional relationships. When the absolute position of the described object changes, the relative positional relationship may also change accordingly.

[0020] The existing cloud platform monitoring methods have the following main technical problems: The troubleshooting process is lengthy and time-consuming: existing methods rely on post-event indicator queries, which cannot capture network anomalies in real time, resulting in delays in fault location and affecting business continuity.

[0021] Lack of tenant-level monitoring: The monitoring system is not associated with the tenant's business, cannot identify the traffic and performance indicators of specific tenants, and is difficult to achieve differentiated service guarantees.

[0022] Lack of real-time monitoring and evaluation: Existing technologies mainly rely on static indicator queries, lacking proactive testing and traffic analysis, and are unable to dynamically assess the impact of network performance on service quality, such as real-time quantification of key indicators like latency and packet loss rate.

[0023] These issues result in insufficient comprehensiveness and reliability of cloud platform monitoring, failing to meet the high availability and SLA (Service Level Agreement) management requirements of modern cloud computing environments.

[0024] To address the aforementioned issues, this application provides a cloud platform monitoring method based on traffic analysis and automated testing. This method integrates active probe testing (automated testing) with passive traffic analysis, deploys multiple types of probes at end-to-end key nodes in the cloud platform network, achieves tenant-level, service-aware panoramic monitoring, and constructs a correlation model between network performance and service quality, supporting SLA verification and intelligent alarms.

[0025] The core of this method lies in achieving multi-dimensional and segmented monitoring of cloud platform network performance by combining active monitoring probes and passive traffic collection. Furthermore, it links network performance indicators with key performance indicators (KPIs) to build a correlation model between network performance and service quality, which is ultimately presented intuitively in the form of service health. This helps operations and maintenance personnel quickly locate faults, assess service status, and improve tenant experience.

[0026] The preferred embodiments of this application are described below with reference to the accompanying drawings. It should be understood that the preferred embodiments described herein are for illustration and explanation only and are not intended to limit this application. Furthermore, the embodiments and features in the embodiments of this application can be combined with each other without conflict.

[0027] Figure 1 This is a flowchart illustrating a cloud platform monitoring method based on traffic analysis and automated testing, provided as an embodiment of this application. Figure 1 As shown, the method may include: Step S110: Based on the cloud platform network topology, deploy active monitoring probes at five key nodes: the user side, the access PE, the provincial network PE, the cloud platform entrance, and the cloud host.

[0028] Deploy proactive monitoring probes at key nodes in the cloud platform network to implement automated testing. Key nodes include: the user side, the access PE (Provider Edge), the provincial network PE, the cloud platform ingress point, and cloud hosts. By deploying probes at these nodes, end-to-end coverage of the network path can be achieved. Specifically: For network components outside the cloud platform network (such as user access network and metropolitan area network): Portable probes are deployed on the user side of the dedicated or public network connecting the user access point to the cloud platform entry point. These probes are lightweight and easy to deploy, supporting basic network testing functions such as Ping testing (network connectivity testing based on ICMP Echo messages, used to calculate network performance metrics such as packet loss rate, latency, and jitter), Traceroute testing (quantifying path changes by probing the packet transmission path and returning intermediate routing information), TCP / UDP testing (testing the packet transmission quality of IP networks, measuring TCP handshake latency and success rate), and business simulation testing (simulating specific business operations such as user transaction behavior to test network and business performance).

[0029] Deploy rack-mounted high-performance probes at the provincial network PE and cloud platform network entry points. These probes are powerful hardware devices that support 10 Gigabit network environments, high-concurrency application stress testing (such as simulating a large number of user requests), and application-layer business performance analysis (such as HTTP transaction analysis).

[0030] For the internal network of the cloud platform (virtualized environment), software probes are used. Software probes are deployed in two forms: Package format: Directly deployed on the host machine (Hypervisor) to collect the host machine's operating system environment parameters (such as CPU model and kernel version) and performance parameters (such as CPU utilization and memory usage).

[0031] Virtual Network Function (VNF) form: Deployed inside the tenant's virtual machine as a lightweight agent to collect performance data inside the virtual machine (such as the virtual machine's own CPU and memory usage).

[0032] Through the above deployment, a complete monitoring chain is formed from the user end to the cloud host.

[0033] Step S120: Through a dedicated data backhaul channel, the test data monitored by the monitoring probe is backhauled to the cloud business monitoring and assurance platform to build a panoramic view of network quality.

[0034] Test data (including latency, packet loss rate, jitter, etc.) detected by the monitoring probes is transmitted back to the cloud service monitoring and assurance platform via a dedicated data backhaul channel (e.g., using a VPN tunnel or dedicated physical link to ensure data security and reliability). Upon receiving the data, the cloud service monitoring and assurance platform performs the following processing to construct a comprehensive view of network quality: Segmentation: The cloud platform network is logically divided into five segments: user side, access PE, provincial network PE, cloud platform entry point, and cloud host. These include the user side-access PE segment, the access PE-provincial network PE segment, the provincial network PE-cloud platform entry point segment, and the cloud platform entry point-cloud host segment. In other words, the starting point (user side) and ending point (cloud host) of the entire path are two independent nodes, plus the three intermediate nodes (access PE, provincial network PE, and cloud platform entry point), making a total of five key nodes. Segmentation refers to the connections between these nodes.

[0035] Test Data Collection: The cloud service monitoring and assurance platform collects test data from each segment boundary node (i.e., the node where probes are deployed). Boundary nodes refer to the start and end points of each segment. It is the proactive monitoring probes deployed on these nodes that provide the data needed for segment performance calculations.

[0036] Performance metrics calculation: For each segment, the data from its ingress and egress nodes are comprehensively analyzed to calculate the performance metrics for that segment. For example, the latency and packet loss rate of the "User Side - Access PE" segment are calculated.

[0037] View Construction: Integrate the performance metrics of each segment and present a panoramic view of network quality in the form of a topology map or dashboard. This view can intuitively display the health status of each part of the network, making it easy for operations and maintenance personnel to quickly locate the segment where performance bottlenecks are located.

[0038] Step S130: Deploy traffic acquisition devices on the host machine or virtual switch in the cloud platform network, collect and analyze traffic in the virtual environment, and identify tenant traffic, service type and corresponding key performance indicators by parsing the five-tuple information or application feature fields in the traffic.

[0039] Key performance indicators (KPIs) for business operations may include: HTTP response latency, transaction success rate, and DNS resolution success rate.

[0040] To conduct in-depth analysis of business traffic, this step involves deploying traffic acquisition devices on the host machine or virtual switch (vSwitch) of the cloud platform network to perform passive traffic collection and analysis. The traffic acquisition solution needs to be adapted to different cloud platform environments, as detailed below: A. The traffic collection scheme may include one or a combination of the following modes, and the specific choice can be made according to the actual cloud environment: Mode 1, Virtual Probe Deployment Mode within Virtual Machine: Deploy a lightweight virtual probe within the tenant's virtual machine to directly capture network interface card traffic.

[0041] Mode 2: Install a collection probe on the host hypervisor: Install a collection probe at the host level to capture all traffic passing through the vSwitch on that host.

[0042] Mode 3, combining vSwitch mirroring with virtual probes on virtual machines: Configure vSwitch port mirroring to copy traffic and send it to a virtual machine (with a virtual probe deployed) specifically for analysis.

[0043] Mode 4, vSwitch Mirroring Out Mode: Using the vSwitch mirroring function, traffic is mirrored to an external physical traffic analysis device.

[0044] (1) When the virtual environment is a Huawei Cloud environment: it is preferable to deploy the collection probe by installing it on the host Hypervisor. This enables in-depth monitoring of the underlying virtualization environment, reduces resource consumption, simplifies management, provides high-performance processing capabilities, ensures real-time data capture and analysis, and improves the overall system security and reliability. Specific steps include: Use the host machine's management port as the probe's management address. Configure the collection probe to capture data from the specified service traffic port using packet capture. Deploy the NPM (Network Performance Management) traffic analysis module in a virtual machine within the cloud platform's public management domain, assigning a management address as the platform's login and maintenance address. The cloud service monitoring and assurance platform issues traffic collection tasks to the NPM module, which then directs the collection probe to capture and analyze packets.

[0045] (2) When the virtual environment is a VMware cloud environment: it is preferable to use vSwitch image combined with virtual machine virtual probe mode or vSwitch image export mode. The specific steps include: allocating two virtual machines and deploying traffic probes and analysis platforms respectively. Configuring the mirroring function of Open vSwitch (OVS) to mirror the traffic of the virtual machine to be monitored to the virtual machine where the probe is located. First, perform local mirroring traffic test (i.e., traffic mirroring test of the host machine where the probe is located) to ensure that the mirroring configuration is correct. Then configure remote OVS mirroring to guide the traffic to the virtual machine where the probe is located to achieve centralized traffic collection.

[0046] This traffic collection method also enhances security by mirroring traffic to isolated probe virtual machines, preventing malicious interference and improving monitoring security. Centralized logging and monitoring not only meet various regulatory requirements but also support compliance and security auditing.

[0047] B. Flow analysis process: The collected traffic data is sent to the analysis engine, which performs the following operations by parsing the five-tuple information (source IP, destination IP, source port, destination port, protocol) or application characteristic fields (such as HTTP headers, DNS query content) in the traffic packets: Traffic identification: Differentiate outgoing tenant traffic from different business types (such as web services, database services, and video streams).

[0048] Business KPI Calculation: For the identified business, calculate its key performance indicators (KPIs), such as: TCP connection establishment latency, success rate, DNS resolution success rate, resolution latency, HTTP business response latency, success rate, throughput, and transaction success rate (for specific applications).

[0049] Fault location: Supports drilling down business KPIs to specific IP addresses or host dimensions. When a KPI is abnormal, the specific virtual machine or physical server that caused the abnormal traffic can be quickly located.

[0050] In addition, traffic analysis can also be used for resource monitoring and anomaly detection, specifically: Collect CPU, memory, disk, and network card utilization data from the host machine and virtual machines to generate resource utilization and overall health information.

[0051] In addition, a session traffic monitoring model is established to collect the number of network sessions and the number of packets per second in real time. This real-time data is compared with the mean calculated based on historical data from the past seven periods (continuously recording the number of sessions and packets per second in each monitoring period, and dynamically maintaining the mean μ and standard deviation σ of the indicators for the most recent seven periods), and a deviation threshold (e.g., μ ± 2σ) is set. When the real-time data deviates from the mean by more than the threshold, it is judged as abnormal, an alarm is generated, and the cloud platform monitoring department is notified to delimit and repair the fault.

[0052] Step S140: Associate the performance indicators of each segment in the cloud platform network with the key performance indicators of different business types, construct a correlation model between network performance and the service quality of different business types, quantify the impact of network performance on service perception based on the model, generate service health, and display the service health view.

[0053] First, the cloud business monitoring and assurance platform ensures that the network performance indicators (such as segmented latency) obtained are synchronized with the business KPI indicators (such as HTTP response latency) obtained in time, laying the foundation for correlation analysis.

[0054] Establish mapping relationships: The cloud service monitoring and assurance platform associates the network segment performance indicators that the service flows through with the service's KPIs based on the service traffic path. For example, the HTTP response latency of a certain web service is mapped to the latency indicators of four segments: "User side - Access PE", "Access PE - Provincial network PE", "Provincial network PE - Cloud platform entry", and "Cloud platform entry - Cloud host", establishing a mapping relationship between network performance indicators and key performance indicators (KPIs) of the service.

[0055] Quantifying Impact Weights: Using data analysis algorithms (such as regression analysis and machine learning models), calculate the degree of impact (i.e., impact weight) of performance changes in each network segment on business KPIs. For example, analysis found that the latency of the "cloud platform entry point - cloud host" segment contributes 40% to the HTTP response latency.

[0056] Calculate the overall health score: Based on the above mapping relationship and influence weights, a comprehensive business health score is calculated for each service. This score is a comprehensive reflection of network performance and business KPIs. For example, the business health score can range from 0 to 100, with a higher score indicating a better business status.

[0057] The business health score is used in the comprehensive perception evaluation system based on four indicators: DNS resolution latency (weight 20%), TCP connection latency (weight 20%), download latency (weight 30%), and throughput (weight 30%). Each indicator is divided into three linear score ranges: 0-50, 50-80, and 80-100, based on the test results. The score obtained by multiplying the interval score of each indicator by the indicator weight and then summing them is the health score.

[0058] Visual presentation: Business health scores can be presented to different roles in an intuitive view (i.e., a business health view), such as: Tenant view: Tenants can view the health status of the services they have subscribed to and understand the operational status of their services without having to worry about the complex underlying network details.

[0059] Special topic view (Operations view): Operations personnel can view the health status of the entire network or specific services, and can drill down to see specific network segment problems or abnormal service KPIs that affect the health status, so as to achieve accurate fault location.

[0060] In some embodiments, the cloud service monitoring and assurance platform can continuously collect cloud platform resource and configuration information, and automatically establish the association between tenants and the services they use based on tenant subscription information. Combined with the service level agreement (SLA) to which the tenant belongs, it implements modular monitoring and differentiated alarm mechanisms for different tenants to meet their varying service quality requirements. For example, stricter alarm thresholds are set for tenant services with high SLA requirements. When resource usage approaches or exceeds the threshold, the platform proactively alerts the tenant or operations personnel via email or SMS. For instance, for tenant services with a "Platinum" SLA requirement, the alarm threshold for HTTP response latency is set to 100 milliseconds; while for "Standard" tenants, the threshold can be set to 200 milliseconds. When monitoring metrics indicate that resource usage (such as CPU and bandwidth) is about to exceed or has already exceeded the upper limit agreed upon in the SLA, the system proactively alerts the tenant and operations personnel via email or SMS, providing early warning.

[0061] To enhance tenant trust, the cloud service monitoring and assurance platform also offers an SLA verification function. This function proactively verifies, through automated testing, whether the cloud platform meets the service levels promised to tenants. Specifically, this includes: Bandwidth Assurance Verification: Actual throughput is measured and compared with the tenant's contracted bandwidth through FTP file download tests or HTTP large file download tests. FTP or HTTP download tests demonstrate the tenant's contracted bandwidth assurance; continuous high-frequency Ping tests demonstrate the tenant's contracted link quality of service level assurance; and customized tests are performed based on different tenant business needs. The test results are then visualized to enhance tenant trust in the cloud platform.

[0062] Link quality verification: Verify whether the link latency and packet loss rate meet the SLA commitment by launching continuous Ping tests to key nodes.

[0063] Test results are visualized in the form of charts and graphs, generating SLA compliance reports.

[0064] The method described in this application enables comprehensive and interconnected monitoring of the cloud platform, from the underlying network to the upper-layer services. This method not only quickly identifies network faults but also assesses the service quality of the cloud platform from a business perspective, effectively ensuring the stable operation of tenant services and improving the management level and user satisfaction of cloud services.

[0065] Corresponding to the above method, embodiments of this application also provide a cloud platform monitoring device based on traffic analysis and automated testing, such as... Figure 2 As shown, the device includes: Deployment unit 210 is used to deploy active monitoring probes at five key nodes—user side, access PE, provincial network PE, cloud platform entry point, and cloud host—in combination with the cloud platform network topology. The backhaul unit 220 is used to transmit the test data monitored by the monitoring probe back to the cloud business monitoring and assurance platform through a dedicated data backhaul channel, so that the cloud business monitoring and assurance platform can process the test data in segments to obtain the performance indicators of each segment in the cloud platform network, so as to construct a panoramic view of network quality. Processing unit 230 is used to deploy traffic acquisition devices on host machines or virtual switches in the cloud platform network to perform traffic acquisition and traffic analysis in the virtual environment; The identification unit 240 is used to identify tenant traffic, service type and corresponding key performance indicators (KPIs) by parsing the five-tuple information or application feature fields in the traffic. Building unit 250 is used to associate the performance indicators of each segment in the cloud platform network with the key performance indicators of different business types, and to build a correlation model between network performance and the service quality of different business types. The generation unit 260 is used to generate service health based on the model by quantifying the impact of network performance on service perception. Display unit 270 is used to display the business health view.

[0066] The functions of each functional unit of the cloud platform monitoring device based on traffic analysis and automated testing provided in the above embodiments of this application can be implemented through the above methods and steps. Therefore, the specific working process and beneficial effects of each unit in the cloud platform monitoring device based on traffic analysis and automated testing provided in the embodiments of this application will not be repeated here.

[0067] This application also provides an electronic device, such as... Figure 3 As shown, it includes a processor 310, a communication interface 320, a memory 330, and a communication bus 340, wherein the processor 310, the communication interface 320, and the memory 330 communicate with each other through the communication bus 340.

[0068] Memory 330 is used to store computer programs; When the processor 310 executes the program stored in the memory 330, it performs the following steps: Based on the cloud platform network topology, proactive monitoring probes are deployed at five key nodes: the user side, the access PE, the provincial network PE, the cloud platform entry point, and the cloud host. Through a dedicated data backhaul channel, the test data monitored by the monitoring probe is transmitted back to the cloud business monitoring and assurance platform, so that the cloud business monitoring and assurance platform can process the test data in segments to obtain the performance indicators of each segment in the cloud platform network, thereby constructing a panoramic view of network quality. Deploy traffic acquisition devices on host machines or virtual switches in the cloud platform network to collect and analyze traffic in the virtual environment, and identify tenant traffic, service types and corresponding key performance indicators by parsing the five-tuple information or application feature fields in the traffic. The performance metrics of each segment in the cloud platform network are correlated with the key performance metrics of different business types to construct a correlation model between network performance and the service quality of different business types. Based on this model, the impact of network performance on service perception is quantified to generate service health and a service health view is displayed.

[0069] The communication bus mentioned above can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This communication bus can be divided into address bus, data bus, control bus, etc. For ease of illustration, only one thick line is used to represent it in the diagram, but this does not mean that there is only one bus or one type of bus.

[0070] The communication interface is used for communication between the aforementioned electronic devices and other devices.

[0071] The memory may include random access memory (RAM) or non-volatile memory (NVM), such as at least one disk storage device. Optionally, the memory may also be at least one storage device located remotely from the aforementioned processor.

[0072] The processors mentioned above can be general-purpose processors, including central processing units (CPUs), network processors (NPs), etc.; they can also be digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components.

[0073] The implementation methods and beneficial effects of the various components of the electronic device in the above embodiments for solving the problem can be found in [reference needed]. Figure 1 The steps in the illustrated embodiments are used to implement the electronic device. Therefore, the specific working process and beneficial effects of the electronic device provided in this application will not be repeated here.

[0074] In another embodiment provided in this application, a computer-readable storage medium is also provided, which stores instructions that, when executed on a computer, cause the computer to perform any of the cloud platform monitoring methods based on traffic analysis and automated testing described in the above embodiments.

[0075] In another embodiment provided in this application, a computer program product containing instructions is also provided, which, when run on a computer, causes the computer to execute any of the cloud platform monitoring methods based on traffic analysis and automated testing described in the above embodiments.

[0076] Those skilled in the art will understand that the embodiments in this application can be provided as methods, systems, or computer program products. Therefore, the embodiments in this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the embodiments in this application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0077] This application describes embodiments of methods, apparatus (systems), and computer program products according to embodiments of this application with reference to flowchart illustrations and / or block diagrams. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0078] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0079] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0080] Although preferred embodiments have been described in this application, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments as well as all changes and modifications falling within the scope of the embodiments of this application.

[0081] Obviously, those skilled in the art can make various modifications and variations to the embodiments of this application without departing from the spirit and scope of the embodiments of this application. Therefore, if these modifications and variations to the embodiments of this application fall within the scope of the claims in this application and their equivalents, then this application also intends to include these modifications and variations.

Claims

1. A cloud platform monitoring method based on traffic analysis and automated testing, characterized in that, The method includes: Based on the cloud platform network topology, proactive monitoring probes are deployed at five key nodes: the user side, the access PE, the provincial network PE, the cloud platform entry point, and the cloud host. Through a dedicated data backhaul channel, the test data monitored by the monitoring probe is transmitted back to the cloud business monitoring and assurance platform, so that the cloud business monitoring and assurance platform can process the test data in segments to obtain the performance indicators of each segment in the cloud platform network, thereby constructing a panoramic view of network quality. Deploy traffic acquisition devices on host machines or virtual switches in the cloud platform network to collect and analyze traffic in the virtual environment, and identify tenant traffic, service types and corresponding key performance indicators by parsing the five-tuple information or application feature fields in the traffic. The performance metrics of each segment in the cloud platform network are correlated with the key performance metrics of different business types to construct a correlation model between network performance and the service quality of different business types. Based on this model, the impact of network performance on service perception is quantified to generate service health and a service health view is displayed.

2. The method as described in claim 1, characterized in that, Deploy active monitoring probes, including: For networks outside the platform network, deploy portable probes on the user side and rack-mounted high-performance probes at the provincial network PE and cloud platform network entrances; Within the cloud platform network, soft probes are used in the virtualization environment, and these soft probes are deployed on the host machine and / or tenant virtual machines.

3. The method as described in claim 2, characterized in that, The portable probe supports Ping testing, Traceroute testing, TCP / UDP testing, and service simulation testing. The rack-mounted high-performance probe supports 10 Gigabit network and high-concurrency application stress testing, as well as application layer service performance analysis. The soft probe is deployed as a software package on the host machine or as a virtual network function (VNF) within a tenant virtual machine to collect operating system environment and performance parameters.

4. The method as described in claim 1, characterized in that, The cloud service monitoring and assurance platform segments the test data to obtain performance metrics for each segment of the cloud platform network, thereby constructing a comprehensive view of network quality, including: The cloud platform network is divided into five segments: user side, access PE, provincial network PE, cloud platform entrance, and cloud host. Collect test data for each segment boundary node; For each segment, the data of its ingress and egress nodes are analyzed comprehensively to calculate the performance index of that segment; By integrating the performance metrics of each segment, a comprehensive view of network quality is constructed.

5. The method as described in claim 1, characterized in that, Traffic analysis includes statistics on at least one of the following key business performance indicators: TCP connection establishment latency, success rate, DNS resolution success rate, HTTP service response latency, and success rate. It also supports drill-down to the IP or host dimension for fault location.

6. The method as described in claim 1, characterized in that, Construct a correlation model between network performance and service quality, including: Perform data synchronization to ensure the time consistency between network performance indicators and key business performance indicators; Establish a mapping relationship between network performance indicators and key business performance indicators; Through data analysis, the impact weight of network performance on service quality is quantified; Based on the above mapping relationship and influence weights, the comprehensive business health score for each business is obtained.

7. The method as described in claim 1, characterized in that, Deploy traffic acquisition devices on host machines or virtual switches in the cloud platform network to perform traffic acquisition and analysis, including: Collect CPU, memory, disk, and network card utilization data from the host machine and virtual machines to generate resource utilization and overall health information; A session traffic monitoring model is established, and the number of sessions and session packets are collected in real time through the session traffic monitoring model. The number of sessions and session packets collected in real time are compared with the corresponding preset deviation thresholds, and an alarm is generated based on the comparison results.

8. A cloud platform monitoring device based on traffic analysis and automated testing, characterized in that, The device includes: The deployment unit is used to deploy active monitoring probes at five key nodes—the user side, the access PE, the provincial network PE, the cloud platform entry point, and the cloud host—in combination with the cloud platform network topology. The backhaul unit is used to transmit the test data monitored by the monitoring probe back to the cloud business monitoring and assurance platform through a dedicated data backhaul channel, so that the cloud business monitoring and assurance platform can process the test data in segments to obtain the performance indicators of each segment in the cloud platform network, and build a panoramic view of network quality. The processing unit is used to deploy traffic acquisition devices on host machines or virtual switches in the cloud platform network to perform traffic acquisition and traffic analysis in the virtual environment. The identification unit is used to identify tenant traffic, service type and corresponding key performance indicators by parsing the five-tuple information or application feature fields in the traffic. The building unit is used to associate the performance indicators of each segment in the cloud platform network with the key performance indicators of different business types, and to build a correlation model between network performance and the service quality of different business types. The generation unit is used to generate service health based on the model by quantifying the impact of network performance on service perception. The display unit is used to show the business health view.

9. An electronic device, characterized in that, The electronic device includes a processor, a communication interface, a memory, and a communication bus, wherein the processor, the communication interface, and the memory communicate with each other through the communication bus; Memory, used to store computer programs; A processor, when executing a program stored in memory, implements the method of any one of claims 1-7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the method described in any one of claims 1-7.

Citation Information

Patent Citations

  • Cloud platform monitoring method and cloud platform monitoring system

    CN115827380A