A method and system for testing network performance on a cloud platform
By testing the network performance differences between Virtual LAN and Virtual Extended LAN, configuring cloud servers with different numbers of network card queues, and using iperf3 and sockperf tools for performance testing, the challenges of VXLAN network performance testing in cloud platforms were solved, achieving accurate network performance testing and bottleneck analysis.
Patent Information
- Application Number
- CN202410844964.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-06-27
- Publication Date
- 2025-10-31
- Estimated Expiration
- 2044-06-27
AI Technical Summary
Existing network performance testing tools are ill-suited to the diverse needs of VXLAN networks in cloud platforms, especially in virtualized environments, where accurately testing the network performance of cloud platforms remains a challenge.
By testing the network performance differences between virtual LANs and virtual extended LANs on physical machines, we created the environment required for cloud servers, configured cloud servers with different numbers of network card queues, used iperf3 and sockperf tools for performance testing, and performed core binding operations on the host machine to intuitively and quickly test the VXLAN network performance of the cloud platform.
It enables accurate testing of network performance in the cloud platform, identifies performance bottlenecks, provides a basis for VXLAN network performance, and helps optimize the cloud environment.
Smart Images

Figure CN118612130B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of network performance testing technology, and in particular to a method and system for testing network performance on a cloud platform. Background Technology
[0002] With the rapid development of cloud computing technology, cloud platforms have been widely used across various industries. Through virtualization technology, cloud platforms achieve centralized management and dynamic allocation of resources, providing enterprises with efficient, flexible, and scalable computing services. In cloud platforms, virtual networking technology is one of the key technologies for achieving multi-tenant isolation and data transmission. VXLAN (Virtual eXtensible Local Area Network), as an advanced virtual networking technology, is widely used in the network architecture of cloud platforms due to its excellent scalability and flexibility.
[0003] VXLAN extends L2 networks to L3 networks by encapsulating Ethernet frames, thus breaking through the limitations of traditional VLANs and achieving network isolation and communication in large-scale virtualization environments. However, with the expansion of cloud platform scale and the growth of business demands, how to effectively test the performance of VXLAN networks in cloud platforms to ensure the stability and efficiency of network transmission has become a technical challenge for the industry.
[0004] Traditional network performance testing methods are typically designed for physical network environments and are ill-suited to the characteristics of virtualized environments. In cloud platforms, the dynamic migration and elastic scaling of virtual machines present even greater challenges to network performance testing. Furthermore, existing network performance testing tools are often designed for specific network protocols or scenarios, making it difficult to meet the diverse needs of VXLAN networks in cloud platforms. Summary of the Invention
[0005] To address the aforementioned technical problems, this invention is proposed. Embodiments of this invention provide a cloud platform network performance testing method and system, which can accurately test the network performance of a cloud platform.
[0006] According to one aspect of the present invention, a cloud platform network performance testing method is provided, comprising: testing the network performance difference between a virtual local area network (VLAN) and a virtual extended local area network (VAN) of a physical machine; when the network performance difference between the VLAN and the VVA is less than a preset performance difference, creating an environment required by a cloud server in the platform environment; wherein the required environment includes the required number of network interface card (NIC) queues; creating multiple sets of cloud servers with different numbers of NIC queues on the same host machine to test the network performance between cloud physical hosts of the same computing node, obtaining multiple sets of first test results; creating multiple sets of cloud servers with different numbers of NIC queues on different host machines to test the network performance between cloud physical hosts of different computing nodes, obtaining multiple sets of second test results; when the multiple sets of first test results are not positively correlated with the number of NIC queues, and the multiple sets of second test results are not positively correlated with the number of NIC queues, repeatedly testing the CPU of the cloud server on the host machine of the cloud server by binding cores, to obtain the performance status of the environment.
[0007] In one embodiment, creating the environment required for a cloud server within a platform environment includes: disabling network performance limitations; configuring a corresponding number of network interface card (NIC) queues based on the attributes of the central processing unit (CPU) and memory after disabling network performance limitations; and creating a virtual machine specification template based on the multiple NIC queues; wherein the virtual machine specification template is used to define the virtual machine type.
[0008] In one embodiment, before testing the network performance difference between the virtual local area network (VLAN) and the virtual extended VLAN of the physical machine, the cloud platform network performance testing method includes: creating an OpenvSwitch bridge on the data network of two computing nodes and adding network interface cards (NICs); adding the NICs to the namespace and simulating network traffic using preset software or hardware tools to test the network performance of the virtual extended VLAN of the physical machine.
[0009] In one embodiment, after testing the network performance difference between the virtual local area network (VLAN) and the virtual extended VLAN of the physical machine, the cloud platform network performance testing method further includes: when the network performance difference between the VLAN and the virtual extended VLAN is greater than or equal to a preset performance difference, analyzing the performance of the created Open vSwitch bridge and checking whether the network card's optimized system performance is missing; when it is determined that the network card's optimized system performance is missing, taking the missing network card's optimized system performance as the reason for the poor network performance test, and ending the network performance test.
[0010] In one embodiment, creating multiple groups of cloud servers with different numbers of network interface card (NIC) queues on the same host machine includes: determining multiple cloud servers with the same number of NIC queues as a group; and deploying the multiple cloud servers in each group on the same compute node; wherein deploying the multiple cloud servers in each group on the same compute node includes: after creating multiple groups of cloud servers, migrating the cloud servers with the same number of NIC queues to the same compute node; or when creating multiple groups of cloud servers, assigning the cloud servers with the same number of NIC queues to the same compute node.
[0011] In one embodiment, multiple groups of cloud servers with different numbers of network interface card (NIC) queues are created on the same host machine to test the network performance between cloud physical hosts on the same computing node and obtain multiple sets of first test results. This includes: performing performance testing on the same group of cloud servers using a tool that tests the maximum available bandwidth on an IP network; if the tool for testing the maximum available bandwidth on an IP network is a single-threaded test, using a process with the same number of NIC queues to perform a flow test simultaneously, obtaining multiple sets of third test results; the flow test refers to simulating network traffic using preset software or hardware tools to test network performance; summing the third test results of multiple processes in each group to obtain a first accumulated result for each group; and performing a latency test on the first accumulated result for each group to obtain multiple sets of first test results.
[0012] In one embodiment, creating multiple groups of cloud servers with different numbers of network interface card (NIC) queues on different host machines includes: determining multiple cloud servers with the same number of NIC queues as a group; and deploying the multiple cloud servers in each group on different computing nodes. Deploying the multiple cloud servers in each group on different computing nodes includes: after creating multiple groups of cloud servers, migrating the cloud servers with the same number of NIC queues to different computing nodes; or when creating multiple groups of cloud servers, forcing each group of cloud servers to be on different computing nodes according to an anti-affinity principle.
[0013] In one embodiment, multiple groups of cloud servers with different numbers of network interface card (NIC) queues are created on different host machines to test the network performance between different computing node cloud physical hosts and obtain multiple sets of second test results. This includes: within the same group of cloud servers, using a tool to test the maximum available bandwidth on an IP network for performance testing; if the tool for testing the maximum available bandwidth on an IP network is a single-threaded test, using a process with the same number of NIC queues to perform a flow test simultaneously, obtaining multiple sets of fourth test results; the flow test refers to simulating network traffic using preset software or hardware tools to test network performance; summing the fourth test results of multiple processes in each group to obtain a second accumulated result for each group; and performing a latency test on the second accumulated result for each group to obtain multiple sets of second test results.
[0014] In one embodiment, on the host machine of the cloud server, the CPU of the cloud server is subjected to repeated core-binding tests, including: modifying the XML file of the cloud server to bind the cores, or using the virshvcpupin command to modify it; comparing the CPU performance of the cloud server before modification with the CPU performance of the cloud server after modification to obtain the performance improvement value; wherein, after the CPU of the cloud server is subjected to repeated core-binding tests, the cloud platform network performance testing method further includes: if the performance improvement value is greater than a preset improvement value, it is determined that the environment required by the cloud server created in the platform environment is unstable, and the cloud servers created in the environment have different network performance.
[0015] According to another aspect of the present invention, a cloud platform network performance testing system is provided, comprising: a testing module for testing the network performance difference between a virtual local area network (VLAN) and a virtual extended VLAN of a physical machine; an environment creation module for creating an environment required by a cloud server in a platform environment when the network performance difference between the VLAN and the virtual extended VLAN is less than a preset performance difference; wherein the required environment includes the required number of network interface card (NIC) queues; a first testing module for creating multiple sets of cloud servers with different numbers of NIC queues on the same host machine to test the network performance between cloud physical hosts of the same computing node and obtaining multiple sets of first test results; a second testing module for creating multiple sets of cloud servers with different numbers of NIC queues on different host machines to test the network performance between cloud physical hosts of different computing nodes and obtaining multiple sets of second test results; and a core-binding testing module for repeatedly testing the central processing unit (CPU) of the cloud server on the host machine of the cloud server when the multiple sets of first test results are not positively correlated with the number of NIC queues and the multiple sets of second test results are not positively correlated with the number of NIC queues, thereby obtaining the performance status of the environment.
[0016] The cloud platform network performance testing method and system provided by this invention first tests the network performance of a physical machine, without considering network card queues or other server performance aspects. Then, it tests the network performance of the physical machine across different / same hosts of the cloud server to examine the impact of the host machine. Finally, it performs core binding tests based on the test results. Each test step allows for the analysis of the cause of any problems, thus providing a direct and rapid assessment of the cloud platform's VXLAN network performance. Attached Figure Description
[0017] The above and other objects, features, and advantages of the present invention will become more apparent from the more detailed description of the embodiments of the invention in conjunction with the accompanying drawings. The drawings are provided to further illustrate the embodiments of the invention and form part of the specification. They are used together with the embodiments of the invention to explain the invention and do not constitute a limitation thereof. In the drawings, the same reference numerals generally represent the same parts or steps.
[0018] Figure 1 This is a flowchart illustrating a cloud platform network performance testing method provided in an exemplary embodiment of the present invention.
[0019] Figure 2 This is a flowchart illustrating a cloud platform network performance testing method provided in another exemplary embodiment of the present invention.
[0020] Figure 3 This is a schematic diagram of the structure of a cloud platform network performance testing system provided in an exemplary embodiment of the present invention. Detailed Implementation
[0021] Hereinafter, exemplary embodiments according to the present invention will be described in detail with reference to the accompanying drawings. Obviously, the described embodiments are merely some embodiments of the present invention, and not all embodiments of the present invention. It should be understood that the present invention is not limited to the exemplary embodiments described herein.
[0022] Figure 1 This is a flowchart illustrating a cloud platform network performance testing method provided in an exemplary embodiment of the present invention, as shown below. Figure 1 As shown, the cloud platform network performance testing methods include:
[0023] S1: Test the network performance difference between the physical machine's virtual LAN and virtual extended LAN.
[0024] The cloud platform of this invention can be taken as an example of an OpenStack environment. OpenStack is an open-source cloud computing management platform project, a combination of a series of open-source software projects designed to provide scalable and elastic cloud computing services. First, the VLAN (Virtual Local Area Network) network performance between physical machines can be tested. VLAN is a technology that establishes logically isolated networks on a physical network. It communicates through virtual links between switch ports; communication between different VLANs requires a router. VLANs can limit broadcast domains and reduce broadcast volume, reduce collision domains, and reduce network congestion, thereby improving network performance. Next, the VXLAN (Virtual Extensible Local Area Network) network performance of the physical machines can be tested. VXLAN is a network virtualization technology that creates virtual networks on existing network infrastructure, enabling communication between virtual machines across physical networks. It encapsulates Layer 2 Ethernet packets in UDP packets for transmission, achieving isolation between the virtual network and the physical network. Testing the physical machine's VLAN and VXLAN allows for a clear view of the performance of the VXLAN network without considering network card queues or other server performance aspects. Furthermore, if the physical machine itself performs poorly in testing, it can easily become a bottleneck in the entire cloud network environment, failing to guarantee network communication between nodes. If the network performance between physical machines is normal, then continue comparing the network performance differences between the physical machine's Virtual LAN (VLAN) and Virtual Extended LAN (VXLAN). If the performance difference between the VXLAN and VLAN networks is significant, it's necessary to check if the network interface card (NIC) is missing any performance features. For example, NIC offload technology (a method to optimize system performance) includes various implementations, such as TSO (TCP Segmentation Offload). Without TSO functionality, TCP bandwidth will be very poor. For the same server, the performance difference between supporting and not supporting TSO is typically around 1:1. Therefore, testing the network performance differences between the physical machine's VLAN and VXLAN can reveal whether the NIC is missing any performance features.
[0025] S2: When the network performance difference between the virtual LAN and the virtual extended LAN is less than the preset performance difference, create the environment required for the cloud server in the platform environment.
[0026] The required environment includes the number of network interface card (NIC) queues.
[0027] Check the power status of the physical server. Servers in power-saving mode will inevitably have poor performance. Therefore, when creating an environment that requires cloud server access within the platform, the physical server's power status needs to be adjusted to normal, and network performance limits should be disabled during the creation of the cloud server environment. For the example OpenStack cloud computing management platform, creating an environment that requires cloud server access is equivalent to creating a Flavor. A Flavor describes the resource configuration of a virtual machine instance. A Flavor is a virtual machine specification template used to define a virtual machine type, such as a virtual machine with a specific number of virtual CPUs (vCPUs), memory size, disk space, and other resources. The basic attributes of a Flavor include: ID, a unique identifier for the Flavor used for identification within OpenStack; Name, the display name of the Flavor, such as "small," "medium," or "large"; Number of vCPUs, the number of virtual CPUs available to the virtual machine, such as 1, 2, or 4; Memory size, the amount of memory allocated to the virtual machine, usually expressed in MB or GB; Disk space, the amount of disk space allocated to the virtual machine, including the root disk and other attached disks; Other options: These may include other extended options, such as network bandwidth and I / O performance. Flavors are created and managed by system administrators to provide ordinary users with the option to select resource specifications when creating virtual machines. Users can choose a suitable flavor to create virtual machine instances according to their needs, thereby ensuring that the virtual machine has sufficient resources to run the required applications or services.
[0028] Taking OpenStack as an example, the process of creating a Flavor includes: 1. Log in to the OpenStack management interface or use the command-line tool; 2. Navigate to Compute or a similar menu option; 3. Go to the Flavors page and view the list of existing Flavors; 4. Click the "Create Flavor" or similar button to start creating a new Flavor; 5. Enter the basic information of the Flavor, such as name, number of vCPUs, memory size, etc.; 6. Add other extended options as needed; 7. Click the "Create" or similar button to complete the creation of the Flavor.
[0029] Different flavor sizes require different numbers of network interface card (NIC) queues depending on the CPU and memory requirements. Generally, the NIC queue ratio is set to 1:2 or 1:4 for CPU. Each 4 cores form one NIC queue. Later, multiple processes are assigned to iperf3 (a tool for actively testing the maximum available bandwidth on an IP network) based on the NIC queue size.
[0030] To create the resources needed for the environment, such as creating a network, flavor, or cloud server, you can use the following command-line format:
[0031] Network creation: openstack network create network-test-c id-f value
[0032] openstack subnet create--subnet-range 40.0.0.0 / 16--gateway 40.0.0.1
[0033] --dhcp--network network-test network-dhcp
[0034] Flavor creation: openstack flavor create--ram=8192--disk=0--vcpus=4
[0035] --public--property SERVICE='ECS'--property SPEC='zqq'--propertyhw:cpu_cores='2'--property hw:cpu_sockets='1'--property hw:cpu_threads='2'--property hw:vif_multiqueue_number='1'ecs_4C8G0G_zqq
[0036] openstack flavor create--ram=16384--disk=0--vcpus=8--public--property SERVICE='ECS'--property SPEC='zqq'--property hw:cpu_cores='4'--propertyhw:cpu_sockets='1'--property hw:cpu_threads='2'--property
[0037] hw:vif_multiqueue_number='2'ecs_8C16G0G_zqq
[0038] openstack flavor create--ram=32768--disk=0--vcpus=16--public
[0039] --property SERVICE='ECS'--property SPEC='zqq'--property hw:cpu_cores='8'--property hw:cpu_sockets='1'--property hw:cpu_threads='2'--propertyhw:vif_multiqueue_number='4'ecs_16C32G0G_zqq
[0040] Open
[0041] Cloud server creation: openstack server create${SERVER_NAME}--flavor
[0042] ${FLAVOR[${FLAVOR_NUMBER}]}--image${IMAGE_ID}--user-data
[0043] / home / user_data.txt\
[0044] --config-drive=true--nicnet-id=${NETWORK_ID},v4-fixed-ip=40.0.0.${IP}\
[0045] --availability-zone$AZ:$1.
[0046] S3: Create multiple cloud servers with different numbers of network interface card queues on the same host machine to test the network performance between cloud physical hosts on the same computing node and obtain multiple sets of first test results.
[0047] A host machine is a physical computer or server that runs a hypervisor. It is responsible for managing physical resources, creating virtual machines, providing a hardware abstraction layer, and management functions. The host machine creates and runs virtual machines on top of itself, providing an isolated execution environment for each virtual machine. A host machine can run multiple virtual machines, each with a different operating system and applications, for testing software compatibility and performance. Several groups of cloud servers of different sizes and specifications are created on the same host machine; for example, each group of cloud servers consists of two machines on the same compute node. Typically, this includes servers with one, two, or four network interface card (NIC) queues. For the same group of cloud servers, performance tests can be performed not only on TCP, UDP, and PPS, but also on latency. The test results are observed to see if they are positively correlated with the number of NIC queues. This data can be recorded for later use in testing VXLAN network performance and latency across different server NIC operating systems.
[0048] S4: Create multiple cloud servers with different numbers of network interface card queues on different host machines to test the network performance between different computing node cloud physical hosts and obtain multiple sets of second test results.
[0049] Create several groups of cloud servers with different host machines and varying sizes. For example, each group could have two machines on different compute nodes. Performance tests for TCP, UDP, and PPS could be performed on the same group of cloud servers, along with latency tests. Observe whether the test results are positively correlated with network interface card (NIC) queue size. Record this data for later use in testing VXLAN network performance and latency on different server NIC operating systems.
[0050] S5: When multiple sets of first test results are not positively correlated with the number of network card queues, and multiple sets of second test results are not positively correlated with the number of network card queues, the central processing unit of the cloud server is repeatedly tested by binding cores on the host machine of the cloud server to obtain the performance status of the environment.
[0051] The execution order of S3 and S4 can be adjusted. If the test results of S3 and S4 do not show a positive correlation, it may be due to reasons such as the server's CPU residing in a NUMA (Non-Uniform Memory Access) node not being on the network card, requiring core binding. Bind all of a server's CPUs to a single NUMA node and check if the test results are stable and show improvement. If so, the cloud environment is unstable, and cloud servers created in this environment cannot guarantee consistent network performance. Servers with significant differences should be explained to users or excluded from the server procurement process.
[0052] Test results based on S1-S5 can better identify performance bottlenecks in the environment, whether in network cards, server architecture, or virtualization software. Simple testing tools such as Iperf3 and sockperf (a network benchmark utility based on socket APIs) can be used to standardize testing. The test results can be used to provide VXLAN network performance data for server procurement, operating system procurement, etc.
[0053] In one embodiment, S2 (creating the environment required for cloud servers in the platform environment) may include: disabling network performance limits; after disabling network performance limits, configuring a corresponding number of network interface card (NIC) queues based on the attributes of the central processing unit (CPU) and memory; and creating a virtual machine specification template based on the multiple NIC queues; wherein the virtual machine specification template is used to define the virtual machine type.
[0054] Check the power status of the physical server; servers in power-saving mode will inevitably perform poorly in tests. Therefore, when creating a flavor (virtual machine specification template), it's necessary to disable network performance limits to minimize their impact on test results. After disabling network performance limits, create the flavor. For flavors of the same size, configure different numbers of network interface card (NIC) queues depending on the CPU and memory requirements; generally, a NIC queue to CPU ratio of 1:2 or 1:4 is used. Later, specify multiple processes for iperf3 based on the NIC queues. Create the image; the image requires the latest versions of sockperf and iperf3 tools. Prioritize the newest versions of the testing software, as earlier versions of iperf3 may produce significantly different test results in the same environment compared to later versions.
[0055] In one embodiment, before S1 (testing the network performance difference between the virtual LAN and virtual extended LAN of the physical machine), the cloud platform network performance testing method may include: creating an OpenvSwitch bridge on the data network of two computing nodes and adding network cards; adding the network cards to the namespace and simulating network traffic using preset software or hardware tools to test the network performance of the virtual extended LAN of the physical machine.
[0056] Create an Open vSwitch bridge on the data network (used to handle VXLAN between servers) of the two compute nodes, add network interface cards (NICs), and add the NICs to the namespace to perform traffic generation. An OVS (Open vSwitch Bridge) is a virtual network switch that allows virtualized entities such as virtual machines or containers to communicate over the network. In the field of network performance testing, traffic generation typically refers to simulating network traffic using specific software or hardware tools to test the performance and reliability of network devices, systems, or services. A common software or hardware tool for simulating network traffic is iperf3, a tool for actively testing the maximum available bandwidth on an IP network. iperf3 can be used to test network throughput to evaluate network performance and bandwidth utilization. It can also be used to test the performance of your applications under various network conditions. During troubleshooting, iperf3 can help identify the cause of network problems. In network optimization and upgrades, iperf3 can be used to test performance differences between different configurations.
[0057] Iperf3 is easy to use, but it is not suitable for testing complex network scenarios. Therefore, multi-processing can be used to overcome the shortcoming of iperf3's single thread not being able to be allocated to multiple network card queues, thus maximizing the utilization of the cloud server.
[0058] VXLAN testing methods may include:
[0059] 1. Create a bridge:
[0060] ovs-vsctl add-brbr-zqq
[0061] 2. Add a VXLAN port to the bridge to enable cross-host virtual network connections. The local IP address is the dataIP of the host machine, and the remote IP address is the dataIP of the other host machine.
[0062] `ip a|grep data` finds the network cards of the current host machine and another host machine.
[0063] ovs-vsctl add-port br-zqq vxlan77--set interface vxlan77type=vxlanoptions:local_ip="100.200.116.117"options:remote_ip="100.200.116.119"options:dst_port="4972"
[0064] 3. Create a tap of type internal to connect to the bridge.
[0065] ovs-vsctl add-port br-test tap1--set interface tap1 type=internal
[0066] 4. Create a namespace and put this tap in it.
[0067] ip netns add test-netns
[0068] ip link set tap1 netns test-netns
[0069] 5. Enter the namespace and assign an IP address to tap.
[0070] ip netns exec test-netns / bin / bash
[0071] ifconfig lo 127.0.0.1up
[0072] ifconfig tap1 192.168.234.111 / 24up
[0073] The same method was executed on another compute node, enabling communication between the two compute nodes. Then, they performed mutual streaming.
[0074] Figure 2 This is a flowchart illustrating a cloud platform network performance testing method provided in another exemplary embodiment of the present invention, as shown below. Figure 2 As shown, after S1 (testing the network performance difference between the virtual LAN and virtual extended LAN of the physical machine), the cloud platform network performance testing method may also include:
[0075] S6: When the network performance difference between the virtual LAN and the virtual extended LAN is greater than or equal to the preset performance difference, analyze the performance of the created Open vSwitch bridge and check whether the network card's optimized system performance is missing.
[0076] Test the network performance between physical machines, then create a VXLAN on the physical machines for testing and compare the results. If the TCP test results differ significantly from those of the physical machines (e.g., 10G on the physical machines, 3G on the VXLAN), check if the network card has performance issues. This could be due to the network card not supporting TSO, resulting in poor server network performance.
[0077] S7: When it is determined that the network card's optimized system performance is lacking, the network performance test will end, taking the lack of optimized system performance of the network card as the reason for the poor network performance test.
[0078] If it is determined that the lack of optimized system performance of the network card is causing a significant performance difference between the VXLAN network and the VLAN network, and that there is a correlation between the two, then the network test should be terminated, and the cause of the network test problem should be accurately determined.
[0079] In one embodiment, S3 (creating multiple groups of cloud servers with different numbers of network interface card queues on the same host machine) may include: determining multiple cloud servers with the same number of network interface card queues as a group; and deploying the multiple cloud servers in each group on the same compute node; wherein, deploying the multiple cloud servers in each group on the same compute node includes: after creating multiple groups of cloud servers, migrating the cloud servers with the same number of network interface card queues to the same compute node; or when creating multiple groups of cloud servers, assigning the cloud servers with the same number of network interface card queues to the same compute node.
[0080] Network interface card (NIC) queues, also known as DMA (Direct Memory Access) queues, are a series of buffers (ring buffers) used by a NIC to receive and send network packets. Multi-queue NICs allow the NIC to have multiple such buffers to achieve higher network I / O throughput and lower latency. For example, in a test with two compute nodes, when creating virtual machines (VMs), ensure they are of the same specifications. Create servers of the same specifications on physical server A and physical server B; this constitutes a group. For instance, group physical server A and physical server B with 4-core 8GB servers each; group physical server A and physical server B with two 8-core 16GB servers each; and group physical server A and physical server B with two 16-core 32GB servers each. The 4-core 8GB server constitutes one NIC queue, and the 8-core 16GB server constitutes two NIC queues. Each group of cloud servers consists of two machines on the same compute node. Avoid assigning all machines to a single compute node, as this will affect the test results due to the compute node's load. You can create the VMs and then migrate them to the same compute node, or you can assign them to the same compute node (cloud platform administrator) during creation.
[0081] In one embodiment, S3 (creating multiple groups of cloud servers with different numbers of network interface card queues on the same host machine to test the network performance between cloud physical hosts of the same computing node and obtain multiple sets of first test results) includes: in the same group of cloud servers, using a tool for testing the maximum available bandwidth on the IP network to perform performance testing; if the tool for testing the maximum available bandwidth on the IP network is a single test thread, using the same number of processes as the number of network interface card queues to perform streaming tests at the same time, obtaining multiple sets of third test results; streaming test means simulating network traffic through preset software or hardware tools to test network performance; summing up the third test results of multiple processes in each group to obtain a first accumulated result for each group; performing a latency test on the first accumulated result for each group to obtain multiple sets of first test results.
[0082] The cloud servers in the same group use iperf3 for TCP, UDP, and PPS performance testing. In multi-NIC queues, since iperf3 is a single-threaded tester, it's difficult to utilize NIC queues effectively. Therefore, multiple iperf3 processes (number of NIC queues = number of iperf3 processes) are used simultaneously for streaming tests. The test results from multiple processes are then summed. Finally, sockperf is used for latency testing. TCP (Transmission Control Protocol) performance testing primarily focuses on the speed of TCP connection establishment, data transmission reliability, throughput, and network latency. UDP (User Datagram Protocol) performance testing primarily focuses on UDP packet transmission efficiency, packet loss rate, and network latency. PPS (Packets-per-second) performance testing primarily focuses on the processing capabilities of network devices (such as switches and routers) under specific network traffic conditions.
[0083] During testing, use the `top` command to view the CPU utilization percentage within the cloud server. PPS is tested using small packets with `-l 16`. Packet loss rate needs to be monitored during testing. If the packet loss rate is too high, `-A` is needed to perform iperf3 core binding. UDP and PPS packet loss should be controlled below 3%. The two cloud servers also need to swap client / server operations, and the test results are averaged.
[0084] For example, three groups of cloud servers on the same compute node are created using the above flavor. After creation, a streaming test is started. The three cloud server types use 1, 2, and 4 iperf3 processes respectively for streaming. The test results are observed to see if they are positively correlated with the network interface card (NIC) queue size. This data can be used later to test VXLAN network performance and latency for different server NIC operating systems.
[0085] In one embodiment, the main command for the test can be the following command:
[0086] TCP bandwidth: Client-side: `iperf3 -c serverIP -p port -t 100`
[0087] Server-side: iperf3-sp port
[0088] UDP bandwidth: Client: iperf3 -c serverIP -p port -ub 0g -P 2 -A 1,1 -t 100
[0089] Server-side: iperf3-sp port
[0090] PPS: Client-side: iperf3 -c serverIP -p port -ub 0g -l 16 -P 2 -A 1,1 -t 100
[0091] Server-side: iperf3-sp port
[0092] Latency: Server: sockperfsr-iip--tcp--daemonize
[0093] Client: sockperf ping-pong-i serverIp--tcp--full-rtt-m 64-t 30
[0094] Bind core: virshvcpupin instance-id--vcpu 0--cpulist 0
[0095] Bind 0 to physical core 0.
[0096] In one embodiment, S4 (creating multiple groups of cloud servers with different numbers of network interface card queues on different host machines) may include: determining multiple cloud servers with the same number of network interface card queues as a group; and deploying the multiple cloud servers in each group on different computing nodes; wherein, deploying the multiple cloud servers in each group on different computing nodes includes: after creating multiple groups of cloud servers, migrating the cloud servers with the same number of network interface card queues to different computing nodes; or when creating multiple groups of cloud servers, forcing each group of cloud servers to be on different computing nodes according to the anti-affinity principle.
[0097] Each cloud server group consists of 2 machines on different compute nodes. You can create the cloud servers and then migrate them to different compute nodes, or you can specify the compute nodes when creating the server group, or create a server group and force each group of servers to be on different compute nodes based on anti-affinity (here, strong anti-affinity is required; if weak anti-affinity is used, you still need to check whether the servers are created on different host machines after creation).
[0098] In one embodiment, S4 (creating multiple groups of cloud servers with different numbers of network interface card queues on different host machines to test the network performance between different computing node cloud physical hosts and obtain multiple sets of second test results) may include: in the same group of cloud servers, using a tool for testing the maximum available bandwidth on the IP network to perform performance testing; if the tool for testing the maximum available bandwidth on the IP network is a single test thread, using the same number of processes as the number of network interface card queues to perform flow testing at the same time, obtaining a fourth test result of multiple processes; flow testing means simulating network traffic through preset software or hardware tools to test network performance; summing the fourth test results of multiple processes in each group to obtain a second accumulated result for each group; performing a latency test on the second accumulated result of each group to obtain multiple sets of second test results.
[0099] During testing, the same number of iperf3 processes as the number of network interface card (NIC) queues are used. Cloud servers in the same group use iperf3 for TCP, UDP, and PPS performance testing. The test results from multiple processes are summed. Then, sockperf is used for latency testing. The `top` command is used to check the CPU utilization percentage within the cloud servers during testing. PPS is tested using small packets with `-l 16`. Packet loss rate needs to be monitored during testing. If the packet loss rate is too high, `-A` is used to bind iperf3 to a specific core. UDP and PPS packet loss should be controlled below 3%. The two cloud servers also need to swap client / server operations, and the test results are averaged. For example, three groups of cloud servers with different compute nodes are created using the above flavor, and streaming tests are started after creation. One, two, and four iperf3 processes are used for streaming on three different cloud server types, respectively. The test results are observed to see if they are positively correlated with the NIC queue size. This data can be used for later testing of VXLAN network performance and latency on different server NIC operating systems.
[0100] In one embodiment, S5 (performing repeated core-binding tests on the cloud server's CPU on the host machine of the cloud server) may include: modifying the cloud server's XML file to perform core binding, or using the virshvcpupin command to perform the modification; comparing the CPU performance of the cloud server before modification with the CPU performance of the cloud server after modification to obtain a performance improvement value; wherein, after S5 (performing repeated core-binding tests on the cloud server's CPU), the above-mentioned cloud platform network performance testing method may further include: when the performance improvement value is greater than a preset improvement value, determining that the environment required by the cloud server created in the platform environment is unstable, and the cloud servers created in the environment have different network performance.
[0101] Considering that different CPU scheduling on different physical servers can lead to performance loss for cloud server CPUs, and that some servers with tightly coupled network cards and NUMA interfaces may experience low performance test results due to cross-NUMA issues, cloud platforms deployed with such servers may exhibit significant performance differences. This is because production environments cannot guarantee that all cloud server CPUs are on the same NUMA node as their network cards. To address this, the cloud server's XML file can be modified to bind CPUs, or the `virshvcpupin` command can be used for modification. Comparing the test results before and after modification reveals that the performance improvement is generally not significant. If a substantial improvement is observed, this should be explained to the user or the server should be excluded from the procurement process. Therefore, if the S3-S4 test results do not show a positive correlation, it can be inferred that the server's CPU is not on the same NUMA node as the network card. Then, S5 should be tested. Whether the test results are stable and show improvement indicates the stability of the cloud environment. If the test results are stable and show improvement, it indicates instability in the cloud environment, and cloud servers created in this environment cannot guarantee consistent network performance.
[0102] The cloud platform network performance testing method provided in this invention, as described above, tests VXLAN performance by creating a bridge on a physical machine, without considering QEMU virtualization (which allows users to simulate and run multiple virtual machines on a single physical machine). It then progressively expands to test cross-host and non-cross-host scenarios on the cloud server. When significant deviations occur, the impact of the host machine's QEMU CPU is considered, allowing for the identification of problems at each step. Furthermore, the testing tools and methods used are simple; Iperf3 and sockperf are easy to install and test, requiring only simultaneous execution on both ends. Finally, the tests are standardized, and the results can be used for server procurement, operating system procurement, etc., providing a basis for VXLAN network performance testing. This helps determine the VXLAN network performance of the cloud platform and then deploy different applications based on the network performance.
[0103] Figure 3This is a schematic diagram of the structure of a cloud platform network performance testing system provided in an exemplary embodiment of the present invention, as shown below. Figure 3 As shown, according to another aspect of the present invention, a cloud platform network performance testing system 8 is provided, comprising: a testing module 81, for testing the network performance difference between a virtual local area network (VLAN) and a virtual extended VLAN of a physical machine; an environment creation module 82, for creating an environment required by a cloud server in a platform environment when the network performance difference between the VLAN and the virtual extended VLAN is less than a preset performance difference; wherein the required environment includes the required number of network interface card (NIC) queues; a first testing module 83, for creating multiple sets of cloud servers with different numbers of NIC queues on the same host machine to test the network performance between cloud physical hosts of the same computing node, and obtaining multiple sets of first test results; a second testing module 84, for creating multiple sets of cloud servers with different numbers of NIC queues on different host machines to test the network performance between cloud physical hosts of different computing nodes, and obtaining multiple sets of second test results; and a core-binding testing module 85, for repeatedly testing the central processing unit (CPU) of the cloud server on the host machine of the cloud server when multiple sets of first test results are not positively correlated with the number of NIC queues, and multiple sets of second test results are not positively correlated with the number of NIC queues, to obtain the performance status of the environment.
[0104] The cloud platform network performance testing system provided by this invention first tests the network performance of a physical machine, without considering network card queues or other server performance aspects. Then, it tests the network performance of the physical machine across different / same hosts of the cloud server to examine the impact of the host machine. Finally, it performs core binding tests based on the test results. Each test step allows for the analysis of the cause of any problems, thus providing a direct and rapid assessment of the cloud platform's VXLAN network performance.
[0105] In one embodiment, the environment creation module 82 can be configured to: disable network performance limitations; after disabling network performance limitations, configure a corresponding number of network interface card (NIC) queues based on the attributes of the central processing unit (CPU) and memory; and create a virtual machine specification template based on the multiple NIC queues; wherein the virtual machine specification template is used to define the virtual machine type.
[0106] In one embodiment, the cloud platform network performance testing system 8 can be configured to include: creating an Open vSwitch bridge on the data network of two computing nodes and adding network cards; adding the network cards to the namespace and simulating network traffic using preset software or hardware tools to test the network performance of the virtual extended local area network of the physical machine.
[0107] In one embodiment, the cloud platform network performance testing system 8 can be configured to: analyze the performance of the created Open vSwitch bridge and check whether the network card's optimized system performance is missing when the network performance difference between the virtual LAN and the virtual extended LAN is greater than or equal to a preset performance difference; when it is determined that the network card's optimized system performance is missing, take the missing network card's optimized system performance as the reason for the poor network performance test and end the network performance test.
[0108] In one embodiment, the first test module 83 can be configured to: determine multiple cloud servers with the same number of network interface card queues as a group; and deploy the multiple cloud servers in each group on the same computing node; wherein, deploying the multiple cloud servers in each group on the same computing node includes: after creating multiple groups of cloud servers, migrating the cloud servers with the same number of network interface card queues to the same computing node; or when creating multiple groups of cloud servers, assigning the cloud servers with the same number of network interface card queues to the same computing node.
[0109] In one embodiment, the first test module 83 can also be configured to: perform performance testing using a tool for testing the maximum available bandwidth on the IP network within the same group of cloud servers; if the tool for testing the maximum available bandwidth on the IP network is a single test thread, use processes with the same number of network card queues to perform streaming tests simultaneously to obtain third test results from multiple processes; streaming test means simulating network traffic using preset software or hardware tools to test network performance; sum up the third test results of multiple processes in each group to obtain a first accumulated result for each group; perform latency testing on the first accumulated result for each group to obtain multiple sets of first test results.
[0110] In one embodiment, the second test module 84 can be configured to: determine multiple cloud servers with the same number of network interface card queues as a group; and deploy the multiple cloud servers in each group on different computing nodes; wherein, deploying the multiple cloud servers in each group on different computing nodes includes: after creating multiple groups of cloud servers, migrating the cloud servers with the same number of network interface card queues to different computing nodes; or when creating multiple groups of cloud servers, forcing each group of cloud servers to be on different computing nodes according to the anti-affinity principle.
[0111] In one embodiment, the second testing module 84 can also be configured to: perform performance testing using a tool for testing the maximum available bandwidth on the IP network within the same group of cloud servers; if the tool for testing the maximum available bandwidth on the IP network is a single test thread, use processes with the same number of network card queues to perform flow testing simultaneously to obtain the fourth test results of multiple processes; flow testing means simulating network traffic through preset software or hardware tools to test network performance; sum up the fourth test results of multiple processes in each group to obtain the second accumulated result of each group; perform latency testing on the second accumulated result of each group to obtain multiple sets of second test results.
[0112] In one embodiment, the core binding test module 85 can be configured to: modify the XML file of the cloud server to bind the core, or use the virshvcpupin command to modify it; compare the CPU performance of the cloud server before modification with the CPU performance of the cloud server after modification to obtain the performance improvement value; wherein, the cloud platform network performance test system 8 can be configured to: when the performance improvement value is greater than the preset improvement value, determine that the environment required by the cloud server created in the platform environment is unstable, and the cloud servers created in the environment have different network performance.
[0113] This invention provides a cloud platform network performance testing system. The system can be implemented through software, hardware, or a combination of both. From a hardware perspective, in addition to the CPU, memory, network interface, and non-volatile memory, the device housing the system in the embodiment typically includes other hardware, such as a forwarding chip responsible for processing packets. Taking software implementation as an example, as a logical system, it is formed by the CPU of the device loading the corresponding computer program instructions from the non-volatile memory into memory for execution.
[0114] According to another aspect of the present invention, a computer-readable storage medium is provided, the storage medium storing a computer program for performing the cloud platform network performance testing method of any of the above embodiments.
[0115] In addition to the methods and devices described above, embodiments of the present invention may also be computer program products, which include computer program instructions that, when executed by a processor, cause the processor to perform the steps in the cloud platform network performance testing methods according to various embodiments of the present invention described in the "Exemplary Methods" section of this specification.
[0116] According to another aspect of the present invention, an electronic device is provided, comprising: a processor; a memory for storing processor-executable instructions; and a processor for executing the cloud platform network performance testing method of any of the above embodiments.
[0117] Furthermore, embodiments of the present invention may also be computer-readable storage media storing computer program instructions thereon, which, when executed by a processor, cause the processor to perform the steps in the cloud platform network performance testing methods according to various embodiments of the present invention described in the "Exemplary Methods" section above.
[0118] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
Claims
1. A method for testing the network performance of a cloud platform, characterized in that, include: Test the network performance differences between the virtual LAN and virtual extended LAN of the physical machine; When the network performance difference between the virtual LAN and the virtual extended LAN is less than a preset performance difference, the environment required by the cloud server is created in the platform environment; wherein, the required environment includes the required number of network interface card queues; Multiple cloud servers with different numbers of network interface card queues were created on the same host machine to test the network performance between cloud physical hosts on the same computing node and obtain multiple sets of first test results; Multiple cloud servers with different numbers of network interface card queues were created on different host machines to test the network performance between different computing node cloud physical hosts and obtain multiple sets of second test results; When multiple sets of the first test results are not positively correlated with the number of network interface card queues, and multiple sets of the second test results are not positively correlated with the number of network interface card queues, the central processing unit of the cloud server is subjected to repeated core binding tests on the host machine of the cloud server to obtain the performance status of the environment.
2. The cloud platform network performance testing method according to claim 1, characterized in that, Create the necessary environment for cloud servers within the platform environment, including: Disable network performance limits; After disabling network performance limits, configure the corresponding number of network interface card queues based on the CPU and memory attributes. Virtual machine specification templates are created based on multiple network interface card (NIC) queues; wherein, the virtual machine specification templates are used to define virtual machine types.
3. The cloud platform network performance testing method according to claim 1, characterized in that, Before testing the network performance differences between the virtual LAN and virtual extended LAN of the physical machine, the cloud platform network performance testing method includes: Create an Open vSwitch bridge on the data network between the two compute nodes and add network interface cards (NICs). Add the network card to the namespace and use preset software or hardware tools to simulate network traffic to test the network performance of the virtual extended LAN on the physical machine.
4. The cloud platform network performance testing method according to claim 3, characterized in that, After testing the network performance differences between the virtual LAN and virtual extended LAN of the physical machine, the cloud platform network performance testing method further includes: When the network performance difference between the virtual LAN and the virtual extended LAN is greater than or equal to the preset performance difference, analyze the performance of the created Open vSwitch bridge and check whether the network card's optimized system performance is missing. When it is determined that the network card's optimized system performance is lacking, the network performance test is terminated because the network card's optimized system performance is lacking, which is taken as the reason for the poor network performance test.
5. The cloud platform network performance testing method according to claim 1, characterized in that, Creating multiple cloud servers with different numbers of network interface card queues on the same host machine, including: Multiple cloud servers with the same number of network interface card queues are grouped together. The multiple cloud servers in each group are deployed on the same computing node; The arrangement of multiple cloud servers in each group on the same computing node includes: After creating multiple cloud server groups, migrate the cloud servers with the same number of network interface card queues to the same compute node; or When creating multiple cloud servers, assign the cloud servers with the same number of network interface card queues to the same computing node.
6. The cloud platform network performance testing method according to claim 5, characterized in that, Multiple cloud servers with different numbers of network interface card queues were created on the same host machine to test the network performance between cloud physical hosts on the same compute node, obtaining multiple sets of first test results, including: In the same group of cloud servers, use a tool to test the maximum available bandwidth on the test IP network to perform performance tests; If the tool for testing the maximum available bandwidth on an IP network is a single-threaded test, then the same number of processes as the number of network interface card queues are used to perform the flow test simultaneously to obtain the third test results of multiple processes; the flow test refers to simulating network traffic through preset software or hardware tools to test network performance; The third test results of multiple processes in each group are summed to obtain the first summation result for each group; Delay testing is performed on the first accumulated result of each group to obtain multiple first test results.
7. The cloud platform network performance testing method according to claim 1, characterized in that, Multiple cloud servers with different numbers of network interface card queues are created on different host machines, including: Multiple cloud servers with the same number of network interface card queues are grouped together. The multiple cloud servers in each group are deployed on different computing nodes; Specifically, the multiple cloud servers in each group are deployed on different computing nodes, including: After creating multiple cloud server groups, migrate the cloud servers with the same number of network interface card queues to different compute nodes; or When creating multiple cloud server groups, each group of cloud servers is forced to reside on different computing nodes based on the anti-affinity principle.
8. The cloud platform network performance testing method according to claim 7, characterized in that, Multiple cloud servers with varying numbers of network interface card queues were created on different host machines to test the network performance between different compute node cloud physical hosts, obtaining multiple sets of secondary test results, including: In the same group of cloud servers, use a tool to test the maximum available bandwidth on the test IP network to perform performance tests; If the tool for testing the maximum available bandwidth on an IP network is a single-threaded test, then the same number of processes as the number of network interface card queues are used to perform the flow test simultaneously to obtain the fourth test result of multiple processes; the flow test refers to simulating network traffic through preset software or hardware tools to test network performance; The fourth test results of multiple processes in each group are summed to obtain the second summation result for each group; Delay testing is performed on the second accumulated result of each group to obtain multiple groups of second test results.
9. The cloud platform network performance testing method according to claim 1, characterized in that, On the host machine of the cloud server, the central processing unit of the cloud server is subjected to repeated core-binding tests, including: Modify the cloud server's XML file to bind the core, or use the virshvcpupin command to modify it; Compare the CPU performance of the cloud server before and after the modification to obtain the numerical value of the performance improvement. The cloud platform network performance testing method, after repeatedly testing the CPU cores of the cloud server, also includes: When the performance improvement value is greater than the preset improvement value, it is determined that the environment required by the cloud server created in the platform environment is unstable, and the cloud server created in the environment has different network performance.
10. A cloud platform network performance testing system, characterized in that, include: The test module tests the network performance differences between the virtual LAN and virtual extended LAN of the physical machine; The environment creation module creates the environment required by the cloud server in the platform environment when the network performance difference between the virtual local area network and the virtual extended local area network is less than a preset performance difference; wherein, the required environment includes the required number of network interface card queues; The first test module creates multiple cloud servers with different numbers of network interface card queues on the same host machine to test the network performance between cloud physical hosts on the same computing node and obtain multiple sets of first test results; The second test module creates multiple cloud servers with different numbers of network card queues on different host machines to test the network performance between different computing node cloud physical hosts and obtain multiple sets of second test results; The core-binding test module performs repeated core-binding tests on the cloud server's central processing unit on the host machine of the cloud server when multiple sets of the first test results are not positively correlated with the number of network card queues, and multiple sets of the second test results are not positively correlated with the number of network card queues, in order to obtain the performance status of the environment.
Citation Information
Patent Citations
Method for realizing multi-tenant network in cloud network environment and device thereof, and medium
CN111478846A
Method for testing Web concurrency performance of multipath server system
CN115904944A