Flow forwarding method and system in cloud platform, terminal and medium
By creating a master and standby instance of the traffic forwarding gateway within the cloud platform and monitoring its status, and configuring a traffic forwarding policy, the traffic forwarding interruption problem caused by a single node failure is solved, high-reliability traffic forwarding is achieved, and the robustness of the cloud platform is improved.
Patent Information
- Application Number
- CN202510292468.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-12
- Publication Date
- 2025-06-03
AI Technical Summary
When using traffic forwarding gateways, existing cloud platforms adopt a single node deployment method, resulting in low reliability. Once a node fails, traffic forwarding will be interrupted, affecting the communication and service functions of the entire link.
Create a primary and secondary instance of the traffic forwarding gateway in the cloud platform, generate a primary gateway and at least one backup gateway to form a traffic forwarding gateway cluster, and monitor the status of the main gateway and backup gateway. Configure the forwarding policy between the main gateway and backup gateway according to the monitoring results to ensure that traffic forwarding is not interrupted.
Through the deployment and status monitoring of the main and spare instances, high reliability of traffic forwarding is achieved, ensuring normal link communication between various virtual private networks within the cloud platform, and improving the robustness of the cloud platform.
Smart Images

Figure CN120090972A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of virtual network traffic forwarding, and particularly relates to a traffic forwarding method, system, terminal and medium in a cloud platform. Background Art
[0002] In recent years, cloud computing technology has been more and more widely used. Its virtualization and abstraction of various resources such as computing, network, and storage provide users with extremely convenient resource usage methods and flexible resource expansion capabilities. In cloud computing, the virtualization of resources enables physical resources such as computing, network, and storage to be abstracted into virtual resources, so that they can be dynamically allocated and managed.
[0003] Virtual Private Cloud (VPC), as a flexible and secure network architecture, can be applied to cloud environments. VPC allows users to create an isolated network environment in the cloud, with its own IP address range, and can independently set security group rules and network topologies. In a cloud platform, communication between multiple VPCs and between a VPC and an external network is achieved through a traffic forwarding gateway to ensure the security and reliability of data transmission.
[0004] However, currently, when using a traffic forwarding gateway in a cloud platform, a single-node deployment method is adopted, resulting in low reliability. Once this node fails, such as hardware damage, software crash, etc., traffic forwarding will be interrupted, thus affecting the communication of the entire link and the business functions. Summary of the Invention
[0005] To solve the above problems, the present invention provides a traffic forwarding method, system, terminal and medium in a cloud platform to ensure that traffic forwarding is not interrupted, thereby ensuring normal link communication between virtual private networks in the cloud platform and improving the robustness of the cloud platform.
[0006] In a first aspect, the technical solution of the present invention provides a traffic forwarding method in a cloud platform, including the following steps: Create a primary and standby instance of a traffic forwarding gateway in the cloud platform to generate a traffic forwarding gateway cluster composed of a primary gateway and at least one standby gateway; Mount the primary gateway and each standby gateway to each virtual private network, and issue the same traffic forwarding rules to the primary gateway and each standby gateway; Monitor the status of the primary gateway and each standby gateway; Configure the forwarding strategy of traffic between the primary gateway and each standby gateway according to the gateway status monitoring result.
[0007] In an optional implementation manner, monitoring the status of the primary gateway and each standby gateway specifically includes: Periodically capture the performance metric data of the primary gateway and each standby gateway through the Prometheus server; the performance metric data includes service survival status, throughput, forwarding delay, CPU usage rate, and memory usage rate. Based on the performance metric data of the gateway, determine whether the gateway is faulty according to the pre-configured detection rules. According to the gateway status monitoring results, configure the forwarding policy of the traffic between the primary gateway and each standby gateway, specifically including: If the primary gateway is faulty, calculate the performance evaluation value of each surviving standby gateway through the following formula: Performance evaluation value = First weight * Throughput + Second weight * Forwarding delay + Third weight * CPU usage rate + Fourth weight * Memory usage rate. Select the standby gateway with the largest performance evaluation value as the new primary gateway. Direct all the traffic directed to the faulty primary gateway to the new primary gateway.
[0008] In an alternative embodiment, determining whether the gateway is faulty according to the pre-configured detection rules specifically includes: Detect whether the identifier of the service survival status is 0. If so, determine that the gateway is faulty. Detect whether the forwarding delay within a unit time exceeds the maximum forwarding delay threshold. If it is the case for N1 consecutive detection cycles, determine that the gateway is faulty. Detect whether the CPU usage rate within a unit time exceeds the maximum CPU usage rate threshold. If it is the case for N2 consecutive detection cycles, determine that the gateway is faulty. Detect whether the memory usage rate within a unit time exceeds the maximum memory usage rate threshold. If it is the case for N3 consecutive detection cycles, determine that the gateway is faulty.
[0009] In an alternative embodiment, according to the primary gateway status monitoring results, configuring the forwarding policy of the traffic between the primary gateway and each standby gateway specifically further includes: If the primary gateway is not faulty, screen the captured performance metric data to obtain the target performance metric data; the target performance metric data includes throughput, forwarding delay, CPU usage rate, and memory usage rate. Preprocess the target performance metric data and construct the preprocessed target performance metric data into an input vector. Transmit the input vector to the pre-trained gateway status prediction model for processing to obtain the predicted type of the future state of the primary gateway. Configure the forwarding policy of the traffic between the primary gateway and each standby gateway according to the predicted type of the future state of the primary gateway.
[0010] In an alternative embodiment, according to the predicted type of the future state of the primary gateway, a forwarding policy for traffic between the primary gateway and each standby gateway is configured, specifically including: Obtain a pre-configured proportional allocation mapping table; According to the predicted type of the future state of the primary gateway, match the allocation ratio of the primary gateway and the sum of the allocation ratios of all standby gateways from the proportional allocation mapping table; Direct the traffic of the virtual private network to the primary gateway according to the allocated ratio of the corresponding primary gateway; Perform proportional allocation of standby gateways on the sum of the allocation ratios of all standby gateways according to the performance evaluation values of each surviving standby gateway. The higher the performance evaluation value, the larger the allocated ratio; Direct the remaining traffic of the virtual private network to the standby gateways according to the allocated ratio of the corresponding standby gateways.
[0011] In an alternative embodiment, the method further includes the following steps: When the primary gateway and / or standby gateway fails, issue a fault alarm.
[0012] In an alternative embodiment, mounting the primary gateway and each standby gateway to each virtual private network specifically includes: Create a subnet in the virtual private network and mount the subnet on the virtual router of the virtual private network; Mount the primary gateway and each standby gateway to the network card of the subnet.
[0013] In a second aspect, the technical solution of the present invention provides a traffic forwarding system within a cloud platform, including: A gateway cluster creation module, configured to create primary and standby instances of traffic forwarding gateways within the cloud platform, and generate a traffic forwarding gateway cluster composed of a primary gateway and at least one standby gateway; A gateway mounting module, configured to mount the primary gateway and each standby gateway to each virtual private network, and issue the same traffic forwarding rules to the primary gateway and each standby gateway; A gateway status monitoring module, configured to monitor the status of the primary gateway and each standby gateway; A traffic forwarding execution module, configured to configure a forwarding policy for traffic between the primary gateway and each standby gateway according to the gateway status monitoring result.
[0014] In a third aspect, the technical solution of the present invention provides a terminal, including: A memory, configured to store a traffic forwarding program within the cloud platform; A processor, configured to implement the steps of the traffic forwarding method within the cloud platform as described in any one of the above when executing the traffic forwarding program within the cloud platform.
[0015] Fourthly, the technical solution of the present invention provides a computer-readable storage medium, on which a traffic forwarding program within a cloud platform is stored. When the traffic forwarding program within the cloud platform is executed by a processor, the steps of the traffic forwarding method within the cloud platform as described in any one of the above are implemented.
[0016] A traffic forwarding method, system, terminal, and medium within a cloud platform provided by the present invention have the following beneficial effects compared with the prior art: creating a traffic forwarding gateway cluster including a primary gateway and at least one standby gateway, and issuing the same traffic forwarding rules to each gateway to enable traffic to be forwarded by any gateway. At the same time, the status of the primary gateway is monitored in real time, and the traffic forwarding strategy between each gateway is configured according to the status of the primary gateway. When the primary gateway fails, all traffic directed to the primary gateway is directed to the new primary gateway to ensure that traffic forwarding is not interrupted, thereby ensuring normal link communication between each virtual private network within the cloud platform and improving the robustness of the cloud platform. BRIEF DESCRIPTION OF THE DRAWINGS In order to more clearly illustrate the technical solution of the present invention, the drawings required for description will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0017] Figure 1 It is a schematic flowchart of a traffic forwarding method for a cloud platform provided by an embodiment of the present invention.
[0018] Figure 2 It is a routing configuration link diagram of a traffic transfer gateway during communication between initial multi-tenant virtual private networks.
[0019] Figure 3 It is a routing configuration link diagram of a traffic transfer gateway during communication between multi-tenant virtual private networks after primary-standby switchover.
[0020] Figure 4 It is a schematic block diagram of the structure of a traffic forwarding system within a cloud platform provided by an embodiment of the present invention.
[0021] Figure 5 It is a schematic diagram of the structure of a terminal provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0022] To make the objectives, features, and advantages of the present invention more obvious and understandable, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings in the specific embodiments. Obviously, the embodiments described below are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in this application without creative efforts shall fall within the scope of protection of this application.
[0023] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which this invention belongs. The terms used in the description of the present invention in this specification are only for the purpose of describing specific embodiments, and are not intended to limit the present invention.
[0024] The following explains the key terms that appear in the present invention.
[0025] VPC: Virtual Private Cloud, a virtual private network.
[0026] Prometheus server: The Prometheus server, which is the core component of the Prometheus open-source monitoring system, is a time series database used to collect, store, and query monitoring data. It collects data in the form of metrics and stores them together with relevant labels, which can be used to classify and filter the metrics for easy querying and analysis.
[0027] Traffic forwarding gateway: A device that realizes the network traffic forwarding function within the cloud platform. In a cloud environment, a Virtual Private Cloud (VPC) allows users to build isolated network environments. The communication between multiple VPCs and between a VPC and the external network depends on the traffic forwarding gateway. It is like the "transportation hub" of the network, receiving data packets from different network regions, processing these data packets according to preset rules, and then accurately forwarding them to the target network region to ensure the accuracy and efficiency of data transmission. The traffic forwarding gateway determines the forwarding path of the data packet by receiving and parsing the network layer information (such as IP address) and transport layer information (such as port number) of the data packet, and based on the pre-configured traffic forwarding rules. In a cloud platform, these rules may be set based on multiple factors, such as the source IP address range, destination IP address range, service type, etc. For example, when a data packet is sent from a host within a certain VPC and the destination address is a host within another VPC, the traffic forwarding gateway will, according to the configured rules, determine the path that the data packet should pass through, and then forward the data packet to the corresponding network interface so that it can reach the target host.
[0028] Figure 1Schematic diagram of a cloud platform traffic forwarding method provided by an embodiment of the present invention. Among them, Figure 1 The execution entity can be a traffic forwarding system within a cloud platform. The cloud platform traffic forwarding method provided by the embodiment of the present invention is executed by a computer device. Correspondingly, the traffic forwarding system within the cloud platform runs in the computer device. According to different requirements, the order of the steps in this flowchart can be changed, and some can be omitted.
[0029] As Figure 1 shown, the method includes the following steps.
[0030] S1. Create the primary and standby instances of the traffic forwarding gateway within the cloud platform, and generate a traffic forwarding gateway cluster composed of a primary gateway and at least one standby gateway.
[0031] In this step, the primary and standby instances of the traffic forwarding gateway are created within the cloud platform, and a traffic forwarding gateway cluster is composed of a primary gateway and at least one standby gateway. By changing the traditional single gateway node deployment mode through the traffic forwarding gateway cluster method, the reliability of cloud platform traffic forwarding can be improved through the redundant design of the primary and standby gateways in the cluster. When the primary gateway fails, such as hardware damage or software crash, the standby gateway can take over the work at any time to ensure that traffic forwarding is not interrupted and maintain the normal communication of the links between virtual private networks within the cloud platform.
[0032] S2. Mount the primary gateway and each standby gateway to each virtual private network, and issue the same traffic forwarding rules to the primary gateway and each standby gateway.
[0033] In this step, the primary gateway and each standby gateway are mounted to each virtual private network, and the same traffic forwarding rules are issued to them at the same time, so that traffic can be forwarded according to the unified rules on any gateway. At the same time, subnets are created in the virtual private network and gateway network cards are mounted, which simplifies the network deployment process, facilitates users to flexibly adjust the network configuration according to actual needs, ensures that each gateway follows consistent rules when processing traffic, and guarantees the standardization and stability of data transmission.
[0034] S3. Monitor the status of the primary gateway and each standby gateway.
[0035] In this step, the Prometheus server is used to periodically capture the performance metric data of the main gateway and each standby gateway, and determine whether the gateway fails according to the pre-set detection rules. The performance metric data may include service survival status, throughput, forwarding delay, CPU usage, and memory usage. By monitoring the running status of the gateway in real time, potential problems can be discovered in a timely manner. Specifically, through the monitoring of multi-dimensional performance metrics, once an abnormality occurs in the gateway, such as an abnormal service survival status or a performance metric exceeding the threshold, it can be detected in a timely manner and provide a basis for subsequent traffic adjustment. On the one hand, the gateway can be switched in time when the gateway fails, and on the other hand, the problem of traffic interruption caused by gateway failure can be prevented in advance, ensuring the stable operation of the cloud platform network.
[0036] S4. According to the gateway status monitoring result, configure the forwarding policy of the traffic between the main gateway and each standby gateway.
[0037] In this step, according to the gateway status monitoring result, the forwarding policy of the traffic between the main gateway and each standby gateway is reasonably configured. Specifically, when the main gateway fails, calculate the performance evaluation value of the surviving standby gateways and select the optimal standby gateway to take over; when the main gateway is normal, allocate traffic according to its future status prediction type. This step realizes the intelligent dynamic allocation of traffic. When the main gateway fails, it can quickly switch to the standby gateway with the best performance to ensure the continuous and stable forwarding of traffic; when the main gateway is normal, adjust the traffic allocation in advance according to the prediction result, make full use of the resources of each gateway, optimize the network transmission efficiency, and improve the overall network performance and robustness of the cloud platform.
[0038] The traffic forwarding method in the cloud platform provided in this embodiment creates a traffic forwarding gateway cluster including a main gateway and at least one standby gateway, and issues the same traffic forwarding rules to each gateway to enable traffic to be forwarded by any gateway. At the same time, the status of the main gateway is monitored in real time, and the forwarding policy of the traffic between the gateways is configured according to the status of the main gateway. When the main gateway fails, all the traffic directed to the main gateway is directed to the new main gateway to ensure that the traffic forwarding is not interrupted, thereby ensuring the normal link communication between the virtual private networks in the cloud platform and improving the robustness of the cloud platform.
[0039] Further, as a refinement and extension of the specific implementation manner of the above embodiment, in order to fully illustrate the specific implementation process in this embodiment, another traffic forwarding method in the cloud platform is provided, and this method includes the following steps.
[0040] SS1. Construct a traffic forwarding gateway cluster.
[0041] Create the primary and standby instances of the traffic forwarding gateway in the cloud platform, and generate a traffic forwarding gateway cluster composed of a main gateway and at least one standby gateway. That is to say, the traffic forwarding gateway cluster includes at least two gateways, one of which is the main gateway and the other gateways are standby gateways.
[0042] SS2, Gateway mounting and traffic forwarding rule distribution.
[0043] Mount the primary gateway and each standby gateway to each virtual private network. Specifically, create subnets in the virtual private network and mount the subnets to the virtual routers of the virtual private network; mount the network cards of the subnets with the primary gateway and each standby gateway.
[0044] Distribute the same traffic forwarding rules to the primary gateway and each standby gateway. The traffic forwarding rules are a set of pre-set instruction sets used to guide the flow of traffic within the cloud platform. In this embodiment, the same traffic forwarding rules are distributed to the primary gateway and the standby gateways, and these rules specify how traffic with different source and destination addresses should be forwarded between virtual private networks (VPCs) and between VPCs and external networks. For example, it is specified how traffic from a specific IP segment of a VPC subnet should follow path selection, port mapping, etc. operations when going to another VPC to ensure the security and reliability of data transmission.
[0045] Distributing the same traffic forwarding rules to the primary gateway and the standby gateways is to ensure that regardless of whether the traffic is forwarded through the primary gateway or the standby gateway, it can follow a unified standard and path policy. This avoids traffic forwarding chaos caused by different gateways and ensures the stability and predictability of network communication. Even in the case of primary-standby gateway switching, the traffic can continue to be normally forwarded according to the established rules without data loss or transmission errors.
[0046] SS3, Monitor the status of each gateway.
[0047] Monitoring the status of each gateway includes monitoring the status of the primary gateway and each standby gateway, and specifically includes the following steps.
[0048] SS3.1, Periodically capture the performance metric data of the primary gateway and each standby gateway through the Prometheus server.
[0049] The performance metric data includes service survival status, throughput, forwarding latency, CPU usage rate, and memory usage rate.
[0050] This embodiment uses the Prometheus server to periodically capture the performance indicator data of the gateway. The Prometheus server is connected to the detection port of the gateway to obtain the performance indicator data of the gateway. The Prometheus server includes three aspects: data collection, data storage, and alarm processing. When collecting data, it regularly captures indicator data from the / metrics path of the proxy or target host according to the capture interval defined in its service configuration file. Specifically, the / metrics path of the gateway is used to provide performance indicator data. The captured data is stored in a time series database and sorted and compressed by timestamp to save storage space. At the same time, Prometheus supports shard storage and automatic data cleaning of data. Prometheus periodically calculates and generates alarm indications based on user-defined alarm rules, and is further processed by the alarm manager component. Specifically, when the main gateway and / or the backup gateway fails, a fault alarm is issued.
[0051] SS3.2, based on the performance indicator data of the gateway, determines whether the gateway is faulty according to the pre-configured detection rules.
[0052] SS3.2.1, check whether the service survival status flag is 0, if so, determine that the gateway is faulty.
[0053] Detect the identity of the gateway service survival status. If the identity is 0, it is directly determined that the gateway is faulty. The service survival status identity is the direct basis for judging whether the gateway is working normally. By detecting the identity of the service survival status, the fault can be quickly identified when a serious fault occurs in the gateway, such as sudden hardware damage or system crash that causes the service to stop completely, so as to buy time for subsequent traffic switching and fault handling, and avoid a large amount of traffic being accumulated or lost for a long time due to gateway failure.
[0054] SS3.2.2, detect whether the forwarding delay in a unit time exceeds the maximum forwarding delay threshold. If it does so for N1 consecutive detection cycles, the gateway is considered to be faulty.
[0055] Monitor the forwarding delay of the gateway in unit time. When the forwarding delay exceeds the maximum forwarding delay threshold and remains in this state for N1 consecutive detection cycles, the gateway is considered to be faulty. Forwarding delay is one of the key indicators for measuring gateway performance. By continuously monitoring the forwarding delay and setting the threshold, it is possible to promptly detect problems with reduced forwarding efficiency of the gateway due to network congestion, hardware performance degradation, and other reasons. Making judgments on such potential failures in advance helps to adjust the traffic distribution strategy in advance, ensure the timeliness of data transmission, and avoid the impact of excessive delay on the normal operation of cloud platform services, such as online business freezes and untimely data transmission.
[0056] SS3.2.3, detect whether the CPU usage per unit time exceeds the maximum CPU usage threshold. If it does so for N2 consecutive detection cycles, the gateway is considered to be faulty.
[0057] Monitor the CPU usage rate of the gateway within a unit time. If it exceeds the maximum CPU usage rate threshold and this condition is met for N2 consecutive detection cycles, the gateway is determined to be faulty. The CPU usage rate reflects the data processing load of the gateway. Continuously monitoring the CPU usage rate can promptly detect whether the gateway has exhausted its CPU resources due to reasons such as an overloaded task or an abnormal process. When the CPU usage rate is too high and persists for a period of time, it indicates that the gateway may face a performance bottleneck or even be on the verge of crashing. Discovering such problems in advance can timely adjust the traffic, avoid gateway failures caused by CPU overload, and ensure the stable operation of the cloud platform network.
[0058] SS3.2.4, detect whether the memory usage rate within a unit time exceeds the maximum memory usage rate threshold. If it is the case for N3 consecutive detection cycles, the gateway is determined to be faulty.
[0059] Monitor the memory usage rate of the gateway within a unit time. When the memory usage rate exceeds the maximum memory usage rate threshold and this is the case for N3 consecutive detection cycles, the gateway is determined to be faulty. The memory usage rate is an important indicator to measure the running state of the gateway. Monitoring the memory usage rate can promptly detect whether the gateway has insufficient memory resources due to problems such as memory leaks and cache overflows. When the memory usage rate persists in being too high, the gateway may experience unstable operation, data processing errors, etc. Through the monitoring of the memory usage rate and threshold judgment, these potential faults can be discovered in advance, ensuring the normal operation of the gateway and maintaining the stability and reliability of the cloud platform network.
[0060] It should be noted that the values of N1, N2, and N3 are pre-configured according to needs. The three can be the same or different.
[0061] SS4, configure the forwarding strategy of traffic between the primary gateway and each standby gateway.
[0062] In this embodiment, according to the gateway status monitoring results, the forwarding strategy of traffic between the primary gateway and each standby gateway is configured. Specifically, there are two cases. The first case is when the primary gateway fails. According to the status of each standby gateway, a new primary gateway is selected, and the primary-standby switch is executed. All traffic directed to the faulty primary gateway is redirected to the new primary gateway. The second case is when the primary gateway is in a normal state. According to the prediction of the future state of the primary gateway, traffic distribution is realized between the primary gateway and each standby gateway, avoiding the primary gateway from failing due to excessive load, reducing the gateway switching frequency, and further ensuring the communication stability within the cloud platform.
[0063] For the first case where the primary gateway fails, configure the forwarding strategy of traffic between the primary gateway and each standby gateway, which specifically includes the following steps.
[0064] SS4.101, if the primary gateway fails, calculate the performance evaluation values of each surviving standby gateway through the following formula: Performance evaluation value = First weight * Throughput + Second weight * Forwarding delay + Third weight * CPU usage rate + Fourth weight * Memory usage rate.
[0065] It should be noted that the first weight, second weight, third weight, and fourth weight are pre-configured according to specific requirements. It should also be noted that the performance evaluation value is directly proportional to the throughput and inversely proportional to the forwarding delay, CPU usage rate, and memory usage rate.
[0066] SS4.102, Select the standby gateway with the largest performance evaluation value as the new primary gateway.
[0067] The performance evaluation value in this embodiment is directly proportional to the throughput and inversely proportional to the forwarding delay, CPU usage rate, and memory usage rate. Selecting the standby gateway with the largest performance evaluation value means selecting the grid with the relatively best performance as the new primary gateway, reducing the performance degradation caused by excessive gateway load, improving the transmission efficiency and response speed of the entire cloud platform network, and ensuring the normal operation of the service.
[0068] SS4.103, Route all traffic directed to the failed primary gateway to the new primary gateway.
[0069] When the old primary gateway fails, route all traffic directed to the old primary gateway to the new primary gateway to ensure normal traffic transmission.
[0070] The following takes one primary gateway and one standby gateway as an example for illustration. As Figure 2 and Figure 3 shown, create a traffic forwarding gateway cluster, including two instance traffic forwarding gateways A and B. As Figure 2 shown, at initial configuration, gateway A is the primary node and gateway B is the standby node. As Figure 3 shown, gateway A fails, switch gateway A to the standby node, and gateway B to the primary node.
[0071] Under the two virtual private networks VPC1 (connected network segment 192.168.1.0 / 24) and VPC2 (connected network segment 192.168.2.0 / 24) of tenant 1, create connection subnets 11.0.103.0 / 24 respectively, associate the gateways to their respective virtual routers, associate eth1(11.0.103.100) and eth2(11.0.103.101) to traffic forwarding gateway A, and associate eth5(11.0.103.200) and eth6(11.0.103.201) to traffic forwarding gateway B.
[0072] Under the two Virtual Private Networks (VPCs), namely VPC3 (connected network segment 192.168.3.0 / 24) and VPC4 (connected network segment 192.168.4.0 / 24) of Tenant 2, connection subnets 11.0.103.0 / 24 are created respectively, and the gateways are associated with their respective virtual routers. For traffic forwarding gateway A, eth3 (11.0.103.102) and eth4 (11.0.103.103) are associated, and for traffic forwarding gateway B, eth7 (11.0.103.202) and eth8 (11.0.103.203) are associated.
[0073] When configuring the communication between the above four connected network segments, a primary and standby cluster composed of traffic forwarding gateways A and B is used for traffic forwarding. Taking the Figure 2 configuration as the initial configuration, gateway A is the primary node and B is the standby node. The four Virtual Private Networks respectively configure routing rules in their respective virtual routers, where the routing destination network segments are other connected network segments, and the next-hop address points to the network card IP address mounted on gateway A. At the same time, forwarding routes are configured in traffic forwarding gateways A and B, with the destination network segments being the connected network segments of each VPC, and the next-hop being the network card of each VPC associated with the traffic forwarding gateway.
[0074] When the status of gateway A is detected to be abnormal in the traffic forwarding gateway cluster, a primary and standby switch is triggered. The specific logic is to change the route in the virtual router of each VPC whose next-hop address points to gateway A to point to gateway B. Since the routing rules configured in gateway B are the same as those in gateway A, when the virtual router forwards traffic to gateway B, it can act the same as A and forward the traffic to the network card mounted on the correct VPC. The switched traffic path is as Figure 3 shown.
[0075] The above process ensures that when a node in the traffic forwarding gateway cluster fails, the end-to-end communication link will not be interrupted.
[0076] In the second case where the primary gateway is normal, configure the forwarding strategy of traffic between the primary gateway and each standby gateway, which specifically includes the following steps.
[0077] SS4.201, if the primary gateway has no fault, screen the captured performance index data to obtain the target performance index data; the target performance index data includes throughput, forwarding delay, CPU usage rate, and memory usage rate.
[0078] SS4.202, preprocess the target performance index data and construct the preprocessed target performance index data into an input vector.
[0079] SS4.203, transmit the input vector to a pre-trained gateway status prediction model for processing to obtain the predicted type of the future status of the primary gateway.
[0080] SS4.204. Configure the forwarding policy of traffic between the primary gateway and each standby gateway according to the predicted type of the future state of the primary gateway.
[0081] SS4.204.1. Obtain the pre-configured proportional allocation mapping table.
[0082] SS4.204.2. According to the predicted type of the future state of the primary gateway, match the allocation ratio of the primary gateway and the sum of the allocation ratios of all standby gateways from the proportional allocation mapping table.
[0083] SS4.204.3. Direct the traffic of the virtual private network to the primary gateway according to the corresponding allocation ratio of the primary gateway in accordance with the matched allocation ratio.
[0084] SS4.204.4. Perform proportional allocation of standby gateways on the sum of the allocation ratios of all standby gateways according to the performance evaluation values of each surviving standby gateway. The higher the performance evaluation value, the larger the allocated ratio.
[0085] SS4.204.5. Direct the remaining traffic of the virtual private network to the standby gateways according to the corresponding allocation ratios of the standby gateways.
[0086] It should be noted that the remaining traffic refers to the total traffic minus the traffic directed to the primary gateway, and the remaining traffic is allocated by the standby gateways.
[0087] In this embodiment, the throughput, forwarding delay, CPU usage rate, and memory usage rate are used as target performance index data. The future state of the primary gateway is predicted through a gateway state prediction model. The gateway state prediction model processes the target performance index data to obtain a future prediction state classification. If the future state is good, the traffic directed to the primary gateway is maintained or increased. If the future state is poor, the traffic directed to the primary gateway is decreased, and the remaining traffic is directed to the standby gateways. The performance of each standby gateway is detected and analyzed. The standby gateway with better performance obtains more remaining traffic, and the standby gateway with poor performance obtains less or no remaining traffic.
[0088] The gateway state prediction model is constructed based on the recurrent neural network (RNN) and its variant, the long short-term memory network (LSTM), and includes an input layer, an LSTM layer, a fully connected layer, and an output layer.
[0089] Input Layer: Responsible for receiving the input vector constructed from the preprocessed target performance metric data. Since the target performance metric data includes throughput, forwarding latency, CPU usage, and memory usage, the number of neurons in the input layer is 4, corresponding to these 4 performance metrics. The dimension of the input vector is [batch_size, sequence_length, 4], where batch_size represents the number of samples input to the model each time, and sequence_length represents the length of the time series (i.e., the time span of historical data).
[0090] LSTM Layer: LSTM is an improved model of RNN that can effectively handle the long-term dependence problem in time series data. Set one layer of LSTM, and the number of neurons can be adjusted according to the actual situation, such as set to 64. The LSTM layer receives the data from the input layer and processes the sequence data through structures such as memory units to capture the time dependence relationships and features in the data.
[0091] Fully Connected Layer: Further processes the output of the LSTM layer. The number of neurons is determined according to the number of types of main gateway future state predictions. Exemplarily, if the types of main gateway future state predictions are divided into normal, possible failure, and failure, then the number of neurons in the fully connected layer is 3. This layer maps the feature vector output by the LSTM layer to the dimension of the prediction types.
[0092] Output Layer: Uses the softmax activation function to convert the output of the fully connected layer into a probability distribution, obtaining the probability values of the main gateway being in different future state prediction types. The output dimension is [batch_size, 3], corresponding to the probabilities of the normal, possible failure, and failure states respectively.
[0093] In this embodiment, the future state of the primary gateway is predicted through the gateway state prediction model, and the traffic allocation strategy is adjusted in advance according to the prediction result, so as to avoid adjusting the traffic only when the primary gateway has performance problems, ensure that network traffic can always be transmitted efficiently and stably, and reduce the risk of service interruption caused by network congestion or gateway performance degradation. At the same time, traffic is allocated according to the future state of the primary gateway and the performance evaluation value of the standby gateway, which can make more reasonable use of network resources, direct more traffic to the gateway with better performance, give full play to the advantages of each gateway, improve the throughput and response speed of the entire network, and reduce network latency. The traffic is allocated to multiple standby gateways to achieve traffic dispersion and redundancy. When the primary gateway fails, the standby gateway can quickly take over the traffic forwarding task to ensure the continuity of network services. At the same time, by dynamically adjusting the traffic allocation, network failures caused by performance bottlenecks of a certain gateway can be detected and avoided in time, enhancing the reliability and stability of the network. In addition, the pre-configured proportional allocation mapping table can be adjusted according to different network environments and service requirements. When the network traffic changes and the business develops, the mapping table can be flexibly modified to make the traffic allocation strategy adapt to the new network conditions and improve the adaptability and scalability of the network.
[0094] In the above text, an embodiment of a traffic forwarding method in a cloud platform is described in detail. Based on the traffic forwarding method in the cloud platform described in the above embodiment, an embodiment of the present invention also provides a traffic forwarding system in the cloud platform corresponding to this method.
[0095] Figure 4 FIG. is a schematic block diagram of the structure of a traffic forwarding system in a cloud platform provided by an embodiment of the present invention. In this embodiment, the traffic forwarding system 400 in the cloud platform can be divided into multiple functional modules according to the functions it performs, such as Figure 4 shown. The module referred to in the present invention means a series of computer program segments that can be executed by at least one processor and can complete fixed functions, and are stored in the memory.
[0096] The gateway cluster creation module 410 is used to create the primary and standby instances of the traffic forwarding gateway in the cloud platform, and generate a traffic forwarding gateway cluster composed of the primary gateway and at least one standby gateway.
[0097] The gateway mounting module 420 is used to mount the primary gateway and each standby gateway to each virtual private network, and send the same traffic forwarding rules to the primary gateway and each standby gateway.
[0098] The gateway state monitoring module 430 is used to monitor the states of the primary gateway and each standby gateway.
[0099] The traffic forwarding execution module 440 is used to configure the traffic forwarding strategy between the primary gateway and each standby gateway according to the gateway state monitoring result.
[0100] In an optional implementation, the gateway status monitoring module 430 monitors the status of the primary gateway and each standby gateway, specifically including: periodically fetching the performance metric data of the primary gateway and each standby gateway through the Prometheus server; the performance metric data includes service survival status, throughput, forwarding delay, CPU usage rate, and memory usage rate; based on the performance metric data of the gateway, it is judged whether the gateway fails according to the pre-configured detection rules.
[0101] In an optional implementation, the traffic forwarding execution module 440 configures the traffic forwarding policy between the primary gateway and each standby gateway according to the gateway status monitoring result, specifically including: if the primary gateway fails, calculate the performance evaluation value of each surviving standby gateway through the following formula Performance evaluation value = first weight * throughput + second weight * forwarding delay + third weight * CPU usage rate + fourth weight * memory usage rate; Select the standby gateway with the largest performance evaluation value as the new primary gateway; Direct all traffic directed to the failed primary gateway to the new primary gateway.
[0102] In an optional implementation, the gateway status monitoring module 430 judges whether the gateway fails according to the pre-configured detection rules, specifically including: detecting whether the identifier of the service survival status is 0, if so, it is determined that the gateway fails; detecting whether the forwarding delay within a unit time exceeds the maximum forwarding delay threshold, if it is the case for consecutive N detection cycles, it is determined that the gateway fails; detecting whether the CPU usage rate within a unit time exceeds the maximum CPU usage rate threshold, if it is the case for consecutive N detection cycles, it is determined that the gateway fails; detecting whether the memory usage rate within a unit time exceeds the maximum memory usage rate threshold, if it is the case for consecutive N detection cycles, it is determined that the gateway fails.
[0103] In an optional implementation, the traffic forwarding execution module 440 configures the traffic forwarding policy between the primary gateway and each standby gateway according to the gateway status monitoring result, and specifically further includes: if the primary gateway is fault-free, screen the fetched performance metric data to obtain the target performance metric data; the target performance metric data includes throughput, forwarding delay, CPU usage rate, and memory usage rate; preprocess the target performance metric data, and construct the preprocessed target performance metric data into an input vector; transmit the input vector to the pre-trained gateway status prediction model for processing to obtain the predicted type of the future state of the primary gateway; configure the traffic forwarding policy between the primary gateway and each standby gateway according to the predicted type of the future state of the primary gateway.
[0104] In an alternative embodiment, the traffic forwarding execution module 440 configures the traffic forwarding policy between the primary gateway and each standby gateway according to the predicted type of the future state of the primary gateway, specifically including: obtaining a pre-configured proportional allocation mapping table; matching the allocation ratio of the primary gateway and the sum of the allocation ratios of all standby gateways from the proportional allocation mapping table according to the predicted type of the future state of the primary gateway; guiding the traffic of the virtual private network to the primary gateway according to the allocated ratio of the primary gateway; performing proportional allocation of the standby gateways on the sum of the allocation ratios of all standby gateways according to the performance evaluation values of each surviving standby gateway, where the higher the performance evaluation value, the larger the allocated ratio; and guiding the remaining traffic of the virtual private network to the standby gateways according to the corresponding allocated ratios of the standby gateways.
[0105] In an alternative embodiment, the gateway status monitoring module 430 is further configured to issue a fault alarm when the primary gateway and / or a standby gateway fails.
[0106] In an alternative embodiment, the gateway mounting module 420 mounts the primary gateway and each standby gateway to each virtual private network and sends the same traffic forwarding rules to the primary gateway and each standby gateway, specifically including: creating a subnet in the virtual private network and hanging the subnet on the virtual router of the virtual private network; and mounting the network cards of the subnet to the primary gateway and each standby gateway.
[0107] The traffic forwarding system within the cloud platform in this embodiment is used to implement the foregoing traffic forwarding method within the cloud platform. Therefore, the specific implementation manners in this system can be seen in the embodiment part of the traffic forwarding method within the cloud platform in the foregoing text. Therefore, its specific implementation manners can refer to the descriptions of the corresponding parts of each embodiment and will not be elaborated here.
[0108] In addition, since the traffic forwarding system within the cloud platform in this embodiment is used to implement the foregoing traffic forwarding method within the cloud platform, its functions correspond to those of the above method and will not be elaborated here.
[0109] Figure 5 The following is a schematic structural diagram of a terminal 500 provided by an embodiment of the present invention, including: a processor 510, a memory 520, and a communication unit 530. When the processor 510 implements the traffic forwarding program within the cloud platform saved in the memory 520, the following steps are implemented: Create a primary and standby instance of the traffic forwarding gateway within the cloud platform, and generate a traffic forwarding gateway cluster composed of a primary gateway and at least one standby gateway; Mount the primary gateway and each standby gateway to each virtual private network, and send the same traffic forwarding rules to the primary gateway and each standby gateway; Monitor the status of the primary gateway and each standby gateway; Configure the forwarding policy of traffic between the primary gateway and each standby gateway according to the monitoring result of the gateway status.
[0110] The present invention also provides a computer storage medium, which can be a magnetic disk, an optical disk, a read-only memory (ROM for short), a random access memory (RAM for short), etc.
[0111] The computer storage medium stores a traffic forwarding program in the cloud platform. When the traffic forwarding program in the cloud platform is executed by a processor, the following steps are implemented: Create the primary and standby instances of the traffic forwarding gateway in the cloud platform, and generate a traffic forwarding gateway cluster composed of the primary gateway and at least one standby gateway; Mount the primary gateway and each standby gateway to each virtual private network, and issue the same traffic forwarding rules to the primary gateway and each standby gateway; Monitor the status of the primary gateway and each standby gateway; Configure the forwarding policy of traffic between the primary gateway and each standby gateway according to the monitoring result of the gateway status.
[0112] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present invention. Various modifications to these embodiments will be obvious to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to the embodiments shown herein, but will be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A method for forwarding traffic within a cloud platform, characterized in that: The following steps are involved: Create primary and backup instances of the traffic forwarding gateway in the cloud platform, generate a primary gateway and at least one backup gateway to form a traffic forwarding gateway cluster; Mount the primary gateway and each backup gateway to each virtual private network, and issue the same traffic forwarding rules to the primary gateway and each backup gateway; Monitor the status of the main gateway and each backup gateway; Configure the traffic forwarding strategy between the primary gateway and each backup gateway based on the gateway status monitoring results.
2. The method for forwarding traffic within a cloud platform according to claim 1, characterized in that: Monitor the status of the main gateway and each backup gateway, including: The Prometheus server is used to periodically capture the performance indicator data of the main gateway and each backup gateway; the performance indicator data includes service survival status, throughput, forwarding delay, CPU usage, and memory usage; Based on the performance indicator data of the gateway, determine whether the gateway is faulty according to the pre-configured detection rules; Based on the gateway status monitoring results, configure the traffic forwarding strategy between the primary gateway and each backup gateway, including: If the main gateway fails, the performance evaluation value of each surviving backup gateway is calculated by the following formula: Performance evaluation value = first weight * throughput + second weight * forwarding delay + third weight * CPU usage + fourth weight * memory usage; Select the backup gateway with the largest performance evaluation value as the new primary gateway; Redirect all traffic directed to the failed primary gateway to the new primary gateway.
3. The method for forwarding traffic within a cloud platform according to claim 2, characterized in that: Determine whether the gateway is faulty based on pre-configured detection rules, including: Check whether the service survival status flag is 0. If so, it is determined that the gateway is faulty. Check whether the forwarding delay in a unit time exceeds the maximum forwarding delay threshold. If it does so for N1 consecutive detection cycles, the gateway is considered to be faulty. Check whether the CPU usage in a unit time exceeds the maximum CPU usage threshold. If it exceeds the maximum CPU usage threshold for N2 consecutive detection cycles, the gateway is considered to be faulty. Check whether the memory usage per unit time exceeds the maximum memory usage threshold. If this is the case for N3 consecutive detection cycles, the gateway is considered to be faulty.
4. The method for forwarding traffic within a cloud platform according to claim 2 or 3, characterized in that: According to the monitoring results of the gateway status, configure the traffic forwarding strategy between the primary gateway and each backup gateway, including: If the main gateway is not faulty, filter the captured performance indicator data to obtain target performance indicator data; the target performance indicator data includes throughput, forwarding delay, CPU usage, and memory usage; Preprocessing the target performance indicator data, and constructing the preprocessed target performance indicator data as an input vector; The input vector is transmitted to the pre-trained gateway state prediction model for processing to obtain the prediction type of the future state of the main gateway; Configure the traffic forwarding strategy between the primary gateway and each backup gateway based on the predicted future status type of the primary gateway.
5. The method for forwarding traffic within a cloud platform according to claim 4, characterized in that: Configure the traffic forwarding strategy between the primary gateway and each backup gateway based on the predicted future status type of the primary gateway, including: Get a pre-configured proportional allocation map; According to the predicted type of the future state of the main gateway, the main gateway allocation ratio and the sum of the allocation ratios of all backup gateways are matched from the proportion allocation mapping table; According to the matched allocation ratio, the traffic of the virtual private network is directed to the main gateway according to the corresponding main gateway allocation ratio; According to the performance evaluation values of each surviving backup gateway, the backup gateway ratio is allocated based on the sum of the allocation ratios of all backup gateways. The higher the performance evaluation value, the greater the allocation ratio. The remaining traffic of the virtual private network is directed to the backup gateway according to the corresponding backup gateway allocation ratio.
6. The method for forwarding traffic within a cloud platform according to claim 5, characterized in that: The method further comprises the following steps: When the main gateway and / or backup gateway fails, a fault alarm is issued.
7. The method for forwarding traffic within a cloud platform according to claim 6, characterized in that: Mount the primary gateway and each backup gateway to each virtual private network, including: Create a subnet in the virtual private network and attach the subnet to the virtual router of the virtual private network; Mount the network card of the subnet on the primary gateway and each backup gateway.
8. A traffic forwarding system within a cloud platform, characterized in that: include: A gateway cluster creation module is used to create a primary and backup instance of a traffic forwarding gateway in a cloud platform, and generate a primary gateway and at least one backup gateway to form a traffic forwarding gateway cluster; The gateway mounting module is used to mount the main gateway and each backup gateway to each virtual private network, and issue the same traffic forwarding rules to the main gateway and each backup gateway; The gateway status monitoring module is used to monitor the status of the main gateway and each backup gateway; The traffic forwarding execution module is used to configure the traffic forwarding strategy between the main gateway and each backup gateway according to the gateway status monitoring result.
9. A terminal, characterized in that: include: Storage, used to store traffic forwarding programs within the cloud platform; A processor, used to implement the steps of the method for forwarding traffic within the cloud platform as described in any one of claims 1 to 7 when executing the traffic forwarding program within the cloud platform.
10. A computer-readable storage medium, characterized in that: The readable storage medium stores a cloud platform traffic forwarding program, and when the cloud platform traffic forwarding program is executed by the processor, the steps of the cloud platform traffic forwarding method as described in any one of claims 1 to 7 are implemented.