Load balancing method for Docker virtual service network

By monitoring container performance and health status and dynamically adjusting traffic allocation and container resources, the problem of load balancers being unable to distinguish container processing capabilities is solved, achieving faster response and higher stability Docker virtual service network load balancing.

CN120956732APending Publication Date: 2025-11-14CHINA NAT BUILDING MATERIALS TECH CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411720653.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-11-27
Publication Date
2025-11-14

AI Technical Summary

Technical Problem

In existing technologies, load balancers cannot effectively differentiate the processing capabilities of containers, resulting in user requests being assigned to containers with slower response times, which affects the user experience.

Method used

By monitoring the real-time performance metrics and health status of containers, traffic allocation strategies are dynamically adjusted, the number of containers and resources are automatically adjusted, response time optimization is combined, an automatic recovery mechanism is set up, and rescheduling and security auditing are performed in the event of container failure.

Benefits of technology

It enables the reasonable allocation of requests to the best-performing container, improving system performance and stability, ensuring user experience and security, and providing real-time alerts and self-healing capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120956732A_ABST
    Figure CN120956732A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of network management, in particular to a load balancing method for a Docker virtual service network. The method comprises the following steps that real-time performance indexes of services are monitored and collected, the health state of a container is checked regularly, and an optimal flow distribution strategy is obtained according to the health state and the real-time performance indexes of the container; automatically adjusting the number of the containers and the resource requests and limits of the containers according to a flow distribution strategy and the health states of the containers in combination with the real-time load condition; when a fault of the container is detected, automatically restarting or rescheduling the container in combination with a flow distribution strategy; and finally executing a flow distribution strategy, and distributing the request to a proper container. According to the method, the container which responds faster can bear more requests, the situation that a certain container or some containers bear too much pressure due to too short response time is avoided, and the fitness of the target node can be evaluated more comprehensively.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of network management technology, and more specifically, to a Docker virtual service network load balancing method. Background Technology

[0002] With the widespread adoption of cloud computing and microservice architectures, container technologies such as Docker have been widely used in enterprises of all sizes due to their lightweight and highly portable characteristics. However, as business volume grows, a single container can hardly meet the demands of high-concurrency access, which has led to an increased demand for efficient and intelligent load balancing solutions. Against this backdrop, a load balancing method for Docker virtual service networks has emerged. This method can not only dynamically and intelligently allocate traffic based on real-time service performance metrics (such as CPU utilization, memory usage, and network latency) and container health status, but also handle sudden access pressures by automatically scaling the number of containers and adjusting resource quotas, while providing security auditing functions to enhance system stability and security. However, existing load balancers cannot effectively differentiate the processing capabilities of containers, which may result in user requests being assigned to slower-responding containers, thus affecting user experience. Therefore, this paper designs a Docker virtual service network load balancing method. Summary of the Invention

[0003] The purpose of this invention is to provide a Dockker virtual service network load balancing method to solve the problem mentioned in the background art that the load balancer cannot effectively distinguish the processing capacity of containers, which may lead to user requests being assigned to containers with slower response times, thereby affecting the user experience.

[0004] To achieve the above objectives, the present invention aims to provide a Dockker virtual service network load balancing method, comprising the following steps:

[0005] S1. Monitor and collect real-time performance metrics of the service and store them in a time-series database. Regularly check the health status of containers and obtain the optimal traffic allocation strategy based on the health status of containers and real-time performance metrics. In the process of obtaining the optimal traffic allocation strategy, the factors affecting response time are introduced for optimization. At the same time, alarm rules are set in advance, and alarm notifications are automatically triggered when performance metrics exceed preset thresholds.

[0006] S2. Based on the traffic allocation strategy, the health status of the containers, and the real-time load, automatically adjust the number of containers and the resource requests and limits of the containers.

[0007] S3. Set up an automatic recovery mechanism. When a container failure is detected, automatically restart or reschedule the container in conjunction with the traffic allocation strategy. Introduce network latency for optimization during the rescheduling process.

[0008] S4. Execute the traffic distribution strategy to distribute requests to appropriate containers, and configure security audit logs to record key operations and events during the traffic distribution process.

[0009] As a further improvement to this technical solution, the performance indicators in S1 include at least CPU utilization, memory utilization, network speed, response time, and error rate.

[0010] As a further improvement to this technical solution, in step S1, the health status of the container is checked, specifically as follows:

[0011]

[0012] Among them, H i The health score for the i-th container; w1 is the weight of CPU utilization; CPU max This represents the maximum CPU utilization of all containers in the system; CPU i w1 represents the CPU utilization of the i-th container; w2 represents the weight of the memory utilization; Mem represents the memory utilization. max Mem represents the maximum memory usage of all containers in the system. i w3 represents the memory usage of the i-th container; w3 is the weight of the network speed; Net avg Net is the average network speed of all containers in the system. i w4 represents the network input or output rate of the i-th container; w4 is the weight of the response time; RespTime avg RespTime is the average response time of all containers in the system. i is the average response time of the i-th container; w5 is the error rate weight; ErrRate i Let be the error rate of the i-th container.

[0013] As a further improvement to this technical solution, in step S1, the optimal traffic allocation strategy is obtained based on the container's health status and real-time performance indicators, as follows:

[0014]

[0015] Among them, Req i H represents the number of requests allocated to the i-th container. j The health score for the j-th container; TotalRequests is the total number of requests; ∑H j The sum of the health scores for all containers.

[0016] As a further improvement to this technical solution, in step S1, the influencing factor of response time is introduced for optimization during the process of obtaining the optimal traffic allocation strategy. The optimized result is as follows:

[0017]

[0018] Where, f(H) i R i R is an adjustment function that combines health score and response time; i Let be the average response time of the i-th container; The average response time for all containers; α is the adjustment factor.

[0019]

[0020] Among them, Req i ′ represents the optimized number of requests allocated to the i-th container.

[0021] As a further improvement to this technical solution, in step S1, after triggering an alarm notification, alarm information is generated and transmitted to staff via SMS, email, and communication tools.

[0022] As a further improvement to this technical solution, in step S3, the container is automatically restarted or rescheduled in conjunction with the traffic allocation strategy, as follows:

[0023] Regularly check the health status of containers. If the health status of a container is not up to standard, trigger a restart operation. If restarting the container is ineffective or the host on which the container is located has a problem, trigger a rescheduling operation to migrate the container to another healthy node.

[0024] As a further improvement to this technical solution, in S3, the principles for rescheduling containers include balanced utilization of node resources after migration, minimizing migration overhead, data consistency, and service continuity; at the same time, fault tolerance mechanisms and rollback schemes are provided during the migration process.

[0025] As a further improvement to this technical solution, the container is migrated to other healthy nodes, as follows:

[0026] S k =w6·(1-R) k )+w7·H k +w8·(1-L k )+w9·A k ;

[0027] Among them, S k w6 is the fitness score for the k-th node; w6 is the weight of the resource utilization score; R kw7 represents the resource utilization rate of node k; w7 represents the weight of the health status score; H k A health status score for node k; w8 is the weight of the load score; L k The load score for node k; w9 is the weight of the availability score; A k Rate the availability of node k.

[0028] As a further improvement to this technical solution, in step S3, network latency is introduced for optimization during the container rescheduling process. The optimization is as follows:

[0029]

[0030] Among them, S k ' is the fitness score of the k-th node after optimization; p6 is the power of resource utilization; p7 is the power of health status; p8 is the power of load; p9 is the power of availability; p10 is the power of network latency; N k Rate the network latency of node k; w 10 The weights for network latency scoring.

[0031] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0032] This Docker virtual service network load balancing method enables faster-responding containers to handle more requests, avoiding excessive pressure on one or more containers due to insufficient response time. It also allows for a more comprehensive assessment of the suitability of target nodes, thereby helping system administrators or automation tools make better container migration decisions. Attached Figure Description

[0033] Figure 1 This is a flowchart of the overall method of the present invention. Detailed Implementation

[0034] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of the present invention.

[0035] Example

[0036] Please see Figure 1 As shown, a Docker virtual service network load balancing method is provided, including the following steps:

[0037] Docker virtual services refer to various services provided within a Docker environment using containerization technology. These services can be web applications, databases, caching services, message queues, etc. Docker, through containerization, enables these services to run consistently across different environments, improving development and deployment efficiency. In this embodiment, the application runs in a Kubernetes cluster, and load balancing is achieved by defining Kubernetes Service objects. A Service object can abstract a group of containers into a logical collection and provide a stable network endpoint for this group of containers.

[0038] S1. Monitor and collect real-time performance metrics of the service and store them in a time-series database. Regularly check the health status of containers and obtain the optimal traffic allocation strategy based on the health status of containers and real-time performance metrics. In the process of obtaining the optimal traffic allocation strategy, the factors affecting response time are introduced for optimization. At the same time, alarm rules are set in advance, and alarm notifications are automatically triggered when performance metrics exceed preset thresholds.

[0039] The performance metrics in S1 include at least CPU utilization, memory utilization, network speed, response time, and error rate;

[0040] In S1, checking the health status of a container first requires defining a set of health check criteria. These criteria can be based on performance metrics such as CPU utilization, memory utilization, and network connectivity. For example, if a container's CPU utilization exceeds 90%, or it fails to respond to heartbeat checks three times consecutively, the container is considered unhealthy; specifically as follows:

[0041]

[0042] Among them, H i The health score for the i-th container; w1 is the weight of CPU utilization; CPU max This represents the maximum CPU utilization of all containers in the system; CPU i w1 represents the CPU utilization of the i-th container; w2 represents the weight of the memory utilization; Mem represents the memory utilization. max Mem represents the maximum memory usage of all containers in the system. i w3 represents the memory usage of the i-th container; w3 is the weight of the network speed; Net avg Net is the average network speed of all containers in the system. i w4 represents the network input or output rate of the i-th container; w4 is the weight of the response time; RespTime avg RespTime is the average response time of all containers in the system. iis the average response time of the i-th container; w5 is the error rate weight; ErrRate i Let be the error rate of the i-th container;

[0043] This formula calculates a health score for each container by comprehensively considering multiple performance metrics (such as CPU utilization, memory utilization, network speed, response time, and error rate) and assigning different weights to each metric. A higher health score indicates better container performance and a greater suitability for handling more requests. This method quantifies the health status of each container, providing a basis for subsequent traffic allocation.

[0044] In S1, the optimal traffic allocation strategy is obtained based on the container's health status and real-time performance metrics, as follows:

[0045]

[0046] Among them, Req i H represents the number of requests allocated to the i-th container. j The health score for the j-th container; TotalRequests is the total number of requests; ∑H j The sum of the health scores for all containers;

[0047] This formula allocates the total number of requests proportionally based on each container's health score. Specifically, the number of requests received by each container is directly proportional to its health score. This means that containers with higher health scores will receive more requests, while containers with lower health scores will receive fewer requests. This dynamic traffic allocation method ensures that requests are rationally distributed to the best-performing containers, thereby improving the overall performance and stability of the system.

[0048] In S1, the response time factor is introduced for optimization during the process of obtaining the optimal traffic allocation strategy, because a short response time usually means high service efficiency and a good user experience. The specific optimization is as follows:

[0049]

[0050] Where, f(H) i R i R is an adjustment function that combines health score and response time; i Let be the average response time of the i-th container; The average response time of all containers is used to standardize the differences in response time among containers; α is an adjustment factor used to control the degree of influence of response time on the final weight.

[0051] The reciprocal of the response time It is added to the calculation as a positive influencing factor. However, directly using... This could lead to containers with very short response times having excessively high weights, resulting in an unbalanced load. Therefore, response times can be normalized, or other transformation functions can be used to ensure a more reasonable weight distribution.

[0052]

[0053] Among them, Req i ′ represents the optimized number of requests allocated to the i-th container;

[0054] This optimization retains the role of the original health score while further optimizing request allocation through response time. This allows containers with faster response times to handle more requests, while avoiding excessive pressure on one or more containers due to excessively short response times.

[0055] After an alarm notification is triggered in S1, alarm information is generated and transmitted to staff via SMS, email and communication tools to ensure that relevant personnel can receive alarm information in a timely manner. Monitoring and alarms are integrated throughout the entire solution, providing real-time data support and fault response mechanisms for intelligent routing, self-healing capabilities, dynamic scaling and zero-downtime updates.

[0056] S2. Based on traffic allocation strategies, container health status, and real-time load conditions, automatically adjust the number of containers and their resource requests and limits. Dynamic scaling relies on monitoring data in intelligent routing and self-healing capabilities to ensure that the system can automatically adjust resources according to real-time load conditions, improving response speed and reducing costs.

[0057] S3. Set up an automatic recovery mechanism. When a container failure is detected, automatically restart or reschedule the container in conjunction with the traffic allocation strategy. Introduce network latency for optimization during the rescheduling process.

[0058] In S3, containers are automatically restarted or rescheduled based on traffic allocation strategies, as follows:

[0059] Regularly check the health status of containers. If the health status of a container is not up to standard (e.g., multiple consecutive health checks fail), trigger a restart operation. If restarting the container is ineffective or the host where the container is located has a problem (e.g., insufficient host resources or host failure), trigger a rescheduling operation to migrate the container to another healthy node.

[0060] In general, restarting a container is suitable when the container is only experiencing a temporary problem. By periodically checking the health status of containers, if an unhealthy container is found, restarting it can be attempted to resolve the issue. If the restart is successful, the container will resume normal operation. Rescheduling a container is suitable when restarting fails or when the host machine hosting the container experiences problems. By using orchestration tools such as Docker Swarm or Kubernetes, containers can be migrated to other healthy nodes, ensuring high availability and stability of the service.

[0061] In S3, the principles for rescheduling containers include balanced utilization of node resources after migration, minimizing migration overhead, data consistency, and service continuity. Simultaneously, fault tolerance mechanisms and rollback schemes are provided during the migration process to address unexpected situations, such as backing up the source node's data and configuration. If the migration fails, a quick rollback to the original state can be achieved to avoid prolonged service interruptions.

[0062] The principle of balanced resource utilization ensures that the resources of the migrated nodes are evenly utilized, avoiding situations where one node is overloaded while others are idle. When selecting target nodes, the current resource utilization of the target nodes (such as CPU, memory, network bandwidth, etc.) should be considered. Nodes with lower resource utilization should be prioritized to achieve a balanced distribution of resources.

[0063] The principle of minimizing migration overhead means reducing resource and time costs during the migration process. The migration process itself consumes resources (such as network bandwidth, CPU, and memory), therefore, a target node close to the source node should be selected to reduce network latency and transmission time. Additionally, incremental migration or snapshot techniques can be used to reduce data transfer volume.

[0064] The data consistency principle aims to ensure the consistency and integrity of data during migration, guaranteeing that data is not lost or corrupted. Data synchronization tools (such as rsync and CRIU) can be used to ensure data consistency. For critical applications such as databases, transaction mechanisms can be used to ensure data consistency.

[0065] The principle of service continuity aims to ensure uninterrupted or minimized downtime during migration, reducing the impact on user services. Strategies such as blue-green deployment or rolling updates can be used to gradually switch traffic to the new node for a seamless migration. Furthermore, preparing the target node's environment in advance can shorten migration time.

[0066] Migrate the container to another healthy node as follows:

[0067] S k =w6·(1-R) k)+w7·H k +w8·(1-L k )+w9·A k ;

[0068] Among them, S k w6 is the fitness score for the k-th node; w6 is the weight of the resource utilization score; R k w7 represents the resource utilization rate of node k; w7 represents the weight of the health status score; H k A health status score for node k; w8 is the weight of the load score; L k The load score for node k; w9 is the weight of the availability score; A k Rate the availability of node k;

[0069] In this way, the best target node can be dynamically selected based on the overall score of the node, and the container can be migrated to the most suitable node, thereby improving the overall performance and reliability of the system.

[0070] In S3, network latency is introduced for optimization during container rescheduling. In network-intensive applications, inter-node network latency has a significant impact on application performance. Lower network latency reduces data transmission time, improves application response speed, and thus enhances user experience. The specific optimizations are as follows:

[0071]

[0072] Among them, S k ' is the fitness score of the k-th node after optimization; p6 is the power of resource utilization, used to adjust the influence of resource utilization on the fitness score; p7 is the power of health status, used to adjust the influence of health status on the fitness score; p8 is the power of load, used to adjust the influence of load on the fitness score; p9 is the power of availability, used to adjust the influence of availability on the fitness score; p10 is the power of network latency, used to adjust the influence of network latency on the fitness score; N k Rate the network latency of node k (between 0 and 1), where 0 represents the minimum network latency and 1 represents the maximum network latency; w 10 The weights for network latency scoring;

[0073] The lower the resource utilization, the more remaining resources a node has, making it more suitable for receiving new containers, through (1-R). k ) p6 To enhance or mitigate the impact of resource utilization; the higher the health status score, the healthier the node, and the more suitable it is to receive new containers, through To enhance or mitigate the impact of health status; the lower the load, the less burden on the node, making it more suitable for receiving new containers, through (1-L k ) p8 To enhance or mitigate the impact of load; the higher the availability score, the more stable the node, and the more suitable it is to receive new containers. To enhance or mitigate the impact of availability; lower network latency results in faster communication between nodes, making it more suitable for receiving new containers, through (1-N) k ) p10 To enhance or mitigate the impact of network latency.

[0074] This optimized formula comprehensively considers the impact of five factors: resource utilization, health status, load, availability, and network latency. It provides a more holistic assessment of the suitability of the target node, thereby helping system administrators or automation tools make better container migration decisions. Weights w6, w7, w8, w9, w 10 The adjustment needs to be based on the importance of the specific application scenario. For example, in applications sensitive to network latency, the w value can be appropriately increased. 10 The value of N. At the same time, N... k The calculations also need to be based on the actual network conditions, which may involve testing with other nodes or other network performance tests.

[0075] S4 executes traffic distribution policies, distributing requests to appropriate containers, and configures security audit logs to record key operations and events during the traffic distribution process. This ensures efficient system operation and service quality, while also providing crucial security and compliance guarantees, enabling administrators to track and review all critical activities and promptly identify and respond to potential security threats and abnormal behaviors.

[0076] The foregoing has shown and described the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The embodiments and descriptions in the specification are merely preferred examples and are not intended to limit the invention. Various changes and modifications can be made to the invention without departing from its spirit and scope, and all such changes and modifications fall within the scope of the claimed invention.

Claims

1. A Docker virtual service network load balancing method, characterized in that, Includes the following steps: S1. Monitor and collect real-time performance metrics of the service and store them in a time-series database. Regularly check the health status of containers and obtain the optimal traffic allocation strategy based on the health status of containers and real-time performance metrics. In the process of obtaining the optimal traffic allocation strategy, the factors affecting response time are introduced for optimization. At the same time, alarm rules are set in advance, and alarm notifications are automatically triggered when performance metrics exceed preset thresholds. S2. Based on the traffic allocation strategy, the health status of the containers, and the real-time load situation, automatically adjust the number of containers and the resource requests and limits of the containers. S3. Set up an automatic recovery mechanism. When a container failure is detected, automatically restart or reschedule the container in conjunction with the traffic allocation strategy. Introduce network latency for optimization during the rescheduling process. S4. Execute the traffic distribution strategy to distribute requests to appropriate containers, and configure security audit logs to record key operations and events during the traffic distribution process.

2. The Docker virtual service network load balancing method according to claim 1, characterized in that: The performance metrics in S1 include at least CPU utilization, memory utilization, network speed, response time, and error rate.

3. The Docker virtual service network load balancing method according to claim 2, characterized in that: In step S1, the health status of the container is checked, specifically as follows: Among them, H i The health score for the i-th container; w1 is the weight of CPU utilization; CPU max This represents the maximum CPU utilization of all containers in the system; CPU i w1 represents the CPU utilization of the i-th container; w2 represents the weight of the memory utilization; Mem represents the memory utilization. max Mem represents the maximum memory usage of all containers in the system. i w3 represents the memory usage of the i-th container; w3 is the weight of the network speed; Net avg Net is the average network speed of all containers in the system. i w4 represents the network input or output rate of the i-th container; w4 is the weight of the response time; RespTime avg RespTime is the average response time of all containers in the system. i is the average response time of the i-th container; w5 is the error rate weight; ErrRate i Let be the error rate of the i-th container.

4. The Docker virtual service network load balancing method according to claim 3, characterized in that: In step S1, the optimal traffic allocation strategy is obtained based on the container's health status and real-time performance metrics, as follows: Among them, Req i H represents the number of requests allocated to the i-th container. j The health score for the j-th container; TotalRequests is the total number of requests; ∑H j The sum of the health scores for all containers.

5. A Docker virtual service network load balancing method according to claim 4, characterized in that: In step S1, the factors affecting response time are introduced for optimization during the process of obtaining the optimal traffic allocation strategy. The optimized result is as follows: Where, f(H) i ,R i R is an adjustment function that combines health score and response time; i Let be the average response time of the i-th container; The average response time for all containers; α is the adjustment factor. Among them, Req i ′ represents the optimized number of requests allocated to the i-th container.

6. A Docker virtual service network load balancing method according to claim 5, characterized in that: After triggering an alarm notification in S1, alarm information is generated and transmitted to staff via SMS, email, and communication tools.

7. A Docker virtual service network load balancing method according to claim 6, characterized in that: In S3, the container is automatically restarted or rescheduled in conjunction with the traffic allocation strategy, as follows: Regularly check the health status of containers. If the health status of a container is not up to standard, trigger a restart operation. If restarting the container is ineffective or the host on which the container is located has a problem, trigger a rescheduling operation to migrate the container to another healthy node.

8. A Docker virtual service network load balancing method according to claim 7, characterized in that: In S3, the principles for rescheduling containers include balanced utilization of node resources after migration, minimizing migration overhead, data consistency, and service continuity; at the same time, fault tolerance mechanisms and rollback schemes are provided during the migration process.

9. A Docker virtual service network load balancing method according to claim 8, characterized in that: The container will be migrated to another healthy node, as follows: S k =w6·(1-R k )+W7·H k +W8·(1-L k )+w9·A k ; Among them, S k w6 is the fitness score for the k-th node; w6 is the weight of the resource utilization score; R k w7 represents the resource utilization rate of node k; w7 represents the weight of the health status score; H k A health status score for node k; w8 is the weight of the load score; L k The load score for node k; w9 is the weight of the availability score; A k Rate the availability of node k.

10. A Docker virtual service network load balancing method according to claim 9, characterized in that: In step S3, network latency is introduced for optimization during the container rescheduling process. The optimization is as follows: Among them, S k ' is the fitness score of the k-th node after optimization; p6 is the power of resource utilization; p7 is the power of health status; p8 is the power of load; p9 is the power of availability; p10 is the power of network latency; N k Rate the network latency of node k; w 10 The weights for network latency scoring.

Citation Information

Patent Citations

  • Container cluster delayed shrinkage scheduling method and system

    CN107395735A

  • Load balancing algorithm based on Docker container cluster

    CN107888708A