Load balancing system and load balancing method

By using multi-processors and multi-network card back-end servers in the load balancing system, and combining the intelligent allocation strategy of the load balancing device, the performance bottleneck problem of a single server and network card in high concurrency scenarios is solved, and the system reliability and stability are achieved.

CN119996410AActive Publication Date: 2025-05-13INSPUR SUZHOU INTELLIGENT TECH CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202411364551.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-09-27
Publication Date
2025-05-13
Estimated Expiration
2044-09-27

AI Technical Summary

Technical Problem

In the prior art, a single server device is difficult to bear high concurrency scenarios, and a single network card leads to tight bandwidth resources, affecting network performance.

Method used

Design a load balancing system, including a load balancer and multiple backend servers, each with at least two processors and two network cards. The load balancer forwards requests to the appropriate processor and network card through a preconfigured policy, and forwards requests to other processors through the network card when the processor load is too high.

Benefits of technology

Load balancing at the hardware and software level is realized, reliability and stability in high concurrency scenarios are ensured, and single point of failure and bandwidth bottlenecks are avoided.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119996410A_ABST
    Figure CN119996410A_ABST
Patent Text Reader

Abstract

The invention provides a load balancing system and a load balancing method, the system comprises at least one back-end server and a load balancer, the back-end server comprises at least two processors and at least two network cards, and each processor is in communication connection with the at least two network cards through a system bus; the load balancer is in communication connection with the back-end server, and is used for responding to a first request sent by a client, determining a first back-end server and a first processor of the first back-end server, and forwarding the first request to a first network card corresponding to the first processor; the first back-end server is used for forwarding a second request in the first request to a second processor through the first network card when it is determined that the processing duration of the first processor for processing the first request exceeds the preset duration, and the second processor is a processor in the first back-end server and is connected with the first network card. The load balancing system provided by the invention not only realizes load balancing of a hardware level, but also realizes load balancing of a software level.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer technology, and in particular to a load balancing system and a load balancing method. Background Art

[0002] As network applications become increasingly complex and data volumes grow dramatically, server processing power and computing intensity also increase accordingly, making it increasingly difficult for a single server device to handle the workload. In addition, in order to improve the server's network throughput and data processing efficiency, modern servers generally use high-performance network interfaces (NICs) as network access interfaces. However, in existing technical architectures, servers are often only equipped with a single NIC to handle all network traffic.

[0003] Although a single network card can provide a certain amount of network bandwidth, in high-concurrency scenarios, all network traffic in and out of the server needs to pass through this network card, which can easily lead to bandwidth resource constraints and affect overall network performance. Summary of the invention

[0004] The present invention provides a load balancing system and a load balancing method, which are used to solve the defects in the prior art that a single server device is difficult to handle high-concurrency scenarios, and a single network card in the server will lead to tight bandwidth resources, thereby affecting the overall network performance, and realize load balancing at both the hardware level and the software level to ensure reliability and stability in high-concurrency scenarios.

[0005] The present invention provides a load balancing system, comprising: a load balancer and at least one back-end server connected to the load balancer in communication, wherein the back-end server comprises at least two processors and at least two network cards, and each of the processors is connected to at least two network cards via a system bus; The load balancer is configured to respond to a first request sent by a client, determine a first backend server and a first processor of the first backend server, and forward the first request to a first network card corresponding to the first processor; The first back-end server is also used to forward the second request in the first request to the second processor through the first network card when it is determined that the processing time of the first processor to process the first request exceeds a preset time. The second processor is a processor in the first back-end server and is connected to the first network card.

[0006] According to a load balancing system provided by the present invention, the load balancer is also used for: In response to a first request sent by a client, determine a request type corresponding to the first request, and select a candidate backend server matching the request type from at least one of the backend servers based on a service type of at least one of the backend servers; Based on a preconfigured load balancing policy, a first backend server and a first processor of the first backend server are determined from the candidate backend servers.

[0007] The present invention provides a load balancing system, wherein the load balancer is further used for: Monitoring the status of the first backend server and the first network card; When it is determined that the first network card is in an unavailable state, the first request is forwarded to a second network card, where the second network card includes other network cards in the first backend server except the first network card or network cards of other backend servers except the first backend server.

[0008] The present invention provides a load balancing system, further comprising: a system service processor, wherein the system service processor is communicatively connected with the load balancer and at least one of the back-end servers; The system service processor is used to: Acquire monitoring data of the load balancing system, the monitoring data including a first performance and a first state of the load balancer and a second performance and a second state of each of the processors in each of the backend servers; Based on the monitoring data, the load balancing system is adjusted, and the adjustment includes at least one of an adjustment of the load balancing strategy and an adjustment of the number of backend servers.

[0009] The present invention provides a load balancing system, wherein the system service processor is further used for: Configuring a virtual IP address of the load balancer so that the client can access the virtual IP address; Configuring a backend server list of the load balancer, wherein the backend server list includes an IP address, a port number, and a service type of at least one of the backend servers; Configure the health check protocol and parameters corresponding to the backend servers in the backend server list; The load balancing strategy of the load balancer is configured, wherein the load balancing strategy includes at least one of a polling mechanism, a weight distribution, a minimum number of connections, and a fastest response time.

[0010] The present invention provides a load balancing system, wherein the load balancing strategy comprises: Determine a backend server pointed to by a first pointer among candidate backend servers that match a first request currently sent by the client, and use the backend server pointed to by the first pointer as a first backend server that matches the first request; Determine the processor pointed to by the second pointer in the current first backend server, and use the processor pointed to by the second pointer as the first processor matched by the first request; When the first processor receives the first request, the first pointer points to the next backend server according to a first polling order, and the second pointer points to the next processor in the first backend server according to a second polling order.

[0011] The present invention provides a load balancing system, wherein the load balancing strategy comprises: Determine a backend server corresponding to the least number of connections among the candidate backend servers that match the first request currently sent by the client, and use the backend server corresponding to the least number of connections as the first backend server that matches the first request; A processor corresponding to the minimum number of threads in the current first backend server is determined, and the processor corresponding to the minimum number of threads is used as a first processor matching the first request.

[0012] The present invention provides a load balancing system, wherein the load balancing strategy comprises: Determine a backend server corresponding to the shortest first response time among the candidate backend servers that match the first request currently sent by the client, and use the backend server corresponding to the shortest first response time as the first backend server that matches the first request; A processor corresponding to the shortest second response time in the current first backend server is determined, and the processor corresponding to the shortest second response time is used as a first processor matched with the first request.

[0013] The present invention provides a load balancing system, wherein the first backend server is further used for: When the first processor receives the first request, starting a timer; When it is determined by the timer that the processing time of the first processor for processing the first request reaches a preset time, determining a current processing progress of the first request; If there are still multiple tasks to be processed in the current first request, the multiple tasks to be processed are parsed to determine at least one target task to be processed that can be processed independently among the multiple tasks to be processed; generating a second request based on at least one of the target tasks to be processed; Determine at least one candidate processor other than the first processor in the current first backend server, wherein the candidate processor is connected to the first network card; Determining a second processor from the candidate processors based on a load state of at least one of the candidate processors; Forwarding the second request in the first request to the second processor through the first network card.

[0014] The present invention also provides a load balancing method, which is applied to any of the load balancing systems described above, and comprises: In response to a first request sent by a client, a load balancer is used to determine a first backend server and a first processor of the first backend server, and the first request is forwarded to a first network card corresponding to the first processor; Through the first back-end server, when it is determined that the processing time of the first processor to process the first request exceeds a preset time, the second request in the first request is forwarded to the second processor through the first network card, wherein the second processor is a processor in the first back-end server and is connected to the first network card.

[0015] According to a load balancing method provided, the method further includes: Responding to the first request sent by the client, by the first backend server, to determine the request type corresponding to the first request, and screening out a candidate backend server matching the request type from at least one of the backend servers based on the service type of at least one of the backend servers; Based on a preconfigured load balancing strategy, a first backend server and a first processor of the first backend server are determined from the candidate backend servers, wherein the load balancing strategy includes at least one of a polling mechanism, weight distribution, a minimum number of connections, and a fastest response time.

[0016] According to a load balancing method provided, the method further includes: Monitoring the status of the first backend server and the first network card through the load balancer; When it is determined that the first network card is in an unavailable state, the first request is forwarded to a second network card, where the second network card includes other network cards in the first backend server except the first network card or network cards of other backend servers except the first backend server.

[0017] The present invention also provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, any of the load balancing methods described above is implemented.

[0018] The present invention also provides a non-transitory computer-readable storage medium on which a computer program is stored. When the computer program is executed by a processor, the load balancing method described in any one of the above is implemented.

[0019] The present invention also provides a computer program product, comprising a computer program, wherein when the computer program is executed by a processor, the load balancing method described above is implemented.

[0020] The load balancing system and load balancing method provided by the present invention establish a communication connection between the load balancer and the back-end server, so that the load balancer can determine to which back-end server and which processor of the server the first request of the client is forwarded. Load balancing at the software level is achieved. In addition, each back-end server is equipped with at least two processors and at least two network cards, and each processor is connected to at least two network cards through a system bus. Specifically, when the processing time of the first processor to process the first request exceeds the preset time, part of the second request in the first request can be forwarded to the second processor, so that the request can be transferred between different processors through the network card to avoid overloading of a single processor, and load balancing is further achieved at the hardware level, which enhances the reliability and stability of the system. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0022] Figure 1 This is one of the structural diagrams of the load balancing system provided by the present invention.

[0023] Figure 2 It is a schematic diagram of the network card configuration connection of the back-end server provided by the present invention.

[0024] Figure 3 This is the second structural diagram of the load balancing system provided by the present invention.

[0025] Figure 4 It is a flow chart of the load balancing method provided by the present invention.

[0026] Figure 5 It is a structural schematic diagram of the electronic device provided by the present invention.

[0027] Reference numerals: 10: Load balancer; 11: Server pool; 20: Backend server; 21: Processor; 22: Network card; 30: System service processor. DETAILED DESCRIPTION

[0028] In order to make the purpose, technical solution and advantages of the present invention clearer, the technical solution of the present invention will be clearly and completely described below in conjunction with the drawings of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0029] It should be noted that the load balancing system and load balancing method provided in the embodiments of the present application can be applied to various high-density server clusters, so that based on the load balancing system and load balancing method provided in the embodiments of the present application, both hardware-level load balancing and software-level load balancing are achieved.

[0030] Figure 1 It is one of the structural diagrams of the load balancing system provided by the present invention; Figure 1 As shown, the system includes a load balancer 10, and a server pool 11 of the load balancer 10 includes at least one backend server 20 that is communicatively connected to the load balancer.

[0031] Here, the main function of the load balancer 10 is to distribute the network traffic (such as HTTP requests) entering the system to the back-end servers 20 in the server pool 11, so as to optimize resource utilization, improve response speed and fault tolerance.

[0032] The server pool 11 is a collection of at least one backend server 20, which jointly process requests forwarded by the load balancer 10. The backend servers 20 in the server pool 11 can be physical servers, virtual machines, or containers, etc., which are configured to process the same type of requests (for example, Web servers may all run the same Web application). They can also be configured to process different types of requests. For example, requests that process a large number of computationally intensive tasks may be assigned to backend servers with high-performance CPUs and memories, while requests that process a large number of IO-intensive tasks may be assigned to backend servers with high-speed storage and network interfaces.

[0033] Specifically, each backend server 20 includes at least two processors 21 and at least two network cards 22, and each processor 21 is connected to at least two network cards 22 via a system bus. The network card types of the at least two network cards 22 in each backend server 20 may be the same or different, and there is no limitation on this.

[0034] The processor 21 is the core computing unit of the backend server 20, responsible for executing instructions, processing data and running programs, such as the CPU (Central Processing Unit) and other chips that perform specific processing tasks (such as GPU, DSP, etc.). These processors 21 contain multiple cores (CPU Core), each of which can execute instructions independently, thereby achieving parallel processing.

[0035] The system bus is a channel for transmitting data and information between various components inside the backend server 20. In this embodiment, each processor 21 is connected to at least two network cards 22 through the system bus, that is, the processor 21 can communicate directly with the network card 22 without going through other intermediate devices. The system bus can be of various types, such as PCI (Peripheral Component Interconnect), PCI-X, PCIe, etc.

[0036] It should be noted that, in this embodiment, a multi-path network connection is achieved by respectively leading multiple links from each processor 21 to connect to different network cards 22. When a link or network card fails, the system can continue to work through another link, thereby improving the reliability of the network connection. In addition, the traffic distribution on different links can be dynamically adjusted according to application requirements, providing greater flexibility.

[0037] As reference Figure 3 A backend server includes two processors: CPU0 and CPU1, and two network cards 22: 100G PCIE network card and 100G OCP network card. Here, a X8 PCIE link can be connected from CPU0 and CPU1 to the 100G PCIE network card, and another X8 PCIE link can be connected from CPU0 and CPU1 to the 100G OCP network card. This configuration allows each CPU to directly access two different network cards.

[0038] Furthermore, since different CPLD versions may support different link configurations and bandwidth allocations, in order to support the above PCIE link allocation, the CPLD firmware on the mainboard needs to be updated to a version that supports this configuration, ensuring that the PCIE link can effectively allocate bandwidth according to actual needs and avoid resource conflicts.

[0039] In addition, the network card firmware is the software running on the network card, which is responsible for controlling the hardware functions and network communications of the network card. Therefore, the network card firmware also needs to be upgraded to this configured version to ensure that the network card can correctly identify and process traffic from multiple CPUs.

[0040] Specifically, the load balancer 10 is used to respond to a first request sent by a client, determine a first backend server and a first processor of the first backend server, and forward the first request to a first network card corresponding to the first processor; The client (such as a browser, API client, etc.) sends a first request to the load balancer 10. The request may be an HTTP request, a database query, or other types of network requests. After receiving the request from the client, the load balancer 10 decides how to distribute the first request according to a predefined load balancing strategy.

[0041] Specifically, the load balancer 10 selects a backend server 20 from the server pool 11 to process the first request. The selected backend server 20 is referred to as a first backend server.

[0042] After selecting the first backend server, the load balancer 10 needs to further determine which processor 21 on the server actually processes the request. For example, a processor 21 is selected from the first backend server based on the current load condition, processing capacity or other configured priority rules of the processor 21 on the first backend server. The selected processor 21 is called the first processor.

[0043] It should be understood that in this embodiment, each processor 21 is connected to at least two network cards 22 via a system bus. Therefore, in this embodiment, after selecting the first processor, it is necessary to further select the first network card corresponding to the first processor.

[0044] In one example, the load balancer 10 is also configured with detailed information about each network card 22 on each backend server 20, including the IP address of each network card 22, the processor connected thereto, and related load information. Therefore, in this embodiment, the load balancer 10 will evaluate the current load of each network card of the first processor (such as network traffic, packet processing rate, error rate, and other indicators). In addition, the current load of other processors connected to each network card will be evaluated, and finally the first network card will be selected based on the current load of each network card and the current load of other processors connected to each network card (such as CPU usage, number of processes and threads, etc.). For example, a network card with low network traffic and low error rate, and whether other connected processors have sufficient remaining processing capacity to process requests is selected as the first network card.

[0045] Finally, the load balancer 10 forwards the request to the first network card corresponding to the first processor responsible for processing the first request on the first backend server. Specifically, through a network protocol (such as TCP / IP), the load balancer 10 sends the data packet of the first request to the IP address and port of the first network card.

[0046] Specifically, the first back-end server is also used to forward the second request in the first request to the second processor through the first network card when it is determined that the processing time of the first processor to process the first request exceeds a preset time. The second processor is a processor in the first back-end server and is connected to the first network card.

[0047] When the first backend server receives the first request distributed by the load balancer 10 and determines that the first processor is to process the request, the first backend server monitors the processing process, especially the time required for the first processor to process the first request.

[0048] If during the processing, the first back-end server finds that the time taken by the first processor to process the first request exceeds a preset time (for example, a preset time set based on system performance requirements, user experience standards, or historical data), in order to reduce the burden on the first processor and speed up the processing of the first request, the first back-end server will forward a part of the first request, namely the second request, to the second processor through the first network card.

[0049] Here, the second request may be a subtask in the first request, a fragment that can be executed independently, or a request that is associated with the first request but can be processed in parallel.

[0050] Here, the second processor is another processor on the first backend server, and it is also connected to the first network card. In other words, the first network card serves multiple processors at the same time, providing redundancy and load distribution capabilities at the hardware level.

[0051] In the system provided by this embodiment, the load balancer establishes a communication connection with the back-end server, so that the load balancer can determine to which back-end server and which processor of the back-end server the first request of the client is forwarded. Load balancing at the software level is achieved. In addition, each back-end server is equipped with at least two processors and at least two network cards, and each processor is connected to at least two network cards through a system bus. Specifically, when the processing time of the first processor to process the first request exceeds the preset time, part of the second request in the first request can be forwarded to the second processor, so that the request can be transferred between different processors through the network card to avoid overloading of a single processor, further realizing load balancing at the hardware level, and enhancing the reliability and stability of the system.

[0052] In some embodiments, the load balancer 10 is further configured to: In response to a first request sent by a client, determine a request type corresponding to the first request, and select a candidate backend server matching the request type from at least one of the backend servers 20 based on a service type of at least one of the backend servers 20; Based on a preconfigured load balancing strategy, a first backend server and a first processor of the first backend server are determined from the candidate backend servers, wherein the load balancing strategy includes at least one of a polling mechanism, weight distribution, a minimum number of connections, and a fastest response time.

[0053] Specifically, the load balancer 10 will first determine the request type of the first request, such as based on the URL path, HTTP type (such as GET, POST), a specific field in the request header, etc.

[0054] Once the request type is determined, the load balancer 10 will screen out candidate backend servers that match the request type based on the service type provided by at least one backend server 20. For example, if the first request is a video streaming request, the load balancer 10 will screen out backend servers 20 that are configured with video streaming services.

[0055] After selecting the candidate backend servers, the load balancer 10 further applies a preconfigured load balancing strategy to determine the first backend server from the candidate backend servers. These load balancing strategies may include but are not limited to the following: Polling mechanism: Requests are assigned to backend servers in turn.

[0056] Weight distribution: Assign different weights to each backend server based on its processing capacity or resource status, and then distribute requests in proportion to the weights.

[0057] Least number of connections: Select the backend server with the least number of current connections (that is, the lightest load) to process the request.

[0058] Fastest response time: Based on the average response time of the backend servers over a period of time or real-time performance monitoring data, select the backend server with the shortest response time.

[0059] Finally, the load balancer 10 determines a backend server from the candidate backend servers as the first backend server according to the load balancing strategy, and may further determine a first processor on the backend server.

[0060] In the system provided in this embodiment, the load balancer not only ensures that requests are distributed to back-end servers that can handle requests of this type, but also optimizes resource utilization through intelligent load balancing strategies, thereby improving the overall performance and reliability of the system.

[0061] In some embodiments, the load balancer 10 is further configured to: Monitoring the status of the first backend server and the first network card; When it is determined that the first network card is in an unavailable state, the first request is forwarded to a second network card, where the second network card includes other network cards in the first backend server except the first network card or network cards of other backend servers except the first backend server.

[0062] Here, the load balancer 10 continuously monitors the status of the first backend server and the first network card connected thereto, including but not limited to the connection status, data flow, error rate, packet loss rate, response time, etc. of the first network card. With this information, the load balancer 10 can evaluate the health of the first network card.

[0063] During the monitoring process, if the load balancer 10 determines whether the status of the first network card is in an unavailable state, the unavailable state means that the first network card can no longer receive and process requests from the load balancer 10, and can no longer return the result corresponding to the request to the load balancer 10.

[0064] When it is detected that the first network card is unavailable, the load balancer 10 will select an alternative second network card to continue processing the first request that should have been forwarded through the first network card. If the first backend server is configured with multiple network cards, the load balancer 10 can select one of the network cards in a normal state as the second network card. If all network cards on the first backend server are unavailable, or in order to achieve higher availability and load balancing, the load balancer 10 can also select network cards of other backend servers of the same service type as the second network card.

[0065] Finally, the load balancer 10 will forward the first request that should have been forwarded through the first network card to the selected second network card. In this way, even if the first network card fails, the first request of the client can be processed in time, thereby ensuring the continuity of the system and the user experience.

[0066] In the system provided in this embodiment, the load balancer ensures that the system can continue to provide services even when some hardware or network components fail, thereby improving the reliability and failover capability of the entire system.

[0067] In some embodiments, reference Figure 3 As shown, the load balancing system further includes a system service processor 30, and the system service processor 30 is communicatively connected with the load balancer 10 and at least one of the backend servers 20; The system service processor 30′ is used to: Acquire monitoring data of the load balancing system, the monitoring data including the first performance and the first state of the load balancer 10 and the second performance and the second state of each of the processors 21 in each of the backend servers 20; Based on the monitoring data, the load balancing system is adjusted, and the adjustment includes at least one of an adjustment of the load balancing strategy and an adjustment of the number of backend servers 20 .

[0068] In this embodiment, the load balancing system further includes a system service processor 30. Specifically, the system service processor 30 collects monitoring data from the load balancer 10 and the backend server 20 in the load balancing system periodically or in real time. Such data includes but is not limited to: The first performance and first status of the load balancer 10: such as the performance indicators of the load balancer 10, such as CPU usage, memory usage, network throughput, number of connections, error rate, and its own health status (such as whether it is online, whether there is a fault report, etc.).

[0069] The second performance and second state of each processor 21 in each backend server 20: such as performance indicators (such as CPU usage, memory usage, processing queue length, etc.) and state information (such as whether it is overloaded, whether an error occurs, etc.) of each processor 21.

[0070] Based on the collected monitoring data, the load balancing monitoring module can dynamically adjust the load balancing strategy. For example, if the processor load of a backend server is too high, the load balancing monitoring module will reduce the subsequent traffic allocated to the backend server to distribute requests more efficiently.

[0071] In addition, the system service processor 30 can also increase or decrease the number of backend servers according to the overall load of the system. If the system load continues to exceed the processing capacity of the current backend server, or if the number of available servers is insufficient due to the failure of some backend servers, the load balancing monitoring module can trigger the automatic expansion mechanism to add new backend servers to the server pool of the load balancer; on the contrary, if the processing load is reduced for a long time, the number of backend servers in the server pool of the load balancer will be automatically reduced to save resources. Alternatively, when a backend server is monitored to have a failure, the load balancer is notified so that the load balancer can remove it from the service pool.

[0072] In the system provided by this embodiment, the system service processor ensures that the load balancing system can self-optimize according to real-time performance data and status information, thereby maintaining high availability, high performance and efficient use of resources.

[0073] In some embodiments, the system service processor 30′ is further configured to: Configuring a virtual IP address of the load balancer 10 so that the client can access the virtual IP address; Configuring a backend server list of the load balancer 10, wherein the backend server list includes an IP address, a port number and a service type of at least one of the backend servers; Configure the health check protocol and parameters corresponding to the backend servers in the backend server list; The load balancing strategy of the load balancer is configured, wherein the load balancing strategy includes at least one of a polling mechanism, a weight distribution, a minimum number of connections, and a fastest response time.

[0074] The system service processor 30 allows the administrator to set one or more virtual IP addresses for the load balancer 10. The client accesses the service through these virtual IP addresses without knowing the actual IP address of the backend server 20. The virtual IP address provides a unified entry point, so that the load balancer 10 can transparently distribute client requests to different backend servers 20.

[0075] The system service processor 30 also allows the administrator to customize a group of backend servers and specify the IP address, port number and service type (such as HTTP, HTTPS, FTP, etc.) for each backend server 20. In this way, the load balancer 10 can correctly forward the request to the corresponding backend server 20 and network card according to the information in the backend server list.

[0076] In addition, to ensure that only healthy backend servers 20 can receive requests, the load balancer configuration module also allows the administrator to configure health check protocols (such as HTTP GET request, TCP connection test, etc.) and parameters (such as request path, expected response code, timeout period, etc.) for each backend server in the backend server list. Through these configurations, the load balancer 10 can automatically detect the health status of the backend server 20 and exclude it from the backend server list when a fault is found.

[0077] Here, the load balancing strategy determines how the load balancer 10 distributes client requests to the backend servers 20. The load balancer configuration module can provide a variety of load balancing strategies for administrators to choose from, including but not limited to: Polling mechanism: Distribute requests to each backend server and corresponding processor in sequence.

[0078] Weight allocation: Assign different weights to backend servers and processors based on their processing capabilities or other factors, and then distribute requests in proportion to the weights.

[0079] Least number of connections: Distributes requests to the backend server with the least number of connections and the corresponding processor to balance the load of each backend server.

[0080] Fastest response time: Measures the response time of each backend server and processor in real time, and distributes requests to the fastest responding backend server and corresponding processor.

[0081] The system provided in this embodiment, through these configurations, ensures that the load balancer can effectively distribute client requests according to established rules and policies, while maintaining the health and efficient operation of the backend server cluster.

[0082] In some embodiments, the load balancing strategy includes: Determine a backend server pointed to by a first pointer among candidate backend servers that match a first request currently sent by the client, and use the backend server pointed to by the first pointer as a first backend server that matches the first request; Determine the processor pointed to by the second pointer in the current first backend server, and use the processor pointed to by the second pointer as the first processor matched by the first request; When the first processor receives the first request, the first pointer points to the next backend server according to a first polling order, and the second pointer points to the next processor in the first backend server according to a second polling order.

[0083] Here, the first polling order is a polling order for backend servers. When the load balancer 10 receives a request from a client, it selects a backend server 20 to process the request according to the first polling order.

[0084] Specifically, each service type is configured with a list that contains all available backend servers. When determining the candidate backend service in the list that matches the first request sent by the client, the first pointer selects the first backend server according to the polling order of the list. Here, the second polling order is the polling order for the processor 21 inside each backend server 20. Once the first backend server is selected, the load balancer 10 selects a first processor on the first backend server to process the request according to the second polling order.

[0085] Further, when it is determined that the selected first processor can process the first request, the first pointer will automatically point to the next backend server 20 according to the first polling order to prepare for the next request. At the same time, the second pointer will also point to the next processor 21 in the first backend server according to the second polling order.

[0086] The system provided in this embodiment can avoid overloading of any single server or processor by distributing requests evenly to different backend servers and processors through a load balancing strategy of a polling mechanism, thereby improving the overall performance and response time of the system.

[0087] In some embodiments, the load balancing strategy includes: Determine a backend server corresponding to the least number of connections among the candidate backend servers that match the first request currently sent by the client, and use the backend server corresponding to the least number of connections as the first backend server that matches the first request; A processor corresponding to the minimum number of threads in the current first backend server is determined, and the processor corresponding to the minimum number of threads is used as a first processor matching the first request.

[0088] Here, the current number of connections refers to the number of connections currently being processed and / or waiting to be processed by the backend server 20. For example, the load balancer 10 can continuously monitor the current number of connections of each backend server 20 by periodically sending health check requests.

[0089] When a new first request arrives, the load balancer 10 compares the number of connections of all candidate backend servers and selects the first backend server with the least number of active connections to process the first request.

[0090] In addition, after selecting the first backend server, the load balancer 10 needs to select the first processor with the least number of threads to process the first request according to the number of threads of each processor 21 in the monitored first backend server.

[0091] Here, the current number of threads of each processor 21 can be determined based on the historical requests sent by the load balancer 10 to each processor 21 and the processing progress of the historical requests. For example, when the load balancer 10 sends a request to a processor 21, the thread count of the processor 21 is increased by 1, and when the result corresponding to the request returned by the processor 21 is received, the thread count of the processor 21 is decreased by 1, and so on.

[0092] The system provided in this embodiment can achieve finer-grained load balancing within the backend server by simultaneously considering the number of connections of the backend server and the number of threads of the processor, thereby balancing not only the load between the backend servers, but also the load of each processor within the backend server.

[0093] In some embodiments, the load balancing strategy includes: Determine a backend server corresponding to the shortest first response time among the candidate backend servers that match the first request currently sent by the client, and use the backend server corresponding to the shortest first response time as the first backend server that matches the first request; A processor corresponding to the shortest second response time in the current first backend server is determined, and the processor corresponding to the shortest second response time is used as a first processor matched with the first request.

[0094] Here, the first response time of each backend server 20 may be determined based on a historical average response time. The historical average response time may be obtained by long-term monitoring and statistics of the response time of all requests processed by the backend server 20.

[0095] Specifically, the load balancer 10 may internally maintain a data list for storing and updating the first response time of each backend server 20, and query the data list when receiving a new first request to select the first backend server according to the first response time.

[0096] Here, the second response time of the processor 21 may be determined based on the workload status of the processor 21, including but not limited to indicators such as CPU usage, memory usage, number of threads, etc. For example, each backend server 20 may periodically feed back the load status of each processor therein to the load balancer 10.

[0097] After the first backend server is determined, the first processor corresponding to the shortest second response time is selected to process the first request according to the second response time of each processor 21 in the first backend server.

[0098] The system provided in this embodiment can reduce the waiting time of the client and improve the response speed of the service by selecting the back-end server and processor with the shortest response time to process the request.

[0099] In some embodiments, the first backend server is further configured to: When the first processor receives the first request, starting a timer; When it is determined by the timer that the processing time of the first processor for processing the first request reaches a preset time, determining a current processing progress of the first request; If there are still multiple tasks to be processed in the current first request, the multiple tasks to be processed are parsed to determine at least one target task to be processed that can be processed independently among the multiple tasks to be processed; generating a second request based on at least one of the target tasks to be processed; Determine at least one candidate processor other than the first processor in the current first backend server, wherein the candidate processor is connected to the first network card; Determining a second processor from the candidate processors based on a load state of at least one of the candidate processors; Forwarding the second request in the first request to the second processor through the first network card.

[0100] When the first processor receives the first request, the first backend server starts a timer. The purpose of this timer is to monitor the time required for processing the first request. Here, the timer can use forward timing or reverse timing, which is not limited.

[0101] Through the timer, the first backend server can monitor and determine the processing time of the first processor to process the first request. If the processing time reaches the preset time, the first processor will further check the processing progress of the current first request.

[0102] If there are multiple pending tasks in the first request, and the processing time has exceeded the preset time, the first backend server will further parse these pending tasks. It will identify at least one pending task that can be processed independently. In other words, it is a pending task that can be executed independently without relying on the completion of other tasks.

[0103] Based on the identified tasks to be processed that can be processed independently, the first server generates a second request that includes tasks that can be separated from the first request and processed in parallel, thereby reducing the complexity and processing time of a single request.

[0104] Next, the first backend server further determines at least one candidate processor connected to the first network card except the first processor, so as to forward the request to the other processors through the first network card.

[0105] It should be noted here that, compared with direct forwarding between processors, forwarding through a network card does not require a specific communication interface or protocol between processors. In addition, compared with direct forwarding which increases the load on the first processor, in this embodiment, through network card forwarding, the network hardware can take over this task and reduce the burden on the processor.

[0106] Finally, based on the load status of the candidate processors (such as CPU usage, memory usage, the number of currently processed tasks, etc.), a most suitable second processor is determined from the candidate processors to process the second request, and the second request is forwarded to the second processor through the first network card for processing.

[0107] The system provided in this embodiment can dynamically allocate and adjust resources based on the network card configuration of the back-end server when processing complex or time-consuming requests, thereby ensuring that the task can be completed in the shortest time possible to avoid overloading of a single processor, further achieving load balancing at the hardware level, and enhancing the reliability and stability of the system.

[0108] Based on the above embodiments, it can be seen that the load balancing system provided in the embodiments of the present application has the following advantages: Improved network bandwidth utilization: Distributing network traffic to multiple backend servers effectively improves network bandwidth utilization.

[0109] Network performance optimization: Reduced the response time of user requests and improved the overall performance of the service.

[0110] Enhanced system reliability: Avoided service interruption caused by single-point network card failure, enhanced system reliability and stability. By implementing network card load balancing technology, successfully solved the network traffic bottleneck problem, and improved service stability and availability.

[0111] The load balancing method provided by the present invention is described below. The load balancing method described below and the load balancing system described above can be referenced to each other.

[0112] Figure 4 is a flow chart of the load balancing method provided by the present invention; the method is executed by the load balancing system provided by the above embodiments; Figure 4 As shown, the method includes: Step 410, responding to the first request sent by the client through the load balancer, determining the first backend server and the first processor of the first backend server, and forwarding the first request to the first network card corresponding to the first processor; Step 420, through the first back-end server, when it is determined that the processing time of the first processor to process the first request exceeds a preset time, the second request in the first request is forwarded to the second processor through the first network card, wherein the second processor is a processor in the first back-end server and is connected to the first network card.

[0113] Here, the client (such as a browser, API client, etc.) sends a first request to the load balancer. The request may be an HTTP request, a database query, or other types of network requests. After the load balancer receives the request from the client, it decides how to distribute the first request based on the predefined load balancing strategy.

[0114] Specifically, the load balancer selects a backend server from the server pool to process the first request. The selected backend server is called the first backend server.

[0115] After selecting the first backend server, the load balancer needs to further determine which processor on the server will actually process the request. For example, a processor is selected from the first backend server based on the current load, processing capacity or other configured priority rules of the processor on the first backend server. The selected processor is called the first processor.

[0116] It should be understood that in this embodiment, each processor is connected to at least two network cards via a system bus. Therefore, in this embodiment, after selecting the first processor, it is necessary to further select the first network card corresponding to the first processor.

[0117] In one example, the load balancer is also configured with detailed information about each network card on each backend server, including the IP address of each network card, the processor to which it is connected, and related load information. Therefore, in this embodiment, the load balancer will evaluate the current load of each network card of the first processor (such as network traffic, packet processing rate, error rate and other indicators). In addition, the current load of other processors connected to each network card will be evaluated, and finally the first network card will be selected based on the current load of each network card and the current load of other processors connected to each network card (such as CPU usage, number of processes and threads, etc.). For example, a network card with low network traffic and low error rate, and whether other connected processors have sufficient remaining processing capacity to process requests is selected as the first network card.

[0118] Finally, the load balancer forwards the request to the first network card corresponding to the first processor responsible for processing the first request on the first backend server. Specifically, through a network protocol (such as TCP / IP), the load balancer sends the data packet of the first request to the IP address and port of the first network card.

[0119] When the first backend server receives the first request distributed by the load balancer and determines that the first processor is to process the request, the first backend server monitors the processing process, especially the time required for the first processor to process the first request.

[0120] If during the processing, the first back-end server finds that the time taken by the first processor to process the first request exceeds a preset time (for example, a preset time set based on system performance requirements, user experience standards, or historical data), in order to reduce the burden on the first processor and speed up the processing of the first request, the first back-end server will forward a part of the first request, namely the second request, to the second processor through the first network card.

[0121] Here, the second request may be a subtask in the first request, a fragment that can be executed independently, or a request that is associated with the first request but can be processed in parallel.

[0122] Here, the second processor is another processor on the first backend server, and it is also connected to the first network card. In other words, the first network card serves multiple processors at the same time, providing redundancy and load distribution capabilities at the hardware level.

[0123] In the method provided in this embodiment, the load balancer establishes a communication connection with the back-end server, so that the load balancer can determine to which back-end server and which processor of the server the first request of the client is forwarded. Load balancing at the software level is achieved. In addition, each back-end server is equipped with at least two processors and at least two network cards, and each processor is connected to at least two network cards through a system bus. Specifically, when the processing time of the first processor to process the first request exceeds the preset time, part of the second request in the first request can be forwarded to the second processor, so that the request can be transferred between different processors through the network card to avoid overloading of a single processor, further realizing load balancing at the hardware level, and enhancing the reliability and stability of the system.

[0124] In some embodiments, step 410 is specifically implemented based on the following steps: Responding to the first request sent by the client, by the first backend server, to determine the request type corresponding to the first request, and screening out a candidate backend server matching the request type from at least one of the backend servers based on the service type of at least one of the backend servers; Based on a preconfigured load balancing strategy, a first backend server and a first processor of the first backend server are determined from the candidate backend servers, wherein the load balancing strategy includes at least one of a polling mechanism, weight distribution, a minimum number of connections, and a fastest response time.

[0125] Specifically, the load balancer first determines the request type of the first request, such as based on the URL path, HTTP type (such as GET, POST), specific fields in the request header, etc.

[0126] Once the request type is determined, the load balancer will filter out candidate backend servers that match the request type based on the service type provided by at least one backend server. For example, if the first request is a video streaming request, the load balancer will filter out backend servers that are configured with video streaming services.

[0127] After selecting the candidate backend servers, the load balancer will further apply the pre-configured load balancing strategy to determine the first backend server from the candidate backend servers. These load balancing strategies may include but are not limited to the following: Polling mechanism: Requests are assigned to backend servers in turn.

[0128] Weight distribution: Assign different weights to each backend server based on its processing capacity or resource status, and then distribute requests in proportion to the weights.

[0129] Least number of connections: Select the backend server with the least number of current connections (that is, the lightest load) to process the request.

[0130] Fastest response time: Based on the average response time of the backend servers over a period of time or real-time performance monitoring data, select the backend server with the shortest response time.

[0131] Finally, the load balancer determines a backend server from the candidate backend servers as the first backend server according to the load balancing strategy, and may further determine the first processor on the backend server.

[0132] In the method provided in this embodiment, the load balancer not only ensures that the request is distributed to the back-end server that can handle the request of this type, but also optimizes the resource utilization through the intelligent load balancing strategy, thereby improving the overall performance and reliability of the system.

[0133] In some embodiments, step 410 may further include the following steps: Monitoring the status of the first backend server and the first network card through the load balancer; When it is determined that the first network card is in an unavailable state, the first request is forwarded to a second network card, where the second network card includes other network cards in the first backend server except the first network card or network cards of other backend servers except the first backend server.

[0134] Here, the load balancer continuously monitors the status of the first backend server and the first network card connected thereto, including but not limited to the connection status, data flow, error rate, packet loss rate, response time, etc. of the first network card. With this information, the load balancer can evaluate the health of the first network card.

[0135] During the monitoring process, the load balancer determines whether the status of the first network card is in an unavailable state. Here, the unavailable state means that the first network card can no longer receive and process requests from the load balancer, and can no longer return the result corresponding to the request to the load balancer.

[0136] When it is detected that the first network card is unavailable, the load balancer will select an alternative second network card to continue processing the first request that should have been forwarded through the first network card. If the first backend server is configured with multiple network cards, the load balancer can select one of the network cards in normal state as the second network card. If all network cards on the first backend server are unavailable, or in order to achieve higher availability and load balancing, the load balancer can also select network cards of other backend servers of the same service type as the second network card.

[0137] Finally, the load balancer will forward the first request that should have been forwarded through the first network card to the selected second network card. In this way, even if the first network card fails, the client's first request can be processed in time, thus ensuring system continuity and user experience.

[0138] The method provided in this embodiment, the load balancer ensures that the system can continue to provide services even when some hardware or network components fail, thereby improving the reliability and failover capability of the entire system.

[0139] The method provided by the present invention is executed by the system based on the above-mentioned method embodiments. Please refer to the above-mentioned embodiments for the specific process and detailed content, which will not be repeated here.

[0140] Figure 5 An example of a physical structure diagram of an electronic device is shown in FIG. Figure 5 As shown, the electronic device may include: a processor 510, a communication interface 520, a memory 530 and a communication bus 540, wherein the processor 510, the communication interface 520 and the memory 530 complete mutual communication through the communication bus 540. The processor 510 may call the logic instructions in the memory 530 to execute the load balancing method, which includes: responding to a first request sent by a client through a load balancer, determining a first backend server and a first processor of the first backend server, and forwarding the first request to a first network card corresponding to the first processor; through the first backend server, when it is determined that the processing time of the first request processed by the first processor exceeds a preset time, forwarding the second request in the first request to the second processor through the first network card, wherein the second processor is a processor in the first backend server and is connected to the first network card.

[0141] In addition, the logic instructions in the above-mentioned memory 530 can be implemented in the form of a software functional unit and can be stored in a computer-readable storage medium when it is sold or used as an independent product. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art or the part of the technical solution, can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk, etc. Various media that can store program codes.

[0142] On the other hand, the present invention also provides a computer program product, which includes a computer program. The computer program can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the load balancing method provided by the above methods, which includes: through a load balancer, responding to a first request sent by a client, determining a first back-end server and a first processor of the first back-end server, and forwarding the first request to a first network card corresponding to the first processor; through the first back-end server, when it is determined that the processing time of the first processor to process the first request exceeds a preset time, forwarding the second request in the first request to the second processor through the first network card, wherein the second processor is a processor in the first back-end server and is connected to the first network card.

[0143] On the other hand, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, is implemented to execute the load balancing method provided by the above-mentioned methods, the method comprising: responding to a first request sent by a client through a load balancer, determining a first back-end server and a first processor of the first back-end server, and forwarding the first request to a first network card corresponding to the first processor; through the first back-end server, when it is determined that the processing time of the first processor to process the first request exceeds a preset time, forwarding the second request in the first request to the second processor through the first network card, wherein the second processor is a processor in the first back-end server and is connected to the first network card.

[0144] The device embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. Ordinary technicians in this field can understand and implement it without paying creative labor.

[0145] Through the description of the above implementation methods, those skilled in the art can clearly understand that each implementation method can be implemented by means of software plus a necessary general hardware platform, and of course, can also be implemented by hardware. Based on this understanding, the above technical solution is essentially or the part that contributes to the prior art can be embodied in the form of a software product, and the computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a disk, an optical disk, etc., including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0146] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A load balancing system, characterized in that: include: A load balancer and at least one back-end server in communication with the load balancer, wherein the back-end server comprises at least two processors and at least two network cards, and each of the processors is connected to at least two of the network cards via a system bus; The load balancer is configured to respond to a first request sent by a client, determine a first backend server and a first processor of the first backend server, and forward the first request to a first network card corresponding to the first processor; The first back-end server is also used to forward the second request in the first request to the second processor through the first network card when it is determined that the processing time of the first processor to process the first request exceeds a preset time. The second processor is a processor in the first back-end server and is connected to the first network card.

2. The load balancing system according to claim 1, characterized in that: The load balancer is also used to: In response to a first request sent by a client, determine a request type corresponding to the first request, and select a candidate backend server matching the request type from at least one of the backend servers based on a service type of at least one of the backend servers; Based on a preconfigured load balancing policy, a first backend server and a first processor of the first backend server are determined from the candidate backend servers.

3. The load balancing system according to claim 1, characterized in that: The load balancer is also used to: Monitoring the status of the first backend server and the first network card; When it is determined that the first network card is in an unavailable state, the first request is forwarded to a second network card, where the second network card includes other network cards in the first backend server except the first network card or network cards of other backend servers except the first backend server.

4. The load balancing system according to claim 1, characterized in that: Also included is a system service processor, the system service processor being communicatively connected to the load balancer and at least one of the backend servers; The system service processor is used to: Acquire monitoring data of the load balancing system, the monitoring data including a first performance and a first state of the load balancer and a second performance and a second state of each of the processors in each of the backend servers; Based on the monitoring data, the load balancing system is adjusted, and the adjustment includes at least one of an adjustment of the load balancing strategy and an adjustment of the number of backend servers.

5. The load balancing system according to claim 4, characterized in that: The system service processor is also used to: Configuring a virtual IP address of the load balancer so that the client can access the virtual IP address; Configuring a backend server list of the load balancer, wherein the backend server list includes an IP address, a port number, and a service type of at least one of the backend servers; Configure the health check protocol and parameters corresponding to the backend servers in the backend server list; Configure the load balancing policy of the load balancer.

6. The load balancing system according to claim 2 or 5, characterized in that: The load balancing strategy includes: Determine a backend server pointed to by a first pointer among candidate backend servers that match a first request currently sent by the client, and use the backend server pointed to by the first pointer as a first backend server that matches the first request; Determine the processor pointed to by the second pointer in the current first backend server, and use the processor pointed to by the second pointer as the first processor matched by the first request; When the first processor receives the first request, the first pointer points to the next backend server according to a first polling order, and the second pointer points to the next processor in the first backend server according to a second polling order.

7. The load balancing system according to claim 2 or 5, characterized in that: The load balancing strategy includes: Determine a backend server corresponding to the least number of connections among the candidate backend servers that match the first request currently sent by the client, and use the backend server corresponding to the least number of connections as the first backend server that matches the first request; A processor corresponding to the minimum number of threads in the current first backend server is determined, and the processor corresponding to the minimum number of threads is used as a first processor matching the first request.

8. The load balancing system according to claim 2 or 5, characterized in that: The load balancing strategy includes: Determine a backend server corresponding to the shortest first response time among the candidate backend servers that match the first request currently sent by the client, and use the backend server corresponding to the shortest first response time as the first backend server that matches the first request; A processor corresponding to the shortest second response time in the current first backend server is determined, and the processor corresponding to the shortest second response time is used as a first processor matched with the first request.

9. The load balancing system according to claim 1, characterized in that: The first backend server is also used for: When the first processor receives the first request, starting a timer; When it is determined by the timer that the processing time of the first processor for processing the first request reaches a preset time, determining a current processing progress of the first request; If there are still multiple tasks to be processed in the current first request, the multiple tasks to be processed are parsed to determine at least one target task to be processed that can be processed independently among the multiple tasks to be processed; generating a second request based on at least one of the target tasks to be processed; Determine at least one candidate processor other than the first processor in the current first backend server, wherein the candidate processor is connected to the first network card; Determining a second processor from the candidate processors based on a load state of at least one of the candidate processors; Forwarding the second request in the first request to the second processor through the first network card.

10. A load balancing method, characterized in that: The method is applied to the load balancing system according to any one of claims 1 to 9, and the method comprises: In response to a first request sent by a client, a load balancer is used to determine a first backend server and a first processor of the first backend server, and the first request is forwarded to a first network card corresponding to the first processor; Through the first back-end server, when it is determined that the processing time of the first processor to process the first request exceeds a preset time, the second request in the first request is forwarded to the second processor through the first network card, wherein the second processor is a processor in the first back-end server and is connected to the first network card.

Citation Information

Patent Citations

  • Cluster load balancing system and achieving method thereof

    CN103731482A

  • Load balancing system, method and device, equipment and medium

    CN111008075A

  • Load balancing method, device and equipment and storage medium

    CN111193773A

  • Four-layer load balancing data processing method and related device

    CN115766729A

  • Method and apparatus for balancing a load among a plurality of servers in a computer system

    US7181524B1