A management method and device of a server, a cluster, a product, and a storage medium
Patent Information
- Application Number
- CN202510337216.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-19
- Publication Date
- 2026-09-22
AI Technical Summary
然而,在网卡发生故障的情况下,由于该多个CPU无法通过网卡收发数据,从而导致该多个CPU不能正常处理业务
[0012] In one possible implementation, when there is a communication connection between multiple CPUs, the candidate second network card includes: the network card corresponding to the CPU that has a communication connection with the first CPU; or, when there is no communication connection between multiple CPUs, and the first CPU has a communication connection with at least two network cards among the multiple network cards, the candidate second network card includes: the network card other than the first network card among the at least two network cards that have a communication connection with the first CPU.
Smart Images

Figure CN122802343A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of servers, and more particularly to a server management method, apparatus, cluster, product, and storage medium. Background Technology
[0002] As is well known, a typical multi-processor server includes multiple central processing units (CPUs) and a network interface card (NIC); these CPUs send and receive data through the NIC. However, if the NIC fails, the CPUs cannot send or receive data, thus preventing them from processing business operations normally. This reduces the stability of the server. Summary of the Invention
[0003] This application provides a server management method, apparatus, cluster, product, and storage medium that can improve server stability.
[0004] To achieve the above objectives, the embodiments of this application adopt the following technical solutions:
[0005] In a first aspect, embodiments of this application provide a server management method. The server includes: multiple CPUs and multiple network interface cards (NICs), wherein the multiple CPUs and the multiple NICs are communicatively connected; wherein a first CPU among the multiple CPUs sends and receives data through a first NIC among the multiple NICs; the method includes: the first CPU acquiring fault information of the first NIC; the fault information indicating a fault in which the NIC cannot function properly; and the first CPU responding to the fault information by sending and receiving data through a second NIC; wherein the second NIC is one of the multiple NICs other than the first NIC that is communicatively connected to the first CPU.
[0006] This application provides a server management method. This method configures multiple network interface cards (NICs) for multiple CPUs in a multi-processor server, ensuring communication connections between the CPUs and these NICs. If the first NIC used by any CPU for sending and receiving data fails, the first CPU sends and receives data through a second NIC (other than the first NIC) that is also connected to it. This solves the problem of all CPUs in the server being unable to process business operations normally due to NIC failure, thus improving server stability.
[0007] In addition, since the method allows the first CPU to send and receive data through the second network card when the first network card corresponding to the first CPU fails, it avoids the problem that the first CPU cannot process business due to the failure of the first network card, thereby reducing the impact of the failure of the first network card on the performance of the entire server (i.e., reducing the explosion radius); thus, it further improves the stability of the server.
[0008] In one possible implementation, the first network card is the primary network card of the first CPU; the second network card is the backup network card of the first CPU; wherein, when the primary network card is not faulty, the first CPU sends and receives data through the primary network card and does not send and receive data through the backup network card.
[0009] In one possible implementation, there is a communication connection between the plurality of CPUs; or, there is no communication connection between the plurality of CPUs, but each of the plurality of CPUs has a communication connection with at least two of the plurality of network cards.
[0010] In one possible implementation, the method further includes: the first CPU obtaining the load of the candidate second network card that has a communication connection with the first CPU among the other network cards; the first CPU determining the network card whose load meets the conditions among the candidate second network cards as the second network card.
[0011] According to the embodiments of this application, the task of sending and receiving data by the first CPU is allocated to the multiple second network cards based on the ratio of available load among the multiple second network cards, so that each of the multiple second network cards can send and receive data for the first CPU based on the allocated task of sending and receiving data, thereby improving the load balancing of the second network cards.
[0012] In one possible implementation, when there is a communication connection between multiple CPUs, the candidate second network card includes: the network card corresponding to the CPU that has a communication connection with the first CPU; or, when there is no communication connection between multiple CPUs, and the first CPU has a communication connection with at least two network cards among the multiple network cards, the candidate second network card includes: the network card other than the first network card among the at least two network cards that have a communication connection with the first CPU.
[0013] In one possible implementation, where there is a communication connection between the aforementioned multiple CPUs, the method further includes: the first CPU, based on the topology of the multiple CPUs, determines the network card corresponding to the second CPU among the multiple CPUs whose communication distance with the first CPU meets the condition as the second network card.
[0014] In the case of a failure of the first network card corresponding to the first CPU, the first CPU uses the topology of multiple CPUs in its server to select the network card corresponding to the second CPU among the multiple CPUs that meets the communication distance condition (e.g., the shortest communication distance) as the second network card. This allows the first CPU to send and receive data based on the second network card with the shortest communication distance, thereby improving the efficiency of the first CPU in sending and receiving data.
[0015] In one possible implementation, when there are multiple second network cards, before sending and receiving data through the second network cards as described above, the process includes: the first CPU allocating the task of sending and receiving data to the multiple second network cards according to the ratio of available load among the multiple second network cards.
[0016] According to the embodiments of this application, the task of sending and receiving data by the first CPU is allocated to the multiple second network cards based on the ratio of available load among the multiple second network cards, so that each of the multiple second network cards can send and receive data for the first CPU based on the allocated task of sending and receiving data, thereby improving the load balancing of the second network cards.
[0017] In one possible implementation, the first CPU obtains the fault information of the first network card by: the first CPU obtaining the fault information of the first network card from the BMC in the server; wherein the BMC is used to read the fault information from the first network card.
[0018] In this embodiment, the fault information of the first network card is obtained from the BMC by the first CPU. Since the BMC reads the fault information from the first network card, the problem of the first network card being unable to send the fault information due to the fault of the first network card is avoided, thus improving the reliability of obtaining the fault information of the first network card.
[0019] Secondly, embodiments of this application provide a server management device. The server includes: multiple CPUs and multiple network interface cards (NICs), wherein the multiple CPUs have communication connections with the multiple NICs; wherein a first CPU among the multiple CPUs sends and receives data through a first NIC among the multiple NICs; the server management device is applied to the first CPU; the server management device includes: a transceiver module and a processing module; the transceiver module is used to obtain fault information of the first NIC; the fault information is used to indicate a fault in which the NIC cannot work properly; the processing module is used to respond to the fault information by sending and receiving data through a second NIC; wherein the second NIC is one of the multiple NICs other than the first NIC that has a communication connection with the first CPU.
[0020] In one possible implementation, the first network interface card (NIC) is the primary NIC of the first CPU; the second NIC is the backup NIC of the first CPU. In this embodiment, when the primary NIC is functioning correctly, the first CPU obtains the fault information of the first NIC from the Base Configuration Management (BMC). Since the BMC reads this fault information from the first NIC, the problem of the first NIC being unable to send fault information due to a fault is avoided, thus improving the reliability of obtaining the fault information of the first NIC. The first CPU sends and receives data through the primary NIC, not through the backup NIC.
[0021] In one possible implementation, there is a communication connection between the multiple CPUs; or, there is no communication connection between the multiple CPUs, but each of the multiple CPUs has a communication connection with at least two of the multiple network interface cards (NICs).
[0022] In one possible implementation, the transceiver module is used to obtain the load of the candidate second network interface card (NIC) that has a communication connection with the first CPU among other NICs; the processing module is used to determine the NIC whose load meets the conditions among the candidate second NICs as the second NIC.
[0023] In one possible implementation, when there is a communication connection between multiple CPUs, the candidate second network interface card (NIC) includes: the NIC corresponding to the CPU that has a communication connection with the first CPU; or, when there is no communication connection between multiple CPUs, and the first CPU has a communication connection with at least two of the NICs, the candidate second NIC includes: the NIC other than the first NIC among the at least two NICs that have a communication connection with the first CPU.
[0024] In one possible implementation, the processing module is used to determine the network card corresponding to the second CPU among the multiple CPUs that meets the communication distance condition with the first CPU as the second network card, based on the topology relationship of multiple CPUs.
[0025] In one possible implementation, the processing module is used to distribute the data transmission and reception tasks of the first CPU to the multiple second network cards according to the ratio of available load among the multiple second network cards.
[0026] In one possible implementation, the transceiver module is used to obtain fault information of the first network card from the BMC in the server; wherein the BMC is used to read fault information from the first network card.
[0027] Thirdly, this application provides a computing device cluster including at least one computing device, each computing device including a processor and a memory; the processor of the at least one computing device is configured to execute instructions stored in the memory of the at least one computing device, such that the computing device cluster performs the method described in the first aspect and any of its possible implementations.
[0028] Fourthly, this application provides a computer-readable storage medium having computer instructions stored thereon, which, when executed on a computing device, cause the computing device to perform the method described in any one of the first aspects and its possible implementations.
[0029] Fifthly, this application provides a computer program product containing instructions that, when run on a computer, cause the computer to perform the method described in any one of the first aspects and their possible implementations.
[0030] It should be understood that the beneficial effects of the technical solutions of the second to fifth aspects of this application and the corresponding possible implementations can be referred to the above-described technical effects of the first aspect and its corresponding possible implementations, and will not be repeated here. Attached Figure Description
[0031] Figure 1 This application provides a schematic diagram of the architecture of a cloud computing system;
[0032] Figure 2 This application provides a schematic diagram of a dual single-path server system architecture;
[0033] Figure 3 A schematic diagram of another dual-single-path server system architecture provided in this application;
[0034] Figure 4 This application provides a schematic diagram of a dual-socket server system architecture.
[0035] Figure 5 A schematic diagram of another dual-socket server system architecture provided in this application;
[0036] Figure 6 One of the flowcharts illustrating a server management method provided in this application;
[0037] Figure 7 A second flowchart illustrating a server management method provided in this application;
[0038] Figure 8 The third flowchart illustrating a server management method provided in this application;
[0039] Figure 9 A schematic diagram of the CPU topology provided in this application;
[0040] Figure 10 The fourth flowchart illustrating a server management method provided in this application;
[0041] Figure 11A schematic diagram of the structure of a server management device provided in this application;
[0042] Figure 12 A hardware schematic diagram of a computing device provided in this application;
[0043] Figure 13 This is a schematic diagram of a computing device cluster provided in this application. Detailed Implementation
[0044] In this article, the term "and / or" is merely a description of the relationship between related objects, indicating that there can be three relationships. For example, A and / or B can represent three situations: A exists alone, A and B exist simultaneously, and B exists alone.
[0045] The terms "first CPU" and "second CPU," etc., used in the specification and claims of this application are used to distinguish different CPUs, not to describe a specific order of CPUs. Similarly, "first network card" and "second network card," etc., are used to distinguish different network cards, not to describe a specific order of network cards.
[0046] In the embodiments of this application, the terms "exemplary" or "for example" are used to indicate that something is an example, illustration, or description. Any embodiment or design that is described as "exemplary" or "for example" in the embodiments of this application should not be construed as being more preferred or advantageous than other embodiments or design. Specifically, the use of the terms "exemplary" or "for example" is intended to present the relevant concepts in a specific manner.
[0047] In the description of the embodiments in this application, unless otherwise stated, "multiple" means two or more. For example, multiple CPUs means two or more CPUs.
[0048] First, some concepts involved in the server management method, apparatus, cluster, product, and storage medium provided in the embodiments of this application will be explained as follows:
[0049] Multi-processor server: refers to a server configured with multiple CPUs.
[0050] Dual-socket server: A server configured with two CPUs, and the CPUs have a communication connection between them.
[0051] Dual-single-processor server: This is a server configured with two CPUs, and there is no communication connection between the CPUs (i.e., the two CPUs run independently); these two CPUs are used to run two independent operating systems.
[0052] PCIe (Peripheral Component Interconnect Express) is a high-speed serial computer expansion bus standard that supports high-speed serial transmission.
[0053] Explosion radius: Used to indicate the range and extent of the impact of a faulty component in a server on the entire server.
[0054] Hop count: Indicates the number of relay nodes required to travel from the source node to the target node.
[0055] With the gradual improvement of photolithography manufacturing technology, the number of physical CPU cores in a single server has increased exponentially, resulting in a significant increase in the overall capabilities of a single server.
[0056] When cloud providers offer elastic computing services, each physical server corresponding to the cloud server hosting the elastic computing service is typically equipped with a smart network interface card (smartNIC). This smartNIC is used to provide data caching and data transmission / reception services for multiple CPUs in the server (i.e., the smartNIC provides data caching and data transmission / reception services for this elastic computing service).
[0057] However, if the smartNIC fails, multiple CPUs in the server will be unable to send or receive data through the smartNIC, causing these CPUs to malfunction and thus reducing the server's stability.
[0058] Based on this, embodiments of this application provide a server management method. This method configures multiple network interface cards (NICs) for multiple CPUs in a multi-processor server, and establishes communication connections between the multiple CPUs and these NICs. If the first NIC used by the first CPU (i.e., any CPU) for sending and receiving data fails, the first CPU sends and receives data through a second NIC (other than the first NIC) that is also connected to it. This solves the problem of all CPUs in the server being unable to process business normally due to NIC failure, thus improving server stability.
[0059] In addition, since the method allows the first CPU to send and receive data through the second network card when the first network card corresponding to the first CPU fails, it avoids the problem that the first CPU cannot process business due to the failure of the first network card, thereby reducing the impact of the failure of the first network card on the performance of the entire server (i.e., reducing the explosion radius); thus, it further improves the stability of the server.
[0060] The server management method provided in this application can be applied to ordinary servers as well as cloud servers. The specific embodiments of this application do not limit the application scenarios of the server management method provided in this application.
[0061] For example, the server management method provided in this application can be applied to Figure 1 The cloud computing system 100 shown is an example. Figure 1 As shown, the cloud computing system 100 includes: a computing server cluster 110, a storage server cluster 120, a management server cluster 130, a network device cluster 140, and a tenant terminal 150. The computing server cluster 110, storage server cluster 120, and management server cluster 130 communicate with the tenant terminal 150 through the network device cluster 140.
[0062] It should be noted that the server management method provided in this application is applicable to any server in the aforementioned cloud computing system 100, including computing servers (i.e., computing nodes) in computing server cluster 110, storage servers (i.e., storage nodes) in storage server cluster 120, management servers (i.e., management nodes) in management server cluster 130, network devices (i.e., network nodes) in network device cluster 140, and tenant terminals 150 (i.e., terminal nodes).
[0063] The computing server cluster 110 serves as a computing resource within the cloud computing system 100. The computing server cluster 110 includes one or more computing servers (such as…). Figure 1 Computing servers 111 and 112 are shown. These computing servers are used to generate and allocate computing resources based on tenant needs using virtualization technology. Storage server cluster 120 serves as storage resources within the cloud computing system 100. Storage server cluster 120 includes one or more storage servers (such as…). Figure 1 Storage servers 121 and 122 are shown. Storage servers can be electronic devices with data storage capabilities. Management server cluster 130 is used to manage all computing services, shared storage, and network of the entire cloud computing system 100, and to provide tenants or administrators with an API for managing the entire node. Management server cluster 130 includes one or more management servers (such as...). Figure 1 The management server 131 and management server 132 are shown.
[0064] Network device cluster 140 includes one or more network devices. These network devices include, but are not limited to, switches and routers. For example, network device cluster 140 includes router 141, switch 142, switch 143, switch 144, and switch 145. Tenant terminal 150 is connected to router 141 via the Internet. Router 141 is connected to switch 143 via switch 142, and switch 143 is connected to each computing server in computing server cluster 110. Switch 144 is connected to each computing server in computing server cluster 110 and each storage server in storage server cluster 120. Switch 145 is connected to each computing server in computing server cluster 110, each storage server in storage server cluster 120, and each management server in management server cluster 130. Tenant terminal 150 includes one or more tenant terminals (e.g., ...). Figure 1 Tenant terminals 151 and 152 are shown. The tenant terminals contain the interfaces and applications required to access the cloud computing system 100.
[0065] It is worth noting that, Figure 1 This is merely an illustration and should not be construed as limiting the scope of this application. The cloud computing system 100 may include more or fewer devices, and this application does not limit this.
[0066] This application provides a method such as Figure 2 The system architecture of the dual single-socket server 200 shown is as follows: the dual single-socket server 200 includes: CPU0, CPU1, network card 0 and network card 1.
[0067] CPU0 has a communication connection with network interface card 0 (e.g., a PCIe bus-based communication connection), and CPU1 has a communication connection with network interface card 1; CPU0 and CPU1 are independent of each other (i.e., they do not have a communication connection). Specifically, CPU0 sends and receives data through network interface card 0; CPU1 sends and receives data through network interface card 1.
[0068] It should be noted that the network card in this application can be a smart network card (smartNIC) or a non-smart network card (i.e., a regular network card). The specific type of network card is not limited in the specific embodiments of this application.
[0069] In this application, both network interface card 0 and network interface card 1 include a control plane and a data plane. The control plane is responsible for generating routing tables, configuring security policies, and scheduling virtualization resources. The data plane is responsible for sending and receiving data.
[0070] In this application, since both CPU0 and CPU1 have network cards that are connected to them, a failure of network card 0 corresponding to CPU0 will not affect CPU1's ability to send and receive data through network card 1; thus, the stability of the server is improved.
[0071] It should be noted that the hierarchical structures corresponding to CPU0 and CPU1 in the dual single-socket server 200 are independent of each other; among them, the hierarchical structure 0 corresponding to CPU0 and the hierarchical structure 1 corresponding to CPU1 both include: hardware layer, virtualization layer and virtual machine layer.
[0072] The following section provides a detailed explanation of the hardware layer, virtualization layer, and virtual machine layer included in layer structure 0:
[0073] The hardware layer includes CPU0 and NIC0. CPU0 receives task requests from virtual machines in the virtualization layer via the virtualization layer and processes these requests. When the task request is for sending or receiving data, CPU0 sends and receives data via NIC0.
[0074] It should be understood that the hardware layer of the dual single-socket server 200 may also include other components (such as memory and / or graphics cards). For details on the server's hardware architecture, please refer to the following... Figure 12 As shown, it will not be elaborated further here.
[0075] The virtualization layer is used to virtualize physical resources in the hardware layer (i.e., CPU0 and NIC0) into virtual resources for use by virtual machines in the virtual machine layer. For example, the virtualization layer virtualizes CPU0 into at least one virtual CPU and NIC0 into a virtual network card.
[0076] The virtual machine layer deploys virtual machines based on the virtual resources virtualized by the virtualization layer. For example, the virtual CPU in virtual machine A deployed by the virtual machine layer is a virtual CPU in at least one virtual CPU obtained by the virtualization layer through performing virtualization operations on CPU0.
[0077] It should be noted that the hardware layer in hierarchy 1 includes CPU1 and network card 1; the functions of each layer in hierarchy 1 are the same as those of each layer in hierarchy 0, and will not be repeated here.
[0078] based on Figure 2 The dual single-processor server 200 shown in this embodiment also provides another system architecture for a dual single-processor server 300. For example... Figure 3The dual single-path server 300 shown, based on the dual single-path server 200, also has a communication connection between CPU0 and network card 1, and a communication connection between CPU1 and network card 0; that is, in the dual single-path server 300, both CPU0 and CPU1 have communication connections with network card 0 and network card 1 respectively.
[0079] It should be noted that in the dual-socket server 300, network interface card 0 (NIC 0) is the primary NIC for CPU 0, and NIC 1 is the backup NIC for CPU 0. When the primary NIC for CPU 0 is functioning normally, CPU 0 sends and receives data through the primary NIC, and will not use the backup NIC.
[0080] Similarly, in the dual-socket server 300, network interface card 1 (NIC 1) is the primary NIC for CPU 1, and NIC 0 is the backup NIC for CPU 1. When the primary NIC for CPU 1 is operating normally, CPU 1 sends and receives data through the primary NIC and will not send or receive data through the backup NIC.
[0081] It should be noted that the communication connection between CPU0 and NIC1 in the dual-single-path server 300 specifically refers to the connection between CPU0 and the transceiver interface of NIC1; similarly, the communication connection between CPU1 and NIC1 in the same dual-single-path server 300 is also the connection between CPU1 and the transceiver interface (usually different transceiver interfaces) of NIC1. Therefore, both CPU0 and CPU1 in the dual-single-path server 300 are connected to the transceiver interface of NIC1. Thus, when CPU0 sends data to NIC1, NIC1 uses the sending module of that transceiver interface to send the data to other devices outside the dual-single-path server 300, and does not send the data to CPU1. In other words, although both CPU0 and CPU1 have a communication connection to NIC1, CPU0 cannot send data to CPU1 through NIC1; similarly, CPU1 cannot send data to CPU0 through NIC1. This means that CPU0 and CPU1 in the dual-single-path server 300 are isolated from each other in communication.
[0082] like Figure 4 This application provides a system architecture for a dual-socket server 400; the dual-socket server 400 in... Figure 2 Based on the dual single-path server 200 shown, there is also a communication connection between CPU0 and CPU1.
[0083] It should be understood that in the dual-socket server 400, since there is a communication connection between CPU0 and CPU1, CPU0 and CPU1 can perform different steps for the same task so that CPU0 and CPU1 can jointly execute a task.
[0084] It should be noted that CPU0 and CPU1 in the dual-socket server 400 belong to the same hardware layer, and CPU0 and CPU1 correspond to the same virtualization layer and the same virtual machine layer. The descriptions of the hardware layer, virtualization layer, and virtual machine layer in this dual-socket server 400 are similar to those in the aforementioned hierarchical structure 0, and will not be repeated here.
[0085] In a dual-socket server 400, network interface card 0 (NIC 0), which has a direct communication connection with CPU 0, is the primary NIC of CPU 0. NIC 1, which has an indirect communication connection with CPU 0 through CPU 1, is the backup NIC of CPU 0. When the primary NIC of CPU 0 is functioning normally, CPU 0 sends and receives data through the primary NIC and will not use the backup NIC.
[0086] Similarly, in the dual-socket server 400, network interface card 1 (NIC 1) is the primary NIC for CPU 1, and NIC 0 is the backup NIC for CPU 1. When the primary NIC for CPU 1 is operating normally, CPU 1 sends and receives data through the primary NIC and will not send or receive data through the backup NIC.
[0087] based on Figure 4 This application also provides another system architecture for a dual-socket server 500; such as Figure 5 The dual-socket server 500 shown is based on the dual-socket server 400, and also has a communication connection between CPU0 and network card 1, and a communication connection between CPU1 and network card 0; that is, in the dual-socket server 500, both CPU0 and CPU1 have communication connections with network card 0 and network card 1 respectively.
[0088] In a dual-socket server 500, network interface card (NIC 0) is the primary NIC for CPU0, and NIC 1 is the backup NIC for CPU0. When the primary NIC for CPU0 is functioning normally, CPU0 sends and receives data through the primary NIC, not through the backup NIC. Similarly, in a dual-socket server 500, NIC 1 is the primary NIC for CPU1, and NIC 0 is the backup NIC for CPU1. When the primary NIC for CPU1 is functioning normally, CPU1 sends and receives data through the primary NIC, not through the backup NIC.
[0089] It should be noted that the dual-single-processor server 200, dual-single-processor server 300, dual-processor server 400, and dual-processor server 500 provided in this application are all acceptable. Figure 1 Devices or nodes in any of the following clusters shown: computing server cluster 110, storage server cluster 120, management server cluster 130, network device cluster 140, and tenant terminal 150.
[0090] This application provides a server management method applied to the CPUs in a server (e.g., a dual-single-socket server and / or a dual-socket server). The server includes multiple CPUs and multiple network interface cards (NICs), with the multiple CPUs and NICs having communication connections (e.g., the connection relationship between the CPUs and NICs in any of the following servers: dual-single-socket server 300, dual-socket server 400, and dual-socket server 500). Figure 6 The method shown includes: S110-S120.
[0091] S110, The first CPU obtains the fault information of the first network card.
[0092] In this application, the first CPU is any one of the multiple CPUs included in the aforementioned server, that is, the multiple CPUs include: the first CPU.
[0093] The first CPU sends and receives data through the first network card among multiple network cards; that is, the first network card is the network card among the multiple network cards that has a connection relationship with the first CPU and is used to send and receive data for the first CPU.
[0094] It should be noted that when the first CPU executes S110 for the first time, the first network card is the main network card of the first CPU; when the first CPU executes S110 for a period of time, the first network card is the network card used to send and receive data when the first CPU previously executed S120.
[0095] It should be understood that, if the primary network card corresponding to the first CPU does not fail, the first CPU will send and receive data through the primary network card and will not send and receive data through any of the multiple network cards other than the primary network card (i.e., the backup network card).
[0096] In this application, the fault information is used to indicate faults that prevent the network card from functioning properly. For example, the fault information indicates faults such as poor network card contact, abnormal network card driver, and abnormal network card system configuration.
[0097] In this application, S110 can be implemented by the first CPU obtaining fault information of the first network card from the first network card, or by the first CPU obtaining fault information of the first network card from the baseboard management controller (BMC) in the aforementioned server; the specific implementation of S110 in this application embodiment is not specifically limited.
[0098] When the first CPU obtains the fault information of the first network card from the first network card, the specific implementation of S110 can be that the first CPU sends a fault request message to the first network card, and the first network card responds to the fault request message by sending the fault information of the fault that occurred in the first network card to the first CPU; or the first network card actively sends the fault information of the fault to the first CPU when the first network card fails; specifically, the embodiments of this application do not limit the specific implementation method of obtaining the above-mentioned fault information from the first network card.
[0099] When the first CPU obtains the fault information of the first network card from the BMC, the specific implementation of S110 includes: when the BMC detects that the first network card has failed, the BMC reads the fault information of the first network card; and then sends the fault information to the first CPU.
[0100] In this embodiment, the fault information of the first network card is obtained from the BMC by the first CPU. Since the BMC reads the fault information from the first network card, the problem of the first network card being unable to send the fault information due to the fault of the first network card is avoided, thus improving the reliability of obtaining the fault information of the first network card.
[0101] S120: The first CPU responds to the fault information and sends and receives data through the second network card.
[0102] The second network interface card (NIC) in this application is the NIC that has a communication connection with the first CPU among the multiple NICs included in the aforementioned server, excluding the first NIC; that is, the second NIC is the NIC that has a communication connection with the first CPU among the multiple NICs and has not malfunctioned.
[0103] It should be understood that, in the case where the first network card is the main network card of the first CPU, the second network card in this application is the backup network card in the first CPU.
[0104] The above server is Figure 3 The dual single-path server 300 shown, or the server is... Figure 5 In the case of the dual-socket server 500 shown, if the first CPU is CPU0, then the first network card is network card 0; the second network card is CPU1.
[0105] The above server is Figure 4 In the case of the dual-socket server 400 shown, if the first CPU is CPU0, then the first network card is network card 0 in the dual-socket server 400 that has a direct communication connection with CPU0; the second network card is network card 1 in the dual-socket server 400 that has an indirect communication connection with CPU0 through CPU1.
[0106] It should be noted that the number of second network interface cards (NICs) in this application can be one or more. When there is only one second NIC, S120 is implemented such that the first CPU directly sends and receives data through the second NIC. When there are multiple second NICs, S120 can be implemented such that the first CPU sends and receives data through multiple second NICs based on a load balancing strategy, as in method 1; or it can be implemented such that the first CPU sends and receives data through multiple second NICs based on an available load percentage strategy, as in method 2.
[0107] Method 1: The first CPU distributes the task of sending and receiving data to multiple second network cards on an equal basis, so that each of the multiple second network cards sends and receives data for the first CPU based on the assigned data sending and receiving task.
[0108] For example, assuming there are 3 second network cards and 6 data transmission and reception tasks for the first CPU, then the first CPU allocates 2 data transmission and reception tasks to each second network card so that each second network card can perform the allocated data transmission and reception tasks.
[0109] Method 2: The first CPU allocates the data transmission and reception tasks of the first CPU to the multiple second network cards according to the ratio of available load among the multiple second network cards, so that each of the multiple second network cards can transmit and receive data for the first CPU based on the allocated data transmission and reception tasks.
[0110] It should be understood that the available load of a network interface card (NIC) is the difference between its rated load and its current load; in other words, the available load of a NIC is its unused load. For example, assuming NIC A has a rated load of 10Gbps and its current load is 6Gbps, then the available load of NIC A is 4Gbps.
[0111] For example, suppose there are multiple second network interface cards (NICs) named NIC A, NIC B, and NIC C; where the available load of NIC A is 5Gbps, the available load of NIC B is 3Gbps, and the available load of NIC C is 2Gbps; then the available load ratio of NIC A, NIC B, and NIC C is 5:3:2.
[0112] Assuming the first CPU has 10 data transmission and reception tasks, then 5 data transmission and reception tasks are assigned to the second network interface card (NIC A) so that NIC A can execute these 5 data transmission and reception tasks; 3 data transmission and reception tasks are assigned to the second network interface card (NIC B) so that NIC B can execute these 3 data transmission and reception tasks; and 2 data transmission and reception tasks are assigned to the second network interface card (NIC C) so that NIC C can execute these 2 data transmission and reception tasks; thereby enabling the first CPU to transmit and receive data through the second network interface cards (NICs A, B, and C).
[0113] According to the embodiments of this application, the task of sending and receiving data by the first CPU is allocated to the multiple second network cards based on the ratio of available load among the multiple second network cards, so that each of the multiple second network cards can send and receive data for the first CPU based on the allocated task of sending and receiving data, thereby improving the load balancing of the second network cards.
[0114] It should be noted that the second network interface card (NIC) in this application can be any of the multiple NICs that have a communication connection with the first CPU and are not malfunctioning (i.e., other than the first NIC), or it can be the NIC that meets the following conditions (e.g., minimum load or shortest communication distance) among the multiple NICs that have a communication connection with the first CPU and are not malfunctioning. Wherein, if the second NIC is the NIC that meets the following conditions (e.g., minimum load or shortest communication distance) among the NICs that have a communication connection with the first CPU and are not malfunctioning, the method for determining the second NIC is described in S210-S220 or S310-S320 below, and will not be repeated here.
[0115] This application provides a server management method. This method configures multiple network interface cards (NICs) for multiple CPUs in a multi-processor server, ensuring communication connections between the CPUs and these NICs. If the first NIC used by any CPU for sending and receiving data fails, the first CPU sends and receives data through a second NIC (other than the first NIC) that is also connected to it. This solves the problem of all CPUs in the server being unable to process business operations normally due to NIC failure, thus improving server stability.
[0116] In addition, since the method allows the first CPU to send and receive data through the second network card when the first network card corresponding to the first CPU fails, it avoids the problem that the first CPU cannot process business due to the failure of the first network card, thereby reducing the impact of the failure of the first network card on the performance of the entire server (i.e., reducing the explosion radius); thus, it further improves the stability of the server.
[0117] It should be noted that if there are multiple network cards (excluding the first network card) that have communication connections with the first network card (e.g., N), then some of these N network cards may be at full load (or about to reach) their rated load. If the first CPU is then sent and received data based on this full-load network card, it will cause the full-load network card to be overloaded, thereby reducing the performance of the server and, in severe cases, causing the server to crash.
[0118] Furthermore, when the server management method provided in this application is applied to the CPU in a multi-processor server (e.g., a dual-processor server), since the N network cards are network cards corresponding to the CPUs that have communication connections with the first CPU, there is an indirect communication connection between the first CPU and the N network cards. If the communication distance between the first CPU and the second network card among the N network cards is large, it will reduce the efficiency of the first CPU in sending and receiving data through the second network card.
[0119] Based on this, the following describes the method for determining the second network card from the perspective of network card load and / or network card communication distance (i.e., the communication distance between the network card and the first CPU), based on the server management method provided above. See Examples 1 to 3 below for details; the main difference between Examples 1 and 3 is:
[0120] Example 1 can be applied to multiple single-processor servers (e.g., dual single-processor servers) as well as multi-processor servers (e.g., dual-processor servers). In Example 1, the first CPU determines the network card among the N network cards that meets the load condition (e.g., minimum load) as the second network card. Thus, the first CPU sends and receives data through this second network card with the minimum load, thereby improving the load balancing of the multiple network cards and thus improving the stability of the server.
[0121] Example 2 applies to a dual-processor server. In Example 2, the first CPU identifies the network interface card (NIC) of the second CPU among multiple CPUs that meets the communication distance requirement (e.g., the shortest communication distance) as the second NIC. Thus, the first CPU sends and receives data based on the second NIC with the shortest communication distance, thereby improving the efficiency of data transmission and reception for the first CPU.
[0122] Example 3 is applicable to a dual-processor server. In Example 3, the first CPU determines the candidate network interface card (NIC) from the aforementioned N NICs that meets the load requirements; then, it determines a second NIC from the candidate NICs that meets the communication distance requirements. In this way, the first CPU sends and receives data based on the second NIC, which meets both the load and communication distance requirements, thus improving not only the performance of the first CPU in sending and receiving data but also the stability of the server.
[0123] Example 1
[0124] Combination Figure 6 This application provides a specific implementation method for server management, such as... Figure 7 The embodiment shown includes S210-S220 before S120.
[0125] S210: The first CPU obtains the load of the candidate second network card that has a communication connection with the first CPU among the network cards other than the first network card.
[0126] In this application, multiple network interface cards (NICs) can be multiple single-processor servers (e.g., ...). Figure 3 The dual-single-path server 300 shown can contain multiple network cards; it can also be a multi-path server (e.g., a single-path server). Figure 4 The dual-socket server 400 shown, or Figure 5 The diagram shows multiple network cards in a dual-socket server (500).
[0127] It should be understood that, in the case where the multiple network interface cards (NICs) are those of a multi-socket server, there is no communication connection between the multiple CPUs in this application, and any one of the multiple CPUs in this application (e.g., the first CPU) has a direct communication connection with at least two of the multiple NICs (e.g., ...). Figure 3 (As shown). In the case where these multiple network cards are network cards in a multi-processor server, there is a communication connection between the multiple CPUs in this application (e.g., Figure 4 or Figure 5 (As shown).
[0128] Method 1: In the case where multiple network cards are network cards of a multi-single-path server, the candidate second network card includes: the network card other than the first network card among at least two network cards that have a direct communication connection with the first CPU; that is, the candidate second network card is the network card that has not failed among at least two network cards that have a direct communication connection with the first CPU.
[0129] Method 2: When multiple network cards are network cards of a multi-processor server, the candidate second network card includes: the network card corresponding to the CPU that has a communication connection with the first CPU among the multiple CPUs in the multi-processor server; that is, the candidate second network card is the network card that has an indirect communication connection with the first CPU among the multiple network cards.
[0130] It should be understood that, due to Figure 5 Each of the multiple CPUs in the multi-processor server 400 shown has a direct communication connection with at least two of the multiple network interface cards (NICs); therefore, when the multiple NICs are NICs of the multi-processor server 400, the candidate second NIC includes: NICs other than the first NIC among the NICs that have a direct communication connection with the first CPU (i.e., NICs that have not failed), and / or, the NIC corresponding to the CPU among the multiple CPUs in the multi-processor server 400 that has a communication connection with the first CPU.
[0131] In this application, S210 can be implemented by the first CPU directly obtaining the load of the candidate second network card from the candidate second network card, or by obtaining the load of the candidate second network card from the BMC; the specific implementation of S210 in this application embodiment does not limit the implementation of S210.
[0132] Optionally, before S210, the method further includes: the first CPU acquiring a candidate second network card; specifically, the first CPU determines a candidate second network card that has a communication connection with the first CPU from a preset mapping table used to indicate the connection relationship between the CPU and the network card.
[0133] S220: The first CPU determines the network card that meets the load conditions from the candidate second network cards as the second network card.
[0134] For example, suppose the candidate second network interface cards (NICs) include NICs 1 to NICs 3; wherein NIC 1 has a load of 2Gbps, NIC 2 has a load of 3Gbps, and NIC 3 has a load of 5Gbps. If the above condition is used to indicate the NIC with the lowest load, then NIC 1 among the candidate second NICs is determined as the second NIC. If the above condition is used to indicate a NIC with a load less than 4Gbps, then NICs 2 and NIC 1 among the candidate second NICs are determined as the second NIC.
[0135] In the case of a failure of the first network card corresponding to the first CPU, this application embodiment obtains the load of candidate second network cards that have a communication connection with the first CPU from among the network cards other than the first network card, and determines the second network card as the network card whose load meets the condition (e.g., the smallest load) among the candidate second network cards; so that the first CPU can send and receive data through the second network card with the smallest load; thereby improving the load balancing of multiple network cards.
[0136] Example 2
[0137] Combination Figure 6 This application provides another specific implementation method for server management, which is applied to CPUs in a multi-processor server where multiple CPUs communicate with each other; such as Figure 8 As shown, the method further includes S310-S320 before S120.
[0138] S310: Obtain the topology relationship of multiple CPUs.
[0139] The topological relationship of multiple CPUs in this application is used to describe the connection relationship between these multiple CPUs (i.e., the relationship of having a communication connection).
[0140] In this application, S310 can be implemented by the first CPU obtaining the topology relationship from the storage medium of the multi-way server, or by the first CPU obtaining the topology relationship from other devices. The specific implementation of this application does not limit the method of obtaining the topology relationship.
[0141] S320: Based on the topology of multiple CPUs, the first CPU determines the network card corresponding to the second CPU among the multiple CPUs that meets the communication distance condition with the first CPU as the second network card.
[0142] In this application, the implementation of S320 includes: determining a third CPU that has a communication connection with the first CPU based on the above topology; then, determining the CPU in the third CPU whose communication distance with the first CPU meets the condition (e.g., the shortest communication distance) as the second CPU, and using the network card corresponding to the second CPU as the second network card.
[0143] In one implementation, the communication distance in this application is used to indicate the number of hops during the communication process.
[0144] For example, suppose the topology of multiple CPUs is as follows: Figure 9 As shown, based on the topology of multiple CPUs, the third CPUs that have communication connections with the first CPU include CPU1 to CPU3. CPU1 has a direct communication connection with the first CPU, CPU2 has an indirect communication connection with the first CPU through CPU1, and CPU3 has an indirect communication connection with the first CPU through both CPU1 and CPU2. Therefore, when the above condition is used to indicate the network interface card (NIC) with the shortest communication distance, since CPU1 is directly connected to the first CPU (i.e., the hop count is 0), the first CPU identifies CPU1 as the second CPU, meaning the NIC corresponding to CPU1 is the second NIC. When the above condition is used to indicate NICs with a communication distance of less than 2 hops, since the hop count between CPU1 and the first CPU is 0, the hop count between CPU2 and the first CPU is 1, and the hop count between CPU3 and the first CPU is 2, the first CPU identifies CPU1 and CPU2 as the second CPUs, meaning the NICs corresponding to CPU1 and CPU2 are each the second NIC.
[0145] In the case of a failure of the first network card corresponding to the first CPU, the first CPU uses the topology of multiple CPUs in its server to select the network card corresponding to the second CPU among the multiple CPUs that meets the communication distance condition (e.g., the shortest communication distance) as the second network card. This allows the first CPU to send and receive data based on the second network card with the shortest communication distance, thereby improving the efficiency of the first CPU in sending and receiving data.
[0146] Example 3
[0147] Combination Figure 6 This application also provides another specific implementation method for server management, which is applied to CPUs in a multi-processor server where multiple CPUs communicate with each other; such as Figure 10As shown, the method further includes S410-S430 before S120.
[0148] S410: The first CPU obtains the load of the candidate second network card that has a communication connection with the first CPU among the network cards other than the first network card.
[0149] It should be understood that the aforementioned candidate second network card is the network card corresponding to the CPU that has a communication connection with the first CPU among multiple CPUs; that is, the candidate second network card is the network card that has an indirect communication connection with the first CPU among multiple network cards.
[0150] S420, the first CPU determines the network card that meets the load conditions from the candidate second network cards as the candidate network card.
[0151] It should be noted that the implementation of S410-S420 is similar to that of S210-S220. For a detailed description of S410-S420, please refer to the above description of S210-S220. It will not be repeated here.
[0152] S430: Based on the topology of multiple CPUs, the first CPU determines the network card corresponding to the second CPU among the CPUs corresponding to the candidate network cards, whose communication distance with the first CPU meets the condition, as the second network card.
[0153] It should be noted that the implementation of S430 is similar to that of S320. For a detailed description of S430, please refer to the above description of S320. It will not be repeated here.
[0154] In this embodiment, when the first network interface card (NIC) corresponding to the first CPU fails, a candidate NIC is determined from the candidate second NICs based on their load, provided that the load meets the specified conditions (e.g., minimum load). Then, based on the topology of multiple CPUs, the NIC corresponding to the second CPU whose communication distance to the first CPU meets the specified conditions is determined as the second NIC. Since this second NIC meets both the load and communication distance conditions (e.g., minimum), the first CPU's data transmission and reception based on this second NIC not only improves the load balancing of the multiple NICs but also increases the data transmission and reception efficiency.
[0155] The foregoing mainly describes the solutions provided by the embodiments of this application from a methodological perspective. To achieve the above functions, the server management device includes hardware structures and / or software modules corresponding to the execution of each function. Those skilled in the art should readily recognize that, based on the modules and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein, this application can be implemented in hardware or a combination of hardware and computer software. Whether a function is executed in hardware or by computer software driving hardware depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0156] This application embodiment can, according to the above method, exemplarily divide the server management device into functional modules. For example, the server management device may include functional units corresponding to each functional division, or two or more functions may be integrated into one processing unit. The integrated unit can be implemented in hardware or as a software functional unit. It should be noted that the module division in this application embodiment is illustrative and only represents one logical functional division; in actual implementation, there may be other division methods.
[0157] When dividing each functional unit according to its corresponding function. Figure 11 This diagram illustrates a possible structural design of the server management device involved in the above embodiments; the server management device is deployed in... Figures 3 to 5 The CPU in the server shown; the management device of the server includes: transceiver module 1101 and processing module 1102.
[0158] The transceiver module 1101 is used to obtain fault information of the first network card; for example, by executing step S110 in the above method embodiment.
[0159] The processing module 1102 is used to send and receive data through the second network card in response to fault information; for example, it executes step S120 in the above method embodiment.
[0160] Optionally, the first network card is the primary network card of the first CPU; the second network card is the backup network card of the first CPU; wherein, if the primary network card does not fail, the first CPU sends and receives data through the primary network card and does not send and receive data through the backup network card.
[0161] Optionally, there is a communication connection between the multiple CPUs in the server; or, there is no communication connection between the multiple CPUs, but each of the multiple CPUs has a communication connection with at least two of the multiple network interface cards (NICs).
[0162] Optionally, the transceiver module 1101 is used to obtain the load of the candidate second network card that has a communication connection with the first CPU among other network cards; for example, by executing step S210 in the above method embodiment.
[0163] Processing module 1102 is used to determine the network card that meets the load condition among the candidate second network cards as the second network card; for example, it executes step S220 in the above method embodiment.
[0164] Optionally, when there is a communication connection between multiple CPUs, the candidate second network interface card (NIC) includes: the NIC corresponding to the CPU that has a communication connection with the first CPU among the multiple CPUs in the server; or, when there is no communication connection between the multiple CPUs, and the first CPU has a communication connection with at least two of the multiple NICs, the candidate second NIC includes: the NIC other than the first NIC among the at least two NICs that have a communication connection with the first CPU.
[0165] Optionally, the processing module 1102 is used to determine the network card corresponding to the second CPU among the multiple CPUs that meets the communication distance condition with the first CPU as the second network card based on the topology relationship of multiple CPUs; for example, by executing step S320 in the above method embodiment.
[0166] Optionally, the processing module 1102 is used to distribute the task of sending and receiving data from the first CPU to the multiple second network cards according to the ratio of available load among the multiple second network cards.
[0167] Optionally, a transceiver module is provided to obtain fault information of the first network card from the BMC in the server.
[0168] Both the transceiver module 1101 and the processing module 1102 can be implemented in software or in hardware. For example, the implementation of the processing module 1102 will be described below. The implementation of the transceiver module 1101 can be referenced from the implementation of the processing module 1102.
[0169] As an example of a software functional unit, processing module 1102 may include code running on a computing instance. The computing instance may include at least one of a physical host (computing device), a virtual machine, and a container. Furthermore, the aforementioned computing instance may be one or more. For example, processing module 1102 may include code running on multiple hosts / virtual machines / containers.
[0170] It should be noted that the multiple hosts / virtual machines / containers used to run this code can be distributed within the same region or in different regions. Furthermore, the multiple hosts / virtual machines / containers used to run this code can be distributed within the same availability zone (AZ) or in different AZs, each AZ comprising one or more geographically proximate data centers. Typically, a region can include multiple AZs.
[0171] Similarly, multiple hosts / virtual machines / containers used to run this code can be distributed within the same Virtual Private Cloud (VPC) or across multiple VPCs. Typically, a VPC is set up within a region. Communication between two VPCs within the same region, as well as between VPCs in different regions, requires a communication gateway to be set up within each VPC to enable interconnection between VPCs.
[0172] As an example of a hardware functional unit, the processing module 1102 may include at least one computing device, such as a server. Alternatively, the processing module 1102 may also be a device implemented using an application-specific integrated circuit (ASIC) or a programmable logic device (PLD). The PLD may be implemented using a complex programmable logical device (CPLD), a field-programmable gate array (FPGA), generic array logic (GAL), or any combination thereof.
[0173] The processing module 1102 includes multiple computing devices that can be distributed in the same region or in different regions. Similarly, the processing module 1102 can be distributed in the same Availability Zone (AZ) or in different AZs. Likewise, the processing module 1102 can be distributed in the same Virtual Private Cloud (VPC) or in multiple VPCs. These multiple computing devices can be any combination of computing devices such as servers, ASICs, PLDs, CPLDs, FPGAs, and GALs.
[0174] It should be noted that, in other embodiments, the processing module 1102 can be used to execute any step in the above-described server management method, and the transceiver module 1101 can be used to execute any step in the above-described server management method. The steps implemented by the transceiver module 1101 and the processing module 1102 can be specified as needed. By implementing different steps in the above-described server management method through the transceiver module 1101 and the processing module 1102 respectively, all functions of the management device can be realized.
[0175] This application also provides a computing device, which is specifically a computing device. Figures 2 to 5 Any server in the network; such as Figure 12 The computing device shown may include multiple processors 301, a memory 302, and a communication interface 303. The processors 301, memory 302, communication interface 303, and multiple physical network interface cards 305 can be connected via a bus 304 or other means. This computing device may be a server or a terminal device. It should be understood that this application does not limit the number of processors and memory in the computing device.
[0176] Each of the multiple processors 301 may include any one or more processors such as a central processing unit (CPU), a graphics processing unit (GPU), a microprocessor (MP), or a digital signal processor (DSP).
[0177] The memory 302 includes, but is not limited to, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), flash memory, or optical memory, disk storage media or other magnetic storage devices, or any other medium capable of carrying or storing desired program code in the form of instructions or data structures and accessible by a computer.
[0178] The memory 302 stores executable program code, and the processor 301 executes the executable program code to implement the functions of the aforementioned transceiver module 1101 and processing module 1102, thereby realizing the server management method. That is, the memory 302 stores instructions for executing the server management method.
[0179] Alternatively, the memory 302 stores executable code, which the processor 301 executes to implement the functions of the aforementioned server management device, thereby implementing the server management method. That is, the memory 302 stores instructions for executing the server management method.
[0180] In one possible implementation, the memory 302 may exist independently of the processor 301. The memory 302 can be connected to the processor 301 via a bus 304 and is used to store data, instructions, or program code. When the processor 301 calls and executes the instructions or program code stored in the memory 302, it can implement the relevant steps in the server management method provided in the embodiments of this application.
[0181] In another possible implementation, the memory 302 can also be integrated with the processor 301.
[0182] The communication interface 303 can be an acquisition module used to communicate with other devices or communication networks, such as Ethernet, RAN, and wireless local area networks (WLAN). The communication interface 303 can receive commands, messages, or data. The acquisition module can be a transceiver or similar device.
[0183] Optionally, the communication interface 303 can also be a transceiver circuit located within the processor 301, used to implement signal input and signal output of the processor 301, such as fault information of the first network card. The communication interface 303 can be a wired interface (port), such as a fiber distributed data interface (FDDI) or a gigabit Ethernet (GE) interface, or the communication interface 303 can also be a wireless interface.
[0184] Each of the multiple physical network interface cards (NICs) 305 includes, but is not limited to, a peripheral component interconnect (PCI) NIC, a peripheral component interconnect express (PCIe) NIC, and a universal serial bus (USB) NIC. In this application, the physical NICs are used to send and receive data for the processor 301.
[0185] Bus 304 can be an industry standard architecture (ISA) bus, a peripheral component interconnect (PCI) bus, or an extended industry standard architecture (EISA) bus. This bus can be divided into address bus, data bus, control bus, etc. The bus can also be divided into serial bus and parallel bus. For ease of representation, Figure 12 The bus is represented by a single thick line, but this does not mean that there is only one bus or one type of bus.
[0186] It should be understood that, Figure 12 The computing device mentioned is merely one example of a computing device; it can have more than Figure 12 The more or fewer components shown can be combined into two or more components, or they can have different component configurations. For example, a computing device can also include a smart network card, such as a data processing unit (DPU).
[0187] This application also provides a computing device cluster. The computing device cluster includes at least one computing device. The computing device can be a server, such as a central server, an edge server, or a local server in a local data center. In some embodiments, the computing device can also be a terminal device such as a desktop computer, a laptop computer, or a smartphone.
[0188] like Figure 13 As shown, the computing device cluster includes at least one computing device 100. The memory 302 of one or more computing devices 100 in the computing device cluster may store the same instructions for executing the aforementioned server management methods.
[0189] In some possible implementations, each computing device 100 in the computing device cluster stores a portion of the instructions for executing the aforementioned server management methods in its memory 302. In other words, each computing device 100 independently executes the instructions for the aforementioned server management methods.
[0190] It should be understood that the processor and bus in the aforementioned computing device 100 are related to... Figure 12 The processor 301 and bus 304 in this computing device 100 are identical. For a detailed description of the processor and bus in this computing device 100, please refer to the description of the processor and bus in this computing device 100. Figure 12 The relevant descriptions will not be repeated here.
[0191] In some possible implementations, one or more computing devices in a computing device cluster can be connected via a network. This network can be a wide area network (WAN) or a local area network (LAN), etc.
[0192] This application also provides a computing device cluster. The connection relationships between the computing devices in this computing device cluster can be similarly referred to... Figure 13 The connection method of the computing device cluster is different in that the memory 302 of one or more computing devices 100 in the computing device cluster can store the same instructions for executing server management methods.
[0193] This application also provides a computer program product containing instructions. The computer program product may be a software or program product containing instructions, capable of running on a computing device or stored on any usable medium. When the computer program product is run on at least one computing device, it causes the at least one computing device to execute a server management method.
[0194] This application also provides a computer-readable storage medium. The computer-readable storage medium can be any available medium that a computing device can store, or a data storage device such as a data center that includes one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid-state drive). The computer-readable storage medium includes instructions that instruct the computing device to execute a server management method.
[0195] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the protection scope of the technical solutions of the embodiments of the present invention.
Claims
1. A server management method, characterized in that, The server includes: multiple CPUs and multiple network interface cards (NICs), wherein the multiple CPUs and the multiple NICs are communicatively connected; wherein a first CPU among the multiple CPUs sends and receives data through a first NIC among the multiple NICs; the method includes: The first CPU acquires fault information of the first network card; the fault information is used to indicate that the network card cannot work properly. In response to the fault information, the first CPU sends and receives data through the second network card; wherein, the second network card is one of the network cards other than the first network card that has a communication connection with the first CPU.
2. The method according to claim 1, characterized in that, The first network card is the primary network card of the first CPU; the second network card is the backup network card of the first CPU; wherein, when the primary network card is not faulty, the first CPU sends and receives data through the primary network card and does not send and receive data through the backup network card.
3. The method according to claim 1 or 2, characterized in that, There are communication connections between the multiple CPUs; Alternatively, there is no communication connection between the plurality of CPUs, but each of the plurality of CPUs has a communication connection with at least two of the plurality of network cards.
4. The method according to claim 3, characterized in that, The method further includes: The first CPU acquires the load of the candidate second network interface card (NIC) that has a communication connection with the first CPU among the other NICs; The first CPU determines the network card that meets the load condition from the candidate second network cards as the second network card.
5. The method according to claim 4, characterized in that, In the case where there is a communication connection between the plurality of CPUs, the candidate second network interface card includes: the network interface card corresponding to the CPU that has a communication connection with the first CPU among the plurality of CPUs; or, When there is no communication connection between the plurality of CPUs, and the first CPU has a communication connection with at least two of the plurality of network cards, the candidate second network card includes: the network card other than the first network card among the at least two network cards that have a communication connection with the first CPU.
6. The method according to claim 3, characterized in that, In the case where there is a communication connection between the plurality of CPUs, the method further includes: Based on the topology of the plurality of CPUs, the first CPU determines the network card corresponding to the second CPU among the plurality of CPUs whose communication distance with the first CPU meets the condition as the second network card.
7. The method according to any one of claims 1-6, characterized in that, When there are multiple second network interface cards (NICs), the process before sending and receiving data through the second NICs includes: The first CPU distributes the data transmission and reception tasks of the first CPU to the multiple second network cards according to the ratio of available load among the multiple second network cards.
8. The method according to any one of claims 1-7, characterized in that, The first CPU obtains fault information of the first network card, including: The first CPU obtains fault information of the first network interface card (NIC) from the baseboard management controller (BMC) in the server; wherein, the BMC is used to read the fault information from the first NIC.
9. A server management device, characterized in that, The server includes: multiple CPUs and multiple network interface cards (NICs), wherein the multiple CPUs and the multiple NICs are communicatively connected; wherein a first CPU among the multiple CPUs sends and receives data through a first NIC among the multiple NICs; the server management device is applied to the first CPU; the server management device includes: a transceiver module and a processing module; The transceiver module is used to acquire fault information of the first network card; the fault information is used to indicate that the network card cannot work properly. The processing module is used to send and receive data via a second network interface card (NIC) in response to the fault information; wherein, the second NIC is one of the NICs other than the first NIC that has a communication connection with the first CPU.
10. The server management device according to claim 9, characterized in that, The first network card is the primary network card of the first CPU; the second network card is the backup network card of the first CPU; wherein, when the primary network card is not faulty, the first CPU sends and receives data through the primary network card and does not send and receive data through the backup network card.
11. The server management device according to claim 9 or 10, characterized in that, There are communication connections between the multiple CPUs; Alternatively, there is no communication connection between the plurality of CPUs, but each of the plurality of CPUs has a communication connection with at least two of the plurality of network cards.
12. The server management device according to claim 11, characterized in that, The transceiver module is used to obtain the load of the candidate second network interface card (NIC) among the other NICs that has a communication connection with the first CPU; The processing module is used to determine the network card that meets the load condition from the candidate second network cards as the second network card.
13. The server management device according to claim 12, characterized in that, In the case where there is a communication connection between the plurality of CPUs, the candidate second network interface card includes: the network interface card corresponding to the CPU that has a communication connection with the first CPU among the plurality of CPUs; or, When there is no communication connection between the plurality of CPUs, and the first CPU has a communication connection with at least two of the plurality of network cards, the candidate second network card includes: the network card other than the first network card among the at least two network cards that have a communication connection with the first CPU.
14. The server management device according to claim 11, characterized in that, The processing module is used to determine the network card corresponding to the second CPU among the multiple CPUs that meets the communication distance condition with the first CPU as the second network card, based on the topology relationship of the multiple CPUs.
15. The server management device according to any one of claims 9-14, characterized in that, The processing module is used to allocate the task of sending and receiving data from the first CPU to the multiple second network cards according to the ratio of available load among the multiple second network cards.
16. The server management device according to any one of claims 9-15, characterized in that, The transceiver module is used to obtain fault information of the first network card from the BMC in the server; wherein the BMC is used to read the fault information from the first network card.
17. A computing device cluster, characterized in that, It includes at least one computing device, each computing device including a processor and memory; The processor of the at least one computing device is configured to execute instructions stored in the memory of the at least one computing device to cause the cluster of computing devices to perform the method as described in any one of claims 1-8.
18. A computer-readable storage medium, characterized in that, The device stores computer instructions that, when executed on a computing device, cause the computing device to perform the method as described in any one of claims 1-8.
19. A computer program product containing instructions, characterized in that, When the instruction is executed by the computing device cluster, the computing device cluster causes the computing device cluster to perform the method as described in any one of claims 1-8.