A hardware fault detection method, system and related device
By utilizing a fault detection controller to create fault detection groups and elect a master component in a cloud computing environment, the problem of high resource consumption in hardware fault detection management is solved, thereby improving resource utilization and reducing operation and maintenance costs.
Patent Information
- Application Number
- CN202011056417.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-09-30
- Publication Date
- 2026-01-27
- Estimated Expiration
- 2040-09-30
AI Technical Summary
Existing technologies for hardware fault detection in cloud computing consume enormous management resources, leading to high complexity and increased operation and maintenance costs, especially in large-scale data centers and massive small site scenarios where effective management is difficult.
The fault detection controller obtains hardware-related information, creates a fault detection group, elects a master fault detection component within the group, and uses heartbeats to detect hardware faults, avoiding the need to deploy separate management nodes and reducing management resource consumption.
It reduced management resource consumption, improved resource utilization, lowered operation and maintenance costs, and ensured the normal operation of business and the reliability of the system.
Smart Images

Figure CN114328036B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of cloud computing technology, and in particular to a hardware fault detection method, system and related equipment. Background Technology
[0002] Cloud computing, as an emerging industry in recent years, has gained widespread attention from the scientific and industrial communities. Its global rise, characterized by its flexible, efficient, low-cost, and energy-saving operation, has made it a crucial engine for promoting green industrial development and a new business platform for the 21st century. Cloud computing distributes computing tasks across a resource pool comprised of numerous servers, enabling various application systems to access computing power, storage space, and various cloud services as needed. When servers or switches fail, the cloud management platform needs to quickly detect hardware failures, rapidly restore computing resources, and ensure the continued operation of application systems.
[0003] With the development of cloud computing, data centers are growing in scale, with an increasing number of servers and switches. To meet real-time requirements, the management system needs to be scaled up. Current solutions, such as hierarchical networking, are complex to configure and consume huge amounts of management resources. Furthermore, when applications require proximity access to computing resources, necessitating the construction of tens of thousands of small-scale edge sites, each edge site requires the separate deployment of fault detection components, further exacerbating the complexity of management configuration and the consumption of management resources.
[0004] Therefore, how to reduce the management resource consumption caused by fault detection, improve resource utilization, and reduce operation and maintenance costs is an urgent problem to be solved. Summary of the Invention
[0005] This invention discloses a hardware fault detection system, method, and related equipment, which can reduce the occupation of management resources, improve resource utilization, and reduce operation and maintenance costs.
[0006] In a first aspect, this application provides a hardware fault detection method, including a fault detection controller and multiple fault detection components, wherein the multiple fault detection components are deployed in multiple physical servers, wherein: the fault detection controller is used to acquire hardware-related information and create a fault detection group based on the hardware-related information, the fault detection group including at least two fault detection components; the fault detection components in the fault detection group elect a master fault detection component, the master fault detection component being used to perform fault detection on the server cluster corresponding to the fault detection group.
[0007] In the solution provided in this application, the fault detection controller creates a fault detection group based on the obtained hardware-related information, and performs fault detection on the server cluster corresponding to the fault detection group by electing a master fault detection component within the fault detection group. This avoids deploying a separate management node to perform fault detection on the server cluster, reduces the occupation of management resources, improves resource utilization, and reduces operation and maintenance costs.
[0008] In conjunction with the first aspect, in one possible implementation of the first aspect, the hardware-related information includes data center location information and topology information, switch topology information, rack location information, and server location information.
[0009] In the solution provided in this application, the fault detection controller ensures that fault detection groups can be created correctly and reasonably by pre-acquiring hardware-related information such as data center location and topology information, switch topology information, rack location information, and server location information.
[0010] In conjunction with the first aspect, in one possible implementation of the first aspect, the fault detection component is deployed in an offload card, which is inserted into the physical server.
[0011] In the solution provided in this application, by deploying the fault detection component in the offload card, the occupation of server management resources can be further reduced and the utilization rate of server resources can be improved.
[0012] In conjunction with the first aspect, in one possible implementation of the first aspect, the fault detection controller divides servers within the same physical location into the same fault detection group; or, the fault detection controller divides servers with the same physical attributes into the same fault detection group; or, the fault detection controller divides servers with the same fault detection requirements into the same fault detection group; or, the fault detection controller divides a preset number of servers into the same fault detection group.
[0013] In the solution provided in this application, the fault detection controller can divide the server to be detected according to the actual detection needs, thereby obtaining different fault detection groups, and further realize fault detection within the fault detection group.
[0014] In conjunction with the first aspect, in one possible implementation of the first aspect, the fault detection components in the fault detection group establish a connection by sending heartbeats to each other and elect the master fault detection component through a preset algorithm; or, the fault detection components in the fault detection group establish a connection by sending heartbeats to each other and elect a master fault detection component cluster through a preset algorithm, wherein the master fault detection component cluster includes at least one master fault detection component.
[0015] In the solution provided in this application, the fault detection components in the fault detection group establish a connection by sending heartbeats to each other, and further elect a master fault detection component or a master fault detection component cluster through a preset algorithm. This enables the master fault detection component to perform fault detection on all servers in the group. When a master fault detection component cluster is elected, it can be guaranteed that if one master fault detection component fails (i.e. cannot perform fault detection normally), other master fault detection components in the master fault detection component cluster can promptly perform fault detection on the servers in the group, ensuring normal business operation and improving system reliability.
[0016] In conjunction with the first aspect, in one possible implementation of the first aspect, when a new fault detection component joins the fault detection group, the main fault detection component receives a heartbeat sent by the new fault detection component and broadcasts it to the fault detection components in the fault detection group, so that the fault detection components in the fault detection group can establish a connection with the new fault detection component.
[0017] In the solution provided in this application, when a new fault detection component needs to be added, the fault detection controller can add the new fault detection component by creating a new fault detection group or by adding the new fault detection component to an existing fault detection group. The controller can also establish a connection between the main fault detection component and the new fault detection component by sending a heartbeat to the main fault detection component in the group, thereby improving the flexibility and scalability of the system.
[0018] Secondly, this application provides a hardware fault detection method. The method includes a fault detection controller acquiring hardware-related information and creating a fault detection group based on the hardware-related information. The fault detection group includes at least two fault detection components, which are deployed in a physical server. The fault detection components in the fault detection group elect a master fault detection component, which is used to perform fault detection on the server cluster corresponding to the fault detection group.
[0019] In the solution provided in this application, the fault detection controller uses the acquired hardware-related information to create a fault detection group, and elects a master fault detection component within the group to perform fault detection on the server cluster corresponding to the fault detection group. This avoids deploying a separate management node to perform fault detection on the server cluster, reduces the occupation of management resources, and improves resource utilization.
[0020] In conjunction with the second aspect, in one possible implementation of the second aspect, the hardware-related information includes data center location information and topology information, switch topology information, rack location information, and server location information.
[0021] In conjunction with the second aspect, in one possible implementation of the second aspect, the fault detection component is deployed in an offloading card, which is inserted into the physical server.
[0022] In conjunction with the second aspect, in one possible implementation of the second aspect, the fault detection controller divides servers within the same physical location into the same fault detection group; or, the fault detection controller divides servers with the same physical attributes into the same fault detection group; or, the fault detection controller divides servers with the same fault detection requirements into the same fault detection group; or, the fault detection controller divides a preset number of servers into the same fault detection group.
[0023] In conjunction with the second aspect, in one possible implementation of the second aspect, the fault detection components in the fault detection group establish connections by sending heartbeats to each other and elect the master fault detection component through a preset algorithm; or, the fault detection components in the fault detection group establish connections by sending heartbeats to each other and elect a master fault detection component cluster through a preset algorithm, wherein the master fault detection component cluster includes at least one master fault detection component.
[0024] In conjunction with the second aspect, in one possible implementation of the second aspect, when a new fault detection component joins the fault detection group, the main fault detection component receives the heartbeat sent by the new fault detection component and broadcasts it to the fault detection components in the fault detection group, so that the fault detection components in the fault detection group can establish a connection with the new fault detection component.
[0025] Thirdly, this application provides a network device, including: an acquisition unit for acquiring hardware-related information;
[0026] A processing unit is configured to create a fault detection group based on the hardware-related information. The fault detection group includes at least two fault detection components, which are deployed in a physical server.
[0027] In conjunction with the third aspect, in one possible implementation of the third aspect, the hardware-related information includes data center location information and topology information, switch topology information, rack location information, and server location information.
[0028] In conjunction with the third aspect, in one possible implementation of the third aspect, the fault detection component is deployed in an offloading card, which is inserted into the network server.
[0029] In conjunction with the third aspect, in one possible implementation of the third aspect, the processing unit is specifically used to: divide servers within the same physical location into the same fault detection group; or, divide servers with the same physical attributes into the same fault detection group; or, divide servers with the same fault detection requirements into the same fault detection group; or, divide a preset number of servers into the same fault detection group.
[0030] Fourthly, this application provides a fault detection device, comprising: a receiving unit for receiving group information, the group information including a fault detection group to which the fault detection device belongs, the fault detection group including at least two fault detection devices, the fault detection devices being deployed in a physical server; and an election unit for electing a master fault detection device from the fault detection group, the master fault detection device being used to perform fault detection on the server cluster corresponding to the fault detection group.
[0031] In conjunction with the fourth aspect, in one possible implementation of the fourth aspect, the fault detection device is deployed in an offloading card, which is inserted into the physical server.
[0032] In conjunction with the fourth aspect, in one possible implementation of the fourth aspect, the election unit is specifically used to: send heartbeats to other fault detection devices in the fault detection group to establish connections, and elect the master fault detection device through a preset algorithm; or, send heartbeats to other fault detection devices in the fault detection group to establish connections, and elect a master fault detection device cluster through a preset algorithm, wherein the master fault detection device cluster includes at least one master fault detection device.
[0033] In conjunction with the fourth aspect, in one possible implementation of the fourth aspect, the receiving unit is further configured to receive a heartbeat sent by the new fault detection device when the new fault detection device joins the fault detection group, and broadcast it to the fault detection devices in the fault detection group so that the fault detection devices in the fault detection group can establish a connection with the new fault detection device.
[0034] Fifthly, this application provides a computing device, the computing device including a processor and a memory, the memory being used to store program code, and the processor being used to call the program code in the memory to execute the method described in the second aspect above and any implementation thereof in conjunction with the second aspect above.
[0035] Sixthly, this application provides a computer-readable storage medium storing a computer program that, when executed by a processor, can implement the process of the method provided in the second aspect above and in combination with any implementation of the second aspect above.
[0036] In a seventh aspect, this application provides a computer program product including instructions that, when executed by a computer, enable the computer to perform the process of the method provided in the second aspect and in combination with any implementation of the second aspect. Attached Figure Description
[0037] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the following description of the embodiments will be briefly introduced. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0038] Figure 1 This is a schematic diagram of a hardware fault detection scenario provided in an embodiment of this application;
[0039] Figure 2 This is a schematic diagram of a hardware fault detection system architecture provided in an embodiment of this application;
[0040] Figure 3 This is a flowchart illustrating a hardware fault detection method provided in an embodiment of this application;
[0041] Figure 4 This is a schematic diagram of a hardware distribution provided in an embodiment of this application;
[0042] Figure 5 This is a schematic diagram illustrating the creation of a fault detection group according to an embodiment of this application;
[0043] Figure 6 This is a schematic diagram of another hardware fault detection architecture provided in the embodiments of this application;
[0044] Figure 7 This is a schematic diagram of the structure of a network device provided in an embodiment of this application;
[0045] Figure 8 This is a schematic diagram of the structure of a fault detection device provided in an embodiment of this application;
[0046] Figure 9 This is a schematic diagram of the structure of a computing device provided in an embodiment of this application. Detailed Implementation
[0047] The technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments in this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of this application.
[0048] In this document, the term "embodiment" means that a particular feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of this application. The appearance of this phrase in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment mutually exclusive with other embodiments. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.
[0049] First, some of the terms and related technologies used in this application will be explained in conjunction with the accompanying drawings to facilitate understanding by those skilled in the art.
[0050] In cloud computing hardware fault detection scenarios, the fault detection system is deployed on a centralized management node. It uses a heartbeat mechanism to detect hardware faults, which are then assessed by the management system, which executes fault recovery strategies. For example... Figure 1 As shown, data center 100 deploys switches 110, 120, 130, a core switch 140, and a management node 150. The management node 150 houses a management system 1510, used for fault detection of servers and switches in data center 100. Switch 110 connects to servers 1110, 1120, and 1130; switch 120 connects to servers 1210, 1220, and 1230; and switch 130 connects to servers 1310, 1320, and 1330. Node 150 can actively collect heartbeat information from all servers and switches (e.g., server 1130 and switch 120), or all servers and switches can actively report their own heartbeat information to management node 150. Management node 150 determines whether the algorithm has a hardware failure by detecting whether the heartbeat of each switch or server is interrupted. For example, if management node 150 detects that the heartbeat of server 1310 is interrupted (i.e., no heartbeat information of server 1310 is collected within a preset time), management node 150 determines that server 1310 has a hardware failure and then executes the corresponding fault recovery strategy.
[0051] As can be seen, the above hardware fault detection method is centralized, that is, it uses a separate management node to detect faults in all servers and switches within the data center. However, as the number of servers increases, for example, when the managed scale increases to 100,000 or more, the management node faces enormous detection pressure and may not be able to complete fault detection in real time. It may be necessary to add a caching layer or hierarchical heartbeat access to meet the detection requirements, which will consume a lot of management resources and increase configuration management complexity. Furthermore, for scenarios with a massive number of distributed sites, a separate management node needs to be deployed for each site, resulting in resource waste and increased operation and maintenance costs and burden.
[0052] Based on the above, this application provides a hardware fault detection system, method, and related equipment. The method is executed by the hardware fault detection system, which may include one or more fault detection controllers. These controllers obtain hardware-related information through a hardware management system and create fault detection groups based on this information. Furthermore, the hardware fault detection system also includes multiple fault detection components. These components form fault detection groups based on grouping information, elect a master fault detection component, and then perform fault detection via heartbeat. By implementing this hardware fault detection method, the occupation of management resources can be reduced, resource utilization can be improved, and operation and maintenance costs can be lowered.
[0053] The technical solutions of this application embodiment can also be applied to various scenarios that require hardware fault detection, including but not limited to hardware fault detection in large-scale data centers and hardware fault detection in a large number of small sites.
[0054] Figure 2 A schematic diagram of a hardware fault detection system provided in an embodiment of this application is shown. Figure 2 As shown, the hardware fault detection system includes a fault detection controller 210, a server 220, a server 230, and a hardware management system 240. A fault detection component 2210 is deployed in server 220. This fault detection component 2210 can also be deployed in a separate offloading card inserted into server 220. Server 230, like server 220, deploys a fault detection component 2310. The hardware management system 240 is responsible for collecting and maintaining hardware-related information, such as data center location and topology information, switch topology information, rack location information, and server location information. The fault detection controller 210 obtains hardware-related information through the hardware management system 240, then creates fault detection groups based on this information and distributes the grouping information to each fault detection component. Fault detection components 2210 and 2310 form a fault detection group based on the grouping information and perform fault detection via heartbeat to detect whether servers 220 and 230 have experienced hardware faults.
[0055] Based on the above, the hardware fault detection method and related equipment provided in the embodiments of this application will be described below. See also Figure 3 , Figure 3 This is a flowchart illustrating a hardware fault detection method provided in an embodiment of this application. Figure 3 As shown, the method includes, but is not limited to, the following steps:
[0056] S301: The fault detection controller acquires hardware-related information.
[0057] Specifically, the fault detection controller can be the above-mentioned Figure 2 The fault detection controller 210 in the system can obtain hardware-related information through the hardware management system, which collects and maintains the hardware-related information in advance.
[0058] Optionally, the hardware-related information obtained by the fault detection controller includes physical location information of the computer room, physical topology information between computer rooms, network switch topology information within the computer room, rack location information, and server location information.
[0059] For example, such as Figure 4 As shown, Figure 4 This is a hardware distribution diagram provided in an embodiment of this application. The hardware management system maintains hardware-related information of a data center 400. The data center 400 includes two server rooms 410 and 420, which are deployed in the same physical area and interconnected. Server room 410 includes racks 4110 and 4120, which are interconnected. Rack 4110 includes a switch 4111, a server 4112, and a server 4113. Servers 4112 and 4113 are connected to switch 4111. Rack 4120 includes a switch... Servers 4121, 4122, and 4123 are connected to switch 4121. The data center 420 includes racks 4210 and 4220, which are interconnected. Rack 4210 includes switch 4211, server 4212, and server 4213, which are connected to switch 4211. Rack 4220 includes switch 4221, server 4222, and server 4223, which are connected to switch 4221.
[0060] S302: The fault detection controller creates a fault detection group based on hardware-related information.
[0061] Specifically, after obtaining the relevant hardware information, the fault detection controller divides all servers in the data center into clusters to obtain multiple fault detection groups. Each fault detection group includes at least two fault detection components. Then, the fault detection group to which each server belongs and its physical location information are written into the configuration file of the installation and deployment system.
[0062] In one possible implementation, the fault detection controller groups servers located in the same physical location into the same fault detection group.
[0063] For example, the fault detection controller divides servers connected to the same computer room (e.g., computer room 410 above), or the same rack (e.g., rack 4110 above), or the same switch (e.g., switch 4111 above) into the same fault detection group, and each fault detection group corresponds to a server cluster.
[0064] In another possible implementation, the fault detection controller groups servers with the same physical attributes into the same fault detection group.
[0065] For example, the fault detection controller groups servers with the same hardware configuration, such as servers with the same central processing unit (CPU), into the same fault detection group.
[0066] In another possible implementation, the fault detection controller groups servers with the same fault detection requirements into the same fault detection group.
[0067] For example, the fault detection controller can group servers with the same fault detection requirements into the same fault detection group based on the fault detection scenario, such as the need to cover multiple switches or racks.
[0068] In another possible implementation, the fault detection controller divides a preset number of servers into the same fault detection group.
[0069] For example, the fault detection controller divides a preset number of servers, such as 100, into the same fault detection group based on the number of servers. Of course, the preset number can also be set to 200 or 300 as needed, and this application does not limit it.
[0070] S303: The fault detection component in the fault detection group elects the master fault detection component.
[0071] Specifically, after the fault detection controller creates fault detection groups, each fault detection group corresponds to a server cluster. Fault detection of the servers is implemented within each fault detection group, eliminating the need to deploy additional management nodes to perform fault detection on each server cluster. This reduces the occupation of management resources and improves resource utilization.
[0072] It should be noted that the fault detection component can be deployed in a separate hardware offload card, using a fully distributed architecture, such as... Figure 5 As shown, servers 510, 520, and 530 are each equipped with a hardware offload card. Hardware offload cards 5110, 5210, and 5310 respectively deploy fault detection components 5111, 5211, and 5311. Fault detection controller 540 is connected to hardware offload cards 5110, 5210, and 5310 and creates fault detection components 5111, 5211, and 5311 into the same fault detection group.
[0073] Furthermore, after power-on, the hardware unloading card loads the corresponding configuration file from the installation and deployment system. When the fault detection component starts, it reads the configuration information to determine its own fault detection group. All fault detection components within the same fault detection group send heartbeats to each other to form an arbitration cluster. Through preset algorithms, such as Consistency, Availability, Partition Tolerance (CAP), Paxos, and Raft, a master fault detection component is elected. The master fault detection component uses heartbeats to detect whether the server cluster corresponding to the fault detection group has experienced a hardware failure. When the master fault detection component detects that the heartbeat of a fault detection component in the group has been interrupted, for example, if the master fault detection component does not receive heartbeat information from a fault detection component within a preset time, the master fault detection component can determine that the server corresponding to that fault detection component has experienced a hardware failure and needs to execute the corresponding fault recovery strategy.
[0074] It is understandable that by using a hardware offload card to run the fault detection component, the hardware offload card can be directly inserted into the server without the need for complex management and configuration. This can reduce the management resource overhead of the server, improve the utilization of server resources, and reduce operation and maintenance costs.
[0075] In one possible implementation, all fault detection components within the same fault detection group send heartbeats to each other to form an arbitration cluster. A master fault detection component cluster is then elected using a preset algorithm. This master fault detection component cluster includes multiple master fault detection components. If the current master fault detection component fails, other master fault detection components in the cluster will take over its work and continue to perform hardware fault detection on the server cluster corresponding to the fault detection group, ensuring the normal operation of the business. It is easy to understand that electing a master fault detection component cluster with multiple master fault detection components and providing redundant backups of the master fault detection components can effectively improve system reliability and ensure the normal operation of the business.
[0076] S304: The main fault detection component performs fault detection on the server cluster corresponding to the fault detection group.
[0077] Specifically, for each fault detection group, a master fault detection component will be elected. After the selection is completed, the master fault detection component can use heartbeat to detect whether the server cluster corresponding to the fault detection group has a hardware failure, and execute the corresponding fault recovery strategy after detecting a hardware failure.
[0078] In one possible implementation, when the data center needs to be expanded, i.e., when new fault detection components and servers need to be added, the fault detection controller can create a new fault detection group to include the newly added fault detection components, thereby enabling hardware fault detection of the new servers. Alternatively, it can add new fault detection components to an existing fault detection group, thereby using the main fault detection component of the fault detection group to enable hardware fault detection of the server corresponding to the new fault detection component.
[0079] Furthermore, when a new fault detection component joins the fault detection group, the new fault detection component broadcasts to the group and registers with the main fault detection component of the group and sends a heartbeat. The main fault detection component then performs hardware fault detection on the server corresponding to the new fault detection component based on the heartbeat sent by the new fault detection component.
[0080] The above details the hardware fault detection method provided in this application. The following will combine... Figure 6 The hardware fault detection method provided in this application will be described in further detail.
[0081] like Figure 6As shown, the fault detection controller 610 obtains hardware-related information from the hardware management system 620 and divides the servers in the cloud environment and edge environment according to the hardware-related information to obtain different fault detection groups. The cloud environment refers to the central computing equipment cluster owned by the cloud service provider, which is used to provide computing, storage and communication resources. The edge environment refers to the edge computing equipment cluster that is geographically close to the data acquisition equipment and is used to provide computing, storage and communication resources. The cloud environment includes core switch 630, switch 6310, switch 6320, and switch 6330. Switch 6310 connects to servers 6311, 6312, and 6313. Switch 6320 connects to servers 6321, 6322, and 6323. Switch 6330 connects to servers 6331, 6332, and 6333. The edge environment includes switch 640, which connects to servers 6410, 642, and 6430. Each server has a hardware offload card with a fault detection component deployed on it. For example, the hardware offload card of server 6311 has a fault detection component 63110 deployed on it. Specifically, the fault detection controller 610 divides switch 6310 and its connected servers into the same fault detection group, switch 6320 and its connected servers into the same fault detection group, switch 6330 and its connected servers into the same fault detection group, and switch 640 and its connected servers into the same fault detection group. After the fault detection group division is completed, the fault detection components in each fault detection group send heartbeats to each other and elect a master fault detection component or a cluster of master fault detection components through a preset algorithm. For example, fault detection components 63110, 63120, and 63130 elect fault detection component 63110 as the master fault detection component of the fault detection group. After electing the master fault detection component of each fault detection group, the master fault detection component performs hardware fault detection on the server cluster in its group through heartbeats. For example, the master fault detection component 63110 performs hardware fault detection on servers 6311, 6312, and 6313 through heartbeats, and executes the corresponding fault recovery strategy after detecting a hardware fault.
[0082] As can be seen, the fault detection controller divides servers in the cloud and edge environments into clusters to obtain different fault detection groups. Fault detection is performed within the fault detection group, eliminating the need for centralized fault detection of all servers through a central management node or setting up an additional management node for each fault detection group. Furthermore, the use of hardware offload cards to deploy fault detection components further reduces management and configuration complexity, reduces management resource consumption, and improves resource utilization.
[0083] It should be noted that, Figure 6 The hardware fault detection method shown is similar to Figure 3 The principle is the same, and you can refer to it. Figure 3 For the sake of brevity, the relevant descriptions in steps S301 to S304 will not be repeated here.
[0084] The methods of the embodiments of this application have been described in detail above. In order to facilitate better implementation of the above solutions of the embodiments of this application, relevant equipment for cooperating in implementing the above solutions is also provided below.
[0085] See Figure 7 , Figure 7 This is a schematic diagram of the structure of a network device provided in an embodiment of this application. The network device can be as described above. Figure 3 The fault detection controller in the method embodiment can execute... Figure 3 The hardware fault detection method embodiments described herein employ a fault detection controller as the executing entity for the methods and steps. For example... Figure 7 As shown, the network device 700 includes an acquisition unit 710 and a processing unit 720. Wherein,
[0086] Acquisition unit 710 is used to acquire hardware-related information;
[0087] The processing unit 720 is configured to create a fault detection group based on the hardware-related information. The fault detection group includes at least two fault detection components, which are deployed in a physical server.
[0088] Specifically, the acquisition unit 710 is used to execute the aforementioned step S301, and optionally executes a method selected in the aforementioned step; the processing unit 720 is used to execute the aforementioned step S302, and optionally executes a method selected in the aforementioned step. The two units can transmit data to each other through a communication path. It should be understood that the units included in the network device 700 can be software units, hardware units, or a combination of both.
[0089] As one example, the hardware-related information includes data center location information and topology information, switch topology information, rack location information, and server location information.
[0090] As one embodiment, the fault detection component is deployed in an offloading card, which is inserted into the network server.
[0091] As one embodiment, the processing unit 720 is specifically used to: divide servers in the same physical location into the same fault detection group; or, divide servers with the same physical attributes into the same fault detection group; or, divide servers with the same fault detection requirements into the same fault detection group; or, divide a preset number of servers into the same fault detection group.
[0092] It is understood that the acquisition unit 710 in the embodiments of this application can be implemented by a transceiver or transceiver-related circuit components, and the processing unit 720 can be implemented by a processor or processor-related circuit components.
[0093] It should be noted that the above-described network device structure is merely an example and should not constitute a specific limitation. The various units within this network device can be added, removed, or combined as needed. Furthermore, the operation and / or function of each unit within this network device are to achieve the above-described... Figure 3 For the sake of brevity, the corresponding process of the described method will not be elaborated here.
[0094] See Figure 8 , Figure 8 This is a schematic diagram of the structure of a fault detection device provided in an embodiment of this application. The fault detection device can be as described above. Figure 3 The fault detection component in the method embodiment can perform... Figure 3 The hardware fault detection method embodiments described herein employ methods and steps with a fault detection control component as the executing entity. For example... Figure 8 As shown, the fault detection device 800 includes a receiving unit 810 and an election unit 820. Among them,
[0095] The receiving unit 810 is used to receive packet information, the packet information including the fault detection group to which the fault detection device belongs, the fault detection group including at least two fault detection devices, and the fault detection devices being deployed in a physical server;
[0096] The election unit 820 is used to elect a master fault detection device in the fault detection group, and the master fault detection device is used to perform fault detection on the server cluster corresponding to the fault detection group.
[0097] Specifically, the receiving unit 810 is used to execute the aforementioned step S303, and optionally executes a method selected in the aforementioned steps. The election unit 820 is used to execute the aforementioned steps S303 and S304, and optionally executes a method selected in the aforementioned steps. The two units can transmit data to each other through a communication path. It should be understood that the units included in the fault detection device 800 can be software units, hardware units, or a combination of both.
[0098] As one embodiment, the fault detection device 800 is deployed in an offloading card, which is inserted into the physical server.
[0099] As one embodiment, the election unit 820 is specifically used to: send a heartbeat to other fault detection devices in the fault detection group to establish a connection, and elect the master fault detection device through a preset algorithm; or, send a heartbeat to other fault detection devices in the fault detection group to establish a connection, and elect a master fault detection device cluster through a preset algorithm, wherein the master fault detection device cluster includes at least one master fault detection device.
[0100] As an example, the receiving unit 810 is further configured to receive a heartbeat sent by a new fault detection device when the new fault detection device joins the fault detection group, and broadcast it to the fault detection devices in the fault detection group so that the fault detection devices in the fault detection group can establish a connection with the new fault detection device.
[0101] It is understood that the receiving unit 810 in the embodiments of this application can be implemented by a transceiver or transceiver-related circuit components, and the election unit 820 can be implemented by a processor or processor-related circuit components.
[0102] It should be noted that the structure of the fault detection device described above is merely an example and should not be construed as a specific limitation. The various units within the fault detection device can be added, removed, or combined as needed. Furthermore, the operation and / or function of each unit within the fault detection device are to achieve the aforementioned... Figure 3 For the sake of brevity, the corresponding process of the described method will not be elaborated here.
[0103] See Figure 9 , Figure 9 This is a schematic diagram of the structure of a computing device provided in an embodiment of this application. Figure 9 As shown, the computing device 900 includes a processor 910, a communication interface 920, and a memory 930, which are interconnected via an internal bus 940. It should be understood that the computing device 900 can be a computing device in cloud computing or a computing device in an edge environment.
[0104] The processor 910 may consist of one or more general-purpose processors, such as a central processing unit (CPU), or a combination of a CPU and hardware chips. The hardware chips may be application-specific integrated circuits (ASICs), programmable logic devices (PLDs), or combinations thereof. The PLDs may be complex programmable logic devices (CPLDs), field-programmable gate arrays (FPGAs), generic array logic (GALs), or any combination thereof.
[0105] Bus 940 can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. Bus 940 can be divided into address bus, data bus, control bus, etc. For ease of representation, Figure 9 The symbol is represented by a single thick line, but this does not mean that there is only one bus or one type of bus.
[0106] Memory 930 may include volatile memory, such as random access memory (RAM); memory 730 may also include non-volatile memory, such as read-only memory (ROM), flash memory, hard disk drive (HDD), or solid-state drive (SSD); memory 730 may also include combinations of the above types.
[0107] It should be noted that the memory 930 of the computing device 900 stores the code corresponding to each unit of the network device 700 or the fault detection device 800. The processor 910 executes this code to implement the functions of each unit of the network device 700 or the fault detection device 800, that is, to execute the methods S301-S304.
[0108] This application also provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program that, when executed by a processor, can implement some or all of the steps described in any of the above method embodiments.
[0109] This invention also provides a computer program that includes instructions that, when executed by a computer, enable the computer to perform some or all of the steps of any method for distributing regional resources.
[0110] In the above embodiments, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments.
[0111] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that this application is not limited to the described order of actions, as some steps may be performed in other orders or simultaneously according to this application. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions and modules involved are not necessarily essential to this application.
[0112] In the several embodiments provided in this application, it should be understood that the disclosed apparatus can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of the units described above is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between devices or units may be electrical or other forms.
[0113] The units described above as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0114] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
Claims
1. A hardware fault detection system, characterized in that, It includes a fault detection controller and multiple fault detection components, which are deployed on multiple physical servers, wherein: The fault detection controller is used to acquire hardware-related information and create a fault detection group based on the hardware-related information. The fault detection group includes at least two fault detection components. The fault detection component in the fault detection group elects a master fault detection component, which is used to perform fault detection on the server cluster corresponding to the fault detection group. When a new fault detection component joins the fault detection group, the main fault detection component receives the heartbeat sent by the new fault detection component and broadcasts it to the fault detection components in the fault detection group so that the fault detection components in the fault detection group can establish a connection with the new fault detection component.
2. The system as described in claim 1, characterized in that, The hardware-related information includes data center location and topology information, switch topology information, rack location information, and server location information.
3. The system as described in claim 1, characterized in that, The fault detection component is deployed across multiple physical servers, including: The fault detection component is deployed in the offloading card, which is inserted into the physical server.
4. The system according to any one of claims 1 to 3, characterized in that, The fault detection controller creates the fault detection group based on the hardware-related information, including: The fault detection controller groups servers located in the same physical location into the same fault detection group; or... The fault detection controller groups servers with the same physical attributes into the same fault detection group; or, The fault detection controller groups servers with the same fault detection requirements into the same fault detection group; or, The fault detection controller divides a preset number of servers into the same fault detection group.
5. The system according to any one of claims 1 to 3, characterized in that, The fault detection components in the fault detection group elect a master fault detection component, including: The fault detection components in the fault detection group establish connections by sending heartbeats to each other, and elect the master fault detection component through a preset algorithm; or... The fault detection components in the fault detection group establish connections by sending heartbeats to each other, and elect a master fault detection component cluster through a preset algorithm. The master fault detection component cluster includes at least one master fault detection component.
6. A hardware fault detection method, characterized in that, include: The fault detection controller acquires hardware-related information and creates a fault detection group based on the hardware-related information. The fault detection group includes at least two fault detection components, which are deployed in a physical server. The fault detection component in the fault detection group elects a master fault detection component, which is used to perform fault detection on the server cluster corresponding to the fault detection group. When a new fault detection component joins the fault detection group, the main fault detection component receives the heartbeat sent by the new fault detection component and broadcasts it to the fault detection components in the fault detection group so that the fault detection components in the fault detection group can establish a connection with the new fault detection component.
7. The method as described in claim 6, characterized in that, The hardware-related information includes data center location and topology information, switch topology information, rack location information, and server location information.
8. The method as described in claim 6, characterized in that, The fault detection component is deployed on a physical server and includes: The fault detection component is deployed in the offloading card, which is inserted into the physical server.
9. The method according to any one of claims 6 to 8, characterized in that, The fault detection controller creates the fault detection group based on the hardware-related information, including: The fault detection controller groups servers located in the same physical location into the same fault detection group; or... The fault detection controller groups servers with the same physical attributes into the same fault detection group; or, The fault detection controller groups servers with the same fault detection requirements into the same fault detection group; or, The fault detection controller divides a preset number of servers into the same fault detection group.
10. The method according to any one of claims 6 to 8, characterized in that, The fault detection components in the fault detection group elect a master fault detection component, including: The fault detection components in the fault detection group establish connections by sending heartbeats to each other, and elect the master fault detection component through a preset algorithm; or... The fault detection components in the fault detection group establish connections by sending heartbeats to each other, and elect a master fault detection component cluster through a preset algorithm. The master fault detection component cluster includes at least one master fault detection component.
11. A network device, characterized in that, include: The acquisition unit is used to acquire hardware-related information; A processing unit is configured to create a fault detection group based on the hardware-related information. The fault detection group includes at least two fault detection components, which are deployed on a physical server. The processing unit is further configured to receive a heartbeat sent by the new fault detection component when a new fault detection component joins the fault detection group, and broadcast it to the fault detection components in the fault detection group so that the fault detection components in the fault detection group can establish a connection with the new fault detection component. The main fault detection component is used to perform fault detection on the server cluster corresponding to the fault detection group.
12. The network device as described in claim 11, characterized in that, The hardware-related information includes data center location and topology information, switch topology information, rack location information, and server location information.
13. The network device as described in claim 11, characterized in that, The fault detection component is deployed in the uninstallation card, which is inserted into the network server.
14. The network device according to any one of claims 11 to 13, characterized in that, The processing unit is specifically used for: Servers located in the same physical location are grouped into the same fault detection group; or... Servers with the same physical properties are grouped into the same fault detection group; or, Servers with the same fault detection requirements are grouped into the same fault detection group; or, Divide a preset number of servers into the same fault detection group.
15. A fault detection device, characterized in that, include: A receiving unit is configured to receive packet information, the packet information including the fault detection group to which the fault detection device belongs, the fault detection group including at least two fault detection devices, and the fault detection devices being deployed in a physical server; An election unit is used to elect a master fault detection device in the fault detection group, and the master fault detection device is used to perform fault detection on the server cluster corresponding to the fault detection group. The receiving unit is further configured to receive a heartbeat sent by a new fault detection device when the new fault detection device joins the fault detection group, and broadcast it to the fault detection devices in the fault detection group so that the fault detection devices in the fault detection group can establish a connection with the new fault detection device.
16. The fault detection device as described in claim 15, characterized in that, The fault detection device is deployed in the unloading card, which is inserted into the physical server.
17. The fault detection device as described in claim 15 or 16, characterized in that, The election unit is specifically used for: Send heartbeats to other fault detection devices in the fault detection group to establish connections, and elect the master fault detection device through a preset algorithm; or... A heartbeat is sent to other fault detection devices in the fault detection group to establish a connection, and a master fault detection device cluster is elected through a preset algorithm. The master fault detection device cluster includes at least one master fault detection device.
18. A computing device, characterized in that, The computing device includes a memory and a processor, the processor executing computer instructions stored in the memory, causing the computing device to perform the method according to any one of claims 6-10.
19. A computer-readable storage medium storing a computer program that, when executed by a processor, performs the method according to any one of claims 6-10.
20. A computer program product comprising instructions that, when executed by a computer, cause the computer to perform the method of any one of claims 6-10.
Citation Information
Patent Citations
Fault detection method and device and related equipment
CN110740072A
Container group POD reconstruction method based on container cluster service and related equipment
CN113872997A
Multiprocessor system
JP1993134998A