A self-checking method and server
By physically partitioning and performing power-on self-tests on the server's resources to be tested, the problem of not being aware of hardware resource failures in the server is solved, ensuring complete testing of the server hardware and avoiding potential hidden dangers and failures.
Patent Information
- Application Number
- CN202210283839.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2017-12-19
- Publication Date
- 2025-10-24
- Estimated Expiration
- 2037-12-19
AI Technical Summary
In existing technologies, some hardware resources of servers cannot perform self-tests when they are not physically partitioned or powered off, leading to the problem of not being able to detect faults.
After receiving the self-test command, the server performs physical partitioning of the resources to be tested, obtains the self-test partition, and performs a self-test on it by powering it on, using the BIOS, OS and self-test tools for comprehensive testing.
This ensures that all hardware resources on the server can be detected during the self-test process, avoiding the problem of not being aware of hardware failures and preventing potential server malfunctions or operational problems.
Smart Images

Figure CN114911655B_ABST
Abstract
Description
[0001] This application is a divisional application of the original application with the application number 201711381216.7 and the original filing date of December 19, 2017, and the entire contents of the original application are incorporated herein by reference. TECHNICAL FIELD
[0002] The present application relates to the field of computers, and in particular to a self-checking method and a server. BACKGROUND
[0003] Physical partitioning, also known as hard partitioning, refers to dividing a single server system into multiple partitions in a physical manner by means of hardware modular design and system flexible configuration capability, each partition having its own dedicated hardware resources, which are electrically isolated and do not affect each other. When the server is scheduled to be powered off or checked, the hardware resources on the server need to be self-checked to exclude faults in advance. As shown in Figure 1 At present, if the server is only divided into partition 1 and partition 2, the hardware resources of the server that are not divided into partition 1 and partition 2 are in an idle state, and the server will not perform self-checking on these hardware resources in the idle state. In addition, as shown in Figure 1 In partition 1, even if the hardware resources have been divided into partition 1, if the strategy configured by partition 1 is that the partition is not powered on when the server is started, the server will not perform self-checking on the hardware resources in partition 1. In the above cases, the server cannot perform self-checking on all hardware resources. SUMMARY
[0004] Embodiments of the present application provide a self-checking method and a server, which are used to solve the problem of fault unawareness of part of the hardware resources in the server and avoid hidden dangers or faulty work of the server.
[0005] The first aspect of the embodiments of the present application provides a self-checking method, comprising:
[0006] The server performs physical partitioning on the to-be-self-checked resources of the server in response to the received self-checking instruction to obtain a self-checking partition. After obtaining the self-checking partition, the server powers on the self-checking partition and performs self-checking on the self-checking partition. In the embodiments of the present application, all the hardware resources mounted on the server can be detected in the self-checking process, the problem of fault unawareness of part of the hardware resources is solved, it is ensured that there is no problem with the hardware in the server, and hidden dangers or faulty work of the server are avoided.
[0007] In a possible design, in the first implementation manner of the first aspect of the embodiment of the present application, the to-be-self-checked resource can include a first hardware resource of the server. The first hardware resource is a hardware resource that has been divided into a historical physical partition but not powered on before the self-check instruction is received. In this implementation manner, the to-be-self-checked resource can also be a resource in the server that has been in a historical physical partition but not powered on, which increases the applicable scenarios of the present application.
[0008] In a possible design, in the second implementation manner of the first aspect of the embodiment of the present application, the server performs physical partitioning on the to-be-self-checked resource of the server to obtain the self-check partition in response to the received self-check instruction, specifically including: the server deletes the historical physical partition of the first hardware resource, and after the historical physical partition of the first hardware resource is deleted, the server re-performs physical partitioning on the first hardware resource from which the historical physical partition is deleted to obtain the self-check partition. In this implementation manner, the specific physical partitioning manner when the to-be-self-checked resource includes the first hardware resource is refined, which makes the embodiment of the present application more operable.
[0009] In a possible design, in the third implementation manner of the first aspect of the embodiment of the present application, before the server performs physical partitioning on the to-be-self-checked resource of the server to obtain the self-check partition in response to the received self-check instruction, the server further includes: the server receives partition information, the partition information containing information of a hardware device. The server queries the position of the hardware device in the affiliation relationship tree, and the server can further determine the to-be-self-checked resource according to the position. In this implementation manner, a possible manner of determining the to-be-self-checked resource is described, which increases the implementable manners of the embodiment of the present application.
[0010] In a possible design, in the fourth implementation manner of the first aspect of the embodiment of the present application, the server determines the to-be-self-checked resource according to the position, specifically including: the server determines the hardware device and the hardware devices belonging to the leaf nodes of the hardware device in the affiliation relationship tree as the to-be-self-checked resource. In this implementation manner, a manner of determining the to-be-self-checked resource according to the position is provided, which makes the embodiment of the present application more operable.
[0011] In a possible design, in the fifth implementation manner of the first aspect of the embodiment of the present application, the server determines the to-be-self-checked resource according to the position, specifically including: the server determines the hardware device and the hardware devices belonging to the same level nodes of the hardware device in the affiliation relationship tree as the to-be-self-checked resource. In this implementation manner, another manner of determining the to-be-self-checked resource according to the position is provided, which increases the implementable manners of the embodiment of the present application.
[0012] In a possible design, in the sixth implementation manner of the first aspect of the embodiment of the present application, when the server receives the normal information of the to-be-inspected resource in the self-check partition, that is, the to-be-inspected resource is normal, the server powers off the self-check partition and deletes the self-check partition. In this implementation manner, the subsequent processing of the self-check partition by the server when the self-check shows that the resource in the to-be-inspected resource that is not partitioned is normal is illustrated.
[0013] In a possible design, in the seventh implementation manner of the first aspect of the embodiment of the present application, before the server deletes the historical physical partition of the first hardware resource, the server backs up the partition record of the historical physical partition. In this implementation manner, when the to-be-inspected resource includes the first hardware resource, the information of the historical physical partition of the first hardware resource also needs to be backed up for use after the self-check, which perfects the operation steps of the embodiment of the present application.
[0014] In a possible design, in the eighth implementation manner of the first aspect of the embodiment of the present application, after the server receives the normal information of the to-be-inspected resource in the self-check partition, the server powers off the self-check partition. Then, the server deletes the self-check partition and restores the historical physical partition of the first hardware resource according to the backed-up partition record of the historical physical partition. In this implementation manner, the subsequent processing of the self-check partition by the server when the self-check shows that the first hardware resource in the to-be-inspected resource is normal is illustrated, which makes the embodiment of the present application more logical.
[0015] In a possible design, in the ninth implementation manner of the first aspect of the embodiment of the present application, the server receives the fault information of the to-be-inspected resource in the self-check partition. In this implementation manner, the case that the to-be-inspected resource in the self-check partition has a fault is illustrated, which increases the application scenarios of the embodiment of the present application.
[0016] The second aspect of the embodiment of the present application provides a server, which includes a partition module and a self-check module. The partition module is configured to perform physical partitioning on to-be-inspected resources of the server to obtain a self-check partition in response to a received self-check instruction. The self-check module is configured to power on the self-check partition and perform self-checking on the self-check partition. In the embodiment of the present application, all hardware resources mounted on the server can be detected in the self-checking process, which solves the problem of fault unawareness of part of the hardware resources and ensures that all hardware in the server is normal, thereby avoiding hidden dangers or fault operation of the server.
[0017] In a possible design, in the first implementation manner of the second aspect of the embodiment of the present application, the to-be-self-checked resource further includes a first hardware resource, and the first hardware resource is a hardware resource that has been divided into a historical physical partition but has not been powered on before the self-check instruction is received. In this implementation manner, the to-be-self-checked resource can also be a resource in a historical physical partition but not powered on in the server, and the application scenario of the present application is increased.
[0018] In a possible design, in the second implementation manner of the second aspect of the embodiment of the present application, the partition module is specifically configured to delete the historical physical partition of the first hardware resource, and perform physical partitioning on the first hardware resource from which the historical physical partition is deleted to obtain the self-check partition. In this implementation manner, the specific physical partitioning manner when the to-be-self-checked resource includes the first hardware resource is refined, and the embodiment of the present application is more operable.
[0019] In a possible design, in the third implementation manner of the second aspect of the embodiment of the present application, the server further includes: a first receiving module configured to receive partition information, the partition information including information of a hardware device; a querying module configured to query a position of the hardware device in a hanging relationship tree; and a determining module configured to determine the to-be-self-checked resource according to the position. In this implementation manner, a possible manner of determining the to-be-self-checked resource is described, and the implementation manner of the embodiment of the present application is increased.
[0020] In a possible design, in the fourth implementation manner of the second aspect of the embodiment of the present application, the determining module is specifically configured to determine the hardware device and a hardware device belonging to a leaf node of the hardware device in the hanging relationship tree as the to-be-self-checked resource. In this implementation manner, a manner of determining the to-be-self-checked resource according to the position is provided, and the embodiment of the present application is more operable.
[0021] In a possible design, in the fifth implementation manner of the second aspect of the embodiment of the present application, the determining module is specifically further configured to determine the hardware device and a hardware device belonging to a same level node of the hardware device in the hanging relationship tree as the to-be-self-checked resource. In this implementation manner, a manner of determining the to-be-self-checked resource according to the position is further provided, and the implementation manner of the embodiment of the present application is increased.
[0022] In a possible design, in the sixth implementation manner of the second aspect of the embodiment of the present application, the server further includes: a second receiving module configured to receive to-be-self-checked resource normal information in the self-check partition; a powering-off module configured to power off the self-check partition; and a deleting module configured to delete the self-check partition. In this implementation manner, when the self-check result is that a resource not partitioned in the to-be-self-checked resource is normal, the subsequent processing of the self-check partition by the server is described, and the embodiment of the present application is more logical.
[0023] In a possible design, in the seventh implementation manner of the second aspect of the embodiment of the present application, the server further includes a backup module, configured to backup the partition record of the historical physical partition before the deletion module deletes the historical physical partition of the first hardware resource. In this implementation manner, when the to-be-inspected resource includes the first hardware resource, the information of the historical physical partition of the first hardware resource is also needed to be backed up for use after the inspection, which perfects the operation steps of the embodiment of the present application.
[0024] In a possible design, in the eighth implementation manner of the second aspect of the embodiment of the present application, the server further includes a recovery module, configured to recover the historical physical partition of the first hardware resource according to the backup partition record after the deletion module deletes the inspection partition. In this implementation manner, the subsequent processing of the server on the inspection partition when the first hardware resource in the to-be-inspected resource is normal after the inspection is illustrated, which makes the embodiment of the present application more logical.
[0025] In a possible design, in the ninth implementation manner of the second aspect of the embodiment of the present application, the server further includes a third receiving module, configured to receive to-be-inspected resource failure information in the inspection partition. In this implementation manner, the case that the to-be-inspected resource in the inspection partition has a failure is illustrated, which increases the application scenario of the embodiment of the present application.
[0026] The third aspect of the present application provides a computer readable storage medium, which stores instructions, and when the instructions are run on a computer, the computer executes the method of the above aspects.
[0027] The fourth aspect of the present application provides a computer program product containing instructions, and when the instructions are run on a computer, the computer executes the method of the above aspects.
[0028] As can be seen from the above technical solutions, the embodiment of the present application has the following advantages: in the embodiment of the present application, after the server receives the inspection instruction, the to-be-inspected resource is physically partitioned to obtain an inspection partition, the inspection partition is powered on, and the inspection partition is inspected, so that all hardware resources or specified hardware resources mounted on the server can be detected in the inspection process, the problem of failure of part of the hardware resources not being perceived is solved, it is ensured that all hardware in the server is normal, and hidden dangers or failure of the server are avoided. BRIEF DESCRIPTION OF DRAWINGS
[0029] Figure 1 A schematic diagram of physical partitioning for a server;
[0030] Figure 2 A flowchart of a possible inspection method provided by the embodiment of the present application;
[0031] Figure 3 a flowchart of another possible self-checking method provided for an embodiment of the present application;
[0032] Figure 4 a schematic diagram of a possible affiliation tree provided for an embodiment of the present application;
[0033] Figure 5 a schematic diagram of a possible structure of a server provided for an embodiment of the present application;
[0034] Figure 6 a schematic diagram of another possible structure of a server provided for an embodiment of the present application;
[0035] Figure 7 a schematic diagram of another possible structure of a server provided for an embodiment of the present application. DETAILED DESCRIPTION
[0036] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only some of the embodiments of the present application, but not all the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative work fall within the scope of protection of the present application.
[0037] The self-checking method in the embodiments of the present application can be applied to a server. To solve the problem of fault unawareness of hardware resources that are not physically partitioned in the server, for the convenience of description, in the embodiments of the present application, the hardware resources that are not physically partitioned are referred to as to-be-self-checked resources. In the present application, the physical partitioning can be a technology of dividing a server into multiple processing units. These processing units can be regarded as small servers, so that a single multi-node server can execute multiple tasks on an independent partitioned operating system.
[0038] In the embodiments of the present application, the to-be-self-checked resources are automatically physically partitioned by management software according to certain principles, and then self-checked by using a basic input output system (BIOS) and / or an operating system (OS), a self-checking tool, a self-checking image, etc., so as to ensure that there is no problem with the hardware in the server, thereby avoiding hidden dangers or faults and eliminating the situation that the hardware cannot be used due to a fault when it needs to be used subsequently.
[0039] The self-checking method in the embodiments of the present application will be described in detail as follows:
[0040] Please refer to Figure 2 An embodiment of the self-checking method of the present application includes:
[0041] 201、the server receives a self-check instruction;
[0042] It can be understood that the self-check instruction received by the server can be triggered by the user actively, or can be triggered by the server according to a preset condition, which is not limited here.
[0043] Optionally, as the condition for triggering the self-check instruction by the server, it can include:
[0044] 1、power-on of the server: that is, after the server is plugged in, since the management software of the server does not save the hardware information of the whole machine, and there can be changes in the hardware during the power-off process of the server, when the server is powered on, it can be necessary to trigger the self-check instruction to scan and self-check the hardware;
[0045] 2、routine check: routine check is divided into two kinds, one is periodic routine check, for example, some routine check plans in the operation and maintenance level at present are to perform self-check operation on the equipment in the time period when the business throughput is low; the other is self-check according to demand, for example, some errors reported by the underlying are detected in the business level, such as cyclic redundancy check (CRC) error on I / O, error check and error correcting code (ECC) error of memory, etc., at this time, relevant personnel will also plan or immediately perform some self-check operation.
[0046] It can be understood that the user can also set other preset self-check triggering conditions according to needs, which are not limited here.
[0047] 202、the server physically partitions the to-be-self-checked resource to obtain a self-check partition;
[0048] It should be noted that the to-be-self-checked resource includes not only the hardware resource that has not been physically partitioned, but also the first hardware resource, wherein the first hardware resource is a hardware resource that has been divided into a historical physical partition before receiving the self-check instruction but has not been powered on. The scope of the to-be-self-checked resource can also be set by the user in advance.
[0049] It can be understood that, since the self-checking is mainly performed by using the BIOS and / or the OS (including the OS itself and some self-checking tools on the OS) in the self-checking process, for example, the BIOS can perform self-checking on the central processing unit (CPU), memory, etc., and the OS can perform self-checking at the level of the operating system, etc., therefore, the BIOS or the OS needs to know the current relevant hardware resource situation, and thus needs to physically partition the to-be-self-checked resource, wherein the physical partition can be a partition that can be directly used after being partitioned, or a virtual partition (i.e., invisible to the user) for self-checking only.
[0050] Optionally, since the first hardware resource has actually been divided into the historical physical partition, as another embodiment of the self-checking method in the present application, when the present step is performed, the server can specifically perform the following steps:
[0051] The server determines whether there is a hardware resource that is not physically partitioned in the to-be-self-checked resource;
[0052] If not, it is determined that the historical physical partition of the first hardware resource is the self-checking partition, wherein the historical physical partition of the first hardware resource is the partition of the first hardware resource before the server receives the self-checking instruction;
[0053] If yes, the hardware resource that is not physically partitioned is physically partitioned, and the obtained physical partition and the historical physical partition of the first hardware resource are taken as the to-be-self-checked partition.
[0054] It can be understood that, in actual application, the first hardware resource can also be directly re-physically partitioned to obtain the self-checking partition, which specifically includes the following steps:
[0055] Step 1, the server backs up the partition record of the historical physical partition of the first hardware resource;
[0056] Step 2, the server deletes the historical physical partition of the first hardware resource;
[0057] Step 3, the server re-physically partitions the first hardware resource from which the historical physical partition is deleted to obtain the self-checking partition.
[0058] In summary, the server can obtain the self-check partition in the following ways: 1. When the to-be-self-checked resource includes a hardware resource that is not physically partitioned, the to-be-self-checked resource is physically partitioned to obtain the self-check partition; 2. When the to-be-self-checked resource further includes a first hardware resource, a historical physical partition of the first hardware resource can be used as the self-check partition; 3. When the to-be-self-checked resource further includes a first hardware resource, a partition record of a historical physical partition of the first hardware resource is backed up, and the historical physical partition is deleted, and the first hardware resource is physically partitioned again to obtain the self-check partition. Therefore, the way in which the server obtains the self-check partition is not limited in the present application.
[0059] In addition, there are many bases for physically partitioning the to-be-self-checked resource, and some of them are described as follows:
[0060] 1. The fastest principle: In actual startup and self-check, each physical partition can be independently and in parallel, and the less the hardware resource, the faster the startup and self-check process. Therefore, if the self-check needs to be completed in the shortest possible time, this principle can be selected. This principle is generally the smallest unit partition supported by the server (for example, the SDX of HP is 2P, which can generally be found in product data).
[0061] 2. The most matching principle: Due to the fact that some servers only support equal partitioning, for example, a 32-way server only supports 8 4P physical partitions, 4 8P physical partitions, 2 16P physical partitions, or 1 32P physical partition. If there are currently 2 8P physical partitions, the remaining resources cannot be divided into 1 16P physical partition or 4 4P physical partitions, but can only be divided into 1 or 2 8P physical partitions.
[0062] 3. The most comprehensive principle: Considering factors such as comprehensive compatibility and whether the IO resource is sufficient, if all the dividable resources are divided into one large physical partition, the system-level compatibility is definitely the most comprehensive.
[0063] 4. The customized principle: This principle is to perform customized physical partitioning on certain resources, for example, a specific GPU card and its supporting resources (such as CPU) are physically partitioned. The method for obtaining the supporting resources in this principle can be referred to the description of the physical partitioning method provided in the embodiments of the present application, which is not described herein.
[0064] It should be noted that in addition to the above, there can be other partitioning bases or principles. In actual application, the corresponding physical partitioning mode can be selected according to the user's selection or predetermined rules, which is not limited herein.
[0065] 203、the server issues a self-checking image and / or a self-checking program for the self-checking partition;
[0066] It should be noted that the self-checking image or the self-checking program can be an image or a program of the BIOS, the OS and the self-checking tool.
[0067] It can be understood that the self-checking is mainly based on one or more of the BIOS, the OS and the self-checking tool, for example, the self-checking can be performed for the CPU, the memory and the like in the BIOS, and the OPROM of the IO board card can also be called; the OS can perform the self-checking at the level of the operating system; in addition, after the OS is started, the related self-checking tool provided by the user, the manufacturer and the like can be loaded to perform the self-checking.
[0068] Optionally, in actual application, the self-checking image and / or the self-checking program can already exist in some to-be-self-checked partitions, therefore, as another embodiment of the self-checking method in the embodiment of the present application, when the present step is executed, the server can specifically execute the following steps:
[0069] The server determines whether the self-checking image and / or the self-checking program exist in each to-be-self-checked partition;
[0070] The server issues the self-checking image and / or the self-checking program for the to-be-self-checked partition in which the self-checking image and / or the self-checking program do not exist.
[0071] Optionally, after the server issues the self-checking image and / or the self-checking program, the server can also perform the pre-self-checking setting of the self-checking image and / or the self-checking program in each to-be-self-checked partition.
[0072] 204、the server powers on the to-be-self-checked partition, and uses the self-checking image and / or the self-checking program to perform the self-checking on the to-be-self-checked partition, to obtain a self-checking result.
[0073] In the present step, the server uses the BIOS, the OS and the self-checking tool contained in the self-checking image and / or the self-checking program of each to-be-self-checked partition to perform the self-checking on each to-be-self-checked partition, to obtain the self-checking result of each to-be-self-checked partition.
[0074] It can be understood that the self-checking result of each to-be-self-checked partition can be summarized and reported after all the self-checking is completed, or can be directly reported in the self-checking process, which is not limited herein.
[0075] Optionally, in the self-checking result, in addition to the information of whether the hardware resource exists, the information of all the hardware resources existing in the server can also be included, and these information can be collected in the self-checking process.
[0076] Optionally, if the self-checking result indicates that the to-be-self-checked resource in the self-checking partition has no fault, the server can execute the following operation:
[0077] When the to-be-self-checked partition includes the hardware resource which is not physically partitioned, the server powers off and deletes the to-be-self-checked partition;
[0078] and / or,
[0079] When the to-be-self-checked resource further includes the first hardware resource, the to-be-self-checked partition is powered off, deleted, and the historical physical partition of the first hardware resource is recovered according to the partition record of the backup historical physical partition of the first hardware resource.
[0080] Optionally, if the self-checking result indicates that there is a faulty hardware resource in the to-be-self-checked resource in the self-checking partition, the server can retain the self-checking partition for further positioning diagnosis and analysis and solution by the user.
[0081] In the embodiment of the application, after the server receives the self-checking instruction, the to-be-self-checked resource is physically partitioned to obtain the to-be-self-checked partition, the to-be-self-checked resource is powered on, the self-checking of the to-be-self-checked partition is performed using the self-checking image and / or the self-checking program issued, and the self-checking result is obtained, so that all the hardware resources mounted on the server can be detected in the self-checking process, the problem of non-awareness of partial hardware resource failure is solved, it is ensured that there is no problem in the hardware in the server, and the server is avoided from working with hidden dangers or faults.
[0082] The above Figure 2 The step 202 shown in the figure provides a plurality of bases for physically partitioning the to-be-self-checked resource, and the following will be based on the customization principle to explain the way of customized physical partitioning according to the partitioning demand of the user, please refer to Figure 3 , an embodiment of a possible self-checking method provided by the application, specifically including:
[0083] 301. The server determines the attachment relationship tree;
[0084] In order to clearly and intuitively represent the hanging relationship of each hardware resource in the server, a tree-like manner can be used for illustration in the present application, which can be defined as a hanging relationship tree, wherein the relationship of each node in the hanging relationship tree can be understood as the relationship of each hardware resource in the server. For example, in the hanging relationship tree, the root node is the server, the server has at least one leaf node, the leaf node of the server includes at least one platform controller hub (PCH), the leaf node of the PCH includes at least one CPU, and the leaf node of the CPU can include a peripheral component interconnect express (PCIE), a dual inline memory module (DIMM) and a disk; therefore, the hanging relationship represented by the hanging relationship tree is the hanging relationship of each hardware resource in the server, which can be obtained by the server according to the constraint relationship of physical partition and hardware running, for example, each partition includes at least one PCH.
[0085] For the convenience of understanding, Figure 4 is a schematic diagram of a hanging relationship tree. In the tree, the resources such as PCIE and DIMM can be hung under the CPU or the leaf node Node of the server, and whether it can be hung depends on the CPU architecture. In addition, the PCIE can be further refined, such as CPU card and network card, so as to facilitate the customization of physical partition according to the demand. Figure 4 The numbers underlined in the middle are set for the convenience of representation in the figure, such as the node and hierarchical relationship, and can be flexibly determined in actual application and naming.
[0086] In addition, the server, node, PCH and CPU and the like can be in a one-to-one or one-to-many relationship, for example, the PCH and the CPU can be one-to-one or one PCH corresponding to multiple CPUs; and the CPU and its lower layer PICE card and the like can be in a one-to-zero, one-to-one or one-to-many relationship, for example, one CPU can only include one PCIE card, or can include multiple DIMMs and multiple PCIE cards, or can not include a PCIE card, which is not limited here.
[0087] 302, the server receives partition information;
[0088] After determining the hanging relationship tree, the server can receive partition information sent by the user, wherein the partition information includes information of a hardware device, and the information of the hardware device is used to indicate the partition demand of the user for the hardware device.
[0089] It can be understood that the partitioning requirement can be a requirement for partitioning according to one or more hardware (such as an I / O card), or a requirement for partitioning according to one or more capabilities (such as GPU capability, virtualization capability, IO capability, etc.), which is not limited here.
[0090] Among them, the capability type that can be used as a partitioning requirement can include:
[0091] 1. The number of processor cores, mainly determined by the CPU. The difference from the current physical partitioning according to the number of CPUs is that the same requirement will produce different results according to the actual CPU model. For example, if the requirement is also 20 cores, if a CPU with a large number of cores (such as a 24-core CPU) is used, a 1P physical partition can meet the requirements, and if a CPU with a small number of cores (such as a 4-core CPU) is used, a 4P or even more physical partition is needed to meet the requirements.
[0092] 2. Virtualization capability, mainly determined by GPU;
[0093] 3. Graphics processing capability, mainly determined by GPU;
[0094] 4. Memory capability, mainly determined by memory;
[0095] 5. Storage capability, mainly determined by hard disk (corresponding to DISK in the resource table) and / or storage card (such as PCIE storage card, etc.);
[0096] 6. Network communication capability, mainly determined by network card;
[0097] In actual application, other capabilities that can be realized by the server can also be used as partitioning requirements, which are not limited here.
[0098] 303. The server determines the hardware device corresponding to the partitioning information;
[0099] After receiving the partitioning information, the server determines the hardware device corresponding to the partitioning information according to the information of the hardware device included in the partitioning information.
[0100] For example, if the partitioning information includes a requirement for corresponding physical partitioning according to one or more hardware devices (such as an I / O card), the server determines the one or more hardware devices.
[0101] If the partitioning information includes a requirement for physical partitioning according to one or more capabilities (such as GPU capability, virtualization capability, IO capability, etc.), the server determines the hardware devices that can meet these capabilities according to the capability requirements.
[0102] 304. The server queries the location of the hardware device in the hanging relationship tree;
[0103] The server finds the position of the determined hardware device in the hanging relationship tree, for example, the server determines the position of the determined GPU card in which slot (PCIE_XXXX) according to the card, so as to determine the to-be-inspected resource in the next step.
[0104] 305、The server determines the to-be-inspected resource according to the position;
[0105] After the server determines the position of the hardware device in the hanging relationship tree, the to-be-inspected resource is determined according to the obtained position, and the determination manner can include:
[0106] The server determines the hardware device and the hardware device belonging to the leaf node of the hardware device in the hanging relationship tree as the to-be-inspected resource; or,
[0107] The server determines the hardware device and the hardware device belonging to the same level node of the hardware device in the hanging relationship tree as the to-be-inspected resource.
[0108] For ease of understanding, the method for determining the to-be-inspected resource will be described below in combination with the hanging relationship tree shown in Figure 4
[0109] In this embodiment, the determination of the to-be-inspected resource can be understood as determining the smallest resource set according to the requirement, for example, in the above Figure 4 embodiment, if it is required to customize the physical partition according to a PCIE card to detect whether the card is normal, the following steps can be performed:
[0110] Step 1, determine the position of the PCIE card in the above Figure 4 , for example, assuming that the card is PCIE_1111+PCIE_2345, or CPU_1NN;
[0111] Step 2, determine from the current node to the upper node of the hanging resource tree until the minimum intersection of the PCH layer and above is found. The explanation of this step is as follows:
[0112] If the specified resource is located below the same PCH, then only the PCH node corresponding to the resource and all lower resource nodes need to be packaged to complete the determination, for example, if the resource (such as a GPU card) determined in step 1 is PCIE_1111 in the above figure, then PCH_11 and CPU_111 and all resources below it and CPU_11N and all resources below it are the to-be-inspected resource of PCIE_1111, that is, even if the specified card is irrelevant to CPU_11N, because CPU_11N and PCIE_1111 are located below the same PCH, they must be divided into the same physical partition;
[0113] If the specified resource is located at the node_1, for example, the node controller, then the resource to be self-checked must include this node and all resources under it;
[0114] If the specified resource is located at multiple nodes, then the resource to be self-checked is determined as the combination of the relevant nodes or even the entire server.
[0115] It can be understood that for multiple cards, combination and specification can be performed, for example, for three GPU cards, the three GPU cards can be specified to be combined into one physical partition, or the three GPU cards are independently divided into three physical partitions, or two cards are specified to be combined into one physical partition and the remaining one card is independently physically partitioned. In these cases, the determined partition quantity and the supporting resource condition of each physical partition are different.
[0116] It can be understood that in actual application, steps 303 to 305 can be executed multiple times, for example, a user wants to set multiple physical partitions, which respectively require virtualization capability and IO capability. If the hardware device corresponding to the partition requirement determined in step 303 is divided into other physical partitions to meet a certain partition requirement in step 305, step 303 can be executed again to determine other idle hardware devices.
[0117] 306、The server receives a self-checking instruction;
[0118] In this embodiment, step 306 is similar to step 201 shown in the above Figure 2 , and details are not repeated here.
[0119] 307、The server physically partitions the resource to be self-checked to obtain a self-checking partition;
[0120] The server actually physically partitions according to the determined partition quantity and the resource to be self-checked of each physical partition.
[0121] 308、The server issues a self-checking image and / or a self-checking program for the self-checking partition;
[0122] 309、The server powers on the resource to be self-checked, uses the self-checking image and / or the self-checking program to perform self-checking on the resource to be self-checked, and obtains a self-checking result.
[0123] In this embodiment, steps 308 to 309 are similar to steps 203 to 204 shown in the above Figure 2 , and details are not repeated here.
[0124] In the embodiment of the application, the server can customize the physical partition according to the user's requirements, simplifying the user's operation and improving the human-computer interaction performance.
[0125] It should be noted that,Figure 3 The physical partition method shown can not only be based on the scenario of the self-checking method of the present application, but also be based on other application scenarios with a physical partition process, for example, a user purchases a board card, and needs to detect the specific hardware of the board card without affecting the current physical partition and business. The physical partition method can be used for division, and the physical partition is deleted after detection is completed. Therefore, the application of the customized physical partition method according to the partition requirements of the user in the present application is not limited here.
[0126] Referring to Figure 5 An embodiment of the server that can execute the self-checking method in the present application includes:
[0127] The server includes a partition module 501 and a self-checking module 502.
[0128] The partition module 501 is configured to perform physical partition on the to-be-self-checked resources of the server in response to a received self-checking instruction, to obtain a self-checking partition.
[0129] The self-checking module 502 is configured to power on the self-checking partition and perform self-checking on the self-checking partition.
[0130] In the embodiment of the present application, the partition module performs physical partition on the to-be-self-checked resources in response to the received self-checking instruction to obtain a self-checking partition, and the self-checking module powers on the self-checking partition and performs self-checking on the self-checking partition. As a result, all hardware resources or specified hardware resources mounted on the server can be detected in the self-checking process, the problem of fault unawareness of part of the hardware resources is solved, it is ensured that there is no problem with the hardware in the server, and hidden dangers or fault work of the server are avoided.
[0131] Referring to Figure 6 Another embodiment of the server that can execute the self-checking method in the present application includes:
[0132] The server includes a partition module 601 and a self-checking module 602.
[0133] The partition module 601 is configured to perform physical partition on the to-be-self-checked resources of the server in response to a received self-checking instruction, to obtain a self-checking partition.
[0134] The self-checking module 602 is configured to power on the self-checking partition and perform self-checking on the self-checking partition.
[0135] Optionally, as another embodiment of the server in the present application, the to-be-self-checked resources further include a first hardware resource, and the first hardware resource is a hardware resource that has been divided into a historical physical partition but has not been powered on before receiving the self-checking instruction.
[0136] Optionally, as another embodiment of the server in the present application, the partition module 601 is specifically used for:
[0137] deleting the historical physical partition, and performing physical partition on the first hardware resource of the deleted historical physical partition to obtain the self-check partition.
[0138] Optionally, as another embodiment of the server in the present application, the server can further include:
[0139] The first receiving module 603 is configured to receive partition information, and the partition information includes information of a hardware device.
[0140] The querying module 604 is configured to query a position of the hardware device in a hanging relationship tree.
[0141] The determining module 605 is configured to determine the to-be-self-checked resource according to the position.
[0142] Optionally, as another embodiment of the server in the present application, the determining module 605 is specifically used for:
[0143] determining the hardware device and a hardware device belonging to a leaf node of the hardware device in the hanging relationship tree as the to-be-self-checked resource.
[0144] Optionally, as another embodiment of the server in the present application, the determining module 605 is specifically further used for:
[0145] determining the hardware device and a hardware device belonging to a same level node of the hardware device in the hanging relationship tree as the to-be-self-checked resource.
[0146] Optionally, as another embodiment of the server in the present application, the server further includes:
[0147] The second receiving module 606 is configured to receive to-be-self-checked resource normal information in the self-check partition.
[0148] The powering-off module 607 is configured to power off the self-check partition.
[0149] The deleting module 608 is configured to delete the self-check partition.
[0150] Optionally, as another embodiment of the server in the present application, the server can further include:
[0151] The backup module 609 is configured to backup partition records of the historical physical partition before the deleting module deletes the historical physical partition of the first hardware resource.
[0152] Optionally, as another embodiment of the server in the present application, the server can further include:
[0153] The recovery module 610 is configured to recover the historical physical partition of the first hardware resource according to the backup partition record after the deletion module 608 deletes the self-check partition.
[0154] Optionally, as another embodiment of the server in the embodiments of the present application, the server can further include:
[0155] The third receiving module 611 is further configured to receive the resource fault information to be checked in the self-check partition.
[0156] In the embodiments of the present application, the server can customize the physical partition according to the user's needs, simplifying the user's operation and improving the human-computer interaction performance.
[0157] Figure 4 The specific implementation of the controller 110 shown in FIG. 11 can refer to the foregoing embodiments, wherein each module can be implemented by a corresponding hardware chip. In another implementation, one or more modules can be integrated on one hardware chip.
[0158] The server in the embodiments of the present application is described from the perspective of unitized functional entities above, and the server in the embodiments of the present application is described from the perspective of hardware processing below. Please refer to Figure 7 The server 700 provided in the embodiments of the present application can have a great difference due to different configurations or performances, and can include one or more central processing units (CPUs) 701 (for example, one or more processors) and a memory 709, one or more storage media 708 (for example, one or more mass storage devices) for storing application programs 709 or data 709. The memory 709 and the storage media 708 can be temporary storage or persistent storage. The programs stored in the storage media 708 can include one or more modules (not shown in the figure), and each module can include a series of instruction operations in the server. Further, the processor 701 can be configured to communicate with the storage media 708 and execute a series of instruction operations in the storage media 708 on the server 700.
[0159] The server 700 can further include one or more power supplies 702, one or more wired or wireless network interfaces 703, one or more input and output interfaces 704, and / or one or more operating systems 705, such as Windows Server, Mac OS X, Unix, Linux, FreeBSD, etc. Those skilled in the art can understand that Figure 7The server structure shown in the figure is not intended to limit the server, and can include more or fewer components than shown, or combine certain components, or arrange different components.
[0160] The following describes the server in detail. Figure 7 The various components of the server are described in detail as follows:
[0161] The processor 701 is the control center of the server and can process according to a set self-checking method. The processor 701 connects various parts of the server through various interfaces and lines, and performs various functions of the server and processes data by running or executing software programs and / or modules stored in the memory 709 and calling data stored in the memory 709.
[0162] The memory 709 can be used to store software programs and modules, and the processor 701 executes various functions and data processing of the server 700 by running the software programs and modules stored in the memory 709. The memory 709 can mainly include a program storage area and a data storage area, wherein the program storage area can store an operating system, at least one application program required by a function, etc.; and the data storage area can store data created according to the use of the server, etc. In addition, the memory 709 can include a high-speed random access memory, and can also include a non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other volatile solid-state memory device. The programs of the self-checking method provided in the embodiments of the present application and the received data stream are stored in the memory, and when needed, the processor 701 calls them from the memory 709.
[0163] The processor 701 is configured to execute the following steps by calling the operation instructions stored in the memory 709:
[0164] In response to the received self-checking instruction, the to-be-self-checked resources of the server are physically partitioned to obtain a self-checking partition;
[0165] Powering on the self-checking partition and performing self-checking on the self-checking partition.
[0166] In some embodiments of the present application, the to-be-self-checked resources include a first hardware resource of the server, and the first hardware resource is a hardware resource that has been divided into a historical physical partition before receiving the self-checking instruction but has not been powered on. The processor 701 is specifically configured to execute the following operations:
[0167] Delete the historical physical partition; and physically partition the first hardware resource on which the historical physical partition is deleted to obtain the self-checking partition.
[0168] In some embodiments of the present application, the input / output interface 704 is further configured to execute the following operations:
[0169] Receive partition information; the partition information contains information of the hardware device;
[0170] The processor 701 is specifically configured to perform the following operations:
[0171] Query the position of the hardware device in the hanging relationship tree, and determine the to-be-self-checked resource according to the position.
[0172] In some embodiments of the application, the processor 701 is specifically configured to perform the following operations:
[0173] Determine the hardware device and the hardware device belonging to the leaf node of the hardware device in the hanging relationship tree as the to-be-self-checked resource.
[0174] In some embodiments of the application, the processor 701 is specifically configured to perform the following operations:
[0175] Determine the hardware device and the hardware device belonging to the same level node of the hardware device in the hanging relationship tree as the to-be-self-checked resource.
[0176] In some embodiments of the application, the input and output interface 704 is further configured to perform the following operations:
[0177] Receive the to-be-self-checked resource normal information in the self-check partition;
[0178] The processor 701 is further configured to perform the following operations:
[0179] Power off the self-check partition, and delete the self-check partition.
[0180] In some embodiments of the application, the processor 701 is further configured to perform the following operations:
[0181] Before deleting the historical physical partition of the first hardware resource, backup the partition record of the historical physical partition.
[0182] The processor 701 is further configured to perform the following operations:
[0183] After deleting the self-check partition, restore the historical physical partition of the first hardware resource according to the backup partition record.
[0184] In some embodiments of the application, the input and output interface 704 is further configured to perform the following operations:
[0185] Receive the to-be-self-checked resource failure information in the self-check partition.
[0186] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working process of the system, device and unit described above can refer to the corresponding process in the foregoing method embodiments, which will not be repeated here.
[0187] In several embodiments provided in the present application, it should be understood that the disclosed system, device and method can be implemented in other manners. For example, the described device embodiments are merely schematic. For example, the division of the units is only a logical function division. There can be another division manner for the actual implementation, for example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the displayed or discussed mutual couplings or direct couplings or communication connections between the units can be indirect couplings or communication connections through some interfaces, devices or units, and can be in electrical, mechanical or other forms.
[0188] The units described as separate components can or can not be physically separate, and the components shown as units can or can not be physical units, i.e., can be located in one place, or can be distributed on multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the embodiment.
[0189] When the integrated unit is implemented in the form of a software function unit and sold or used as an independent product, it can be stored in a computer readable storage medium. Based on such understanding, the technical solutions of the present application essentially or the part that contributes to the prior art or the whole or part of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium, and includes several instructions to make a computer device (which can be a personal computer, a server, or a network device, etc.) execute all or part of the steps of the methods described in the various embodiments of the present application. The foregoing storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), magnetic disk or optical disk, and various program codes that can be stored in the medium.
[0190] The above described and above embodiments are only used to illustrate the technical solutions of the present application, and not to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or make equivalent replacements for some technical features; and these modifications or replacements do not make the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application.
Claims
1. A self-checking method, characterized by, The method comprises: The server responds to the received self-checking instruction, and physically partitions the to-be-self-checked resources of the server to obtain a self-checking partition, wherein the to-be-self-checked resources comprise first hardware resources of the server, and the first hardware resources are hardware resources that have been divided into historical physical partitions but not powered on before receiving the self-checking instruction; The server powers on the self-checking partition, and performs self-checking on the self-checking partition; The server responds to the received self-checking instruction, and physically partitions the to-be-self-checked resources of the server to obtain the self-checking partition, and specifically comprises: The server deletes the historical physical partition; The server physically partitions the first hardware resources deleted from the historical physical partition to obtain the self-checking partition.
2. The method of claim 1, wherein, Before the server responds to the received self-checking instruction, and physically partitions the to-be-self-checked resources of the server to obtain a self-checking partition, the method further comprises: The server receives partition information, wherein the partition information comprises information of a hardware device; The server queries a position of the hardware device in a hanging relationship tree; The server determines the to-be-self-checked resources according to the position.
3. The method of claim 2, wherein, The server determines the to-be-self-checked resources according to the position, and specifically comprises: The server determines the hardware device and hardware devices belonging to leaf nodes of the hardware device in the hanging relationship tree as the to-be-self-checked resources.
4. The method of claim 3, wherein, The server determines the to-be-self-checked resources according to the position, and specifically comprises: The server determines the hardware device and hardware devices belonging to the same level nodes of the hardware device in the hanging relationship tree as the to-be-self-checked resources.
5. The method according to any one of claims 1 to 4, characterized in that, The method further comprises: The server receives to-be-self-checked resource normal information in the self-checking partition; The server powers off the self-checking partition; The server deletes the self-checking partition.
6. The method of claim 5, wherein, The method further comprises: Before the server deletes the historical physical partition of the first hardware resources, the method further comprises:
7. The method of claim 6, wherein, The server restores the historical physical partition of the first hardware resources according to the backup partition record after deleting the self-checking partition. The method further comprises:
8. The method according to any one of claims 1 to 4, characterized in that, The server receives to-be-self-checked resource failure information in the self-checking partition. The server comprises a partition module and a self-checking module; 9. A server, characterized by The partition module is configured to respond to a received self-checking instruction, and physically partition to-be-self-checked resources of the server to obtain a self-checking partition, wherein the to-be-self-checked resources comprise first hardware resources of the server, and the first hardware resources are hardware resources that have been divided into historical physical partitions but not powered on before receiving the self-checking instruction; The self-checking module is configured to power on the self-checking partition, and perform self-checking on the self-checking partition; The partition module is specifically configured to: delete the historical physical partition; and physically partition the first hardware resources deleted from the historical physical partition to obtain the self-checking partition. The server further comprises:
10. The server of claim 9, wherein, A first receiving module configured to receive partition information, wherein the partition information comprises information of a hardware device; The query module is configured to query a position of the hardware device in the hanging relationship tree. The determination module is configured to determine the to-be-self-checked resource according to the position.
11. The server of claim 10, wherein, The determination module is specifically configured to: determine the hardware device and a hardware device belonging to a leaf node of the hardware device in the hanging relationship tree as the to-be-self-checked resource.
12. The server of claim 11, wherein, The determination module is specifically configured to: determine the hardware device and a hardware device belonging to a same level node of the hardware device in the hanging relationship tree as the to-be-self-checked resource.
13. The server of any one of claims 9-12, wherein, The server further comprises: The second receiving module is configured to receive to-be-self-checked resource normal information in the self-checking partition. The power-off module is configured to power off the self-checking partition. The deletion module is configured to delete the self-checking partition.
14. The server of claim 13, wherein, The server further comprises: The backup module is configured to backup partition records of the historical physical partition of the first hardware resource before the deletion module deletes the historical physical partition.
15. The server of claim 14, wherein, The server further comprises: The recovery module is configured to recover the historical physical partition of the first hardware resource according to the backup partition records after the deletion module deletes the self-checking partition.
16. The server of any one of claims 9-12, wherein, The server further comprises: The third receiving module is configured to receive to-be-self-checked resource failure information in the self-checking partition.
17. A computer readable storage medium comprising instructions which, when executed on a computer, cause the computer to carry out the method of any one of claims 1-8.
18. A computer program product comprising instructions which, when executed on a computer, cause the computer to carry out the method of any one of claims 1-8.
Citation Information
Patent Citations
Information processing device, failure processing method, program and recording medium thereof
CN101211283A
Cache optimized logical partitioning a symmetric multi-processor data processing system
CN1604041A