Server startup method, computer device, readable storage medium, and program product
The baseboard management controller identifies and evaluates the hard disk topology and health status, generates a prioritized list of repairs, and dynamically adjusts the server startup strategy, solving the problems of low startup efficiency and failure risks caused by the multi-level unit cascade architecture and achieving efficient and reliable server startup.
Patent Information
- Application Number
- CN202511023166.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-24
- Publication Date
- 2025-09-23
- Estimated Expiration
- 2045-07-24
AI Technical Summary
The multi-level unit cascade architecture increases the depth of the PCIe topology layer, resulting in reduced server startup efficiency, exponentially increased enumeration time, increased risk of cascading faulty disks, and rigid resource allocation, which cannot meet the rapid elastic scaling requirements of AI clusters.
The baseboard management controller obtains the historical and current topological fingerprints of the hard disk, identifies the changed hard disk, performs health information scanning and scoring, generates a list of hard disks to be repaired, and initializes them according to priority and processing methods, dynamically adjusting the startup sequence and strategy.
It improves the reliability and efficiency of server startup, reduces the risk of cascading disk failures, and meets the rapid elastic scaling requirements of AI clusters.
Smart Images

Figure CN120523760B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of server startup, and in particular to a server startup method, computer equipment, readable storage medium, and program product. Background Art
[0002] With the development of artificial intelligence training and high-frequency trading, modern data centers need a large number of storage devices to meet the massive storage needs of AI training, real-time trading and other businesses. To improve storage density, the industry generally adopts a multi-level unit cascade architecture, such as the JBOF (Just a Bunch of Flash) unit cascade architecture. The multi-level unit cascade architecture will significantly increase the depth of the PCIe topology layer, complicate storage device initialization, and reduce server startup efficiency. Summary of the Invention
[0003] The present application provides a server startup method, a computer device, a readable storage medium, and a program product to solve the technical problem of low server startup efficiency.
[0004] The present application provides a server startup method, comprising: in response to the server being powered on, a baseboard management controller obtains a historical topological fingerprint of a hard disk and a current topological fingerprint of the hard disk, and uses a hard disk whose historical topological fingerprint and current topological fingerprint do not match as a changed hard disk; the baseboard management controller scans the changed hard disk, obtains health information of the changed hard disk, and determines a health score of the changed hard disk based on the health information of the changed hard disk; obtains a hard disk status and a hard disk processing method of the changed hard disk corresponding to the health score of the changed hard disk, determines a list of hard disks to be repaired according to the hard disk status and the hard disk processing method of the changed hard disk, and the baseboard management controller sends the list of hard disks to be repaired to a basic input / output system; in response to the basic input / output system obtaining the list of hard disks to be repaired, determines the priorities of multiple hard disks to be repaired in the list of hard disks to be repaired, and initializes the multiple hard disks to be repaired according to the priority levels of the multiple hard disks to be repaired and the hard disk processing methods of the hard disks to be repaired, so as to realize server startup.
[0005] The present application also provides a computer device, comprising: a memory for storing a computer program; and a processor for implementing the steps of the server startup method in the following embodiments when executing the computer program:
[0006] In response to the server being powered on, the baseboard management controller obtains the historical topology fingerprint and the current topology fingerprint of the hard disk, and takes the hard disk whose historical topology fingerprint and the current topology fingerprint do not match as the changed hard disk; the baseboard management controller scans the changed hard disk, obtains the health information of the changed hard disk, and determines the health score of the changed hard disk based on the health information of the changed hard disk; obtains the hard disk status and hard disk processing method of the changed hard disk corresponding to the health score of the changed hard disk, determines the list of hard disks to be repaired according to the hard disk status and the hard disk processing method of the changed hard disk, and sends the list of hard disks to be repaired to the basic input and output system; in response to the basic input and output system obtaining the list of hard disks to be repaired, determines the priorities of multiple hard disks to be repaired in the list of hard disks to be repaired, and initializes the multiple hard disks to be repaired according to the priority levels of the multiple hard disks to be repaired and the hard disk processing methods of the hard disks to be repaired, so as to realize server startup.
[0007] The present application also provides a computer-readable storage medium, wherein a computer program is stored in the computer-readable storage medium, wherein when the computer program is executed by a processor, the steps of the server startup method in the following embodiment are implemented:
[0008] In response to the server being powered on, the baseboard management controller obtains the historical topology fingerprint and the current topology fingerprint of the hard disk, and takes the hard disk whose historical topology fingerprint and the current topology fingerprint do not match as the changed hard disk; the baseboard management controller scans the changed hard disk, obtains the health information of the changed hard disk, and determines the health score of the changed hard disk based on the health information of the changed hard disk; obtains the hard disk status and hard disk processing method of the changed hard disk corresponding to the health score of the changed hard disk, determines the list of hard disks to be repaired according to the hard disk status and the hard disk processing method of the changed hard disk, and sends the list of hard disks to be repaired to the basic input and output system; in response to the basic input and output system obtaining the list of hard disks to be repaired, determines the priorities of multiple hard disks to be repaired in the list of hard disks to be repaired, and initializes the multiple hard disks to be repaired according to the priority levels of the multiple hard disks to be repaired and the hard disk processing methods of the hard disks to be repaired, so as to realize server startup.
[0009] The present application also provides a computer program product, including a computer program, which, when executed by a processor, implements the steps of the server startup method in the following embodiments:
[0010] In response to the server being powered on, the baseboard management controller obtains the historical topology fingerprint and the current topology fingerprint of the hard disk, and takes the hard disk whose historical topology fingerprint and the current topology fingerprint do not match as the changed hard disk; the baseboard management controller scans the changed hard disk, obtains the health information of the changed hard disk, and determines the health score of the changed hard disk based on the health information of the changed hard disk; obtains the hard disk status and hard disk processing method of the changed hard disk corresponding to the health score of the changed hard disk, determines the list of hard disks to be repaired according to the hard disk status and the hard disk processing method of the changed hard disk, and sends the list of hard disks to be repaired to the basic input and output system; in response to the basic input and output system obtaining the list of hard disks to be repaired, determines the priorities of multiple hard disks to be repaired in the list of hard disks to be repaired, and initializes the multiple hard disks to be repaired according to the priority levels of the multiple hard disks to be repaired and the hard disk processing methods of the hard disks to be repaired, so as to realize server startup.
[0011] The server startup method provided by the present application, in response to the server being powered on, the baseboard management controller obtains the historical topology fingerprint of the hard disk and the current topology fingerprint of the hard disk, and takes the hard disk whose historical topology fingerprint and current topology fingerprint do not match as the changed hard disk; the baseboard management controller scans the changed hard disk, obtains the health information of the changed hard disk, and determines the health score of the changed hard disk based on the health information of the changed hard disk; obtains the hard disk status and hard disk processing method of the changed hard disk corresponding to the health score of the changed hard disk, determines a list of hard disks to be repaired according to the hard disk status and hard disk processing method of the changed hard disk, and the baseboard management controller sends the list of hard disks to be repaired to the basic input and output system; in response to the basic input and output system obtaining the list of hard disks to be repaired, determines the priorities of multiple hard disks to be repaired in the list of hard disks to be repaired, and initializes the multiple hard disks to be repaired according to the priority levels of the multiple hard disks to be repaired and the hard disk processing methods of the hard disks to be repaired, so as to realize server startup. In this way, by comparing the current topological fingerprint and historical topological fingerprint of the hard disk, the changed hard disk can be located quickly and accurately, the health status of the changed hard disk can be classified, a list of hard disks to be repaired can be generated, and the hard disks to be repaired in the list can be prioritized. Differentiated initialization can be performed according to the hard disk processing method corresponding to the priority of the hard disk to be repaired and the health status of the hard disk to be repaired, which can improve the reliability of server startup. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] In order to more clearly illustrate the embodiments of the present application, the following is a brief introduction to the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0013] Figure 1A schematic diagram of the structure of a server startup system provided in one embodiment of the present application;
[0014] Figure 2 A flowchart of a server startup method provided in one embodiment of the present application;
[0015] Figure 3 A flowchart of a server startup method provided in another embodiment of the present application;
[0016] Figure 4 A flowchart of a server startup method provided in another embodiment of the present application;
[0017] Figure 5 A schematic diagram of the structure of a server startup device provided in one embodiment of the present application;
[0018] Figure 6 This is a diagram of the internal structure of a computer device provided in one embodiment of the present application. DETAILED DESCRIPTION
[0019] The following will be combined with the accompanying drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of them. Based on the embodiments of this application, all other embodiments obtained by ordinary technicians in this field without making any creative work are within the scope of protection of this application.
[0020] It should be noted that, in the description of this application, the terms "comprises," "includes," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. The terms "first," "second," etc., in this application are used to distinguish similar objects, and are not used to describe a particular order or sequence.
[0021] The multi-level unit cascade architecture will significantly increase the depth of the PCIe topology layer and also lead to three core problems such as exponential growth in enumeration time: First, the boot time deteriorates exponentially. The traditional BIOS (Basic Input Output System) needs to enumerate PCIe devices layer by layer, resulting in an exponential increase in enumeration time. The startup time of a cluster with a large number of disks will exceed expectations and cannot meet the rapid elastic scaling requirements of AI clusters; second, the risk of cascading faulty disks is exacerbated: the surge in the number of hard drives has expanded the health detection blind spot, and competition among multiple controllers has caused link training conflicts. Single-point failures can easily spread to the entire cascade path; it also leads to rigid resource allocation: fixed retry strategies cause key device initialization to be blocked in time-sensitive scenarios (such as real-time trading systems), and system reliability has plummeted.
[0022] In order to enable those skilled in the art to better understand the present application, the present application is further described in detail below with reference to the accompanying drawings and specific implementation methods.
[0023] like Figure 1 As shown, an embodiment of the present application provides a server boot system, which specifically includes: a basic input and output system, a baseboard management controller, a complex programmable logic device, a hard disk expansion device, and multiple hard disks.
[0024] The Basic Input / Output System (BIOS) is the system that allows the host computer to communicate with the outside world. It consists of two parts: peripheral devices and an input / output control system, and is a crucial component of a computer system. Peripheral devices include input and output devices, as well as disk storage, tape storage, and optical disk storage. The BIOS is connected to the baseboard management controller.
[0025] The baseboard management controller (BMC) is an independent service processor based on the Intelligent Platform Management Interface (IPMI) specification. This device monitors server hardware status data, including power, CPU, memory, hard disk, and environmental parameters, in real time through a built-in sensor network. It also supports out-of-band communication with the motherboard via the baseband management interface. It provides remote power on / off control, firmware updates, event logging, and fault diagnosis capabilities, making it widely used in scenarios such as data center server cluster monitoring and remote maintenance of distributed equipment. The baseboard management controller connects to the complex programmable logic device (CPLD) via a dedicated bus, which can be an I2C serial communication bus.
[0026] The baseboard management controller and complex programmable logic device are connected to multiple hard drive expansion devices (hard drive expansion device 1 and hard drive expansion device n in the figure). Each hard drive expansion device is connected to multiple hard drives. A complex programmable logic device (CPLD) is a high-density programmable logic device with an integration density greater than 1,000 gates and a large number of input / output signals, product terms, and macrocells. The hard drive expansion device can be a JBOF (Just a Bunch of Flash) storage device in a server that connects multiple hard drives. The multiple hard drives connected to the hard drive expansion device can be NVME drives.
[0027] like Figure 2 As shown, an embodiment of the present application provides a server startup method, which specifically includes the following steps:
[0028] Step 101: In response to the server being powered on, the baseboard management controller obtains the historical topology fingerprints and the current topology fingerprints of the hard disks, and regards the hard disks whose historical topology fingerprints and current topology fingerprints do not match as changed hard disks.
[0029] Specifically, the baseboard management controller and the complex programmable logic device jointly scan the hard disk expansion device to obtain the current cascade path of the hard disk between the hard disk expansion device and the hard disk; the baseboard management controller obtains the current hard disk key information, which includes the communication rate of the system accessing the hard disk, the parameters for determining the negotiation rate when the system accesses the hard disk, and the hard disk number; the current topology fingerprint of the hard disk is determined based on the current cascade path of the hard disk and the current hard disk key information; the baseboard management controller obtains the historical topology fingerprint of the hard disk, and compares the historical topology fingerprint of the hard disk with the current topology fingerprint of the hard disk.
[0030] When the server powers on, the baseboard management controller (BMC) and complex programmable logic device (CPLD) collaborate to scan the system's drive expansion devices, perform physical layer detection, and read the PCIe switches within each drive expansion device (JBOF). They record the mapping between downstream ports and drives to determine the cascade path between the drive expansion device and the drives. The BMC then collects key drive information, which can be obtained through the link status register. Specifically, it includes the current negotiated rate mode (the communication rate at which the system accesses the drives), the transmit Tx equalization preset value (0-10) read from the vendor-specific space, which determines the negotiated rate when the system accesses the drives, and the drive ID, which includes the drive serial number (the drive's unique factory ID) and the drive part number (the drive's factory batch number). The BMC then uses a hash algorithm to calculate the current cascade path and key drive information, resulting in the current topology fingerprint value for each drive.
[0031] After obtaining the current topology fingerprint value of each hard disk in the system, the current topology fingerprint of each hard disk in the system can be constructed. Specifically, it can include concatenating the current cascade path of the hard disk and the current hard disk key information into a string; performing hash calculation on the string to obtain the current topology fingerprint value of the hard disk; storing the current cascade path of the hard disk, the current hard disk key information and the current topology fingerprint value of the hard disk in a tree topology structure in the target storage area of the baseboard management controller to generate the current topology fingerprint of the hard disk.
[0032] In this application, each hard disk is regarded as a region, and the topological fingerprint of the hard disk is stored by region. As shown in Table 1, Root Port-SWITCH-JBOFn-NVMEn is regarded as a region, and the topological fingerprint of the hard disk is saved in the designated area of the baseboard management controller. The topological fingerprint of each area is obtained in turn, which can be accurate to Root Port-SWITCH-JBOF-NVME. Here, the Root Port device is used to connect the CPU / memory subsystem and the I / O device.
[0033] Table 1
[0034]
[0035] The topology fingerprint in this application can be stored in the form of a PCIe topology. PCIe (Peripheral Component Interconnect Express) is a high-speed serial computer expansion bus standard with a tree-like topology. In this structure, the Root Complex serves as the root of the tree and is responsible for communicating with the CPU and other PCIe devices. The Root Complex typically contains multiple subcomponents, such as processor interfaces and memory controllers, and may support multiple PCIe ports. These ports are used to connect external PCIe devices, such as graphics cards and hard drives. In addition, switches and bridges are also used to expand connections, allowing more devices to be connected to a single PCIe port.
[0036] The hash calculation method used in this application can specifically be the SHA-256 algorithm. The SHA-256 algorithm is used to encrypt the hard disk's Root Port, Switch downstream port mapping relationship, average negotiation rate, Tx balance Preset value average, and NVME PN serial number parameters to generate a unique hash result as the hard disk's topological fingerprint. Any slight change in any parameter will result in a completely different hash value. By comparing the hard disk's current topological fingerprint and historical topological fingerprint, the hard disk changes can be quickly and accurately identified, thereby determining the hard disk processing method for the changed hard disk based on the health score of the changed hard disk, which can improve the server startup rate.
[0037] In addition, SHA-256 is a one-way hash, the original parameters cannot be deduced from the result, and it is tamper-proof.
[0038] See also Figure 3In this application, when comparing the historical topological fingerprint of the hard disk with the current topological fingerprint, the historical topological fingerprint can be compared one by one with the information in the current topological fingerprint, or the historical topological fingerprint value of the historical topological fingerprint can be compared with the current topological fingerprint value of the current topological fingerprint. The description is made by taking the comparison of the information in the historical topological fingerprint with the current topological fingerprint as an example. In response to the historical topological fingerprint of the hard disk not matching the current topological fingerprint, the hard disk whose historical topological fingerprint and current topological fingerprint do not match is treated as a changed hard disk, and a first mark is made. The first mark indicates that the hard disk topology has changed, and the hard disk with the first mark needs to be scanned. The changed hard disk is added to the queue to be scanned. The subsequent baseboard management controller obtains the changed hard disk from the scan queue for health status scanning. In response to the historical topological fingerprint of the hard disk matching the current topological fingerprint, the hard disk whose historical topological fingerprint and current topological fingerprint match is marked with a second mark. The second mark indicates that the hard disk topology has not changed, that is, the cascade relationship of the hard disk has not changed, and the historical health information of the hard disk with the second mark is loaded.
[0039] Step 102: The baseboard management controller scans the changed hard disk, obtains health information of the changed hard disk, and determines a health score of the changed hard disk based on the health information of the changed hard disk; obtains the hard disk status and hard disk processing method of the changed hard disk corresponding to the health score of the changed hard disk, determines a list of hard disks to be repaired based on the hard disk status and hard disk processing method of the changed hard disk, and the baseboard management controller sends the list of hard disks to be repaired to the basic input and output system.
[0040] Specifically, the baseboard management controller obtains the health information of the changed hard disk, and the health information includes the hard disk temperature stability, hard disk failure rate, and hard disk life status. Among them, the hard disk temperature stability can be obtained by continuously monitoring a period of time, such as the hard disk slot temperature fluctuation within 72 hours, recording the number of temperature mutations in a single day, and recording the number of fluctuations that exceed expectations; the hard disk failure rate can be obtained by counting the number of command timeouts that occur within a certain period of time, such as 7 days, recording data transmission verification failure events, and tracking the number of abnormal power outages of the hard disk; the hard disk life status can be obtained by reading the cumulative power-on time of the hard disk.
[0041] The baseboard management controller obtains target scoring conditions corresponding to the health information of the changed hard disk, and the target scoring conditions include a target temperature stability scoring condition, a target hard disk failure rate scoring condition, and a target hard disk life scoring condition; the baseboard management controller obtains the target temperature stability scoring condition, the target hard disk failure rate scoring condition, and the target hard disk life scoring condition respectively corresponding to the target temperature stability scoring condition, the target hard disk failure rate scoring condition, and the target hard disk life scoring condition; and calculates the health score of the changed hard disk based on the target temperature stability scoring calculation scheme, the target hard disk failure rate scoring calculation scheme, and the target hard disk life scoring calculation scheme.
[0042] Specifically, obtain the target temperature stability score interval corresponding to the target temperature stability scoring condition, and calculate the temperature stability health score of the changed hard disk based on the upper limit value of the target temperature stability score interval, the lower limit value of the target temperature stability scoring condition, and the hard disk temperature stability of the changed hard disk; obtain the target hard disk failure rate score interval corresponding to the target hard disk failure rate scoring condition, and calculate the hard disk failure rate health score of the changed hard disk based on the upper limit value of the target hard disk failure rate score interval and the hard disk failure rate of the changed hard disk; obtain the target hard disk life score interval corresponding to the target hard disk life scoring condition, and calculate the hard disk life health score of the changed hard disk based on the upper limit value of the target hard disk life score interval, the lower limit value of the target hard disk life scoring condition, and the hard disk life of the changed hard disk; and take the sum of the temperature stability health score of the changed hard disk, the hard disk failure rate health score of the changed hard disk, and the hard disk life health score of the changed hard disk as the health score of the changed hard disk.
[0043] This application provides a pre-set health scoring rule table, which includes multiple scoring conditions, multiple scoring intervals, and multiple scoring calculation schemes corresponding to hard disk temperature stability, hard disk failure rate, and hard disk life status.
[0044] The pre-set health scoring rule table can be shown in Table 2:
[0045] Table 2
[0046]
[0047] In a specific example, as shown in Table 2, the multiple scoring conditions corresponding to the hard disk temperature stability include a first temperature stability scoring condition (72-hour fluctuation ≤ 3°C), a second temperature stability scoring condition (72-hour fluctuation 3-5°C), and a third temperature stability scoring condition (72-hour fluctuation > 5°C); the multiple scoring conditions corresponding to the hard disk failure rate include a first hard disk failure rate scoring condition (7 days without error), a second hard disk failure rate scoring condition (1-3 errors), and a third hard disk failure rate scoring condition (≥4 errors); the multiple scoring conditions corresponding to the hard disk life include a first hard disk life scoring condition (power-on < 20,000 hours), a second hard disk life scoring condition (20,000-30,000 hours), and a third hard disk life scoring condition (>30,000 hours).
[0048] The hard disk temperature stability of the changed hard disk can be obtained, and the hard disk temperature stability of the changed hard disk can be compared with the multiple scoring conditions corresponding to the hard disk temperature stability in the health scoring rule table, and the hard disk temperature stability scoring condition that matches the hard disk temperature stability of the changed hard disk can be used as the target temperature stability scoring condition; the hard disk failure rate of the changed hard disk can be obtained, and the hard disk failure rate of the changed hard disk can be compared with the multiple scoring conditions corresponding to the hard disk failure rate in the health scoring rule table, and the hard disk failure rate scoring condition that matches the hard disk failure rate of the changed hard disk can be used as the target hard disk failure rate scoring condition; the hard disk life of the changed hard disk can be obtained, and the hard disk life of the changed hard disk can be compared with the multiple scoring conditions corresponding to the hard disk life in the health scoring rule table, and the hard disk life scoring condition that matches the hard disk life of the changed hard disk can be used as the target hard disk life scoring condition, that is, depending on which scoring condition the actual hard disk temperature stability of the changed hard disk, the hard disk failure rate of the changed hard disk, and the hard disk life of the changed hard disk fall within, this scoring condition can be used as the target dimension scoring condition.
[0049] As can be seen from Table 2, each scoring condition corresponds to a scoring range and calculation method. By obtaining the target temperature stability scoring condition, target hard drive failure rate scoring condition, and target hard drive life scoring condition, the target temperature stability scoring calculation scheme, target hard drive failure rate scoring calculation scheme, and target hard drive life scoring calculation scheme corresponding to the target temperature stability scoring condition, target hard drive failure rate scoring calculation scheme, and target hard drive life scoring calculation scheme, respectively, the health score of the changed hard drive is calculated based on the target temperature stability scoring calculation scheme, target hard drive failure rate scoring calculation scheme, and target hard drive life scoring calculation scheme.
[0050] In response to the target temperature stability scoring condition being the first temperature stability scoring condition, a first temperature stability score interval (35-40) corresponding to the first temperature stability scoring condition is obtained, and a subtraction operation is performed with the first temperature stability score interval upper limit value (40) of the first temperature stability score interval as the minuend and double the hard disk temperature stability of the changed hard disk as the subtrahend (fluctuation value × 2) to obtain the temperature stability health score of the changed hard disk.
[0051] In response to the target temperature stability scoring condition being the second temperature stability scoring condition, a second temperature stability score interval (20-34) corresponding to the second temperature stability scoring condition is obtained, and the difference between the hard disk temperature stability of the changed hard disk and the lower limit value (3) of the second temperature stability scoring condition is calculated. The upper limit value (34) of the second temperature stability score interval is used as the minuend, and six times the difference between the hard disk temperature stability of the changed hard disk and the lower limit value of the second temperature stability scoring condition is used as the subtrahend to perform a subtraction operation (fluctuation value - 3) × 6 to obtain the temperature stability health score of the hard disk to be repaired.
[0052] In response to the target temperature stability scoring condition being the third temperature stability scoring condition, a third temperature stability score interval (0-19) corresponding to the third temperature stability scoring condition is obtained, and the difference between the hard disk temperature stability of the changed hard disk and the lower limit value (5) of the third temperature stability scoring condition is calculated. The upper limit value (19) of the third temperature stability score interval of the third temperature stability score interval is used as the minuend, and a subtraction operation is performed with three times the difference between the hard disk temperature stability of the changed hard disk and the lower limit value of the third temperature stability scoring condition (fluctuation value - 5) × 3 as the subtrahend to obtain the temperature stability health score of the hard disk to be repaired.
[0053] In response to the target hard disk failure rate scoring condition being the first hard disk failure rate scoring condition, a first hard disk failure rate score interval (35-40) corresponding to the first hard disk failure rate scoring condition is obtained, and a first hard disk failure rate score (40) corresponding to the first hard disk failure rate score interval is obtained as the hard disk failure rate health score of the changed hard disk.
[0054] In response to the target hard disk failure rate scoring condition being the second hard disk failure rate scoring condition, a second hard disk failure rate score interval (20-34) corresponding to the second hard disk failure rate scoring condition is obtained, and a subtraction operation is performed using the upper limit value (34) of the second hard disk failure rate score interval as the minuend and five times the hard disk failure rate of the changed hard disk (number of errors × 5) as the subtrahend to obtain the hard disk failure rate health score of the changed hard disk.
[0055] In response to the target hard drive failure rate scoring condition being the third hard drive failure rate scoring condition, a third hard drive failure rate score interval (0-19) corresponding to the third hard drive failure rate scoring condition is obtained, a difference between the hard drive failure rate of the changed hard drive and the first value (number of errors - 3) is calculated, and a subtraction operation is performed using the upper limit of the third hard drive failure rate score interval as the minuend and four times the difference between the hard drive failure rate of the changed hard drive and the first value as the subtrahend (number of errors - 3) × 4 to obtain a hard drive failure rate health score for the hard drive to be repaired.
[0056] In response to the target hard disk life scoring condition being the first hard disk life scoring condition, a first hard disk life score interval (18-20) corresponding to the first hard disk life scoring condition is obtained, the hard disk life of the changed hard disk is used as the dividend, and a division operation is performed with the second value as the divisor to obtain a first operation result (number of hours / 2000), and a subtraction operation is performed with the upper limit value of the first hard disk life score interval as the minuend (20) and the first operation result as the subtrahend to obtain the hard disk life health score of the changed hard disk.
[0057] In response to the target hard disk life scoring condition being the second hard disk life scoring condition, a second hard disk life score interval (10-17) corresponding to the second hard disk life scoring condition is obtained, the hard disk life of the changed hard disk is used as the minuend, and the lower limit value of the second hard disk life scoring condition is used as the subtrahend to perform a subtraction operation to obtain a second operation result (number of hours - 20000), the second operation result is used as the dividend, and the third value is used as the divisor to perform a division operation to obtain a third operation result, the upper limit value (17) of the second hard disk life score interval is used as the minuend, and seven times the third operation result is used as the minuend ((number of hours - 20000) / 1000×7) to perform a subtraction operation to obtain the hard disk life health score of the changed hard disk.
[0058] In response to the target hard disk life scoring condition being the third hard disk life scoring condition, a third hard disk life score interval (0-9) corresponding to the third hard disk life scoring condition is obtained, and a subtraction operation is performed with the hard disk life of the changed hard disk as the minuend and the lower limit value of the third hard disk life scoring condition as the subtrahend to obtain a fourth operation result (number of hours - 30000), and a division operation is performed with the fourth operation result as the dividend and the third value as the divisor to obtain a fifth operation result, and a subtraction operation is performed with the upper limit value of the third hard disk life score interval as the minuend and three times the fifth operation result as the minuend ((number of hours - 30000) / 1000×3) to obtain the hard disk life health score of the changed hard disk.
[0059] The health score of the changed hard drive is calculated as the sum of its temperature stability health score, its failure rate health score, and its lifespan health score. The contributions of these three factors to the health score can be divided based on their respective impacts on the health score.
[0060] A hard disk processing rule table is also set up in this application. The baseboard management controller obtains a pre-set hard disk processing rule table, which stores multiple health score ranges, multiple hard disk states, and multiple hard disk processing methods; the baseboard management controller obtains the health score of the changed hard disk, compares the health score of the changed hard disk with the multiple health score ranges in the hard disk processing rule table, and uses the health score range that matches the health score of the changed hard disk as the target health score range.
[0061] The baseboard management controller obtains the target hard disk status corresponding to the target health score range, and the target hard disk status is one of the healthy status, warning status, repair status, and fault status; among them, if the target hard disk status of the changed hard disk is a healthy status, the hard disk processing method is to load the hard disk historical configuration for initialization; if the target hard disk status of the changed hard disk is a warning status, the hard disk processing method is to shorten the hard disk detection cycle; if the target hard disk status of the changed hard disk is a repair status, the hard disk processing method is to execute the repair strategy after initializing the hard disk; if the target hard disk status of the changed hard disk is a fault status, the hard disk processing method is to physically isolate the hard disk; the baseboard management controller updates the list of hard disks to be repaired according to the hard disk status of the changed hard disk and the hard disk processing method.
[0062] Specifically, the hard disk processing rules table is shown in Table 3:
[0063] Table 3
[0064]
[0065] Step 103: In response to the basic input and output system obtaining the list of hard disks to be repaired, the priorities of multiple hard disks to be repaired in the list are determined, and the multiple hard disks to be repaired are initialized according to the priority levels of the multiple hard disks to be repaired and the hard disk processing methods of the multiple hard disks to be repaired to realize server startup.
[0066] Specifically, the basic input and output system determines whether the hard disks to be repaired in the list of hard disks to be repaired include the system disk. If the hard disks to be repaired in the list of hard disks to be repaired include the system disk, the system disk in the list of hard disks to be repaired will be marked with the first priority; the basic input and output system determines whether the hard disks to be repaired in the list of hard disks to be repaired include hard disks to be repaired whose hard disk status is healthy. If the list of hard disks to be repaired includes hard disks to be repaired whose hard disk status is healthy, the hard disks to be repaired whose hard disk status is healthy will be marked with the second priority; the basic input and output system determines whether the hard disks to be repaired in the list of hard disks to be repaired include hard disks to be repaired whose hard disk status is warning status, repair status, or failure status. If the list of hard disks to be repaired includes hard disks to be repaired whose hard disk status is warning status, repair status, or failure status, the hard disks to be repaired whose hard disk status is warning status, repair status, or failure status will be marked with the third priority.
[0067] The BIOS locates the NVME device where the EFI system partition is located based on the GUID in the UEFI boot item and marks it as the first priority. Healthy disks (80-100 points) are marked as the second priority. Hard disks in the warning, repair, and failure states (0-79 points) are marked as the third priority. The priority order is: system disk (first priority) > data disk (second priority) > high-risk disk (third priority).
[0068] The basic input and output system determines the startup order of multiple hard disks to be repaired according to the priorities of the multiple hard disks to be repaired; obtains the hard disk processing methods corresponding to the multiple hard disks to be repaired, and initializes them based on the hard disk processing methods corresponding to the multiple hard disks to be repaired and the startup order of the multiple hard disks to be repaired, thereby improving the server startup efficiency.
[0069] In one embodiment, a multi-level repair strategy is set for a hard disk in a repair state in the present application, and the basic input and output system executes the multi-level repair strategy on the hard disk to be repaired whose target hard disk state is a repair state. The multi-level repair strategy includes a link recovery strategy, a hardware recovery strategy, and an isolation strategy; wherein the link recovery strategy is to adjust the parameters for determining the negotiation rate when the system accesses the hard disk and trigger link retraining; the hardware recovery strategy is to trigger a power-on and power-off reset of the hard disk; and the isolation strategy is to physically isolate the hard disk.
[0070] Specifically, the repair strategy includes three levels. The first level is link-level recovery, which can be link Tx balance value parameter adjustment and link retraining. For example, access the manufacturer-specific configuration space of the NVME device, adjust the Tx balance Preset value to the historical optimal parameter, and set the Retrain Link bit of the PCIe link control register to trigger physical layer retraining; the second level is hardware-level recovery, which can be triggering a power-on and power-off reset of the hard disk; the third level is physical isolation, which can be to completely isolate the hardware fault disk to prevent cascading failures.
[0071] The basic input / output system in this application implements a multi-level repair strategy for a hard drive to be repaired, which is in the repair state. This strategy also includes obtaining the required boot time for the hard drive to be repaired and calculating a repair threshold for the hard drive to be repaired based on the upper limit of the required boot time. In this repair strategy, due to boot time constraints, the number of repair attempts is dynamically calculated based on the boot time. For example, the repair threshold for the hard drive to be repaired can be calculated using the formula: Maximum number of retries (repair threshold) = Tremaining / α, where Tremaining is the remaining standard boot time (unit: milliseconds) and α is an adjustment factor (default value: P0=150, P1=300, P2=600). Dynamically allocating repair attempts based on the required system boot time can improve system reliability.
[0072] In response to the hard disk startup time upper limit value of the startup time requirement of the hard disk to be repaired being less than the preset initialization threshold, the initialization of the hard disk to be repaired that is not the first priority is suspended, and the hard disk to be repaired with the first priority is placed in the fast recovery mode.
[0073] Quick recovery mode is an emergency mode. When the remaining time Tremaining is less than 10% Ttotal, the initialization of non-critical devices (hard disks other than the system disk) is suspended, and quick recovery mode is enabled (skipping the firmware verification step). Only the minimum system disk set (≥1 boot device) is protected. In this way, switching to emergency mode only protects the minimum system disk, ensuring that the boot is completed within the specified time, thereby increasing boot reliability.
[0074] This application uses the baseboard management controller to perform a topological scan of all JBOFs (hard disk expansion devices) and hard disks, obtain the current topological fingerprint, and compare it with the historical topological fingerprint. If the fingerprints are consistent, the historical list of hard disks to be repaired (i.e., the historical topological fingerprint data stored by the BMC before the system was operational) is directly sent to the BIOS for booting. If the fingerprints are inconsistent, the inconsistent hard disks are rescanned to obtain the health status of all disks in the changed area (hard disks with inconsistent fingerprints), and the current list of hard disks to be repaired is updated and synchronized to the basic input and output system. After receiving the list of hard disks to be repaired from the baseboard management controller, the basic input and output system performs differentiated initialization according to the device priority strategy: healthy disks quickly load the historical configuration, and hard disks to be repaired use a three-level repair strategy. Within the repair strategy, the number of fault repairs can be dynamically calculated based on the system boot time limit, and different fault repair levels can be implemented based on the health status score. When the boot time is insufficient, the system switches to emergency mode, protecting only the minimum system disk, ensuring that the boot is completed within the specified time, thereby increasing boot reliability.
[0075] This application implements the coordinated optimization of dynamic PCIe topology perception and hard disk health pre-check during server startup. The BMC constructs a topology fingerprint through physical layer detection and link feature extraction, and compares it with historical records to determine changes in the topology fingerprint. If there is a change, a health scan is triggered, and a multi-dimensional health score is generated based on temperature, error records, and performance parameters, dividing the health status (healthy / warning / suspicious / faulty). After the BIOS receives the list of hard disks to be repaired, it performs differentiated initialization according to device priority (P0 system disk>P1 data disk>P2 hard disk to be repaired): the healthy disk quickly loads the historical configuration, the hard disk to be repaired uses a three-level repair strategy (link retraining, hard disk power on and off, physical isolation), and dynamically allocates the number of repairs based on the system startup time requirement; when the startup time is insufficient, the emergency mode is switched to protect only the minimum system disk, ensuring that the startup is completed within the specified time, thereby increasing the reliability of the startup.
[0076] In a feasible implementation, a TPM module can be configured in a baseboard management controller. After the baseboard management controller obtains the historical topology fingerprint of the hard disk and the current topology fingerprint of the hard disk, the historical topology fingerprint of the hard disk and the current topology fingerprint of the hard disk can be transmitted to a complex programmable logic device. In response to the complex programmable logic device receiving the historical topology fingerprint of the hard disk and the current topology fingerprint of the hard disk, the port mapping list is returned to the baseboard management controller. The baseboard management controller reads the SMART data and expands the PCR (16, SHA-256 topology fingerprint). Here, the historical topology fingerprint of the hard disk and the current topology fingerprint of the hard disk are calculated by SHA-256 respectively, and the current value of the platform configuration register (PCR value) is obtained. A composite fingerprint including the PCR value is generated based on the current value of the platform configuration register, the historical topology fingerprint and the current topology fingerprint of the hard disk, a dedicated signature key is generated in the TPM module, the composite fingerprint value is signed, and a signature data packet is obtained. The baseboard management controller signature data packet is sent to the UEFI environment of the basic input and output system for verification. Specifically, the PCR value in the signature data packet is compared with the PCR value currently obtained from the TPM module to see if they are consistent. If they are consistent, the verification is considered successful. If they are inconsistent, the verification is considered failed and an alarm is issued. In this way, by checking whether the PCR value in the signature data packet is consistent with the PCR value currently obtained from the TPM module, it can be ensured that the topology fingerprint has not been tampered with, which is beneficial to improving the security of the system.
[0077] See also Figure 4 In order to better describe the server startup method provided by this application, a specific example is used for illustration:
[0078] S1, after the server is powered on, the baseboard management controller works in conjunction with the complex programmable logic device through a dedicated bus to poll the status of all JBOF backplanes in place and scan the current hard disk cascade path between the current hard disk expansion device and the hard disk.
[0079] S2, scan all hard disk key information, save it as the hard disk current cascade path and current hard disk key information in the form of PCIe topology, store it in the designated area, and generate the current topology fingerprint.
[0080] S3, the BMC accesses the designated area (the designated area is the area where the historical topology fingerprint is stored) to check whether the historical topology fingerprint exists. If so, step S4 is executed; if not, the process jumps to step S5.
[0081] S4, check whether the current topology fingerprint of the hard disk is consistent with the historical topology fingerprint of the hard disk. If they are consistent, it proves that the PCIe topology has not changed during this startup, and then skip this health status check, directly load the historical health data of the partition, and use the historical list of hard disks to be repaired; if the topology fingerprints are inconsistent, add it to the queue to be scanned, treat the hard disk with inconsistent topology fingerprints as the changed hard disk, and jump to step S5 to perform a health status scan.
[0082] S5: The baseboard management controller scans the health status of the hard disk, obtains the hard disk health score, and updates the list of hard disks to be repaired according to the scoring rules.
[0083] S6: The baseboard management controller sends a list of hard disks to be repaired to the basic input and output system.
[0084] S7, the basic input and output system is guided according to the list of hard disks to be repaired sent by the baseboard management controller, guides loading historical data on healthy disks, and executes retry and repair strategies for hard disks in repair status.
[0085] S8, complete server startup.
[0086] The embodiment of the present application provides a server startup device, which is specifically as follows: Figure 5 As shown, the server startup device includes: an acquisition module 20, a determination module 21 and a startup module 22.
[0087] An acquisition module 20 is configured to, in response to the server being powered on, cause the baseboard management controller to acquire the historical topology fingerprints and the current topology fingerprints of the hard disks, and to select hard disks whose historical topology fingerprints and current topology fingerprints do not match as changed hard disks;
[0088] Determination module 21, configured for the baseboard management controller to scan the changed hard disk, obtain health information of the changed hard disk, determine a health score of the changed hard disk based on the health information of the changed hard disk, obtain a hard disk status and hard disk processing method corresponding to the health score of the changed hard disk, determine a list of hard disks to be repaired based on the hard disk status and hard disk processing method of the changed hard disk, and the baseboard management controller sends the list of hard disks to be repaired to the basic input and output system;
[0089] The startup module 22 is used to obtain a list of hard disks to be repaired in response to the basic input and output system, determine the priorities of multiple hard disks to be repaired in the list of hard disks to be repaired, and initialize the multiple hard disks to be repaired according to the priority levels of the multiple hard disks to be repaired and the hard disk processing methods of the multiple hard disks to be repaired to realize server startup.
[0090] like Figure 6As shown, an embodiment of the present application further provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor is configured to run the computer program to execute the steps in any of the above server startup method embodiments.
[0091] An embodiment of the present application further provides a computer-readable storage medium, in which a computer program is stored. The computer program is configured to execute the steps of any of the above-mentioned server startup method embodiments when running.
[0092] In an exemplary embodiment, the computer-readable storage medium may include, but is not limited to, various media that can store computer programs, such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk, or an optical disk.
[0093] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the above description has generally described the components and steps of each example according to their functions. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians can use different methods to implement the described functions for each specific application, but such implementation should not be judged to be beyond the scope of this application.
[0094] The above describes in detail a server startup method provided by the present application. This document uses specific examples to illustrate the principles and implementation methods of the present application. The description of the above embodiments is intended only to help understand the method and core concept of the present application. It should be noted that, for those skilled in the art, without departing from the principles of the present application, various improvements and modifications may be made to the present application, and such improvements and modifications also fall within the scope of protection of the present application.
Claims
1. A server startup method, characterized in that: The server startup method includes: In response to the server being powered on, the baseboard management controller and the complex programmable logic device collaboratively scan the hard disk expansion device to obtain the current hard disk cascade path between the hard disk expansion device and the hard disk; the baseboard management controller obtains current hard disk key information, the hard disk key information including the communication rate of the system accessing the hard disk, the parameter for determining the negotiation rate when the system accesses the hard disk, and the hard disk number; the current topology fingerprint of the hard disk is determined based on the current hard disk cascade path and the current hard disk key information; the baseboard management controller obtains the historical topology fingerprint of the hard disk, and treats the hard disk whose historical topology fingerprint and the current topology fingerprint do not match as a changed hard disk; The baseboard management controller scans the changed hard disk to obtain health information of the changed hard disk, and determines a health score of the changed hard disk based on the health information of the changed hard disk; obtains a hard disk status and a hard disk processing method of the changed hard disk corresponding to the health score of the changed hard disk, determines a list of hard disks to be repaired based on the hard disk status and the hard disk processing method of the changed hard disk, and sends the list of hard disks to be repaired to the basic input and output system; In response to the basic input and output system obtaining a list of hard disks to be repaired, the priorities of multiple hard disks to be repaired in the list of hard disks to be repaired are determined, and the multiple hard disks to be repaired are initialized according to the priority levels of the multiple hard disks to be repaired and the hard disk processing methods of the hard disks to be repaired to realize server startup.
2. The server startup method according to claim 1, wherein: After the baseboard management controller obtains the historical topology fingerprint of the hard disk, the following steps are performed: The historical topological fingerprint of the hard disk is compared with the current topological fingerprint of the hard disk.
3. The server startup method according to claim 1, wherein: The determining of the current topology fingerprint of the hard disk based on the current cascade path of the hard disk and the current key information of the hard disk includes: Concatenate the current cascade path of the hard disk and the key information of the current hard disk into a character string; Perform hash calculation on the string to obtain the current topological fingerprint value of the hard disk; The current cascade path of the hard disk, the current key information of the hard disk and the current topology fingerprint value of the hard disk are stored in a target storage area of the baseboard management controller in a tree topology structure to generate the current topology fingerprint of the hard disk.
4. The server startup method according to claim 2, characterized in that: The comparing the historical topological fingerprint of the hard disk with the current topological fingerprint of the hard disk includes: In response to a mismatch between the historical topology fingerprint of the hard disk and the current topology fingerprint, the hard disk whose historical topology fingerprint and the current topology fingerprint do not match is used as a changed hard disk, a first mark is performed, and the changed hard disk is added to a queue to be scanned; In response to a match between the historical topology fingerprint of the hard disk and the current topology fingerprint, a second mark is performed on the hard disk that matches the historical topology fingerprint and the current topology fingerprint, and historical health information of the hard disk with the second mark is loaded.
5. The server startup method according to claim 1, wherein: The acquiring health information of the changed hard disk and determining the health score of the changed hard disk based on the health information of the changed hard disk includes: The baseboard management controller obtains health information of the changed hard disk, wherein the health information includes hard disk temperature stability, hard disk failure rate, and hard disk life status; The baseboard management controller obtains target scoring conditions corresponding to the changed hard disk health information, wherein the target scoring conditions include a target temperature stability scoring condition, a target hard disk failure rate scoring condition, and a target hard disk life scoring condition; The baseboard management controller obtains the target temperature stability scoring condition, the target hard disk failure rate scoring condition, and the target hard disk life scoring condition corresponding to the target temperature stability scoring condition, the target hard disk failure rate scoring condition, and the target hard disk life scoring condition respectively; The baseboard management controller calculates the health score of the changed hard disk based on the target temperature stability score calculation scheme, the target hard disk failure rate score calculation scheme, and the target hard disk life score calculation scheme.
6. The server startup method according to claim 5, characterized in that: The target scoring conditions for obtaining and changing the health information of the hard disk include: Obtain a pre-set health scoring rule table, which includes multiple scoring conditions, multiple scoring intervals, and multiple scoring calculation schemes corresponding to hard drive temperature stability, hard drive failure rate, and hard drive life status; Obtaining the hard disk temperature stability of the changed hard disk, comparing the hard disk temperature stability of the changed hard disk with multiple scoring conditions corresponding to hard disk temperature stability in the health scoring rule table, and using the hard disk temperature stability scoring condition that matches the hard disk temperature stability of the changed hard disk as the target temperature stability scoring condition; Obtaining the hard disk failure rate of the changed hard disk, comparing the hard disk failure rate of the changed hard disk with multiple scoring conditions corresponding to the hard disk failure rate in the health scoring rule table, and using the hard disk failure rate scoring condition that matches the hard disk failure rate of the changed hard disk as the target hard disk failure rate scoring condition; Obtain the hard disk life of the changed hard disk, compare the hard disk life of the changed hard disk with multiple scoring conditions corresponding to the hard disk life in the health scoring rule table, and use the hard disk life scoring condition that matches the hard disk life of the changed hard disk as the target hard disk life scoring condition.
7. The server startup method according to claim 5, characterized in that: The baseboard management controller obtains the target temperature stability scoring condition, the target hard disk failure rate scoring condition, and the target hard disk life scoring condition, respectively, corresponding to the target temperature stability scoring condition, the target hard disk failure rate scoring condition, and the target hard disk life scoring condition; the baseboard management controller calculates the health score of the changed hard disk based on the target temperature stability scoring calculation scheme, the target hard disk failure rate scoring calculation scheme, and the target hard disk life scoring calculation scheme, including: Obtain the target temperature stability score range corresponding to the target temperature stability score condition, and calculate the temperature stability health score of the changed hard drive based on the upper limit of the target temperature stability score range, the lower limit of the target temperature stability score condition, and the hard drive temperature stability of the changed hard drive; Obtaining a target hard disk failure rate score range corresponding to the target hard disk failure rate scoring condition, and calculating a hard disk failure rate health score of the changed hard disk based on an upper limit of the target hard disk failure rate score range and the hard disk failure rate of the changed hard disk; Obtaining a target hard disk lifespan score interval corresponding to the target hard disk lifespan score condition, and calculating a hard disk lifespan health score for the changed hard disk based on an upper limit value of the target hard disk lifespan score interval, a lower limit value of the target hard disk lifespan score condition, and the hard disk lifespan of the changed hard disk; The health score of the changed hard disk is calculated as the sum of the temperature stability health score of the changed hard disk, the hard disk failure rate health score of the changed hard disk, and the hard disk life health score of the changed hard disk.
8. The server startup method according to claim 1, wherein: The step of obtaining the hard disk status and hard disk processing method of the changed hard disk corresponding to the health score of the changed hard disk, and determining a list of hard disks to be repaired according to the hard disk status and hard disk processing method of the changed hard disk includes: The baseboard management controller obtains a preset hard disk processing rule table, wherein the hard disk processing rule table stores multiple health score intervals, multiple hard disk states, and multiple hard disk processing methods; The baseboard management controller obtains the health score of the changed hard disk, compares the health score of the changed hard disk with multiple health score intervals in the hard disk processing rule table, and uses the health score interval that matches the health score of the changed hard disk as the target health score interval; The baseboard management controller obtains the target hard disk status and target hard disk processing method corresponding to the target health score range; The target hard disk state is one of a healthy state, a warning state, a repair state, and a fault state. The hard disk processing method corresponding to the healthy state is to load the hard disk historical configuration for initialization. The hard disk processing method corresponding to the warning state is to shorten the hard disk detection cycle. The hard disk processing method corresponding to the repair state is to execute the repair strategy after initialization. The hard disk processing method corresponding to the fault state is to physically isolate the hard disk. The baseboard management controller updates the hard disk list to be repaired according to the target hard disk status and the target hard disk processing method.
9. The server startup method according to claim 1, wherein: In response to the basic input and output system acquiring the hard disk list to be repaired, determining the priorities of multiple hard disks to be repaired in the hard disk list to be repaired includes: The basic input and output system determines whether the hard disks to be repaired in the hard disk list to be repaired include a system disk, and if so, marks the system disk in the hard disk list to be repaired with a first priority; The basic input and output system determines whether the hard disk to be repaired in the hard disk to be repaired list includes a hard disk to be repaired in a healthy state, and if the hard disk to be repaired list includes a hard disk to be repaired in a healthy state, marks the hard disk to be repaired in a healthy state as a second priority; The basic input and output system determines whether the hard disks to be repaired in the hard disk list to be repaired include hard disks to be repaired whose hard disk status is a warning state, a repair state, or a fault state. If the hard disk list to be repaired includes hard disks to be repaired whose hard disk status is a warning state, a repair state, or a fault state, the hard disks to be repaired whose hard disk status is a warning state, a repair state, or a fault state are marked as the third priority.
10. The server startup method according to claim 1, characterized in that: Initializing the multiple hard disks to be repaired according to the priority levels of the multiple hard disks to be repaired and the hard disk processing methods of the hard disks to be repaired to realize server startup includes: The basic input and output system determines the startup order of multiple hard disks to be repaired according to the priorities of the multiple hard disks to be repaired; obtains the hard disk processing methods corresponding to the multiple hard disks to be repaired, and initializes based on the hard disk processing methods corresponding to the multiple hard disks to be repaired and the startup order of the multiple hard disks to be repaired to realize server startup.
11. The server startup method according to claim 8, characterized in that: The repair strategy is a multi-level repair strategy, and the server startup method further includes: The basic input and output system executes a multi-level repair strategy for the hard disk to be repaired whose target hard disk state is a repair state, wherein the multi-level repair strategy includes a link recovery strategy, a hardware recovery strategy, and an isolation strategy; Among them, the link recovery strategy is to adjust the parameters that determine the negotiation rate when the system accesses the hard disk and trigger link retraining; the hardware recovery strategy is to trigger the hard disk power on and off reset; the isolation strategy is to physically isolate the hard disk.
12. The server startup method according to claim 11, characterized in that: The basic input and output system executes a multi-level repair strategy for a hard disk to be repaired whose target hard disk state is a repair state, including: Obtaining a startup time requirement for the hard disk to be repaired, and calculating a repair count threshold for the hard disk to be repaired based on an upper limit of the hard disk startup time required for the startup time of the hard disk to be repaired; In response to the hard disk startup time upper limit value of the startup time requirement of the hard disk to be repaired being less than the preset initialization threshold, the initialization of the non-first priority hard disk to be repaired is suspended, and the first priority hard disk to be repaired is placed in fast recovery mode.
13. A computer device, characterized in that: include: memory for storing computer programs; A processor, configured to implement the steps of the server startup method according to any one of claims 1 to 12 when executing the computer program.
14. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, wherein when the computer program is executed by a processor, the steps of the server startup method according to any one of claims 1 to 12 are implemented.
15. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the server startup method according to any one of claims 1 to 12 are implemented.
Citation Information
Patent Citations
Server NVME hard disk starting method and system and related equipment
CN115756620A
Server startup item starting method and device
CN119917180A