Hard disk alarm method and device of server, server and medium

By introducing complex programmable logic devices into the server to read the hard disk presence signal and status, the problem of insufficient fault identification when the baseboard management controller relies on the RAID card is solved, accurate alarms and rapid location of hard disk failures are achieved, and the accuracy of hard disk fault identification and system reliability are improved.

CN120704931APending Publication Date: 2025-09-26INSPUR (SHANDONG) COMPUTER TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510883904.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-27
Publication Date
2025-09-26

AI Technical Summary

Technical Problem

In the prior art, when a baseboard management controller relies on a redundant array of independent disks card to monitor a hard disk, it is unable to accurately identify a hard disk failure when the hard disk cannot be identified by the RAID card, resulting in insufficient fault identification accuracy.

Method used

By introducing complex programmable logic devices into the server, the hard disk's presence signal and hard disk status are read, and combined with the hard disk's physical layer status, it is determined whether the hard disk is a faulty hard disk, and an alarm message is output when necessary, thereby enhancing the accuracy of fault identification.

Benefits of technology

The accuracy of hard disk fault identification has been improved, and accurate alarms can be issued even when the RAID card cannot monitor the hard disk, quickly locating the faulty hard disk and improving product competitiveness and business continuity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120704931A_ABST
    Figure CN120704931A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of servers, in particular to a hard disk alarm method and device of a server, the server and a medium, the method is applied to a substrate management controller in the server, and the server further comprises a backboard and a plurality of hard disks; a complex programmable logic device is arranged on the back plate; the complex programmable logic device and the backboard are connected with the substrate management controller. The method comprises the following steps: reading a target hard disk in-place signal of the complex programmable logic device, wherein the target hard disk is any one of a plurality of hard disks; reading the hard disk state of the target hard disk; determining whether the target hard disk is a fault hard disk according to the in-place signal and the hard disk state of the target hard disk; if the target hard disk is the fault hard disk, alarm information of the target hard disk is output, and the alarm information is used for fault alarm of the target hard disk; the problem that the substrate management controller can still give an alarm under the condition that the card end of the redundant array of independent disks of the hard disk is not recognized is solved, and the accuracy of determining the faulty hard disk is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of servers, and in particular to a hard disk alarm method, device, server and medium for a server. Background Art

[0002] In the big data era, hard drive health is directly related to data security and business continuity. Real-time monitoring of hard drive health information on a management platform allows for rapid identification of faulty drives, avoiding the delays of manual inspections. Monitoring hard drive health information is particularly crucial for server manufacturers. The commonly used solution in the industry relies on the library interface of the RAID (Redundant Array of Independent Disks) card, managing hard drives in real time through the card. The BMC (Baseboard Management Controller) then calls the RAID card's interface to generate drive alarms. However, if a drive is no longer recognized by the RAID card, the BMC only monitors hard drive information through the card, resulting in inaccurate fault identification.

[0003] It can be seen that how to improve the accuracy of hard disk fault identification is a problem that needs to be solved by those skilled in the art. Summary of the Invention

[0004] The present invention aims to provide a server hard disk alarm method, device, server and medium, which can improve the accuracy of hard disk fault identification.

[0005] In a first aspect, a server hard disk alarm method is provided, which is applied to a baseboard management controller in the server, wherein the server further includes: a backplane and multiple hard disks; a complex programmable logic device is provided on the backplane; the complex programmable logic device and the backplane are both connected to the baseboard management controller;

[0006] The hard disk alarm method of the server includes: reading the target hard disk presence signal of the complex programmable logic device, where the target hard disk is any hard disk among the multiple hard disks; reading the hard disk status of the target hard disk; determining whether the target hard disk is a faulty hard disk based on the presence signal and the hard disk status of the target hard disk; if the target hard disk is a faulty hard disk, outputting alarm information of the target hard disk, where the alarm information is used to issue a fault alarm for the target hard disk.

[0007] In a preferred example, the present invention can be further configured as follows: the server further includes: an independent disk redundant array card; correspondingly, reading the hard disk status of the target hard disk includes: reading the hard disk status of the independent disk redundant array card, and reading the hard disk physical layer status of the target hard disk through out-of-band management; correspondingly, determining whether the target hard disk is a faulty hard disk based on the in-place signal and hard disk status of the target hard disk includes: determining whether the target hard disk is a faulty hard disk based on the in-place signal, hard disk physical layer status and hard disk status of the target hard disk.

[0008] In a preferred example, the present invention can be further configured as follows: determining whether the target hard disk is a faulty hard disk based on the target hard disk's in-place signal, the hard disk physical layer status, and the hard disk status, including: determining whether the target hard disk is physically disconnected based on the target hard disk's hard disk in-place signal; if the target hard disk is physically disconnected, determining that the target hard disk is a faulty hard disk; if the target hard disk is physically connected, determining whether the hard disk physical layer status is in the on state; if the hard disk physical layer status is not in the on state, determining that the target hard disk is a normal hard disk; if the hard disk physical layer status is in the on state, determining whether the hard disk status is that hard disk information exists; if the hard disk status is that hard disk information exists, determining that the target hard disk is a normal hard disk; otherwise, determining that the target hard disk is a faulty hard disk.

[0009] In a preferred example, the present invention can be further configured as follows: determining whether the target hard disk is physically disconnected based on the hard disk presence signal of the target hard disk, including: determining the level change of the presence signal of the target hard disk; if the presence signal changes from the connected level to the unplugged level, determining that the target hard disk is physically disconnected; if the presence signal changes from the maintained connected level, determining that the target hard disk is physically connected.

[0010] In a preferred example, the present invention can be further configured as follows: it also includes: obtaining the current backplane data of the backplane, the current backplane data including: backplane voltage, backplane temperature, and backplane signal quality; based on the backplane data of the previous period, using a health prediction model, predicting the health status of the backplane to obtain first prediction data; if the similarity between the first preset data and the current backplane data reaches a preset similarity threshold, then using a health prediction model based on the current backplane data to predict the health status of the backplane to obtain second prediction data; determining the health status of the target hard disk based on the second prediction data; if the health status is abnormal, automatically switching to the backup backplane through the CPLD.

[0011] In a preferred example, the present invention can be further configured as follows: it also includes: if the target hard disk is not a faulty hard disk and the physical layer status of the hard disk is in the on state, then obtaining the link signal quality of the target hard disk, the link signal quality including signal strength, bit error rate, and number of repetitions; evaluating the link health score of the target hard disk based on the signal strength, bit error rate, and number of repetitions; determining the link alarm information of the target hard disk based on the link health score and the preset level classification mechanism of the target hard disk; wherein the preset level classification mechanism of the target hard disk is set according to the performance of the target hard disk and the criticality of the business.

[0012] In a preferred example, the present invention can be further configured to include: triggering an adaptive reconnection strategy when the link health score is lower than a preset score threshold, wherein the adaptive reconnection strategy is a strategy with the highest reconnection success rate determined from historical reconnection strategies.

[0013] In a second aspect, a hard disk alarm device for a server is provided, comprising:

[0014] a reading module, configured to read a target hard disk presence signal of a complex programmable logic device; and to read a hard disk status of a target hard disk; the target hard disk being any one of a plurality of hard disks;

[0015] a determination module, configured to determine whether the target hard disk is a faulty hard disk according to a presence signal and a hard disk status of the target hard disk;

[0016] The output module is used to output the alarm information of the target hard disk if the target hard disk is a faulty hard disk, and the alarm information is used to issue a fault alarm of the target hard disk.

[0017] In a third aspect, a server is provided, comprising a baseboard management controller, an independent disk redundant array card, a backplane, and multiple hard disks; the backplane is provided with a complex programmable logic device; the independent disk redundant array card is connected to multiple hard disks via the backplane; the complex programmable logic device and the independent disk redundant array card are both connected to the baseboard management controller; the baseboard management controller comprises: a memory and a processor, the memory storing a computer program, and the processor executing the server hard disk alarm method described in any one of the first aspects when running the computer program.

[0018] In a fourth aspect, a computer-readable storage medium is provided, wherein at least one program code is stored in the computer-readable storage medium, and the program code is loaded and executed by a processor to implement the hard disk alarm method of the server as described in any one of the first aspects.

[0019] In a fifth aspect, a computer program product is provided, comprising a computer program or instructions, which, when executed by a processor, implements the hard disk alarm method for a server as described in any one of the first aspects.

[0020] In summary, the server hard disk alarm method provided by the present invention has the following beneficial technical effects:

[0021] The server hard disk alarm method provided by this technical solution is applied to the baseboard management controller in the server. The server also includes: a backplane, multiple hard disks; a complex programmable logic device is provided on the backplane; the complex programmable logic device and the backplane are both connected to the baseboard management controller; the server hard disk alarm method includes: reading the target hard disk presence signal of the complex programmable logic device, where the target hard disk is any of the multiple hard disks; reading the hard disk status of the target hard disk; determining whether the target hard disk is a faulty hard disk based on the presence signal and the hard disk status of the target hard disk; if the target hard disk is a faulty hard disk, outputting the target hard disk alarm information, which is used to generate a fault alarm for the target hard disk. This solves the problem that the baseboard management controller has long relied too much on the independent disk redundant array card to realize the hard disk alarm function. When the independent disk redundant array card can no longer realize the monitoring function in existing scenarios, the baseboard management controller can no longer realize the alarm function. The baseboard management controller obtains the target hard disk presence signal of the complex programmable logic device; reads the hard disk status of the target hard disk; and then combines the above key information to determine whether the target hard disk is a faulty hard disk. This solves the problem that the baseboard management controller can still issue an alarm even when the hard disk is not recognized on the independent disk redundant array card side, improves the accuracy of faulty hard disk identification, and also facilitates rapid on-site positioning, greatly increasing the competitiveness of the product and greatly improving business.

[0022] In addition, the present invention also provides a hard disk alarm device for a server, a server, and a medium, all of which have the above-mentioned beneficial technical effects. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] In order to more clearly illustrate the embodiments of the present invention, the following is a brief introduction to the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0024] Figure 1 This is a flow chart of a server hard disk alarm method provided by an embodiment of the present invention;

[0025] Figure 2 This is a schematic diagram of the structure of a server provided by an embodiment of the present invention;

[0026] Figure 3is a schematic diagram of an alarm situation provided by an embodiment of the present invention;

[0027] Figure 4 is a schematic diagram of another alarm situation provided by an embodiment of the present invention;

[0028] Figure 5 This is a structural diagram of a hard disk alarm device for a server provided by an embodiment of the present invention;

[0029] Figure 6 It is a structural diagram of a part of the structure of a server provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0030] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making any creative efforts shall fall within the scope of protection of the present invention.

[0031] The terms "including" and "having," as used in the present description and accompanying drawings, and any variations thereof, are intended to cover non-exclusive inclusions. For example, a process, method, system, product, or apparatus comprising a series of steps or elements is not limited to the listed steps or elements and may include steps or elements that are not listed.

[0032] In order to enable those skilled in the art to better understand the present invention, the present invention will be further described in detail below with reference to the accompanying drawings and specific implementation methods.

[0033] Figure 1 A flow chart of a server hard disk alarm method provided by an embodiment of the present invention is applied to a baseboard management controller in the server. The server also includes: a backplane and multiple hard disks; a complex programmable logic device is provided on the backplane; the complex programmable logic device and the backplane are both connected to the baseboard management controller;

[0034] Server hard disk alarm methods include:

[0035] S101, reading a target hard disk presence signal of a complex programmable logic device;

[0036] The target hard disk is any hard disk among multiple hard disks;

[0037] The hard drive's presence signal (PRESENT#) records the drive's presence status. This signal is implemented via a physical GPIO pin on the hard drive backplane or slot (which contacts the hard drive connector). When a hard drive is connected, this pin is grounded, and the PRESENT# signal can be low. When a hard drive is removed, the PRESENT# signal can be high. The CPLD monitors the slot's PRESENT# signal in real time. When a hard drive is inserted or removed, the CPLD converts the physical signal into a stable logic level to generate a presence signal. This signal is then transmitted to the BMC via I2C. In other words, the CPLD monitors the hard drive slot's PRESENT# signal, and when the level changes from low to high, the BMC polls the CPLD. When the level changes from low to high, the BMC issues a hard drive alarm.

[0038] It is understandable that a change from low to high in the hard disk slot PRESENT# signal indicates a fault, and the physical connection is considered to be disconnected. If the hard disk slot PRESENT# signal is always at a high level, it may be that no hard disk is inserted, and it is not considered to be a physical connection disconnection, and no alarm information is generated. If the hard disk slot PRESENT# signal changes from high to low, it may be that the hard disk was manually inserted. Therefore, there is a change from high to low, and there is no subsequent change from low to high, and it is considered that the physical connection is not disconnected.

[0039] In this step, the presence signal of the target hard disk of the complex programmable logic device is read to determine the power-on status of the hard disk.

[0040] S102, reading the hard disk status of the target hard disk;

[0041] S103, determining whether the target hard disk is a faulty hard disk based on the target hard disk presence signal and hard disk status;

[0042] In one possible scenario, the monitoring server is not connected to the topology configuration of the CPU directly without the RAID card. At this time, the status of the target hard disk is read through out-of-band management. Since the RAID card is not connected, the physical layer status (phy status) of the hard disk is all in the on state. Accordingly, it is only necessary to determine whether the target hard disk is a faulty hard disk based on the hard disk presence signal and the hard disk status. Specifically, based on the hard disk presence signal of the target hard disk, determine whether the target hard disk is physically disconnected; if the target hard disk is physically disconnected, determine that the target hard disk is a faulty hard disk; if the target hard disk is physically connected, determine whether the hard disk status is that hard disk information exists; if the hard disk status is that hard disk information exists, determine that the target hard disk is a normal hard disk; otherwise, determine that the target hard disk is a faulty hard disk. Among them, the hard disk information refers to the current working status of the hard disk, that is, whether the hard disk is in a normal working state.

[0043] In another possible scenario, monitor the server RAID card, see Figure 2 , Figure 2 This is a structural diagram of a server provided by an embodiment of the present invention. The server includes: a baseboard management controller, a backplane, multiple hard disks, and an independent disk redundant array card; a complex programmable logic device is provided on the backplane; and the complex programmable logic device and the backplane are both connected to the baseboard management controller. The present invention adopts an out-of-band management method to obtain the hard disk status. Out-of-band management (OOB) is a hardware management method independent of the host operating system. Even if the host is down or the operating system crashes, the hardware status can still be accessed. The PHY status is independently detected by the PHY controller of the RAID card and does not rely on the protocol response of the hard disk. Correspondingly, reading the hard disk status of the target hard disk includes: reading the hard disk status of the independent disk redundant array card, and reading the hard disk physical layer status of the target hard disk through out-of-band management; correspondingly, determining whether the target hard disk is a faulty hard disk based on the target hard disk's in-place signal and hard disk status includes: determining whether the target hard disk is a faulty hard disk based on the target hard disk's in-place signal, hard disk physical layer status, and hard disk status. Determine whether the target hard disk is a faulty hard disk based on the target hard disk's presence signal, the hard disk physical layer status, and the hard disk status, including: determining whether the target hard disk is physically disconnected based on the target hard disk's hard disk presence signal; if the target hard disk is physically disconnected, determining that the target hard disk is a faulty hard disk; if the target hard disk is physically connected, determining whether the hard disk physical layer status is in the on state; if the hard disk physical layer status is not in the on state, determining that the target hard disk is a normal hard disk; if the hard disk physical layer status is in the on state, determining whether the hard disk status indicates that hard disk information exists; if the hard disk status indicates that hard disk information exists, determining that the target hard disk is a normal hard disk; otherwise, determining that the target hard disk is a faulty hard disk.

[0044] It is understandable that there are two reasons why hard drives are not recognized by RAID cards: 1) because of unstable links and direct physical connection disconnection; 2) because of hard drive failure that leads to data communication failure. The first scenario: unstable links and physical connection disconnection. When the hard drive encounters a physical connection disconnection, that is, the physical connection between the backplane and the hard drive is disconnected, it will be accompanied by a PRESENT# signal level conversion. At this time, the BMC can directly access the CPLD through i2c, and can monitor the hard drive unrecognition situation, thereby generating an alarm. For details, see Figure 3. The second scenario: the hard disk fails in data communication due to a fault in the hard disk itself. In this case, the hard disk presence signal can no longer be used to determine the condition of the hard disk, because the physical link is normal at this time, but it only affects the transmission of the data signal. At this time, the BMC can determine whether the disk is normal and whether an alarm is needed by judging the hard disk PRESENT# signal, phy status and hard disk status. The BMC can obtain the hard disk presence information from the CPLD end through i2c. At the same time, the BMC can obtain the phy status and hard disk status through the Out-Of-band (out-of-band management) of the RAID card, and comprehensively judge whether to report the hard disk alarm. When the BMC obtains a low level of hard disk PRESENT#, the phy status is enable, but the hard disk status is no hard disk information. At this time, the BMC alarms. For details, see Figure 4 .

[0045] S104: If the target hard disk is a faulty hard disk, output alarm information of the target hard disk, where the alarm information is used to generate a fault alarm of the target hard disk.

[0046] Alarm information is used to alert users or administrators of hard drive failures, allowing them to promptly understand the abnormal state of the hard drive and take appropriate measures to address it. Through the above process, the present invention can detect failures on multiple hard drives. If the target hard drive is not a faulty hard drive, no action is taken.

[0047] Furthermore, a message queue service can be established. When a hard disk failure is detected, the hard disk information can be published as a message in the message queue. After the alarm processing program subscribes to the message queue, the alarm processing program will obtain the information of the failed hard disk from the message queue, and then notify the corresponding management personnel through email notification / SMS notification / interface push, so that the management personnel can deal with the failed hard disk in time and reduce the impact on the business.

[0048] It can be seen that in the embodiments of the present invention, the problem of the baseboard management controller being overly dependent on the independent disk redundant array card to implement the hard disk alarm function for a long time has been solved. When the independent disk redundant array card can no longer realize the monitoring function in existing scenarios, the baseboard management controller can no longer realize the alarm function. The baseboard management controller obtains the target hard disk presence signal of the complex programmable logic device; reads the hard disk status of the target hard disk; and then determines whether the target hard disk is a faulty hard disk based on the above key information. This solves the problem of the baseboard management controller still being able to realize the alarm when the hard disk is not recognized on the independent disk redundant array card side, improves the accuracy of the faulty hard disk identification, and also facilitates rapid on-site positioning, greatly increasing the competitiveness of the product and greatly improving the business.

[0049] A possible implementation of an embodiment of the present invention is to determine whether the target hard disk is physically disconnected based on the hard disk presence signal of the target hard disk, including: determining the level change of the target hard disk presence signal; if the presence signal changes from the connected level to the unplugged level, determining that the target hard disk is physically disconnected; if the presence signal changes from the maintained connected level, determining that the target hard disk is physically connected.

[0050] In this embodiment of the present invention, a disconnection is only indicated when the target hard drive's presence signal changes from the connected level to the unplugged level. If it remains at the unplugged level, it is not considered an abnormal disconnection but rather a normal sign of the hard drive not being inserted. Only when the presence signal remains at the connected level is the physical connection considered intact. Determining physical disconnection through this level change is simple and efficient.

[0051] A possible implementation method of an embodiment of the present invention also includes: obtaining current backplane data of the backplane, the current backplane data including: backplane voltage, backplane temperature, and backplane signal quality; predicting the health status of the backplane based on the backplane data of the previous period using a health prediction model to obtain first prediction data; if the similarity between the first preset data and the current backplane data reaches a preset similarity threshold, predicting the health status of the backplane based on the current backplane data using a health prediction model to obtain second prediction data; determining the health status of the target hard disk based on the second prediction data; if the health status is abnormal, automatically switching to the backup backplane through the CPLD.

[0052] In an embodiment of the present invention, the I2C protocol is used to exchange data with the backplane, thereby obtaining the current backplane data, including key information such as backplane voltage, backplane temperature, and backplane signal quality. The backplane voltage reflects the stability of the backplane power supply; the backplane temperature reflects the heat generated by the backplane during operation. Excessive temperature may affect device performance or even cause failure. Low-power sensors (such as voltage sensors, temperature sensors, and signal integrity detection chips) are integrated on the backplane PCB; the sensor data is transmitted to the BMC via I2C or SMBus, where the backplane signal quality is specifically determined by monitoring the signal eye diagram quality.

[0053] To achieve accurate predictions, a health prediction model is used to predict the backplane's health status based on the previous period's backplane data, generating first prediction data. The health prediction model is trained on a large amount of historical backplane data and can analyze correlations and patterns between data to predict the backplane's health status. The previous period's backplane data contains various operating parameters of the backplane over a period of time. This data is input into the health prediction model, which performs a preliminary assessment of the backplane's current health status and outputs first prediction data. If the similarity between the first preset data and the current backplane data reaches a preset similarity threshold, it indicates that the current backplane data and the preset data are similar in characteristics and are likely in similar operating states. The health prediction model's prediction is accurate and the current prediction can be executed. Furthermore, the health prediction model is used to predict the backplane's health status based on the current backplane data, generating second prediction data. The second prediction data reflects the future health status of the backplane under the current circumstances. Each parameter can then be compared with the corresponding threshold to determine the health status of the target hard drive. For example, unstable backplane voltage, excessively high temperature, or low signal quality may affect the target hard drive's normal power supply and operating environment, causing the target hard drive to malfunction. Therefore, based on the second prediction data, it is possible to accurately predict whether the target hard drive will have future health issues. If the health status is abnormal, the CPLD automatically switches to the backup backplane. The backup backplane is similar in design and configuration to the primary backplane, and can quickly take over in the event of a primary backplane failure, ensuring continuous system operation.

[0054] It can be seen that in the embodiment of the present invention, the similarity calculation is performed between the predicted data of the historical period and the current data. Only when the similarity meets the standard, it indicates that the prediction ability is accurate, and then the current data is substituted into the model to estimate the future data. Then, the future health status is predicted based on the future data, so that when there is a health problem, the CPLD can be used in time to automatically switch to the backup backplane to ensure the continuous operation of the system.

[0055] A possible implementation method of an embodiment of the present invention also includes: if the target hard disk is not a faulty hard disk and the physical layer status of the hard disk is in the on state, then obtaining the link signal quality of the target hard disk, the link signal quality including signal strength, bit error rate, and number of repetitions; evaluating the link health score of the target hard disk based on the signal strength, bit error rate, and number of repetitions; determining the link alarm information of the target hard disk based on the link health score and the preset level classification mechanism of the target hard disk; wherein the preset level classification mechanism of the target hard disk is set according to the performance of the target hard disk and the criticality of the business.

[0056] Specifically, if the target hard drive is a healthy hard drive and its physical layer status is on, meaning it is functioning normally, the link signal quality of the target hard drive is constantly monitored to obtain signal strength, bit error rate, and number of repetitions. A weighted calculation is then performed to obtain a link health score, for example, using LHS = (Wsignal × Ssignal) + (Wber × Sber) + (Wretry × Sretry), where LHS is the link health score, Wsignal, Wber, and Wretry are the weights of signal strength, bit error rate, and number of repetitions (for example, 40%, 30%, and 30%, respectively), and Ssignal, Sber, and Sretry are the standardized scores of signal strength, bit error rate, and number of repetitions, respectively.

[0057] For signal strength normalization: The ideal range is -10dBm to -30dBm (the stronger the signal, the higher the score). Normalization formula: ; Example: When the signal strength is -25dBm, Ssignal=100; when the signal strength is -35dBm, Ssignal=50 (linear interpolation).

[0058] For bit error rate normalization: ideal range: BER <1e-12 (the lower the bit error rate, the higher the score). Normalization formula: .

[0059] For retry count normalization: Ideal range: Retry Count ≤ 5 (the fewer the retry count, the higher the score). Normalization formula: .

[0060] Then, the corresponding level is set according to the performance of the target hard disk and the criticality of the business.

[0061] If the target hard disk is a high-performance hard disk and / or a mission-critical hard disk, the first preset grading mechanism is adopted; otherwise, the second preset grading mechanism is adopted.

[0062] For example, the first preset grading mechanism is as follows: Excellent: LHS ≥ 85; Good: 70 ≤ LHS < 85; Warning: 55 ≤ LHS < 70; Dangerous: LHS < 55. The second preset grading mechanism is as follows: Excellent: LHS ≥ 80; Good: 65 ≤ LHS < 80; Warning: 50 ≤ LHS < 65; Dangerous: LHS < 50. Each grading mechanism corresponds to a standard template for link alarm information. After the grading mechanism is determined, link alarm information is obtained based on the template and hard disk information.

[0063] It can be seen that in the embodiment of the present invention, the link signal quality of the target hard disk that is not a faulty hard disk and whose physical layer status is open is obtained, including key indicators such as signal strength, bit error rate, and number of repetitions; combining these three indicators, the link health score of the target hard disk can be comprehensively evaluated; a preset level classification mechanism is set according to the hard disk performance and business criticality; according to the preset level classification mechanism of the target hard disk, the link health score is mapped to specific link alarm information; so that the administrator can quickly understand the actual status of the link and take corresponding maintenance measures.

[0064] A possible implementation of the embodiment of the present invention further includes: when the link health score is lower than a preset score threshold, triggering an adaptive reconnection strategy, where the adaptive reconnection strategy is a strategy with the highest reconnection success rate determined from historical reconnection strategies.

[0065] When the link health score falls below the preset threshold, an adaptive reconnection strategy is automatically triggered. Reconnection strategies include, but are not limited to: Strategy 1: Increase signal transmission power (e.g., from -20dBm to -15dBm). Strategy 2: Switch communication protocols (e.g., from SATA 6Gbps to 3Gbps). Strategy 3: Reset the PHY layer connection between the RAID card and the hard drive.

[0066] In an embodiment of the present invention, each time an adaptive reconnection strategy is triggered, the success rate, reconnection time, reconnection result, etc. of the reconnection strategy are recorded; then, the strategy with the highest reconnection success rate is screened out from the historical reconnection strategies. If the current reconnection based on the highest success strategy reaches a preset number of times, such as 3 times, and is still unsuccessful, the strategy with the second success rate is selected for reconnection until it succeeds, and the reconnection information of that time is recorded. At this time, the information of multiple historical reconnection strategies is updated. Through continuous optimization, each time a reconnection is required, it is based on the latest data, which makes the effect more excellent. Of course, the strategy with the highest success rate can be selected based on the data of the reconnection strategy within a certain period of time. The embodiment of the present invention is no longer limited, and the user can set it according to actual needs.

[0067] In an embodiment of the present invention, when the link health score is lower than the preset score threshold, an adaptive reconnection strategy is automatically triggered. The design of this strategy selects the strategy with the highest reconnection success rate based on the success rate of historical reconnection strategies, which significantly improves the recovery capability of the server hard disk link under abnormal conditions.

[0068] Figure 5 A schematic structural diagram of a server hard disk alarm device provided in an embodiment of the present invention includes:

[0069] The reading module 210 is used to read the target hard disk presence signal of the complex programmable logic device; and read the hard disk status of the target hard disk; the target hard disk is any hard disk among the multiple hard disks;

[0070] The determination module 220 is configured to determine whether the target hard disk is a faulty hard disk based on the target hard disk presence signal and the hard disk status;

[0071] The output module 230 is configured to output alarm information of the target hard disk if the target hard disk is a faulty hard disk. The alarm information is used to generate a fault alarm of the target hard disk.

[0072] In a preferred example, the present invention can be further configured as follows: the server further includes: an independent disk redundant array card; correspondingly, a reading module 210 is used to: read the hard disk status of the independent disk redundant array card, and read the hard disk physical layer status of the target hard disk through out-of-band management; correspondingly, a determination module 220 is used to: determine whether the target hard disk is a faulty hard disk based on the target hard disk's presence signal, hard disk physical layer status, and hard disk status.

[0073] In a preferred example, the present invention can be further configured as: a determination module 220, used to: determine whether the target hard disk is physically disconnected based on the hard disk presence signal of the target hard disk; if the target hard disk is physically disconnected, determine that the target hard disk is a faulty hard disk; if the target hard disk is physically connected, determine whether the hard disk physical layer status is in the on state; if the hard disk physical layer status is not in the on state, determine that the target hard disk is a normal hard disk; if the hard disk physical layer status is in the on state, determine whether the hard disk status is that hard disk information exists; if the hard disk status is that hard disk information exists, determine that the target hard disk is a normal hard disk; otherwise, determine that the target hard disk is a faulty hard disk.

[0074] In a preferred example, the present invention can be further configured as: a determination module 220, used to: determine the level change of the in-place signal of the target hard disk; if the in-place signal changes from the connected level to the unplugged level, it is determined that the physical connection of the target hard disk is disconnected; if the in-place signal changes from the maintained connected level, it is determined that the physical connection of the target hard disk is not disconnected.

[0075] In a preferred example, the present invention can be further configured as follows: it also includes: an acquisition module, used to obtain the current backplane data of the backplane, the current backplane data including: backplane voltage, backplane temperature, and backplane signal quality; a prediction module, used to predict the health status of the backplane based on the backplane data of the previous period using a health prediction model, and obtain first prediction data; if the similarity between the first preset data and the current backplane data reaches a preset similarity threshold, the health status of the backplane is predicted based on the current backplane data using a health prediction model, and obtain second prediction data; a health status determination module, used to determine the health status of the target hard disk based on the second prediction data; a switching module, used to automatically switch to the backup backplane through the CPLD if the health status is abnormal.

[0076] In a preferred example, the present invention can be further configured as follows: it also includes: a signal quality alarm module, which is used to: if the target hard disk is not a faulty hard disk and the physical layer status of the hard disk is on, then obtain the link signal quality of the target hard disk, the link signal quality includes signal strength, bit error rate, and number of repetitions; based on the signal strength, bit error rate, and number of repetitions, evaluate the link health score of the target hard disk; based on the link health score and the preset level classification mechanism of the target hard disk, determine the link alarm information of the target hard disk; wherein, the preset level classification mechanism of the target hard disk is set according to the performance of the target hard disk and the criticality of the business.

[0077] In a preferred example, the present invention can be further configured to include: a reconnection module, which is used to trigger an adaptive reconnection strategy when the link health score is lower than a preset score threshold. The adaptive reconnection strategy is a strategy with the highest reconnection success rate determined from historical reconnection strategies.

[0078] Figure 5 The description of the features in the corresponding embodiment can be found in Figure 1 The relevant descriptions of the corresponding embodiments will not be repeated here one by one.

[0079] Figure 6 A structural diagram of a server provided in an embodiment of the present invention, such as Figure 6 As shown, a server includes a baseboard management controller, an independent disk redundant array card, a backplane, and multiple hard disks; the backplane is provided with a complex programmable logic device; the independent disk redundant array card is connected to the multiple hard disks via the backplane; the complex programmable logic device and the independent disk redundant array card are both connected to the baseboard management controller; the baseboard management controller includes: a memory 60 for storing computer programs;

[0080] The processor 61 is configured to implement the steps of the hard disk alarm method of the server in the above embodiment when executing a computer program.

[0081] The memory 60 may include one or more computer-readable storage media, which may be non-transitory. The memory 60 may also include a high-speed random access memory, and a non-volatile memory, such as one or more disk storage devices, flash memory storage devices. In this embodiment, the memory 60 is at least used to store the following computer program 601, wherein, after the computer program is loaded and executed by the processor 61, it can implement the relevant steps of the hard disk alarm method of the server disclosed in any of the aforementioned embodiments. In addition, the resources stored in the memory 60 may also include an operating system 602 and data 603, etc., and the storage method may be temporary storage or permanent storage. Among them, the operating system 602 may include Windows, Unix, Linux, etc.

[0082] In some embodiments, the server may further include a display screen 62 , an input / output interface 63 , a communication interface 64 , a power supply 65 , and a communication bus 66 .

[0083] Those skilled in the art will understand that Figure 6 The structure shown in the figure does not constitute a limitation to the server, and may include more or fewer components than shown in the figure.

[0084] It is understandable that if the server hard disk alarm method in the above embodiment is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the current technology, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and executes all or part of the steps of the various embodiments of the present invention. The aforementioned storage medium includes: USB flash drives, mobile hard drives, read-only memories (ROM), random access memories (RAM), electrically erasable programmable ROMs, registers, hard drives, removable disks, CD-ROMs, magnetic disks, or optical disks, and other media that can store program code.

[0085] Based on this, an embodiment of the present invention further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the above method are implemented.

[0086] Based on this, an embodiment of the present invention further provides a computer program product, including a computer program / instruction, which implements the steps of the above method when executed by a processor.

[0087] The above describes in detail the server hard disk alarm method, device, server, and medium provided by the embodiments of the present invention. The various embodiments are described in a progressive manner throughout this specification, with each embodiment focusing on its differences from other embodiments. Similar or identical parts between the various embodiments can be referenced for reference only. The device disclosed in the embodiments corresponds to the method disclosed in the embodiments, so the description is relatively brief. For relevant details, refer to the method description.

[0088] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the above description has generally described the components and steps of each example according to their functions. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the present invention.

[0089] The above describes in detail the server hard disk alarm method, device, server, and medium provided by the present invention. This article uses specific examples to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only intended to help understand the method and core concept of the present invention. It should be noted that for those skilled in the art, various improvements and modifications can be made to the present invention without departing from the principles of the present invention, and such improvements and modifications also fall within the scope of protection of the present invention.

Claims

1. A server hard disk alarm method, characterized in that: A baseboard management controller used in a server, the server further comprising: a backplane and a plurality of hard disks; a complex programmable logic device is provided on the backplane; the complex programmable logic device and the backplane are both connected to the baseboard management controller; The server hard disk alarm method includes: Reading a target hard disk presence signal of the complex programmable logic device, where the target hard disk is any one of the multiple hard disks; Read the hard disk status of the target hard disk; Determining whether the target hard disk is a faulty hard disk according to the in-position signal and hard disk status of the target hard disk; If the target hard disk is a faulty hard disk, alarm information of the target hard disk is output, and the alarm information is used to issue a fault alarm of the target hard disk.

2. The server hard disk alarm method according to claim 1, characterized in that: The server also includes: an independent disk redundant array card, Correspondingly, read the hard disk status of the target hard disk, including: Reading the hard disk status of the redundant array of independent disks card, and reading the hard disk physical layer status of the target hard disk through out-of-band management; Correspondingly, determining whether the target hard disk is a faulty hard disk according to the in-place signal and hard disk status of the target hard disk includes: Determine whether the target hard disk is a faulty hard disk according to a presence signal, a hard disk physical layer status, and a hard disk status of the target hard disk.

3. The server hard disk alarm method according to claim 2, characterized in that: Determining whether the target hard disk is a faulty hard disk according to a presence signal, a physical layer status of the hard disk, and a hard disk status of the target hard disk includes: Determining whether the target hard disk is physically disconnected according to the hard disk presence information of the target hard disk; If the target hard disk is physically disconnected, determining that the target hard disk is a faulty hard disk; If the target hard disk is physically connected, determining whether the physical layer state of the hard disk is open; If the physical layer state of the hard disk is not an open state, determining that the target hard disk is a normal hard disk; If the hard disk physical layer state is an open state, determining whether the hard disk state is hard disk information existence; If the hard disk status indicates that hard disk information exists, the target hard disk is determined to be a normal hard disk; otherwise, the target hard disk is determined to be a faulty hard disk.

4. The server hard disk alarm method according to claim 3, characterized in that: Determining whether the target hard disk is physically disconnected according to the hard disk presence information of the target hard disk includes: Determining a level change of a presence signal of the target hard disk; If the presence signal changes from a connected level to a disconnected level, it is determined that the target hard disk is physically disconnected; If the presence signal remains at the access level, it is determined that the physical connection of the target hard disk is not disconnected.

5. The server hard disk alarm method according to claim 1, characterized in that: Also includes: Acquire current backplane data of the backplane, wherein the current backplane data includes: backplane voltage, backplane temperature, and backplane signal quality; Based on the backplane data of the previous period, the health status of the backplane is predicted using a health prediction model to obtain first prediction data; If the similarity between the first preset data and the current backplane data reaches a preset similarity threshold, predicting the health status of the backplane using a health prediction model based on the current backplane data to obtain second prediction data; determining a health status of the target hard disk according to the second prediction data; If the health status is abnormal, the system automatically switches to the backup backplane via the CPLD.

6. The server hard disk alarm method according to any one of claims 1 to 5, characterized in that: Also includes: If the target hard disk is not a faulty hard disk and the physical layer status of the hard disk is on, then obtaining the link signal quality of the target hard disk, where the link signal quality includes signal strength, bit error rate, and number of repetitions; Evaluate the link health score of the target hard disk based on the signal strength, bit error rate, and number of repetitions; Determining link alarm information of the target hard disk based on the link health score and a preset grading mechanism of the target hard disk; The preset classification mechanism of the target hard disk is set according to the performance of the target hard disk and the criticality of the business.

7. The server hard disk alarm method according to claim 6, characterized in that: Also includes: When the link health score is lower than a preset score threshold, an adaptive reconnection strategy is triggered, where the adaptive reconnection strategy is a strategy with the highest reconnection success rate determined from historical reconnection strategies.

8. A server hard disk alarm device, characterized in that: include: A reading module, used to read the target hard disk presence signal of the complex programmable logic device; and, reading the hard disk status of the target hard disk; The target hard disk is any hard disk among multiple hard disks; a determination module, configured to determine whether the target hard disk is a faulty hard disk according to a presence signal and a hard disk status of the target hard disk; The output module is used to output the alarm information of the target hard disk if the target hard disk is a faulty hard disk, and the alarm information is used to issue a fault alarm of the target hard disk.

9. A server, characterized in that: include: Baseboard management controller, independent disk redundant array card, backplane, multiple hard disks; The backplane is provided with a complex programmable logic device; The independent disk redundant array card is connected to a plurality of hard disks via the backplane; the complex programmable logic device and the independent disk redundant array card are both connected to the baseboard management controller; The baseboard management controller includes: a memory for storing a computer program; A processor is used to execute the computer program to implement the steps of the hard disk alarm method of the server as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the hard disk alarm method of the server according to any one of claims 1 to 7 are implemented.

Citation Information

Cited By

  • Hard disk information management system and hard disk information management method

    CN121879690A