Equipment firmware upgrading method and device, storage medium and electronic equipment
By evaluating the environment and performance status before the hard disk firmware upgrade and delaying the upgrade until the conditions are appropriate, the problem of low success rate of hard disk firmware upgrade is solved and the stability of the storage system is improved.
Patent Information
- Application Number
- CN202511020083.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-23
- Publication Date
- 2025-08-26
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing hard disk firmware upgrade methods may cause data consistency problems when reading and writing large-scale data in hard disks or busy states, or even damaging firmware, affecting the stability and reliability of the storage system.
Before sending the firmware upgrade command, receive the environmental health information and performance status information of the external device to determine whether the preset status conditions are met; if it is not met, record the current time point and delay the upgrade, wait for the delay cycle to re-evaluate the status before performing firmware upgrade.
The success rate of firmware upgrades is improved under unsuitable environmental conditions, and the stability of the storage system is enhanced, and the upgrade failure and potential hardware damage are avoided.
Smart Images

Figure CN120540679A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of device upgrades, and in particular to a method and apparatus for upgrading device firmware, a storage medium, and an electronic device. Background Art
[0002] In today's data centers and enterprise-class server environments, hard drives are critical components supporting the storage and rapid access of massive amounts of data. As the core of its internal logic control, hard drive firmware plays a crucial role in optimizing read and write speeds, improving data processing efficiency, enhancing error recovery capabilities, and adapting to new storage technologies. Therefore, regular hard drive firmware upgrades are essential to ensure the continued efficient operation of storage systems.
[0003] In the existing firmware upgrade method, the Small Computer System Interface (SCSI) protocol instructions are directly called to update the hard disk firmware. However, if the hard disk is performing large-scale data reading and writing or is in a busy state, forcibly performing a firmware upgrade may cause data consistency issues due to internal processing interruptions in the hard disk, or even damage the firmware, resulting in the hard disk being unable to boot normally. If the upgrade fails, not only may the stored data be lost, but complex troubleshooting and repair work may also be required. In severe cases, the hard disk may even need to be replaced, which in turn affects the overall stability and reliability of the storage system. In other words, the hard disk firmware upgrade method in the related art suffers from the problem of low storage system stability. Summary of the Invention
[0004] The present application provides a device firmware upgrade method and apparatus, a storage medium, and an electronic device to at least solve the problem of low stability of the storage system in the hard disk firmware upgrade method in the related art.
[0005] The present application provides a device firmware upgrade method, comprising: sending a firmware upgrade instruction to an external device, and receiving first status data sent by the external device, wherein the first status data is used to indicate environmental health information and performance status information of the external device;
[0006] If the first state data does not meet the preset state condition, the current time point is recorded as the first moment;
[0007] At a second moment that is separated from the first moment by a delay period, receiving second status data sent by the external device;
[0008] When the second status data meets the preset status condition, the device firmware of the external device is upgraded.
[0009] The present application also provides a device firmware upgrade apparatus, comprising: a first data receiving module, configured to send a firmware upgrade instruction to an external device and receive first status data sent by the external device, wherein the first status data is used to indicate environmental health information and performance status information of the external device;
[0010] A time recording module, configured to record the current time point as the first moment when the first state data does not satisfy a preset state condition;
[0011] A second data receiving module is configured to receive second status data sent by an external device at a second moment that is spaced apart from the first moment by a delay period;
[0012] The firmware upgrade module is used to upgrade the device firmware of the external device when the second state data meets the preset state condition.
[0013] The present application also provides an electronic device, comprising: a memory for storing a computer program; and a processor for implementing the steps of any of the above-mentioned device firmware upgrade methods when executing the computer program.
[0014] The present application also provides a computer-readable storage medium, in which a computer program is stored. When the computer program is executed by a processor, the steps of any of the above-mentioned device firmware upgrade methods are implemented.
[0015] The present application also provides a computer program product, including a computer program, which implements the steps of any of the above-mentioned device firmware upgrade methods when executed by a processor.
[0016] Through the present application, a firmware upgrade instruction is sent to an external device, and first status data sent by the external device is received, wherein the first status data is used to indicate the environmental health information and performance status information of the external device. If the first status data does not meet the preset status conditions, the current time point is recorded as the first moment. At a second moment that is separated from the first moment by a delay period, second status data sent by the external device is received. If the second status data meets the preset status conditions, the device firmware of the external device is upgraded. When upgrading the external device, the current status data of the external device can be read. If the status data does not meet the preset status conditions, the upgrade process is aborted and the upgrade time is delayed. After the delay period, the latest status data of the external device is reread and re-determined whether it meets the preset status conditions. If it does, the device firmware upgrade of the external device is started. In this way, firmware upgrades can be avoided under environmental conditions that are likely to cause upgrade failures. Therefore, the technical problem of the low success rate of current hard disk firmware upgrade methods can be solved, achieving the technical effect of improving the stability of the storage system. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] In order to more clearly illustrate the embodiments of the present application, the following is a brief introduction to the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0018] Figure 1 is a schematic diagram of a hardware environment of an optional device firmware upgrade method according to an embodiment of the present application;
[0019] Figure 2 is a flowchart of an optional device firmware upgrade method according to an embodiment of the present application;
[0020] Figure 3 is a schematic diagram of optional temperature log data according to an embodiment of the present application;
[0021] Figure 4 is a schematic diagram of another optional temperature log data according to an embodiment of the present application;
[0022] Figure 5 is a schematic diagram of optional voltage log data according to an embodiment of the present application;
[0023] Figure 6 is a schematic diagram of optional performance log data according to an embodiment of the present application;
[0024] Figure 7 is a schematic diagram of an optional device firmware upgrade method according to an embodiment of the present application;
[0025] Figure 8 is a schematic diagram of another optional device firmware upgrade method according to an embodiment of the present application;
[0026] Figure 9 This is a structural block diagram of an optional device firmware upgrade apparatus according to an embodiment of the present application. DETAILED DESCRIPTION
[0027] The following will be combined with the accompanying drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0028] It should be noted that, in the description of this application, the terms "comprises," "includes," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. The terms "first," "second," etc., in this application are used to distinguish similar objects, and are not used to describe a particular order or sequence.
[0029] In order to enable those skilled in the art to better understand the present application, the present application is further described in detail below with reference to the accompanying drawings and specific implementation methods.
[0030] According to one aspect of the embodiments of the present application, a method for upgrading device firmware is provided. As an optional implementation, the method for upgrading device firmware can be applied to, but is not limited to, Figure 1 The device firmware upgrade system in the hardware environment shown in FIG. The device firmware upgrade system may include but is not limited to the terminal device 102, the network 110, the server 112, the database 114, and the external device 118. The terminal device 102 runs a target client (such as Figure 1 (As shown, the target client is a client capable of performing device firmware upgrades.) The terminal device 102 includes a display 108, a processor 106, and a memory 104. The display 108 can be used to display device information, etc., and also provides a human-computer interaction interface to receive user input on the interface and touch controls. The processor is configured to generate interaction instructions in response to these human-computer interaction operations and send the interaction instructions to the server. The memory is configured to store firmware files.
[0031] In addition, the server 112 includes a processing engine 116 , which is configured to perform a storage or read operation on the database 114 . Specifically, the processing engine 116 reads a firmware file from the database 114 .
[0032] Assumptions Figure 1 The terminal device 102 in the embodiment runs a client for performing firmware upgrades. The specific process of this embodiment is as follows: In step S102, the terminal device 102 sends the firmware file to the server 112 via the network 110. The user or system administrator must first upload the new version of the firmware file to the storage system. This is usually done in a secure and controlled directory on the system to ensure that only verified firmware files can be used.
[0033] Server 112 then executes step S104 to send a firmware upgrade instruction to external device 118. External device 118 then executes step S106 to send first status data to server 112. Server 112 then executes step S108 to record the current time point if the first status data does not meet the preset status condition. External device 118 then executes step S110 to send second status data to server 112. Server 112 then executes steps S112-S114 to receive the second status data and, if the second status data meets the preset status condition, upgrades the device firmware of the external device.
[0034] Optionally, in this embodiment, the terminal device 102 may be a terminal device configured with a target client, including but not limited to at least one of the following: a mobile phone (such as an Android phone or iOS phone), a laptop, a tablet computer, a PDA, a MID (Mobile Internet Device), a PAD, a desktop computer, a smart TV, etc. The target client may be a client that supports firmware upgrades. The network may include but is not limited to a wired network or a wireless network. The wired network may include a local area network, a metropolitan area network, and a wide area network, and the wireless network may include Bluetooth, Wi-Fi, and other networks that enable wireless communication. The server may be a single server, a server cluster consisting of multiple servers, or a cloud server. The above is merely an example and is not intended to be limiting in this embodiment.
[0035] Optionally, in a computing and network environment, server 112 generally refers to a central computing node responsible for coordinating and running various services, applications, or processes. In the firmware upgrade scenario, the execution subject server 112 refers to a server or storage system that has sufficient authority and capabilities to manage and implement the firmware upgrade process. Its main responsibilities include, but are not limited to: receiving, verifying, storing, and distributing firmware update files. Initiate the firmware upgrade process, monitor the upgrade status, handle exceptions during the upgrade process, and roll back when necessary. Before and during the upgrade, check the health status, environmental conditions, and load conditions of the device to ensure that the upgrade is performed at the best time. Communicate with the hardware device and send necessary control instructions, such as silencing the hard disk, issuing the WRITE BUFFER instruction, activating the new firmware, and releasing the silent state. Record all important data and operations during the upgrade process to facilitate subsequent auditing and troubleshooting.
[0036] Server 112 is typically equipped with a high-performance processor, ample memory, and a specialized storage management system to ensure efficient and stable upgrade operations. Furthermore, it is equipped with interfaces necessary for communicating with storage devices, such as Serial Attached SCSI (SAS), Serial Advanced Technology Attachment (SATA), and Non-Volatile Memory Express (NVMe), as well as a software stack capable of performing complex I / O control and status monitoring.
[0037] External devices 118 refer to various hardware components connected to servers or storage arrays via external interfaces (such as SAS, USB, and Thunderbolt). Among external devices, hard disk drives (HDDs) and modern solid-state drives (SSDs) are of particular interest, serving as the primary carriers of data storage. These devices offer non-volatile storage for persistent data storage and support corresponding communication protocols (such as SCSI, SAS, and SATA) for data exchange and status communication with the host system. Many modern hard disks have firmware upgrade capabilities, allowing for remote or on-site upgrades via host software to fix bugs, enhance performance, or add new features. The ability to return device status information, such as temperature, voltage, and load conditions, via protocols such as SCSI is crucial for the host to make intelligent decisions.
[0038] During the firmware upgrade process, external devices such as hard drives must comply with server instructions, such as suspending I / O operations, receiving new firmware data, and activating the new firmware. Furthermore, the device itself must possess a certain level of intelligence to process received upgrade instructions and protect itself when necessary, avoiding upgrade operations under adverse conditions to ensure data security and device stability.
[0039] The embodiment of the present application provides a device firmware upgrade method for a server, Figure 2 Flowchart of an optional device firmware upgrade method according to an embodiment of the present application; Figure 2 As shown, the device firmware upgrade method includes:
[0040] Step S202: sending a firmware upgrade instruction to the external device and receiving first status data sent by the external device, wherein the first status data is used to indicate environmental health information and performance status information of the external device;
[0041] It should be noted that firmware upgrade instructions can be used in storage system management to update the firmware of external devices (such as hard drives). The system sends specific instructions to the target device, instructing it to prepare to receive new firmware and perform the upgrade process. These instructions are usually based on the communication protocol supported by the device, such as the instructions in the SCSI protocol.
[0042] The first status data refers to the information that the external device returns to the system after receiving the upgrade instruction to reflect its current health and performance status. These data include but are not limited to key indicators such as temperature, voltage, read and write latency, which are used to assess whether the device is in an environmental condition suitable for the upgrade. Environmental health information involves the physical environment when the device is running, such as temperature and voltage levels. Excessive temperature or unstable voltage may cause device performance degradation or even hardware damage. Therefore, it is necessary to ensure that these conditions are within a safe range before upgrading the firmware. Performance status information refers to the current working status and performance indicators of the device, such as read and write latency, I / O operation frequency, etc. It is not suitable to upgrade the firmware when the device is under high load or has large latency, otherwise the risk of upgrade failure may increase.
[0043] During the firmware upgrade process, the storage system or server sends a firmware upgrade instruction to the external device and receives the first status data returned by the device. This data is used to evaluate the environmental health information and performance status information of the device to ensure that the upgrade operation is performed when the device is in a stable and appropriate state.
[0044] In an optional implementation, after confirming the correctness of the firmware file and preparing to begin the upgrade, the system sends an upgrade instruction to the target external device. Simultaneously, the system proactively collects device status information, known as first status data. This data includes the device's current environmental health indicators (such as temperature and voltage) and performance status (such as read and write latency and load). The storage system analyzes this status data based on preset thresholds to determine whether the device is at the optimal time for an upgrade. If both environmental health and performance status meet the requirements, the system allows the firmware upgrade to proceed. Conversely, if any status data exceeds a safe range, the system automatically delays the upgrade and issues an alert, preventing the risks associated with upgrading the firmware under adverse conditions.
[0045] Step S204: if the first state data does not meet the preset state condition, record the current time point as the first moment;
[0046] It should be noted that the preset status conditions can be a set of system-defined conditions before a firmware upgrade to determine whether the current time is optimal. If the current status data exceeds the set range of these conditions, the firmware upgrade is considered unsuitable. The first moment refers to the time point recorded by the system when the status data does not meet the preset conditions. This time point is used for subsequent logical judgments, such as starting a timer to wait for conditions to improve before retrying the upgrade.
[0047] When the system detects that the current status of the hard drive (such as temperature, voltage, or read / write latency) is not suitable for an immediate firmware upgrade, the system automatically records this time point and marks it as the "first moment". Subsequent logic will use this time point to decide when to retry the firmware upgrade operation.
[0048] Intelligent decision-making during the firmware upgrade process is key to ensuring upgrade success and mitigating business risks. When the system checks environmental conditions before an upgrade (for example, by using the LOG SENSE command to query temperature, voltage, and read / write latency), if the initial status data does not meet pre-defined conditions (e.g., excessive temperature, unstable voltage, or excessive hard drive load), the system will not immediately perform the firmware upgrade. Instead, it will record the current time as the "first moment," indicating that the system considers the current environment unsuitable for the upgrade.
[0049] Step S206, receiving second status data sent by the external device at a second moment that is separated from the first moment by a delay period;
[0050] It should be noted that the delay period can be a time interval set by the system to recheck pre-upgrade conditions to ensure that environmental conditions are suitable before executing the firmware upgrade. The second moment refers to the time point after the first moment, after the delay period, at which the system will recheck the status of the external device to ensure that the conditions still meet the upgrade requirements. The second status data refers to the latest status information sent by the external device to the execution server at the second moment to reconfirm whether the firmware upgrade conditions are met.
[0051] In an optional embodiment, after the initial or previous check (the first moment), if environmental conditions (e.g., temperature, voltage, load) are found to be unsuitable for an immediate firmware upgrade, the system will automatically delay the upgrade and start a timer for a set delay period (e.g., 30 minutes). At the end of the delay period (the second moment), the system will receive the latest status data from the external device and reassess whether the conditions required for the firmware upgrade are met.
[0052] In an optional embodiment, the delay period may be determined by a timer, which is an electronic circuit or software function module used to trigger a specific event or operation after a specified time interval.
[0053] It's important to note that timing mechanisms can be implemented at the hardware level, commonly found in chips like microcontrollers (MCUs) and digital signal processors (DSPs). Hard timers typically offer high accuracy and reliability, operating independently of the processor and ensuring timely triggering even when the processor is busy. Timing mechanisms can also be implemented in software. Soft timers are widely used in operating systems, applications, and scripting languages. They rely on the underlying hardware clock and the operating system's time management features. While their accuracy may be affected by system load compared to hard timers, they are generally sufficient for most applications.
[0054] In an optional embodiment, when using a timer, a preset value (i.e., delay period) and an interrupt handler (i.e., the action to be executed after the timer expires) can be set for the timer. The timer is then started at the first moment, and the countdown begins. The timer begins counting down from the preset value until the count reaches zero. When the count reaches zero, the timer triggers a preset event or interrupt, executing the associated action, such as checking the hard drive status, sending a network packet, or updating the screen display. After the event is processed, the timer can be reset (restarting) or stopped, and then reactivated as needed.
[0055] Timers can be used to periodically monitor device status (such as temperature, voltage, and load) to determine if conditions are suitable for an upgrade. For example, hard drive status can be automatically checked every 30 minutes. If the system detects that the current environment or status is not suitable for an immediate upgrade, a timer can be set to delay the upgrade attempt, while continuously monitoring device status changes.
[0056] During the firmware upgrade process, the timer can also be used to monitor the progress of the operation. Once it detects that the upgrade is stalled for more than a preset time, it can trigger a timeout response, such as terminating the upgrade, saving the current status, or issuing an alarm.
[0057] Step S208 : When the second status data satisfies a preset status condition, the device firmware of the external device is upgraded.
[0058] It should be noted that the second status data represents the latest status information obtained by the system from the external device at a certain point in time (the second moment) after the delay period. This information is used to reassess whether the device's health and performance meet the ideal conditions for the firmware upgrade. The preset status conditions are a set of environmental and performance standards set by the system that the external device must meet before the firmware upgrade can proceed. These conditions typically include safe thresholds for parameters such as temperature, voltage, and load.
[0059] In an optional embodiment, after a series of environmental health information and performance status information checks, if the initial check (generating first status data) indicates that the device is in a state unsuitable for upgrade, the system will record the current time point as the first moment and initiate a delay period (e.g., 30 minutes). After the delay period, the system will re-initiate the status check at the second moment and collect the second status data. At this time, if all collected data indicates that the device status has returned to a safe range within the preset status conditions, the system will determine that it is the best time to upgrade. The system will cancel the delay state and begin the specific firmware upgrade process, including but not limited to silencing the hard disk, writing the new firmware using the WRITE BUFFER command, activating the new firmware, and other steps, ultimately completing the device firmware update.
[0060] Example 1:
[0061] Assume that during the initial test, the hard drive temperature is too high (65°C, which is 80% higher than the reference temperature of 60°C, or 48°C). The system records the current time as the first moment and starts a 30-minute delay. During this time, the heat dissipation system around the hard drive operates, gradually reducing the temperature.
[0062] 30 minutes later, at the second moment, the system automatically resends the LOG SENSE command (Page Code: 0x0D, Subpage Code: 0x00) to the hard drive, requesting the latest temperature status. The second status data returned by the hard drive indicates that the current temperature has dropped to 45°C, which is 80% lower than the reference temperature of 60°C, or the calculated value of 48°C. Furthermore, the voltage and read / write latency are within the preset safety range.
[0063] Given this, the storage system determines that all conditions have been met and immediately executes the firmware upgrade process without delaying it. The system first quiesces the hard drive and then begins using the WRITE BUFFER command, reading 512KB of data from the firmware file at a time and writing it to the hard drive in a loop until the entire firmware file has been transferred. After writing the firmware, the system uses the TEST UNIT READY command to check the hard drive's status. Once the upgrade is successful, the system updates the hard drive's firmware version number and unquiesces, allowing the hard drive to resume normal data read and write operations.
[0064] Through the present application, a firmware upgrade instruction is sent to an external device, and first status data sent by the external device is received, wherein the first status data is used to indicate the environmental health information and performance status information of the external device. If the first status data does not meet the preset status conditions, the current time point is recorded as the first moment. At a second moment that is separated from the first moment by a delay period, second status data sent by the external device is received. If the second status data meets the preset status conditions, the device firmware of the external device is upgraded. When upgrading the external device, the current status data of the external device can be read. If the status data does not meet the preset status conditions, the upgrade process is aborted and the upgrade time is delayed. After the delay period, the latest status data of the external device is reread and re-determined whether it meets the preset status conditions. If it does, the device firmware upgrade of the external device is started. In this way, firmware upgrades can be avoided under environmental conditions that are likely to cause upgrade failures. Therefore, the technical problem of the low success rate of current hard disk firmware upgrade methods can be solved, achieving the technical effect of improving the stability of the storage system.
[0065] In an optional implementation, it is necessary to receive status data sent by an external device (such as a hard drive). Sending hard drive status information is an important part of storage system monitoring and maintenance. This information is typically based on SCSI (Small Computer System Interface) or similar standard communication protocols, allowing the hard drive to exchange status information with the storage system or other devices. Here are some common methods for sending hard drive status information, especially for SCSI hard drives:
[0066] The SCSI INQUIRY command is used to obtain basic identification information and the current status of a hard drive, including device type, version, and rotational speed. The storage system sends an INQUIRY command to a hard drive, and the hard drive responds with a response packet containing its ID information and status flags.
[0067] The SCSI READ / REPORT LUNS command reports the status of logical units (LUNs) on a hard drive. This is crucial for hard drives or storage arrays with multiple LUNs. The storage system sends a READ / REPORT LUNS command, and the hard drive returns a response containing the status of each LUN.
[0068] The SCSI LOG SENSE command is used to read detailed log and sensor information from the drive, such as temperature (Page Code: 0x0D), voltage (Page Code: 0x38), and performance statistics (Page Code: 0x37). The system requests this information by sending a specific LOG SENSE command, and the drive responds with a predefined data format, providing insight into the drive's internal environment and operating conditions.
[0069] SMART attributes are a self-monitoring technology built into hard drives that reports on their health and predicts potential failures. The storage system can query SMART attributes, such as error rate, number of starts, and seek errors, by sending the SMART READ DATA or SMART READ THRESHOLD commands. The hard drive will then provide the current values and thresholds for these attributes.
[0070] In addition, some modern hard drives support active event notifications, which proactively send alerts or status updates to the system when specific events occur (such as excessive temperature or signs of hard drive failure). This is usually achieved by configuring the hard drive's event notification configuration, allowing the hard drive to send a signal to the system via a dedicated event notification channel of the SCSI or SATA protocol when an anomaly is detected.
[0071] In an optional embodiment, before receiving the first status data sent by an external device, it includes: sending at least one query instruction to the external device, wherein the query instruction is used to indicate a type of status sub-data; receiving the first status data sent by the external device, wherein the first status data includes status sub-data corresponding to each of the at least one query instruction.
[0072] It should be noted that query instructions can be specific commands sent by the storage system to external devices (such as hard drives) during the firmware upgrade process to request different types of status sub-data, such as temperature, voltage, or performance indicators. These instructions are based on the SCSI protocol or other related interface specifications, such as the LOG SENSE command.
[0073] Status sub-data can be information returned by a query command about a specific aspect of the device, such as the current temperature or voltage. Status sub-data typically includes key parameters used to assess the device's suitability for firmware upgrades.
[0074] In an alternative implementation, a server or storage system first sends a series of query commands, each querying a specific device status parameter. For example, one command might request temperature, while another might query voltage. These commands are designed to collect different types of status information so that the system can comprehensively assess whether the device is in a suitable condition for a firmware upgrade.
[0075] The storage system or server not only needs to know the overall status data of the device, but also needs to have a detailed understanding of the environmental health information and performance status of the device. Therefore, before officially receiving the first status data, the system will send multiple query instructions to the external device in advance. Each instruction focuses on obtaining a certain category of status sub-data, such as temperature, voltage, or read and write delays. After receiving these query instructions, the device will measure or retrieve its own status parameters in real time and return the results to the storage system in the form of status sub-data. After the system collects all the status sub-data, it will integrate them into a complete first status data set, and then analyze and judge it according to the preset firmware upgrade conditions to decide whether the firmware upgrade process should be started.
[0076] It's important to note that in addition to temperature, voltage, and latency, the SCSI protocol and its extensions, such as SAS (Serial Attached SCSI) and SATA (Serial ATA), also provide numerous log pages and commands for querying and monitoring various hard drive information. This information provides a more comprehensive understanding of the hard drive's health and performance, and is invaluable for making decisions about firmware upgrades and other maintenance activities. For example:
[0077] SMART data (Self-Monitoring Analysis and Reporting Technology): SMART data contains a wealth of information about the drive's health, such as error rate, number of boots, operating hours, firmware version, pre-read error rate, and seek error rate. SMART monitoring is an important component of drive health testing, providing a more comprehensive health assessment before firmware upgrades.
[0078] Error Log: Use the `READ LOG EXTENDED` or `LOG SENSE` command to query the hard drive's error log, understand the type and number of recent errors, and determine if there are any potential problems with the hard drive that may affect the stability of the firmware upgrade.
[0079] Performance statistics: In addition to latency, you can use the `LOG SENSE` command to query more performance-related data, such as IOPS (input and output operations per second), throughput, and read-write ratio. This data can help determine whether the hard drive is under continuous high load and determine the best time to upgrade.
[0080] Device Status: Query the online / offline status of the hard drive, whether it is in idle or busy mode, and the hard drive's current task queue to determine whether the hard drive is suitable for interrupting for firmware upgrade.
[0081] Firmware update capability: Use the `INQUIRY` or `READ FIRMWARE REV` command to learn whether the hard drive supports firmware updates and its limitations, such as whether it can be upgraded online or whether there is a specific upgrade mode.
[0082] Hardware wear: For SSDs, you can evaluate the SSD lifespan and remaining capacity by querying the wear leveling information (WearLevelingCount) in `LOG SENSE` or the number of erase / write cycles of the NAND flash memory block. This helps determine whether a firmware upgrade is not worthwhile because the hardware is nearing the end of its lifespan.
[0083] Power status: In addition to voltage, you can also check the hard drive's power management status, such as whether it is in energy-saving mode, the number of power cycles, and the continuous power supply time to ensure that the power status is stable during the upgrade process.
[0084] Cooling fan status: For some hard drives with built-in cooling fans, you can query the fan speed and status to ensure that the hard drive cooling system is operating normally.
[0085] Through the above-described implementation of this application, by sending different types of query commands, device status information is collected item by item, and ultimately a decision is made based on this information whether to initiate a firmware upgrade. This approach ensures that the firmware upgrade operation is performed based on a full understanding and assessment of the device's current status, increasing the success rate and security of the upgrade process.
[0086] In an optional embodiment, receiving first status data sent by an external device includes: receiving at least one log data sent by the external device, wherein the log data includes status data and the log data corresponds one-to-one to the query instruction; reading header information of at least one log data, and determining the data type corresponding to each log data based on the header information; determining at least one status sub-data from at least one log data based on the data type corresponding to each log data, wherein the current status sub-data is determined from the current log data based on the data type of the current log data.
[0087] It should be noted that log data can be generated by external devices and sent to the system, containing detailed information about the device's operating status and performance indicators. Log data is categorized by type (such as temperature, voltage, and performance), with each type corresponding to a specific LOG SENSE command in the SCSI protocol. Header information is the portion of the log data that identifies the data type, format, and length. By reading this header information, the storage system can determine the type of log data being received and how to correctly parse it.
[0088] Status sub-data is specific status information parsed from log data, such as current temperature, voltage level, or read / write latency. Status sub-data is a key indicator used by the storage system to assess whether an external device is suitable for firmware upgrades.
[0089] In an optional embodiment, the system receives at least one log data item containing environmental health information and performance status information about the device. The system then reads the header information of each log data item to determine the specific data type, ensuring that each received data item can be correctly parsed and understood. Finally, based on the data type indicated by the header information, the system extracts and determines at least one status sub-data item from the log data item for subsequent firmware upgrade decisions.
[0090] Each log data is a record of a specific aspect of the status of an external device, corresponding to the query command issued previously, and each command requests a specific status information. After the system receives the log data, the first step is to parse the header information of the log data. The header information contains the type identifier of the data, which enables the system to identify whether the current data is about temperature, voltage or performance status. After the identification is completed, the system will extract the exact status sub-data from the log data based on the data type, such as the current temperature value, voltage level or the latest read and write delay. These status sub-data will then be used to determine whether the device meets the preset conditions for firmware upgrades and are important input information in the entire decision-making process.
[0091] Example 2:
[0092] In preparation for a firmware upgrade for a data center's SAS hard drives, the storage system first issues a series of LOG SENSE query commands, each requesting a different type of log data. Specifically, these include:
[0093] 1. Temperature log data query: The system sends a LOG SENSE command (Page Code: 0x0D, Subpage Code: 0x00) to request the hard disk to provide current temperature information.
[0094] 2. Voltage history data query: Use the LOG SENSE command (Page Code: 0x38, Subpage Code: 0x00) to request the hard disk to return its current voltage level and voltage fluctuation history.
[0095] 3. Performance statistics log query: Use the LOG SENSE command (Page Code: 0x37, Subpage Code: 0x05) to obtain the hard disk's read and write latency statistics for the past five minutes.
[0096] In response to these queries, the hard drive sends back log data containing relevant status information. Upon receiving this data, the storage system first reads the header information from the first line of each log data to determine the data type and length. For example, for temperature data, the system reads the first few bytes of the log data and identifies the Page Code as 0x0D and the Subpage Code as 0x00, thus confirming that this is temperature log data.
[0097] After identifying the data type, the system extracts the status sub-data from the corresponding log data based on the Page Code and Subpage Code in the header information. For example, it reads Byte[9] from the temperature log data to obtain the current temperature, reads Byte[4] and Byte[5] from the voltage history log data to obtain the current voltage and maximum voltage, and reads Byte
[132] to Byte
[139] and Byte
[156] to Byte
[163] from the performance statistics log data to obtain the read and write latency.
[0098] The system compares the extracted status sub-data (current temperature, voltage, read / write latency, etc.) with the preset upgrade conditions. For example, if the current temperature is less than 80% of the device's reference temperature, the voltage is no more than 80% of the maximum operating voltage, and the read / write response time is less than 500ms, the system determines that the device is in a suitable state for a firmware upgrade.
[0099] Through the above-mentioned implementation of the present application, through this series of queries and data processing, the storage system can not only actively obtain real-time status information about the external device, but also intelligently determine whether the device is in an environment where firmware upgrades can be performed safely, thereby effectively avoiding the risks and problems that may arise from upgrading under adverse conditions, and improving the success rate of firmware upgrades and the stability of the device.
[0100] In an optional embodiment, determining at least one status sub-data from at least one log data according to the data types corresponding to the respective log data includes: when the data type of the log data is temperature data, reading indication information from a first preset location of the log data, wherein the indication information is used to indicate the type of the temperature data stored in the first preset location; when the indication information is the first indication information, determining the current temperature data from the first preset location, and determining the reference temperature data from the second preset location;
[0101] When the first state data does not meet the preset state condition, recording the current time point as the first moment includes: when the ratio of the current temperature data to the reference temperature data is greater than the first ratio, recording the current time as the first moment.
[0102] It should be noted that the first preset location can be a specific location in the storage device log for storing current temperature data, which can be defined and follow the provisions of the SCSI protocol. The indication information can be information in the log data that is used to identify the type of data stored at the specific location. For example, a specific parameter code can indicate that the data stored at the first preset location is current temperature data. The first indication information indicates the identification of the current temperature data in the log data, such as the parameter code 0000h. The second preset location can be a location in the log for storing reference temperature data, which is different from the location of the current temperature data and can be customized.
[0103] The current temperature data is the actual temperature value measured by the storage device and is used to assess whether the device is within suitable temperature conditions. The reference temperature data is the recommended upper temperature limit or normal operating temperature range, which is used to compare with the current temperature data to determine whether a firmware upgrade is suitable. The first ratio can be a preset temperature safety threshold, used to determine whether the ratio of the current temperature data to the reference temperature data is too high, such as 80%.
[0104] In an optional embodiment, during the status check phase before a firmware upgrade, the system needs to accurately extract status sub-data related to temperature, voltage, performance, and other factors from the log data. First, the storage system retrieves the log data using the LOG SENSE command. Then, based on the data type, the system locates the first preset location in the log that stores temperature data and identifies the data type by reading the indication information (parameter code) at that location. If the indication information is confirmed to be the first indication information (parameter code 0000h), the system reads the current temperature data at that location and reads reference temperature data from a second preset location for comparison.
[0105] If the ratio of the current temperature to the reference temperature exceeds a preset first value (e.g., 80%), this indicates that the drive's temperature may be at or near an unsafe threshold, potentially causing the firmware upgrade to fail or damaging the drive. In this case, the system records the current time as the "first moment" and initiates a delay mechanism to prevent the upgrade from proceeding under unsuitable temperature conditions.
[0106] In an optional implementation, the query instruction LOG SENSE (opareation code=0x4D) may be:
[0107] -Page Code:0x0D(Temperature);
[0108] -Subpage Code:0x00;
[0109] -Allocation Length: 255 bytes;
[0110] Operation Code = 0x4D is the hexadecimal operation code of the `LOG SENSE` command in the SCSI protocol, representing a specific instruction set used to request log information from an external device (such as a hard disk).
[0111] Page Code: 0x0D is the page code, which specifies the type of log requested. In this example, `0x0D` is a code that points to the temperature log page, instructing the device to return temperature-related information. The SCSI standard defines several different log pages, each corresponding to a specific data log type.
[0112] Subpage Code: 0x00 is the subpage code, which is used to further refine the data type within the log page. In the temperature log page, `0x00` generally means requesting data for the entire page without any specific subpage information. If the device supports multiple subpages for the temperature log, using different subpage codes can obtain more detailed or specific temperature data.
[0113] Allocation Length: 255 This parameter specifies the maximum length of log data the host expects to receive from the device. When sending the `LOG SENSE` command, the host informs the device of the amount of data it can accept. For example, `255 bytes` means the host provides a buffer of 255 bytes to receive temperature log data. The device will fill the buffer with data to this length. If the actual amount of data is less than the allocated length, the remaining space may be filled with zeros or not filled at all.
[0114] Figure 3 This is a schematic diagram of an optional temperature log data according to an embodiment of the present application; the format of the temperature type log data can be as follows Figure 3 As shown, bits 0 to 5 of byte 0 can indicate the page code, which can be 0Dh corresponding to the previous instruction. Byte 1 can indicate the subpage code, which can be 00h. Bytes 2-3 can indicate the valid data length. Bytes 4 to n can be used to indicate temperature log parameters (temperature data). The temperature log parameters in bytes 4 to n can be determined by querying other log formats.
[0115] Figure 4 is a schematic diagram of another optional temperature log data according to an embodiment of the present application; Figure 4 As shown in (a) in the figure, the parsing format of the current temperature data is given, such as Figure 4 As shown in (b) in the figure, the analytical format of the reference temperature data is given. Figure 4As shown in (a), bytes 0-1 can indicate parameter codes (indication information), and byte 5 indicates current temperature data. Figure 4 As shown in (b), bytes 0-1 can indicate the parameter code, and byte 5 indicates the reference temperature data.
[0116] Example 3: Specific process of obtaining hard disk temperature information through SCSI instructions:
[0117] The server or storage controller constructs a LOG SENSE instruction that contains the following key parameters:
[0118] Opcode: Use the LOG SENSE instruction in the SCSI instruction set, whose Opcode is 0x4D.
[0119] Page Code: Set to 0x0D, which means requesting the temperature-related log page.
[0120] Subpage Code: Set to 0x00 to obtain basic temperature information.
[0121] Allocation Length: Defined as the size of the receive buffer, which must be large enough to accommodate the returned temperature data.
[0122] The constructed LOG SENSE command is sent to the target hard disk device (external device) through the SCSI interface.
[0123] After receiving the command, the hard disk device returns log data containing temperature information.
[0124] The server or storage controller first confirms that the Opcode and Page Code of the returned data are consistent with the sent command to ensure a correct response. At the beginning of Byte [4] of the data, look for the Parameter Code. Since the focus is on temperature information, the target Parameter Code should be 0000h (current temperature) and 0001h (reference temperature).
[0125] If the Parameter Code obtained by parsing Byte[4] to Byte[5] is 0000h (the first indication information), the current temperature information is read from Byte[9] to Byte
[10] . Temperature information is usually two bytes of data and requires appropriate conversion. If the Parameter Code obtained by parsing Byte
[10] to Byte
[11] is 0001h, the reference temperature data is read from Byte
[15] to Byte
[16] . The reference temperature is also two bytes of data and is used for comparison with the current temperature.
[0126] Temperature data is typically stored in binary format and needs to be converted to a readable temperature value. This conversion method depends on the specific device's temperature data encoding. Sometimes, you might need to convert the byte data to an integer and then convert the integer to Celsius or Fahrenheit, depending on the device's specifications.
[0127] The current temperature is compared with the reference temperature to determine whether the hard drive is within suitable temperature conditions for the firmware upgrade. If the current temperature exceeds the reference temperature by a certain percentage (e.g., 80%), the environment is considered unsuitable for the upgrade and the upgrade operation needs to be delayed.
[0128] If the current temperature is below the reference temperature threshold, the system can safely perform the firmware upgrade. If the current temperature is too high, the system should delay the upgrade and may need to send an alert to the administrator to indicate that the firmware upgrade is not suitable at this time.
[0129] Through the above-described implementation of the present application, the real-time temperature information of the hard drive can be obtained through the SCSI LOG SENSE instruction. During implementation, it is necessary to ensure that the instruction is correctly constructed and the response data is parsed to accurately obtain the current temperature and reference temperature, and make firmware upgrade decisions based on this information. The parsing process involves identifying the parameter code at a specific byte position, reading the temperature data from the specified byte, and performing conversion and comparison. Accurately obtain temperature information through a preset format, and by detecting temperature information, improve the success rate of firmware upgrades.
[0130] In an optional embodiment, determining at least one state sub-data from at least one log data according to the data type corresponding to each log data includes: when the data type of the log data is voltage data, reading the data at the third preset position of the log data as the current voltage; and reading the data at the fourth preset position of the log data as the reference voltage;
[0131] When the first state data does not meet the preset state condition, recording the current time point as the first moment includes: when the ratio of the current voltage to the reference voltage is greater than the second ratio, recording the current time as the first moment.
[0132] It should be noted that the third preset location can be a fixed location in the voltage history log for storing the current voltage value, determined by the SCSI protocol or customized, and typically located at the beginning of the log data. The fourth preset location can be a location in the log for storing a reference voltage value, typically immediately following the current voltage value, for comparison with the current voltage value to assess whether the device's voltage status is safe.
[0133] The current voltage can be the actual voltage value measured by the external device upon receiving a voltage status query command, used to monitor the device's power status in real time. The reference voltage is the manufacturer's recommended or maximum allowable voltage value, used as a standard for assessing whether the current voltage is safe. The second ratio can be a preset voltage safety threshold, used to compare the current voltage to the reference voltage to determine whether the firmware upgrade should be delayed.
[0134] The system determines that the log data is of a voltage type. It then reads the data from the third preset location (typically storing the current voltage) and the fourth preset location (typically storing the reference voltage) in the log data. Next, the system compares the ratio of the current voltage to the reference voltage. If the ratio is greater than a second preset ratio (for example, 80%), the voltage status is deemed to not meet the upgrade requirements. The system then records the current time as the "first moment" and delays the firmware upgrade until the voltage status meets the preset requirements.
[0135] In an optional implementation, to ensure data security and device stability during the firmware upgrade process, the storage system must comprehensively assess the external device's environment and performance, with voltage status being a key indicator. The system retrieves historical voltage log data by sending the LOG SENSE command (Page Code: 0x38). The structured log data returned by this command includes current voltage and reference voltage information. By reading and parsing the data in the third and fourth preset locations of the log data, the system can determine the actual power supply status of the external device and its manufacturer's recommended voltage range. The system then evaluates the ratio of the current voltage to the reference voltage to determine whether the device is within a suitable voltage range for the upgrade. If the current voltage exceeds the recommended voltage by a certain percentage (i.e., the second ratio), the system deems the voltage unstable and an immediate firmware upgrade inappropriate. The system then records the current time as the starting point for a delayed upgrade, known as the "first moment." This automatically initiates a delay period, waiting for the device's voltage to return to normal before reinitiating the firmware upgrade decision process.
[0136] In an optional implementation, the query instruction LOG SENSE (opareation code=0x4D) may be:
[0137] Page Code: 0x38 (Voltage History Page, vendor specific, custom implementation);
[0138] Subpage Code:0x00;
[0139] Allocation Length: 255 bytes;
[0140] Operation Code 0x4D is the hexadecimal operation code for the LOG SENSE command, which is used to request log information from a storage device. This command is part of the SCSI command set and is used to obtain diagnostic and status data from a device.
[0141] Page Code 0x38 specifies the type of log requested. In this example, Page Code 0x38 points to the voltage history page, a custom implementation by the hard drive manufacturer that records changes in the drive's power supply voltage. Because Page Code 0x38 falls within the vendor-specific range (0x30 to 0x3F), its content and format do not conform to the unified SCSI standard but are instead designed by the specific device manufacturer based on their needs.
[0142] The subpage code 0x00 further refines the retrieval of the page content. In the voltage history page, a subpage code of 0x00 usually means requesting the data of the entire page, with no special subpage information specified.
[0143] The Allocation Length specifies the size of the host's buffer for receiving data. When sending the LOGSENSE command, the host informs the device of the amount of data it can accept. For example, a value of 255 bytes means the host has prepared a 255-byte buffer to receive the voltage history log data. The device will attempt to return sufficient data within the allocated length. If the returned data is shorter than the allocated length, the remaining space may be padded with zeros or left blank.
[0144] Figure 5 This is a schematic diagram of optional voltage log data according to an embodiment of the present application. Similar to the temperature log data format, byte 0 of the voltage log data format can also indicate a page code (38h). Byte 1 can be a reserved bit, and bytes 2-3 can indicate the valid data length. Byte 4 can indicate the current voltage value, byte 5 can indicate the reference voltage value, byte 6 can indicate the maximum voltage value, and byte 7 can indicate the minimum voltage value.
[0145] Example 4: The specific process of obtaining hard disk voltage information through SCSI commands mainly includes the steps of sending commands, receiving responses, parsing data, etc. The following uses the LOG SENSE command to obtain voltage information as an example to explain this process in detail:
[0146] First, you need to build a LOG SENSE instruction packet, which includes the following fields:
[0147] Operation Code: Set to 0x4D, which represents the LOG SENSE instruction.
[0148] Page Code: Set to 0x38, indicating that the voltage history page is to be read.
[0149] Subpage Code: If the device supports subpages, it can be set to 0x00, indicating that the main page information is obtained.
[0150] Allocation Length: Specifies the buffer size for receiving data, such as 255 bytes.
[0151] Then the constructed LOG SENSE instruction packet is sent to the target hard disk (external device) through the SCSI interface.
[0152] After receiving the LOG SENSE command, the hard drive returns data from the voltage history page as requested. This data typically includes information such as the current voltage, reference voltage, maximum voltage, and minimum voltage.
[0153] The response packet is parsed according to the SCSI protocol format. First, the first two bytes of the response packet are verified to ensure successful command execution. If the returned status code indicates failure, log the error and retry or take other measures. The page code is read from the third byte of the response packet to confirm it is 0x38 (voltage history page). The fourth and fifth bytes are then read to parse the length of the returned data.
[0154] Specifically for page 0x38, its data structure is as follows:
[0155] Current Voltage: Usually located at a specific byte position in the instruction return data, such as Byte[4].
[0156] Reference Voltage: Usually follows the current voltage, such as Byte[5].
[0157] Maximum Voltage: May also be included in the page data to understand the highest voltage the hard drive can withstand.
[0158] Minimum Voltage: Used to understand the minimum voltage requirement for hard disk operation.
[0159] Because voltage information is stored in binary format, the following parsing is typically required: Starting from Byte [4], parse the current voltage value based on the actual data format (byte count and encoding method). Starting from Byte [5], parse the reference voltage value according to the data format requirements. Parsing from the corresponding position will determine the voltage tolerance of the hard drive.
[0160] The parsed voltage value is compared with the preset threshold to assess whether the current voltage conditions are suitable for the firmware upgrade. If the current voltage approaches or exceeds 80% of the maximum voltage, the upgrade should be delayed and an alarm should be reported.
[0161] Through the above-mentioned implementation of the present application, not only can firmware upgrades in an unstable voltage environment be effectively avoided, reducing the risk of hardware failure and data loss, but it can also achieve automated management and operation and maintenance, reduce the burden of manual monitoring, and improve the operational efficiency and reliability of the data center.
[0162] In an optional embodiment, determining at least one status sub-data from at least one log data according to the data types corresponding to the respective log data includes: when the data type of the log data is response time data, reading data from a fifth preset position of the log data as a first response time, and reading data from a sixth preset position of the log data as a second response time, wherein the first response time is used to indicate a response time of a read command of the external device within a preset period, and the second response time is used to indicate a response time of a write command of the external device within a preset period;
[0163] When the first state data does not meet the preset state condition, recording the current time point as the first moment includes: when the first response time or the second response time is greater than the preset response time, recording the current time as the first moment.
[0164] It should be noted that response time data can be statistical information about the time it takes an external device (such as a hard drive) to process read and write commands within a specific period (e.g., the last 5 minutes) and is a key indicator of device performance. The fifth preset location can be a specific location in the log data for storing the response time of read commands within the most recent preset period, typically defined by the device manufacturer or industry standards. The sixth preset location can be a location for storing the response time of write commands within the most recent preset period. This location, separate from the fifth preset location, provides a comprehensive view of the device's read and write performance.
[0165] The preset period can be a time window for collecting read and write response time data in order to evaluate device performance, such as the past 5 minutes, which represents the working efficiency and load status of the device in the recent period.
[0166] The first response time can be the average response time for the device to process read commands within a preset period, used to determine whether the device is in a low-load state and suitable for firmware upgrades. The second response time can be the average response time for the device to process write commands within a preset period, and together with the first response time, it constitutes a comprehensive evaluation of the device's I / O performance. The preset response time can be a system-set threshold used to determine whether the response time data is abnormally high. If it exceeds this threshold, the device is considered to be busy and unsuitable for firmware upgrades.
[0167] The system determines that the received log data belongs to response time data. Then, by reading and parsing the first response time at the fifth preset position and the second response time at the sixth preset position, it evaluates whether the read and write performance of the device meets the requirements for the firmware upgrade. If any response time exceeds the preset response time (for example, 500 milliseconds), the system believes that the device is currently under high load or busy state, and the firmware upgrade may bring additional risks or affect business continuity. At this time, the system will record the current time point as the "first moment" and trigger the delay strategy until the device status is restored to within the preset status conditions.
[0168] In an optional embodiment, during the critical phase of intelligent firmware upgrade decision-making, the storage system must evaluate the device's real-time performance, particularly read and write response times, to ensure that upgrades are not performed while the device is under high load. This is achieved through the LOG SENSE command in the SCSI protocol, which the storage system sends to retrieve the device's performance log data (Page Code: 0x37, Subpage Code: 0x05). When parsing this log data, the system focuses on the first response time at the fifth preset location and the second response time at the sixth preset location, which respectively reflect the device's read and write performance within a preset period. If the analysis results indicate that the average response time for either read or write commands exceeds the preset response time limit, the system records the current time as the "first moment" and implements a delayed upgrade strategy to avoid firmware upgrades during periods of poor device performance until the device returns to normal status during subsequent testing.
[0169] In an optional implementation, the query instruction LOG SENSE (opareation code=0x4D) may be:
[0170] Page Code: 0x37 (SSD Performance Statistics, vendor specific, custom implementation);
[0171] Subpage Code: 0x05, specifies to return the statistical data of the last 5 minutes;
[0172] Allocation Length: 255 bytes;
[0173] Operation Code: The LOG SENSE command's Opcode is 0x4D, which is fixed in the SCSI command set. When using the LOG SENSE command, a Page Code must be specified to request a specific type of log data. In this case, Page Code 0x37 is used for SSD performance statistics. It's worth noting that 0x37 can be vendor-specific, meaning different manufacturers may have different implementations specific to their devices. Subpage Codes allow for more granular data requests. Specifying 0x05 means that statistics are requested for the last five minutes. This is also based on the hard drive manufacturer's custom implementation; different SSDs may have different Subpage Codes to represent data for different time periods. Allocation Length specifies the size of the receive buffer used to store log data returned from the device. If set to 255 bytes, the device will return up to 255 bytes of log data. This value should be set based on the expected data volume to ensure that all required information is correctly received.
[0174] Figure 6 is a schematic diagram of optional performance log data according to an embodiment of the present application; Figure 6 This figure shows the format of performance log data for an external device. By reading data at different byte positions, corresponding performance data can be obtained. Similarly, byte 0 is the page code (37h), and byte 1 is the subpage code (05h), indicating data within 5 minutes. Bytes 2-3 indicate the valid data length. Bytes 132-139 indicate the read command response time, and bytes 156-163 indicate the write command response time.
[0175] When the device responds to the LOG SENSE command, it returns a data packet containing specific performance statistics (log data). The structure and content of this data packet are defined by the manufacturer. In the case of Page Code: 0x37, it usually includes:
[0176] Average Host Read Command Response Time is the average host read command response time, located at a specific position in the data packet, usually Byte
[132] to Byte
[139] , a total of 8 bytes.
[0177] Average Host Write Command Response Time is the average host write command response time, located from Byte
[156] to Byte
[163] , also 8 bytes.
[0178] These time data are usually stored in microseconds (μs) and need to be properly converted to understand their true value. For example, the data read from Byte
[132] to Byte
[139] should be converted to an integer to represent the average response time of all host read commands during this period; similarly, the data from Byte
[156] to Byte
[163] represents the average response time of host write commands.
[0179] In the SCSI protocol, the time data in the SSD Performance Statistics Page returned by the LOG SENSE command is typically encoded as a 64-bit floating-point number to ensure high-precision time measurement and storage. Therefore, when the description "the average read IO command response time of the hard disk in the last 5 minutes" is parsed with an 8-byte length starting from Byte
[132] , this is because a 64-bit floating-point number takes up exactly 8 bytes (1 byte = 8 bits).
[0180] Performance statistics, especially time-related data, often require high precision to reflect the operating status and performance characteristics of the drive. Using 64-bit floating-point numbers provides higher precision than 32-bit floating-point numbers or integers, which is crucial for monitoring microsecond-level response times. 64-bit floating-point numbers support a wider data range, which is necessary for response times that can span multiple orders of magnitude (from microseconds to milliseconds, and even seconds in extreme cases).
[0181] Example 5: Specific process of obtaining hard disk latency information and data analysis:
[0182] Send the LOG SENSE command to request performance statistics from an external device, specifically the response times of recent read and write I / O commands, to assess latency. Upon receiving the LOG SENSE command, the external device returns a log page containing performance statistics. This data is organized in a pre-defined format, including average read and write command response times.
[0183] The returned data must be parsed according to the SCSI standard and vendor-specific definitions to extract the average read and write command response times. Read the page code and subpage code to confirm that the returned data originates from page code 0x37 and subpage code 0x05 to ensure data validity and relevance. Based on the above structure, determine the starting positions of the Average Host Read Command Response Time and Average Host Write Command Response Time. Extract the 8 bytes from Byte 132 to Byte 139 and, considering whether the data is stored in big-endian or little-endian format, convert them to integers in the correct byte order to obtain the average read command response time (in microseconds). Similarly, extract the data from Byte 156 to Byte 163 to obtain the average write command response time. Compare the obtained average read and write response times to pre-set thresholds. If either value exceeds the threshold (for example, 500 milliseconds), the hard drive is considered to be under heavy load and unsuitable for a firmware upgrade.
[0184] The above-mentioned implementation of this application not only ensures that firmware upgrades are performed in optimal device conditions, effectively reducing the risk of upgrade failures due to busy devices, but also enables intelligent and automated operations and maintenance, significantly improving data center maintenance efficiency and security capabilities. This intelligent decision-making based on hard drive performance data can significantly improve the security and success rate of firmware upgrades, reducing the impact on business operations.
[0185] In an optional embodiment, after receiving the second status data sent by the external device at the second moment whose interval with the first moment meets the delay period, it includes: when the i-th status data does not meet the preset status condition, recording the current time point as the i-th moment; at the i+1-th moment whose interval with the i-th moment meets the delay period, receiving the i+1-th status data sent by the external device, where i is an integer greater than 1; when i+1 is greater than or equal to the preset number of times, stopping the upgrade of the external device.
[0186] It should be noted that if the device status does not meet the upgrade conditions during the first status check, the system records this time point as the "first moment" and starts a delay cycle. After the delay period ends (i.e., the "second moment"), the system will receive the device status data again and evaluate the conditions again. This process will be repeated. Each time the device status does not meet the conditions, the system will record a time point as the "i-th moment" and restart the delay cycle. Until the "i+1-th moment" (i.e., the end of the i-th round of delay cycle), if the system finds that the device status still does not meet the upgrade conditions, and the i+1 value at this time exceeds the preset retry limit (i.e., the "preset number"), the system will stop subsequent firmware upgrade attempts to avoid continued invalid or risky operations.
[0187] In an optional implementation, firmware upgrades are performed while the device is in a stable and secure state, avoiding the risk of failure or device damage caused by performing the upgrade under adverse conditions. When the storage system detects during initial testing that the device status (such as temperature, voltage, response time, etc.) does not meet preset status conditions, the system automatically records the current time as the starting point for a delayed upgrade, known as "time one," and begins a preset delay period, for example, 30 minutes. Once the delay period ends, at "time two," the system receives and evaluates the device's status data again. If the device's status still does not meet the upgrade conditions, the system repeats this process, recording the "time i" of each test and receiving "time i+1" status data after the delay period for a new round of evaluation. This cycle continues until the "time i+1" value reaches or exceeds the preset retry limit, at which point the storage system stops attempting the firmware upgrade, terminating the entire upgrade process. This mechanism not only improves the success rate of firmware upgrades but also ensures stable device operation while avoiding unnecessary resource waste and service interruption.
[0188] In an optional embodiment, when the i-th state data does not meet the preset state condition, recording the current time point as after the i-th moment includes: when a conditional instruction sent by an external device is received, receiving the (i+1)-th state data sent by the external device, wherein the conditional instruction is used to indicate that the current state data of the external device meets the preset state condition.
[0189] It should be noted that when the storage system records that the device status does not meet the upgrade requirements at time i, it initiates a delay period, waiting for the device status to recover. During this delay period, if a conditional instruction is received from an external device, this indicates that the device's current status may have improved to meet the preset status requirements. At this point, the storage system receives the (i+1)th status data from the external device based on the conditional instruction and reconfirms whether the device status is suitable for the firmware upgrade. This eliminates the need to re-query the status data during the delay period.
[0190] In an alternative embodiment, suppose that during a data center firmware upgrade, the system performs its first status check at 2:00 PM and discovers that the voltage level of the SAS hard drive does not meet the preset status conditions (e.g., the ratio of the current voltage to the reference voltage exceeds 85%). The system then records 2:00 PM as "time 1" and initiates a 30-minute delay period. However, halfway through the delay period (i.e., 2:15 PM), the storage system receives a conditional command from the hard drive, indicating that the device status may have improved to meet the upgrade conditions.
[0191] In response to this conditional instruction, the storage system immediately receives the "second status data" sent by the hard drive, which contains information about the current voltage status. After analyzing this data, the storage system finds that the hard drive's voltage level has returned to a safe range, with the ratio of the current voltage to the reference voltage falling below 80%. This indicates that the device's status has naturally recovered within the delay period, meeting the pre-set status conditions, and the firmware upgrade can proceed.
[0192] In an optional implementation, after receiving the first status data sent by the external device, the method includes: if the first status data meets a preset status condition, upgrading the device firmware of the external device.
[0193] It should be noted that when the system obtains the device's real-time status information (i.e., first status data) through a status query command and determines whether this data meets the preset status conditions, if so, the device is deemed to be in a suitable state and can undergo a firmware upgrade. Conversely, if the status data does not meet the preset conditions, the system will delay or interrupt the upgrade operation according to the aforementioned intelligent decision-making process.
[0194] After the storage system completes a comprehensive check of device status (such as temperature, voltage, and response time), if the system determines that all status data is within an acceptable safety range—that is, the temperature is stable, the voltage is normal, and the response time is below the preset threshold—this indicates that the device is in a suitable environment for a firmware upgrade. At this point, the system automatically initiates the firmware upgrade process and updates the firmware for the external device without any additional delays or alerts. This mechanism ensures that the firmware upgrade is performed when the device is in optimal condition, significantly improving the upgrade success rate and reducing potential risks during the upgrade process, such as hardware failure, data loss, or business interruption.
[0195] Through the above-mentioned implementation of the present application, when the first status data meets the conditions, the firmware upgrade is performed. When it does not meet the conditions, multiple rounds of queries can be performed to determine the status of the external device and make multiple judgments. In addition, in one process, by using conditional instructions, the storage system no longer waits for the preset delay period to end, but immediately terminates the current delay timing and restarts the firmware upgrade process. The firmware upgrade can be performed immediately when the device status is restored to the optimal moment, avoiding unnecessary waiting and improving the efficiency and timeliness of the upgrade. Through such a mechanism, the system can respond to real-time changes in device status more flexibly, ensuring that firmware upgrades are performed under the safest and optimal conditions, while reducing interference with the daily operations of the data center.
[0196] Figure 7 Schematic diagram of an optional device firmware upgrade method according to an embodiment of the present application; Figure 7 As shown, before performing a firmware upgrade on an external device, it is necessary to detect the status data of the external device. First, step S702 is executed to receive the temperature data of the external device, and then step S704 is executed to determine whether the temperature data meets the preset status conditions. If not, step S706-2 is executed to determine whether the number of delays has reached the preset number; if it does, step S706-1 is executed to receive voltage data. Then, step S708 is executed to determine whether the voltage data meets the preset status conditions. If not, step S710-2 is executed to determine whether the number of delays has reached the preset number; if it does, step S710-1 is executed to receive delay data, and then step S712 is executed to determine whether the delay data meets the preset status conditions. If not, step S714-2 is executed to determine whether the number of delays has reached the preset number; if it does, step S714-1 is executed to perform the firmware upgrade. After the judgment of S706-2, S710-2 and S714-2, if the preset number of times is not reached, execute steps S716, S718 or S720, and re-execute step S702 after the delay period to receive temperature data; if the preset number of times is not reached, execute step S722 to stop the upgrade process.
[0197] In an alternative embodiment, Figure 7 The process is an environmental check before upgrading the hard drive firmware. It is designed to ensure that the upgrade is carried out under appropriate conditions to reduce the risk of upgrade failure and avoid unnecessary impact on business. The following is a more detailed description of the process:
[0198] First, the system sends a SCSI LOG SENSE command to the target hard drive to obtain its current temperature information. This command uses the temperature log page (Page Code: 0x0D). The hard drive receives its response data and then begins parsing it to obtain the temperature information. During the parsing process, the system checks the parameter code. For temperature, a code of 0000h indicates the current temperature, and 0001h indicates the reference temperature. The system compares the parsed current temperature with the reference temperature to determine whether it is below 80% of the reference temperature. If the temperature is acceptable, the process continues; if it is too high, the system proceeds to a delayed escalation step and reports an alarm.
[0199] If the temperature check fails, the system starts a 30-minute timer. This timer waits for environmental conditions to improve before trying the temperature check again. When the timer expires, the system automatically executes the callback function and restarts the environmental check process, starting with the temperature check.
[0200] If the temperature check passes, or if a retry attempt occurs after the timer expires, the system continues to send a LOGSENSE command to the hard drive, this time to obtain voltage information (Page Code: 0x38). The system parses the voltage data returned by the hard drive, reading the current voltage and the maximum operating voltage. Based on the parsed voltage data, the system determines whether the current voltage is less than 80% of the maximum voltage. If the voltage is appropriate, the process continues to the next step. Conversely, if the voltage is too high, the system will delay the escalation and report an alarm.
[0201] When the voltage is appropriate, the system sends the LOG SENSE command again, this time to obtain the drive's read and write latency (Page Code: 0x37, Subpage Code: 0x05). The system parses the returned data to determine the average read and write command response time over the past five minutes. If the average read or write command response time is less than 500 milliseconds, the drive is considered to be in a low-latency state and the environment is suitable for a firmware upgrade. If either response time exceeds the threshold, the drive is considered to be in a high-latency state and is not suitable for an immediate upgrade.
[0202] After checking temperature, voltage, and latency, the system will comprehensively determine whether all conditions meet the upgrade requirements. If all checks pass, the hard drive will enter a silent state and prepare for the firmware upgrade process. If any check fails, the system will delay the upgrade and issue an alarm according to the preset mechanism.
[0203] Through the above steps, the system can perform a comprehensive environmental condition check on the hard drive to be upgraded, ensuring that the firmware upgrade is performed under the safest and most appropriate conditions, thereby improving the upgrade success rate, reducing the impact on business, and realizing intelligent operation and maintenance management.
[0204] In an optional embodiment, upgrading the device firmware of an external device includes: establishing a command queue, wherein the command queue is used to store read and write instructions, and the read and write instructions include read commands or write commands; when the external device receives the read and write instructions, storing the read and write instructions in the command queue; when the device firmware upgrade of the external device is completed, sending at least one read and write instruction in the command queue to the external device.
[0205] It should be noted that the command queue is a data structure used to temporarily store all read and write commands received from the host during the firmware upgrade process, ensuring that data transmission between the host and the storage device does not interfere with or cause data loss during the upgrade. Read and write commands can be read or write commands issued by the host to the storage device to perform data read and write operations. During the firmware upgrade, these commands are temporarily stored in the command queue.
[0206] When the storage system decides to upgrade the firmware of an external device, it first establishes a command queue to temporarily store all read and write commands to the device. During the upgrade process, any read and write commands sent from the host to the device are stored in this queue rather than being sent directly to the device. This ensures that the device can focus entirely on the firmware upgrade without being interrupted by external I / O operations. Once the device completes the firmware upgrade and returns to normal operation, the storage system resends all read and write commands in the command queue to the device, resuming normal data read and write operations.
[0207] In an optional embodiment, before starting the upgrade, the system establishes a command queue to collect and store all read and write instructions issued by the host to external devices during this period. This is to avoid the impact that real-time IO operations may have on the upgrade process, such as data errors, upgrade interruptions, and other issues. During the upgrade, all read and write instructions will be temporarily stored in the command queue instead of being forwarded directly to the device. After the device firmware upgrade is completed and it is verified that the device has returned to normal working status, the storage system will send the read and write instructions in the command queue to the device one by one, so that the device can quickly resume normal data processing capabilities. This mechanism effectively balances the needs of upgrade operations and business continuity, ensuring efficient operation and maintenance of the data center and data security.
[0208] In an optional embodiment, before the firmware upgrade begins, the storage system performs a "silent operation," which is a pre-set security mechanism designed to protect the upgrade process from external read and write operations, thereby increasing the success rate of the upgrade and ensuring data security.
[0209] The system places the hard drive to be upgraded into a quiesced state, meaning it temporarily stops processing external I / O (input / output) commands. This state provides the drive with the stable environment necessary for the upgrade, avoiding potential interruptions or data consistency issues caused by data read and write operations. During this period, although the hard drive itself will not respond to new I / O requests, the storage system will continue to receive I / O commands from the host. These commands are not sent directly to the drive but are instead collected and temporarily stored in an internal queue called the "IO quiescing queue" (command queue).
[0210] The storage system maintains this queue to ensure that all IO commands received during the silent period are recorded so that they can be reapplied to the hard drives in order after the upgrade is complete. This is done to avoid a large number of random IO requests immediately after the upgrade, which would overload the hard drives and affect their performance and stability.
[0211] After the firmware upgrade is complete, it is necessary to verify that the drive can safely resume normal I / O operations. Once the firmware upgrade is deemed complete and the system receives a successful response, it removes the drive from quiesced state, reactivating it as an active drive and beginning to receive and execute queued I / O commands. Furthermore, the system checks all I / O commands that were temporarily stored in the quiesced queue to ensure they are correctly replayed in sequence, restoring normal service to the drive without missing any operations.
[0212] Through the aforementioned implementations of this application, by employing a silent operation mechanism, the storage system can intelligently manage and control the firmware upgrade process, improving upgrade security while minimizing the impact on user services and ensuring immediate availability and high performance of the upgraded hard drives. This process is a key component in achieving automated and intelligent storage management, and is particularly well-suited for the frequent firmware updates required in large data centers and enterprise-level storage environments.
[0213] In an optional embodiment, upgrading the device firmware of an external device includes: dividing the target device firmware into multiple target sub-firmwares according to a preset data volume, wherein the data volume of the target sub-firmwares is the preset data volume; reading the target sub-firmwares one by one, and sending the read target sub-firmwares to the external device.
[0214] It should be noted that the target device firmware can be a new firmware version planned to be installed or updated on an external device, typically a complete firmware file. The preset data size is the maximum amount of firmware data that the system sets for a single transfer to ensure efficient and reliable data transmission. This setting prevents performance bottlenecks or errors that may be caused by excessive data volumes during data transmission. The target sub-firmware is the firmware data segmented into segments by the target device firmware according to the preset data size, facilitating segment-by-segment transfer and updates.
[0215] The storage system divides the target device firmware into segments according to the preset data size, generating a series of target sub-firmware, each with a data size equal to the preset size. The system then sequentially reads these sub-firmware segments and sends them to the external device using the WRITE BUFFER command or other appropriate data transfer mechanism, ensuring stable data transmission and continuous device upgrades.
[0216] During the firmware upgrade process, to overcome transmission issues that may arise from the firmware file's excessive data volume, the storage system splits the complete target device firmware file into multiple target sub-firmware files according to a preset data size. This preset data size is typically set based on the device's write cache size and the system's transmission capacity to ensure that each sub-firmware file is of appropriate size, neither too large to burden the device's processing nor too small to affect transmission efficiency. Next, the system reads the split target sub-firmware files one by one and uses appropriate firmware upgrade instructions (for example, the WRITE BUFFER instruction) to send each sub-firmware file to the external device until all segmented firmware data has been successfully written to the device. This segmented upgrade and segment-by-segment sending mechanism not only improves the efficiency and success rate of firmware upgrades, but also reduces the risk of device performance bottlenecks and data transmission errors that may be encountered during the upgrade process.
[0217] Example 6:
[0218] During a firmware upgrade operation in a data center, an administrator planned to upgrade the firmware of a SAS hard drive. The total size of the target device firmware file was 16MB, and the preset data size was set to 512KB per transfer, which meant that the entire firmware file needed to be split into 32 target sub-firmwares.
[0219] Before the upgrade begins, the storage system first splits the firmware file into 32 512KB target sub-firmware files. The system then reads each sub-firmware file one by one and sends it to the hard drive using the WRITE BUFFER command. During the transfer process, the system monitors the device status to ensure that the transfer occurs when the device is in a suitable condition, avoiding firmware upgrades when the device is under heavy load or unstable.
[0220] For example, the system will read the first 512KB data segment and send it to the drive's write buffer using the WRITE BUFFER command. This process repeats until all target sub-firmware data has been successfully transferred. After each transfer, the system will also use the TEST UNIT READY command to check the drive's status to confirm that it is ready to receive the next firmware data segment or complete the post-upgrade reboot process.
[0221] Through the above-described implementation of this application, the stability and success rate of firmware upgrades can be ensured even when the firmware file is large, avoiding problems such as device overload or data transmission errors caused by the one-time transmission of firmware data. Furthermore, by monitoring device status and adjusting upgrade strategies in a timely manner, the system can effectively reduce the impact of firmware upgrades on daily data center operations and enhance the intelligent level of operation and maintenance.
[0222] In an optional embodiment, upgrading the device firmware of an external device includes: sending a firmware activation instruction and a status detection instruction to the external device, wherein the firmware activation instruction is used to activate the device firmware, and the status detection instruction is used to verify the upgrade status of the external device; upon receiving a first signal sent by the external device, the device firmware upgrade of the external device is completed.
[0223] It should be noted that the firmware activation command is a special command used to activate the new version of firmware after an external device firmware upgrade, making it the main firmware used by the device during operation. The status detection command is used to query and verify the status of the external device during the firmware upgrade process and after the upgrade is completed, ensuring that the upgrade operation was completed smoothly and the device is operating normally. The first signal is a confirmation signal received from the external device, indicating that the device firmware upgrade is complete and the new firmware has been successfully activated, and the device is ready to be put back into service.
[0224] During the final phase of the firmware upgrade, the storage system takes two important steps to ensure the integrity of the upgrade operation and device availability. First, the system initiates the new firmware image by sending a firmware activation command, enabling the device to recognize and switch to the new firmware version. This command is crucial in the firmware upgrade process, triggering the device's internal mechanisms to set the new firmware version written to the cache as the active version. Next, the storage system sends a status check command to query the device's current status, specifically the upgrade status, to confirm whether the new firmware has been successfully activated and that the device is ready to handle read and write operations and meet business needs.
[0225] Receiving the first signal is a crucial indicator of a successful upgrade. This signal typically contains information about the device's new status after the upgrade, such as the new firmware version number and device operating status, proving that the device has successfully transitioned to the new firmware version and is functioning properly. Only after receiving the first signal will the storage system unquiesce the device, allowing it to resume data reading, writing, and other operations, marking the end of the firmware upgrade process.
[0226] Example 7:
[0227] In a data center, an administrator is performing a firmware upgrade on a SAS hard drive. After the new firmware data is written to the drive cache, the storage system first invokes a specific mode of the WRITE BUFFER command (the firmware activation command), namely activation mode (for example, downloading microcode and activating it), to activate the new firmware image. The system then verifies the firmware upgrade status by issuing a LOG SENSE command, specifically monitoring log pages related to the firmware upgrade, such as SSD performance statistics.
[0228] Assume that after the system sends the firmware activation command, it immediately issues a status check command (such as TEST UNIT READY) to confirm whether the hard drive has been successfully activated and is running the new firmware, and that the device is ready to accept read and write operations. If the hard drive responds with a first signal, indicating that the TEST UNIT READY command returned a success, this means that the hard drive firmware upgrade is complete, the new firmware has been successfully activated, and the hard drive is in stable working condition, able to meet data read and write requirements.
[0229] After receiving the first signal, the storage system will perform the following operations: update the firmware version information of the hard disk to reflect the upgrade that has just been completed; release the hard disk from silent state, allowing it to re-participate in normal business activities such as data reading and writing; send the read and write instructions in the previously temporarily stored command queue to the hard disk one by one, and quickly restore the hard disk's data processing function.
[0230] Through the above-described implementation of the present application, by sending a firmware activation instruction and a status detection instruction, and receiving a first signal, the entire upgrade process is ensured to proceed smoothly and the device quickly resumes normal operation. Through this mechanism, the data center can effectively improve the reliability of firmware upgrades and the level of intelligent equipment maintenance.
[0231] Figure 8Schematic diagram of another optional device firmware upgrade method according to an embodiment of the present application; after the environmental detection is completed, during the firmware upgrade of the external device, the server 802 can execute step S802 to send the target sub-firmware one by one. After the target sub-firmware is sent, step S804 is executed to send a firmware activation instruction to activate the firmware file. Then, step S806 is executed to send a status detection instruction to determine whether the external device 804 has completed the firmware upgrade. If the external device 804 executes step S808 and sends a first signal, the firmware upgrade is determined to be complete.
[0232] In an optional implementation, after determining that the upgrade conditions are met, the storage system will start the firmware upgrade process. This process reads data from the firmware file and transfers it to the target hard disk in batches. The size of the firmware data read each time is set to 512KB. This size selection balances transmission efficiency and data integrity, ensuring that a large amount of data can be transferred in a shorter time, while also facilitating system management and error checking. The storage system will use the WRITE BUFFER instruction to cyclically write the firmware data to the cache area of the hard disk in units of 512KB. This process will continue until all the data in the firmware file has been transferred to the hard disk. The advantage of doing this is that the cache mechanism inside the hard disk can be utilized to optimize the data writing speed and reduce the number of direct accesses to the physical layer of the hard disk, thereby reducing potential operational risks.
[0233] The WRITE BUFFER instruction is not only used to write ordinary data, but can also be used specifically for downloading microcode (another name for firmware). When using this mode, the instruction will be marked as having microcode download functionality, which means that the written data will be considered part of the firmware update. During the firmware update process, "with offsets" allows you to specify the write location of the firmware image in the hard drive's internal cache area, ensuring that each piece of data is accurately written to the predetermined address, which is very important for ensuring the integrity and stability of the firmware upgrade. One of the most important features is "save, and activated." After the firmware data is written, the "save" function ensures that the data is persistently stored rather than just existing in the cache. Subsequently, the "activated" function automatically activates the newly downloaded firmware image, making it the running firmware of the hard drive without the need for additional manual intervention.
[0234] Specifically, the system first decompresses the firmware file prepared for the hard drive upgrade. This decompression process ensures that the firmware file is in a usable state and is ready to be written to the hard drive. The system then opens the decompressed firmware file and reads the firmware data from it for the upgrade. The system then checks to see if the file is successfully opened. If the file opening fails (for example, due to permission issues or file corruption), the system returns a firmware upgrade failure message and logs detailed error information.
[0235] If the file is successfully opened, the system reads a 512KB data block from the firmware file, which is the data transfer unit commonly used for firmware upgrades in the SCSI protocol. The system then checks whether the data block is successfully read. If the read fails, the system reports a firmware upgrade failure and logs the appropriate error message. For SAS hard drives, the system uses the WRITEBUFFER command (Opcode 0x3B) in the SCSI command set and sets the mode to "Download microcode with offsets, save, and activate" (Mode code 0x07). This mode allows the system to download the firmware data to the hard drive's buffer and automatically activate the new firmware upon completion.
[0236] Before sending the WRITE BUFFER command, the system initializes a buffer and writes a 512KB data block into it, setting the command data buffer length to match the data block size. The system then sends the encapsulated WRITE BUFFER command to the hard drive, beginning to write the firmware data to the hard drive's buffer. Once a 512KB data block (the target sub-firmware) has been successfully written, the system updates the firmware file's read position (offset) to point to the next data block. The system then checks whether the read data block is the last block in the firmware file. If not, it continues reading and writing subsequent data blocks. If so, the process continues to the next step.
[0237] Once the entire firmware file's data blocks have been written to the drive's buffer, the system will encapsulate and send a TEST UNITREADY command (status check command) to verify that the drive has completed the internal firmware upgrade process. The result returned by the TEST UNITREADY command indicates the drive's firmware upgrade status. If Success (the first signal) is returned, the drive's internal firmware upgrade is complete and the drive is ready to resume I / O operations. After confirming the firmware upgrade is successful, the system closes the firmware file, concluding the upgrade process. This is the state of the drive after a successful upgrade, at which point it is ready to resume data read and write operations with the new firmware version.
[0238] Through the above process, the storage system can complete hard drive firmware upgrades in an automated and intelligent manner, ensuring not only the stability of the upgrade process but also reducing the need for human intervention, improving upgrade efficiency, and mitigating potential risks. This process is particularly suitable for large data centers and enterprise storage environments, ensuring efficient and secure firmware upgrades without disrupting business operations.
[0239] It's important to note that multiple hard drives in a system can be upgraded collaboratively. By globally sensing the real-time status of all hard drives in the storage system, including but not limited to load, temperature, voltage, and business impact, the system prioritizes and optimally timed firmware upgrades through algorithmic evaluation, minimizing disruption to business operations while ensuring upgrade success and storage system stability.
[0240] In an alternative implementation, a module is required to continuously monitor the status of all hard drives in the storage system. This includes real-time collection of information such as drive load, temperature, and voltage. The LOG SENSE command in the SCSI protocol can be used to periodically retrieve required data from each drive, while also recording and analyzing the system's critical data processing periods to better plan upgrades.
[0241] The collected drive status data needs to be standardized and converted into metrics that can be directly compared and analyzed. For example, temperature and voltage values can be compared with pre-set thresholds to determine whether they are in a suitable state for upgrade; load status can be converted into a percentage to quantify the current level of hard drive activity.
[0242] Develop an upgrade priority assessment algorithm based on real-time drive status data (load, temperature, and voltage). Consider the current drive load level, with lower drives receiving higher priority. Also, prioritize drives with more stable temperature and voltage conditions. Avoid critical system data processing periods, choosing to perform firmware upgrades during off-peak hours. Based on the algorithm's results, select the most suitable drive for upgrade as the next target. If multiple drives have similar status, further prioritize them based on their role in the storage system and business importance.
[0243] Perform standard pre-upgrade preparations on the selected hard drives, including data backup and I / O quiescence. At the optimal time, update the hard drive firmware using the SCSI WRITE BUFFER command according to the previously outlined firmware upgrade process. After completing the firmware upgrade for one drive, the system should reassess the status of all drives and adjust the upgrade plan based on the latest data. If a new drive with better performance becomes available or if environmental conditions change, the intelligent scheduling algorithm should dynamically adjust the upgrade order and timing.
[0244] If, at any point in time, a sudden deterioration in the environmental conditions of a drive is detected, such as a sudden temperature rise, voltage fluctuation, or load surge, upgrades to that drive will be immediately suspended and marked as pending. Adjust the upgrade strategy based on the actual situation, such as relocating the pending drives to the back of the queue, recalculating the optimal upgrade sequence, or postponing the entire upgrade plan until environmental conditions stabilize again. If any potential threat to system stability is detected, such as multiple drives reporting errors simultaneously or storage redundancy dropping to a dangerous level, all firmware upgrades will be terminated immediately and an emergency notification will be sent to the administrator so that remedial measures can be taken promptly.
[0245] Through the above implementation, the storage system can intelligently and securely perform firmware upgrades in a multi-hard disk environment, minimizing the impact on services and improving the overall success rate of firmware upgrades and the availability of the storage system.
[0246] Through the description of the above implementation methods, those skilled in the art can clearly understand that the method according to the above embodiment can be implemented by means of software plus the necessary general hardware platform, and of course it can also be implemented by hardware, but in many cases the former is a better implementation method.
[0247] The embodiment of the present application also provides a device firmware upgrade apparatus, Figure 9 This is a structural block diagram of an optional device firmware upgrade device according to an embodiment of the present application. Figure 9 As shown, the device includes:
[0248] A first data receiving module 902 is configured to send a firmware upgrade instruction to an external device and receive first status data sent by the external device, wherein the first status data is used to indicate environmental health information and performance status information of the external device;
[0249] A time recording module 904 is configured to record the current time point as the first moment when the first state data does not meet the preset state condition;
[0250] A second data receiving module 906 is configured to receive second status data sent by an external device at a second moment that is spaced from the first moment by a delay period;
[0251] The firmware upgrade module 908 is configured to upgrade the device firmware of the external device when the second status data meets a preset status condition.
[0252] Optionally, the above-mentioned first data receiving module 902 is also used to: send at least one query instruction to an external device, wherein the query instruction is used to indicate a type of status sub-data; receive first status data sent by the external device, wherein the first status data includes status sub-data corresponding to each of the at least one query instruction.
[0253] Optionally, the above-mentioned first data receiving module 902 is also used to: receive at least one log data sent by an external device, wherein the log data includes status data, and the log data corresponds to the query instruction one-to-one; read the header information of at least one log data, and determine the data type corresponding to each log data based on the header information; determine at least one status sub-data from at least one log data based on the data type corresponding to each log data, wherein the current status sub-data is determined from the current log data based on the data type of the current log data.
[0254] Optionally, the above-mentioned first data receiving module 902 is also used to: when the data type of the log data is temperature data, read the indication information of the first preset position of the log data, wherein the indication information is used to indicate the category of the temperature data stored at the first preset position; when the indication information is the first indication information, determine the current temperature data from the first preset position, and determine the reference temperature data from the second preset position; the above-mentioned time recording module 904 is also used to: when the ratio of the current temperature data to the reference temperature data is greater than the first ratio, record the current time as the first moment.
[0255] Optionally, the above-mentioned first data receiving module 902 is also used to: when the data type of the log data is voltage data, the data read from the third preset position of the log data is the current voltage; the data read from the fourth preset position of the log data is the reference voltage; the above-mentioned time recording module 904 is also used to: when the ratio of the current voltage to the reference voltage is greater than the second ratio, record the current time as the first moment.
[0256] Optionally, the above-mentioned first data receiving module 902 is also used to: when the data type of the log data is response time data, read the data of the fifth preset position of the log data as the first response time, and read the data of the sixth preset position of the log data as the second response time, wherein the first response time is used to indicate the response time of the external device to read the command within the preset period, and the second response time is used to indicate the response time of the external device to write the command within the preset period; the above-mentioned time recording module is also used to: when the first response time or the second response time is greater than the preset response time, record the current time as the first moment.
[0257] Optionally, the above-mentioned second data receiving module 906 is also used to: when the i-th state data does not meet the preset state conditions, record the current time point as the i-th moment; at the i+1-th moment whose interval with the i-th moment meets the delay period, receive the i+1-th state data sent by the external device, where i is an integer greater than 1; when i+1 is greater than or equal to the preset number of times, stop upgrading the external device.
[0258] Optionally, the second data receiving module 906 is further configured to receive the (i+1)th status data sent by the external device upon receiving a conditional instruction sent by the external device, wherein the conditional instruction is configured to indicate that the current status data of the external device satisfies a preset status condition.
[0259] Optionally, the first data receiving module 902 is further configured to upgrade the device firmware of the external device when the first status data meets a preset status condition.
[0260] Optionally, the above-mentioned firmware upgrade module 908 is also used to: establish a command queue, wherein the command queue is used to store read and write instructions, and the read and write instructions include read commands or write commands; when the external device receives the read and write instructions, the read and write instructions are stored in the command queue; when the device firmware upgrade of the external device is completed, at least one read and write instruction in the command queue is sent to the external device.
[0261] Optionally, the firmware upgrade module 908 is further configured to: divide the target device firmware into multiple target sub-firmware according to a preset data volume, wherein the data volume of the target sub-firmware is the preset data volume; read the target sub-firmware one by one, and send the read target sub-firmware to the external device.
[0262] Optionally, the above-mentioned firmware upgrade module 908 is also used to: send firmware activation instructions and status detection instructions to the external device, wherein the firmware activation instruction is used to activate the device firmware, and the status detection instruction is used to verify the upgrade status of the external device; when the first signal sent by the external device is received, the device firmware upgrade of the external device is completed.
[0263] For the description of the features in the embodiment corresponding to the device firmware upgrade apparatus, reference may be made to the relevant description of the embodiment corresponding to the device firmware upgrade method, which will not be described in detail here.
[0264] An embodiment of the present application further provides an electronic device, comprising a memory and a processor, wherein the memory stores a computer program, and the processor is configured to run the computer program to execute the steps in any of the above-mentioned device firmware upgrade method embodiments.
[0265] An embodiment of the present application further provides a computer-readable storage medium, in which a computer program is stored, wherein the computer program is configured to execute the steps of any of the above-mentioned device firmware upgrade method embodiments when running.
[0266] In an exemplary embodiment, the computer-readable storage medium may include, but is not limited to, various media that can store computer programs, such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk, or an optical disk.
[0267] An embodiment of the present application further provides a computer program product, which includes a computer program. When the computer program is executed by a processor, the steps of any of the above-mentioned device firmware upgrade method embodiments are implemented.
[0268] An embodiment of the present application also provides another computer program product, including a non-volatile computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it implements the steps in any of the above-mentioned device firmware upgrade method embodiments.
[0269] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the above description has generally described the components and steps of each example according to their functions. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0270] The above is a detailed introduction to the device firmware upgrade method and apparatus, storage medium and electronic device provided by the present application. This article uses specific examples to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only used to help understand the method of the present application and its core idea. It should be pointed out that for ordinary technicians in this technical field, without departing from the principles of the present application, several improvements and modifications can be made to the present application, and these improvements and modifications also fall within the scope of protection of the claims of the present application.
Claims
1. A method for upgrading device firmware, characterized in that: include: Sending a firmware upgrade instruction to an external device, and receiving first status data sent by the external device, wherein the first status data is used to indicate environmental health information and performance status information of the external device; If the first state data does not meet the preset state condition, recording the current time point as the first moment; receiving second status data sent by the external device at a second moment that is spaced apart from the first moment by a delay period; When the second status data meets a preset status condition, the device firmware of the external device is upgraded.
2. The method according to claim 1, characterized in that Before receiving the first status data sent by the external device, the method includes: Sending at least one query instruction to the external device, wherein the query instruction is used to indicate a type of status sub-data; The first status data sent by the external device is received, wherein the first status data includes the status sub-data corresponding to each of the at least one query instructions.
3. The method according to claim 2, characterized in that The receiving the first status data sent by the external device includes: receiving at least one log data sent by the external device, wherein the log data includes the status data, and the log data corresponds to the query instruction one by one; Reading header information of the at least one log data, and determining the data type corresponding to each of the log data according to the header information; At least one state sub-data is determined from the at least one log data according to the data types corresponding to the respective log data, wherein the current state sub-data is determined from the current log data according to the data type of the current log data.
4. The method according to claim 3, characterized in that The determining of at least one state sub-data from the at least one log data according to the data types corresponding to the respective log data includes: In a case where the data type of the log data is temperature data, reading indication information of a first preset location of the log data, wherein the indication information is used to indicate a type of the temperature data stored in the first preset location; In a case where the indication information is the first indication information, determining current temperature data from the first preset position, and determining reference temperature data from the second preset position; When the first state data does not meet the preset state condition, recording the current time point as the first moment includes: When the ratio of the current temperature data to the reference temperature data is greater than a first ratio, the current time is recorded as the first moment.
5. The method according to claim 3, characterized in that The determining of at least one state sub-data from the at least one log data according to the data types corresponding to the respective log data includes: In a case where the data type of the log data is voltage data, the data read from the third preset position of the log data is the current voltage; The data at the fourth preset position of the log data is read as a reference voltage; When the first state data does not meet the preset state condition, recording the current time point as the first moment includes: When the ratio of the current voltage to the reference voltage is greater than a second ratio, the current time is recorded as the first moment.
6. The method according to claim 3, characterized in that The determining of at least one state sub-data from the at least one log data according to the data types corresponding to the respective log data includes: In a case where the data type of the log data is response time data, the data read from the fifth preset position of the log data is a first response time, and the data read from the sixth preset position of the log data is a second response time, wherein the first response time is used to indicate a response time of a read command of the external device within a preset period, and the second response time is used to indicate a response time of a write command of the external device within a preset period; When the first state data does not meet the preset state condition, recording the current time point as the first moment includes: When the first response time or the second response time is greater than the preset response time, the current time is recorded as the first moment.
7. The method according to any one of claims 1 to 6, characterized in that After receiving the second status data sent by the external device at a second moment that is spaced apart from the first moment by a delay period, the method includes: If the i-th state data does not meet the preset state condition, the current time point is recorded as the i-th moment; At the i+1th moment, which is separated from the i-th moment by a delay period, receiving the i+1th status data sent by the external device, where i is an integer greater than 1; When the number i+1 is greater than or equal to a preset number, the upgrade of the external device is stopped.
8. The method according to claim 7, characterized in that When the i-th state data does not meet the preset state condition, recording the current time point as after the i-th moment includes: In case of receiving a conditional instruction sent by the external device, receiving the (i+1)th status data sent by the external device, wherein the conditional instruction is used to indicate that the current status data of the external device meets a preset status condition.
9. The method according to any one of claims 1 to 6, characterized in that After receiving the first status data sent by the external device, the method further includes: When the first status data meets a preset status condition, the device firmware of the external device is upgraded.
10. The method according to claim 1, characterized in that The upgrading of the device firmware of the external device includes: Establishing a command queue, wherein the command queue is used to store read and write instructions, and the read and write instructions include read commands or write commands; When the external device receives a read / write instruction, storing the read / write instruction in the command queue; When the device firmware of the external device is upgraded, at least one read / write instruction in the command queue is sent to the external device.
11. The method according to claim 10, characterized in that The upgrading of the device firmware of the external device includes: Dividing the target device firmware into a plurality of target sub-firmware according to a preset data volume, wherein the data volume of the target sub-firmware is the preset data volume; The target sub-firmware is read one by one, and the read target sub-firmware is sent to the external device.
12. The method according to claim 10, characterized in that The upgrading of the device firmware of the external device includes: Sending a firmware activation instruction and a status detection instruction to the external device, wherein the firmware activation instruction is used to activate the device firmware, and the status detection instruction is used to verify the upgrade status of the external device; When the first signal sent by the external device is received, the device firmware upgrade of the external device is completed.
13. An electronic device, characterized in that: include: Memory for storing computer programs; A processor, configured to implement the steps of the device firmware upgrade method according to any one of claims 1 to 12 when executing the computer program.
14. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, wherein when the computer program is executed by a processor, the steps of the device firmware upgrade method according to any one of claims 1 to 12 are implemented.
15. A computer program product comprising computer instructions, characterized in that When the computer instructions are executed by a processor, the steps of the device firmware upgrade method according to any one of claims 1 to 12 are implemented.
Citation Information
Patent Citations
Hard disk firmware upgrading method and device, storage medium and electronic equipment
CN119025141A
Hard disk firmware upgrading method and device, electronic equipment and storage medium
CN119127254A
Log processing method and system, log management platform, and electronic device
WO2025103085A1