On-site generation of CoreDump file based on firmware stored during abnormal power failure

By identifying abnormal power outages of storage devices, using persistent storage media space to store data blocks on the firmware field, and generating CoreDump files, it solves the problem of difficulty in saving firmware field data in the prior art, and improves the efficiency of fault analysis and debugging.

CN120122989APending Publication Date: 2025-06-10MEMBLAZE TECH BEIJING
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510043058.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-10
Publication Date
2025-06-10

AI Technical Summary

Technical Problem

In the case of abnormal power outage of external power supply, the prior art is difficult to effectively save the on-site data of the firmware, resulting in low fault analysis and debugging efficiency.

Method used

By identifying abnormal power outages in the storage device, using the fixed-sized continuous space in the persistent storage medium space, the data blocks that constitute the firmware field are read and stored in turn, and the CoreDump file is generated for debugging.

Benefits of technology

After abnormal power failure occurs in the storage device, the key firmware field information can be saved in a short time, ensuring that even if the firmware field is not saved incompletely, debugging and analysis can be carried out, greatly improving the debugging efficiency of the storage device.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120122989A_ABST
    Figure CN120122989A_ABST
Patent Text Reader

Abstract

The invention relates to a method for generating a CoreDump file on the basis of a firmware field stored during abnormal power failure, and the method is applied to a host and comprises the steps that a Telemetry Log is read from a storage device, and the Telemetry Log comprises firmware field data; analyzing and unpacking the firmware field data to generate a packaging configuration file and a data sub-block file; and obtaining a core dump file based on the packaging configuration file and the data sub-block file. By the adoption of the technical scheme, when the storage device is powered on again due to abnormal power failure, debugging analysis can be conducted even if firmware is not completely stored on site, and the debugging efficiency of the storage device is greatly improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computer technologies, and particularly to a method for generating a CoreDump file based on a firmware context saved during abnormal power-off. Background Art

[0002] Figure 1 A block diagram of a solid-state storage device is shown. The solid-state storage device 102 is coupled to a host and is used to provide storage capabilities for the host. The host and the solid-state storage device 102 can be coupled in various ways, including but not limited to coupling the host and the solid-state storage device 102 through, for example, SATA (Serial Advanced Technology Attachment), SCSI (Small Computer System Interface), SAS (Serial Attached SCSI), IDE (Integrated Drive Electronics), USB (Universal Serial Bus), PCIE (Peripheral Component Interconnect Express), NVMe (NVM Express), Ethernet, Fibre Channel, a wireless communication network, etc. The host can be an information processing device capable of communicating with the storage device through the above-mentioned ways. For example, the host can be a personal computer, a tablet computer, a server, a portable computer, a network switch, a router, a cellular phone, a personal digital assistant, etc. The storage device 102 includes an interface 103, a control component 104, one or more NVM chips 105, and DRAM (Dynamic Random Access Memory) 110.

[0003] Common NVMs include NAND flash memory, phase change memory, FeRAM (Ferroelectric RAM), MRAM (Magnetic Random Access Memory), RRAM (Resistive Random Access Memory), etc.

[0004] The interface 103 can be adapted to exchange data with the host through, for example, SATA, IDE, USB, PCIE, NVMe, SAS, Ethernet, Fibre Channel, etc.

[0005] The control component 104 is used to control data transfer between the interface 103, the NVM chip 105, and the DRAM 110, and is also used for storage management, mapping of host logical addresses to flash physical addresses, wear leveling, bad block management, etc. The control component 104 can be implemented in a variety of ways through software, hardware, firmware, or a combination thereof. For example, the control component 104 can be in the form of an FPGA (Field Programmable Gate Array), an ASIC (Application Specific Integrated Circuit), or a combination thereof. The control component 104 can also include a processor or a controller, and software is executed in the processor or controller to manipulate the hardware of the control component 104 to process IO (Input / Output) commands. The control component 104 can also be coupled to the DRAM 110 and can access the data in the DRAM 110. The DRAM can store the FTL table and / or the data of the cached IO commands.

[0006] The control component 104 includes a flash interface controller (or also known as a media interface controller, a flash channel controller). The flash interface controller is coupled to the NVM chip 105 and issues commands to the NVM chip 105 in a manner that follows the interface protocol of the NVM chip 105 to operate the NVM chip 105 and receives the command execution results output from the NVM chip 105. Known NVM chip interface protocols include "Toggle", "ONFI", etc.

[0007] Generally, when a storage device faces abnormal power-off of external power supply, it is necessary to save the currently cached IO data, various metadata used by the firmware (such as the FTL table), etc. to the NVM chip. For this purpose, components such as capacitors are usually used as backup power supplies to provide a short internal power supply time for the storage device. However, this time is often very short and the amount of data that can be saved is limited. Therefore, it is required that the firmware has a relatively complex abnormal power-off handling process to achieve the purpose of saving all the data that needs to be saved within a short time.

[0008] However, the complex abnormal power-off handling process brings a higher firmware failure rate. In order to discover and solve these firmware failures, it is necessary to analyze the running situation when the firmware occurs. However, when an abnormal power-off occurs, the backup power supply can only supply power for a very short time, which limits the ways of investigating firmware failures, such as online debugging and other methods cannot be used. Therefore, the firmware needs to save its own firmware scene within an extremely short time, or try to retain key contents so that they can be externally obtained when powered on next time for debugging.

[0009] It is understandable that if the previous power-off did not end properly, the firmware might not have had time to save the firmware context for debugging. Therefore, it is necessary to restore the firmware context for debugging. Summary of the Invention

[0010] To solve the above technical problems or at least partially solve the above technical problems, the present disclosure provides a method for generating a CoreDump file based on the firmware context saved during abnormal power-off.

[0011] An embodiment of the present disclosure provides a method for obtaining the firmware context of abnormal power-off, which is applied to a storage device. The method includes: when the storage device is powered on, if it is recognized that the abnormal power-off handling process did not end properly when the storage device was powered off last time, based on a plurality of fixed-size regions in a continuous space of a fixed size in the persistent storage medium space, sequentially read the data blocks corresponding to the first fixed-size sub-regions of one, part, or all of the plurality of fixed-size regions; if a legal data block cannot be read from the first fixed-size sub-region of the current fixed-size region, read the next fixed-size region, and if a legal data block is read from the first fixed-size sub-region of the current fixed-size region, continue to read the current fixed-size region to obtain all the data blocks constituting the firmware context; determine the firmware context length based on the total sum of the number of all the data blocks; in response to the host reading the Telemetrylog, determine the size of the Telemetry Log based on the firmware context length, and use one, multiple, or all of the all data blocks as the content of the Telemetry Log.

[0012] An embodiment of the present disclosure further provides a method for generating a core dump file using the firmware context of abnormal power-off, which is applied to a host. The host is connected to the storage device. The method includes: reading the Telemetry Log from the storage device, where the Telemetry Log includes firmware context data; parsing and unpacking the firmware context data to generate a packaging configuration file and a data sub-block file; obtaining a core dump file based on the packaging configuration file and the data sub-block file.

[0013] The embodiments of the present disclosure also provide a method for obtaining a core dump file generated from the firmware scene of an abnormal power failure, including: when the storage device is powered on, if it is recognized that the abnormal power failure handling process did not end normally when the storage device was powered off last time, based on multiple fixed-size regions in a continuous space of a fixed size in the persistent storage medium space, sequentially read the data blocks corresponding to the first fixed-size sub-regions of each of one, part or all of the fixed-size regions among the multiple fixed-size regions; if a legal data block cannot be read from the first fixed-size sub-region of the current fixed-size region, read the next fixed-size region, and if a legal data block is read from the first fixed-size sub-region of the current fixed-size region, continue to read the current fixed-size region to obtain all the data blocks constituting the firmware scene; determine the firmware scene length based on the total sum of the number of all the data blocks; the host reads the TelemetryLog from the storage device, and the Telemetry Log includes firmware scene data; in response to the host reading the Telemetrylog, the storage device determines the size of the Telemetry Log based on the firmware scene length, and uses one, multiple or all of the all data blocks as the content of the Telemetry Log; the host parses and unpacks the firmware scene data to generate a packaged configuration file and a data sub-block file; the host obtains a core dump file based on the packaged configuration file and the data sub-block file.

[0014] The embodiments of the present disclosure also provide a system for obtaining a firmware scene of abnormal power-off and generating a core dump file, including a host and a storage device; when the storage device is powered on, if it is recognized that the abnormal power-off handling process did not end normally when the storage device was powered off last time, based on multiple fixed-size regions in a continuous space of a fixed size in the persistent storage medium space, read the data blocks corresponding to the first fixed-size sub-regions of each of one, part or all of the multiple fixed-size regions in sequence; if a legal data block cannot be read from the first fixed-size sub-region of the current fixed-size region, read the next fixed-size region, and if a legal data block is read from the first fixed-size sub-region of the current fixed-size region, continue to read the current fixed-size region to obtain all the data blocks constituting the firmware scene; determine the firmware scene length based on the total sum of the number of all the data blocks; the host is used to read the Telemetry Log from the storage device, and the Telemetry Log includes firmware scene data; in response to the host reading the Telemetry log, the storage device determines the size of the Telemetry Log based on the firmware scene length, and uses one, multiple or all of the all data blocks as the content of the Telemetry Log; the host is also used to parse and unpack the firmware scene data to generate a packaging configuration file and a data sub-block file, and obtain a core dump file based on the packaging configuration file and the data sub-block file.

[0015] The embodiments of the present disclosure also provide an electronic device, which includes: a processor; a memory for storing executable instructions of the processor; the processor is used to read the executable instructions from the memory and execute the executable instructions to implement the method for obtaining a firmware scene of abnormal power-off or the method for generating a core dump file using the firmware scene of abnormal power-off or the method for obtaining a firmware scene of abnormal power-off and generating a core dump file provided by the embodiments of the present disclosure.

[0016] The embodiments of the present disclosure also provide a computer-readable storage medium, where the storage medium stores a computer program, and the computer program is used to execute the method for obtaining a firmware scene of abnormal power-off or the method for generating a core dump file using the firmware scene of abnormal power-off or the method for obtaining a firmware scene of abnormal power-off and generating a core dump file provided by the embodiments of the present disclosure.

[0017] The embodiments of the present disclosure also provide a computer program product, including a computer program, where the computer program, when executed by a processor, implements the method for obtaining a firmware scene of abnormal power-off or the method for generating a core dump file using the firmware scene of abnormal power-off or the method for obtaining a firmware scene of abnormal power-off and generating a core dump file provided by the embodiments of the present disclosure.

[0018] The technical solution provided by the embodiments of the present disclosure has the following advantages compared with the prior art: The solution for saving the firmware scene for abnormal power-off provided by the embodiments of the present disclosure includes a host, which reads the Telemetry Log from a storage device, where the Telemetry Log includes firmware scene data; parses and unpacks the firmware scene data to generate a packaging configuration file and a data sub-block file; and obtains a core dump file based on the packaging configuration file and the data sub-block file. By adopting the above technical solution, when the storage device experiences a power-off anomaly and is powered on again, even if the firmware scene is not completely saved, debugging and analysis can still be performed, greatly improving the efficiency of debugging the storage device. Description of the Drawings

[0019] In combination with the accompanying drawings and with reference to the following specific embodiments, the above and other features, advantages, and aspects of the various embodiments of the present disclosure will become more apparent. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic, and the original components and elements are not necessarily drawn to scale.

[0020] Figure 1 Schematic diagram of a solid-state storage device;

[0021] Figure 2 Block diagram of a high-brightness storage device provided by an embodiment of the present disclosure;

[0022] Figure 3 Flow schematic diagram of a method for saving the firmware scene for abnormal power-off provided by an embodiment of the present disclosure;

[0023] Figure 4 Example diagram of splitting the firmware scene into scene sub-blocks provided by an embodiment of the present disclosure;

[0024] Figure 5 Example diagram of the structure of a scene sub-block provided by an embodiment of the present disclosure;

[0025] Figure 6 Example flow diagram of saving the firmware scene provided by an embodiment of the present disclosure;

[0026] Figure 7 Example diagram of the connection between a host and a storage device provided by an embodiment of the present disclosure;

[0027] Figure 8 Flow schematic diagram of a method for obtaining the firmware scene of abnormal power-off provided by an embodiment of the present disclosure;

[0028] Figure 9A Schematic diagram of the structure of the persistent storage medium space provided by an embodiment of the present disclosure;

[0029] Figure 9BStructural schematic diagram of a fixed-size area provided by an embodiment of the present disclosure;

[0030] Figure 9C Structural schematic diagram of a field sub-block provided by an embodiment of the present disclosure;

[0031] Figure 10 Flow example diagram of firmware in-situ restoration provided by an embodiment of the present disclosure;

[0032] Figure 11 Flow example diagram of firmware in-situ restoration provided by an embodiment of the present disclosure;

[0033] Figure 12 Flow schematic diagram of a method for generating a core dump file using the firmware in-situ under abnormal power-off provided by an embodiment of the present disclosure;

[0034] Figure 13 Example diagram of obtaining the firmware in-situ provided by an embodiment of the present disclosure;

[0035] Figure 14 Example diagram of firmware in-situ restoration provided by an embodiment of the present disclosure;

[0036] Figure 15 Flow example diagram of firmware in-situ restoration provided by an embodiment of the present disclosure;

[0037] Figure 16 Flow example diagram of firmware in-situ restoration provided by an embodiment of the present disclosure;

[0038] Figure 17 Structural schematic diagram of a device for saving the firmware in-situ under abnormal power-off provided by an embodiment of the present disclosure;

[0039] Figure 18 Structural schematic diagram of a device for obtaining the firmware in-situ under abnormal power-off provided by an embodiment of the present disclosure;

[0040] Figure 19 Structural schematic diagram of a device for generating a core dump file using the firmware in-situ under abnormal power-off provided by an embodiment of the present disclosure. Detailed implementation manners

[0041] Embodiments of the present disclosure will be described in more detail with reference to the accompanying drawings. Although some embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. On the contrary, these embodiments are provided to more thoroughly and completely understand the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are only for exemplary purposes and are not used to limit the protection scope of the present disclosure.

[0042] It should be understood that the various steps recited in the method embodiments of the present disclosure may be executed in different orders and / or in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present disclosure is not limited in this regard.

[0043] As used herein, the term "comprising" and its variations are open-ended, i.e., "including but not limited to". The term "based on" means "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". The relevant definitions of other terms will be given in the following description.

[0044] It should be noted that the concepts such as "first" and "second" mentioned in the present disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order or interdependence of the functions performed by these devices, modules or units.

[0045] It should be noted that the modifications of "one" and "plural" mentioned in the present disclosure are illustrative rather than restrictive. Those skilled in the art should understand that unless otherwise clearly indicated in the context, it should be understood as "one or more".

[0046] The names of the messages or information exchanged between multiple devices in the embodiments of the present disclosure are only for illustrative purposes and are not used to limit the scope of these messages or information.

[0047] Based on the foregoing description of the background art, Figure 2 represents a block diagram of a typical storage device to which the embodiments of the present disclosure are applied. Figure 2 Also Figure 1 is a detailed block diagram of the control component 104 of

[0048] The control component includes multiple processor cores (e.g., processor cores 0-3). Each processor core has its own memory (e.g., memories 0-3) for storing the firmware running on the corresponding processor core. The firmware running on each processor core is different. The storage device of each processor core also stores the data generated or used during the operation of the firmware. The software running on an embedded device is usually referred to as firmware.

[0049] The control component further includes a cache / shared memory. The shared memory can be used by the aforementioned processor cores.

[0050] The control component further includes an SRAM, which is mainly used as a cache for storing IO data and is also used in the present application to store the in-situ sub-blocks generated by each processor core.

[0051] Communication can be carried out between the processor cores, either through an inter-core queue (not shown) or shared memory.

[0052] The tasks are divided among multiple processor cores, and one of the cores is selected as the control core.

[0053] Generally, the firmware running on one processor core is called a logic module. For example, the control core is logic module 0. However, multiple logic modules may run on some processor cores, and these logic modules can be regarded as different firmware that can run concurrently. For example, a logic module can be a process, a thread, or a coroutine.

[0054] Firmware context: For subsequent debugging purposes, the firmware context that the processor core / logic module needs to save includes various status information during abnormal power-off, mainly including: register status, data in the independent storage area of the core, data in the inter-core shared fast storage area, data in the inter-core shared slow storage area, stack information, SQ / CQ queues, Admin queue, memory management information, memory information, other status information of the processor, status information of the hardware accelerator, cache content, inter-core communication queue, inter-core shared memory, etc.

[0055] Memory information: It mainly includes the memory usage of the application program, including the stack, code segment, data segment, and heap, etc.

[0056] Register status: It includes the values of the architecture registers of the processor core when the program crashes, such as the program pointer, stack pointer, etc.

[0057] Stack information: It includes the stack pointer and function call stack information, which helps to locate the function call chain when the program crashes.

[0058] Memory management information: It involves the memory allocation and usage of the program, which helps to check problems such as memory leaks.

[0059] Other processor and operating system status information: In addition to memory information, there are also some key program running statuses.

[0060] The firmware context mainly exists in the registers inside the processor core, the memory corresponding to the processor core, and also includes the content of the inter-core queue, SQCQ and / or AdminQ defined by the NVMe protocol. The status of the flash channel controller can also be used as the firmware context. The flash channel controller is used to operate the NAND flash.

[0061] Outside the control component, there are a backup power supply, DRAM, and multiple NAND flashes connected.

[0062] The backup power supply is used to provide short-term power during abnormal power-off, enabling the control component to execute the processing flow for handling abnormal power-off. It can supply power for 10 - 20 ms.

[0063] The NAND flash memory is a "persistent storage medium". The storage space it provides is divided into two parts. One part stores user data (related to IO commands), and the other part stores system data. A space is reserved in the space for storing system data as the "persistent storage medium space" of the storage context sub-block of the present invention. This is further divided into multiple regions. The context sub-blocks are to be written into one or some of these regions.

[0064] The main function of the backup power supply is to write the user data cached in the memory SRAM or DRAM into the NAND flash memory in case of abnormal power-off. And the storage context sub-block is an additional requirement for it. Therefore: (1) The backup power supply is unreliable for saving the context sub-blocks, so the process of saving the context sub-blocks needs to be optimized, including multi-core collaborative work, avoiding affecting the normal abnormal power-off handling process, and preferentially saving the more important context sub-blocks, etc.; (2) What the backup power supply is used to save is not only the context sub-blocks but also other data.

[0065] Each processor core generates several context sub-blocks respectively. Here, it is further divided into context sub-blocks that can only be generated by the processor core itself and context sub-blocks that can be generated by any processor core. For the internal registers of the processor core and the memory it exclusively occupies, only the processor core itself knows this information, and this part of the context sub-blocks can only be generated by itself. And for contents such as SQ / CQ, any processor core can obtain and generate the corresponding context sub-blocks.

[0066] The method for saving the firmware context for abnormal power-off proposed in the embodiments of the present disclosure detects abnormal power-off; after a predetermined time of detecting abnormal power-off, the control core among multiple processor cores detects the power-off completion flag. If the power-off completion flag does not exist, the control core sends an interrupt to the multiple processor cores; each of the multiple processor cores generates context sub-blocks based on the corresponding data sub-blocks obtained from the interrupt, obtaining multiple context sub-blocks; the control core stores some or all of the multiple context sub-blocks into the target region of the persistent storage medium space based on the sorting result of sorting the multiple context sub-blocks according to the possibility of the occurrence of the exception. By adopting the above technical solution, it is possible to save the critical context in an extremely short time in case of the failure of the abnormal power-off process, and be externally obtained after the next power-on to restore the context for maximum debugging, thereby greatly improving the debugging efficiency.

[0067] Figure 3 It is a flowchart of a method for saving the firmware context for abnormal power-off provided by the embodiments of the present disclosure. The method... as Figure 3 shown, this method is applied to a storage device and includes:

[0068] Step 201, detect abnormal power-off.

[0069] Step 202: After a predetermined time of detecting abnormal power-off by the control core among multiple processor cores, the control core detects a power-off completion flag. If the power-off completion flag does not exist, the control core sends an interrupt to the multiple processor cores.

[0070] Specifically, the storage device is connected to the host. Under normal circumstances, the user shuts down the host through a key press or operating system operation. During the host shutdown process, the storage device is notified to power off, and the host provides sufficient power during the power-off period of the storage device, so that the storage device has enough time to execute the power-off processing flow.

[0071] In the embodiments of the present disclosure, abnormal power-off refers to an event in which the storage device loses power without the user performing a normal shutdown operation such as pressing a shutdown button.

[0072] In some embodiments, a preset time greater than the average duration is set, where the average duration is determined according to multiple historical durations for the storage device to complete power-off processing.

[0073] Specifically, under normal circumstances (such as the average duration when the power-off completion flag exists), the preset time is slightly greater than the average duration, so that each processor core can complete the abnormal power-off process. Thus, the power-off completion flag is checked through the preset time. If it does not exist, it means that one or some logic modules have problems. Only then does it make sense to save the firmware scene.

[0074] It should be noted that the preset time can also be selected and set according to actual application needs. For example, a timer with a fixed time is set, and when the timing time arrives, the power-off completion flag is checked. The power-off completion flag is used to identify the completion of power-off.

[0075] In the embodiments of the present disclosure, the firmware includes multiple logic modules running on multiple processor cores. The abnormal power-off processing flow of each logic module itself includes multiple power-off steps. In response to detecting abnormal power-off, each logic module executes its own abnormal power-off processing flow, and records the number of the power-off steps that have been completed during the execution of its own abnormal power-off processing flow. After all the multiple logic modules have completed their own abnormal power-off processing flows, a power-off completion flag is generated. Thus, each logic module in the firmware divides the power-off processing flow into multiple sub-steps, records the numbers of the completed sub-steps in real time, analyzes the possibility of each logic module having an abnormality according to the numbers of the completed sub-steps of each logic module and the logical dependency relationships between the sub-steps, sorts the logic modules according to the magnitude of the possibility of having an abnormality, and subsequently, the on-site data of the logic modules with a greater possibility of having an abnormality can be preferentially saved, further ensuring the subsequent debugging efficiency and effect.

[0076] The power-down completion flag cannot be detected, that is, the power-down completion flag does not exist, and it is necessary to start the firmware in-situ preservation. That is to say, when abnormal power-down starts, the firmware will start a timer with a fixed time, check the power-down completion flag when the timing time arrives, and if the power-down completion flag does not exist, start the firmware in-situ preservation. It should be noted that the abnormal power-down handling processes of each logic module still continue to run.

[0077] Among them, to start the firmware in-situ preservation, the control core in multiple processor cores can send an interrupt to multiple processor cores (including the control core) to notify multiple processor cores to save the firmware in-situ; among them, one processor core in multiple processor cores can be used as the control core according to actual application needs, and the remaining processor cores perform corresponding operations based on the interrupt sent by the control core.

[0078] Step 103: Each of the multiple processor cores obtains a corresponding data sub-block based on the interrupt to generate an in-situ sub-block, and obtains multiple in-situ sub-blocks.

[0079] Step 104: The control core stores some or all of the multiple in-situ sub-blocks in the target area of the persistent storage medium space based on the sorting result of sorting the multiple in-situ sub-blocks according to the possibility of occurrence of the exception.

[0080] In the embodiment of the present disclosure, after each processor core receives an interrupt, it packages the data sub-block corresponding to each processor core into the corresponding in-situ sub-block, and writes the completion mark information when the packaging of the in-situ sub-block is completed. After the control core generates the in-situ sub-block of its own core, it exits the interrupt. For each core that is not the control core, after generating the in-situ sub-block of its own core, it also assists in processing the unfinished packaging from the data sub-block to the in-situ sub-block. After the control core exits the interrupt, it stores some or all of the multiple in-situ sub-blocks in the target area of the persistent storage medium space based on the sorting result of sorting the multiple in-situ sub-blocks according to the possibility of occurrence of the exception.

[0081] Among them, the firmware includes multiple logic blocks running on multiple processor cores, the power-down completion flag includes the power-down completion flag generated by one or more of the multiple logic blocks, and the possibility of occurrence of an exception of one or more of the multiple logic blocks is judged based on the power-down completion flag generated by one or more of the multiple logic blocks; or, the possibility of occurrence of an exception of one or more of the multiple logic blocks is judged based on the historical exception occurrence situations of one or more of the multiple logic blocks. That is to say, if the progress of the abnormal power-down handling process is slower than expected, there is a greater possibility of problems, so it is necessary to save its firmware in-situ first; or, based on experience or historical data, which logic modules are more likely to have problems during abnormal power-down.

[0082] In some embodiments, the abnormal power-off handling current processes of each logical module in the firmware are sorted according to the process progress, and the abnormal possibility sorting of each logical module in the firmware is set according to the sorting result; and / or, the abnormal possibility of each logical module in the firmware is set according to the firmware context content processed by each logical module. That is to say, the sorting of the possibility of occurrence of an abnormality is related to the order of the logical modules. For example, for a logical module with a slow progress in the abnormal power-off handling process, there is a high possibility of problems, and the corresponding on-site sub-block of this logical module is arranged in the front; the sorting of the possibility of occurrence of an abnormality is related to the order of the firmware context content. For example, the register status and stack information are in the front, and the memory information is in the back.

[0083] In some embodiments, in response to detecting an abnormal power-off, each processing core processes its own abnormal power-off handling process in parallel, and multiple processor cores pause their own abnormal power-off handling processes in response to an interrupt, generate multiple on-site sub-blocks, and then resume processing their own abnormal power-off handling processes.

[0084] Specifically, the firmware includes multiple logical modules running on multiple processor cores. The abnormal power-off handling process of each logical module itself includes multiple power-off steps. In response to detecting an abnormal power-off, each logical module executes its own abnormal power-off handling process, and records the number of the power-off steps that have been completed during the execution of its own abnormal power-off handling process. After all the multiple logical modules have completed their own abnormal power-off handling processes, a power-off completion flag is generated.

[0085] In some embodiments, the firmware context includes multiple logical parts, each logical part is divided into one or more data sub-blocks, and on-site sub-blocks are generated according to the data sub-blocks.

[0086] In some embodiments, the logical parts include one or more of the in-core register status, data in the independent storage area of the core, inter-core shared fast storage, inter-core shared slow storage, stack information, SQ or CQ queues, Admin queues, memory management information, memory information, other status information of the processor, status information of the hardware accelerator, cache content, inter-core communication queues, and inter-core shared memories.

[0087] In some embodiments, as Figure 4 and Figure 5 shown, each on-site sub-block contains:

[0088] Magic number, used to identify the new on-site sub-block;

[0089] Number, the number of the firmware context packaged into the on-site sub-block;

[0090] The size of the on-site sub-block;

[0091] Data sub - block descriptor, which is used to describe the information of data sub - blocks included in the on - site sub - block, including: a plurality of data sub - block descriptor entries, and each data sub - block descriptor entry corresponds to one of the data sub - blocks;

[0092] Among them, each data sub - block descriptor entry includes: the type of the data sub - block, which is used to describe the logical part of the data sub - block from any target, and the logical part includes one or more of the in - core register status, in - core independent storage area data, inter - core shared fast storage, inter - core shared slow storage, stack information, SQ or CQ queue, Admin queue, memory management information, memory information, other status information of the processor, status information of the hardware accelerator, cache content, inter - core communication queue, and inter - core shared memory;

[0093] Log area, including: the valid number of log entries, which is used to indicate the number of valid log entries; the total number of log entries, which is used to indicate the number of log entries, including valid and blank ones; the size of each log entry, and the size of each log entry is the same; log entry area, which stores all log entries in an equal - size manner;

[0094] Data sub - block area, which is used to store a plurality of data sub - blocks, and each data sub - block corresponds to one of the data sub - block descriptor entries.

[0095] In an optional implementation, the on - site sub - block includes more or less content. For example, the on - site sub - block includes a plurality of data sub - block descriptors and corresponding data sub - blocks. The data sub - block descriptors and the data sub - blocks are in one - to - one correspondence. The data sub - block descriptor is used to record the metadata of the corresponding data sub - block. For example, the type of the firmware on - site recorded in the data sub - block. Optionally, the data sub - block descriptor has a specified size and is stored continuously; and the data sub - block also has a specified size and is stored continuously. Thus, it is convenient to obtain a specified data sub - block in the on - site sub - block.

[0096] In some embodiments, when the control core determines that the power - down is not completed, according to the power - down step numbers completed by each of the multiple logical modules, it determines the possibility of an exception occurring for each of the multiple logical modules; sorts them according to the possibility of an exception occurring to obtain a sorting result.

[0097] In some embodiments, after a predetermined time when the control core detects an abnormal power - down, based on the existence of the power - down completion flag, without the need to start the abnormal power - down processing flow, it does not send an interrupt to the multiple processor cores and does not save the firmware on - site.

[0098] In some embodiments, each of the multiple processor cores generates multiple snapshot sub-blocks based on interrupts to obtain corresponding data sub-blocks, including: after each processor core receives an interrupt, it wraps its corresponding data sub-block into the corresponding snapshot sub-block and adds marker information; after the control core generates its own corresponding snapshot sub-block, it exits the interrupt handling process; after other processor cores that are not the control core generate their own corresponding snapshot sub-blocks, they generate snapshot sub-blocks including marker information for other data sub-blocks according to the sorting result and exit the interrupt handling process.

[0099] Specifically, after each processor core receives an interrupt, it wraps data sub-blocks such as its own registers and on-core independent storage areas into the corresponding snapshot sub-block and writes a completion marker when the snapshot sub-block is completed. Except for the control core, after each processor core finishes its own snapshot sub-block, it saves the remaining unprocessed snapshot sub-blocks in the order of the sorting result, with the ones ranked higher (i.e., the most important remaining ones) being saved first, and writes a completion marker when the snapshot sub-block is completed until all snapshot sub-blocks are completed and the interrupt is exited. After the control core finishes its own snapshot sub-block (including writing the completion marker), it directly exits the interrupt.

[0100] In the embodiments of the present disclosure, a target area can be determined in advance in the persistent storage medium space, and the snapshot sub-blocks are stored in the target area in order of the sorting result of the exception probabilities of the respective snapshot sub-blocks, for example, from the largest to the smallest exception probability.

[0101] In the embodiments of the present disclosure, a continuous space of a fixed size in the persistent storage medium space is obtained. The continuous space of the fixed size is divided into multiple regions of the fixed size, and a target area is determined from the multiple regions of the fixed size. The target area is used to record the snapshot sub-blocks written into the persistent storage medium space. The target serial number corresponding to the target area is obtained, and the target serial number is stored in the target position of the persistent storage medium space when the abnormal power-off process is completed, so that when the storage device is powered on next time, it can obtain the target serial number from the target position and then find the saved firmware snapshot.

[0102] Specifically, the firmware selects a continuous space of a fixed size from the persistent storage medium space, divides this space into multiple regions of the fixed size, selects an empty region from them, records the serial number of the selected region as the target serial number, which is used as the "dirty region serial number". When the abnormal power-off process is completed normally, the "dirty region serial number" will be saved in the target position of the persistent storage medium.

[0103] It should be understood that during the abnormal power-off handling process, the backup power supply cannot always ensure that the target serial number is written to the target position. Therefore, when the storage device is powered on, it searches for the previously saved firmware context in various ways. Reading the 'dirty area serial number' at the target position to obtain the saved firmware context is one of the ways. If the 'dirty area serial number' fails to be successfully saved, it is possible to traverse each area of a continuous space of a fixed size in the persistent storage medium space, and the position of the target data block should be recorded in each area to search for data that conforms to the characteristics of the target data block, so as to identify whether the firmware context is stored in this area.

[0104] In the embodiments of the present disclosure, the control core stores some or all of the multiple context sub-blocks in the target area of the persistent storage medium space based on the sorting result of sorting the multiple context sub-blocks according to the likelihood of an exception occurring, including: dividing each context sub-block into small data blocks of a fixed size, and adding description information to each small data block to obtain target data blocks, and sequentially storing all the target data blocks in the target area according to the sorting result; wherein, after each target data block is stored, the power-off completion flag is detected, and based on the power-off completion flag, it is determined that when the power-off is completed, the operation of storing some or all of the multiple context sub-blocks in the target area of the persistent storage medium space is terminated, and when it is determined based on the power-off completion flag that the power-off is not completed, the storage operation of storing some or all of the multiple context sub-blocks in the target area of the persistent storage medium space is continued.

[0105] Specifically, after the control core exits the interrupt, it serially processes multiple context sub-blocks according to the sorting result: dividing the aforementioned selected target area into multiple sub-areas of a fixed size, and each sub-area is divided into multiple blocks of a fixed size; waiting for the context sub-block to be saved to the buffer to be completed; dividing the context sub-block into small data blocks of a fixed size, and adding a section of description information to each small data block, the description information includes the identifier and serial number of the data block, and the size of "small data block + data description" is equal to the size of the aforementioned fixed size sub-area; sequentially storing all "small data block + data description" into the aforementioned selected area one by one; after each small data block is stored, the power-off completion flag is checked, and if the power-off has been completed, the process is terminated, and the remaining small data blocks are not saved, nor are the remaining other context sub-blocks processed.

[0106] In this application, in order to save key content in the shortest time, first, the firmware context is divided according to the logical relationship, and then the divided logical parts are segmented in the form of restricting the data size to form data sub-blocks; then the data sub-blocks are supplemented with necessary information and packaged into numbered context sub-blocks and placed in the buffer. After that, each context sub-block is stored in the persistent storage medium in order from the highest to the lowest importance. In this way, even if the internal power supply of the storage device ends prematurely for a short time, the most important part can be retained as much as possible for subsequent investigation and analysis.

[0107] It should be noted that if there are multiple data sub - blocks in a certain on - site sub - block, they can also be stored from front to back according to the degree of importance. For example, the most important register is placed first, and so on. In this way, even if this on - site sub - block is not completely saved, important debugging information can be recovered as much as possible from one or more successfully saved data sub - blocks.

[0108] Figure 6 This is a flowchart example of firmware on - site saving provided by an embodiment of the present disclosure.

[0109] Step 4.1: After detecting abnormal power - off, start the timer. The time period can be set according to the average duration for each processor core to complete the abnormal power - off process, and the timer is set according to the time period.

[0110] Step 4.2: Determine whether the timer reaches the timing time; if not, continue to execute Step 4.2; if so, execute Step 4.3.

[0111] Step 4.3: Determine whether the power - off process has been completed. Specifically, the control core among multiple processor cores detects the power - off completion flag. If the power - off completion flag does not exist, it means the power - off process has not been completed, and execute Step 4.4.

[0112] Step 4.4: The control core selects an empty area from the holding medium storage space, stores the area serial number into the 'dirty area serial number'. Specifically, the control core obtains a continuous space of a fixed size in the persistent storage medium space, divides the continuous space of the fixed size into multiple areas of the fixed size, determines the target area from multiple areas of the fixed size, obtains the target serial number corresponding to the target area as the 'dirty area serial number', and stores the target serial number to the target position of the persistent storage medium space when the abnormal power - off process is completed.

[0113] Step 4.5: The control core among multiple processor cores sorts the on - site sub - blocks according to the degree of importance. Specifically, when the control core determines that the power - off is not completed, according to the power - off step numbers completed by each of the multiple logic modules, determine the possibility of an exception occurring for each of the multiple logic modules, sort according to the possibility of an exception occurring to obtain a sorting result, and sort the multiple on - site sub - blocks according to the degree of importance according to the sorting result.

[0114] Step 4.6: The control core issues an interrupt to multiple processor cores (including itself).

[0115] In response to the interrupt, each processor core executes a part of Steps 4.7 to 4.12 respectively.

[0116] Step 4.7 Each processor core wraps its own data sub - blocks such as registers and on - core independent storage into corresponding context sub - blocks. Specifically, after each processor core receives an interrupt, it wraps its corresponding data sub - blocks into the corresponding context sub - blocks and adds marker information.

[0117] Step 4.8 Determine whether the processor core executing the current processing flow itself is the control core; if not, execute Step 4.9.

[0118] Step 4.9 Determine whether there are still unfinished context sub - blocks; if not, execute Step 4.10 to exit the interrupt; if there are unfinished context sub - blocks, execute Step 4.11.

[0119] Step 4.11 Process the context sub - blocks with a higher sort order. After other processor cores that are not the control core generate their own corresponding context sub - blocks, they generate context sub - blocks including marker information for other data sub - blocks according to the sorting result. After these processor cores generate context sub - blocks, they also return to Step 4.9 to determine whether there are still context sub - blocks to be generated.

[0120] In Step 4.8, if it is the control core itself, execute Step 4.12 to exit the interrupt; specifically, after the control core generates its own corresponding context sub - block, it exits the interrupt handling process and continues to execute Steps 4.13 - 4.17. Steps 4.13 - 4.17 are executed by the control core, and other processor cores do not execute these steps.

[0121] Step 4.13 The control core determines whether there are context sub - blocks that have not been saved to the persistent storage medium; if so, execute Step 4.14. Otherwise, the abnormal power - off handling process ends. Optionally, if there are no longer context sub - blocks that have not been saved to the persistent storage medium, the target sequence number is also written to the target location here, or the target sequence number is written to the target location through Step 4.20.

[0122] Step 4.14 The control core determines whether the context sub - blocks with a higher sort order that have not been saved are completed; if not, continue to execute Step 4.14 to wait for the context sub - blocks with a higher sort order to appear; if so, execute Step 4.15. It should be understood that in Step 4.14, the control core loops to execute Step 4.14 to wait for the context sub - blocks with a higher sort order to appear. At this time, other processor cores may be executing Step 4.11 to generate context sub - blocks. After the control core waits for the context sub - blocks with a higher sort order to appear in Step 4.14, it turns to Step 4.15 and subsequent steps to write the context sub - blocks with a higher sort order to the persistent storage medium. It still needs to be understood that there can be multiple context sub - blocks with a higher sort order. The control core does not need to wait for all context sub - blocks with a higher sort order to be completed in Step 4.14, but after a context sub - block with a higher sort order appears, it writes these context sub - blocks to the persistent storage medium through Step 4.15 and subsequent steps.

[0123] Step 4.15: The control core divides the on-site sub-blocks into small data blocks; Step 4.16: Explanatory information is added to the small data blocks and saved to the area corresponding to the 'dirty area serial number'; Step 4.17: Determine whether there are still unsaved small data blocks; if so, execute Step 4.16; if not, execute Step 4.13. It should be understood that returning to Step 4.13 here is because other processor cores outside the control core may still be executing Step 4.11 and generating further on-site sub-blocks. These newly emerged on-site sub-blocks are processed by returning to Step 4.13.

[0124] In addition, after Step 4.1, it further includes Step 4.18: Each logic module starts to process the abnormal power-off process; Step 4.19: The control core determines whether the power-off process is completed normally; if so, execute Step 4.20 to save the 'dirty area serial number' to the persistent storage medium.

[0125] It can be understood that taking the failure to complete the power-off process within the specified time as the condition for triggering the firmware on-site preservation. During the preservation of the firmware on-site, the firmware still continues the normal power-off process, and the action of preserving the on-site does not affect the normal logic operation of the firmware. Once the power-off process is completed normally, the action of preserving the firmware on-site will be aborted, and the space of the permanent storage medium it uses can also be recycled and continue to be used, further improving the flexibility of preserving the firmware on-site for abnormal power-off.

[0126] The scheme for preserving the firmware on-site for abnormal power-off provided by the embodiments of the present disclosure detects abnormal power-off. After a predetermined time of detecting abnormal power-off, the control core in multiple processor cores detects the power-off completion flag. If the power-off completion flag does not exist, the control core sends an interrupt to multiple processor cores, and each of the multiple processor cores obtains the corresponding data sub-blocks based on the interrupt to generate on-site sub-blocks, obtaining multiple on-site sub-blocks. The control core stores some or all of the multiple on-site sub-blocks to the target area of the persistent storage medium space based on the sorting result of sorting the multiple on-site sub-blocks according to the possibility of the occurrence of the abnormality. By adopting the above technical solution, when the device has a power-off abnormality, the key information of the firmware on-site can be retained, so that the problem can be quickly located according to the saved on-site data, greatly improving the debugging efficiency.

[0127] It can also be understood that after a power failure and then power-on again, the firmware can identify whether the previous power failure process ended normally; if the previous power failure ended normally, the firmware reads the "dirty area number" from the target location of the permanent storage medium. If the number is a legal area number, it indicates that part or all of the on-site sub-block data may have been written to this area. The firmware erases this area, and after the erasure is completed, this area can continue to be used during operation; if the previous power failure did not end normally, the host will sense the loss of the storage device or the storage device being in an abnormal mode. At this time, for the purpose of debugging the storage device, the on-site sub-block data can be read to restore the abnormal scene.

[0128] In some embodiments, after a power failure, the method further includes: when detecting that the storage device is powered on, if it is recognized that the abnormal power failure process did not end normally when the storage device was powered off last time, query from a fixed-size continuous space to obtain the firmware scene saved when the storage device was powered off last time, and set the target index corresponding to the firmware scene, so that the host can read the firmware scene based on the target index.

[0129] Thus, the present application realizes that after an abnormal power failure occurs in the storage device, the most critical firmware scene information for debugging can be saved in a short time. Moreover, even if the firmware scene is not completely saved, debugging and analysis can still be carried out, so that when a power failure anomaly occurs, the key information of the firmware scene can be retained, and the problem can be quickly located according to the saved on-site data, greatly improving the debugging efficiency of the storage device.

[0130] As Figure 7 shown, the host is connected to the storage device, including:

[0131] When the storage device is powered on, if it is recognized that the abnormal power failure handling process did not end normally when the storage device was powered off last time, based on multiple fixed-size areas in the fixed-size continuous space of the persistent storage medium, read in sequence the data blocks corresponding to the first fixed-size sub-areas of one, part or all of the multiple fixed-size areas; if a legal data block cannot be read from the first fixed-size sub-area of the current fixed-size area, read the next fixed-size area, and if a legal data block is read from the first fixed-size sub-area of the current fixed-size area, continue to read the current fixed-size area to obtain all the data blocks constituting the firmware scene; determine the firmware scene length based on the total sum of the number of all data blocks.

[0132] A host for reading a Telemetry Log from a storage device, where the Telemetry Log includes firmware in-situ data; the storage device, in response to the host reading the Telemetry log, determines the size of the Telemetry Log based on the firmware in-situ length, and determines one, multiple, or all of all data blocks as the content of the Telemetry Log; unpacks and parses the firmware in-situ data to generate a packaging configuration file and a data sub-block file, and obtains a core dump file based on the packaging configuration file and the data sub-block file.

[0133] Figure 8 Shows the processing flow of a method for obtaining the firmware in-situ of abnormal power-off, including:

[0134] Step 301: When the storage device is powered on, if it is recognized that the abnormal power-off processing flow did not end normally during the previous power-off of the storage device, based on multiple fixed-size regions in a fixed-size continuous space in the persistent storage medium, read the data blocks corresponding to the first fixed-size sub-regions of one, part, or all of the multiple fixed-size regions in sequence.

[0135] Step 302: If a legal data block cannot be read from the first fixed-size sub-region of the current fixed-size region, read the next fixed-size region, and if a legal data block is read from the first fixed-size sub-region of the current fixed-size region, continue to read the current fixed-size region to obtain all data blocks that make up the firmware in-situ.

[0136] Step 303: Determine the firmware in-situ length based on the total number of all data blocks.

[0137] Step 304: In response to the host reading the Telemetry log, determine the size of the Telemetry Log based on the firmware in-situ length, and determine one, multiple, or all of all data blocks as the content of the Telemetry Log.

[0138] In some embodiments, when reading each data block, determine whether the data block is legal based on the data description of each data block.

[0139] It should be noted that when reading multiple fixed-size regions in sequence, if a legal data block is read from the first fixed-size sub-region of the current fixed-size region, there is no need to read data from other fixed-size regions; and if a valid data block has been read from the fixed-size sub-region of the current fixed-size region, only read data blocks from the current fixed-size region and no longer read other fixed-size regions, further improving the processing efficiency.

[0140] In some embodiments, if it is recognized that the abnormal power-off handling process ended normally when the storage device was powered off last time, the target serial number is read from the target address in the persistent storage medium space. If the target serial number is a legal area serial number, the target serial number is deleted, and the fixed-size area corresponding to the target serial number is erased.

[0141] In the embodiments of the present disclosure, each fixed-size area corresponds to multiple fixed-size sub-areas, and each fixed-size sub-area includes multiple fixed-size data blocks. When the firmware context was written into the selected fixed-size area during the last power-off of the storage device, the context sub-blocks were split into multiple small data blocks, and a data description was added to each small data block. The size of each data block in the fixed-size area is equal to the combined size of the small data block and the data description.

[0142] In the embodiments of the present disclosure, the data description includes a data block identifier and a data block serial number. The data block identifier is used to indicate that the content of the small data block is the backed-up firmware context, and the data block serial number is used to indicate the serial number of the small data block in the current context sub-block or in all context sub-blocks.

[0143] It can be understood that based on the description of the foregoing Figures 1 to 6 description of the embodiment of saving the firmware context, which "fixed-size area" is selected is determined by the control core and is uncertain. The control core will write the selection result as the target serial number, that is, the "dirty area serial number" at a specific target location. When the storage device is powered on again, there may or may not be a real "dirty area serial number" at the specific target location. That is to say, the previous power-off may not be an abnormal power-off, so the control core did not write the "dirty area serial number", or the previous power-off may be during the abnormal power-off handling process, and the control core did not have time to successfully write the "dirty area serial number". Therefore, it is necessary to identify the target serial number during the power-on process.

[0144] Exemplarily, as Figure 9A shown, the multiple fixed-size areas corresponding to the fixed-size continuous space in the persistent storage medium space include a fixed-size area (serial number 0), a fixed-size area (serial number 1), a fixed-size area (serial number 2), and a fixed-size area (serial number 3), and the target serial number, that is, the "dirty area serial number", is stored at the target address in the persistent storage medium space. As an example, "dirty area serial number (=1)" means that the fixed-size area (serial number 1) is the target area. The target area is used to record the context sub-blocks written into the persistent storage medium space. The target serial number corresponding to the target area is obtained, and when the abnormal power-off process is completed, the target serial number is stored at the target location in the persistent storage medium space, so that the storage device can obtain the target serial number from the target location when it is powered on next time, and then find the saved firmware context.

[0145] Specifically, the fixed-size area is further divided into multiple fixed-size sub-areas, and each fixed-size sub-area includes multiple fixed-size data blocks. Exemplarily, as Figure 9B shown, the fixed-size area includes four fixed-size sub-areas, and each fixed-size sub-area includes multiple fixed-size data blocks.

[0146] It can be understood that when the control core writes the firmware live into the selected fixed-size area, the live sub-blocks need to be split into multiple small data blocks, and a data description is added to each small data block to form a data block. Thus, when the storage device is powered on again, it can identify whether the stored content is the backed-up firmware live and reconstruct the live sub-blocks.

[0147] Based on the description of the foregoing embodiments, the live sub-block includes a "number" field, and the "data block descriptor" records the type of the data sub-block. Among them, the "number" field, for example, represents the number of the processor core corresponding to the live sub-block.

[0148] Exemplarily, as Figure 9C shown, for the structure of the live sub-block 0, the live sub-block is split into multiple small data blocks, that is, Figure 9C the multiple "small data blocks" shown, and a "data description" is added to each "small data block", that is, Figure 9C the data block identifier and data block sequence number shown. The identifier is used to indicate that the content of the data block is the backed-up firmware live, and the sequence number represents its sequence number in the current live sub-block or the sequence number in all the live sub-blocks written by the control core; the multiple fixed-size data blocks included in each fixed-size sub-area can be continuous or discontinuous. For example, in the example after writing one or more live sub-blocks into the fixed-size sub-area, the red "data blocks" record the "small data blocks" obtained from the live sub-blocks with the "data description" added. The non-red "data blocks" can be blank data blocks or other data except the live sub-blocks; as Figure 9C shown, the red "data blocks" in the left fixed-size sub-area are continuous, and the non-red "data blocks" in the right fixed-size sub-area are discontinuous (because while the control core writes the live sub-blocks into the fixed-size sub-area, other cores may still perform the power-off process and may generate data written into the fixed-size sub-area).

[0149] To read the saved live sub-block, if the last power failure did not end properly, the firmware scans all fixed-size areas of the fixed-size persistent storage medium space selected in the save process to search for the first live sub-block data. That is, the firmware reads all fixed-size areas of this space one by one, and for each fixed-size area, only reads its first fixed-size sub-area. For the first fixed-size sub-area, reads all its "data blocks" one by one; after reading each "data block", checks the data description in the "data block". If the data description conforms to the characteristics of the live sub-block data, it means the required data block (also called a legal data block or valid data block) has been found; starts reading the subsequent "data blocks" from the address of the first found data block; after reading each "data block", checks the data description. If it is a legal data block, the number of valid data blocks is incremented by 1; continues until all fixed-size sub-areas have been read; calculates the firmware live length using the total sum of the number of valid data blocks. All valid data blocks constitute the firmware live saved during the last abnormal power failure of the storage device. Optionally, the saved firmware live is stored in one or more fixed-size sub-areas, rather than all fixed-size sub-areas. For example, it is recognized that the subsequent data blocks will not store the firmware live based on the data blocks in the fixed-size sub-areas that have not been written with data. Optionally, along with the "dirty area number", the size of the firmware live is also recorded to identify the number of data blocks that record the firmware live.

[0150] If no live sub-block data is found after reading all of the first fixed-size sub-area, read the next fixed-size area. If no live sub-block data is still found after reading all fixed-size areas, it indicates a scan failure and there is no firmware live.

[0151] If there is a firmware live, add a telemetry log index. The first telemetry log area, telemetry log area 1, points to the address of the first data block where this firmware live is located, and the second Telemetry log area, telemetry log area 2, and the third Telemetry log area, telemetry log area 3, are set to empty.

[0152] It can be obtained from https: / / nvmexpress.org / wp-content / uploads / NVM_Express_Revision_ 1.3.pdf The telemetry log is defined in the NVMe 1.3 standard manual obtained. For its detailed description and usage method, please refer to https: / / zhuanlan.zhihu.com / p / 399501400 .

[0153] The host can obtain the previously saved firmware snapshot by reading the telemetry log. Specifically, the host obtains the telemetry log as follows: when the host reads the telemetry log header, the storage device responds to the host by setting the Last Block of area1, area 2, and area 3 of the telemetry log to the length of the previously obtained firmware snapshot plus the length of the telemetry log header; when the host reads the telemetry log area 1, the firmware starts from the address of the first data block where the firmware snapshot is located, reads the data, and transfers the read data to the host side; since the Last Block of area1, area 2, and area 3 are equal, the host side does not read the telemetry log area 2 and the telemetry log area 3; the host forms the entire telemetry log into a continuous file, which is called the snapshot file.

[0154] It should be understood that the operation of the host reading the telemetry log may not occur, and the timing of its occurrence is determined by the host. To respond to the potential operation of the host reading the telemetry log, the storage device obtains the positions of multiple data blocks recording the firmware snapshot and the length of the firmware snapshot in one or more fixed-size areas when powering on. These information are used to provide responses when the host reads the telemetry log. Before the host reads the telemetry log, the storage device does not need to read out the data blocks representing the firmware snapshot.

[0155] Figure 10 Shows the processing flow after the storage device powers on, including:

[0156] Step 10.1 The storage device powers on after abnormal power-off.

[0157] Step 10.2 Determine whether the previous abnormal power-off processing flow of the storage device was completed normally; if it was completed normally, execute Step 10.3. At this time, it is known that the previous abnormal power-off processing of the storage device was successful, so there is no need to obtain the firmware snapshot to analyze the fault. However, some firmware snapshots may have been generated during the previous abnormal power-off processing based on Figure 6 the processing method, and these firmware snapshots need to be deleted to avoid confusion with other firmware snapshots generated during subsequent abnormal power-off processing.

[0158] Step 10.3 Determine whether the "dirty area serial number" is valid. It can be judged whether the target serial numbers of each fixed-size area are valid according to the number of fixed-size areas. For example, if the number of fixed-size areas is 4, the "dirty area serial number" of 1 is considered valid, and the "dirty area serial number" of 5 or more is invalid; if so, execute Step 10.4.

[0159] Step 10.4 Erase the dirty area corresponding to the "dirty area serial number", that is, the fixed-size area corresponding to the target serial number. The erased fixed-size area can be used continuously, so as to ensure that the firmware scene stored in the dirty area corresponding to the "dirty area serial number" is the firmware scene stored during the last abnormal power failure.

[0160] Step 10.5 The storage device is used normally.

[0161] Figure 11 Another processing flow after the storage device is powered on is shown, including:

[0162] Step 11.1 The storage device is powered on after an abnormal power failure.

[0163] Step 11.2 Determine whether the last abnormal power failure processing flow of the storage device is completed normally; if the abnormal power failure processing flow is not completed normally, the on-site data can be obtained for debugging, and Step 11.3 is executed. Optionally, Step 11.2 and Figure 10 Step 10.2 of can be implemented in the same processing flow. For example, if it is judged as no in Step 10.2, it will turn to Step 11.3 and continue to execute.

[0164] Step 11.3 Set the fixed-size area to be read as the first fixed-size area, and set the data block to be read as the first data block.

[0165] Step 11.4 Read the data block to be read in the fixed-size area to be read.

[0166] Step 11.5 Determine whether the read data block is on-site sub-block data; if so, execute Step 11.6; if the read data is not on-site sub-block data, execute Step 11.10.

[0167] Step 11.6 Increment the number of read data blocks by 1.

[0168] Step 11.7 Determine whether the fixed-size area to be read has been completely read or the number of read data blocks has reached the expectation; if not, execute Step 11.8; if so, execute Step 11.9.

[0169] Step 11.8 Set the data block to be read as the next data block and then execute Step 11.4.

[0170] Step 11.9 Add a telemetry log. Point its area 1 to the first data block found, and set the Last Block of area 1 to the length of the previously obtained firmware snapshot plus the length of the telemetry log header.

[0171] Step 11.10 Determine whether the first sub-region has been completely read; if so, execute Step 11.11; if not, execute Step 11.12. Set the block to be read as the next data block and then execute Step 11.4.

[0172] Step 11.11 Determine whether all regions have been completely read; if so, execute Step 11.9; if not, execute Step 11.13. Set the fixed-size region to be read as the next fixed-size region, set the data block to be read as the first data block of the next fixed-size region, and then execute Step 11.4.

[0173] Specifically, after the storage device is powered on, if it is determined that the last abnormal power-off processing procedure of the storage device has not been completed normally, the storage device needs to obtain the firmware snapshot. For example, based on the "dirty region number", determine the fixed-size region that stores the firmware snapshot, and identify the data block that records the firmware snapshot from it. However, the "dirty region number" may be invalid. For example, the recording of the "dirty region number" was not completed during the last abnormal power-off processing procedure. At this time, it is necessary to traverse the fixed-size regions to identify whether they store the firmware snapshot. More specifically, read the first fixed-size sub-region of each fixed-size region, and read each "data block" in the first fixed-size sub-region. If the "data block" read records the firmware snapshot, it is considered that the fixed-size sub-region where it is located stores the firmware snapshot; if no "data block" in the first fixed-size sub-region stores the firmware snapshot, it is considered that this fixed-size region does not store the firmware snapshot, and do not read its other fixed-size sub-regions, but traverse another fixed-size region. For each "data block" in the fixed-size sub-region, identify whether it stores the firmware snapshot. More specifically, read out the "data block" and check whether the position corresponding to the data description of the "data block" records an identifier. Use the existence of the identifier representing the firmware snapshot as the judgment basis; for the block without the identifier representing the firmware snapshot, it may be other types of data (it may be other data backed up during power-off) or illegal data (indicating that there is no other valid data behind, and at this time, the subsequent blocks may not be read anymore).

[0174] It should be noted that all the on-site sub-blocks backed up to the persistent storage medium space during the last abnormal power failure are found from the fixed-size area. Among them, all the on-site sub-blocks refer to the backed-up on-site sub-blocks, rather than all the generated on-site sub-blocks, because all the on-site sub-blocks may not have had time to be backed up; moreover, some on-site sub-blocks may not have been fully backed up, but only part of the small data blocks have been recorded.

[0175] Specifically, based on the small data blocks representing all the on-site sub-blocks backed up to the persistent storage medium space during the last abnormal power failure, the firmware on-site length is obtained; among them, there are many ways to determine the firmware on-site length, for example Figure 9C In the case of the left fixed-size sub-region in, the address of the first red "data block" and the address of the last red "data block" can be recorded to determine the firmware on-site length; for another example Figure 9C In the case of the right fixed-size sub-region in, the address of each red "data block" is recorded to determine the firmware on-site length, further improving the flexibility of determining the firmware on-site length.

[0176] It can be understood that the firmware of the storage device will record the firmware on-site length, and when receiving a host read telemetry log subsequently, the recorded firmware on-site length is used to respond to the host, and the response to the host also includes the starting address of the telemetry log representing the firmware on-site; for example, Figure 9C The address of the first red "data block" on the right in is used as the starting address of the telemetry log; for another example, a virtual address mapped to Figure 9C The first red "data block" on the right in is used to respond to the host, so as to prevent the host from knowing the internal physical address of the storage device and further improving the security of the storage device.

[0177] Figure 12 FIG. is a schematic flow chart of a method for generating a core dump file using the firmware on-site of abnormal power failure provided by an embodiment of the present disclosure. The method. As Figure 12 shown, this method is applied to a host and includes:

[0178] Step 401, obtain a Telemetry Log recording the firmware on-site from a storage device.

[0179] Specifically, the host side can splice the data received from the storage device into a on-site file by reading the telemetry log.

[0180] When the host reads the firmware on-site, the firmware on-site is obtained. As Figure 13 shown, all the red "data blocks" are the firmware on-site.

[0181] Exemplarily, as Figure 14As shown, the host obtains the starting address (e.g., Area 1) of the firmware scene in the Telemetry Log and the firmware scene length (e.g., obtained from the LastBlock of Area 1) from the Telemetry Log Header provided by the storage device; and obtains the firmware scene from the Telemetry Log based on the starting address and length. From the perspective of the storage device, it obtains multiple small data blocks that store the firmware scene in the process shown in Figure 11 (also see Figure 13 , where the "data blocks" in red). The data (small data blocks) representing the firmware scene in these data blocks are mapped to Telemetry Log Area 1. In response to the host reading Area 1 of the Telemetry Log, the storage device reads out multiple small data blocks that store the firmware scene and sends them to the host as a response to the host's reading of Area 1 of the Telemetry Log.

[0182] Step 402: Parse and unpack the firmware scene data obtained from the Telemetry Log to generate a packaging configuration file and a data sub-block file.

[0183] In some embodiments, parsing and unpacking the firmware scene data to obtain a packaging configuration file and a data sub-block file includes: checking the magic number for each scene sub-block in the firmware scene data, and reading the number after the magic number matches; parsing the data sub-block descriptor and the data sub-block area; for each data sub-block descriptor entry, reading the storage area type and number of the data sub-block together to determine the data sub-block identifier, reading the offset of the data sub-block in the scene sub-block and the data sub-block size, and reading the offset of the data sub-block in the storage area; generating a data sub-block file based on the data sub-block identifier, the offset of the data sub-block in the scene sub-block, and the data sub-block size; and generating a packaging entry based on the data sub-block identifier, the data sub-block size, and the offset of the data sub-block in the storage area and adding it to the packaging configuration file.

[0184] The host generates multiple scene sub-blocks based on the firmware scene data. The scene sub-blocks are composed of multiple consecutive small data blocks. The data description part of the small data block does not belong to the firmware scene, while the small data block belongs to the firmware scene. For each small data block of the firmware scene data, when it is recognized that a specific position therein includes the "magic number" of the scene sub-block, it is recognized as the first data block of the scene sub-block, and the number of small data blocks of the scene sub-block to which the data block belongs is recognized based on the "size of the scene sub-block" therein, and the specified number of subsequent small data blocks are used to construct the scene sub-block. Repeat this process to obtain all the scene sub-blocks of the firmware scene.

[0185] In the embodiments of the present disclosure, the storage area type of the data sub - block is used to describe which logical part the data sub - block belongs to; where the logical part includes one or more of in - core register status, in - core independent storage area data, inter - core shared fast storage, inter - core shared slow storage, stack information, SQ or CQ queue, Admin queue, memory management information, memory information, other status information of the processor, status information of the hardware accelerator, cache content, inter - core communication queue, and inter - core shared memory.

[0186] In some embodiments, parsing and unpacking the firmware on - site data to obtain an encapsulated configuration file and a data sub - block file includes: obtaining the data sub - block identifier corresponding to all data sub - blocks, the offset of the data sub - block in the on - site sub - block, and the data sub - block size based on the log area of each on - site sub - block in the firmware on - site data; generating an encapsulation entry based on the data sub - block identifier, the data sub - block size, and the offset of the data sub - block in the storage area and adding it to the encapsulated configuration file.

[0187] It can be understood that after the host reads the on - site file, it parses, unpacks, and encapsulates the Coredump file.

[0188] As an example, as Figure 15 shown, it includes:

[0189] Step 15.1: For the on - site file, starting from the beginning of the telemetry log area 1 as the first on - site sub - block, start parsing. If the check fails or there is insufficient remaining data in the file, etc., the parsing and unpacking immediately stop and directly enter the next stage of encapsulating the Coredump file. At this time, the data sub - block file, encapsulation entries, and log files that have been normally parsed can still be normally encapsulated into the Coredump file.

[0190] Specifically, for each on - site sub - block, with the Figure 2 and Figure 3 shown data structure: check the magic number. Only if the check passes is it a legal on - site sub - block, otherwise an error occurs; read the size of the on - site sub - block. If it is greater than the remaining data size of the current file, then there is a next on - site sub - block; otherwise, the current on - site sub - block is the last one; the remaining data size of the current file being less than the size of the on - site sub - block does not represent an error and parsing can continue to be attempted.

[0191] Step 15.2: Determine whether there is still an on - site sub - block; if so, execute Step 15.3.

[0192] Step 15.3: Determine whether the magic numbers match; if so, execute Step 15.4.

[0193] Step 15.4 reads the size of the on-site sub-block, determines whether there is a next on-site sub-block and the location of the next on-site. Step 15.5.

[0194] Step 15.5 reads the number.

[0195] Step 15.6 reads the valid number of data sub-block descriptor entries, the total number of data sub-block descriptor entries, and the size of the data sub-block descriptor entries.

[0196] Step 15.7 determines whether there are still data sub-block descriptor entries; if so, execute Step 15.8.

[0197] Step 15.8 reads the storage area type of the data sub-block and the number together to determine the identifier of the data sub-block.

[0198] Step 15.9 reads the offset of the data sub-block in the on-site sub-block and the size of the data sub-block, extracts a section of data from the on-site file, saves it as a new data sub-block file, and the file name is the identifier of the data sub-block, obtaining the data sub-block file.

[0199] Step 15.10 reads the offset of the data sub-block in the storage area and the identifier of the data sub-block together to form a packaging entry, adds it to the end of the packaging configuration file, obtaining the packaging configuration file. And returns to Step 15.7 to continue execution.

[0200] Specifically, in Step 15.5, for each on-site sub-block, reads the number; parses the data sub-block descriptor and the data sub-block area: reads the valid number of data sub-block descriptor entries, the total number of data sub-block descriptor entries, and the size of the data sub-block descriptor entries; the product of the total number of data sub-block descriptor entries and the size of the data sub-block descriptor entries determines the starting position of the log area (in the on-site sub-block, the log area is after the data sub-block area); the valid number of data sub-block descriptor entries and the size of the data sub-block descriptor entries determine the traversal method of each subsequent data sub-block descriptor entry.

[0201] Among them, for each data sub-block descriptor entry: reads the storage area type of the data sub-block, and the number obtained from the on-site sub-block before together determine the identifier of the data sub-block; for example, if the storage area type of the data sub-block is the in-core register, combined with number 0, it can be combined into the data sub-block identifier: the in-core register of core 0; and for another example, if the storage area type of the data sub-block is the inter-core shared fast storage combined with number 10, it can be combined into the data sub-block identifier: block A of the inter-core shared fast storage (assuming that it is agreed to use the on-site sub-block with number 10 to save the data sub-block of block A of the inter-core shared fast storage). Reads the offset of the data sub-block in the on-site sub-block and the size of the data sub-block, extracts this section of data from the on-site sub-block, and saves it as a new data sub-block file, and the file name is the identifier of the data sub-block mentioned above.

[0202] The data sub - block file participates in the subsequent encapsulation of the Coredump file. Reading the offset of the data sub - block in the storage area, together with the identifier and size of the data sub - block, constitutes an encapsulation entry, which is added to the end of the encapsulation configuration file. The encapsulation configuration file participates in the subsequent encapsulation of the Coredump file.

[0203] It should be noted that in step 15.7, if there is no data sub - block descriptor entry to process, step 15.11 is executed.

[0204] Step 15.11 reads the valid number of log entries, the total number of log entries, and the size of the log entries.

[0205] Step 15.12 determines whether there are still log entries; if so, step 15.13 is executed.

[0206] In step 15.13, each log entry is added to the end of the log file to obtain the log file.

[0207] Specifically, parsing the log area includes: reading the valid number of log entries, the total number of log entries, and the size of the log entries. The valid number of log entries and the size of the log entries determine the traversal method for each subsequent log entry; for each log entry, it is added to the end of the log file. The log file does not participate in the subsequent encapsulation of the Coredump file.

[0208] The host needs to generate a Coredump file based on the read telemetry log (field file or all valid small data blocks). Among them, the Coredump file is a file with a specific format used to record the program field. In the embodiments of the present disclosure, after generating the Coredump file, software tools of the prior art can be used to analyze the firmware field, so as to know the situation when an abnormality occurs in the abnormal power - off handling process, which is convenient for finding the cause of the abnormality.

[0209] In the embodiments of the present disclosure, a data block sub - file and an encapsulation configuration file are generated from the telemetry log, and then a coredump file is generated based on the data block sub - file and the encapsulation configuration file. A file is a specific form of data. In addition to using the data block sub - file and the encapsulation configuration file to record the useful data obtained from the firmware field, other forms of data carriers can also be used. For example, the data is retained in the memory of the host in a specified format.

[0210] In the embodiments of the present disclosure, based on the telemetry log, the host can obtain multiple "data blocks", and each "data block" includes "small data block + data description". The host first reconstructs one or more on-site sub-blocks from the multiple "data blocks". Reconstructing the on-site sub-block can be understood as obtaining the "size of the on-site sub-block" from the front of the on-site sub-block, so as to obtain how many subsequent "data blocks" constitute an on-site sub-block. The on-site sub-block of the telemetry log may be incomplete. For example, it has 100 data sub-blocks with a length of 500, but due to power failure, not all of them are recorded. The data length belonging to this on-site sub-block in the telemetry log is 400, and then these 400 are used to construct this on-site sub-block. There are multiple generated data sub-block files. For example, for each data sub-block in the on-site sub-block, a corresponding data sub-block file is generated.

[0211] In the embodiments of the present disclosure, in the encapsulation configuration file, for each data sub-block file, there is 1 corresponding encapsulation entry. The key information of each encapsulation entry in the encapsulation configuration file lies in the "identifier of the data sub-block". Through the "identifier of the data sub-block", the corresponding data sub-block file can be obtained, and the meaning of the data sub-block can also be obtained.

[0212] It should be noted that traversing all data sub-block files can replace the encapsulation configuration file. Among them, there are multiple data sub-block descriptors and multiple data sub-blocks in the on-site sub-block, and the data sub-block descriptors and data sub-blocks are in one-to-one correspondence.

[0213] Step 403: Generate a core dump file (Coredump) based on the encapsulation configuration file and the data sub-block file.

[0214] In some embodiments, encapsulation is performed based on the encapsulation configuration file and the data sub-block file to obtain a core dump file, including: obtaining the encapsulation entry of the in-core register of the current core and the corresponding data sub-block file from the encapsulation configuration file, and generating corresponding program headers (Program Header) and sections (Sector) according to the format of the core dump file (for example, the http: / / www.skyfree.org / linux / references / ELF_ Format.pdf obtainable ELF file format); obtaining the encapsulation entries of all in-core independent storages, all inter-core shared fast storages, and all inter-core shared slow storages of the current core from the encapsulation configuration file, and the corresponding data sub-block files, and generating corresponding program headers and sections according to the target file format of the core dump type; generating corresponding program headers (Program Header) and sections (Sector) according to the core dump file format.

[0215] Next, based on the multiple groups of program headers and sectors generated previously, generate the ELF header of the ELF file.

[0216] Finally, combine the generated ELF header with multiple groups of program headers and sectors to generate a core dump file.

[0217] After obtaining multiple data sub-block files and encapsulation configuration files in the foregoing embodiments, a coredump file is generated. It can be understood that there are multiple processor cores in the control component, and a corecump file is generated for each processor core. The coredump file of each processor core includes a program header and a sector representing the in-core registers of the current core (the programheader and the sector are fields defined in the ELF file format of the prior art as the coredump file); a program header and a sector representing the in-core independent storage area of the current core; a program header and a sector representing the inter-core shared fast storage of the current core; a program header and a sector representing the inter-core shared slow storage of the current core; an ELF header, which is generated according to the above content. The "identifier of the data sub-block" in the data sub-block file or the encapsulation configuration file provides specific information about its corresponding coredump file. For example, the "identifier of the data sub-block" represents "the in-core registers of processor core 0" and "the inter-core shared fast storage of processor core 1". The main content of the above program header comes from the "identifier of the data sub-block" and the data sub-block descriptor, while the content of the sector comes from the data sub-block.

[0218] Furthermore, the generated coredump file is provided to the debugging software of the prior art (such as GDB), and the debugging software can use the coredump file to analyze the fault scene of the firmware.

[0219] As an example, as Figure 16 shown, it includes:

[0220] Step 16.1 Determine whether there is still a core for which a Coredump file has not been generated; if not, it means that the required coredump file has been generated, and execute Step 16.2 to directly end and wait for subsequent GDB (GNU Debugger) debugging and analysis; if so, execute Step 16.3 to process the next core.

[0221] Step 16.4 Find the encapsulation entry of the kernel registers of the current core from the encapsulation configuration file, find the corresponding data sub-block file, and package them into the corresponding Program Header and Sector.

[0222] Specifically, for each core, find the encapsulation entry of the in-core registers of the current core from the encapsulation configuration file, and find the corresponding data sub-block file by using the identifier of the data sub-block obtained from the encapsulation entry; according to the ELF file format of the Coredump type, package the content of the data sub-block file into the Program Header and Sector for the in-core registers of the current core.

[0223] Specifically, set the p_type field in the program header representing the in-core registers of the CPU core to PT_NOTE; fill in the corresponding register values read from the data sub-block file at the position of each register in the section representing the in-core registers of the CPU core; that is, set the p_type field in the Program Header to PT_NOTE, fill in the other fields of the Program Header according to the actual situation, and fill in the corresponding register values read from the data sub-block file at the position of each register in the Sector.

[0224] Step 16.5 Find the encapsulation entries of all in-core independent storages and the encapsulation conditions of all inter-core shared fast (slow) storages of the current core from the encapsulation configuration file, and then find the corresponding data sub-block files, and package them into the corresponding Program Header and Sector.

[0225] Specifically, set the p_type field in the Program Header to PT_LOAD. Calculate the base address of the storage area type according to the storage area type in the identifier of the data sub-block, and add the offset of the data sub-block in the storage area to obtain the start address of the data sub-block. Set the p_vaddr and p_paddr fields in the Program Header to this start address, where the base address of the storage area type is determined by the storage area type in the identifier of the data sub-block; add the offset of the data sub-block in the storage area to the base address to determine the start address of the data sub-block. Set the p_filesz field and p_memsz field in the Program Header to the data sub-block file size. The Sector is the content read out from the complete data sub-block file.

[0226] Step 16.6 generates a Header in the ELF (Executable and Linkable Format) file format. Specifically, the e_type field in the ELF Header is set to ET_CORE, and other fields in the ELF Header are generated according to the actual situation.

[0227] Step 16.7 combines the ELF Header, all Program Headers, and all Sectors to generate an ELF file, obtaining a Coredump file.

[0228] Combine the ELF Header generated in Step 16.6, all Program Headers and all Sectors generated in Steps 16.4 and 16.5 to generate an ELF file, which is the Coredump file. The ELF file generated by each core can be debugged and analyzed through GDB.

[0229] It should be noted that it is not necessary to have all data sub - blocks when each core encapsulates the Coredump file. If the saved power - off scene is incomplete, a Coredump file can still be encapsulated from the existing data sub - blocks for GDB debugging and analysis.

[0230] Thus, the most critical debugging information can be saved in a short time, and even if the save is incomplete, debugging and analysis can still be carried out, ensuring that the key information of the firmware scene is retained when a power - off exception occurs, and the problem can be quickly located based on the saved on - site data, greatly improving the debugging efficiency.

[0231] Figure 17 The figure is a schematic structural diagram of a device for saving the firmware scene in case of abnormal power - off provided by an embodiment of the present disclosure. This device can be implemented by software and / or hardware and is generally integrated in an electronic device. As Figure 17 shown, the device includes:

[0232] A detection module 501 for detecting abnormal power - off.

[0233] A sending module 502 for the control core among multiple processor cores to detect a power - off completion flag after a predetermined time of detecting abnormal power - off. If the power - off completion flag does not exist, the control core sends an interrupt to the multiple processor cores.

[0234] A generation module 503 for each of the multiple processor cores to obtain corresponding data sub - blocks based on the interrupt to generate on - site sub - blocks, obtaining multiple on - site sub - blocks.

[0235] A storage module 504, configured to store some or all of the multiple live sub - blocks to a target area of a persistent storage medium space by the control core according to a sorting result of sorting the multiple live sub - blocks based on the likelihood of an exception occurrence.

[0236] Figure 18 FIG. is a schematic structural diagram of a device for obtaining a firmware live scene in case of abnormal power - off provided by an embodiment of the present disclosure. This device can be implemented by software and / or hardware and is generally integrated in an electronic device. As Figure 18 shown, when applied to a storage device, the device includes:

[0237] A reading and judging module 601, configured to, when the storage device is powered on, if it is recognized that the abnormal power - off processing flow did not end normally during the previous power - off of the storage device, based on multiple fixed - size areas in a fixed - size continuous space in the persistent storage medium space, sequentially read data blocks corresponding to the first fixed - size sub - areas of each of one, some or all of the multiple fixed - size areas.

[0238] A reading and obtaining module 602, configured to, if a legal data block cannot be read from the first fixed - size sub - area of the current fixed - size area, read the next fixed - size area, and if a legal data block is read from the first fixed - size sub - area of the current fixed - size area, continue to read the current fixed - size area to obtain all data blocks constituting the firmware live scene.

[0239] A determining module 603, configured to determine the firmware live scene length based on the total sum of the number of all data blocks.

[0240] A response module 604, configured to, in response to the host reading a Telemetry log, determine the size of the Telemetry Log based on the firmware live scene length, and use one, multiple or all of the all data blocks as the content of the Telemetry Log.

[0241] Figure 19 FIG. is a schematic structural diagram of a device for generating a core dump file using a firmware live scene in case of abnormal power - off provided by an embodiment of the present disclosure. This device can be implemented by software and / or hardware and is generally integrated in an electronic device. As Figure 19 shown, when applied to a host, the device includes:

[0242] A reading module 701, configured to read a Telemetry Log from the storage device, where the Telemetry Log includes firmware live scene data.

[0243] An analysis module 702, configured to analyze and unpack the firmware live scene data to generate a packaging configuration file and a data sub - block file.

[0244] An encapsulation module 703 is configured to obtain a core dump file based on the encapsulation configuration file and the data sub-block file.

[0245] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a portion of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that, in some alternative implementations, the functions noted in the blocks may occur in a different order than that noted in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented by a dedicated hardware-based system that performs the specified functions or operations, or by a combination of dedicated hardware and computer instructions.

[0246] The units described in the embodiments of the present disclosure can be implemented in software or in hardware. In some cases, the name of the unit does not constitute a limitation on the unit itself.

[0247] The functions described above herein can be performed, at least in part, by one or more hardware logic components. By way of example, and without limitation, exemplary types of hardware logic components that may be used include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on a chip (SOCs), complex programmable logic devices (CPLDs), and the like.

[0248] In the context of the present disclosure, a machine-readable medium may be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0249] The above description is only a preferred embodiment of the present disclosure and an explanation of the applied technical principles. Those skilled in the art should understand that the scope of disclosure involved in the present disclosure is not limited to the technical solutions formed by the specific combination of the above technical features, and should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above disclosure concept. For example, the technical solutions formed by mutually replacing the above features with the technical features (but not limited to) having similar functions disclosed in the present disclosure.

[0250] In addition, although the operations are depicted in a particular order, this should not be construed as requiring that the operations be performed in the particular order shown or in sequential order. In certain environments, multitasking and parallel processing may be advantageous. Similarly, although several specific implementation details are included in the above discussion, these should not be construed as limiting the scope of the present disclosure. Certain features described in the context of separate embodiments may also be implemented combinatorially in a single embodiment. Conversely, the various features described in the context of a single embodiment may also be implemented separately or in any suitable sub-combination in multiple embodiments.

[0251] Although the subject matter has been described in language specific to structural features and / or methodological logical acts, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. On the contrary, the specific features and acts described above are merely example forms for implementing the claims.

Claims

1. A method for obtaining firmware scene of abnormal power failure, characterized in that: Applied to a storage device, the method comprises: When the storage device is powered on, if it is identified that the abnormal power-off processing flow of the storage device when it was powered off last time did not end normally, based on multiple fixed-size areas of fixed-size continuous space in the persistent storage medium space, read the data blocks corresponding to the first fixed-size sub-area of ​​each of one, part or all of the multiple fixed-size areas in sequence; If no legal data block can be read from the first fixed-size sub-region of the current fixed-size region, then read the next fixed-size region; and if a legal data block is read from the first fixed-size sub-region of the current fixed-size region, then continue to read the current fixed-size region to obtain all data blocks constituting the firmware scene; Determine the firmware field length based on the sum of the quantities of all data blocks; In response to the host reading the Telemetry log, the size of the Telemetry Log is determined based on the firmware field length, and one, more than one, or all of the data blocks are used as the content of the Telemetry Log.

2. The method according to claim 1, characterized in that: Each of the fixed-size regions includes a plurality of fixed-size sub-regions; Each of the fixed-size sub-regions includes a plurality of fixed-size data blocks; When the storage device was powered off for the last time so that the firmware scene was written into the selected fixed-size area, the scene sub-block was split into multiple small data blocks, and the data description was added to each of the small data blocks; wherein the size of each of the data blocks in the fixed-size area is equal to the size of the small data block and the data description combined.

3. The method according to claim 1 or 2, characterized in that: Multiple small data blocks form a live sub-block; Each of the on-site sub-blocks includes a data sub-block descriptor and a plurality of data sub-blocks, and the data sub-block descriptor is used to describe information of the data sub-blocks included in the on-site sub-block; The data sub-block descriptor includes a plurality of data sub-block descriptor entries, each data sub-block descriptor entry corresponding to one of the data sub-blocks.

4. The method according to any one of claims 1 to 3, characterized in that: The telemetry log includes a first telemetry log area telemetry log area 1; wherein the telemetry log area 1 points to the first data block address where the firmware scene is located; When the host reads the Telemetry log header, the storage device sets the Last Block of the telemetry logarea 1 to the firmware field length plus the telemetry log header length to respond to the host.

5. The method according to any one of claims 1 to 4, characterized in that: When the host reads the Telemetry log area 1, the storage device provides all the data blocks constituting the firmware scene to the host.

6. A method for generating a core dump file on-site using firmware that has abnormal power failure, characterized in that: Applied to a host, the host is connected to a storage device, the method comprises: Reading a Telemetry Log from the storage device, wherein the Telemetry Log includes firmware field data; Parsing and unpacking the firmware field data to generate a packaging configuration file and a data sub-block file; A core dump file is obtained based on the encapsulation configuration file and the data sub-block file.

7. The method according to claim 6, characterized in that: The step of parsing and unpacking the firmware field data to generate a packaging configuration file and a data sub-block file includes: Reading a number from each field sub-block in the firmware field data; Get the data sub-block descriptor from the live sub-block; For each data sub-block descriptor entry, obtaining a storage area type of the data sub-block from the data sub-block descriptor entry, determining a data sub-block identifier based on the storage area type and the number, obtaining a data sub-block size from the data sub-block descriptor entry, and obtaining a position of the data sub-block within the field sub-block; generating the data sub-block file based on the data sub-block identifier, the position of the data sub-block within the on-site sub-block, and the size of the data sub-block; A packaging entry is generated based on the data sub-block identifier, the data sub-block size and the position of the data sub-block in the on-site sub-block and added to the packaging configuration file.

8. The method according to claim 6 or 7, characterized in that: The obtaining of a core dump file based on the encapsulation configuration file and the data sub-block file comprises: Obtaining encapsulation entries representing the intra-core registers of the CPU core from the encapsulation configuration file, and corresponding data sub-block files, and generating program headers and sections representing the intra-core registers of the CPU core according to a target file format of a core dump type; Obtaining from the encapsulation configuration file encapsulation entries of all independent storages within the core, encapsulation entries of fast storages shared between all cores, encapsulation entries of slow storages shared between all cores, and corresponding data sub-block files of the current core, and generating corresponding program headers and sections according to the target file format of the core dump type; Generate a target file header according to the target file format of the core dump type; The target file header, all program headers and all sections are combined to generate the core dump file.

9. A method for obtaining a core dump file generated on-site by firmware of abnormal power failure, characterized in that: include: When the storage device is powered on, if it is identified that the abnormal power-off processing flow of the storage device when it was powered off last time did not end normally, based on multiple fixed-size areas of fixed-size continuous space in the persistent storage medium space, read the data blocks corresponding to the first fixed-size sub-area of ​​each of one, part or all of the multiple fixed-size areas in sequence; If no legal data block can be read from the first fixed-size sub-region of the current fixed-size region, then read the next fixed-size region; and if a legal data block is read from the first fixed-size sub-region of the current fixed-size region, then continue to read the current fixed-size region to obtain all data blocks constituting the firmware scene; Determine the firmware field length based on the sum of the quantities of all data blocks; The host reads a Telemetry Log from the storage device, wherein the Telemetry Log includes firmware field data; in response to the host reading the Telemetry Log, the storage device determines a size of the Telemetry Log based on the firmware field length, and uses one, more or all of the data blocks as the content of the Telemetry Log; The host parses and unpacks the firmware field data to generate a packaging configuration file and a data sub-block file; The host obtains a core dump file based on the encapsulation configuration file and the data sub-block file.

10. An electronic device, characterized in that: The electronic device comprises: processor; a memory for storing instructions executable by the processor; The processor is used to read the executable instructions from the memory and execute the executable instructions to implement the method for obtaining the firmware scene of abnormal power failure described in any one of claims 1-5, or the method for generating a core dump file using the firmware scene of abnormal power failure described in any one of claims 6-8, or the method for obtaining the firmware scene of abnormal power failure described in claim 9.

Citation Information

Cited By

  • Method for firmware context preservation in event of unexpected power loss, and method for firmware context acquisition

    WO2026124207A1