Fault reporting method of storage array and electronic device
Patent Information
- Application Number
- CN202610877863.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-17
- Publication Date
- 2026-09-22
- Estimated Expiration
- 2046-06-17
AI Technical Summary
[0003]本申请提供了一种存储阵列的故障上报方法及电子设备,以至少解决相关技术中存储阵列的故障上报效率较低的问题
[0006]通过本申请,存储系统中的目标业务节点获取在存储阵列上执行接收到的用于请求访问存储阵列上的目标存储空间的第一访问请求的第一访问结果,在第一访问结果用于指示目标存储空间的第一运行状态为故障状态的情况下,根据在存储阵列上在第一访问请求之后执行的用于访问目标存储空间的参考访问操作的参考访问结果确认目标存储空间的目标上报信息,即目标业务节点在访问到故障状态的目标存储空间的情况下,不会立刻上报目标存储空间故障,而是根据参考访问操作确认目标存储空间的目标上报信息,再向控制节点上报目标上报信息,减少了对存储空间的临时故障的上报,减少了存储空间的临时故障信息在节点之间的不必要传输,因此,可以解决相关技术中存储阵列的故障上报效率较低的技术问题,达到提高存储阵列的故障上报效率的技术效果。
Smart Images

Figure CN122431939B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of storage technology, and in particular to a fault reporting method and electronic device for a storage array. Background Technology
[0002] Storage array technology is commonly used in storage systems to improve data security and I / O (Input / Output) performance. In related technologies, when a media error occurs while accessing data in the storage array, a storage space failure is reported to the storage system. The storage system then synchronizes this failure information to all service nodes within the storage system. However, due to the existence of temporary storage failures, the failure reporting efficiency of the storage array is not high. Summary of the Invention
[0003] This application provides a fault reporting method and electronic device for a storage array, so as to at least solve the problem of low fault reporting efficiency of storage arrays in the related art.
[0004] This application provides a fault reporting method for a storage array, applied to a target service node in a storage system. The method includes: obtaining a first access result of executing a received first access request on the storage array, wherein the storage system includes: a control node, multiple service nodes, and a storage array, the multiple service nodes and the control node are all connected to the storage array, the control node is also connected to the multiple service nodes, the multiple service nodes include a target service node, the storage array includes multiple storage spaces, the control node stores array status information of the storage array, the array status information records the operating status of the multiple storage spaces, the first access request is used to request access to a target storage space on the storage array, the multiple storage spaces include the target storage space; if the first access result indicates that the first operating status of the target storage space is a fault state, confirming target reporting information of the target storage space based on a reference access operation executed on the storage array, wherein the reference access operation is used to access the target storage space and is executed after the first access request; and reporting the target reporting information to the control node, wherein the control node updates the array status information based on the target reporting information.
[0005] This application also provides an electronic device, including: a memory for storing a computer program; and a processor for implementing the fault reporting method of any of the above-described memory arrays when executing the computer program.
[0006] Through this application, the target service node in the storage system obtains the first access result of the first access request received on the storage array for requesting access to the target storage space on the storage array. When the first access result indicates that the first operating state of the target storage space is a fault state, the target service node confirms the target reporting information of the target storage space according to the reference access result of the reference access operation executed on the storage array after the first access request for accessing the target storage space. That is, when the target service node accesses the target storage space in a fault state, it will not immediately report the target storage space fault, but will confirm the target reporting information of the target storage space according to the reference access operation, and then report the target reporting information to the control node. This reduces the reporting of temporary faults of the storage space and reduces the unnecessary transmission of temporary fault information of the storage space between nodes. Therefore, it can solve the technical problem of low fault reporting efficiency of storage arrays in related technologies and achieve the technical effect of improving the fault reporting efficiency of storage arrays. Attached Figure Description
[0007] To more clearly illustrate the embodiments of this application, the accompanying drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0008] Figure 1 This is a hardware structure block diagram of the fault reporting method for the storage array according to an embodiment of this application;
[0009] Figure 2 This is a flowchart of a fault reporting method for a storage array according to an embodiment of this application;
[0010] Figure 3 This is a schematic diagram of a storage system according to an embodiment of this application;
[0011] Figure 4 This is a schematic diagram of a metadata storage structure according to an embodiment of this application;
[0012] Figure 5 This is a schematic diagram of a storage set according to an embodiment of this application;
[0013] Figure 6 This is a structural block diagram of a fault reporting device for a storage array according to an embodiment of this application. Detailed Implementation
[0014] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the protection scope of this application.
[0015] It should be noted that, in the description of this application, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. The terms "first," "second," etc., in this application are used to distinguish similar objects and are not used to describe a specific order or sequence.
[0016] To enable those skilled in the art to better understand the present application, the present application will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0017] The specific application environment architecture or specific hardware architecture on which the execution of the fault reporting method of the storage array depends is described here.
[0018] The methods and embodiments provided in this application can be executed on a server device or a similar computing device. Taking running on a server device as an example, Figure 1 This is a hardware structure block diagram of a fault reporting method for a storage array according to an embodiment of this application. Figure 1 As shown, the server device may include one or more ( Figure 1 Only one is shown in the image. A processor 102 (which may include, but is not limited to, a central processing unit (CPU), microprocessor (MCU), or programmable logic device (FPGA), etc.) and a memory 104 for storing data are also shown. The server device may further include a transmission device 106 for communication functions and an input / output device 108. Those skilled in the art will understand that... Figure 1 The structure shown is for illustrative purposes only and does not limit the structure of the server equipment described above. For example, the server equipment may also include components that are more... Figure 1 The more or fewer components shown, or having the same Figure 1 The different configurations shown.
[0019] The memory 104 can be used to store computer programs, such as application software programs and modules, like the computer program corresponding to the fault reporting method of the storage array in this embodiment. The processor 102 executes various functional applications and data processing by running the computer program stored in the memory 104, thus implementing the above-described method. The memory 104 may include high-speed random access memory and non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 104 may further include memory remotely located relative to the processor 102, and these remote memories can be connected to server devices via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.
[0020] The transmission device 106 is used to receive or send data via a network. Specific examples of the network described above may include a wireless network provided by a communication provider for the server device. In one example, the transmission device 106 includes a Network Interface Controller (NIC), which can connect to other network devices via a base station to communicate with the Internet. In another example, the transmission device 106 may be a Radio Frequency (RF) module used for wireless communication with the Internet.
[0021] The embodiments of this application provide a fault reporting method for a storage array. The method is described in detail below, taking into account the execution flow of the fault reporting method for a storage array.
[0022] This embodiment provides a method for reporting faults in a storage array. Figure 2 This is a flowchart of a fault reporting method for a storage array according to an embodiment of this application, such as... Figure 2 As shown, the method includes the following steps:
[0023] Step S202: Obtain the first access result of executing the received first access request on the storage array. The storage system includes: a control node, multiple service nodes, and a storage array. The multiple service nodes and the control node are all connected to the storage array. The control node is also connected to the multiple service nodes. The multiple service nodes include a target service node. The storage array includes multiple storage spaces. The control node stores the array status information of the storage array. The array status information records the operating status of the multiple storage spaces. The first access request is used to request access to the target storage space on the storage array. The multiple storage spaces include the target storage space.
[0024] Step S204: If the first access result indicates that the first operating state of the target storage space is a fault state, the target reporting information of the target storage space is confirmed according to the reference access result of the reference access operation performed on the storage array, wherein the reference access operation is used to access the target storage space and is performed after the first access request.
[0025] Step S206: Report target reporting information to the control node, wherein the control node is used to update the array status information based on the target reporting information.
[0026] Based on the above, the target service node in the storage system obtains the first access result of the first access request received on the storage array for requesting access to the target storage space on the storage array. When the first access result indicates that the first operating state of the target storage space is a fault state, the target service node confirms the target reporting information of the target storage space according to the reference access result of the reference access operation executed on the storage array after the first access request for accessing the target storage space. That is, when the target service node accesses the target storage space in a fault state, it will not immediately report the target storage space fault. Instead, it confirms the target reporting information of the target storage space according to the reference access operation and then reports the target reporting information to the control node. This reduces the reporting of temporary faults of the storage space and reduces the unnecessary transmission of temporary fault information of the storage space between nodes. Therefore, it can solve the technical problem of low fault reporting efficiency of storage arrays in related technologies and achieve the technical effect of improving the fault reporting efficiency of storage arrays.
[0027] In the embodiment provided in step S202, Figure 3 This is a schematic diagram of a storage system according to an embodiment of this application. For example... Figure 3 As shown, the storage system may include, but is not limited to, a control node, multiple service nodes, and a storage array. The multiple service nodes may, but are not limited to, all be connected to the storage array, and the control node may, but is not limited to, also be connected to the storage array. The aforementioned fault reporting method for the storage array may, but is not limited to, be applied to a target service node in the storage system. The target service node may, but is not limited to, be any one of the service nodes in the storage array.
[0028] Optionally, in this embodiment, the control node may be, but is not limited to, a centralized management entity in the storage system responsible for global state management, metadata coordination, and resource scheduling. The control node may not directly participate in the read / write services of business data, but instead centrally maintains the global operational status information of the storage array. This global operational status information may include, but is not limited to, metadata such as the operational status of each storage space, fault records, redundancy distribution, and spare block mapping. The control node may, but is not limited to, receive fault reports from each business node, verify and update its internal array status information, ensuring consistency and high availability in the management of storage resources throughout the entire storage system.
[0029] Optionally, in this embodiment, the service node may be, but is not limited to, a processing unit in the storage system that directly responds to I / O (Input / Output) requests from clients or upper-layer applications. The service node may be, but is not limited to, used for physical or logical data interaction with the storage array and performing actual read / write operations. Each service node may, but is not limited to, operate independently, handle concurrent requests from different applications, and, when a media error occurs while accessing the storage array, determine whether to report the fault recovery status through local judgment and secondary verification.
[0030] Optionally, in this embodiment, besides configuring a control node separately from multiple service nodes, it is also possible, but not limited to, selecting one service node from multiple service nodes as the control node when multiple service nodes are configured. For example, in a resource-constrained or lightweight deployment scenario, a service node with high reliability, low load, high network bandwidth, and infrequent I / O operations can be temporarily or logically assigned the function of a control node. For example, in a small cloud storage cluster, a service node mainly runs metadata caching services, log collection systems, or monitoring agents, with extremely low disk I / O load, stable network connections, and redundant and reliable connections to the storage array. In this case, this node can be logically upgraded to a control node, and a fault metadata service can be deployed on it, while other high-load service nodes (such as nodes running databases) still focus on handling I / O.
[0031] Optionally, in this embodiment, the service node may include, but is not limited to, a storage server, a virtual machine, or an SSD (Solid State Drive) controller, etc., and the control node may include, but is not limited to, a metadata server, a cluster monitoring service, or a cloud platform control plane, etc.
[0032] Optionally, in this embodiment, the storage array may be, but is not limited to, a unified storage unit with data protection and performance scalability capabilities, logically aggregated from multiple physical storage devices (such as HDDs, SSDs, and NVMe disks) through striping, redundant coding, or replication mechanisms. The storage array may, but is not limited to, appear externally as one or more logical storage spaces (such as LUNs, volumes, or object buckets), and may, but is not limited to, support block-level or file-level access.
[0033] Optionally, in this embodiment, the storage array may include, but is not limited to, RAID (Redundant Array of Independent Disks), RAID 5 / 6, Erasure Coding system, multi-replica storage pool, or logical volume of distributed object storage, etc.
[0034] Optionally, in this embodiment, the storage space may be, but is not limited to, the smallest manageable logical unit that can be independently marked, independently recovered, and independently reported. The storage space may be, but is not limited to, a sector (the finest granularity), or a data block, stripe, or PACK (a set of stripes) (the coarsest granularity) (a PACK may be, but is not limited to, a set of multiple stripes, and a PACK may include, but is not limited to, multiple stripes), etc.
[0035] Optionally, in this embodiment, the control node may, but is not limited to, store the array status information of the storage array, and each service node may, but is not limited to, also store the array status information of the storage array. When the control node updates the array status information stored on the control node, it may, but is not limited to, also synchronously update the array status information stored on each service node.
[0036] Optionally, in this embodiment, the array status information may, but is not limited to, record the operating status of multiple storage spaces in the storage array. Optionally, the array status information may, but is not limited to, record the operating status of the storage space in the form of a bitmap, and may, but is not limited to, setting the bitmap corresponding to the storage space to 0 when the operating status of the storage space is normal, and setting the bitmap corresponding to the storage space to 1 when the operating status of the storage space is faulty.
[0037] Optionally, in this embodiment, the array status information may be used, but is not limited to, for maintaining the storage array, and may be used, but is not limited to, to determine whether each storage space needs to be replaced by a spare storage space based on the operating status of each storage space recorded in the array status information.
[0038] Optionally, in this embodiment, the first access request may refer to, but is not limited to, a raw I / O operation initiated by an upper-layer application or client, aimed at accessing a specific target storage space (such as a specific block or sector in a LUN) in the storage array. The first access request may be, but is not limited to, a read request (such as a database reading a page) or a write request (such as a virtual machine writing to disk).
[0039] Optionally, in this embodiment, the first access result of the first access request may be used, but is not limited to, to indicate the first operating state of the target storage space. For example, the first access result may include, but is not limited to, whether the first access request successfully read the data stored in the target storage space, or it may include, but is not limited to, whether the first access request successfully stored the data that needs to be stored in the target storage space, etc.
[0040] In the embodiment provided in step S204, if the first access result indicates that the first operating state of the target storage space is a fault state, the specific details of the current fault state of the target storage space can be further confirmed, such as whether the current fault state of the target storage space is a temporary fault, a permanent fault, or the probability that the current fault state is a permanent fault, etc.
[0041] Optionally, in this embodiment, the reference access operation may be, but is not limited to, the access operation corresponding to the access request for requesting access to the target storage space received by the target service node other than the first access request, or may be, but is not limited to, the access operation to the target storage space constructed by the target service node, or may be, but is not limited to, a set of the above-mentioned multiple access operations.
[0042] Optionally, in this embodiment, similar to the previous one, the reference access result can be used, but is not limited to, to indicate the operating status of the target storage space. For example, the reference access result can include, but is not limited to, whether the data stored in the target storage space was successfully read through the reference access operation, or it can also include, but is not limited to, whether the constructed data was successfully stored in the target storage space through the reference access operation, etc.
[0043] Optionally, in this embodiment, the operating state may include, but is not limited to, a fault state and a normal state. Optionally, a fault state may refer to a storage space being detected with a media error during read / write operations, requiring it to be marked as "unreliable," but its recoverability has not yet been confirmed. A normal state may refer to a storage space where the physical media has not experienced an error or the physical media has recovered, and the storage space can be safely used for subsequent read / write operations.
[0044] Optionally, in this embodiment, the target reporting information may be used, but is not limited to, to indicate the operating status of the target storage space, changes in the operating status of the target storage space, etc. Optionally, the target reporting information may be, but is not limited to, empty. Empty target reporting information may be used, but is not limited to, to indicate that the target service node has not detected a storage block that has experienced a fault / temporary fault. It may be, but is not limited to, not reporting any information to the control node or reporting an empty message to the control node when the target reporting information is empty.
[0045] Optionally, in this embodiment, confirming the target reporting information of the target storage space based on the reference access result of performing a reference access operation on the storage array may include, but is not limited to, sending a historical acquisition request to the control node, wherein the historical acquisition request is used to request the acquisition of the first number of times the target storage space was marked as having a fault state in a historical time period before the current time and the second number of times the target storage space was marked as having changed its operating state from a fault state to a normal state in the historical time period; receiving the first number and the second number; comparing the first number with a first number threshold; if the first number is greater than or equal to the first number threshold, generating a test access request, wherein the test access request is used to request the storage of test data to the target storage space; executing the test access request; and confirming the target reporting information based on the access result of the test access request, wherein the reference access result is used to confirm the target reporting information of the target storage space. The reference access operation includes the access operation corresponding to the test access request; if the first number of times is less than the first number threshold and the difference between the first number and the second number is greater than or equal to the preset recovery confidence threshold, obtain the second access result of the second access request received after the first access request is executed on the storage array; determine the target reporting information based on the second access result, wherein the reference access operation includes the access operation corresponding to the second access request; if the first number of times is less than the first number threshold and the difference between the first number and the second number is less than the preset recovery confidence threshold, generate a third access request, wherein the third access request is used to request the recovery of data stored in the target storage space; execute the third access request; determine the target reporting information based on the third access result of the third access request, wherein the reference access operation includes the access operation corresponding to the third access request.
[0046] Optionally, in this embodiment, confirming the target reporting information based on the access result of the test access request may include, but is not limited to: confirming that the target reporting information indicates an unrecoverable failure of the target storage space when the access result of the test access request indicates that the test running status of the target storage space is a fault state; and confirming that the probability of the target reporting information indicating an unrecoverable failure of the target storage space is greater than a probability threshold when the access result of the test access request indicates that the test running status of the target storage space is a normal state. The control node is configured to: update the running status of the target storage space recorded in the array status information to indicate an unrecoverable failure when the number of reported information indicating that the probability of an unrecoverable failure of the target storage space is greater than the probability threshold is greater than or equal to a quantity threshold.
[0047] Through the above steps, by introducing the first number (i.e., failure frequency) and the second number (i.e., recovery count), and combining the threshold judgment, the response strategy is divided into three categories: relying on natural business I / O verification when there is mild fluctuation (i.e., confirming the target reported information through the second access request), triggering redundant reconstruction when there is moderate degradation (i.e., confirming the target reported information through the third access request), and directly injecting test write when there is severe high risk (i.e., confirming the target reported information through the test access request, and further determining the probability of an unrecoverable failure by the test access result, so that the control node can dynamically update the array status based on the probability model).
[0048] Through the above steps, on the one hand, performance and reliability are precisely balanced. Low-risk blocks avoid ineffective active testing and reconstruction, saving bandwidth, computing, and I / O resources; medium-risk blocks delay the consumption of spare blocks through self-healing, extending media lifespan; and high-risk blocks complete status confirmation and replacement at the fastest speed, preventing business avalanche caused by hesitation. On the other hand, by upgrading the target reporting information from whether recovery is possible to the probability of failure, a quantifiable decision basis is provided for the more complex scheduling of the subsequent storage array, enabling the entire storage system to have advanced autonomous capabilities that are predictable, assessable, and optimizable.
[0049] Optionally, in this embodiment, confirming the target reporting information of the target storage space based on the reference access results of the reference access operation performed on the storage array may also include, but is not limited to: detecting the ratio between the number of reference access results indicating that the target storage space is in a normal operating state and the number of reference access results indicating that the target storage space is in a fault operating state; if the ratio is greater than a ratio threshold and the target storage space is in a fault operating state as recorded in the array status information, determining that the target reporting information is used to indicate that the target storage space's operating state has changed from a fault state to a normal state; if the ratio is greater than or equal to the ratio threshold and the target storage space is in a fault operating state as recorded in the array status information, determining that the target reporting information is empty; if the ratio is less than or equal to the ratio threshold, determining that the target reporting information is used to indicate that the target storage space is in a fault operating state.
[0050] Optionally, in this embodiment, after each business node discovers a storage space suspected of being faulty, it further confirms whether the operating status of the storage space is faulty. This allows multiple business nodes to simultaneously perform confirmation operations on different storage spaces, improving the efficiency of further confirming the operating status of the storage space.
[0051] In the embodiment provided in step S206, after confirming the target reporting information of the target storage space, the target reporting information can be reported to the control node. After receiving the target reporting information, the control node can update the array status information (or array status information and its mirror image) on the control node (or the control node and the service node) according to the target reporting information. Optionally, to ensure the global consistency and reliability of the array status information in the storage system, in the storage system, only the control node may have the authority to modify the array status information and its mirror image, and each service node may only be able to report status change suggestions to the control node, and may not directly modify the global array status information / its mirror image.
[0052] Optionally, in this embodiment, target reporting information may be reported to the control node alone, or target reporting information may be reported to the control node while reporting reporting information to other storage spaces.
[0053] As an optional implementation, the target reporting information of the target storage space can be confirmed based on the reference access result of the reference access operation performed on the storage array in the following ways: obtaining the second access result of the second access request received after the execution of the first access request on the storage array, wherein the second access request is used to request access to the target storage space, and the reference access operation includes the second access request; when the second access result indicates that the second operating state of the target storage space is a normal state, extracting the third operating state of the target storage space from the mirror information of the stored array state information; and generating target reporting information based on the third operating state and the second operating state.
[0054] Optionally, in this embodiment, if the second access result indicates that the second operating state of the target storage space is a fault state, the target reported information can be determined to indicate that the operating state of the target storage space is a fault state.
[0055] Optionally, in this embodiment, the accuracy of the detected fault status of the target storage space can be determined by the access results of other access requests received after the first access request, in order to avoid cross-node data transmission caused by reporting inaccurate fault status.
[0056] As an optional implementation, the target reporting information can be generated based on the third operating state and the second operating state in the following ways, but not limited to: when the third operating state is a fault state, the target reporting information is determined to indicate that the operating state of the target storage space changes from the third operating state to the second operating state; when the third operating state is a normal state, the target reporting information is determined to be empty.
[0057] Optionally, in this embodiment, it is possible, but not limited to, that when the third operating state is a fault state, it is determined that the operating state of the target storage space has recovered from the fault state recorded in the array status information / array status information to the normal state, and it is possible, but not limited to, that the target reported information is used to indicate that the operating state of the target storage space has changed from the fault state to the normal state.
[0058] Optionally, in this embodiment, if the third operating state is a normal state, it can be determined that the operating state of the target storage space is actually still the normal state recorded in the array status information / array status information. In this case, it can be determined that the target storage space is operating normally, saving cross-node data transmission overhead.
[0059] As an optional implementation, the target reporting information of the target storage space can also be confirmed based on the reference access result of the reference access operation performed on the storage array in the following ways: obtaining the third access result of the third access request performed on the storage array after the first access request was performed, wherein the third access request is used to request the recovery of data stored in the target storage space, and the reference access operation includes the third access request; and generating target reporting information based on the third access result.
[0060] Optionally, in this embodiment, it is possible, but not limited to, to search for the storage space where the data associated with the original data stored in the target storage space is located, and to access these searched storage spaces to obtain the data associated with the original data stored in the target storage space. The original data originally stored in the target storage space is calculated using the obtained data, and an access request is created to request that the original data be written to the target storage space, thereby obtaining a third access request.
[0061] Optionally, in this embodiment, the operating status of the target storage space can be reconfirmed by creating a third access request and using the access result of the third access request, which can avoid the waiting time for receiving other access requests to the target storage space.
[0062] As an optional implementation, the generation of target reporting information based on the third access result can be achieved, but is not limited to, by the following means: if the third access result indicates that the data in the target storage space has been successfully recovered, the third operating state of the target storage space is extracted from the mirror information of the stored array status information; target reporting information is generated based on the third operating state; if the third access result indicates that the data in the target storage space has failed to be recovered, the target reporting information is determined to indicate that the operating state of the target storage space is a fault state.
[0063] Optionally, in this embodiment, when the third access result indicates that the data in the target storage space has been successfully restored, it is possible, but not limited to, to determine whether it is necessary to report information that the target storage space is operating normally by confirming the status of the target storage space recorded in the stored image information, thereby reducing unnecessary cross-node information transmission.
[0064] Optionally, in this embodiment, the operating status recorded in the array status information may include, but is not limited to, an unrecoverable fault status in addition to fault status and normal status. An unrecoverable fault status may, but is not limited to, represent that the physical medium of the storage space has suffered irreversible damage, and the storage space must be replaced. It may, but is not limited to, determine that the target reported information is empty if the operating status of the target storage space is confirmed to be an unrecoverable fault status.
[0065] Optionally, in this embodiment, the control node may, but is not limited to, extract the third operating state of the target storage space from the stored array status information when it receives target reporting information indicating that the operating state of the target storage space is in a fault state. If the third operating state is in a normal state, the control node updates the number of times the target storage space has been detected to have a fault. If the number of times is equal to the fault count threshold, the control node changes the third operating state of the target storage space recorded in the array status information to an unrecoverable fault state. If the number of times is less than the fault count threshold, the control node changes the third operating state of the target storage space recorded in the array status information to a fault state.
[0066] As an optional implementation, the generation of target reporting information based on the third operating state can be achieved, but is not limited to, in the following ways: when the third operating state is a fault state, the target reporting information is determined to indicate that the operating state of the target storage space is a normal state; when the third operating state is a normal state, the target reporting information is determined to be empty.
[0067] Optionally, in this embodiment, the control node may, but is not limited to, extract the third operating state of the target storage space from the stored array status information when it receives target reporting information indicating that the operating state of the target storage space is normal, and change the third operating state of the target storage space recorded in the array status information to normal if the third operating state is faulty.
[0068] As an optional implementation, before obtaining the third access result of executing the third access request on the storage array after executing the first access request, it is possible, but not limited to, to search for the associated storage space of the target storage space from the storage array, wherein the associated storage space is a storage space whose stored data is associated with the data stored in the target storage space; to generate target data using the data stored in the associated storage space; and to generate a third access request carrying the target data, wherein the third access request is used to request that the target data be written to the target storage space.
[0069] Optionally, in this embodiment, the associated storage space in the storage array that is associated with the target storage space may include, but is not limited to, other storage spaces in the stripe where the target storage space is located, and the storage space may be, but is not limited to, blocks under the stripe.
[0070] Optionally, in this embodiment, by constructing a third access request carrying the target data, the operating status of the target storage space can be confirmed while avoiding any impact on the data stored in the target storage space.
[0071] As an optional implementation, the third access result of executing a third access request on the storage array after executing a first access request can be obtained in the following ways: detecting whether the operating status of the target storage space has changed from a fault state to a normal state by executing the access request; if the operating status of the target storage space has not changed from a fault state to a normal state by executing the access request, the third access result is obtained.
[0072] Optionally, in this embodiment, the operating status of the target storage space can be confirmed by simultaneously utilizing both the access operation corresponding to the received access request and the access operation corresponding to the constructed access request. This utilizes existing access requests without causing continuous waiting due to excessive reliance on receiving access requests to confirm the operating status of the storage space.
[0073] Optionally, in this embodiment, the period for statistically detecting fault conditions and fault recovery status on each service node can be set, and the operating status of the target storage space can be determined by the access results of the received access requests to the target storage space within each period, and the storage spaces whose operating status is indicated as faulty by the access results detected by the target service node within the next period can be traversed when the next period arrives. For these storage spaces, it can be queried whether the operating status has changed from the detected faulty state to the normal state through the execution of subsequent access requests. If the operating status has not changed from the detected faulty state to the normal state through the execution of subsequent access requests, then access requests for data recovery are constructed respectively, and the operating status is determined by the constructed access requests.
[0074] As an optional implementation, the detection of whether the operating status of the target storage space has changed from a fault state to a normal state by executing an access request can be achieved in the following ways, including: searching for the target storage space from a stored first list, wherein the first list records storage spaces that have changed from a fault state to a normal state by executing an access request; if the target storage space is found, determining that the operating status of the target storage space has changed from a fault state to a normal state by executing an access request; if the target storage space is not found, determining that the operating status of the target storage space has not changed from a fault state to a normal state by executing an access request.
[0075] Optionally, in this embodiment, if the operating status of the storage space changes from a fault state to a normal state after the operation requested by the access request is executed, the storage space can be recorded in the first list. Then, the subsequent processing operation of the access request can be executed. This avoids the timeout waiting caused by the message of changing the operating status of the storage space from a fault state to a normal state being transmitted to all business nodes and control nodes when executing the access request. This allows the access request to end quickly and greatly improves the performance of the storage system.
[0076] Optionally, in this embodiment, in addition to recording storage spaces that were found to be in a faulty state by the target service node and then recovered to a normal state, the first list may also record, but is not limited to, storage spaces that were found to be in a faulty state by other service nodes but recovered to a normal state by the target service node. The changes in the operating status of the storage spaces recorded in the first list that were found to be in a faulty state by other service nodes but recovered to a normal state by the target service node may also be reported to the control node.
[0077] As an optional implementation, the target reporting information of the target storage space can also be confirmed based on the reference access result of the reference access operation performed on the storage array when the first access result indicates that the first operating state of the target storage space is a fault state: when the first access result indicates that the first operating state of the target storage space is a fault state, the target space identifier of the target storage space is stored in a second list, wherein the second list records the fault space identifier of the fault storage space whose operating state of the requested access storage space is a fault state when the access result of the received access request performed by the target service node on the storage array indicates that the access result of the requested access storage space is a fault state; the fault space identifiers recorded in the second list are traversed; when the target space identifier is encountered, the target reporting information is confirmed based on the reference access result; and the target space identifier is deleted from the second list.
[0078] Optionally, in this embodiment, the fault space identifier of the faulty storage space that indicates the operating status of the requested storage space is faulty, which is the result of the access request received by the target service node on the storage array, can be recorded in the second list, but is not limited to. That is, the space identifier of the storage space suspected of being faulty, which is discovered by the target service node, can be recorded in the second list.
[0079] Optionally, in this embodiment, by traversing the fault space identifiers recorded in the second list to find the storage space whose operating status needs to be further confirmed, each target business node can confirm the operating status of the storage space within its processing scope, avoiding repeated confirmation of the operating status of the same storage space by multiple different business nodes.
[0080] Optionally, in this embodiment, Figure 4 This is a schematic diagram of a metadata storage structure according to an embodiment of this application. For example... Figure 4 As shown, cluster bad block metadata (i.e., the aforementioned array status information) can be stored on node 3 (i.e., the aforementioned control node), and cluster bad block metadata mirror (i.e., the aforementioned array status information mirror information), local bad block metadata (i.e., the aforementioned second list) and bad block to be cleared metadata (i.e., the aforementioned first list) can be stored on node 0 and node 1 (i.e., the aforementioned business node).
[0081] As an optional implementation, the target space identifier of the target storage space can be stored in the second list in the following manner, but not limited to: searching for the target space identifier in the second list, wherein the second list records faulty space identifiers in the order in which the storage spaces are found to be in a faulty state; if the target space identifier is not found, storing the target space identifier at the end of the faulty space identifiers already recorded in the second list.
[0082] Optionally, in this embodiment, it is possible, but not limited to, before storing the target space identifier in the second list, to first determine whether the target space identifier has already been stored in the second list. Only if the target space identifier is not stored in the second list will the target space identifier be stored in the second list. This avoids repeated entry of the target space identifier and also avoids multiple confirmations of the target storage space's operating status when further confirming the operating status of the storage space suspected of being faulty based on the information recorded in the second list.
[0083] As an optional implementation, after confirming the target reporting information based on the reference access results, other storage spaces in the target space set where the target storage space is located can be searched from the storage array, but not limited to cases where the target reporting information indicates that the operating status of the target storage space has changed from a fault state to a normal state; other space identifiers of other storage spaces are searched in the second list to obtain search results; and other reporting information of other storage spaces is generated based on the search results.
[0084] Optionally, in this embodiment, if it is determined that the target reported information is used to indicate that the target storage space has recovered from the fault state to the normal state, it is possible, but not limited to, to determine whether other storage spaces in the same space set as the target storage space are in the normal state. The above operations prepare for the overall reporting of the recovery of the entire space set from the fault state to the normal state.
[0085] Optionally, in this embodiment, multiple consecutive storage spaces in the storage array can be aggregated into a space set, for example, 16 consecutive stripes can be defined as a space set (PACK).
[0086] Optionally, in this embodiment, Figure 5 This is a schematic diagram of a storage set according to an embodiment of this application. For example... Figure 5 As shown, a PACK may include, but is not limited to, 16 stripes, each stripe may include, but is not limited to, multiple blocks, and each data block may include, but is not limited to, multiple sectors. Optionally, each stripe may manage, but is not limited to, the bad blocks in the blocks under each stripe, and each block structure may manage, but is not limited to, a set of bitmaps to identify which sectors in the current block are bad blocks. If all sectors in a block are bad blocks, the block may be marked as a fully bad block.
[0087] Optionally, in this embodiment, storage spaces that logically belong to the same application, the same file, or the same database tablespace can be aggregated into a single space set, even if these storage spaces are not contiguous in physical address. Through the above, the complete usage status of the storage space of an application, a file, or a database tablespace can be determined as quickly as possible, effectively avoiding false availability misjudgments caused by partial storage space recovery.
[0088] As an optional implementation, other reporting information for other storage spaces can be generated based on the search results in the following ways, but not limited to: if other space identifiers are found, other reporting information for other storage spaces is confirmed based on the access results corresponding to the other storage spaces; if other space identifiers are not found, other reporting information is determined to indicate that the operating status of other storage spaces is normal.
[0089] Optionally, in this embodiment, the method of confirming other reported information of other storage spaces based on the access results corresponding to other storage spaces may be, but is not limited to, similar to the aforementioned method of confirming the target reported information of the target storage space based on the reference access results of executing a reference access request on the storage array, and will not be described again here.
[0090] As an optional implementation, target reporting information can be reported to the control node in the following way, but not limited to: reporting set reporting information to the control node, wherein the set reporting information includes: target reporting information and other reporting information, and the control node is used to update the operating status of the storage space in the target space set stored in the array status information according to the set reporting information.
[0091] Optionally, in this embodiment, a method of reporting information from a collection can be used instead of reporting information from each storage space individually. This can reduce interaction with the control node and further reduce interaction between the control node and other service nodes, thereby improving the overall efficiency of maintaining the array status information of the storage system.
[0092] Optionally, in this embodiment, it should be noted that the aggregated reporting information includes target reporting information and other reporting information. However, the aggregated reporting information is not a simple concatenation of target reporting information and other reporting information, but a complete reporting information that integrates the semantics of target reporting information and other reporting information.
[0093] Optionally, in this embodiment, when reporting information indicating a faulty state of the storage space, or when determining the storage space's operating state during the first access, the reporting can be done at the sector level, but is not limited to supporting bad block submission at the block level. When an I / O operation (i.e., the aforementioned access request) encounters a media error while reading one or more sectors, it can be confirmed, but is not limited to, at the sector level, that a storage space suspected of being faulty exists. If an I / O operation encounters a media error while reading an entire stripe block, it can be recorded, but is not limited to, at the block level. When reporting a faulty state of the storage space, it can be reported that the entire stripe block is faulty when it is, but is not limited to, instead of reporting a faulty sector individually when a faulty sector is found. Through the above, frequent bitmap operations can be avoided.
[0094] Optionally, in this embodiment, each service node may, but is not limited to, initiate a background bad block cleanup task. The background bad block cleanup tasks of all service nodes in the storage system may, but are not limited to, be performed concurrently and independently of each other.
[0095] Based on the bad block records in the local bad block metadata and bad block pending cleanup metadata of the PACK to which the bad block belongs, the following three priorities apply: First priority: Locate the PACK to which the bad block belongs in the bad block pending cleanup metadata. Check if all bad blocks on that PACK are recorded in the bad block pending cleanup metadata. If all bad block records for that PACK are in the bad block pending cleanup metadata, since the bad blocks in the bad block pending cleanup metadata have been successfully written by I / O, they are considered normal blocks. Therefore, only a bad block elimination request (i.e., the aforementioned reporting information) needs to be sent to the control node. After receiving the bad block elimination request, the control node first deletes the bad block metadata stored in the bad block metadata image of each node cluster. Then, it deletes the bad block metadata stored in the cluster bad block metadata. Finally, it notifies the service nodes to delete the corresponding metadata in the bad block pending cleanup metadata. Second priority: Locate the PACK to which the bad block belongs in the bad block pending cleanup metadata. Check if all bad blocks on that PACK are recorded in the bad block pending cleanup metadata. If some bad blocks in that PACK are not recorded in the bad block pending cleanup metadata but exist in the local bad block metadata... At this point, for all stripes not included in the bad block pending metadata, read the other data blocks of the first stripe, calculate the data of the segment containing the bad block, and write the calculated data to the segment containing the bad block. Process all stripes containing bad blocks within the PACK according to the above steps. After processing, send a bad block elimination request to the control center on a PACK basis. After receiving the bad block elimination request, the control node first deletes the bad block metadata stored in the bad block metadata image of each node cluster. Then deletes the bad block metadata stored in the cluster bad block metadata. Finally, delete the corresponding metadata in the bad block pending metadata and the corresponding metadata in the local bad block metadata. For the third priority, establish a background bad block elimination task based on the local bad block metadata. First, find the first PACK of the local bad block metadata, and starting from the first stripe with bad blocks in the PACK, read the other data blocks of the stripe, calculate the data of the segment containing the bad block, and write the calculated data to the segment containing the bad block. Then start processing the next stripe in the PACK. After all stripe processing within a PACK is completed, a bad block elimination request is sent to the control center on a PACK-by-PACK basis. Upon receiving the bad block elimination request, the control center first deletes the bad block metadata stored in the bad block metadata image of each node cluster. Then, it deletes the bad block metadata stored in the cluster bad block metadata. Finally, it deletes the corresponding metadata from the local bad block metadata. This optimizes the granularity of bad block elimination operations (i.e., the aforementioned confirmation operation of whether the storage space is actually in a faulty state) from sector-by-sector to PACK-by-PACK, significantly reducing interaction with the control node and improving the speed of bad block elimination.
[0096] Because access requests exhibit spatial and temporal locality, most bad blocks and the I / O operations involving their access are highly likely to occur on the same node. Therefore, bad block elimination is performed concurrently across nodes using local bad block metadata and bad block metadata to be cleared, instead of relying on the cluster's bad block metadata for sequential bad block cleanup. This significantly speeds up the bad block cleanup process.
[0097] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods according to the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method.
[0098] Embodiments of this application also provide a fault reporting device for a storage array, which can be applied to, but is not limited to, target service nodes. Figure 6 This is a structural block diagram of a fault reporting device for a storage array according to an embodiment of this application, such as... Figure 6 As shown, the device includes:
[0099] The acquisition module 602 is used to acquire the first access result of executing the received first access request on the storage array. The storage system includes: a control node, multiple service nodes and a storage array. The multiple service nodes and the control node are all connected to the storage array. The control node is also connected to the multiple service nodes. The multiple service nodes include a target service node. The storage array includes multiple storage spaces. The control node stores the array status information of the storage array. The array status information records the operating status of the multiple storage spaces. The first access request is used to request access to the target storage space on the storage array. The multiple storage spaces include the target storage space.
[0100] The confirmation module 604 is used to confirm the target reporting information of the target storage space based on the reference access result of the reference access operation performed on the storage array when the first access result indicates that the first operating state of the target storage space is a fault state. The reference access operation is used to access the target storage space and is performed after the first access request.
[0101] The reporting module 606 is used to report target reporting information to the control node, wherein the control node is used to update the array status information based on the target reporting information.
[0102] Through the above device, the target service node in the storage system obtains the first access result of the first access request received on the storage array for requesting access to the target storage space on the storage array. When the first access result indicates that the first operating state of the target storage space is a fault state, the target reporting information of the target storage space is confirmed according to the reference access result of the reference access operation executed on the storage array after the first access request for accessing the target storage space. That is, when the target service node accesses the target storage space in a fault state, it will not immediately report the target storage space fault, but will confirm the target reporting information of the target storage space according to the reference access operation, and then report the target reporting information to the control node. This reduces the reporting of temporary faults of the storage space and reduces the unnecessary transmission of temporary fault information of the storage space between nodes. Therefore, it can solve the technical problem of low fault reporting efficiency of storage arrays in related technologies and achieve the technical effect of improving the fault reporting efficiency of storage arrays.
[0103] In some embodiments, the confirmation module includes: a first acquisition unit, configured to acquire a second access result of a second access request received after the execution of a first access request on the storage array, wherein the second access request is used to request access to a target storage space, and the reference access operation includes the second access request; an extraction unit, configured to extract a third operating state of the target storage space from the mirror information of the stored array state information when the second access result indicates that the second operating state of the target storage space is a normal state; and a first generation unit, configured to generate target reporting information based on the third operating state and the second operating state.
[0104] In some embodiments, the first generating unit is further configured to: determine that target reporting information is used to indicate that the operating state of the target storage space changes from the third operating state to the second operating state when the third operating state is a fault state; and determine that target reporting information is empty when the third operating state is a normal state.
[0105] In some embodiments, the confirmation module includes: a second acquisition unit, configured to acquire a third access result of executing a third access request on the storage array after executing a first access request, wherein the third access request is used to request the recovery of data stored in the target storage space, and the reference access operation includes the third access request; and a second generation unit, configured to generate target reporting information based on the third access result.
[0106] In some embodiments, the second generation unit is further configured to: extract the third operating state of the target storage space from the mirror information of the stored array status information when the third access result indicates that the data in the target storage space has been successfully recovered; generate target reporting information based on the third operating state; and determine that the target reporting information indicates the operating state of the target storage space as a fault state when the third access result indicates that the data in the target storage space has failed to be recovered.
[0107] In some embodiments, the second generation unit is further configured to: determine that the target reporting information is used to indicate that the operating state of the target storage space is normal when the third operating state is a fault state; and determine that the target reporting information is empty when the third operating state is a normal state.
[0108] In some embodiments, the confirmation module further includes: a first lookup unit, configured to look up an associated storage space of the target storage space in the storage array before obtaining the third access result of executing a third access request on the storage array after executing a first access request, wherein the associated storage space is a storage space whose stored data is associated with the data stored in the target storage space; a third generation unit, configured to generate target data using the data stored in the associated storage space; and a fourth generation unit, configured to generate a third access request carrying the target data, wherein the third access request is used to request that the target data be written to the target storage space.
[0109] In some embodiments, the second acquisition unit is further configured to: detect whether the operating state of the target storage space has changed from a fault state to a normal state by executing an access request; and if it is detected that the operating state of the target storage space has not changed from a fault state to a normal state by executing an access request, acquire a third access result.
[0110] In some embodiments, the second acquisition unit is further configured to: search for a target storage space from a stored first list, wherein the first list records storage spaces that have changed from a fault state to a normal state by executing an access request; if a target storage space is found, determine that the operating state of the target storage space has changed from a fault state to a normal state by executing an access request; if a target storage space is not found, determine that the operating state of the target storage space has not changed from a fault state to a normal state by executing an access request.
[0111] In some embodiments, the confirmation module includes: a storage unit, configured to store a target space identifier of the target storage space into a second list when the first access result indicates that the first operating state of the target storage space is a fault state, wherein the second list records fault space identifiers of fault storage spaces whose access results on the storage array, when executed by the target service node, indicate that the operating state of the requested accessed storage space is a fault state; a traversal unit, configured to traverse the fault space identifiers recorded in the second list; and a determination unit, configured to confirm the target reported information based on the reference access result when the target space identifier is traversed; and to delete the target space identifier from the second list.
[0112] In some embodiments, the storage unit is further configured to: search for a target space identifier in a second list, wherein the second list records fault space identifiers in the order in which the storage spaces are found to be in a faulty state; and if the target space identifier is not found, store the target space identifier at the end of the fault space identifiers already recorded in the second list.
[0113] In some embodiments, the confirmation module further includes: a second search unit, configured to, after confirming the target reported information based on the reference access result, search for other storage spaces in the target space set where the target storage space is located from the storage array when the target reported information indicates that the operating status of the target storage space has changed from a fault state to a normal state; a third search unit, configured to search for other space identifiers of other storage spaces in a second list to obtain search results; and a fifth generation unit, configured to generate other reported information of other storage spaces based on the search results.
[0114] In some embodiments, the fifth generation unit is further configured to: if other space identifiers are found, confirm other reported information of other storage spaces based on the access results corresponding to other storage spaces; if other space identifiers are not found, determine that other reported information is used to indicate that the operating status of other storage spaces is normal.
[0115] In some embodiments, the reporting module includes: a reporting unit, configured to report set reporting information to the control node, wherein the set reporting information includes: target reporting information and other reporting information, and the control node is configured to update the operating status of the storage space in the target space set stored in the array status information according to the set reporting information.
[0116] For a description of the features in the embodiment corresponding to the fault reporting device of the storage array, please refer to the relevant description of the embodiment corresponding to the fault reporting method of the storage array, which will not be repeated here.
[0117] Embodiments of this application also provide an electronic device, including a memory and a processor, wherein the memory stores a computer program, and the processor is configured to run the computer program to perform the steps in any of the above embodiments of the fault reporting method for a memory array.
[0118] Embodiments of this application also provide a computer-readable storage medium storing a computer program, wherein the computer program is configured to execute the steps in the embodiments of the fault reporting method for any of the above-described storage arrays when it is run.
[0119] In one exemplary embodiment, the aforementioned computer-readable storage medium may include, but is not limited to, various media capable of storing computer programs, such as a USB flash drive, read-only memory (ROM), random access memory (RAM), portable hard disk, magnetic disk, or optical disk.
[0120] Embodiments of this application also provide a computer program product, which includes a computer program that, when executed by a processor, implements the steps in any of the above-described storage array fault reporting method embodiments.
[0121] Embodiments of this application also provide another computer program product, including a non-volatile computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps in any of the above-described embodiments of the fault reporting method for a storage array.
[0122] Any of the components, modules, units, parts, methods, and operations described herein can be implemented using software, firmware, hardware (e.g., fixed logic circuitry), manual processing, or any combination thereof. Alternatively or additionally, any functionality described herein can be executed at least in part by one or more hardware logic components, such as, but not limited to, a central processing unit (CPU), a field-programmable gate array (FPGA), an application-specific integrated circuit (ASIC), an application-specific standard product (ASSP), a system-on-a-chip (SoC), a complex programmable logic device (CPLD), a microprocessor (MCU), etc. The terms "system," "computing device," or "apparatus" as used herein encompass various means, devices, and machines for processing data, including, for example, one or more programmable processors, computers, SoCs, or combinations thereof. The apparatus may also include code that creates an execution environment for the computer program in question, such as code constituting processor firmware, a protocol stack, a database management system, an operating system, a cross-platform runtime environment, a virtual machine, or one or more combinations thereof. The aforementioned computer program (also known as a program, software, software application, app, script, or code) can be written in any form of programming language, including compiled or interpreted languages, declarative or procedural languages, and can be deployed in any form, including as a standalone program or as a module, component, subroutine, object, or other unit suitable for a computing environment.
[0123] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0124] The foregoing has provided a detailed description of a fault reporting method and electronic device for a storage array provided in this application. Specific examples have been used to illustrate the principles and implementation methods of this application. The descriptions of the embodiments above are merely for the purpose of helping to understand the method and its core ideas. It should be noted that those skilled in the art can make various improvements and modifications to this application without departing from its principles, and these improvements and modifications also fall within the protection scope of the claims of this application.
Claims
1. A method for reporting faults in a storage array, characterized in that, The method, applied to a target service node in a storage system, includes: Obtain a first access result of executing a received first access request on a storage array, wherein the storage system includes: a control node, multiple service nodes, and the storage array, the multiple service nodes and the control node are all connected to the storage array, the control node is also connected to the multiple service nodes, the multiple service nodes include the target service node, the storage array includes multiple storage spaces, the control node stores array status information of the storage array, the array status information records the operating status of the multiple storage spaces, and the first access request is used to request access to the target storage space on the storage array, the multiple storage spaces include the target storage space; When the first access result indicates that the first operating state of the target storage space is a fault state, the target reporting information of the target storage space is confirmed according to the reference access result of the reference access operation performed on the storage array, wherein the reference access operation is used to access the target storage space and the reference access operation is performed after the first access request; The target reporting information is reported to the control node, wherein the control node is used to update the array status information based on the target reporting information; The step of confirming the target reporting information of the target storage space based on the reference access result of a reference access operation performed on the storage array when the first access result indicates that the first operating state of the target storage space is a fault state includes: storing the target space identifier of the target storage space in a second list when the first access result indicates that the first operating state of the target storage space is a fault state within the current period, wherein the second list records the fault space identifier of the fault storage space whose operating state of the requested access space is the fault state when the access result of the received access request performed by the target service node on the storage array indicates that the operating state of the requested access storage space is the fault state; traversing the fault space identifiers recorded in the second list when the next period of the current period arrives; confirming the target reporting information based on the reference access result when the target space identifier is encountered, wherein the reference access operation includes the access operation corresponding to the access request for requesting access to the target storage space received by the target service node other than the first access request and the access operation for the target storage space constructed by the target service node; and deleting the target space identifier from the second list. The step of confirming the target reporting information based on the reference access result includes: searching for the target storage space from a first list, wherein the first list records storage spaces that have changed from the fault state to the normal state by executing an access request; if the target storage space is not found, determining that the operating state of the target storage space has not changed from the fault state to the normal state by executing an access request; if the target storage space is found, determining that the operating state of the target storage space has changed from the fault state to the normal state by executing an access request; if the operating state of the target storage space has not changed from the fault state to the normal state by executing an access request, obtaining a third access result of a third access request, wherein the third access request is used to request the recovery of data stored in the target storage space, and the reference access operation includes the third access request; and generating the target reporting information based on the third access result.
2. The fault reporting method for a storage array according to claim 1, characterized in that, The step of generating the target reporting information based on the third access result includes: If the third access result indicates that the data in the target storage space has been successfully recovered, the third operating state of the target storage space is extracted from the mirror information of the stored array status information; the target reporting information is generated based on the third operating state. If the third access result indicates that data recovery in the target storage space has failed, the target reported information is determined to indicate that the operating status of the target storage space is the fault state.
3. The fault reporting method for a storage array according to claim 2, characterized in that, The step of generating the target reporting information based on the third operating state includes: If the third operating state is the fault state, the target reported information is determined to indicate that the operating state of the target storage space is normal. If the third operating state is the normal state, it is determined that the target reported information is empty.
4. The fault reporting method for a storage array according to claim 1, characterized in that, Before obtaining the third access result of the third access request, the method further includes: Locate the associated storage space of the target storage space from the storage array, wherein the associated storage space is a storage space whose stored data is related to the data stored in the target storage space; The target data is generated using the data stored in the associated storage space; A third access request carrying the target data is generated, wherein the third access request is used to request that the target data be written to the target storage space.
5. The fault reporting method for a storage array according to claim 1, characterized in that, The step of storing the target space identifier of the target storage space into the second list includes: The target space identifier is searched in the second list, wherein the faulty space identifiers are recorded in the second list in the order in which the storage spaces are found to be in a faulty state; If the target space identifier is not found, the target space identifier is stored at the end of the fault space identifier already recorded in the second list.
6. The fault reporting method for a storage array according to claim 1, characterized in that, After confirming the target reported information based on the reference access result, the method further includes: When the target reported information indicates that the operating status of the target storage space has changed from the fault state to the normal state, other storage spaces in the target space set where the target storage space is located are searched from the storage array; Search for other space identifiers of the other storage spaces in the second list to obtain the search results; Based on the search results, other reporting information for the other storage spaces is generated.
7. The fault reporting method for a storage array according to claim 6, characterized in that, The step of generating other reporting information for the other storage spaces based on the search results includes: If the other space identifiers are found, the other reported information of the other storage spaces is confirmed based on the access results corresponding to the other storage spaces; If no other space identifier is found, the other reported information is determined to indicate that the operating status of the other storage space is normal.
8. The fault reporting method for a storage array according to claim 6, characterized in that, The step of reporting the target reporting information to the control node includes: The control node reports set reporting information, wherein the set reporting information includes: target reporting information and other reporting information, and the control node is used to update the operating status of the storage space in the target space set stored in the array status information according to the set reporting information.
9. An electronic device, characterized in that, include: Memory, used to store computer programs; A processor, configured to implement the steps of the fault reporting method for the storage array as described in any one of claims 1 to 8 when executing the computer program.
Citation Information
Patent Citations
A memory array monitor system and method
CN109460194A
Storage system disk fault information acquisition method and device, electronic equipment and medium
CN113934581A