Abnormality processing system, method and equipment and storage medium

By designing an exception handling system in the data center, and using the coordinated work of distributed storage systems and operation and maintenance nodes to early warning and mark abnormal application data nodes, the service interruption problem caused by passive handover of abnormal handling in the data center in the prior art is solved, and the effect of reducing the risk of service interruption and ensuring service quality is achieved.

CN119945882AActive Publication Date: 2025-05-06HANGZHOU ALICLOUD FEITIAN INFORMATION TECH CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202311473009.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-06
Publication Date
2025-05-06
Estimated Expiration
2043-11-06

AI Technical Summary

Technical Problem

The existing data center exception handling method is passive switching, and it takes a long time to determine the data center exception, resulting in service interruption and affecting service quality.

Method used

Design an exception handling system, through the coordinated work of distributed storage systems, operation and maintenance nodes and user terminals, warning and marking abnormal application data nodes, the user terminal actively avoids abnormal nodes, switches to other normal nodes, and reduces the risk of service interruption.

Benefits of technology

By predicting and marking abnormal application data nodes in advance, user terminals can switch access traffic to other normal nodes without perception, reducing the risk of service interruption, or even without service interruption, thereby ensuring the service quality of the distributed storage system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119945882A_ABST
    Figure CN119945882A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides an exception handling system and method, equipment and a storage medium. In the embodiment of the invention, the metadata node can predict the application data node which is about to be subjected to the service interruption before the service interruption, and stores the identifier of the predicated abnormal application data node which is about to be subjected to the service interruption. The user terminal and the metadata node interact the identifier of the abnormal application data node, so that the user terminal can actively avoid the abnormal application data node in the process of accessing the distributed storage system, the access flow is switched to other normal application data nodes under the condition that a user does not perceive, and the user experience is improved. The risk of service interruption is reduced, even no service interruption exists, and therefore the service quality of the distributed storage system is guaranteed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of cloud service technology, and in particular to an exception handling system, method, device and storage medium. Background Art

[0002] As information infrastructure, data centers are a strong support for the development of the digital economy. Data centers sometimes experience abnormalities such as temperature and humidity, power outages, or network disconnections, which can cause upper-layer services in the data center to be unavailable. Services are deployed across data centers, and when one data center fails, services can be switched to other data centers, which will continue to provide services.

[0003] However, the existing method for handling data center anomalies is a passive switching method, which takes a long time to determine whether the data center is abnormal. During the period of determining whether the data center is abnormal, service interruption may occur, affecting service quality. Summary of the invention

[0004] Multiple aspects of the present application provide an exception handling system, method, device and storage medium to reduce the risk of service interruption.

[0005] In a first aspect, an embodiment of the present application provides an exception handling system, comprising: a distributed storage system, a user terminal, and an operation and maintenance node of the distributed storage system; the distributed storage system comprises: a storage cluster; the storage cluster comprises: a plurality of storage units; the plurality of storage units are deployed in a plurality of physical spaces; a storage unit corresponds to a physical space one-to-one; the storage unit comprises: a metadata node and an application data node; the metadata nodes in the plurality of storage units are backed up with each other and used to store metadata of the storage cluster; the application data nodes in the plurality of storage units are backed up with each other and used to store application data;

[0006] The operation and maintenance node is used to send abnormal prompt information to the target metadata node in the multiple physical spaces in response to the service interruption warning information for the target physical space in the multiple physical spaces; the abnormal prompt information includes the identifier of the target physical space;

[0007] The target metadata node is used to mark the application data node in the target physical space as an abnormal application data node in response to the abnormal prompt information; and store the identifier of the abnormal application data node;

[0008] The user terminal is used to obtain the identifier of the abnormal application data node from the target metadata node; and send an access request to other application data nodes except the application data node corresponding to the identifier of the abnormal application data node.

[0009] In a second aspect, an embodiment of the present application further provides an exception handling method, which is applicable to an operation and maintenance node of a distributed storage system, wherein the distributed storage system comprises: a storage cluster; the storage cluster comprises: a plurality of storage units; the plurality of storage units are deployed in a plurality of physical spaces; the storage units correspond to the physical spaces one by one; the storage units comprise: a metadata node and an application data node; the metadata nodes in the plurality of storage units are backed up with each other and used to store metadata of the storage cluster; the application data nodes in the plurality of storage units are backed up with each other and used to store application data;

[0010] The method comprises:

[0011] Obtaining service interruption warning information for a target physical space among the multiple physical spaces;

[0012] In response to the service interruption warning information, abnormal prompt information is sent to the target metadata nodes in the multiple physical spaces, so that the target metadata nodes can respond to the abnormal prompt information, mark the application data nodes in the target physical space as abnormal application data nodes, and store the identifiers of the abnormal application data nodes, so that the user terminal can send access requests to other application data nodes except the application data nodes corresponding to the identifiers of the abnormal application data nodes.

[0013] In a third aspect, an embodiment of the present application further provides an exception handling method, which is applicable to a target metadata node in a distributed storage system, wherein the distributed storage system comprises: a storage cluster; the storage cluster comprises: a plurality of storage units; the plurality of storage units are deployed in a plurality of physical spaces; the storage units correspond to the physical spaces one by one; the storage units comprise: a metadata node and an application data node; the metadata nodes in the plurality of storage units are backed up with each other and used to store the metadata of the storage cluster; the application data nodes in the plurality of storage units are backed up with each other and used to store application data;

[0014] The method comprises:

[0015] Acquire abnormal prompt information sent by the operation and maintenance node of the distributed storage system; the abnormal prompt information is generated by the operation and maintenance node in response to service interruption warning information for a target physical space among the multiple physical spaces, and the abnormal prompt information includes an identifier of the target physical space;

[0016] In response to the abnormal prompt information, marking the application data node in the target physical space as an abnormal application data node;

[0017] The identifier of the abnormal application data node is stored so that the user terminal can obtain the identifier of the abnormal application data node, and an access request is sent to other application data nodes except the application data node corresponding to the identifier of the abnormal application data node.

[0018] In a fourth aspect, an embodiment of the present application further provides an exception handling method, which is applicable to a user terminal, wherein a distributed storage system comprises: a storage cluster; the storage cluster comprises: a plurality of storage units; the plurality of storage units are deployed in a plurality of physical spaces; a storage unit corresponds to a physical space one by one; the storage unit comprises: a metadata node and an application data node; the metadata nodes in the plurality of storage units are backed up with each other and used to store metadata of the storage cluster; the application data nodes in the plurality of storage units are backed up with each other and used to store application data;

[0019] The method comprises:

[0020] Obtaining an identifier of an abnormal application data node from a target metadata node in the multiple physical spaces; the abnormal application data node is an application data node in a target physical space warned by the service interruption warning information; the target physical space is a portion of the multiple physical spaces;

[0021] The access request is sent to other application data nodes except the application data node corresponding to the identifier of the abnormal application data node.

[0022] In a fifth aspect, an embodiment of the present application further provides a computing device, comprising: a memory, a processor, and a communication component; the memory is used to store a computer program;

[0023] The processor is coupled to the memory and the communication component, and is used to execute the computer program to perform the steps in the exception handling method provided in the second aspect, the third aspect and / or the fourth aspect.

[0024] In the sixth aspect, an embodiment of the present application also provides a computer-readable storage medium storing computer instructions, which, when executed by one or more processors, causes the one or more processors to execute the steps in the exception handling method provided in the second aspect, the third aspect and / or the fourth aspect.

[0025] In an embodiment of the present application, the metadata node can predict the application data node that will experience a service interruption before the service is interrupted, and store the identification of the abnormal application data node that is predicted to experience a service interruption. By exchanging the identification of the abnormal application data node between the user terminal and the metadata node, the user terminal can actively avoid the abnormal application data node during access to the distributed storage system, and switch the access traffic to other normal application data nodes without the user's perception, reducing the risk of service interruption, or even eliminating service interruption, thereby ensuring the service quality of the distributed storage system. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:

[0027] Figure 1 A schematic diagram of the structure of an exception handling system provided in an embodiment of the present application;

[0028] Figure 2 A schematic diagram of the structure of the physical space provided in the embodiment of the present application;

[0029] Figure 3-Figure 5 A flowchart of an exception handling method provided in an embodiment of the present application;

[0030] Figure 6 A schematic diagram of the structure of a computing device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0031] In order to make the purpose, technical solution and advantages of the present application clearer, the technical solution of the present application will be clearly and completely described below in combination with the specific embodiments of the present application and the corresponding drawings. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present application.

[0032] Distributed storage systems store data in a distributed manner. Data can be file data, image data, or data tables, etc. Distributed storage systems can implement distributed storage of data based on block storage services. For example, in a distributed storage system, data is distributed and stored in the form of data blocks (Chunk) on different disks (such as disks D0-Dn). For most distributed storage systems, metadata and application data are usually stored separately, that is, the control flow and the data flow are separated, so as to obtain higher policy scalability and input / output (I / O) concurrency. Based on this, distributed storage systems are generally divided into two types of service nodes: metadata (Meta) nodes and application data (Data) nodes.

[0033] The metadata node is used to store the metadata of the distributed storage system. The metadata includes: the address mapping relationship of the application data, that is, the address information of the application data node where the application data is located. The application data node is used to store application data, that is, service data. The metadata node can control the location of the application data node where the application data is written, and inform the user end of the location of the application data node where the data to be read is located when reading data.

[0034] When the distributed storage system is deployed across data centers, the metadata node can automatically identify the application data nodes in each data center, so that the metadata node can establish a heartbeat connection with each application data node. When an application data node in a data center is abnormal, the metadata node can determine that the application data node is abnormal by detecting multiple heartbeat abnormalities of the application data node. If the number of abnormal application data nodes in the data center exceeds the set number threshold, the data center is determined to be abnormal. From then on, subsequent data reading and writing are switched to other normal data centers and are no longer directed to abnormal data centers, thereby achieving the purpose of avoiding abnormal data centers.

[0035] The above data center abnormality handling method requires multiple heartbeat anomalies to determine data node abnormalities. This process takes minutes to complete and is time-consuming. Service interruptions may occur during the determination of whether the data center is abnormal, affecting service quality. Furthermore, if the data center has abnormal temperature and humidity or unexpected disasters (such as fire, flood, etc.), not all application data nodes will be abnormal at the same time. The above method will take a longer time to determine the abnormality of the entire data center, resulting in data reading and writing being smoothly diverted to the abnormal data center, affecting service quality.

[0036] In order to reduce the risk of service interruption and improve service quality, in some embodiments of the present application, the metadata node can predict the application data node that will experience service interruption before the service interruption, and store the identification of the abnormal application data node that is predicted to experience service interruption. By exchanging the identification of abnormal application data nodes between the user terminal and the metadata node, the user terminal can actively avoid abnormal application data nodes during access to the distributed storage system, and switch the access traffic to other normal application data nodes without the user's perception, reducing the risk of service interruption, or even eliminating service interruption, thereby ensuring the service quality of the distributed storage system.

[0037] The technical solutions provided by various embodiments of the present application are described in detail below in conjunction with the accompanying drawings.

[0038] It should be noted that the same reference numerals denote the same objects in the following drawings and embodiments, and therefore, once an object is defined in one drawing or embodiment, it does not need to be further discussed in the subsequent drawings and embodiments.

[0039] Figure 1 This is a schematic diagram of the structure of the exception handling system provided in the embodiment of the present application. Figure 1 As shown, the exception handling system includes: a distributed storage system 10, an operation and maintenance node 20 of the distributed storage system, and a user terminal 30.

[0040] like Figure 1 As shown, the distributed storage system 10 may include: a storage cluster 101. The storage cluster may be one or more. More than one means two or more. The distributed storage system 10 may use the storage cluster 101 to store data in a distributed manner.

[0041] The storage cluster 101 includes: a plurality of storage units 102. In order to increase data security, the plurality of storage units 102 are deployed in a plurality of physical spaces 40, and the storage units 102 correspond to the physical spaces 40 one by one. That is, one physical space 40 includes one storage unit 102 in the same storage cluster 101. In each embodiment of the present application, a plurality refers to two or more.

[0042] In the embodiment of the present application, the physical space 40 refers to the actual carrier space for storing electronic devices such as servers, and the physical space 40 can provide resources such as power and network for the servers. The distributed storage system 10 is deployed across the physical space 40. The distributed storage system 10 is deployed across the physical space 40, and can connect servers in multiple physical spaces 40 to provide services. When a problem occurs in a physical space, servers in other physical spaces can take over the services in the abnormal physical space to continue to ensure the normal operation of the services.

[0043] The implementation form of the physical space 40 is determined by the distributed storage granularity of the distributed storage system 10. For example, if the distributed storage system 10 is deployed across data centers, the physical space 40 may be a data center. A data center may be used as a physical space. For another example, if the distributed storage system 10 is deployed across computer rooms, the physical space 40 may be a computer room. One or more computer rooms may also constitute a physical space. For another example, if the distributed storage system 10 is deployed across cabinets, the physical space 40 may be a cabinet. One or more cabinets may constitute a physical space. Figure 2 As shown, a data center may include one or more computer rooms; and one computer room may include multiple cabinets. As the infrastructure of the data center, the cabinet can be used as a carrier for storing servers and can also provide resources such as power and communication for the servers.

[0044] In this embodiment, the distributed storage system 10 stores metadata and application data independently, that is, separates the control flow from the data flow. Figure 1 As shown, the storage unit 102 may include: a metadata node 103 and an application data node 104 .

[0045] In this embodiment, the metadata nodes 103 in the multiple storage units 102 in the same storage cluster 101 back up each other and are used to store the metadata of the storage cluster. The application data nodes 104 in the multiple storage units 102 in the same storage cluster 101 back up each other and are used to store application data. Among them, the metadata includes: the address mapping relationship of the application data in the storage cluster 101, that is, the address information of the application data node where the application data is located. The application data node is used to store application data, that is, service data. The metadata node can control the location of the application data node where the application data is written, and inform the user terminal 30 of the location of the application data node where the data to be read is located when reading data.

[0046] In this embodiment, the user terminal 30 refers to a computer device occupied by the user and having the functions of computing, surfing the Internet, communicating, etc. required by the user, such as a mobile phone, a tablet computer, a personal computer, or a wearable device. The user terminal 30 can obtain the address information of the application data node to be accessed from the metadata node 103; and based on the address information, access the application data node. Among them, the address information of the application data node can be the address information required by the network protocol followed by the access request. For example, if the network protocol followed by the access request is the Transmission Control Protocol (TCP), the address information of the application data node can be the Internet Protocol (IP) address and port number.

[0047] The operation and maintenance node 20 is a server-side device where the operation and maintenance system of the distributed storage system is located, and generally has the ability to undertake and guarantee services. The server-side device can be a single server device, a cloud-based server array, or a virtual machine (VM) running in a cloud-based server array. In addition, the server-side device can also refer to other computing devices with corresponding service capabilities, such as computers and other terminal devices (running service programs), etc.

[0048] In this embodiment, if Figure 1 As shown in step 1, the operation and maintenance node 20 obtains the service interruption warning information of the above-mentioned physical space 40. In the embodiment of the present application, the cause of the service interruption of the physical space 40 is not limited. The abnormality of the physical space 40 will cause the service interruption of the physical space 40. The abnormality of the physical space 40 is generally predictable in advance. For example, the air parameters (temperature and / or humidity, etc.) of the physical space 40 change gradually, rather than changing sharply in an instant. The abnormality of the physical space 40 may be the abnormality of the air parameters of the physical space 40, such as the temperature and / or humidity of the physical space 40. The abnormality of the air parameters of the physical space 40 may cause the server corresponding to the storage unit in the physical space 40 to be abnormal, thereby causing the service interruption of the physical space 40. In this embodiment, an air parameter monitoring device (not shown in the accompanying drawings) can be set in the physical space 40. The air parameter monitoring device can monitor the air parameters of the physical space 40, and when the air parameters are monitored to exceed the set normal range, an air parameter abnormality warning (such as temperature abnormality warning and / or humidity abnormality warning, etc.) is issued, and the air parameter abnormality warning information is sent to the operation and maintenance node 20. The operation and maintenance node 20 obtains abnormal warning information of air parameters for the physical space as service interruption warning information.

[0049] For another example, the abnormality of the physical space 40 may be the abnormality of the power resources of the physical space 40, such as the power outage of the physical space 40 within a certain period of time. The power outage of the physical space 40 will also cause the interruption of the physical space service. In this embodiment, the power department will generally notify the operation and maintenance personnel of the physical space 40 in advance of the time point at which one or some physical spaces 40 will be powered off due to power shortages and other reasons. The operation and maintenance node 20 can also obtain power outage warning information before the physical space 40 is powered off. Accordingly, the operation and maintenance node 20 can obtain power outage warning information for the physical space 40 as service interruption warning information.

[0050] For another example, the abnormality of the physical space 40 may be a network abnormality of the physical space 40, such as the physical space 40 will be disconnected from the network for a certain period of time. The disconnection of the physical space 40 will also cause the physical space service to be interrupted. In this embodiment, the network operation and maintenance department generally notifies the operation and maintenance personnel of the physical space 40 in advance of the time point at which one or some physical spaces 40 will be disconnected from the network. The operation and maintenance node 20 can also obtain network disconnection warning information before the physical space 40 is powered off. Accordingly, the operation and maintenance node 20 can obtain network disconnection warning information for the physical space 40 as service interruption warning information.

[0051] The reasons for the service interruption caused by the abnormality of the physical space shown in the above embodiment are only exemplary and do not constitute a limitation. In short, the operation and maintenance node 20 can obtain abnormal warning information for the physical space 40 as service interruption warning information. Among them, the abnormal warning information may include: the cause of the abnormality and the time information when the abnormality will occur.

[0052] In some embodiments of the present application, operation and maintenance of the physical space 40 may also cause service interruption of the physical space 40. Operation and maintenance of the servers in the physical space 40 can also be predicted in advance, and the operation and maintenance personnel will announce in advance the time point at which the physical space 40 will be operated and maintained. Among them, operation and maintenance of the physical space 40 includes but is not limited to: service upgrades of the equipment in the physical space 40, and / or, network switching of the physical space 40 (such as changing network operators, etc.). Correspondingly, the operation and maintenance node 20 can obtain operation and maintenance notice information for the physical space 40 as service interruption warning information, etc. Among them, the operation and maintenance notice information may include: operation and maintenance matters and operation and maintenance time, etc.

[0053] Since the service interruption of the physical space 40 can be predicted in advance, therefore, in the embodiment of the present application, based on the predicted service interruption, intervention can be made in advance to reduce the risk of service interruption of the distributed storage system. The scheme is described in detail below. In the embodiment of the present application, for the convenience of description and distinction, the physical space warned by the service interruption warning information is defined as the target physical space, that is, the target physical space will experience a service interruption. The service interruption warning information may include: the identification of the target physical space warned and the time information of the service interruption, etc.

[0054] like Figure 1 As shown in step 2, when the operation and maintenance node 20 obtains the service interruption warning information for the target physical space 40, in response to the service interruption warning information, it sends abnormal prompt information to the target metadata nodes 103 in the multiple physical spaces 40. The abnormal prompt information is used to prompt that the target physical space 40 will have a service interruption and trigger the target metadata node 103 to respond accordingly. The abnormal prompt information may include the identifier of the target physical space.

[0055] In an embodiment of the present application, the target metadata node 103 may be all metadata nodes in each storage cluster 101. Since the metadata nodes 103 in the same storage cluster 101 are backed up with each other, in some embodiments, a leader node (Leader) may be selected from multiple metadata nodes 103 in the same storage cluster 101 by means of leader election, and the leader node interacts with the operation and maintenance node 20. Correspondingly, the target metadata node may also be a leader node generated by leader election of the metadata nodes 103 in each storage cluster 101. Among them, there is one leader node in one storage cluster 101. Preferably, the target metadata node 103 is a leader node generated by leader election of the metadata nodes in the storage cluster 101.

[0056] In the embodiment of the present application, the method of electing a leader for the multiple metadata nodes in the storage cluster 101 is not limited. In some embodiments, the multiple metadata nodes in the storage cluster 101 may use a distributed consistency algorithm, such as the Raft algorithm, to elect a leader. Specifically, the Raft algorithm divides time into terms and performs a leader election operation at the beginning of each term. Among them, the Raft algorithm uses a heartbeat to trigger the leader election operation. When the multiple metadata nodes in the storage cluster 101 are started, they are initialized as standby nodes (Followers). The leader node periodically sends heartbeat information to other followers. When the standby node receives the heartbeat information sent by other nodes, it confirms that the node that sends the heartbeat information is the master node.

[0057] For the target metadata node 103, in response to the above abnormal prompt information, the application data node in the target physical space can be marked as an abnormal application data node, and the identifier of the abnormal application data node (corresponding to Figure 1 Step 3). Optionally, the identifier of the abnormal application data node may be cached in the memory of the target metadata node 103, or may be stored in a persistent storage medium (such as a disk, hard disk, etc.) of the target metadata node 103.

[0058] In this embodiment, the identifier of the abnormal application data node may be information that uniquely identifies an application data node, such as address information, a number or an identity document (ID) of the application data node.

[0059] like Figure 1As shown in step 4, the user terminal 30 may obtain the identifier of the abnormal application data node from the target metadata node 103. Optionally, the user terminal 30 may periodically obtain the identifier of the abnormal application data node from the target metadata node 103. Specifically, the user terminal 30 may periodically send a query request to the target metadata node 103. In some embodiments, the user terminal 30 may install a software development kit (SDK) corresponding to the distributed storage system 10, and call the storage service through the SDK. Accordingly, the user terminal 30 may periodically send a query request to the target metadata node 103 through the SDK. The query request is used to request the identifier of the abnormal application data node stored in the target metadata node 103.

[0060] Accordingly, the target metadata node 103 may respond to the query request and return the identifier of the stored abnormal application data node to the user terminal 30. The user terminal 30 receives the identifier of the abnormal application data node returned by the target metadata node 103. Specifically, the target metadata node 103 may respond to the query request and return the identifier of the stored abnormal application data node to the SDK on the user terminal 30 side. The SDK on the user terminal 30 side receives the identifier of the abnormal application data node. Further, as Figure 1 As shown in step 5, the user terminal 30 (specifically, SDK) can send the access request to the distributed storage system to other application data nodes except the identifier of the abnormal application data node. Among them, the other application data nodes are application data nodes belonging to the same storage cluster as the abnormal application data node. Since SDK is a background thread running on the user terminal 30, the user terminal 30 resends the access request to other application data nodes except the abnormal application data node through SDK without the user's perception.

[0061] In this embodiment, the metadata node can predict the application data node that will experience service interruption before the service interruption, and store the identification of the abnormal application data node that is predicted to experience service interruption. By exchanging the identification of the abnormal application data node between the user terminal and the metadata node, the user terminal can actively avoid the abnormal application data node during access to the distributed storage system, and switch the access traffic to other normal application data nodes without the user's perception, thereby reducing the risk of service interruption or even eliminating service interruption.

[0062] On the other hand, there is a certain delay because the user terminal obtains the latest identifier of the abnormal application data node from the metadata node. During the period when the user terminal has not yet updated to the latest abnormal application data node, the access traffic will still be directed to the application data node marked as abnormal. In this embodiment, the metadata node only marks the application data node of the target physical space as abnormal, and does not stop the service of these application data nodes. Therefore, before all the access traffic of the user terminal is directed to other application data nodes that are not marked as abnormal, if the application data nodes accessed by the user terminal include application data nodes marked as abnormal, these application data nodes marked as abnormal can still provide services, which can ensure the service quality of the distributed storage system.

[0063] In the embodiment of the present application, the specific implementation method of the user terminal 30 sending the access request to other application data nodes except the application data node corresponding to the identifier of the abnormal application data node is not limited. In some embodiments, the user terminal 30 can detect whether the application data node currently to be accessed includes the abnormal application data node based on the obtained identifier of the abnormal application data node.

[0064] Specifically, when accessing the distributed storage system 10, the user terminal 30 may obtain the identifier of the application data node to be accessed, such as the address information of the application data node to be accessed, from the metadata node 103. Further, the user terminal 30 may match the obtained identifier of the abnormal application data node with the identifier of the current application data node to be accessed; if the identifier of the current application data node to be accessed includes the identifier of the abnormal application data node, it is determined that the current application data node to be accessed includes the abnormal application data node.

[0065] Further, in the case where the current application data node to be accessed includes an abnormal application data node, the user terminal 30 stops sending access requests to the abnormal application data node; and determines a new application data node to be accessed from other application data nodes belonging to the same storage cluster as the abnormal application data node. In the embodiment of the present application, for the convenience of description and distinction, the application data node originally to be accessed by the user terminal 30 is defined as the first application data node; and the determined new application data node to be accessed is defined as the second application data node.

[0066] Wherein, the access operation of the access request is different, and the implementation method of determining the new application data node to be accessed from other application data nodes belonging to the same storage cluster as the abnormal application data node is also different. In some embodiments, the access operation of the access request is a write operation, and the access request is a write request. When the user terminal 30 determines the new application data node to be accessed from other application data nodes belonging to the same storage cluster as the abnormal application data node, it can re-request the application data node to be written to the target metadata node 103.

[0067] Accordingly, the target metadata node 103 can determine a new application data node to be written as the second application data node from other application data nodes belonging to the same storage cluster as the abnormal application data node in response to the request of the user terminal 30 for the application data node to be written.

[0068] In the embodiment of the present application, the specific implementation method of the target metadata node 103 determining the new application data node to be written from other application data nodes belonging to the same storage cluster as the abnormal application data node is not limited.

[0069] In some embodiments, the amount of data to be written in a write request has requirements on the free storage resources of the application data node, that is, the amount of free storage resources of the application data node should meet the storage resource requirements of the write request, such as the amount of free storage resources of the application data node is greater than the amount of storage resources required by the write request. Accordingly, the target metadata node 103 can determine the application data node whose free storage resource amount meets the storage resource requirements of the write request from other application data nodes belonging to the same storage cluster as the abnormal application data node. If there is one application data node whose free storage resource amount meets the storage resource requirements of the write request, the application data node whose free storage resource amount meets the storage resource requirements of the write request is determined to be the second application data node.

[0070] If there are multiple application data nodes whose idle storage resources meet the storage resource requirements of the write request, the application data nodes whose idle storage resources meet the storage resource requirements of the write request can be selected from the application data nodes whose idle storage resources meet the storage resource requirements of the write request, and the application data nodes whose distance from the user terminal 30 meets the set distance condition, as the second application data node. The set distance condition can be that the distance is less than or equal to the set distance threshold. Correspondingly, the target metadata node 103 selects the application data node whose distance from the user terminal 30 meets the set distance threshold from the application data nodes whose idle storage resources meet the storage resource requirements of the write request, as the second application data node. Alternatively, the distance condition is set to the shortest distance. Correspondingly, the target metadata node 103 selects the application data node whose distance from the user terminal 30 meets the shortest distance from the application data nodes whose idle storage resources meet the storage resource requirements of the write request, as the second application data node, and so on.

[0071] In other embodiments, the target metadata node 103 may select, based on the storage water levels of other application data nodes belonging to the same storage cluster as the abnormal application data node, an application data node whose storage water level meets the set water level conditions from the application data nodes whose free storage resources meet the storage resource requirements of the write request, as the second application data node.

[0072] The storage water level of the application data node 104 may be the water level of the occupied storage resources in the application data node 104, that is, the proportion of the occupied storage resources in the application data node 104 to the total storage resources of the application data node 104. In some embodiments, in order to minimize storage resource fragmentation, the application data node with the largest storage water level may be selected from the application data nodes whose idle storage resources meet the storage resource requirements of the write request as the second application data node. In other embodiments, in order to achieve load balancing, the application data node with the smallest storage water level may be selected from the application data nodes whose idle storage resources meet the storage resource requirements of the write request as the second application data node.

[0073] Alternatively, the storage water level of the application data node 104 may also be the water level of the unoccupied storage resources (i.e., free storage resources) in the application data node 104, that is, the proportion of the unoccupied storage resources (i.e., free storage resources) in the application data node 104 to the total storage resources of the application data node 104. Accordingly, in some embodiments, in order to minimize storage resource fragmentation, the application data node with the smallest storage water level may be selected from the application data nodes whose free storage resources meet the storage resource requirements of the write request as the second application data node. In other embodiments, in order to achieve load balancing, the application data node with the largest storage water level may be selected from the application data nodes whose free storage resources meet the storage resource requirements of the write request as the second application data node.

[0074] The method for determining the second application data node shown in the above embodiment is only an exemplary description and does not constitute a limitation. After determining the identifier of the second application data node, the target metadata node 103 can return the identifier of the second application data node to the user terminal 30. Accordingly, the user terminal 30 can receive the identifier of the application data node returned by the target metadata node 103; and determine that the application data node corresponding to the identifier of the application data node returned by the target metadata node 103 is the second application data node.

[0075] Further, the user terminal 30 may send a write request to the second application data node, and the second application data node responds to the write request and writes the to-be-written data carried in the write request to the second application data node.

[0076] In other embodiments, the access operation of the access request is a read operation, and the access request is a read request. When the user terminal 30 determines the second application data node to be accessed from other application data nodes belonging to the same storage cluster as the abnormal application data node, it can determine the application data node belonging to the first application data node originally to be read from other application data nodes belonging to the same storage cluster as the abnormal application data node as the second application data node. That is, for a read request, the user terminal 30 only reads data from other application data nodes in the application data nodes to be read by the read request except the abnormal application data node, thereby avoiding directing the read request to the abnormal application data node.

[0077] The above embodiment takes the access request being a read request or a write request as an example to exemplify the implementation method of determining the second application data node to be accessed from other application data nodes belonging to the same storage cluster as the abnormal application data node, but does not constitute a limitation.

[0078] After determining the second application data node to be accessed by the user terminal 30, the user terminal 30 may send an access request to the second application data node. The second application data node may respond to the access request and perform an access operation according to the access request. For example, if the access request is a write request, the second application data node responds to the write request and writes the data to be written carried by the write request to the second application data node. For another example, if the access request is a read request, the second application data node responds to the read request and reads the data requested by the read request from the second application data node; further, the data requested by the read request is returned to the user terminal 30.

[0079] In an embodiment of the present application, the service of the target physical space where the service interruption occurs will also be restored. For the above-mentioned embodiment in which the service interruption occurs due to the abnormality of the target physical space 40, the target physical space 40 resumes service after the abnormality of the target physical space 40 is repaired. For example, after the air parameters of the target physical space 40 return to normal, the target physical space 40 resumes service. For another example, after the target physical space 40 is powered on again, the target physical space 40 resumes service, and so on. For the above-mentioned embodiment in which the service interruption occurs due to the operation and maintenance of the target physical space 40, the target physical space 40 resumes service after the operation and maintenance of the target physical space 40 is completed. For example, after the equipment upgrade of the target physical space 40 is completed, the target physical space 40 resumes service. For another example, after the network change of the target physical space 40 is completed, the target physical space 40 resumes service, and so on.

[0080] In the case where the target physical space resumes service, the embodiment of the present application can also redirect the access traffic to the application data node of the target physical space. Specifically, the operation and maintenance node 20 can send a service recovery prompt message to the target metadata node 103 when the target physical space resumes service; the service recovery prompt message includes the identifier of the target physical space. Accordingly, the target metadata node 103 can delete the identifier of the application data node in the target physical space from the identifiers of the stored abnormal application data nodes in response to the service recovery prompt message. In this way, when the user terminal subsequently accesses the application data node, the access request of the user terminal can be redirected to the application data node in the target physical space.

[0081] For example, when the user terminal 30 generates a new access request to the distributed storage system 10, if the identifier of the application data node in the target physical space no longer exists in the identifier of the abnormal application data node obtained from the target metadata node 103, the new access request can be sent to the application data node in the target physical space to restore access to the target physical space.

[0082] For write requests, when there is new data to be written, the user terminal 30 may request the target metadata node 103 for the application data node to be written; the target metadata node 103 may exclude the application data node corresponding to the identifier of the stored abnormal application data node from the application data nodes of the storage cluster to which it belongs, obtain the normal application data node, and determine the application data node to be written from the normal application data nodes. Among them, the specific implementation method of the target metadata node 103 determining the application data node to be written from the normal application data nodes can be referred to the above-mentioned target metadata node 103 determining the second application data node from other application data nodes belonging to the same storage cluster as the abnormal application data node. The details will not be repeated here. Further, the target metadata node 103 may return the identifier of the application data node to be written to the user terminal 30. Based on the identifier of the application data node to be written, the user terminal 30 sends the write request to the application data node to be written.

[0083] Since the identifier of the abnormal application data node stored in the target metadata node 103 has deleted the identifier of the application data node in the target physical space of service recovery, when the target metadata node 103 determines the application data node to be written, it can include the application data node in the target physical space in the selection range, and can normally direct the write request to the target physical space of service recovery, thereby restoring the write operation of the target physical space.

[0084] In response to a read request, when reading data, the user terminal 30 may request the target metadata node 103 for the application data node to be read; the target metadata node 103 may exclude the application data node corresponding to the identifier of the abnormal application data node stored in the application data nodes of the storage cluster to which it belongs, obtain normal application data nodes, and determine the application data node where the data to be read corresponding to the read request is located from the normal application data nodes. Further, the target metadata node 103 may return the identifier of the application data node where the data to be read corresponding to the read request is located to the user terminal 30. Based on the identifier of the application data node where the data to be read corresponding to the read request is located, the user terminal 30 sends the read request to the application data node where the data to be read corresponding to the read request is located. The application data node where the data to be read corresponding to the read request is located may execute the read request, read the above-mentioned data to be read, and send the read data to the user terminal 30.

[0085] Since the identifier of the abnormal application data node stored in the target metadata node 103 has deleted the identifier of the application data node in the target physical space of service recovery, when the target metadata node 103 determines the application data node where the data to be read corresponding to the read request is located, it can include the application data node in the target physical space in the selection range, and can normally direct the read request to the target physical space of service recovery, thereby restoring the read operation of the target physical space.

[0086] In addition to the above-mentioned system embodiments, the embodiments of the present application also provide an exception handling method. The exception recovery method provided in the embodiments of the present application is exemplarily described below.

[0087] Figure 3 The flowchart of the exception handling method provided in the embodiment of the present application is shown in FIG. The exception handling method is applicable to the operation and maintenance node of the distributed storage system. Figure 3 As shown, the exception handling method includes:

[0088] 301. Obtain service interruption warning information for a target physical space among multiple physical spaces.

[0089] 302. In response to the service interruption warning information, abnormal prompt information is sent to the target metadata nodes in multiple physical spaces, so that the target metadata nodes can respond to the abnormal prompt information, mark the application data nodes in the target physical space as abnormal application data nodes, and store the identifiers of the abnormal application data nodes, so that the user terminal can send access requests to other application data nodes except the application data nodes corresponding to the identifiers of the abnormal application data nodes.

[0090] Figure 4 FIG. 1 is a flow chart of another exception handling method provided in an embodiment of the present application. The exception handling method is applicable to a target metadata node in a distributed storage system. Figure 4 As shown, the exception handling method includes:

[0091] 401. Obtain abnormal prompt information sent by an operation and maintenance node of a distributed storage system; the abnormal prompt information is generated by the operation and maintenance node in response to service interruption warning information for a target physical space among multiple physical spaces, and the abnormal prompt information includes an identifier of the target physical space.

[0092] 402. In response to the abnormal prompt information, mark the application data node in the target physical space as an abnormal application data node.

[0093] 403. Store the identifier of the abnormal application data node so that the user terminal can obtain the identifier of the abnormal application data node, and send an access request to other application data nodes except the application data node corresponding to the identifier of the abnormal application data node.

[0094] Figure 5 A flowchart of another exception handling method provided in an embodiment of the present application. The method is mainly applicable to a user terminal. The user terminal can access a distributed storage system. Figure 5 As shown, the exception handling method includes:

[0095] 501. Obtain identifications of abnormal application data nodes from target metadata nodes in multiple physical spaces; the abnormal application data nodes are application data nodes in the target physical space warned by service interruption warning information; the target physical space is part of the multiple physical spaces.

[0096] 502. Send the access request to other application data nodes except the application data node corresponding to the identifier of the abnormal application data node.

[0097] In each embodiment of the present application, the distributed storage system includes: a storage cluster. The storage cluster may be one or more. Multiple means two or more. The distributed storage system may use the storage cluster to store data in a distributed manner.

[0098] The storage cluster includes: multiple storage units. In order to increase the security of data, multiple storage units are deployed in multiple physical spaces, and the storage units correspond to the physical spaces one by one. That is, one physical space includes one storage unit in the same storage cluster. In each embodiment of the present application, multiple refers to 2 or more.

[0099] The implementation form of the physical space is determined by the distributed storage granularity of the distributed storage system. For example, if the distributed storage system is deployed across data centers, the physical space can be a data center. A data center can be used as a physical space. For another example, if the distributed storage system is deployed across computer rooms, the physical space can be a computer room. One or more computer rooms can also constitute a physical space. For another example, if the distributed storage system is deployed across cabinets, the physical space can be a cabinet. One or more cabinets can constitute a physical space.

[0100] In this embodiment, the storage unit may include: a metadata node and an application data node. The metadata nodes in multiple storage units in the same storage cluster back up each other and are used to store metadata of the storage cluster. The application data nodes in multiple storage units in the same storage cluster back up each other and are used to store application data.

[0101] In this embodiment, for the operation and maintenance node, in step 301, service interruption warning information for a target physical space among multiple physical spaces may be obtained. In some embodiments, abnormal warning information for the target physical space may be obtained as service interruption warning information. The abnormal warning information may include: abnormal cause and time information when the abnormality will occur.

[0102] In other embodiments, maintenance notice information for the physical space may be obtained as service interruption warning information. For the description of the abnormal cause and maintenance of the target physical space, please refer to the relevant content of the above system embodiment, which will not be repeated here.

[0103] Since the service interruption of the physical space can be predicted in advance, therefore, in the embodiment of the present application, based on the predicted service interruption, intervention can be made in advance to reduce the risk of service interruption of the distributed storage system. The scheme is described in detail below. In the embodiment of the present application, for the convenience of description and distinction, the physical space warned by the service interruption warning information is defined as the target physical space, that is, the target physical space will experience a service interruption. The service interruption warning information may include: the identification of the target physical space warned and the time information of the service interruption, etc.

[0104] When the operation and maintenance node obtains the service interruption warning information for the target physical space, in step 302, in response to the service interruption warning information, an abnormal prompt information may be sent to the target metadata nodes in the multiple physical spaces. The abnormal prompt information is used to prompt that the target physical space will have a service interruption and trigger the target metadata node to respond accordingly. The abnormal prompt information may include an identifier of the target physical space.

[0105] In an embodiment of the present application, the target metadata node may be all metadata nodes in each storage cluster. Since the metadata nodes in the same storage cluster are backed up with each other, in some embodiments, a leader node (Leader) can be selected from multiple metadata nodes in the same storage cluster through a leader election method, and the leader node interacts with the operation and maintenance node. Correspondingly, the target metadata node may also be a leader node generated by leader election of metadata nodes in each storage cluster. Among them, there is one leader node in a storage cluster. Preferably, the target metadata node is a leader node generated by leader election of metadata nodes in the storage cluster. For the specific implementation method of leader election of multiple metadata nodes, please refer to the relevant content of the above-mentioned system embodiment, which will not be repeated here.

[0106] For the target metadata node, in step 401, the above-mentioned abnormal prompt information can be obtained, and in step 402, in response to the above-mentioned abnormal prompt information, the application data node in the target physical space can be marked as an abnormal application data node. Further, in step 403, the identifier of the abnormal application data node can be stored for the user terminal to obtain the identifier of the abnormal application data node, and the access request can be sent to other application data nodes except the application data node corresponding to the identifier of the abnormal application data node. Optionally, the identifier of the abnormal application data node can be cached in the memory of the target metadata node, and of course, the identifier of the abnormal application data node can also be stored in the persistent storage medium (such as a disk, hard disk, etc.) of the target metadata node.

[0107] In this embodiment, the identifier of the abnormal application data node may be information that uniquely identifies an application data node, such as address information, number or ID of the application data node.

[0108] Accordingly, for the user terminal, in step 501, the identifier of the abnormal application data node can be obtained from the target metadata node. Optionally, the identifier of the abnormal application data node can be periodically obtained from the target metadata node. Specifically, a query request can be periodically sent to the target metadata node. The query request is used to request to obtain the identifier of the abnormal application data node stored in the target metadata node.

[0109] Accordingly, the target metadata node may respond to the query request and return the identifier of the stored abnormal application data node to the user terminal. The user terminal receives the identifier of the abnormal application data node returned by the target metadata node.

[0110] Furthermore, for the user terminal, in step 502, the access request to the distributed storage system may be sent to other application data nodes except the identifier of the abnormal application data node, wherein the other application data nodes are application data nodes belonging to the same storage cluster as the abnormal application data node.

[0111] In this embodiment, the metadata node can predict the application data node that will experience service interruption before the service interruption, and store the identification of the abnormal application data node that is predicted to experience service interruption. By exchanging the identification of the abnormal application data node between the user terminal and the metadata node, the user terminal can actively avoid the abnormal application data node during access to the distributed storage system, and switch the access traffic to other normal application data nodes without the user's perception, thereby reducing the risk of service interruption or even eliminating service interruption.

[0112] On the other hand, there is a certain delay because the user terminal obtains the latest identifier of the abnormal application data node from the metadata node. During the period when the user terminal has not yet updated to the latest abnormal application data node, the access traffic will still be directed to the application data node marked as abnormal. In this embodiment, the metadata node only marks the application data node of the target physical space as abnormal, and does not stop the service of these application data nodes. Therefore, before all the access traffic of the user terminal is directed to other application data nodes that are not marked as abnormal, if the application data nodes accessed by the user terminal include application data nodes marked as abnormal, these application data nodes marked as abnormal can still provide services, which can ensure the service quality of the distributed storage system.

[0113] In the embodiment of the present application, the specific implementation method of sending the access request to other application data nodes except the application data node corresponding to the identifier of the abnormal application data node in step 502 is not limited. In some embodiments, it is possible to detect whether the application data node currently to be accessed includes an abnormal application data node based on the obtained identifier of the abnormal application data node.

[0114] Specifically, when accessing the distributed storage system, the user terminal may obtain the identifier of the application data node to be accessed from the metadata node, such as the address information of the application data node to be accessed. Further, the user terminal may match the obtained identifier of the abnormal application data node with the identifier of the current application data node to be accessed; if the identifier of the current application data node to be accessed includes the identifier of the abnormal application data node, it is determined that the current application data node to be accessed includes the abnormal application data node.

[0115] Furthermore, when the current application data node to be accessed includes an abnormal application data node, stop sending access requests to the abnormal application data node; and determine a new application data node to be accessed from other application data nodes that belong to the same storage cluster as the abnormal application data node. In the embodiment of the present application, for the convenience of description and distinction, the original application data node to be accessed is defined as the first application data node; and the determined new application data node to be accessed is defined as the second application data node.

[0116] Among them, the access operation of the access request is different, and the implementation method of determining the new application data node to be accessed from other application data nodes belonging to the same storage cluster as the abnormal application data node is also different. In some embodiments, the access operation of the access request is a write operation, and the access request is a write request. The above-mentioned determination of the new application data node to be accessed from other application data nodes belonging to the same storage cluster as the abnormal application data node can be implemented as follows: the application data node to be written can be re-requested from the target metadata node; and the identifier of the application data node returned by the target metadata node is received; and the application data node corresponding to the identifier of the application data node returned by the target metadata node is determined to be the second application data node.

[0117] Accordingly, the target metadata node can determine a new application data node to be written as the second application data node from other application data nodes belonging to the same storage cluster as the abnormal application data node in response to the user terminal's request for the application data node to be written.

[0118] In the embodiment of the present application, the specific implementation method of determining a new application data node to be written from other application data nodes belonging to the same storage cluster as the abnormal application data node is not limited.

[0119] In some embodiments, the amount of data to be written in a write request has requirements on the idle storage resources of the application data node, that is, the idle storage resource amount of the application data node should meet the storage resource requirements of the write request, such as the idle storage resource amount of the application data node is greater than the storage resource amount required by the write request. Accordingly, the application data node whose idle storage resource amount meets the storage resource requirements of the write request can be determined from other application data nodes belonging to the same storage cluster as the abnormal application data node. If there is one application data node whose idle storage resource amount meets the storage resource requirements of the write request, the application data node whose idle storage resource amount meets the storage resource requirements of the write request is determined to be the second application data node.

[0120] If there are multiple application data nodes whose idle storage resources meet the storage resource requirements of the write request, the application data node whose idle storage resources meet the storage resource requirements of the write request can be selected from the application data nodes whose idle storage resources meet the storage resource requirements of the write request, and whose distance from the user terminal meets the set distance condition, as the second application data node. The set distance condition can be that the distance is less than or equal to the set distance threshold. Accordingly, the application data node whose distance from the user terminal meets the set distance threshold can be selected from the application data nodes whose idle storage resources meet the storage resource requirements of the write request, as the second application data node. Alternatively, the distance condition can be set to the shortest distance. Accordingly, the application data node whose idle storage resources meet the storage resource requirements of the write request can be selected from the application data nodes whose idle storage resources meet the storage resource requirements of the write request, and whose distance from the user terminal meets the set distance threshold, as the second application data node, and so on.

[0121] In other embodiments, based on the storage water levels of other application data nodes belonging to the same storage cluster as the abnormal application data node, an application data node whose storage water level meets the set water level conditions can be selected from the application data nodes whose free storage resources meet the storage resource requirements of the write request as the second application data node.

[0122] The method for determining the second application data node shown in the above embodiment is only an exemplary description and does not constitute a limitation. After determining the identifier of the second application data node, the target metadata node can return the identifier of the second application data node to the user terminal. Accordingly, the user terminal can receive the identifier of the application data node returned by the target metadata node; and determine that the application data node corresponding to the identifier of the application data node returned by the target metadata node is the second application data node.

[0123] Furthermore, for the user terminal, a write request may be sent to the second application data node, and the second application data node responds to the write request and writes the to-be-written data carried in the write request to the second application data node.

[0124] In other embodiments, the access operation of the access request is a read operation, and the access request is a read request. When the user terminal determines the second application data node to be accessed from other application data nodes belonging to the same storage cluster as the abnormal application data node, the user terminal can determine the application data node belonging to the first application data node originally to be read from other application data nodes belonging to the same storage cluster as the abnormal application data node as the second application data node. That is, for a read request, the user terminal only reads data from other application data nodes in the application data nodes to be read by the read request except the abnormal application data node, thereby avoiding directing the read request to the abnormal application data node.

[0125] The above embodiment takes the access request being a read request or a write request as an example to exemplify the implementation method of determining the second application data node to be accessed from other application data nodes belonging to the same storage cluster as the abnormal application data node, but does not constitute a limitation.

[0126] After determining the second application data node to be accessed by the user terminal, the user terminal may send an access request to the second application data node. The second application data node may respond to the access request and perform an access operation according to the access request.

[0127] In an embodiment of the present application, the service of the target physical space where the service interruption occurs will be restored. In the case where the target physical space restores the service, the embodiment of the present application can also re-direct the access traffic to the application data node of the target physical space. Specifically, for the operation and maintenance node, when the target physical space restores the service, a service recovery prompt information can be sent to the target metadata node; the service recovery prompt information includes the identifier of the target physical space. Accordingly, the target metadata node can delete the identifier of the application data node in the target physical space from the stored identifiers of the abnormal application data node in response to the service recovery prompt information. In this way, when the user terminal subsequently accesses the application data node, the access request of the user terminal can be directed to the application data node in the target physical space to restore the service of the target physical space.

[0128] For example, when a user terminal generates a new access request to a distributed storage system, if the identifier of the application data node in the target physical space no longer exists in the identifier of the abnormal application data node obtained from the target metadata node, the new access request can be sent to the application data node in the target physical space to restore access to the target physical space.

[0129] It should be noted that the execution subject of each step of the method provided in the above embodiment can be the same device, or the method can be executed by different devices. For example, the execution subject of steps 401 and 402 can be device A; for another example, the execution subject of step 401 can be device A, and the execution subject of step 402 can be device B; and so on.

[0130] In addition, in some of the processes described in the above embodiments and the accompanying drawings, multiple operations appearing in a specific order are included, but it should be clearly understood that these operations may not be executed in the order in which they appear in this document or may be executed in parallel, and the sequence numbers of the operations, such as 401, 402, etc., are only used to distinguish between different operations, and the sequence numbers themselves do not represent any execution order. In addition, these processes may include more or fewer operations, and these operations may be executed in sequence or in parallel.

[0131] Accordingly, an embodiment of the present application also provides a computer-readable storage medium storing computer instructions. When the computer instructions are executed by one or more processors, the one or more processors are caused to execute the steps in the exception handling method provided in the above embodiments.

[0132] Figure 6 This is a schematic diagram of the structure of a computing device provided in an embodiment of the present application. Figure 6 As shown, the computing device includes: a memory 60a, a processor 60b and a communication component 60c. The memory 60a is used to store computer programs.

[0133] The processor 60b is coupled to the memory 60a and the communication component 60c, and is used to execute the computer program to execute the steps of the exception handling method provided in the above embodiments. The detailed description of each method step can be found in the above embodiments, which will not be repeated here.

[0134] In some optional embodiments, such as Figure 6 As shown, the computing device may also include optional components such as a power component 60d, a display component 60e, and an audio component 60f. Figure 6 The components are shown schematically only and do not necessarily include Figure 6 The components shown do not necessarily mean that the computing device can only include Figure 6 Components shown.

[0135] in addition, Figure 6 The components in the dashed box are optional components, not mandatory components, and may depend on the product form of the electronic device. The computing device of this embodiment can be implemented as a terminal device such as a desktop computer, a laptop computer, a mobile phone, or an Internet of Things device; it can also be a traditional server, a cloud server, or a server cluster.

[0136] In an embodiment of the present application, the memory is used to store a computer program and can be configured to store various other data to support operations on the device where it is located. Among them, the processor can execute the computer program stored in the memory to implement the corresponding control logic. The memory can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (Static Random-Access Memory, SRAM), electrically erasable programmable read only memory (Electrically Erasable Programmable Read Only Memory, EEPROM), erasable programmable read only memory (Electrical Programmable Read Only Memory, EPROM), programmable read only memory (Programmable Read Only Memory, PROM), read only memory (Read Only Memory, ROM), magnetic memory, flash memory, magnetic disk or optical disk.

[0137] In the embodiment of the present application, the processor can be any hardware processing device that can execute the logic of the above method. Optionally, the processor can be a central processing unit (CPU), a graphics processing unit (GPU) or a microcontroller unit (MCU); it can also be a field programmable gate array (FPGA), a programmable array logic device (PAL), a general array logic device (GAL), a complex programmable logic device (CPLD) and other programmable devices; or an advanced reduced instruction set (RISC) processor (Advanced RISC Machines, ARM) or a system on chip (SoC), etc., but not limited to this.

[0138] In an embodiment of the present application, the communication component is configured to facilitate wired or wireless communication between the device in which it is located and other devices. The device in which the communication component is located can access a wireless network based on a communication standard, such as Wireless Fidelity (WiFi), 2G or 3G, 4G, 5G or a combination thereof. In an exemplary embodiment, the communication component receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component can also be based on Near Field Communication (NFC) technology, Radio Frequency Identification (RFID) technology, Infrared Data Association (IrDA) technology, Ultra Wide Band (UWB) technology, Bluetooth (BT) technology or other technologies.

[0139] In an embodiment of the present application, the display component may include a liquid crystal display (LCD) and a touch panel (TP). If the display component includes a touch panel, the display component may be implemented as a touch screen to receive input signals from a user. The touch panel includes one or more touch sensors to sense touch, slide, and gestures on the touch panel. The touch sensor may not only sense the boundaries of a touch or slide action, but also detect the duration and pressure associated with the touch or slide operation.

[0140] In an embodiment of the present application, a power supply component is configured to provide power to various components of the device in which it is located. The power supply component may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power to the device in which the power supply component is located.

[0141] In an embodiment of the present application, the audio component may be configured to output and / or input audio signals. For example, the audio component includes a microphone (Microphone, MIC), and when the device where the audio component is located is in an operating mode, such as a call mode, a recording mode, and a speech recognition mode, the microphone is configured to receive an external audio signal. The received audio signal may be further stored in a memory or sent via a communication component. In some embodiments, the audio component also includes a speaker for outputting an audio signal. For example, for a device with a language interaction function, voice interaction with a user can be achieved through an audio component.

[0142] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, occupation and processing of relevant data must comply with the relevant laws, regulations and standards of the relevant countries and regions, and provide corresponding operation entrances for users to choose to authorize or refuse.

[0143] It should also be noted that the descriptions such as "first" and "second" in this article are used to distinguish different messages, devices, modules, etc., and do not represent the order of precedence, nor do they limit "first" and "second" to different types.

[0144] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the present application may adopt the form of a computer program product implemented in one or more computer-usable storage media (including but not limited to disk storage, compact disc read-only memory (Compact Disc Read-Only Memory, CD-ROM), optical storage, etc.) that contain computer-usable program code.

[0145] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0146] These computer program instructions may also be stored in a computer readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture including an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.

[0147] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process in the computer or other programmable device. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.

[0148] In a typical configuration, a computing device includes one or more processors (CPU, etc.), input / output interfaces, network interfaces, and memory.

[0149] Memory may include non-permanent storage in a computer-readable medium, random-access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. Memory is an example of a computer-readable medium.

[0150] The storage medium of a computer is a readable storage medium, which may also be referred to as a readable medium. The readable storage medium includes permanent and non-permanent, removable and non-removable media that can be used to store information by any method or technology. The information can be a computer-readable instruction, a data structure, a module of a program, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, CD-ROM, digital versatile disc (DVD) or other optical storage, magnetic cassettes, disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include temporary computer-readable media (transitory media), such as modulated data signals and carrier waves.

[0151] It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of further restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the process, method, commodity or device including the above elements.

[0152] The above contents are only embodiments of the present application and are not intended to limit the present application. For those skilled in the art, the present application may have various changes and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application should be included within the scope of the claims of the present application.

Claims

1. An exception handling system, characterized in that: include: A distributed storage system, a user terminal, and an operation and maintenance node of the distributed storage system; The distributed storage system comprises: a storage cluster; the storage cluster comprises: a plurality of storage units; the plurality of storage units are deployed in a plurality of physical spaces; the storage units correspond to the physical spaces one by one; the storage units comprise: a metadata node and an application data node; the metadata nodes in the plurality of storage units are backed up with each other and used to store the metadata of the storage cluster; the application data nodes in the plurality of storage units are backed up with each other and used to store application data; The operation and maintenance node is used to send abnormal prompt information to the target metadata node in the multiple physical spaces in response to the service interruption warning information for the target physical space in the multiple physical spaces; the abnormal prompt information includes the identifier of the target physical space; The target metadata node is used to mark the application data node in the target physical space as an abnormal application data node in response to the abnormal prompt information; and store the identifier of the abnormal application data node; The user terminal is used to obtain the identifier of the abnormal application data node from the target metadata node; and send an access request to other application data nodes except the application data node corresponding to the identifier of the abnormal application data node.

2. The system according to claim 1, characterized in that The physical space is a data center, a computer room or a cabinet.

3. An exception handling method, applicable to an operation and maintenance node of a distributed storage system, characterized in that: The distributed storage system comprises: a storage cluster; the storage cluster comprises: a plurality of storage units; the plurality of storage units are deployed in a plurality of physical spaces; the storage units correspond to the physical spaces one by one; the storage units comprise: a metadata node and an application data node; the metadata nodes in the plurality of storage units are backed up with each other and used to store the metadata of the storage cluster; the application data nodes in the plurality of storage units are backed up with each other and used to store application data; The method comprises: Obtaining service interruption warning information for a target physical space among the multiple physical spaces; In response to the service interruption warning information, abnormal prompt information is sent to the target metadata nodes in the multiple physical spaces, so that the target metadata nodes can respond to the abnormal prompt information, mark the application data nodes in the target physical space as abnormal application data nodes, and store the identifiers of the abnormal application data nodes, so that the user terminal can send access requests to other application data nodes except the application data nodes corresponding to the identifiers of the abnormal application data nodes.

4. The method according to claim 3, characterized in that The service interruption warning information includes: abnormal warning information for the target physical space, and / or, Operation and maintenance forecast information for the target physical space.

5. The method according to claim 3, characterized in that: Also includes: When the target physical space resumes service, sending service resumption prompt information to the target metadata node; The service recovery prompt information includes the identifier of the target physical space, so as to delete the identifier of the application data node in the target physical space from the identifier of the abnormal application data node.

6. An exception handling method, applicable to a target metadata node in a distributed storage system, characterized in that: The distributed storage system comprises: a storage cluster; the storage cluster comprises: a plurality of storage units; the plurality of storage units are deployed in a plurality of physical spaces; the storage units correspond to the physical spaces one by one; the storage units comprise: a metadata node and an application data node; the metadata nodes in the plurality of storage units are backed up with each other and used to store the metadata of the storage cluster; the application data nodes in the plurality of storage units are backed up with each other and used to store application data; The method comprises: Acquire abnormal prompt information sent by the operation and maintenance node of the distributed storage system; the abnormal prompt information is generated by the operation and maintenance node in response to service interruption warning information for a target physical space among the multiple physical spaces, and the abnormal prompt information includes an identifier of the target physical space; In response to the abnormal prompt information, marking the application data node in the target physical space as an abnormal application data node; The identifier of the abnormal application data node is stored so that the user terminal can obtain the identifier of the abnormal application data node, and an access request is sent to other application data nodes except the application data node corresponding to the identifier of the abnormal application data node.

7. The method according to claim 6, characterized in that The access request is a write request; the method further comprises: In response to a request by the user terminal for an application data node to be written, determining a second application data node from the other application data nodes; The identifier of the second application data node is returned to the user terminal, so that the user equipment sends the access request to the second application data node.

8. The method according to claim 7, characterized in that The determining the second application data node from the other application data nodes comprises: Determine, from the other application data nodes, an application data node whose idle storage resource amount meets the storage resource requirement of the write request; and, from the application data nodes that meet the storage resource requirement of the write request, determine, as the second application data node, an application data node whose distance to the user terminal meets a set distance condition; or, From the other application data nodes, determine an application data node whose idle storage resource amount meets the storage resource requirement of the write request; from the application data nodes whose idle storage resource amount meets the storage resource requirement of the write request, determine an application data node whose storage water level meets the set water level condition as the second application data node.

9. The method according to any one of claims 6 to 8, characterized in that: Also includes: Obtain service recovery prompt information sent by the operation and maintenance node; The service recovery prompt information is sent by the operation and maintenance node when the target physical space recovers the service; In response to the service recovery prompt information, the identifier of the application data node in the target physical space is deleted from the identifier of the abnormal application data node, so that when the user terminal accesses the application data node, the access request of the user terminal is directed to the application data node in the target physical space.

10. An exception handling method, applicable to a user terminal, characterized in that: The distributed storage system includes: a storage cluster; the storage cluster includes: a plurality of storage units; the plurality of storage units are deployed in a plurality of physical spaces; the storage units correspond to the physical spaces one by one; the storage units include: a metadata node and an application data node; the metadata nodes in the plurality of storage units are backed up with each other and used to store the metadata of the storage cluster; the application data nodes in the plurality of storage units are backed up with each other and used to store application data; The method comprises: Obtaining an identifier of an abnormal application data node from a target metadata node in the multiple physical spaces; the abnormal application data node is an application data node in a target physical space warned by the service interruption warning information; the target physical space is a portion of the multiple physical spaces; The access request is sent to other application data nodes except the application data node corresponding to the identifier of the abnormal application data node.

11. The method according to claim 10, characterized in that The step of sending the access request to other application data nodes except the application data node corresponding to the identifier of the abnormal application data node includes: According to the identifier of the abnormal application data node, detecting whether the first application data node currently to be accessed includes the abnormal application data node; In a case where the first application data node includes the abnormal application data node, stopping sending access requests to the abnormal application data node; A second application data node is determined from the other application data nodes, and the access request is sent to the second application data node.

12. The method according to claim 11, characterized in that The access request is a write request; and determining the second application data node from the other application data nodes comprises: Re-requesting the target metadata node for the application data node to be written; and receiving the identifier of the application data node returned by the target metadata node; the application data node is determined by the target metadata node from the other application data nodes; Determine that the application data node corresponding to the identifier of the application data node returned by the target metadata node is the second application data node.

13. The method according to claim 11, characterized in that The access request is a read request; and determining the second application data node from the other application data nodes comprises: An application data node belonging to the first application data node is determined from the other application data nodes as the second application data node.

14. A computing device, characterized in that include: Memory, processor and communication components; The memory is used to store computer programs; The processor is coupled to the memory and the communication component, and is configured to execute the computer program to perform the steps in the method according to any one of claims 3 to 13.

15. A computer-readable storage medium storing computer instructions, characterized in that: When the computer instructions are executed by one or more processors, the one or more processors are caused to execute the steps in the method according to any one of claims 3 to 13.

Citation Information

Patent Citations

  • Method and system for distributed data storage with enhanced security, resilience, and control

    CN113994626A

  • Server switching method, MooseFS system and storage medium

    CN115145782A

  • Thermal maintenance method, system and related device

    CN115269325A

  • Fault recovery method and device for distributed file system supporting additional writing

    CN116010149A

  • Techniques for Reading From and Writing to Distributed Data Stores

    US20180285263A1