Alarm method, device and storage medium
By clustering exception 10 requests in the calculation separation architecture, an exception 10 request group is formed, and alarm information is output in groups, the processing pressure problem caused by massive alarm information is solved, and the timely processing and efficiency improvement of alarm information is achieved.
Patent Information
- Application Number
- PCT/IB2024/063187
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-01-25
- Filing Date
- 2024-12-26
- Publication Date
- 2025-07-31
AI Technical Summary
In the separation architecture of the storage and computing, a large number of abnormal 10 requested alarm information leads to high processing pressure and cannot be processed in time, affecting the alarm efficiency.
By determining the request path of the exception 10 request, clustering the exception 10 request group based on the request path, and outputting alarm information in units of the exception 10 request group, reducing duplicate alarms and improving processing efficiency.
It effectively reduces the number of alarm information, ensures the timely processing of alarm information, and improves the efficiency of alarm processing.
Smart Images

Figure IB2024063187_31072025_PF_FP_ABST
Abstract
Description
[0001]TECHNICAL FIELD The present disclosure relates to the field of cloud storage technology, and more particularly to an alarm method, device, and storage medium. Background: Cloud storage can be understood as a type of online storage (Cloud storage). With the development of cloud storage technology, a storage-computing separation architecture, also known as the storage-computing separation architecture, has been proposed based on this technology. Cloud storage technology can be used to implement the storage layer in the storage-computing separation architecture. The storage-computing separation architecture also includes a computing layer. The computing and storage layers are decoupled and connected via a network. Both the computing and storage layers can be implemented as independent distributed systems. Each computing node in the computing layer can access the storage layer via I / O requests to read and write data from the storage nodes in the storage layer. Currently, a full-link anomaly monitoring system is typically deployed in the storage-computing separation architecture to detect and automatically diagnose abnormal I / O requests and generate alarms based on abnormal I / O requests. However, as the number of I / O requests continues to rise, the number of alarms is also increasing. The massive amount of alarm information leads to an alarm accumulation, which creates a significant pressure on alarm processing and prevents timely alarm processing. SUMMARY OF THE INVENTION Various aspects of the present disclosure provide an alarm method, device, and storage medium for improving alarm processing efficiency. An embodiment of the present disclosure provides an alarm method, comprising: upon the occurrence of an alarm triggering event, determining a request path corresponding to each abnormal 10 request, wherein a single request path represents the connectivity between the physical nodes traversed by the corresponding abnormal 10 request; clustering the abnormal 10 requests based on the request path to generate at least one abnormal 10 request group, wherein abnormal 10 requests corresponding to request paths traversing the same physical node are within the same abnormal 10 request group; and outputting alarm information for each abnormal 10 request group. Furthermore, determining the request path corresponding to each abnormal 10 request comprises: obtaining abnormal monitoring information corresponding to each abnormal 10 request; and parsing, from the abnormal monitoring information, identification information and the order of the physical nodes traversed by the corresponding abnormal 10 request to determine the request path corresponding to each abnormal 10 request. Further, obtaining the abnormality monitoring information corresponding to each of the abnormality 10 requests includes: sending an abnormality monitoring information acquisition request to the abnormality monitoring system, wherein the acquisition request carries a target abnormality type, wherein the abnormality monitoring system has diagnosed the abnormality type corresponding to each of the abnormality 10 requests; and receiving the abnormality monitoring information corresponding to the abnormality 10 requests diagnosed as the target abnormality type returned by the abnormality monitoring system.Furthermore, the abnormality monitoring information utilizes trace information. The trace information corresponds one-to-one with each I / O request. The trace information includes multiple span items in a sequential relationship. The span items correspond one-to-one with the physical nodes traversed by the I / O request. The span items include identification information of the corresponding physical nodes. The sequential relationship between the span items in the trace information is used to characterize the order of the physical nodes traversed by the I / O request. Furthermore, abnormal I / O requests are clustered based on request paths to generate at least one abnormal I / O request group. This includes: clustering abnormal I / O requests corresponding to each request path that can be connected through a physical node into an abnormal I / O request group; or, searching for physical nodes that can connect multiple request paths as clustering nodes; and clustering abnormal I / O requests corresponding to each request path that can be connected by a single clustering node into an abnormal I / O request group. Furthermore, clustering the abnormal 10 requests corresponding to each request path that can be connected through the physical node into an abnormal 10 request group includes: constructing an undirected graph with the physical nodes on each request path as vertices and the connectivity relationships between the physical nodes on each request path as edges; searching for connected components in the undirected graph; and clustering the abnormal 10 requests corresponding to each request path contained in a single connected component into an abnormal 10 request group. Furthermore, searching for connected components in the undirected graph includes: upon traversing to a target vertex in the undirected graph, searching for the connected component where the target vertex is located; deleting the connected component where the target vertex is located from the undirected graph; continuously determining the next target vertex from the remaining vertices in the undirected graph and searching for and deleting the corresponding connected component until no remaining vertices exist in the undirected graph; and outputting the searched connected components. Furthermore, outputting alarm information based on the abnormal 10 request group as a unit includes: after deduplicating the physical nodes passed by each abnormal 10 request in the target abnormal 10 request group, determining the remaining physical nodes as the target nodes; based on the identification information of the target node, the identification information of the cluster to which the target node belongs and / or the abnormality type involved in the target abnormal 10 request group, outputting alarm information for the target abnormal 10 request group; wherein the target abnormal 10 request group is any abnormal 10 request group.Furthermore, after deduplicating the physical nodes traversed by each abnormal 10 request within the target abnormal 10 request group, determining the remaining physical nodes as target nodes includes: if an undirected graph is constructed based on each request path and connected components are searched from the undirected graph to cluster the target abnormal 10 request group, then determining the physical nodes represented by each vertex contained in the connected components corresponding to the target abnormal 10 request group as the target nodes. Furthermore, the physical nodes include at least computing nodes and storage nodes. After outputting the alarm information, the method further includes: in response to the alarm processing instruction, analyzing the connectivity structure formed between computing nodes and storage nodes in the target abnormal 10 request group corresponding to the target alarm information; and inferring the abnormal node causing the abnormality in the target abnormal 10 request group based on the directional relationship between the connectivity structure and the abnormal node. Furthermore, based on the directional relationship between the connectivity structure and the abnormal node, inferring the abnormal node causing the abnormality in the target abnormal I / O request group includes: if a first type of connectivity structure exists in the target abnormal I / O request group, inferring a computing node in the first type of connectivity structure as an abnormal node, wherein the first type of connectivity structure is one computing node connected to multiple storage nodes; or if a second type of connectivity structure exists in the target abnormal I / O request group, inferring a storage node in the second type of connectivity structure as an abnormal node, wherein the second type of connectivity structure is one storage node connected to multiple computing nodes; or if a third type of connectivity structure exists in the target abnormal I / O request group, inferring an intermediate node in the third type of connectivity structure as an abnormal node, wherein the third type of connectivity structure is multiple computing nodes and multiple storage nodes connected via an intermediate node. Furthermore, the target abnormality type includes an I / O unavailable type or an I / O damaged type, the physical node includes a computing node in the computing system, a storage node in the storage system, and / or an intermediate node used for network connection, and the abnormal I / O request is an I / O request initiated by a computing node in the computing system to a storage node in the storage system and in which an abnormality has occurred. An embodiment of the present disclosure further provides an electronic device comprising a memory and a processor; the memory is configured to store one or more computer instructions; and the processor is coupled to the memory and configured to execute the one or more computer instructions to perform the aforementioned alarm method. An embodiment of the present disclosure further provides a computer-readable storage medium storing computer instructions. When the computer instructions are executed by one or more processors, the one or more processors are configured to perform the aforementioned data processing method.Embodiments of the present disclosure also provide a computer program product, including a computer program / instructions. When executed by a processor, the computer program causes the processor to implement the aforementioned alarm method. In embodiments of the present disclosure, when an alarm triggering event occurs, a request path is determined for each abnormal request, and the abnormal requests are clustered based on the request path to generate at least one abnormal request group. Abnormal requests corresponding to request paths that pass through the same physical node are included in the same abnormal request group. Furthermore, alarm information is output based on the abnormal request group. Thus, from the perspective of a single physical node where an exception occurs, the request paths of all abnormal I0 requests caused by that physical node all pass through that physical node. This allows the abnormal I0 requests corresponding to these request paths to be clustered into the same abnormal I0 request group. This ensures that abnormal I0 requests caused by the same abnormal cause can be alerted in a single alert message. This avoids duplicate alerts for the same abnormal cause, thereby reducing the number of alert messages, enabling timely processing of alert messages, and improving alert processing efficiency. BRIEF DESCRIPTION OF THE DRAWINGS The drawings described herein are provided to provide a further understanding of the present disclosure and constitute a part of this disclosure. The illustrative embodiments of this disclosure and their descriptions are provided to explain the disclosure and are not intended to unduly limit the disclosure. In the accompanying drawings: Figure 1 is a flowchart of an alarm method provided by an exemplary embodiment of the present disclosure; Figure 2 is a schematic diagram of an exemplary request path for an abnormal 10 request provided by an exemplary embodiment of the present disclosure; Figure 3 is a logical diagram of an exemplary clustering scheme provided by an exemplary embodiment of the present disclosure; Figure 4 is a flowchart of another alarm method provided by an exemplary embodiment of the present disclosure; Figure 5 is a flowchart of yet another alarm method provided by an exemplary embodiment of the present disclosure; Figure 6 is a schematic diagram of several exemplary connectivity structures provided by an exemplary embodiment of the present disclosure; and Figure 7 is a schematic diagram of the structure of an electronic device provided by another exemplary embodiment of the present disclosure. DETAILED DESCRIPTION To further clarify the objectives, technical solutions, and advantages of the present disclosure, the technical solutions of the present disclosure will be described clearly and completely below in conjunction with the specific embodiments of the present disclosure and the corresponding drawings. Obviously, the described embodiments are only some of the embodiments of the present disclosure, and are not exhaustive. All other embodiments derived by persons of ordinary skill in the art based on the embodiments of the present disclosure without inventive effort are within the scope of protection of the present disclosure. As mentioned in the background art, in the current alarm solution for IQ requests, alarm information is usually output in units of 10 requests. Therefore, as the number of IQ requests continues to increase, the resulting alarm information also increases massively.The ever-increasing volume of alarm information places tremendous pressure on alarm processing, hindering timely processing and leading to inefficient alarm handling. To address this issue, embodiments of the present disclosure propose a new alarm method. The basic concept is to cluster abnormal requests to generate abnormal request groups and output alarm information for each abnormal request group. This effectively reduces the number of alarms, ensures timely delivery of alarm information, and improves alarm handling efficiency. The following, combined with the accompanying drawings, details the technical solutions provided by various embodiments of the present disclosure. Figure 1 is a schematic flow chart of an alarm method provided by an exemplary embodiment of the present disclosure. This method can be executed by an alarm device, which can be implemented as software, hardware, or a combination of software and hardware. The alarm device can be integrated into an electronic device. Referring to Figure 1 , the method may include: Step 100: When an alarm triggering event occurs, determining the request path corresponding to each abnormal I / O request, where a single request path represents the connectivity between the physical nodes traversed by the corresponding abnormal I / O request; Step 101: Clustering the abnormal I / O requests based on the request path to generate at least one abnormal I / O request group, wherein abnormal I / O requests corresponding to request paths traversing the same physical node are within the same abnormal I / O request group; Step 102: Outputting alarm information per abnormal I / O request group. The alarm method provided in this embodiment is applicable to various scenarios requiring abnormal I / O request alarms. In different scenarios, the initiator and recipient of an I / O request may differ. For example, in the cloud storage scenario mentioned in the background, the initiator of an I / O request is typically a compute node in a computing system, and the recipient is typically a storage node in a storage system. For I / O requests in other scenarios, further examples of initiators and recipients are not provided here. It should be understood that this embodiment does not limit the application scenario. In application scenarios where I / O requests traverse multiple physical nodes, the alarm method provided in this embodiment can be used to improve alarm processing efficiency. Referring to Figure 1 , in step 100, the alarm triggering event can be the receipt of an alarm triggering instruction, the reaching of a preset alarm period, or other types of events. This embodiment does not limit the event type of the alarm triggering event, and accordingly, does not limit the timing for initiating the alarm method. The abnormal I / O request in step 100 refers to an I / O request that has been diagnosed as abnormal. This embodiment does not limit the abnormality diagnosis process for I / O requests. This embodiment supports the use of existing or future I / O request abnormality monitoring methods to monitor I / O requests for abnormalities, thereby promptly detecting abnormal I / O requests.For example, in cloud storage scenarios, an anomaly monitoring system based on full-link tracing has been deployed. This anomaly monitoring system can generate trace information for each request and automatically diagnose the anomaly type corresponding to the request. Of course, this is merely exemplary; this embodiment also supports the use of other anomaly monitoring methods to proactively detect abnormal requests. In this embodiment, considering that only abnormal requests require alerts, abnormal requests are screened in step 100 as the targets for processing in this embodiment. Based on this, in step 100, the request path corresponding to each abnormal request is determined. The request path represents the connectivity between the physical nodes traversed by the corresponding abnormal request. It should be understood that a request path includes at least the physical node initiating the abnormal request and the physical node responding to the abnormal request. Of course, the request path may also include physical nodes used for intermediate forwarding of the abnormal request. Figure 2 is a schematic diagram of an exemplary request path for an abnormal request, provided in an exemplary embodiment of the present disclosure. Referring to Figure 2, a storage-computing separation architecture may include a computing system and a storage system. The computing system may include computing nodes, and the storage system may include storage nodes. Different cloud storage products may employ different node organizational structures within their provided storage systems. For example, as shown in Figure 2, for a block storage product, at least two types of storage nodes may be deployed within the provided storage system: the first type of storage nodes may be equipped with block servers for scheduling storage resources, while the second type of storage nodes may be equipped with chunk servers for data storage. In other words, data related to requests is stored on the second type of storage nodes in the storage system, while the first type of storage nodes are primarily responsible for scheduling and managing the storage resources on the second type of storage nodes and do not perform data storage. Of course, the storage-computing separation architecture shown in Figure 2 is merely exemplary; the storage system may also not include the first type of storage nodes, and this is not intended to be limiting. Continuing with Figure 2, the exception request 10 shown in Figure 2 is issued by compute node A in the computing system, forwarded by storage node 1 in the storage system, and ultimately reaches storage node 3 in the storage system, where it is responded to by storage node 3. Based on this, the request path corresponding to the exception request 10 can be determined as: compute node A - storage node 1 - storage node 2. It will be appreciated that the request path not only represents the physical nodes that the exception request 10 traverses, but also represents the connectivity between the physical nodes along the way.Referring to Figure 2, the request path corresponding to the abnormal 10 request indicates that computing node A is connected to storage node 1, and storage node 1 is connected to storage node 2. During research, the inventors discovered that the request paths corresponding to different abnormal 10 requests may pass through the same physical node, and the request paths may be connected based on these physical nodes. Based on this, this embodiment proposes, in step 101, clustering the abnormal 10 requests based on the request paths to generate at least one abnormal 10 request group. The abnormal 10 requests corresponding to the request paths that pass through the same physical node are included in the same abnormal 10 request group. It should be understood that the abnormal 10 requests are clustered in step 102 to generate at least one abnormal 10 request group. The clustering is based on the connectivity between the request paths. As mentioned above, the request paths may be connected based on the same physical node. Therefore, the request paths that are connected based on the same physical node can be clustered together, ensuring that the abnormal 10 requests that pass through the same physical node are clustered in the same abnormal 10 request group. In this way, based on the connectivity between request paths, each request path determined in step 100 will be assigned to a clustered abnormal 10 request group. During research, the inventors discovered that a first request path may be connected to multiple other request paths, and the physical nodes through which these multiple request paths connect to the first request path may be different. In a preferred implementation, in step 101, the abnormal 10 requests corresponding to each request path that can be connected based on a physical node are clustered into abnormal 10 request groups. In this preferred implementation, multiple other request paths that can connect to the first request path based on different physical nodes can be clustered into the same abnormal 10 request group. The first request path can be any request path determined in step 100. For example, the first request path may be AB. Based on physical node A, other request paths connected to the first request path include AC and ADF. Based on physical node B, other request paths connected to the first request path include B-Go. Therefore, although request paths BG and AC do not pass through the same physical node, they can be clustered into the same abnormal request group because they both connect to the first request path. In this preferred implementation, various clustering schemes can be used to cluster abnormal requests. Figure 3 is a logical diagram of an exemplary clustering scheme provided in an exemplary embodiment of the present disclosure.Referring to Figure 3 , in this exemplary clustering scheme, an undirected graph is constructed using the physical nodes on each request path as vertices and the connectivity relationships between the physical nodes on each request path as edges. Connected components are searched for within the undirected graph, and the abnormal 10 requests corresponding to each request path contained within a single connected component are clustered into abnormal 10 request groups. This exemplary clustering scheme introduces an undirected graph data structure to represent each request path. A graph is a nonlinear structure consisting of vertices and edges, with edges representing the connectivity between vertices. A graph with undirected edges is considered an undirected graph. In this exemplary clustering scheme, the connectivity relationships between physical nodes in the request paths can be represented using the undirected graph, and whether the request paths are connected. Based on this, this exemplary clustering scheme proposes searching for connected components within the constructed undirected graph. A connected component refers to a maximal subgraph in an undirected graph that is connected. That is, a reachable path exists between any two vertices in a connected component. Figure 3 shows the undirected graph constructed in this exemplary clustering scheme. It should be understood that not all vertices in this undirected graph have a reachable path. For example, there is no reachable path between vertex 6 and vertex 7. Figure 3 also shows three connected components searched from the undirected graph. In this exemplary clustering scheme, the implementation logic for searching for connected components may be: upon traversing to a target vertex in the undirected graph, search for the connected component where the target vertex is located; delete the connected component where the target vertex is located from the undirected graph; continue to determine the next target vertex from the remaining vertices in the undirected graph and search and delete the corresponding connected component until no vertices remain in the undirected graph; and output the searched connected components. In practical applications, the IP address of a physical node can be used as identification information. In this way, each vertex in an undirected graph can be labeled as an IP address. Based on this, each IP address in the undirected graph can be traversed. When the target IP address is reached, the connected component corresponding to the target vertex can be searched. The search algorithm used to search for the connected component where the target vertex is located is not limited here; for example, a depth-first search (DFS) algorithm can be used. The search principle is not detailed here. The searched connected component can then be deleted from the undirected graph, and the next target IP address can be determined from the remaining IP addresses. This cycle can be repeated to search for all connected components in the undirected graph.In this way, the connectivity between each request path can be accurately characterized based on the undirected graph, and the clustering problem between request paths can be converted into the problem of searching for connected components in the undirected graph, thereby effectively improving the clustering efficiency between request paths. Clustering between request paths is essentially clustering between abnormal 10 requests. Therefore, by introducing the undirected graph, the clustering efficiency of abnormal 10 requests can be effectively improved, and abnormal 10 requests that pass through the same physical node can be clustered into the same abnormal 10 request group. It is worth noting that the clustering scheme provided in Figure 3 is merely exemplary. In this embodiment, other clustering schemes can also be used to implement the clustering of abnormal 10 requests in step 101. For example, each request path can be traversed. When traversing to the target request path, other request paths that share the same physical node as the target request path are searched, and the connectivity relationships between the target request path and these other request paths are recorded. Then, the next request path is traversed again. This cycle is repeated to obtain the request paths that each request path connects to. On this basis, a request path can be selected as the starting path, and this starting path and its connected request paths can be added to a path group. Subsequently, each request path connected to this request path can be selected as a starting path, and the request paths connected to each starting path can be added to the path group. Request paths already in the path group do not need to be added repeatedly. This cycle can be repeated to cluster all connected request paths into the same path group. This clustering scheme can achieve the same clustering effect as shown in Figure 3. Further examples of clustering schemes are not provided here, and this embodiment is not limited to this. In the above preferred implementation, abnormal 10 requests corresponding to each directly or indirectly connected request path can be clustered into the same abnormal 10 request group. This not only ensures that abnormal 10 requests passing through the same physical node are clustered into the same abnormal 10 request group, but also ensures that each abnormal 10 request appears in only one abnormal 10 request group. Therefore, repeated analysis of the same abnormal 10 request can be avoided, thereby further reducing the number of clustered abnormal 10 request groups. In addition to the preferred implementations described above, in this embodiment, other implementations may be used in step 101 to cluster abnormal 10 request groups. For example, in another optional implementation, physical nodes that can connect multiple request paths are queried as clustering nodes; the abnormal 10 requests corresponding to each request path that can be connected by a single clustering node are clustered into abnormal 10 request groups. This implementation also ensures that abnormal 10 requests that pass through the same physical node are clustered into the same abnormal 10 request group. This allows subsequent alarm information to be directed to a single abnormal cause, thereby facilitating alarm processing.However, the number of clustered abnormal 10 request groups will be greater than that of the aforementioned preferred implementation. Furthermore, in this optional implementation, the aforementioned undirected graph approach can also be used to cluster heterogeneous 10 request groups. For example, when traversing to a target vertex in the undirected graph, if multiple request paths pass through the target vertex, the abnormal 10 requests corresponding to the multiple request paths passing through the target vertex are clustered into an abnormal 10 request group. The multiple request paths passing through the target vertex are deleted from the undirected graph. From the remaining vertices in the undirected graph, the next target vertex is determined, and the multiple request paths passing through the target vertex are searched and deleted until no vertices remain in the undirected graph, thereby obtaining at least one abnormal 10 request group. It will be understood that regardless of the implementation approach, in this embodiment, at least one abnormal 10 request group can be clustered in step 101. Moreover, from the perspective of a single physical node, abnormal 10 requests passing through the same physical node can be clustered into the same abnormal 10 request group. During research, the inventors discovered that when a physical node experiences an anomaly, I / O requests that must pass through that physical node are likely to experience an anomaly. In step 101 of this embodiment, the anomaly I / O requests that pass through that physical node are clustered into the same anomaly I / O request group. This essentially means that the anomaly I / O requests resulting from the same anomaly cause are clustered into the same anomaly I / O request group. Based on this, with reference to FIG1 , this embodiment proposes, in step 102, outputting alarm information based on an anomaly I / O request group. That is, one alarm message is output for each anomaly I / O request group. In this embodiment, the alarm content included in the alarm information is not limited; it only needs to provide the necessary content required for alarm processing within the anomaly I / O request group. This embodiment supports on-demand configuration of the content field in the alarm information. In step 102, corresponding alarm content is generated based on the required content field in the alarm information and encapsulated in the corresponding content field in the alarm information to generate and output the alarm information. The alarm information construction scheme will be exemplified later and will not be detailed here. Following the aforementioned clustering of 10 abnormal requests resulting from the same abnormal cause into the same 10-request group, in step 102, all 10 abnormal requests resulting from the same abnormal cause can be issued in a single alarm message, thus avoiding duplicate alarms for the same abnormal cause. During research, the inventors discovered that the number of 10-request groups clustered in step 101 of this embodiment is far smaller than the number of 10 abnormal requests. Therefore, the number of alarm messages output in step 102 is far smaller than the number of alarm messages output per 10-request unit in conventional methods.In summary, this embodiment proposes determining request paths for each abnormal request in the event of an alarm triggering event, clustering the abnormal requests based on the request paths to generate at least one abnormal request group. The abnormal requests corresponding to request paths that pass through the same physical node are placed in the same abnormal request group. Furthermore, alarm information is output based on abnormal request groups. Thus, from the perspective of a single abnormal physical node, the request paths of the abnormal requests caused by that physical node all pass through that physical node. This allows the abnormal requests corresponding to these request paths to be clustered in the same abnormal request group. Consequently, abnormal requests caused by the same abnormal cause can be reported in a single alarm message. This avoids duplicate alarms for the same abnormal cause, reduces the number of alarm messages, and enables timely processing of alarm messages, improving alarm processing efficiency. FIG4 is a flow diagram of another alarm method provided by an exemplary embodiment of the present disclosure. Referring to FIG4 , the method may include: Step 400: Upon occurrence of an alarm triggering event, obtaining abnormality monitoring information corresponding to each abnormal 10 request; Step 401: Parsing the abnormality monitoring information for identification information of the physical nodes traversed by the corresponding abnormal 10 request and the order of their traversal to determine the request path corresponding to each abnormal 10 request; Step 402: Clustering the abnormal 10 requests based on the request path to generate at least one abnormal 10 request group, wherein abnormal 10 requests corresponding to request paths traversing the same physical node are within the same abnormal 10 request group; and Step 403: Outputting alarm information for each abnormal 10 request group. Steps 402 and 403 may be referred to the relevant descriptions in the preceding embodiments and are not repeated here. This embodiment provides an optional implementation method for determining the request path corresponding to the abnormal 10 request based on Steps 400 and 401. This optional implementation can be combined with the implementations provided for other steps in the above or following embodiments to create a new technical solution. Referring to Figure 4 , in this optional implementation, abnormality monitoring information corresponding to each abnormal IQ request can be obtained. As mentioned above, this embodiment supports various abnormality monitoring methods for IQ requests, all of which can generate abnormality monitoring information. Preferably, this embodiment employs an abnormality monitoring method that can generate abnormality monitoring information in units of IQ requests. Of course, this is merely a preference. When abnormality monitoring data is generated in other units, this embodiment supports organizing this abnormality monitoring data into abnormality monitoring information in units of IQ requests.In one exemplary solution, the exception monitoring information uses trace information. The trace information corresponds one-to-one with each I0 request and contains multiple sequentially related span items. Each span item corresponds one-to-one with each physical node traversed by the I0 request. Each span item contains identification information for the corresponding physical node. The sequential relationship between the span items in the trace information is used to characterize the order of the physical nodes traversed by the I0 request. In this exemplary solution, in step 400, trace information corresponding to each abnormal I0 request can be obtained from an exception monitoring system based on full-link tracing. As mentioned above, this exception monitoring system can generate trace information per I0 request and automatically diagnose the type of exception corresponding to the I0 request. Furthermore, this exception monitoring system constructs a span item for each physical node traversed by the I0 request to describe the I0 processing process on the physical node. The sequential relationship between the span items in the trace information can serve as a basis for determining the order of the physical nodes traversed by the I0 request. It should be understood that this is merely exemplary. In this embodiment, other types of exception monitoring information may be used, and the present invention is not limited thereto. Based on this, referring to FIG4 , step 401 proposes parsing the identification information and the order of the physical nodes traversed by the corresponding abnormal IOC request from the exception monitoring information to determine the request path corresponding to each abnormal IOC request. During research, the inventors discovered that the exception monitoring information typically records IOC processing descriptions on each physical node traversed by the abnormal IOC request. This IOC processing description typically includes information such as the identification information of the physical node, the identification information of the cluster to which the physical node belongs, the IOC processing time on the physical node, the identification information of the previous hop node of the physical node, the identification information of the next hop node of the physical node, and the type of IOC processing operation performed on the physical node. For example, the span item mentioned above records this IOC processing description information on the physical node. Based on this, in step 401, the identification information and path sequence of the physical nodes traversed by the corresponding abnormal I / O request can be parsed from the abnormality monitoring information, thereby determining the request path corresponding to the abnormal I / O request. Furthermore, during research, the inventors discovered that the exception types corresponding to different abnormal I / O requests may not be exactly the same. In practical applications, I / O request exception types may include, but are not limited to, I / O unavailable and I / O damaged. The I / O unavailable type typically indicates an incomplete I / O request response, while the I / O damaged type typically indicates an I / O request response completed but at a slow rate.It should be understood that the several exception types provided here are merely exemplary and the present embodiment is not limited thereto. Furthermore, the inventors have discovered that the exception type serves as an important reference for subsequent alarm processing. Therefore, in this optional implementation, an exemplary acquisition scheme is proposed for obtaining the exception monitoring information corresponding to each exception request. An exception monitoring information acquisition request may be sent to an exception monitoring system deployed for the storage system. The acquisition request carries a target exception type, wherein the exception monitoring system has already diagnosed the exception type corresponding to each exception request. The exception monitoring system then returns the exception monitoring information corresponding to each exception request diagnosed as the target exception type. In this exemplary acquisition scheme, the target exception type is included in the exception monitoring information acquisition request sent to the exception monitoring system. Since the exception monitoring system has already diagnosed the exception type corresponding to each exception request, the exception monitoring system can filter out the exception requests diagnosed as the target exception type and return the exception monitoring information corresponding to the filtered exception requests. Thus, based on this exemplary acquisition scheme, the alarm solution provided in this embodiment first performs a one-level clustering of abnormal requests based on the abnormality type, clustering abnormal requests diagnosed as belonging to the same abnormality type. Secondly, following the concept of clustering based on the request path provided in step 402, abnormal requests with different abnormality types can be further clustered. This allows for clustering of abnormal request groups based on different abnormality types. In other words, each abnormal request within the same abnormal request group will correspond to the same abnormality type. Consequently, the single alarm message output in step 403 will only involve one abnormality type. This provides a reference for abnormality type information in subsequent alarm processing steps. Of course, in this optional implementation, other exemplary schemes can also be adopted to support the display of abnormality types in alarm messages. For example, abnormality monitoring data corresponding to all abnormal requests can be obtained from the abnormality monitoring system and processed uniformly according to steps 401 through 403. However, in step 403, the individual exception request groups can be re-clustered by exception type, and the exception requests associated with different exception types can be recorded in the alarm information. This can also provide a reference for the exception type in subsequent alarm processing. Further examples of solutions that support displaying the exception type in the alarm information are not provided here, and this embodiment is not limited thereto.In summary, this embodiment can obtain the exception monitoring information corresponding to each of the 10 abnormal requests, and accurately determine the request path corresponding to each of the 10 abnormal requests based on the exception monitoring information, thereby providing an accurate basis for clustering the 10 abnormal requests. Furthermore, it is proposed that the 10 abnormal requests can be clustered one level at a time based on the abnormality type, thereby enabling reasonable presentation of the abnormality type in the alarm information and providing a reference for subsequent alarm processing, further improving alarm processing efficiency. In the above and following embodiments, alarm information can be constructed using various implementations. Since the logic for constructing alarm information for each 10 abnormal request group is consistent, for ease of description, the following description of the alarm information construction scheme uses the target 10 abnormal request group as an example. It should be understood that the target 10 abnormal request group can be any clustered 10 abnormal request group. In one optional implementation, after deduplication of the physical nodes traversed by each abnormal request within a target abnormal request group, the remaining physical nodes are identified as target nodes. Based on the identification information of the target node, the identification information of the cluster to which the target node belongs, and / or the abnormality type involved in the target abnormal request group, an alarm message is output for the target abnormal request group. As mentioned above, different request paths may traverse the same physical node. Therefore, this optional implementation proposes deduplication of the physical nodes. Through deduplication, the physical nodes involved in the target abnormal request group can be determined. In an exemplary deduplication scheme, if an undirected graph is constructed based on each request path and connected components are searched in the undirected graph to cluster the target abnormal request group, the physical nodes represented by each vertex contained in the connected component corresponding to the target abnormal request group are used as the target nodes. It will be appreciated that in this optional implementation, the alarm information may include at least a content field for recording the identification information of the target node, identification information of the cluster to which the target node belongs, and / or a content field for recording the type of exception involved in the target abnormal 10 request group. Of course, these content fields are merely exemplary, and this embodiment is not limited thereto. The aforementioned identification information of the target node and identification information of the cluster to which the target node belongs can be obtained from the abnormality monitoring information corresponding to each abnormal 10 request included in the target abnormal 10 request group, for example, from the span item mentioned above.Regarding the exception types involved in the target exception 10 requests, reference can be made to the exemplary solution provided above. Before clustering the exception 10 requests based on the request path, the exception 10 requests can first be clustered based on the exception type. In this case, each exception 10 request clustered based on the request path under the target exception type can be marked as the target exception type, and the exception type marked for the target exception 10 request group can be included in the alarm information. Furthermore, if, after clustering the exception 10 requests based on the request path, further clustering is performed within the target exception 10 request group based on the exception type, the clustered exception types can be marked for the target exception 10 request group, and the exception 10 requests under each exception type can be recorded separately. The exception type marked for the target exception 10 request group and the exception 10 requests under each exception type can then be included in the alarm information. It should be understood that other implementations can also be used to construct the alarm information in this embodiment, and the content fields included in the alarm information are not limited to the exemplary content fields provided above. Based on the alarm information construction solution provided in this embodiment, alarm information can be used to indicate which clusters and / or nodes have experienced which type of I / O anomaly. It can be seen that the alarm information in this embodiment does not indicate anomalies based on I / O requests, but rather on anomaly type, cluster, and node. This facilitates anomaly location during subsequent alarm processing, thereby further improving alarm processing efficiency. Figure 5 is a flowchart of another alarm method provided in an exemplary embodiment of the present disclosure. Referring to Figure 5, the method may include: Step 500: When an alarm triggering event occurs, determine the request path corresponding to each abnormal I / O request, where a single request path represents the connectivity between the physical nodes traversed by the corresponding abnormal I / O request; Step 501: Cluster the abnormal I / O requests based on the request path to generate at least one abnormal I / O request group, where abnormal I / O requests corresponding to request paths traversing the same physical node are within the same abnormal I / O request group; Step 502: Output alarm information based on the abnormal I / O request group. Step 503: In response to the alarm handling instruction, analyze the connectivity structure between the computing nodes and the storage nodes in the target abnormal 10 request group corresponding to the target alarm information. Step 504: Infer the abnormal node that caused the abnormality in the target abnormal 10 request group based on the directional relationship between the connectivity structure and the abnormal node. Steps 500 to 502 can be found in the relevant descriptions of the previous embodiment and are not repeated here. This embodiment provides an optional implementation solution after the alarm information is output based on steps 503 and 504.This optional implementation scheme can be combined with the implementation schemes provided for other steps in the above or following embodiments to create a new technical solution. Referring to Figure 5 , after outputting the alarm information, the alarm processing logic for each output alarm information can be initiated in response to an alarm processing instruction. Since the alarm processing logic implemented for different alarm information is consistent, for ease of description, this embodiment uses the target alarm information as an example to illustrate the alarm processing logic. It should be understood that the target alarm information can be any alarm information output in step 502 of this embodiment. Referring to Figure 5 , in this embodiment, considering that the nature of the I0 request can be understood as a data read / write request, there is at least one physical node used for data storage. In this embodiment, the node used for data storage is described as a storage node. Since the root cause of data read / write is typically for computing, the physical node initiating the I0 request is described as a computing node. Thus, in this embodiment, the physical nodes in the request path can include at least computing nodes and storage nodes. As mentioned above, the request path can also include intermediate nodes used for intermediate forwarding of the I0 request, etc. In a storage-computing separation architecture, compute nodes can be located in a computing system, while storage nodes can be located in a storage system. Of course, in other scenarios, the deployment locations of compute nodes and storage nodes are not limited to this, and the deployment locations of compute nodes and storage nodes are not limited here. Based on this, in step 503, it is proposed to analyze the connectivity structure formed between compute nodes and storage nodes for the target abnormal I / O request group corresponding to the target alarm information. In this embodiment, the connectivity structure can be understood as the structure formed based on the request paths connected by a physical node. This physical node can be a compute node, a storage node, or an intermediate node for I / O request forwarding. During research, the inventors discovered that multiple connectivity structures may be analyzed for a single abnormal I / O request group. Figure 6 is a schematic diagram of several exemplary connectivity structures provided in an exemplary embodiment of the present disclosure. Referring to Figure 6, in this embodiment, at least three types of connectivity structures may exist: the first type of connectivity structure is one compute node connected to multiple storage nodes; the second type of connectivity structure is one storage node connected to multiple compute nodes; and the third type of connectivity structure is multiple compute nodes and multiple storage nodes connected via an intermediate node. The intermediate node may be a network device used for network transit, for example. To this end, in step 504, a preconfigured directional relationship between the connectivity structure and the abnormal node is proposed. This directional relationship can be used to guide the location of the abnormal node under different connectivity structures. Thus, in step 504, the abnormal node causing the abnormality can be inferred based on this directional relationship within the target abnormal IO request group.If multiple connectivity structures are analyzed under the target abnormal 10 request group, the abnormal node can be inferred based on the directional relationships under each of the multiple connectivity structures. In one exemplary inference scheme: if the target abnormal 10 request group has a first-type connectivity structure, the compute nodes in the first-type connectivity structure are inferred as abnormal nodes; alternatively, if the target abnormal 10 request group has a second-type connectivity structure, the storage nodes in the second-type connectivity structure are inferred as abnormal nodes; alternatively, if the target abnormal 10 request group has a third-type connectivity structure, the intermediate nodes in the third-type connectivity structure are inferred as abnormal nodes. Referring to Figure 6, for the first-type connectivity structure, if multiple storage nodes are connected to the same compute node, the cause of the anomaly can be preliminarily located in that compute node. Typically, an anomaly in that compute node causes anomalies in the 10 requests corresponding to multiple storage nodes connected to that compute node. Similarly, for the second-type connectivity structure, the cause of the anomaly can be preliminarily located in the storage nodes within that connectivity structure. For the third type of connectivity structure, since all I / O requests between multiple compute nodes and multiple storage nodes have experienced anomalies, the cause of the anomaly can be initially located in the intermediate node connecting the multiple compute nodes and the multiple storage nodes. Typically, an anomaly in the intermediate node causes all I / O requests passing through the intermediate node to experience anomalies. It is worth noting that the alarm processing logic provided in this embodiment is merely exemplary. Furthermore, based on this exemplary alarm processing logic, the abnormal node can be inferred, and the inferred abnormal node can serve as a reference for operation and maintenance. In actual applications, more inference dimensions can be introduced to further refine the inference results provided in this embodiment. Of course, manual analysis and other processing steps can also be incorporated to ensure accurate identification of the anomaly cause. These other inference dimensions and manual analysis logic are not limited herein. In summary, this embodiment provides alarm processing logic after outputting alarm information. This alarm processing logic fully utilizes data such as the request path and abnormal I / O request group generated during the alarm process in this embodiment as analytical basis for the alarm processing logic. Based on this, the connectivity structure can be analyzed under each alarm information, and then the abnormal nodes can be inferred based on the connectivity structure, providing a reference basis for operation and maintenance, so that alarm processing can be completed faster and more accurately, which can further improve the efficiency of alarm processing.It should be noted that some of the processes described in the above embodiments and accompanying drawings include multiple operations that appear in a specific order. However, it should be understood that these operations may be executed in a different order than the order in which they appear herein or in parallel. Operation numbers such as 101 and 102 are merely used to distinguish between different operations and do not represent any specific order of execution. Furthermore, these processes may include more or fewer operations, and these operations may be executed sequentially or in parallel. It should be noted that terms such as "first" and "second" are used herein to distinguish between different connected devices and do not represent a specific order or limit the "first" and "second" to different types. Figure 7 is a schematic structural diagram of an electronic device provided by another exemplary embodiment of the present disclosure. As shown in FIG7 , the electronic device may include a memory 70 and a processor 71. The processor 71 is coupled to the memory 70 and configured to execute a computer program in the memory 70 to: determine, upon an alarm triggering event, a request path corresponding to each abnormal 10 request, wherein a single request path represents the connectivity between the physical nodes traversed by the corresponding abnormal 10 request; cluster the abnormal 10 requests based on the request path to generate at least one abnormal 10 request group, wherein abnormal 10 requests corresponding to request paths traversing the same physical node are within the same abnormal 10 request group; and output alarm information per abnormal 10 request group. In an optional embodiment, when determining the request path corresponding to each abnormal 10 request, the processor 71 may specifically be configured to: obtain abnormality monitoring information corresponding to each abnormal 10 request; and parse, from the abnormality monitoring information, identification information and a path order of the physical nodes traversed by the corresponding abnormal 10 request to determine the request path corresponding to each abnormal 10 request. In an optional embodiment, when the processor 71 obtains the abnormality monitoring information corresponding to each of the abnormality 10 requests, it can be specifically used to: send an abnormality monitoring information acquisition request to the abnormality monitoring system used for performing 10 abnormality monitoring, wherein the acquisition request carries a target abnormality type, wherein the abnormality monitoring system has diagnosed the abnormality type corresponding to each of the abnormality 10 requests; receive the abnormality monitoring information corresponding to the abnormality 10 request that has been diagnosed as the target abnormality type returned by the abnormality monitoring system.In an optional embodiment, the exception monitoring information utilizes trace information. The trace information corresponds one-to-one with each I / O request. The trace information includes multiple span items in a sequential relationship. The span items correspond one-to-one with the physical nodes traversed by the I / O request. The span items include identification information of the corresponding physical nodes. The sequential relationship between the span items in the trace information is used to characterize the order of the physical nodes traversed by the I / O request. In an optional embodiment, when clustering the abnormal I / O requests based on the request paths to generate at least one abnormal I / O request group, the processor 71 may specifically be configured to: cluster the abnormal I / O requests corresponding to each request path that can be connected through a physical node into an abnormal I / O request group; or, alternatively, search for physical nodes that can connect multiple request paths as clustering nodes; and cluster the abnormal I / O requests corresponding to each request path that can be connected by a single clustering node into an abnormal I / O request group. In an optional embodiment, when clustering the abnormal IQ requests corresponding to each request path that can be connected through a physical node into an abnormal IQ request group, the processor 71 may be specifically configured to: construct an undirected graph using the physical nodes on each request path as vertices and the connectivity relationships between the physical nodes on each request path as edges; search for connected components in the undirected graph; and cluster the abnormal IQ requests corresponding to each request path contained in a single connected component into an abnormal IQ request group. In an optional embodiment, when searching for connected components in the undirected graph, the processor 71 may be specifically configured to: upon traversing to a target vertex in the undirected graph, search for the connected component where the target vertex is located; delete the connected component where the target vertex is located from the undirected graph; continue to determine the next target vertex from the remaining vertices in the undirected graph and search for and delete the corresponding connected component until no more vertices remain in the undirected graph; and output the searched connected components. In an optional embodiment, when the processor 71 outputs alarm information in units of abnormal 10 request groups, it can be specifically used to: deduplicate the physical nodes passed by each abnormal 10 request in the target abnormal 10 request group, and then determine the remaining physical nodes as target nodes; based on the identification information of the target node, the identification information of the cluster to which the target node belongs, and / or the abnormality type involved in the target abnormal 10 request group, output alarm information for the target abnormal 10 request group; wherein the target abnormal 10 request group is any abnormal 10 request group.In an optional embodiment, after deduplicating the physical nodes traversed by each abnormal 10 request within the target abnormal 10 request group, the processor 71 may be specifically configured to: If an undirected graph is constructed based on each request path and connected components are searched within the undirected graph to cluster the target abnormal 10 request group, then the physical nodes represented by each vertex contained in the connected components corresponding to the target abnormal 10 request group are selected as the target nodes. In an optional embodiment, the physical nodes include at least computing nodes and storage nodes. After outputting the alarm information, the processor 71 may be further configured to: In response to the alarm processing instruction, analyze the connectivity structure formed between computing nodes and storage nodes in the target abnormal 10 request group corresponding to the target alarm information; and infer the abnormal node causing the abnormality in the target abnormal 10 request group based on the directional relationship between the connectivity structure and the abnormal node. In an optional embodiment, when the processor 71 infers the abnormal node causing the abnormality under the target abnormal 10 request group based on the directional relationship between the connectivity structure and the abnormal node, it can be specifically used to: if there is a first type of connectivity structure under the target abnormal 10 request group, the computing node in the first type of connectivity structure is inferred to be the abnormal node, and the first type of connectivity structure is that one computing node is connected to multiple storage nodes; or, if there is a second type of connectivity structure under the target abnormal 10 request group, the storage node in the second type of connectivity structure is inferred to be the abnormal node, and the second type of connectivity structure is that one storage node is connected to multiple computing nodes; or, if there is a third type of connectivity structure under the target abnormal 10 request group, the intermediate node in the third type of connectivity structure is inferred to be the abnormal node, and the third type of connectivity structure is that multiple computing nodes and multiple storage nodes are connected through intermediate nodes. In an optional embodiment, the target exception type includes an unavailable or damaged class. The physical node includes a computing node in a computing system, a storage node in a storage system, and / or an intermediate node used for network connection. The abnormal I / O request is an I / O request initiated by a computing node in the computing system to a storage node in the storage system and in which an abnormality has occurred. Furthermore, as shown in FIG7 , the electronic device also includes other components, such as a communication component 72 and a power supply component 73. FIG7 only schematically illustrates some components and does not mean that the electronic device only includes the components shown in FIG7 . It is worth noting that the technical details of the above-mentioned electronic device embodiments can be found in the relevant descriptions of the aforementioned method embodiments. To save space, they will not be repeated here, but this should not affect the scope of protection of the present disclosure.Accordingly, embodiments of the present disclosure also provide a computer-readable storage medium storing a computer program. When executed, the computer program can implement the steps of the above-described method embodiments. Accordingly, embodiments of the present disclosure also provide a computer program product. When executed, the computer program in the computer program product can implement the steps of the above-described method embodiments. The memory in FIG. 7 is used to store the computer program and can be configured to store various other data to support operations on the computing platform. Examples of such data include instructions for any application or method operating on the computing platform, contact data, phone book data, messages, images, videos, etc. The memory can be implemented by any type of volatile or non-volatile storage device, or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk, or optical disk. The communication component in Figure 7 is configured to facilitate wired or wireless communication between the device containing the communication component and other devices. The device containing the communication component can access a wireless network based on a communication standard, such as WiFi, 2G, 3G, 4G / LTE, 5G, or other mobile communication networks, or a combination thereof. In one exemplary embodiment, the communication component receives broadcast signals or broadcast-related information from an external broadcast management system via a broadcast channel. In one exemplary embodiment, the communication component also includes a near-field communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on radio frequency identification (RFID), infrared data association (IrDA), ultra-wideband (UWB), Bluetooth (BT), or other technologies. The power supply component in Figure 7 provides power to various components of the device containing the power supply component. The power supply component may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power to the device containing the power supply component. Those skilled in the art will appreciate that embodiments of the present disclosure may be provided as methods, systems, or computer program products. Therefore, the present disclosure may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present disclosure may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.The present disclosure is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present disclosure. It should be understood that each process and / or block in the flowcharts and / or block diagrams, as well as combinations of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, such that execution of the instructions by the processor of the computer or other programmable data processing device produces means for implementing the functions specified in one or more processes in the flowcharts and / or one or more blocks in the block diagrams. These computer program instructions can also be stored in a computer-readable memory capable of directing the computer or other programmable data processing device to operate in a specific manner, such that the instructions stored in the computer-readable memory produce an article of manufacture including instruction means that implement the functions specified in one or more processes in the flowcharts and / or one or more blocks in the block diagrams. These computer program instructions can also be loaded onto a computer or other programmable data processing device, causing the computer or other programmable device to execute a series of operational steps to produce a computer-implemented process. The instructions executed on the computer or other programmable device thus provide steps for implementing the functions specified in one or more flow charts and / or one or more blocks in a block diagram. It should also be noted that the terms "comprise," "include," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, product, or apparatus comprising a series of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, product, or apparatus. Without further limitation, the phrase "comprises a..." does not preclude the presence of additional identical elements in the process, method, product, or apparatus comprising the elements. It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, storage, and display) involved in this disclosure are all information and data authorized by the user or fully authorized by all parties. The collection, use, and processing of relevant data must comply with the relevant laws, regulations, and standards of relevant countries and regions, and corresponding operation portals are provided for users to choose to authorize or reject. The above description is merely an embodiment of the present disclosure and is not intended to limit the present disclosure. Those skilled in the art will appreciate that various modifications and variations of the present disclosure are possible.Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of this disclosure should be included in the scope of protection of this disclosure.
Claims
Claims 1. An alarm method, comprising: In the case of an alarm trigger event, determine the request paths corresponding to the respective abnormal 10 requests. A single request path represents the connectivity relationship between the physical nodes passed by the corresponding abnormal 10 request. Cluster the abnormal 10 requests based on the request paths to generate at least one abnormal 10 request group, where the abnormal 10 requests corresponding to the request paths passing through the same physical node are located in the same abnormal 10 request group. Output the alarm information in units of the abnormal 10 request groups.
2. The method according to claim 1, wherein Determining the request paths corresponding to the respective abnormal 10 requests includes: obtaining the abnormal monitoring information corresponding to the respective abnormal 10 requests; parsing, from the abnormal monitoring information, the identification information and passing order of the physical nodes passed by the corresponding abnormal 10 requests to determine the request paths corresponding to the respective abnormal 10 requests.
3. The method according to claim 2, wherein Obtaining the abnormal monitoring information corresponding to the respective abnormal 10 requests includes: sending an abnormal monitoring information acquisition request to an abnormal monitoring system for 10 abnormal monitoring, where the acquisition request carries a target abnormal type, and the abnormal monitoring system has diagnosed the abnormal types corresponding to the respective abnormal 10 requests; receiving the abnormal monitoring information corresponding to the abnormal 10 requests diagnosed as the target abnormal type returned by the abnormal monitoring system.
4. The method according to claim 2 or 3, wherein, The abnormal monitoring information uses trace information. The trace information corresponds one-to-one with the 10 requests. The trace information contains multiple span items with an order relationship. The span items correspond one-to-one with the physical nodes passed by the 10 requests. The span items contain the identification information of the corresponding physical nodes. The order relationship between the span items included in the trace information is used to represent the passing order between the physical nodes passed by the 10 requests.
5. The method according to claim 1, wherein Clustering the abnormal 10 requests based on the request paths to generate at least one abnormal 10 request group includes: clustering the abnormal 10 requests corresponding to the request paths that can be connected through physical nodes into an abnormal 10 request group, or querying the physical nodes that can connect multiple request paths as clustering nodes; clustering the abnormal 10 requests corresponding to the request paths that can be connected by a single clustering node into an abnormal 10 request group.
6. The method according to claim 5, wherein Clustering the abnormal 10 requests corresponding to the request paths that can be connected through physical nodes into an abnormal 10 request group includes: constructing an undirected graph with the physical nodes on each request path as vertices and the connectivity relationship between the physical nodes on each request path as edges; searching for connected components from the undirected graph. Clustering the abnormal IO requests corresponding to the respective request paths included in a single connected component into an abnormal IO request group.
7. The method according to claim 6, wherein Search for connected components in the undirected graph, including: when traversing to a target vertex in the undirected graph, search for the connected component where the target vertex is located; delete the connected component where the target vertex is located from the undirected graph; continue to determine the next target vertex from the remaining vertices in the undirected graph and search for and delete the corresponding connected component until there are no remaining vertices in the undirected graph; output the searched connected components.
8. The method according to claim 1, wherein Output alarm information in units of abnormal 10 request groups, including: after de-duplicating the physical nodes passed by each abnormal 10 request in the target abnormal 10 request group, determine the remaining physical nodes as target nodes; based on the identification information of the target nodes, the identification information of the clusters to which the target nodes belong, and / or the abnormal types involved in the target abnormal 10 request group, output alarm information for the target abnormal 10 request group; where the target abnormal 10 request group is any abnormal 10 request group.
9. The method according to claim 8, wherein After de-duplicating the physical nodes passed by each abnormal 10 request in the target abnormal 10 request group, determine the remaining physical nodes as target nodes, including: if an undirected graph is constructed based on each request path and connected components are searched from the undirected graph to cluster the target abnormal 0 request group, then use the physical nodes represented by each vertex included in the connected component corresponding to the target abnormal 10 request group as target nodes.
10. The method according to claim 1, wherein The physical nodes at least include computing nodes and storage nodes. After outputting the alarm information, the method further includes: in response to an alarm processing instruction, analyze the connected structure formed between the computing nodes and the storage nodes under the target alarm information corresponding to the target abnormal 10 request group; based on the pointing relationship between the connected structure and the abnormal nodes, infer the abnormal nodes that cause the abnormality under the target abnormal 10 request group.
11. The method according to claim 10, wherein, Based on the pointing relationship between the connected structure and the abnormal nodes, infer the abnormal nodes that cause the abnormality under the target abnormal 10 request group, including: if there is a first type of connected structure under the target abnormal 10 request group, then infer the computing node in the first type of connected structure as an abnormal node, where the first type of connected structure is one computing node connected to multiple storage nodes; or, if there is a second type of connected structure under the target abnormal 10 request group, then infer the storage node in the second type of connected structure as an abnormal node, where the second type of connected structure is one storage node connected to multiple computing nodes; or, if there is a third type of connected structure under the target abnormal 10 request group, then infer the intermediate node in the third type of connected structure as an abnormal node, where the third type of connected structure is multiple computing nodes and multiple storage nodes connected through an intermediate node.
12. The method according to claim 3, wherein, The target abnormal type includes 10 unavailable types and / or 10 affected A loss class, where the physical nodes include computing nodes in a computing system, storage nodes in a storage system, and / or intermediate nodes for network connection, and the abnormal I / O request is an I / O request that has an abnormality and is initiated by a computing node in the computing system to a storage node in the storage system.
13. An electronic device, comprising a memory and a processor; the memory is used for storing one or more computer instructions; the processor is coupled to the memory and is used for executing the one or more computer instructions to execute the alarm method according to any one of claims 1-12.
14. A computer-readable storage medium storing computer instructions, which when executed by one or more processors cause the one or more processors to execute the alarm method according to any one of claims 1-12.
15. A computer program product comprising a computer program / instructions, wherein, When a computer program is executed by a processor, it causes the processor to implement the alarm method according to any one of claims 1-12.
Citation Information
Patent Citations
Abnormal IO request positioning method and system
CN116610480A
Alarm clustering root cause analysis method and device, equipment and storage medium
CN116668264A
Association rule determination method and device and storage medium
CN117221078A
Request processing method and device, equipment and readable storage medium
CN117311619A