Fault monitoring method and device and related equipment

By creating a background data stream that is not affected by the client in a distributed system, and combining it with the monitoring indicators of the service data stream, the problem of low fault monitoring accuracy in the prior art is solved, and accurate fault monitoring of message queue middleware is achieved.

CN120034454APending Publication Date: 2025-05-23SHENZHEN HUAWEI CLOUD COMPUTING TECHNOLOGIES CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510108873.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-21
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

The existing fault monitoring methods have low accuracy in message queue middleware in distributed systems, which are prone to misdiagnosis problems, especially it is difficult to rule out the impact of client and network links.

Method used

By creating a background data stream that is not affected by the client, and combining the monitoring indicators of the background data stream and service data stream to make fault judgments, troubleshoot interference between the client and network links, and accurately monitor the failure of the message queue middleware.

Benefits of technology

Accurate fault monitoring of message queue middleware is realized, avoiding the error of client and network link failures as message queue middleware failures, and improving the accuracy of monitoring results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120034454A_ABST
    Figure CN120034454A_ABST
Patent Text Reader

Abstract

The invention provides a fault monitoring method. According to the method, interference of a part irrelevant to the message queue middleware can be filtered out through the background data stream. The fault monitoring method comprises the steps that a background data stream is created, and the background data stream is not affected by a client; acquiring a monitoring index of the background data stream and monitoring indexes of a plurality of service data streams; and judging whether the message queue middleware has a fault or not according to the monitoring index of the background data stream and the monitoring indexes of the plurality of service data streams. In addition, the invention further provides a corresponding device, a computing device cluster, a computer readable storage medium and a computer program product.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of processing cluster technology, and in particular to a fault monitoring method, device and related equipment. Background Art

[0002] A distributed system can include producers and consumers. Producers are the roles that generate messages to be processed. Consumers are the roles that process messages. In order to reasonably distribute messages to different consumers, distributed systems often also include message queue middleware. Message queue middleware is used to manage and distribute messages.

[0003] Specifically, the message from the producer will be sent to the message queue middleware first. The message queue middleware can assign the appropriate consumer to the message and add the task to the queue corresponding to the consumer. The consumer can get the message from the queue and process the message. The message queue middleware can realize the functions of flexible message allocation and peak load shifting, and is an important module in the distributed system.

[0004] During operation, message queue middleware may fail. Therefore, it is necessary to monitor the message queue middleware for failure. However, existing fault monitoring methods have the problems of low accuracy and misdiagnosis. Summary of the invention

[0005] In view of this, the present application provides a fault monitoring method for accurately determining whether a message queue middleware fails. The present application also provides a corresponding apparatus, a computing device cluster, a computer-readable storage medium, and a computer program product.

[0006] In the first aspect, the present application provides a fault monitoring method. The method can filter out interference from parts that are not related to the message queue middleware through background data streams. Specifically, before monitoring, a background data stream can be created first. The background data stream will not be interfered by the client. Therefore, if the message queue middleware has a problem when scheduling the background data stream, it can be considered that the problem occurs on the server side of the distributed system, and the message queue middleware can be used for troubleshooting. After creating the background data stream, the background data stream can be monitored to obtain monitoring indicators of the background data stream. In addition, the service data stream of the distributed system can also be monitored to obtain monitoring indicators of multiple service data streams. Based on the monitoring indicators of the background data stream and the monitoring indicators of the service data stream, it can be determined whether the message queue middleware has a fault. In this way, a background data stream that is not affected by the client is created, and the background data stream and the service data stream are considered during fault monitoring, so that the influence of the client and the network link can be eliminated, and the message queue middleware can be accurately monitored for faults. In this way, it is possible to avoid identifying the faults of the client and the network link as the faults of the message queue middleware, avoid the faults of other modules in the distributed system from affecting the fault monitoring results, and accurately monitor the message queue middleware for faults.

[0007] In some possible implementations, when creating a background data stream, you can first select a suitable client, and then create a data stream that meets the requirements on the client as the background data stream. Specifically, you can first select a target client whose network stability to the server of the distributed system meets the stability condition. Since the stability of the network from the target client to the server meets the stability condition, the data from the target client will not be abnormal due to network fluctuations. Then, you can create a background data stream on the target client without business logic and whose stability meets the stability condition. Since the network from the target client that initiates the background data stream to the server is stable, the background data stream is stable and has no business logic, so the background data stream will not cause the message queue middleware to fail. In this way, it is guaranteed that the background data stream is not affected, and the accuracy of the results of fault detection combined with the background data stream and the service data stream can be guaranteed.

[0008] In some possible implementations, fault judgment can be performed based on the background data stream and the service data stream, respectively. Specifically, it can be judged whether the message queue middleware has a fault based on the monitoring indicators of the background data stream, and it can be judged whether the message queue middleware has a fault based on the monitoring indicators of the service data stream. If the monitoring indicators of the background data stream and the monitoring indicators of at least one of the multiple service data streams both indicate that the message queue middleware has a fault, it can be determined that the message queue middleware has a fault.

[0009] In some possible implementations, it is also possible to determine whether there are faults in the client and the network link based on the background data stream. Specifically, if the monitoring indicator of the first service data stream among multiple service data streams indicates that there is a fault in the message queue middleware, and the background data stream corresponding to the same message queue as the first service data stream indicates that there is no fault in the message queue middleware, it can be considered that the first service data stream has a fault before being transmitted to the server. Accordingly, it can be determined that there is a fault in the client corresponding to the first service data stream or in the network link from the client to the server.

[0010] In some possible implementations, in addition to fault monitoring, the cause of the message queue middleware failure can also be analyzed based on the background data stream and the service data stream. Specifically, after determining that the message queue middleware has a fault, the pre-trained fault cause classification model can be called to combine the monitoring indicators of the background data stream and the identification of the fault cause of the monitoring indicators of multiple service data streams to obtain the fault cause of the message queue middleware. In this way, when analyzing the cause of the fault, the background data stream is also introduced for analysis, which can eliminate the interference of other modules in the distributed system and obtain the accurate cause of the fault.

[0011] In some possible implementations, multiple models can also be introduced to analyze the cause of the fault. Specifically, a fault cause classification model based on machine learning and a rule classification model based on preset rules can be used. The fault cause classification model can obtain the first fault cause identification result of the message queue middleware based on the monitoring indicators of the background data flow and the monitoring indicators of multiple service data flows. The rule classification model can obtain the first fault cause identification result of the message queue middleware based on the monitoring indicators of the background data flow and the monitoring indicators of multiple service data flows. Combining the first fault cause identification result and the second fault cause identification result, the fault cause of the message queue middleware can be determined. In this way, by identifying the cause of the fault through the machine learning model and the rule model, the cause of the message queue middleware failure can be accurately identified, which is helpful for fault recovery.

[0012] In a second aspect, the present application provides a fault monitoring device, the device is used to perform abnormal monitoring on a message queue middleware in a distributed system, the device comprising:

[0013] A creating unit, used for creating a background data stream, wherein the background data stream is not affected by the client;

[0014] A monitoring unit, used to obtain monitoring indicators of the background data flow and monitoring indicators of multiple service data flows;

[0015] The analyzing unit is used to determine whether the message queue middleware has a fault according to the monitoring indicators of the background data flow and the monitoring indicators of the multiple service data flows.

[0016] In some possible implementations, the creation unit is specifically used to determine a target client, and the stability of the network from the target client to the server of the distributed system meets the stability condition; on the target client, a background data stream is created that has no business logic and whose stability meets the stability condition.

[0017] In some possible implementations, the analysis unit is specifically used to determine whether the message queue middleware has a fault based on the monitoring indicators of the background data flow and the monitoring indicators of the service data flow, respectively; in response to the monitoring indicators of the background data flow and at least one of the multiple service data flows both indicating that the message queue middleware has a fault, determine whether the message queue middleware has a fault.

[0018] In some possible implementations, the analysis unit is further used to determine that there is a fault in the client or network corresponding to the first service data flow in response to the monitoring indicator of the background data flow indicating that there is no fault in the queue middleware, and the monitoring indicator of the first service data flow among the multiple service data flows indicating that there is a fault in the queue middleware.

[0019] In some possible implementations, the analysis unit is further used to obtain the cause of the failure of the message queue middleware by combining the monitoring indicators of the background data flow and the monitoring indicators of the multiple service data flows through a pre-trained fault cause classification model.

[0020] In some possible implementations, the analysis unit is specifically used to input the monitoring indicators of the background data flow and the monitoring indicators of the multiple service data flows into the fault cause classification model to obtain a first fault cause identification result of the message queue middleware; input the monitoring indicators of the background data flow and the monitoring indicators of the multiple service data flows into the rule classification model to obtain a second fault cause identification result of the message queue middleware; and combine the first fault cause identification result and the second fault cause identification result to obtain the fault cause of the message queue middleware.

[0021] In a third aspect, the present application provides a computing device, the computing device comprising at least one processor and at least one memory; the at least one memory is used to store instructions, and the at least one processor executes the instructions stored in the at least one memory, so that the computing device executes the method in the above-mentioned first aspect or any possible implementation of the first aspect. It should be noted that the memory can be integrated into the processor or can be independent of the processor. The at least one computing device may also include a bus. The processor is connected to the memory via a bus. The memory may include a readable memory and a random access memory.

[0022] In a fourth aspect, the present application provides a computing device cluster, wherein the computing device includes at least one computing device, and the at least one computing device includes at least one processor and at least one memory; the at least one memory is used to store instructions, and the at least one processor executes the instructions stored in the at least one memory, so that the computing device cluster executes the method in the above-mentioned first aspect or any possible implementation of the first aspect. It should be noted that the memory can be integrated into the processor or can be independent of the processor. The at least one computing device may also include a bus. The processor is connected to the memory via a bus. The memory may include a readable memory and a random access memory.

[0023] In a fifth aspect, the present application provides a computer-readable storage medium, wherein instructions are stored in the computer-readable storage medium, and when the computer-readable storage medium is executed on at least one computing device, the at least one computing device executes the method described in the first aspect or any one of the implementations of the first aspect.

[0024] In a sixth aspect, the present application provides a computer program product comprising instructions, which, when executed on at least one computing device, enables the at least one computing device to execute the method described in the first aspect or any one of the implementations of the first aspect.

[0025] Based on the implementations provided in the above aspects, the present application can also be further combined to provide more implementations. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in the present application. For ordinary technicians in this field, other drawings can also be obtained based on these drawings.

[0027] Figure 1 A schematic diagram of an application scenario provided in an embodiment of the present application;

[0028] Figure 2 A schematic diagram of a flow chart of a fault monitoring method provided in an embodiment of the present application;

[0029] Figure 3 A schematic diagram of a flow chart of a fault monitoring and recovery method provided in an embodiment of the present application;

[0030] Figure 4 A schematic diagram of the structure of a fault monitoring device provided in an embodiment of the present application;

[0031] Figure 5A schematic diagram of a structure of a computing device provided in an embodiment of the present application;

[0032] Figure 6 A schematic diagram of a structure of a computing device cluster provided in an embodiment of the present application;

[0033] Figure 7 A schematic diagram of an implementation method of a computing device cluster provided in an embodiment of the present application. DETAILED DESCRIPTION

[0034] The scheme in the embodiments provided in this application will be described below in conjunction with the drawings in this application.

[0035] The terms "first", "second", etc. in the specification and claims of this application and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the terms used in this way can be interchangeable under appropriate circumstances. This is just a way of distinguishing objects with the same attributes when describing the embodiments of this application.

[0036] First, some nouns involved in this application are introduced.

[0037] Message queue middleware: Message queue middleware is a role used to implement message distribution in distributed systems. Messages produced by producers will be sent to message queue middleware. Message queue middleware can select appropriate consumers as consumers to process messages. Failures or exceptions in message queue middleware will have a significant impact on distributed systems.

[0038] Producer: A producer is a role that generates messages in a distributed system.

[0039] Consumer: A consumer is a role in a distributed system that consumes messages. Consuming messages may include, for example, processing the messages.

[0040] Message: For message queue middleware, messages refer to data from producers. For example, if a distributed system is used to process tasks, messages can be tasks published by producers; for distributed systems with microservice architecture, service callers can be considered producers, called services can be considered consumers, and messages can be service call requests; for big data systems, producers can be systems that generate data, consumers can be systems that clean and analyze data, and messages can be data to be processed.

[0041] Data stream: In some scenarios, the message sent by the producer can be a data stream. For example, if the task needs to process real-time collected data, the message can be a data stream, and the producer can send data to the consumer through the message queue middleware in the form of a data stream. For another example, if the task needs to process large files and the network bandwidth is limited, the producer can send the large file to be processed to the consumer through the message middleware in the form of a data stream.

[0042] Message queue middleware is an important component of distributed systems and is used to decouple producers and consumers. If a message queue middleware fails, it will have a significant impact on the distributed system. For example, for common message accumulation failures, if the producer's message production rate is higher than the consumer's message consumption rate for a long time, messages will accumulate in the message queue middleware, resulting in message accumulation failures. Message accumulation failures will increase the processing time of new messages and may cause messages to be discarded, affecting the performance of the distributed system.

[0043] In order to control the impact within a certain range, a fault monitoring mechanism can be configured for the message queue middleware. Specifically, the fault monitoring mechanism can collect various data generated by the message queue middleware during its operation, and analyze the working conditions of the message queue middleware based on the data, so as to determine whether the message queue middleware has a fault. In this way, the fault of the message queue middleware can be discovered and investigated in time. In addition, by analyzing the specific conditions of the messages in the message queue, the cause of the fault can also be determined.

[0044] However, this method of fault monitoring is not accurate.

[0045] In fact, in addition to the message queue middleware, problems in other parts of the distributed system will also affect the performance of the distributed system. Therefore, fault monitoring of the message queue middleware may be interfered with by other modules in the distributed system. If a fault is detected, it is impossible to determine whether the fault is caused by the message queue middleware. In other words, the traditional fault monitoring method may be affected by the client and network links in the distributed system, and other faults will be attributed to the fault of the message queue middleware, thereby erroneously diagnosing the fault.

[0046] For example, in a distributed system, the producer of a message is often the client, and the consumer of the message is often located on the server. The client and the server are connected through a network. If the network fluctuates, the data stream corresponding to the message may be transmitted to the server normally. In this way, the processing time of the data stream will be extended, and the message queue to which the data stream belongs may be blocked. If the traditional fault monitoring method is used, it may be considered that the message queue middleware has a fault when scheduling the queue based on the long processing time and message queue blocking.

[0047] For another example, if the data to be processed is very long and the bandwidth from the client to the message queue middleware is limited, then a large amount of data may need to be transmitted to the message queue middleware in a short period of time, which may extend the message processing time.

[0048] For another example, if the client freezes or crashes for some reason, the client cannot send the data in the data stream to the message queue middleware in a timely manner, which may also affect the message processing time.

[0049] In the above three cases, the message queue middleware may have queue blocking and long message processing time. However, the failures in these three cases are caused by the client and / or the network between the client and the server, and have nothing to do with the message queue middleware. The message queue middleware did not fail.

[0050] Based on this, the present application provides a fault monitoring method. The method can filter out interference from parts that are not related to the message queue middleware through background data flow. Specifically, before monitoring, a background data flow can be created first. The background data flow will not be interfered by the client. Therefore, if the message queue middleware has a problem when scheduling the background data flow, it can be considered that the problem occurs on the server side of the distributed system, and the message queue middleware can be used for troubleshooting. After creating the background data flow, the background data flow can be monitored to obtain monitoring indicators of the background data flow. In addition, the service data flow of the distributed system can also be monitored to obtain monitoring indicators of multiple service data flows. Based on the monitoring indicators of the background data flow and the monitoring indicators of the service data flow, it can be determined whether the message queue middleware has a fault. In this way, a background data flow that is not affected by the client is created, and the background data flow and the service data flow are considered during fault monitoring, so that the influence of the client and the network link can be excluded, and the message queue middleware can be accurately monitored for faults. In this way, it can avoid identifying the faults of the client and the network link as the faults of the message queue middleware, avoid the faults of other modules in the distributed system from affecting the fault monitoring results, and accurately monitor the message queue middleware for faults.

[0051] Next, various non-limiting specific implementations of the fault monitoring process are described in detail.

[0052] First, an exemplary application scenario is introduced. The fault monitoring method provided in the present application can be applied to a server, which can be a server of a language generation model system.

[0053] See also Figure 1 , Figure 1 A schematic diagram of an application scenario of the fault monitoring method provided in an embodiment of the present application. Figure 1The application scenario shown includes a client 11, a client 12, a client 13 and a server 20. The server 20 includes a message queue middleware 21, a processing cluster 22 and a fault monitoring device 23.

[0054] Among them, users can use clients 11, 12, and 13 to call the services provided by the server 20. Clients 11, 12, and 13 can send messages to the server 20 as producers. The processing cluster 22 may include one or more consumers for consuming the messages produced by clients 11, 12, and 13. The message queue middleware 21 can schedule messages from clients 11, 12, and 13, and allocate appropriate consumers from the processing cluster 22 for consumption. In the above process, there is a data flow A from the client 11 through the message queue middleware 21 to the processing cluster 22, a data flow B from the client 12 through the message queue middleware 21 to the processing cluster 22, and a data flow C from the client 13 through the message queue middleware 21 to the processing cluster 22. Since data flows A, B, and C are all used to provide services to users, data flows A, B, and C can be called service data flows.

[0055] The fault monitoring device 23 includes a creation unit 231, a monitoring unit 232 and an analysis unit 233. The creation unit 231 is used to create a background data stream that is independent of the client. The monitoring unit 232 is used to obtain monitoring indicators of the background data stream and monitoring indicators of multiple service data streams (including data stream A, data stream B and data stream C). The analysis unit is used to combine the monitoring indicators of the background data stream and the monitoring indicators of the multiple service data streams to determine whether the message queue middleware 21 has a fault.

[0056] The background data stream created by the creation unit 231 needs to be unaffected by the client and the network link between the client and the server 20. To this end, the creation unit 231 can create the background data stream on a device that is stable and has a stable network connection with the server 20. Figure 1 The implementation shown also includes a device 30. The device 30 is a stable device, and there is no network fluctuation between the device 30 and the server 20. The creation unit 231 can create a background data stream on the device 30.

[0057] In an actual application scenario, client 11, client 12, and client 13 may be cloud service clients. Server 20 is a server for providing cloud services. Processing cluster 22 in server 20 may be executed by one or more processing devices. Processing cluster 22 may include multiple consumers. Each consumer may be a virtual machine or a container in processing cluster 22, or a processing device in processing cluster 22. Device 30 may be a device independent of server 20. Optionally, device 30 may be a data processing device that is deployed at the same location as processing cluster 22 but does not belong to processing cluster 22. In this way, the stability of the network link between device 30 and server 20 can be ensured, the influence of the network link can be eliminated, and the influence of internal faults of processing cluster 22 on fault monitoring results can be avoided.

[0058] It should be noted that the above application scenarios are only examples, and the fault monitoring method provided in the embodiments of the present application can be applied to any application scenario of fault monitoring of message queue middleware.

[0059] The specific implementation method of the fault monitoring method is introduced in detail below.

[0060] See also Figure 2 , Figure 2 A flow chart of the fault monitoring method provided in this application. The method can be applied to Figure 1 The application scenarios shown may also be applied to other applicable application scenarios.

[0061] Specifically, Figure 2 The fault monitoring method shown may specifically include:

[0062] S201: Create a background data stream.

[0063] In order to accurately monitor the message queue middleware, the data monitoring device first needs to create a background data stream that is independent of the client, wherein the background data stream will not be affected by the client and is used to provide a reference for fault monitoring.

[0064] Specifically, the background data stream can be a smooth and stable data stream. Stability means that the data volume of the background data stream is within a relatively constant range without large fluctuations. Stability means that the characteristics of the data in the background data stream reaching the server will not be affected, and there will be no jitter. In addition, the background data stream has no business attributes, that is, the background business stream does not correspond to one or more specific cloud services.

[0065] Because the background data stream has a smooth and stable nature and has no business attributes, the server can continuously and evenly obtain data without business attributes from the background data stream. Accordingly, the message queue middleware can continuously and stably dispatch data from the background data stream to consumers. In such an application scenario, it can be assumed that the background data stream will not fail in the process of reaching the message queue middleware from the client. In this way, if an abnormality is detected in the processing of the background data stream, it can be ruled out that the abnormality comes from the background data stream itself and the process from the client to the message queue middleware, and the abnormality is considered to come from the message queue middleware.

[0066] When creating a background data stream, you can first select a suitable client as the target client, and then create a data stream that meets the requirements on the target client as the background data stream.

[0067] Specifically, the stability of the network between the target client and the server of the distributed system satisfies the stability condition. The stability condition includes the requirements for the fluctuation of network parameters such as bandwidth and latency. Optionally, the stability condition may include multiple thresholds, indicating that the fluctuation of various parameters of the network between the target client and the server is less than the threshold.

[0068] Optionally, the target client may be pre-configured for the distributed system. For example, in order to monitor the working condition of the distributed system, the service provider may pre-set a stable and reliable client as the target client. For example, a computing device independent of the distributed system may be set at the location where the distributed system is located as a device for running the target client. For another example, a dedicated physical device or virtual device may be divided in the distributed system for running the target client. For another example, a container or a virtual machine may be configured in the fault monitoring device as a target client to create a background data stream.

[0069] After determining the target client, a background data stream can be created on the target client. Specifically, a data stream that does not include business logic and whose stability satisfies the stability condition can be created as the background data stream. Not including business logic means that the data in the background data stream does not correspond to the actual business, and the law of data transmission in the background data stream is also different from the actual business. The stability of the background data stream satisfies the stability condition, which means that the data in the background data stream will be sent from the target client evenly and smoothly. Because the background data stream satisfies the stability condition and the target client satisfies the stability condition, the data from the background data stream will arrive at the message queue middleware on the server evenly and smoothly. In this way, by restricting the client and the data stream when creating the background data stream, the influence of the client and the network link from the client to the server can be filtered out.

[0070] S202: Acquire monitoring indicators of a background data flow and monitoring indicators of a plurality of service data flows.

[0071] During the operation of the message queue middleware, the data stream carried by the message queue middleware can be monitored, and the monitoring indicators of the data stream can be collected so as to determine whether the message queue middleware fails based on the monitoring indicators of the data stream. Among them, the data stream carried by the message queue middleware includes the background data stream created in step S201, and also includes the service data stream initiated by the client for calling the service. In other words, the fault monitoring device can obtain the monitoring indicators of the background data stream and the monitoring indicators of multiple service data streams. Optionally, if the distributed system only carries one data stream, the fault monitoring device can also obtain the monitoring indicators of one service data stream.

[0072] In an embodiment of the present application, monitoring indicators may include indicators such as queue traffic, bandwidth utilization, message sending time, message pull time interval and message pull volume. Among them, queue traffic refers to the traffic of one or more queues in the message queue middleware per unit time. Bandwidth utilization refers to the bandwidth utilization of the server. Message sending time refers to the time taken for a message to be added to the message queue and taken out of the message queue. Message pull time interval refers to the time interval between two times a consumer takes out a message from the message queue. Message pull volume refers to the amount of data extracted by a consumer each time it obtains data from a message queue.

[0073] It is understandable that in actual application scenarios, more or fewer monitoring indicators may be configured according to fault monitoring requirements.

[0074] S203: judging whether the message queue middleware has a fault according to the monitoring indicators of the background data flow and the monitoring indicators of the plurality of service data flows.

[0075] After obtaining the monitoring indicators of the background data flow and the monitoring indicators of multiple service data flows, fault monitoring can be performed based on the monitoring indicators of the background data flow and the monitoring indicators of the service data flow to determine whether the message queue middleware has a fault. If the message queue middleware has a fault, an alarm can be issued and / or the fault can be automatically repaired.

[0076] Optionally, fault monitoring can be performed through a fault monitoring model. Specifically, the monitoring indicators of the background data flow and the monitoring indicators of the service data flow can be respectively extracted, and the extracted features are input into the fault monitoring model, and the fault monitoring model determines whether the message queue middleware fails based on the features. Optionally, the fault monitoring model can be a classification model based on machine learning technology. For example, the fault monitoring model can be an extreme gradient boosting model (Extreme Gradient Boosting, XGB Boost) model.

[0077] Optionally, in order to improve accuracy, the fault can be analyzed in combination with historical monitoring indicators. Specifically, the collected monitoring indicators and historical data of the monitoring indicators can be combined to calculate the high-order parameters such as the week-on-week, day-on-day, volatility ratio and moving average of each monitoring indicator. Then, features can be extracted from the monitoring indicators and high-order parameters based on the isolation forest algorithm, the 3σ principle, etc., and the extracted features can be input into the fault monitoring model, and the features can be analyzed through the fault monitoring model to determine whether there is a fault in the message queue middleware.

[0078] In some possible implementations, when judging whether a queue middleware fails, a judgment can be made based on the monitoring indicators of the background data flow and the monitoring indicators of the service data flow respectively, and the judgment result of the background data flow and the judgment result of the service data flow can be combined to determine whether the message queue middleware fails.

[0079] Specifically, assuming that there are N service data flows, judgments can be made based on the monitoring indicators of the N service data flows. If the monitoring indicators of one (or more) of the N service data flows indicate that the message queue middleware has a fault, further judgment can be made in combination with the monitoring indicators of the background data flow. If the monitoring indicators of the background data flow also indicate that the message queue middleware has a fault, it can be determined that the message queue middleware has a fault.

[0080] Message queue middleware is often used to manage multiple message queues. In actual application scenarios, it is possible that one of the message queues has a fault. That is, if the monitoring indicator of a service data stream among N service data streams is used to indicate that a message queue has a fault, then the monitoring indicator of the background data stream can be combined to determine whether the message queue has a fault. Optionally, the message queue middleware can be controlled to schedule the background data stream to the message queue that may have a fault, and collect the monitoring indicators of the background data stream after scheduling, so as to determine whether the message queue has a fault based on the newly collected monitoring indicators of the background data stream.

[0081] If both the background data stream and the service data stream indicate that a message queue is faulty, it can be considered that the message queue is faulty. Since the background data stream is introduced and the background data stream is not interfered by the client and the network from the client to the server, the result obtained by combining the background data stream and the service data stream is more accurate and will not be affected by other factors.

[0082] If the monitoring indicator of the first service data flow among the multiple service data flows indicates that the message queue middleware has a fault, and the background data flow corresponding to the same message queue as the first service data flow indicates that the message queue middleware has no fault, it can be considered that the first service data flow has a fault before being transmitted to the server. Accordingly, it can be determined that the client corresponding to the first service data flow or the network link from the client to the server has a fault.

[0083] The present application provides a fault monitoring method. The method can filter out interference from parts that are not related to the message queue middleware through background data flow. Specifically, before monitoring, a background data flow can be created first. The background data flow will not be interfered by the client. Therefore, if there is a problem with the message queue middleware when scheduling the background data flow, it can be considered that the problem occurs on the server side of the distributed system, and the message queue middleware can be used for troubleshooting. After creating the background data flow, the background data flow can be monitored to obtain monitoring indicators of the background data flow. In addition, the service data flow of the distributed system can also be monitored to obtain monitoring indicators of multiple service data flows. Based on the monitoring indicators of the background data flow and the monitoring indicators of the service data flow, it can be determined whether the message queue middleware has a fault. In this way, a background data flow that is not affected by the client is created, and the background data flow and the service data flow are considered during fault monitoring, so that the influence of the client and the network link can be eliminated, and the message queue middleware can be accurately monitored for faults. In this way, it is possible to avoid identifying the faults of the client and the network link as the faults of the message queue middleware, avoid the faults of other modules in the distributed system from affecting the fault monitoring results, and accurately monitor the faults of the message queue middleware.

[0084] In actual application scenarios, it is necessary not only to monitor the faults of the message queue middleware, but also to analyze the causes of the faults so as to conduct timely troubleshooting. To this end, in some traditional distributed systems, not only a fault monitoring mechanism is configured, but also a fault root cause analysis mechanism is configured. Through the fault root cause analysis mechanism, various indicators in the service data flow can be combined for analysis to determine the cause of the message queue middleware failure, so as to perform fault recovery based on the cause and then issue an alarm.

[0085] However, similar to the situation described above, the various indicators in the service data flow will not only be affected by various parts of the server (such as the message queue middleware and consumers), but also by the client and the network link from the client to the server. This will lead to inaccurate analysis of the cause of the fault, and inaccurate fault recovery and alarm.

[0086] Therefore, based on the above fault monitoring method, the embodiment of the present application also provides a corresponding fault cause analysis and fault recovery mechanism, which will be described in detail below in conjunction with the accompanying drawings of the specification.

[0087] See also Figure 3 , Figure 3 A flow chart of a fault monitoring and recovery method provided in an embodiment of the present application. The method can be applied to Figure 1 The application scenarios shown may also be applied to other applicable application scenarios.

[0088] Specifically, Figure 3 The fault monitoring and recovery method shown may specifically include:

[0089] S301: Create a background data stream.

[0090] In order to eliminate the interference of client and other modules and accurately analyze the cause of the message queue middleware failure, you can first create a background data stream. For an introduction to creating a background data stream, see Figure 2 , I will not go into details here.

[0091] S302: Acquire monitoring indicators of a background data flow and monitoring indicators of a plurality of service data flows.

[0092] In order to determine whether the message queue middleware has a fault, you can obtain the monitoring indicators of the background data flow and the monitoring indicators of multiple service data flows. For an introduction to this part, see Figure 2 , I will not go into details here.

[0093] It should be noted that compared with Figure 2 The implementation shown in Figure 3 In the implementation shown, it is necessary not only to determine whether the message queue middleware has a fault based on the monitoring indicators of the data flow, but also to perform a root cause analysis of the fault based on the monitoring indicators of the data flow. Therefore, the monitoring indicators obtained in step S302 can be the same as the monitoring indicators obtained in step S202, or they can be different.

[0094] S303: Determine that a fault exists in the message queue middleware according to the monitoring indicators of the background data flow and the monitoring indicators of the plurality of service data flows.

[0095] After obtaining the monitoring indicators of the background data flow and the monitoring indicators of multiple service data flows, it can be determined whether the message queue middleware has a fault based on the monitoring indicators. Figure 3 In the implementation shown, the situation where the message queue middleware has a fault is introduced. If it is determined that the message queue middleware has a fault, the following steps S304 and S305 may be continued to be executed to analyze the cause of the fault.

[0096] For an introduction on how to determine whether the message queue middleware has a fault based on monitoring indicators, see Figure 2 , I will not go into details here.

[0097] S304: Obtain the fault cause of the message queue middleware through the pre-trained fault cause classification model, combined with the monitoring indicators of the background data flow and the monitoring indicators of the plurality of service data flows.

[0098] After determining that the message queue middleware fails, the cause of the failure of the message queue middleware can be analyzed through a failure cause classification model. The failure cause classification model is a pre-trained model, and its input includes monitoring indicators of background data flows and monitoring indicators of multiple service data flows. The failure cause classification model can analyze the cause of the failure of the message queue middleware based on the monitoring indicators to obtain the failure cause of the message queue middleware.

[0099] When training a fault cause classification model, a training data set can be obtained first. If a supervised learning method is used to train the fault cause classification model, the training data set includes monitoring indicators with pre-labeled fault causes. In addition, the monitoring indicators in the training data set also include monitoring indicators for background data flows and monitoring indicators for service data flows. For example, historical data of the message queue middleware can be collected, and the monitoring indicators of the background data flow and the monitoring indicators of the service data flow when the message queue middleware fails can be manually or automatically labeled to obtain a training data set. By training with the training data set, the fault cause classification model can learn the mapping rules between monitoring indicators and fault causes. When analyzing the fault cause of the message queue middleware, the fault cause classification model can combine the monitoring indicators of the background data flow and the monitoring indicators of multiple service data flows to analyze and obtain the fault cause of the message queue middleware.

[0100] When analyzing the cause of the failure of the message queue middleware, not only the monitoring indicators of the service data flow are considered, but also the monitoring indicators of the background data flow are introduced. In this way, through the background data flow, the interference of other parts of the distributed system can be eliminated, and the cause of the failure of the message queue middleware can be analyzed more accurately.

[0101] Considering that there may be deviations in the analysis using the fault cause classification model, in some implementations, a rule model may be introduced to assist in analyzing the fault cause.

[0102] Specifically, in addition to the fault cause classification model, a rule classification model can also be configured. The rule classification model is used to classify according to the pre-configured classification rules in combination with the monitoring indicators. Among them, the pre-configured classification rules can be classification rules configured based on expert experience, indicating the association between the monitoring indicators and the classification results. The classification results can include the fault cause, and can also include the intermediate results in the process of determining the fault cause.

[0103] For example, the classification rules may include flow control classification rules. If the flow of the queue exceeds the threshold, the cause of the failure can be determined according to the flow control classification rules to be that the consumer is under flow control. For another example, the classification rules may include pull request classification rules and processing time classification rules. According to the pull request classification rules, it can be determined whether the consumer's message pull volume has reached the maximum. If so, it can be further combined with the processing time classification rules to determine whether the time consumption for processing the message is lower than the threshold. If it is lower, it can be considered that the current maximum pull volume setting is inappropriate, resulting in the consumer pulling too few messages from the message queue each time. The cause of the failure is that the amount of messages pulled is too small.

[0104] Accordingly, after obtaining the monitoring indicators of the background data flow and the monitoring indicators of the multiple service data flows, the monitoring indicators of the background data flow and the monitoring indicators of the multiple service data flows can be respectively input into the fault cause classification model and the rule classification model. The fault cause classification model can output the fault cause of the message queue middleware by combining the monitoring indicators of the background data flow and the monitoring indicators of the multiple service data flows through a machine learning method. The rule classification model can classify according to the pre-configured classification rules, combining the monitoring indicators of the background data flow and the monitoring indicators of the multiple service data flows, and output the fault cause of the message queue middleware. Among them, the fault cause output by the fault cause classification model can be called the first fault cause identification result. The fault cause output by the rule classification model can be called the second fault cause identification result. Combining the first fault cause identification result and the second fault cause identification result, the fault cause of the message queue middleware can be obtained. In this way, by combining the classification model based on machine learning and the rule model for identification, the cause of the message queue middleware failure can be accurately analyzed and the accurate fault cause can be obtained.

[0105] Optionally, when the first fault cause identification result and the second fault cause identification result are the same, the first fault cause identification result may be determined as the fault cause.

[0106] Alternatively, optionally, the first fault cause identification result and the second fault cause identification result may include the probability corresponding to the fault cause, so the fault cause of the message queue middleware can be determined in combination with the pre-configured weight. Among them, the probability represents the probability that the fault of the message queue middleware is caused by a certain fault cause. The pre-configured weight may include a weight pre-configured for the first fault cause identification result and a weight pre-configured for the second fault cause identification result. The probability corresponding to each fault cause can be obtained by multiplying the first fault cause identification result and the second fault cause identification result by the corresponding weights respectively, and adding the two products. The fault cause with the highest probability can be used as the fault cause of the message queue middleware.

[0107] S305: Perform fault recovery and / or generate an alarm according to the fault cause of the message queue middleware.

[0108] After the cause of the failure of the message queue middleware is determined, failure recovery and / or an alarm may be performed based on the cause of the failure of the message queue middleware.

[0109] Optionally, an automatic recovery mechanism for faults with different fault causes can be pre-configured. After the fault cause is determined, the corresponding automatic recovery mechanism can be used to automatically recover the fault. Optionally, an alarm mechanism can also be configured. Based on the alarm mechanism, after a fault is detected, an alarm message can be generated and sent. The alarm message may include the time when the fault occurred and the cause of the fault. Based on the alarm message, the management personnel can quickly troubleshoot and respond to the fault.

[0110] Optionally, a judgment may be made in combination with the actual cause of the fault. Specifically, for faults caused by some fault causes, automatic recovery of the fault may be performed. For faults caused by other fault causes, such as faults caused by insufficient capacity on the server side, the recovery method requires capacity expansion on the server side, and the capacity expansion on the server side requires user confirmation, so automatic recovery of the fault cannot be performed. Accordingly, after determining the cause of the fault, a judgment may be made based on the cause of the fault. If automatic recovery is possible, an automatic recovery mechanism may be used to recover from the fault. If automatic recovery is not possible, an alarm may be used to remind the user to recover from the fault.

[0111] The present application also provides a fault monitoring device, wherein the fault monitoring device can be applied to Figure 1 The server 20 in the implementation is shown to implement Figure 2 The function of the fault monitoring device in the implementation shown. Specifically, Figure 4 As shown, the fault monitoring device 400 includes:

[0112] A creating unit 410, configured to create a background data stream, wherein the background data stream is not affected by the client;

[0113] A monitoring unit 420, configured to obtain monitoring indicators of the background data flow and monitoring indicators of multiple service data flows;

[0114] The analyzing unit 430 is used to determine whether the message queue middleware has a fault according to the monitoring indicators of the background data flow and the monitoring indicators of the multiple service data flows.

[0115] Among them, the creation unit 410, the monitoring unit 420 and the analysis unit 430 can all be implemented by software, or can be implemented by hardware. Exemplarily, the implementation of the analysis unit 430 is introduced below by taking the analysis unit 430 as an example. Similarly, the implementation of the creation unit 410 and the monitoring unit 420 can be the implementation of the analysis unit 430.

[0116] As an example of a software functional unit, the analysis unit 430 may include code running on a computing instance. Among them, the computing instance may include at least one of a physical host (computing device), a virtual machine, and a container. Further, the above-mentioned computing instance may be one or more. For example, the analysis unit 430 may include code running on multiple hosts / virtual machines / containers. It should be noted that the multiple hosts / virtual machines / containers used to run the code may be distributed in the same region (region) or in different regions. Furthermore, the multiple hosts / virtual machines / containers used to run the code may be distributed in the same availability zone (AZ) or in different AZs, each AZ including a data center or multiple data centers with close geographical locations. Among them, usually a region may include multiple AZs.

[0117] Similarly, multiple hosts / virtual machines / containers used to run the code can be distributed in the same virtual private cloud (VPC) or in multiple VPCs. Usually, a VPC is set up in a region. For cross-region communication between two VPCs in the same region and between VPCs in different regions, a communication gateway needs to be set up in each VPC to achieve interconnection between VPCs through the communication gateway.

[0118] As an example of a hardware functional unit, the analysis unit 430 may include at least one computing device, such as a server, etc. Alternatively, the analysis unit 430 may also be a device implemented using an application-specific integrated circuit (ASIC) or a programmable logic device (PLD). The PLD may be a complex programmable logical device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL) or any combination thereof.

[0119] The multiple computing devices included in the analysis unit 430 can be distributed in the same region or in different regions. The multiple computing devices included in the analysis unit 430 can be distributed in the same AZ or in different AZs. Similarly, the multiple computing devices included in the monitoring unit 420 can be distributed in the same VPC or in multiple VPCs. The multiple computing devices can be any combination of computing devices such as servers, ASICs, PLDs, CPLDs, FPGAs, and GALs.

[0120] It should be noted that, in other embodiments, the creation unit 410 can be used to execute any step in the fault monitoring method, the monitoring unit 420 can be used to execute any step in the fault monitoring method, and the analysis unit 430 can be used to execute any step in the fault monitoring method. The steps that the creation unit 410, the monitoring unit 420, and the analysis unit 430 are responsible for implementing can be specified as needed, and the creation unit 410, the monitoring unit 420, and the analysis unit 430 respectively implement different steps in the fault monitoring method to achieve the full functions of the fault monitoring device.

[0121] The present application also provides a computing device 100. Figure 5 As shown, the computing device 100 includes: a bus 102, a processor 104, a memory 106, and a communication interface 108. The processor 104, the memory 106, and the communication interface 108 communicate with each other through the bus 102. The computing device 400 can be a server or a terminal device. It should be understood that the present application does not limit the number of processors and memories in the computing device 100.

[0122] The bus 102 may be a peripheral component interconnect (PCI) bus or an extended industry standard architecture (EISA) bus. The bus may be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 5 The bus 102 is represented by only one line, but it does not mean that there is only one bus or one type of bus. The bus 102 may include a path for transmitting information between various components of the computing device 100 (eg, the memory 106, the processor 104, and the communication interface 108).

[0123] The processor 104 may include any one or more of a central processing unit (CPU), a graphics processing unit (GPU), a microprocessor (MP), or a digital signal processor (DSP).

[0124] The memory 106 may include a volatile memory, such as a random access memory (RAM). The processor 104 may also include a non-volatile memory, such as a read-only memory (ROM), a flash memory, a hard disk drive (HDD), or a solid state drive (SSD).

[0125] The memory 106 stores executable program codes, and the processor 104 executes the executable program codes to respectively implement the functions of the aforementioned creation unit 410, monitoring unit 420 and analysis unit 430, thereby implementing the fault monitoring method. That is, the memory 106 stores instructions for executing the storage method.

[0126] Alternatively, the memory 106 stores executable codes, and the processor 104 executes the executable codes to implement the functions of the aforementioned fault monitoring device, thereby implementing the fault monitoring method. That is, the memory 106 stores instructions for executing the fault monitoring method.

[0127] The communication interface 108 uses a transceiver module such as, but not limited to, a network interface card or a transceiver to implement communication between the computing device 100 and other devices or a communication network.

[0128] The embodiment of the present application also provides a computing device cluster. The computing device cluster includes at least one computing device. The computing device can be a server, such as a central server, an edge server, or a local server in a local data center. In some embodiments, the computing device can also be a terminal device such as a desktop computer, a laptop computer, or a smart phone.

[0129] like Figure 6 As shown, the computing device cluster includes at least one computing device 100. The memory 106 in one or more computing devices 100 in the computing device cluster may store the same instructions for executing the fault monitoring method.

[0130] In some possible implementations, the memory 106 of one or more computing devices 100 in the computing device cluster may also store partial instructions for executing the fault monitoring method. In other words, the combination of one or more computing devices 100 may jointly execute instructions for executing the fault monitoring method.

[0131] It should be noted that the memory 106 in different computing devices 100 in the computing device cluster can store different instructions, which are respectively used to execute part of the functions of the fault monitoring device. That is, the instructions stored in the memory 106 in different computing devices 100 can implement the functions of one or more modules in the creation unit 410, the monitoring unit 420 and the analysis unit 430.

[0132] In some possible implementations, one or more computing devices in the computing device cluster may be connected via a network, which may be a wide area network or a local area network. Figure 7 A possible implementation is shown. Figure 7 As shown, two computing devices 100A and 100B are connected via a network. Specifically, the network is connected via a communication interface in each computing device. In this type of possible implementation, the memory 106 in the computing device 100A stores instructions for executing the functions of the creation unit 410. At the same time, the memory 106 in the computing device 100B stores instructions for executing the functions of the monitoring unit 420 and the analysis unit 430.

[0133] It should be understood that Figure 7 The functions of the computing device 100A shown in FIG. 1 may also be completed by multiple computing devices 100. Similarly, the functions of the computing device 100B may also be completed by multiple computing devices 100.

[0134] The present application embodiment also provides another computing device cluster. The connection relationship between the computing devices in the computing device cluster can be similar to that of Figure 6 and Figure 7 The connection mode of the computing device cluster is different in that the memory 106 in one or more computing devices 100A in the computing device cluster may store the same instructions for executing the fault monitoring method.

[0135] In some possible implementations, the memory of one or more computing devices 100B in the computing device cluster may also store partial instructions for executing the fault monitoring method. In other words, a combination of one or more computing devices may jointly execute instructions for executing the fault monitoring method.

[0136] It should be noted that the memory 106 in different computing devices 100A in the computing device cluster may store different instructions for executing part of the functions of the fault monitoring device. That is, the instructions stored in the memory 106 in different computing devices 100A may implement the functions of one or more devices in the cloud service system.

[0137] The embodiment of the present application also provides a computer program product including instructions. The computer program product may be software or a program product including instructions that can be run on a computing device or stored in any available medium. When the computer program product is run on at least one computing device, the at least one computing device executes the fault monitoring method.

[0138] The embodiment of the present application also provides a computer-readable storage medium. The computer-readable storage medium can be any available medium that can be stored by a computing device or a data storage device such as a data center that contains one or more available media. The available medium can be a magnetic medium (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state hard disk). The computer-readable storage medium includes instructions that instruct a computing device to execute a fault monitoring method, or instructs a computing device to execute a fault monitoring method.

[0139] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the protection scope of the technical solutions of the embodiments of the present invention.

Claims

1. A fault monitoring method, characterized in that: The method is used to perform abnormal monitoring on a message queue middleware in a distributed system, and the method comprises: Creating a background data stream, where the background data stream is not affected by the client; Obtaining monitoring indicators of the background data flow and monitoring indicators of multiple service data flows; It is determined whether the message queue middleware has a fault according to the monitoring index of the background data flow and the monitoring indexes of the plurality of service data flows.

2. The method according to claim 1, characterized in that: The creating background data stream comprises: Determine a target client, wherein the stability of a network from the target client to a server of the distributed system satisfies a stability condition; On the target client, a background data stream is created which has no business logic and whose stability satisfies a stability condition.

3. The method according to claim 1 or 2, characterized in that: The determining, based on the monitoring indicators of the background data flow and the monitoring indicators of the plurality of service data flows, whether there is a fault in the message queue middleware comprises: Determining whether the message queue middleware has a fault according to the monitoring indicators of the background data flow and the monitoring indicators of the service data flow respectively; In response to the monitoring indicator of the background data flow and the monitoring indicator of at least one of the multiple service data flows both indicating that the message queue middleware has a fault, it is determined whether the message queue middleware has a fault.

4. The method according to claim 3, characterized in that The method further comprises: In response to the monitoring indicator of the background data flow indicating that the queue middleware has no fault, and the monitoring indicator of the first service data flow among the multiple service data flows indicating that the queue middleware has a fault, it is determined that the client or network corresponding to the first service data flow has a fault.

5. The method according to any one of 1 to 4, characterized in that After determining that the message queue middleware has a fault, the method further includes: The fault cause of the message queue middleware is obtained by using a pre-trained fault cause classification model and combining the monitoring indicators of the background data flow and the monitoring indicators of the multiple service data flows.

6. The method according to claim 5, characterized in that The fault cause classification model pre-trained, combined with the monitoring indicators of the background data flow and the monitoring indicators of the plurality of service data flows, obtains the fault cause of the message queue middleware, including: Inputting the monitoring index of the background data flow and the monitoring index of the plurality of service data flows into the fault cause classification model to obtain a first fault cause identification result of the message queue middleware; Inputting the monitoring indicators of the background data flow and the monitoring indicators of the plurality of service data flows into a rule classification model to obtain a second fault cause identification result of the message queue middleware; The first fault cause identification result and the second fault cause identification result are combined to obtain the fault cause of the message queue middleware.

7. A fault monitoring device, characterized in that: The device is used to perform abnormal monitoring on a message queue middleware in a distributed system, and the device includes: A creating unit, used for creating a background data stream, wherein the background data stream is not affected by the client; A monitoring unit, used to obtain monitoring indicators of the background data flow and monitoring indicators of multiple service data flows; The analyzing unit is used to determine whether the message queue middleware has a fault according to the monitoring indicators of the background data flow and the monitoring indicators of the multiple service data flows.

8. The device according to claim 7, characterized in that The creation unit is specifically used to determine a target client, and the stability of the network from the target client to the server of the distributed system meets the stability condition; on the target client, create a background data flow without business logic and whose stability meets the stability condition.

9. The device according to claim 7 or 8, characterized in that The analysis unit is specifically used to determine whether the message queue middleware has a fault based on the monitoring indicators of the background data flow and the monitoring indicators of the service data flow respectively; in response to the monitoring indicators of the background data flow and the monitoring indicators of at least one of the multiple service data flows both indicating that the message queue middleware has a fault, determine whether the message queue middleware has a fault.

10. The device according to claim 9, characterized in that The analysis unit is further used to determine that there is a fault in the client or network corresponding to the first service data flow in response to the monitoring indicator of the background data flow indicating that there is no fault in the queue middleware, and the monitoring indicator of the first service data flow among the multiple service data flows indicating that there is a fault in the queue middleware.

11. The device according to any one of claims 7 to 10, characterized in that The analysis unit is further used to obtain the fault cause of the message queue middleware by combining the monitoring indicators of the background data flow and the monitoring indicators of the multiple service data flows through a pre-trained fault cause classification model.

12. The device according to claim 11, characterized in that The analysis unit is specifically used to input the monitoring indicators of the background data flow and the monitoring indicators of the multiple service data flows into the fault cause classification model to obtain a first fault cause identification result of the message queue middleware; input the monitoring indicators of the background data flow and the monitoring indicators of the multiple service data flows into the rule classification model to obtain a second fault cause identification result of the message queue middleware; and obtain the fault cause of the message queue middleware by combining the first fault cause identification result and the second fault cause identification result.

13. A computing device, characterized in that: The computing device includes a processor and a memory; The processor is configured to execute instructions stored in the memory so that the computing device performs the operation steps of the method according to any one of claims 1 to 6.

14. A computing device cluster, characterized in that: The computing device cluster includes at least one computing device, each computing device including a processor and a memory: The memory is used to store instructions; The processor is configured to cause the computing device cluster to execute the operation steps of any one of claims 1 to 6 according to the instructions.

15. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores instructions, which, when executed on a computing device, enable the computing device to perform the operating steps of the method according to any one of claims 1 to 6.

16. A computer program product comprising instructions, which, when executed on a computing device, causes the computing device to perform the operating steps of the method according to any one of claims 1 to 6.