Fault processing method and related device
By using the first operation and maintenance model in the cloud management platform to generate summary alarm information and interactively determine the cause of the fault, the problem of low fault handling efficiency in intelligent operation and maintenance technology is solved, and efficient fault handling is achieved.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-06-27
- Publication Date
- 2026-04-02
AI Technical Summary
Current intelligent operation and maintenance technologies suffer from low fault handling efficiency, require manual intervention, and are time-consuming when faced with complex service dependencies and alarm storms.
By applying the first operation and maintenance model in the cloud management platform, summative alarm information is generated based on historical fault database and service deployment information. The cause of the fault is determined through natural language interaction, thereby reducing the number of alarm information and improving fault handling efficiency.
It effectively suppresses alarm storms, improves the efficiency of fault cause identification, reduces manual intervention, and enhances fault handling efficiency.
Smart Images

Figure CN2025104544_02042026_PF_FP_ABST
Abstract
Description
A fault processing method and related device
[0001] The present application claims priority to the Chinese patent application No. 202411375287.6, filed on September 29, 2024, with the State Intellectual Property Office of China, and the Chinese patent application No. 202411375287.6 has the title of “A fault processing method and related device”, the content of which is incorporated herein by reference in its entirety. TECHNICAL FIELD
[0002] The present application relates to the field of cloud computing, and more particularly, to a fault processing method, a computing device, a computing device cluster, a computer program product and a computer readable storage medium. BACKGROUND
[0003] With the continuous development of AI technology, intelligent operation and maintenance technology emerges as the times require. The intelligent operation and maintenance technology centrally manages operation and maintenance data through an AI model, thereby facilitating timely alarm when a service fails, and helping to determine the cause of the service failure. Due to the low maturity of the current intelligent operation and maintenance technology, in the case of a large number of services and complex dependency relationships between services, the current intelligent operation and maintenance technology still requires operation and maintenance personnel to have high professional ability. For example, when a service fails, the current intelligent operation and maintenance technology still needs the operation and maintenance personnel to manually assist in troubleshooting, thereby realizing fault diagnosis and root cause positioning, resulting in a long time required for processing the fault and low processing efficiency. At the same time, for the alarm storm in the operation and maintenance scenario, i.e., the case of issuing a large number of alarms in a short time, due to the flooding of alarm information, the operation and maintenance personnel need to spend a lot of time to determine valuable alarm information from the large number of alarm information, thereby resulting in low efficiency of fault processing.
[0004] Therefore, how to improve the efficiency of fault processing becomes a problem to be solved. SUMMARY
[0005] The present application provides a fault processing method, a computing device, a computing device cluster, a computer program product and a computer readable storage medium, which can effectively suppress the alarm storm and improve the efficiency of determining the cause of the fault, thereby improving the efficiency of fault processing.
[0006] In a first aspect, a fault processing method is provided. The method is applied to a cloud management platform configured to manage an infrastructure providing cloud services, the infrastructure comprising at least one computing node configured to run M services, M being a positive integer. The method comprises: in a case where at least one service of the M services fails and a plurality of alarm information exists, providing first alarm information for a tenant, the first alarm information being configured to indicate that the at least one service fails, the first alarm information being generated by summarizing alarm information triggered by a same root fault from the plurality of alarm information according to a first operation and maintenance model, the first operation and maintenance model being obtained by training according to a historical fault database and deployment information of the M services, the plurality of alarm information being generated according to a set of running information of the M services, the set of running information comprising at least one running information; receiving first input information of the tenant, the first input information being configured to request to obtain a cause of the root fault corresponding to the first alarm information; and in response to the first input information, providing first output information for the tenant, the first output information being configured to indicate the cause of the root fault corresponding to the first alarm information, the first output information being generated according to the first operation and maintenance model.
[0007] In the embodiments of the present application, in a case where at least one service of the M services fails and a plurality of alarm information exists, the cloud management platform can process at least one piece of alarm information triggered by the same root fault from the plurality of alarm information according to the first operation and maintenance model, and generate one piece of valuable alarm information corresponding to each root fault, thereby reducing the number of alarm information provided for the tenant, and effectively suppressing the alarm storm. Meanwhile, the cloud management platform can interact with the tenant in natural language through the first operation and maintenance model, thereby providing the tenant with the cause of the failure according to the input information of the tenant, improving the efficiency of determining the cause of the failure, and thereby facilitating the tenant to solve the failure according to the cause of the failure, and improving the efficiency of fault processing.
[0008] In combination with the first aspect, in some implementations, the historical fault database comprises at least one of: at least one historical operation and maintenance case, a fault mode library, or a diagnosis tree. The historical operation and maintenance case comprises at least one of: historical fault alarm information, a location of a fault corresponding to the historical fault alarm information, a cause of the fault corresponding to the historical fault alarm information, a processing measure of the fault corresponding to the historical fault alarm information. The fault mode library comprises at least one of: identification information of the fault, a location of the fault, a cause of the fault, a processing measure of the fault, a level of the fault, and an association relationship between faults. The diagnosis tree comprises a first diagnosis tree and / or a second diagnosis tree, the first diagnosis tree being configured to determine a root cause of a fault according to fault alarm information, and the second diagnosis tree being configured to determine a processing measure of the fault according to the root cause of the fault.
[0009] In some implementations of the first aspect, the deployment information of the M services includes at least one of the following: information of a computing node cluster in which computing nodes running the M services are located, an association relationship between the M services, a service topology of the M services.
[0010] In the embodiments of the present application, the first operation and maintenance model is obtained by training the historical fault database and the deployment information of the M services as training data, so that the first operation and maintenance model can classify, interpret or summarize the plurality of alarm information, thereby generating the first alarm information.
[0011] In some implementations of the first aspect, the running information set includes at least one of the following: a data index information set, a call link information set, and a log information set. The data index information set includes at least one data index in the running process of the M services, the call link information set includes call link information in the processing process of at least one request in the running process of the M services, and the log information set includes at least one log information in the running process of the M services.
[0012] In some implementations of the first aspect, the first abnormal information set is determined according to the running information set of the M services and the second operation and maintenance model, the second operation and maintenance model is obtained by training the historical running information set of the M services, the first abnormal information set includes at least one abnormal information, and each abnormal information in the at least one abnormal information is running information in the running information set that meets an abnormal characteristic; and the plurality of alarm information is determined according to the first abnormal information set.
[0013] In the embodiments of the present application, the real-time running information set of the M services is analyzed by the second operation and maintenance model, thereby determining abnormal running information in the running information set and generating at least one alarm information, and further facilitating the provision of important and valuable alarm information for tenants and effectively suppressing the alarm storm.
[0014] In some implementations of the first aspect, the second operation and maintenance model includes at least one of the following: a first sub-model, a second sub-model or a third sub-model. The method further includes: training a first initial model according to a historical data index information set of the M services to obtain the first sub-model. And / or, training a second initial model according to a historical call link information set of the M services to obtain the second sub-model. And / or, training a third initial model according to a historical log information set of the M services to obtain the third sub-model.
[0015] The first sub-model is configured to determine a first data indicator information set from the data indicator information set of the M services, the first data indicator information set including at least one abnormal first data indicator, and the first data indicator information set belonging to the first abnormal information set. The second sub-model is configured to determine a first call link information set from the call link information set of the M services, the first call link information set including at least one abnormal call link information, and the first call link information set belonging to the first abnormal information set. The third sub-model is configured to determine a first log information set from the log information set of the M services, the first log information set including at least one abnormal log information, and the first log information set belonging to the first abnormal information set. The first initial model, the second initial model, and the third initial model are foundation models (FM).
[0016] In the embodiments of the application, the historical data indicator information set, the call link information set, and the log information set of the M services are respectively taken as training data to train the foundation model, so that at least one of the first sub-model for extracting abnormal data indicator information, the second sub-model for extracting abnormal call link information, or the third sub-model for extracting abnormal log information is obtained, and then a large amount of running data can be analyzed to generate alarm information.
[0017] In combination with the first aspect, in some implementations, the initial large model is a large language model (LLM), and the first operation and maintenance model is obtained by training the initial large model according to the historical fault database and the deployment information of the M services.
[0018] In the embodiments of the application, the LLM is fine-tuned and enhanced in the vertical field knowledge by adding the historical fault database and the deployment information of the M services to the LLM training corpus, so that the first operation and maintenance model for generating the first alarm information and the first output information is obtained, and then the alarm storm can be effectively suppressed, and the efficiency of fault processing can be improved.
[0019] In combination with the first aspect, in some implementations, the first operation and maintenance model performs retrieval in the historical fault database and / or the running information set of the M services according to the first input information to determine a first response information set, the first response information set including at least one response information related to the first input information; and the first operation and maintenance model generates the first output information according to the first response information set.
[0020] In the embodiments of the present application, the first operation and maintenance model searches according to the first input information, determines the response information related to the first alarm information in the historical fault database and / or the running information set, generates the first output information, and further provides the tenant with the information related to the root cause fault corresponding to the first alarm information, so as to help the tenant to determine the root cause fault as soon as possible.
[0021] In combination with the first aspect, in some implementations, in the case that the root cause fault corresponding to the first alarm information is not included in the historical fault database, at least one of the following is added to the historical fault database: the first alarm information, the cause of the root cause fault corresponding to the first alarm information, the location of the root cause fault corresponding to the first alarm information, and the processing measure of the root cause fault corresponding to the first alarm information.
[0022] In the embodiments of the present application, when the first alarm information or the root cause fault corresponding to the first alarm information is included in the historical fault database, the first operation and maintenance model determines the related information of the root cause fault corresponding to the first alarm information by querying the historical fault database, and directly feeds back to the tenant, so as to improve the efficiency of fault processing. When the first alarm information or the root cause fault corresponding to the first alarm information is not included in the historical fault database, the first operation and maintenance model stores the related information of the root cause fault corresponding to the first alarm information to the historical fault database after generating the related information, so that when the first alarm information occurs again, the related information of the root cause fault corresponding to the first alarm information can be fed back to the tenant as soon as possible.
[0023] In combination with the first aspect, in some implementations, the first output information further includes: the location of the root cause fault corresponding to the first alarm information and / or the processing measure of the root cause fault corresponding to the first alarm information.
[0024] In the embodiments of the present application, the cloud management platform can provide the tenant with the cause and / or location of the root cause fault corresponding to the first alarm information, and can also provide the tenant with repair suggestions or generate a fault report, so as to improve the efficiency of fault processing.
[0025] The second aspect provides a computing device. The device includes modules for implementing the first aspect or any possible implementation of the first aspect.
[0026] The third aspect provides a computing device cluster including at least one computing device, each computing device including a processor and a memory; the processor of the at least one computing device is configured to execute instructions stored in the memory of the at least one computing device, so that the computing device cluster executes the method of the first aspect or any implementation of the first aspect.
[0027] In a fourth aspect, a computer program product including instructions, which when executed by a computer device cluster, cause the computer device cluster to perform the method of the first aspect or any of the implementations of the first aspect.
[0028] In a fifth aspect, a computer-readable storage medium includes computer program instructions, which when executed by a computer device cluster, cause the computer device cluster to perform the method of the first aspect or any of the implementations of the first aspect. BRIEF DESCRIPTION OF DRAWINGS
[0029] FIG. 1 is a schematic structural diagram of a fault processing system according to an embodiment of the present application.
[0030] FIG. 2 is a schematic flowchart of a fault processing method according to an embodiment of the present application.
[0031] FIG. 3 is a schematic diagram of a first diagnostic tree according to an embodiment of the present application.
[0032] FIG. 4 is a schematic diagram of a second diagnostic tree according to an embodiment of the present application.
[0033] FIG. 5 is a schematic diagram of a second graphical interface according to an embodiment of the present application.
[0034] FIG. 6 is a schematic flowchart of a method for determining first alarm information according to an embodiment of the present application.
[0035] FIG. 7 is a schematic structural block diagram of a computing device according to an embodiment of the present application.
[0036] FIG. 8 is a schematic structural diagram of a computing device according to an embodiment of the present application.
[0037] FIG. 9 is a schematic structural diagram of a computer device cluster according to an embodiment of the present application.
[0038] FIG. 10 is a schematic diagram of a connection between computing devices 800A and 800B through a network according to an embodiment of the present application. DETAILED DESCRIPTION
[0039] The technical solutions in the present application will be described below with reference to the accompanying drawings.
[0040] Embodiments of the present application will present various aspects, embodiments or features around a system including a plurality of devices, components, modules, etc. It should be understood and appreciated that each system can include additional devices, components, modules, etc., and / or can not include all of the devices, components, modules, etc. discussed in connection with the accompanying drawings. In addition, combinations of these solutions can also be used.
[0041] In addition, in the embodiments of the present application, the words "example" and "for example" are used to mean serving as an example or illustration. Any embodiment or design presented as an "example" in the embodiments of the present application should not be construed as preferred or advantageous over other embodiments or design. In fact, the word "example" is used to present concepts in a concrete manner.
[0042] The business scenarios described in the embodiments of the present application are used to more clearly illustrate the technical solutions of the embodiments of the present application, and do not constitute a limitation on the technical solutions provided by the embodiments of the present application. Those skilled in the art can know that, with the evolution of technology and the appearance of new business scenarios, the technical solutions provided by the embodiments of the present application are also applicable to similar technical problems.
[0043] In this specification, the reference "one embodiment" or "some embodiments" means that a particular feature, structure, or characteristic described in connection with the embodiment is included in at least one embodiment of the application. Thus, the appearances of the phrases "in one embodiment", "in some embodiments", "in other embodiments", "in additional embodiments", and so on, in various places in the specification are not necessarily all referring to the same embodiment, unless otherwise specified. The terms "comprise", "comprising", "have", "having", "include", "including", and "contain", "containing" mean "including but not limited to", unless otherwise specified.
[0044] In the embodiments of the present application, "at least one" means one or more, and "multiple" means two or more. The "and / or" describes the association relationship of the associated objects, which means that there can be three relationships, for example, A and / or B, which can represent the following cases: A exists alone, A and B exist together, and B exists alone, where A and B can be singular or plural. The character " / " generally represents an "or" relationship between the associated objects before and after it. "At least one of the following" or similar expressions means any combination of these items, including any combination of single item or multiple items. For example, at least one of a, b, or c can represent a, b, c, a-b, a-c, b-c, or a-b-c, where a, b, and c can be single or multiple.
[0045] The method in the embodiments of the present application can be applied to various cloud management platforms. The cloud management platform is used to manage an infrastructure providing cloud services, and the infrastructure includes at least one computing node. The computing node is, for example, a processing unit, a processor, a virtual machine, a container, a computing device, etc. The at least one computing node is used to run M services, where M is a positive integer. At least one of the M services is a service provided by a cloud vendor or a service deployed by a tenant.
[0046] FIG. 1 is a schematic structural diagram of a fault processing system 100 according to an embodiment of the present application. The fault processing system 100 in FIG. 1 includes a cloud management platform 110. The cloud management platform 110 can be used to manage an infrastructure providing cloud services, which includes at least one data center (for example, a data center 120). The data center 120 can include at least one computing node cluster, and each computing node cluster includes at least one computing node. For example, the data center 120 includes a computing node cluster 130 and a computing node cluster 140, the computing node cluster 130 includes a computing node 131 and a computing node 132, and the computing node cluster 140 includes a computing node 141 and a computing node 142. A tenant can apply for using resources in the data center 120 through the cloud management platform 110. The tenant is a tenant of a public cloud that registers a public cloud account and purchases public cloud resources of the public cloud.
[0047] In some embodiments, the plurality of computing nodes in the computing node cluster are directly connected or connected through a network, for example, a wide area network or a local area network, etc.
[0048] In some embodiments, the data center 120 can further include at least one storage node cluster, and each storage node cluster includes at least one storage node. The type of storage node is not limited in the embodiments of the present application, for example, the storage node is a centralized storage node or a distributed storage node. The storage node is used to store data of the tenant, or the storage node is used to store data required when the method in the embodiments of the present application is executed and / or data generated. Exemplarily, the plurality of storage nodes in the storage node cluster are directly connected or connected through a network, for example, a wide area network or a local area network, etc.
[0049] In some embodiments, the at least one data center is used to run M services, and M is a positive integer. The M services belong to a cloud vendor and / or a tenant.
[0050] The cloud management platform 110 is used to provide a first alarm information for the tenant in a case that at least one service of the M services fails and there are a plurality of alarm information. The first alarm information is used to indicate that at least one service fails, and the first alarm information is generated by summarizing alarm information triggered according to a same root fault from the plurality of alarm information according to a first operation and maintenance model. The first operation and maintenance model is obtained by training according to a historical fault database and deployment information of the M services. The plurality of alarm information is generated according to a running information set of the M services, and the running information set includes at least one running information.
[0051] In some embodiments, when a root cause failure occurs in at least one of the M services, the root cause failure can cause at least one other failure to occur, resulting in at least one alarm information corresponding to each failure. In other words, the root cause failure is the root cause of the at least one failure or at least one alarm information.
[0052] For example, in the case of a first root cause failure occurring in a computing device used to run a service, at least one of the following can occur: a computing instance (e.g., a virtual machine, a container, etc.) running in the computing device fails, a service running in the computing device or computing instance fails, another service that invokes the service fails, and thus at least one of the following alarm information can be triggered: alarm information indicating that the computing device fails, alarm information indicating that a computing instance (e.g., a virtual machine, a container, etc.) running in the computing device fails, alarm information indicating that a service running in the computing device fails, alarm information indicating that another service that invokes the service to implement a certain function fails, etc.
[0053] In some embodiments, the cloud management platform 110 inputs a plurality of alarm information into the first operation and maintenance model to obtain an output of the first operation and maintenance model: at least one first alarm information. Each of the at least one first alarm information is generated according to alarm information corresponding to the same root cause failure in the plurality of alarm information. For example, the first operation and maintenance model classifies, interprets, or summarizes the plurality of alarm information to generate the at least one first alarm information. In other words, the first operation and maintenance model is used to process the plurality of alarm information and summarize alarm information corresponding to the same root cause failure into one alarm information, thereby providing valuable alarm information to the tenant and reducing the number of alarm information provided to the tenant, thereby effectively suppressing alarm storms. The historical failure database and the deployment information of the M services are described in step 210.
[0054] In some embodiments, the cloud management platform 110 trains the initial large model according to the historical failure database and the deployment information of the M services to obtain the first operation and maintenance model. In other words, the cloud management platform 110 trains the initial large model using the historical failure database and the deployment information of the M services as training data, so that the obtained first operation and maintenance model can classify, interpret, or summarize a plurality of alarm information to generate first alarm information.
[0055] For example, the initial large model is an LLM.
[0056] In some embodiments, the cloud management platform 110 determines a first set of abnormal information according to the set of running information of the M services and a second operation and maintenance model. The second operation and maintenance model is obtained by training according to a set of historical running information of the M services. The first set of abnormal information includes at least one abnormal information, and each abnormal information in the at least one abnormal information is running information in the set of running information that meets an abnormal feature. The cloud management platform 110 further determines a plurality of alarm information according to the first set of abnormal information.
[0057] In some embodiments, the running information in the set of running information that meets the abnormal feature includes: a data index whose value does not belong to a preset range, a call link information of a call exception, log information that meets a preset format, etc. The log information that meets the preset format, for example, includes: log information containing a preset keyword, log information used to indicate a running error, etc. The preset keyword is not limited by the embodiments of the present application, and for example, includes error, warning, etc. The abnormal feature can be predefined by the cloud management platform, the tenant, or the operation and maintenance personnel.
[0058] In some embodiments, the cloud management platform 110 trains an initial model according to the set of historical running information of the M services to obtain the second operation and maintenance model. In other words, the cloud management platform 110 trains the initial model by taking the set of historical running information of the M services as training data, so that the obtained second operation and maintenance model can extract running information in the set of running information that meets the abnormal feature. For details, refer to the description in FIG. 2 or FIG. 6.
[0059] The cloud management platform 110 is further configured to receive first input information of a tenant. The first input information is used to request a cause of occurrence of a root cause failure corresponding to first alarm information. The cloud management platform 110 is further configured to provide first output information for the tenant in response to the first input information. The first output information is used to indicate the cause of occurrence of the root cause failure corresponding to the first alarm information. The first output information is generated according to the first operation and maintenance model. In other words, the cloud management platform 110 can perform natural language dialogue with the tenant through the first operation and maintenance model, so as to provide the tenant with the cause of occurrence of the root cause failure corresponding to the first alarm information.
[0060] In some embodiments, the cloud management platform 110 inputs the first input information into the first operation and maintenance model, so that the first operation and maintenance model performs retrieval in the historical fault database and / or the set of running information of the M services according to the first input information, determines a first set of response information, and generates the first output information according to the first set of response information. The first set of response information includes at least one response information related to the first input information.
[0061] In some embodiments, the first output information comprises at least one of the following: a cause of the root fault corresponding to the first alarm information, a location of the root fault corresponding to the first alarm information, a processing measure of the root fault corresponding to the first alarm information.
[0062] In some embodiments, the first operation and maintenance model queries a historical fault database according to the first input information, and if the first alarm information or the root fault corresponding to the first alarm information does not exist in the historical fault database, the first operation and maintenance model queries the running information set of the M services, so as to determine the first response information set according to the historical fault database and / or the running information set. If the first alarm information or the root fault corresponding to the first alarm information exists in the historical fault database, the first operation and maintenance model determines the first response information set according to the related information of the root fault corresponding to the first alarm information included in the historical fault database. The first response information set comprises the related information of the root fault corresponding to the first alarm information. The related information of the root fault corresponding to the first alarm information comprises at least one of the following: alarm information corresponding to the root fault, a cause of the root fault, a location of the root fault, a processing measure corresponding to the root fault, and the like.
[0063] In some embodiments, in the case that the first alarm information or the root fault corresponding to the first alarm information does not exist in the historical fault database, at least one of the following is added to the historical fault database: the first alarm information, a cause of the root fault corresponding to the first alarm information, a location of the root fault corresponding to the first alarm information, a processing measure of the root fault corresponding to the first alarm information, the first output information, and the like.
[0064] Optionally, the cloud management platform 110 can also receive second input information of the tenant, and provide second output information for the tenant in response to the second input information. The second input information is used to request to obtain a cause of the root fault corresponding to the first alarm information. The second input information comprises different content from the first input information. The second output information is used to indicate the cause of the root fault corresponding to the first alarm information, and the second output information is generated according to the first operation and maintenance model. The second output information comprises the same or different content from the first output information, and the embodiments of the present application are not limited thereto. In other words, the cloud management platform 110 can perform multiple rounds of natural language dialogue with the tenant through the first operation and maintenance model, so as to gradually determine the cause of the root fault corresponding to the alarm information according to the alarm information provided by the first operation and maintenance model, thereby improving the efficiency of fault processing.
[0065] The fault processing system 100 in FIG. 1 can process at least one piece of alarm information triggered by the same root fault from the multiple pieces of alarm information according to the first operation and maintenance model, generate one piece of valuable alarm information corresponding to each root fault, and thus reduce the number of alarm information provided to the tenant, thereby effectively suppressing the alarm storm. Meanwhile, the fault processing system 100 can interact with the tenant in natural language through the first operation and maintenance model, and thus provide the tenant with the cause of the fault according to the input information of the tenant, improve the efficiency of determining the cause of the fault, and thus facilitate the tenant to solve the fault according to the cause of the fault and improve the efficiency of fault processing.
[0066] FIG. 2 is a schematic flowchart of a fault processing method according to an embodiment of the present application. The method in FIG. 2 can be performed by a cloud management platform, for example, the cloud management platform 110 in FIG. 1. The method in FIG. 2 includes the following steps.
[0067] 210. In a case where at least one service of the M services fails and there are multiple pieces of alarm information, providing a first piece of alarm information for the tenant.
[0068] When at least one root fault occurs in at least one service of the M services, each root fault of the at least one root fault can cause at least one other fault in addition to the root fault, so that the cloud management platform generates multiple pieces of alarm information. That is, each piece of alarm information of the multiple pieces of alarm information corresponds to one fault, or each piece of alarm information is triggered by one fault. The multiple pieces of alarm information correspond to at least one root fault. The cloud management platform generates a first piece of alarm information according to the multiple pieces of alarm information and the first operation and maintenance model, and provides the first piece of alarm information to the tenant. The first piece of alarm information is used to indicate that at least one service fails, and the first piece of alarm information is generated by summarizing alarm information triggered by the same root fault from the multiple pieces of alarm information according to the first operation and maintenance model. In other words, the cloud management platform processes the multiple pieces of alarm information through the first operation and maintenance model, thereby summarizing alarm information corresponding to the same root fault into one piece of alarm information, and then providing valuable alarm information to the tenant, reducing the number of alarm information provided to the tenant, and effectively suppressing the alarm storm.
[0069] In some embodiments, an alarm storm refers to a large number of alarm information generated in a short period of time, so that the number of alarm information exceeds the upper limit that can be processed by an operation and maintenance personnel.
[0070] In some embodiments, the M services belong to a cloud vendor and / or a tenant. The M services run in at least one computing node.
[0071] Optionally, before step 210, the cloud management platform obtains a first operation and maintenance model. The first operation and maintenance model is obtained according to a historical fault database and deployment information of the M services.
[0072] In some embodiments, the historical fault database comprises at least one of: at least one historical operation and maintenance case, a fault mode library, or a diagnosis tree. The historical operation and maintenance case comprises at least one of: historical fault alarm information, a location of a fault corresponding to the historical fault alarm information, a cause of the fault corresponding to the historical fault alarm information, and a processing measure of the fault corresponding to the historical fault alarm information. The fault mode library comprises at least one of: identification information of the fault, alarm information corresponding to the fault, a location of the fault, a cause of the fault, a processing measure of the fault, a level of the fault, and an association relationship between faults. The diagnosis tree comprises a first diagnosis tree and / or a second diagnosis tree, the first diagnosis tree being used to determine a root cause of a fault according to fault alarm information, and the second diagnosis tree being used to determine a processing measure of the fault according to the root cause of the fault.
[0073] For example, the level of the fault is used to indicate a severity of the fault. The severity of the fault is determined according to at least one of: a semantic of the alarm information of the fault, an alarm frequency of the alarm information of the fault, an alarm periodicity of the alarm information of the fault, an alarm time of the alarm information of the fault, and a type of the fault. For example, the lower the alarm frequency, the higher the severity of the fault. Or, the stronger the periodicity of the alarm, the lower the severity of the fault. Or, the higher the severity of the fault corresponding to the alarm information that occurs when the business is busy. Or, the higher the severity of the fault corresponding to the alarm information that suddenly occurs after the service runs for a period of time. Or, the higher the severity of the application-level fault relative to the operating system-level fault, the network-level fault, and the memory-level fault. Or, the severity of the fault is predefined by the cloud management platform, the tenant, or the operation and maintenance personnel.
[0074] For example, the association relationship between the faults is used to indicate a hierarchical relationship between the multiple faults and / or whether the multiple faults belong to the same type. For example, if the occurrence of fault 1 has a direct or indirect relationship with fault 2, there is a hierarchical relationship between fault 1 and fault 2. If the occurrence of fault 1 has nothing to do with fault 2, there is no hierarchical relationship between fault 1 and fault 2.
[0075] For example, the first diagnosis tree can be represented as a tree structure. The root node of the first diagnosis tree is used to indicate the fault alarm information, the first-level leaf node is a direct cause of triggering the fault alarm information, the second-level leaf node is a direct cause of triggering the fault cause represented by the first-level leaf node, and so on. That is, the leaf node of the first diagnosis tree represents a direct or indirect cause of triggering the fault alarm information.
[0076] For example, the first diagnostic tree is shown in FIG. 3. FIG. 3 is a schematic diagram of a first diagnostic tree 300 according to an embodiment of the present application. In the first diagnostic tree 300, the root node is fault alarm information 1, the first level leaf nodes include three leaf nodes, and the second level leaf nodes include three leaf nodes. The three leaf nodes in the first level leaf nodes are respectively used to indicate fault cause 1, fault cause 2, and fault cause 3, which are direct causes of triggering the fault alarm information 1. The three leaf nodes in the second level leaf nodes are respectively used to indicate fault cause 1-1, fault cause 1-2, and fault cause 1-3, which are direct causes of causing the fault cause 1. That is, the fault cause 1-1, the fault cause 1-2, and the fault cause 1-3 are indirect causes of triggering the fault alarm information 1. It can also be seen from FIG. 3 that the fault alarm information 1 corresponds to a root cause of a fault, and the root cause of the fault is at least one of the following: the fault cause 1-1, the fault cause 1-2, the fault cause 1-3, the fault cause 2, and the fault cause 3.
[0077] Exemplarily, the second diagnostic tree can be represented as a tree structure. The root node of the second diagnostic tree is used to represent a fault cause, the first level leaf nodes are direct handling measures for solving the fault cause, and the second level leaf nodes are handling measures for further solving the fault cause on the basis of the first level leaf nodes, and so on. That is, the leaf nodes of the second diagnostic tree represent handling measures for directly or indirectly solving the fault cause.
[0078] For example, the second diagnostic tree is shown in FIG. 4. FIG. 4 is a schematic diagram of a second diagnostic tree 400 according to an embodiment of the present application. In the second diagnostic tree 400, the root node is fault cause 1, the first level leaf nodes include three leaf nodes, and the second level leaf nodes include three leaf nodes. The three leaf nodes in the first level leaf nodes are respectively used to indicate handling measure 1, handling measure 2, and handling measure 3, which are direct handling measures for solving the fault cause 1. The three leaf nodes in the second level leaf nodes are respectively used to indicate handling measure 1-1, handling measure 1-2, and handling measure 1-3, which are handling measures for further solving the fault cause 1 on the basis of the handling measure 1. That is, the handling measure 1-1, the handling measure 1-2, and the handling measure 1-3 are indirect handling measures for solving the fault cause 1.
[0079] It should be understood that FIG. 3 and FIG. 4 are only exemplary descriptions, and the number of leaf nodes at each level or the number of leaf nodes of each leaf node in the first diagnostic tree or the second diagnostic tree is not limited according to the embodiments of the present application.
[0080] In some embodiments, the deployment information of the M services comprises at least one of: information of a computing node cluster in which the computing nodes running the M services are located, an association relationship between the M services, a service topology of the M services.
[0081] Exemplarily, the information of the computing node cluster in which the computing nodes running the M services are located comprises at least one of: identification information of the computing node cluster in which the computing node running each of the M services is located, information of the computing nodes belonging to the same computing node cluster among the computing nodes running the M services, and the like.
[0082] Exemplarily, the computing node cluster is, for example, a computing device cluster at a region level or an availability zone (AZ) level.
[0083] Exemplarily, the association relationship between the M services comprises at least one of: a dependency relationship between the M services, a hierarchical relationship between the M services. The dependency relationship is used to indicate information of other services that need to be called when running one service. The hierarchical relationship is used to indicate information of at least one sub-service included in one service or information of a service to which one sub-service belongs. The sub-service belongs to the M services.
[0084] Exemplarily, the service topology of the M services is used to describe the dependency relationship between the services. The service topology comprises at least one node and at least one edge connecting the at least one node, each node in the at least one node is used to represent one service, and the edge connecting two nodes is used to represent a request between two services. The health status of the node or the edge can also be represented by different colors in the service topology.
[0085] In some embodiments, the first operation and maintenance model is used to classify or denoise the plurality of alarm information according to the levels of the faults corresponding to the plurality of alarm information, thereby generating the first alarm information. The first alarm information is generated according to at least one alarm information corresponding to the same type or the same level of fault in the plurality of alarm information.
[0086] In some embodiments, the first operation and maintenance model is used to determine an association relationship between the plurality of alarm information according to the semantics of each alarm information in the plurality of alarm information, thereby generating the first alarm information according to the association relationship between the plurality of alarm information. The first alarm information is generated according to at least one alarm information having the association relationship in the plurality of alarm information.
[0087] Exemplarily, the association relationship between the alarm information is determined according to at least one of a triggering time, a triggering position, and semantic correlation of the alarm information. For example, the first operation and maintenance model associates at least one alarm information in a preset time window into a set describing a specific problem in a current scene according to the association scene and the triggering time, the triggering position, and the semantic correlation, so as to generate the first alarm information according to the set.
[0088] In some embodiments, the cloud management platform directly obtains the first operation and maintenance model that has been trained. Alternatively, the cloud management platform trains an initial large model according to the historical fault database and the deployment information of the M services to obtain the first operation and maintenance model.
[0089] Exemplarily, the cloud management platform trains an initial large model according to the historical fault database and the deployment information of the M services as training data by a supervised fine-tuning (SFT) technology, so as to obtain the first operation and maintenance model. The first operation and maintenance model is used for classifying, interpreting, summarizing, and the like of the at least one alarm information, so as to aggregate at least one alarm information triggered according to the same root fault into one alarm information output.
[0090] Exemplarily, the cloud management platform trains an initial large model according to at least one of the historical fault database, the deployment information of the M services, and a historical running information set of the M services to obtain the first operation and maintenance model.
[0091] Exemplarily, the initial large model is an LLM. The specific type of the LLM is not limited in the embodiments of the present application, and any LLM that can realize the above functions in the prior art can be applied to the method provided in the embodiments of the present application.
[0092] In some embodiments, the historical fault database or the deployment information of the M services can be in various forms, such as text, image, and the like. The specific form of the historical fault database or the deployment information of the M services is not limited in the embodiments of the present application. When the historical fault database and / or the deployment information of the M services is in the form of an image, the cloud management platform performs image-to-text processing on the image to obtain text information in the image and an association relationship between the text information, so as to train the initial large model by taking the text information in the image and the association relationship between the text information as training data.
[0093] Optionally, the plurality of alarm information is generated according to a running information set of the M services, and the running information set includes at least one running information.
[0094] In some embodiments, the running information set comprises at least one of the following: a data metric information set, a call link information set, and a log information set. The data metric information set comprises at least one data metric in the running of the M services. The call link information set comprises at least one call link information in the processing of a request in the running of the M services. The log information set comprises at least one log information in the running of the M services.
[0095] In some embodiments, the cloud management platform determines a first exception information set according to the running information set of the M services and a second operation and maintenance model. The second operation and maintenance model is obtained by training historical running information sets of the M services. Each exception information in the first exception information set is running information in the running information set that meets an exception feature. The cloud management platform further determines a plurality of alarm information according to the first exception information set. See the description in step 610.
[0096] In some embodiments, the running information in the running information set that meets the exception feature comprises a data metric whose value is out of a preset range, a call link information that is abnormal, or log information that meets a preset format. The log information that meets the preset format comprises, for example, log information that contains a preset keyword or log information that indicates a running error. The preset keyword is not limited in the embodiments of the present application, and may comprise, for example, error and warning.
[0097] In some embodiments, before step 210, the cloud management platform obtains a second operation and maintenance model that has been trained. Alternatively, the cloud management platform trains an initial model according to historical running information sets of the M services to obtain the second operation and maintenance model.
[0098] For example, the initial model is FM. The specific type of the FM is not limited in the embodiments of the present application, and any FM that can realize the above functions in the prior art can be applied to the method provided in the embodiments of the present application.
[0099] Optionally, the cloud management platform provides a first graphical interface for the tenant and displays the first alarm information in the first graphical interface, so that the tenant can view the first alarm information through the first graphical interface. The specific form of the first graphical interface is not limited in the embodiments of the present application.
[0100] 220, receiving first input information of the tenant.
[0101] The cloud management platform receives first input information of the tenant, and the first input information is used to request the cause of the root cause failure corresponding to the first alarm information.
[0102] Optionally, the cloud management platform provides a second graphical interface for the tenant, so that the tenant can select, input or upload the first input information in the second graphical interface. Embodiments of the present application do not limit the specific forms of the second graphical interface.
[0103] 230, in response to the first input information, providing the tenant with first output information.
[0104] After the cloud management platform receives the first input information, the cloud management platform provides the tenant with first output information in response to the first input information. The first output information is used to indicate the cause of the root cause failure corresponding to the first alarm information. The first output information is generated according to the first operation and maintenance model.
[0105] For example, the first input information is a question input by the tenant, and the first output information is an answer to the question.
[0106] Optionally, the cloud management platform provides a third graphical interface for the tenant, and displays the first output information in the third graphical interface, so that the tenant can view the first output information through the third graphical interface. Embodiments of the present application do not limit the specific forms of the third graphical interface. Alternatively, the cloud management platform displays the first output information in the second graphical interface, so that the tenant can view the first output information through the second graphical interface.
[0107] Exemplarily, the second graphical interface is shown in FIG. 5. FIG. 5 is a schematic diagram of a second graphical interface 500 provided by an embodiment of the present application. As shown in FIG. 5, the second graphical interface 500 includes a control 510 and a control 520. The control 510 is used to display the first input information input by the tenant, and the control 520 is used to display the first output information in response to the first input information. Exemplarily, the second graphical interface 500 can also include a control 530, which is used to receive the input of the tenant, such as the first input information, the second input information, etc.
[0108] Optionally, the cloud management platform inputs the first input information into the first operation and maintenance model, thereby obtaining the output of the first operation and maintenance model, i.e. the first output information. The first operation and maintenance model searches in the historical fault database and / or the running information set of the M services according to the first input information, to determine a first response information set. The first response information set includes at least one response information related to the first input information. The first operation and maintenance model also generates the first output information according to the first response information set.
[0109] In some embodiments, the first operation and maintenance model determines the first set of response information according to a retrieval-augmented generation (RAG) technology, and performs retrieval in the historical fault database and / or the set of running information of the M services.
[0110] In some embodiments, the first operation and maintenance model queries the historical fault database according to the first input information, and if the first alarm information does not exist in the historical fault database or the fault corresponding to the first alarm information does not exist, the first operation and maintenance model queries the set of running information of the M services to determine the first set of response information. At least one response information in the first set of response information is determined according to information related to the first alarm information in the historical fault database and / or the set of running information. For example, the at least one response information is information related to the first alarm information in the historical fault database and / or the set of running information, or the at least one response information is determined by summarizing information related to the first alarm information in the historical fault database and / or the set of running information. The information related to the first alarm information in the historical fault database and / or the set of running information includes at least one of the following: alarm information similar to the first alarm information in the historical fault database, related information of the alarm information similar to the first alarm information in the historical fault database, running information in the set of running information for generating the first alarm information, and alarm information related to the running information for generating the first alarm information in the set of running information. If the first alarm information exists in the historical fault database or the fault corresponding to the first alarm information exists, the first operation and maintenance model determines the first set of response information according to related information of the first alarm information included in the historical fault database. The first set of response information includes the related information of the first alarm information. The related information of the first alarm information includes at least one of the following: a root fault corresponding to the first alarm information, a cause of the root fault, a location of the root fault, a treatment measure corresponding to the root fault, and the like.
[0111] In some embodiments, in the case that the first alarm information or the fault corresponding to the first alarm information does not exist in the historical fault database, the cloud management platform adds at least one of the following to the historical fault database: the first alarm information, a root fault corresponding to the first alarm information, a cause of the root fault corresponding to the first alarm information, a location of the root fault corresponding to the first alarm information, a treatment measure of the root fault corresponding to the first alarm information, first output information, and the like. In other words, the cloud management platform can expand the historical fault database, so that when the first alarm information occurs again subsequently, the cloud management platform can directly provide the cause of the root fault corresponding to the first alarm information according to information included in the historical fault database, thereby improving the efficiency of fault handling.
[0112] Optionally, the first output information comprises at least one of the following: a cause of the root fault corresponding to the first alarm information, a location of the root fault corresponding to the first alarm information, and a processing measure of the root fault corresponding to the first alarm information.
[0113] Optionally, the cloud management platform can further receive second input information of the tenant, and in response to the second input information, provide second output information for the tenant. The second input information is used to request a cause of the root fault corresponding to the first alarm information. The second input information is different from the content included in the first input information. The second output information is used to indicate the cause of the root fault corresponding to the first alarm information, and the second output information is generated according to the first operation and maintenance model. The second output information is the same as or different from the content included in the first output information, and the embodiments of the present application are not limited thereto. In other words, the cloud management platform can perform multiple rounds of natural language dialogue with the tenant through the first operation and maintenance model, so as to gradually determine the cause of the root fault corresponding to the alarm information according to the alarm information provided by the first operation and maintenance model, thereby improving the efficiency of fault handling.
[0114] In some embodiments, after the cloud management platform inputs the first input information into the first operation and maintenance model, the first operation and maintenance model generates a first diagnosis tree and / or a second diagnosis tree corresponding to the first alarm information according to the deployment information of the M services and the historical fault database. According to the first diagnosis tree, the first operation and maintenance model determines the cause of the root fault corresponding to the first alarm information through multiple rounds of natural language dialogue with the tenant, starting from the root node (i.e., the first alarm information) of the first diagnosis tree, and gradually diagnoses layer by layer, and provides the tenant. And / or, according to the second diagnosis tree, the first operation and maintenance model determines the processing measure of the root fault corresponding to the first alarm information through multiple rounds of natural language dialogue with the tenant, starting from the root node (i.e., the cause of the root fault corresponding to the first alarm information) of the second diagnosis tree, and gradually investigates layer by layer, and provides the tenant.
[0115] The method in FIG. 2 can be used to process at least one alarm information triggered by the same root fault among multiple alarm information according to the first operation and maintenance model when at least one service among the M services fails, thereby generating one valuable alarm information corresponding to each root fault, reducing the number of alarm information provided for the tenant, and effectively suppressing the alarm storm. At the same time, the cloud management platform can interact with the tenant through natural language according to the first operation and maintenance model, thereby providing the tenant with the cause of the fault according to the input information of the tenant, improving the efficiency of determining the cause of the fault, and thereby facilitating the tenant to solve the fault according to the cause of the fault, and improving the efficiency of fault handling.
[0116] In some embodiments, the determination of the first alarm information in step 210 is as shown in FIG. 6.
[0117] FIG. 6 is a schematic flowchart of a method for determining the first alarm information according to an embodiment of the present application. The method in FIG. 6 can be performed by the cloud management platform, for example, the cloud management platform 110 in FIG. 1. The method in FIG. 6 includes the following steps.
[0118] 610, obtaining the first operation and maintenance model, the second operation and maintenance model, and a set of running information of the M services.
[0119] Optionally, the cloud management platform directly obtains the trained first operation and maintenance model. Alternatively, the cloud management platform trains the first operation and maintenance model according to the historical fault database and the deployment information of the M services. Alternatively, the cloud management platform trains the first operation and maintenance model according to the historical fault database, the deployment information of the M services, and a set of historical running information of the M services. M is a positive integer. The implementation of the cloud management platform training the first operation and maintenance model is described in step 210. The first operation and maintenance model, the historical fault database, the deployment information of the M services, and the set of running information of the M services are described in step 210.
[0120] Optionally, the cloud management platform directly obtains the trained second operation and maintenance model. Alternatively, the cloud management platform trains the second operation and maintenance model according to the set of historical running information of the M services. The second operation and maintenance model is described in step 210. The set of historical running information of the M services is a set of running information generated by the M services in the historical running process.
[0121] In some embodiments, the second operation and maintenance model includes at least one of the following: a first sub-model, a second sub-model, or a third sub-model. The first sub-model is used to determine a first set of data indicator information from the set of data indicator information of the M services, the first set of data indicator information including at least one abnormal first data indicator, i.e., each first data indicator in the first set of data indicator information meets the abnormal characteristic. The second sub-model is used to determine a first set of call link information from the set of call link information of the M services, the first set of call link information including at least one abnormal call link information, i.e., each call link information in the first set of call link information meets the abnormal characteristic. The third sub-model is used to determine a first set of log information from the set of log information of the M services, the first set of log information including at least one abnormal log information, i.e., each log information in the first set of log information meets the abnormal characteristic.
[0122] Exemplarily, the data indicator conforming to the abnormal feature includes a data indicator whose value does not belong to a preset range. The call link information conforming to the abnormal feature includes call link information of a call exception. The call exception has a similar meaning to call failure or call error, and can be replaced by each other. The log information conforming to the abnormal feature includes log information conforming to a preset format. The log information conforming to the preset format includes, for example, log information containing a preset keyword, log information used to indicate a running error, and the like. The preset keyword is not limited in the embodiments of the present application, and includes, for example, error, warning, and the like.
[0123] Exemplarily, the abnormal feature described above can be predefined by the cloud management platform, the tenant, or the operation and maintenance personnel, and the meaning of the abnormal feature is not limited in the embodiments of the present application.
[0124] In some embodiments, the cloud management platform trains the first initial model according to the historical data indicator information set of the M services, to obtain the first sub-model. And / or, the cloud management platform trains the second initial model according to the historical call link information set of the M services, to obtain the second sub-model. And / or, the cloud management platform trains the third initial model according to the historical log information set of the M services, to obtain the third sub-model.
[0125] Exemplarily, the first initial model, the second initial model, and the third initial model are FM. The specific type of the FM is not limited in the embodiments of the present application, and any FM that can realize the functions described above in the prior art can be applied to the method provided in the embodiments of the present application.
[0126] Optionally, the cloud management platform directly obtains the real-time running information set of the M services. Alternatively, the cloud management platform manages the running of the M services, to generate the real-time running information set of the M services.
[0127] 620, determining a first abnormal information set according to the second operation and maintenance model and the running information set of the M services.
[0128] After obtaining the second operation and maintenance model and the real-time running information set of the M services, the cloud management platform inputs the real-time running information set of the M services into the second operation and maintenance model, to obtain the output of the second operation and maintenance model, that is, to obtain the first abnormal information set. In other words, in the real-time running of the M services, the cloud management platform determines the abnormal information in the running information set according to the running information set of the M services, to facilitate the generation of multiple alarm information.
[0129] In some embodiments, the cloud management platform inputs a set of real-time data indicator information of the M services into the first sub-model, obtains the output of the first sub-model, i.e., obtains a first set of data indicator information. And / or, the cloud management platform inputs a set of real-time call link information of the M services into the second sub-model, obtains the output of the second sub-model, i.e., obtains a first set of call link information. And / or, the cloud management platform inputs a set of real-time log information of the M services into the third sub-model, obtains the output of the third sub-model, i.e., obtains a first set of log information. The first set of abnormal information includes at least one of the following: the first set of data indicator information, the first set of call link information, and the first set of log information.
[0130] 630, according to the first set of abnormal information, determining at least one alarm information.
[0131] After the cloud management platform obtains the first set of abnormal information, the cloud management platform analyzes and processes at least one abnormal information included in the first set of abnormal information, thereby generating at least one alarm information.
[0132] In some embodiments, the cloud management platform inputs the first set of abnormal information into an artificial intelligence for operations (AIOps) model, obtains the output of the AIOps model, i.e., obtains at least one alarm information. The AIOps model is a model that empowers traditional internet technology (IT) operation and maintenance management with big data, artificial intelligence or machine learning technology. The AIOps model is used for intelligent analysis of data, thereby realizing functions such as abnormal detection, root cause analysis of faults, optimization suggestions, etc. The specific type of the AIOps model is not limited in the embodiments of the present application, and the AIOps model that can realize the above functions in the prior art can be applied to the method provided in the embodiments of the present application.
[0133] 640, according to the first operation and maintenance model and the at least one alarm information, determining a first alarm information.
[0134] After the cloud management platform obtains the at least one alarm information, the cloud management platform inputs the at least one alarm information into the first operation and maintenance model, obtains the output of the first operation and maintenance model, i.e., obtains a first alarm information. The first alarm information is generated by summarizing the alarm information triggered by the same root fault in the at least one alarm information according to the first operation and maintenance model.
[0135] In some embodiments, the first operation and maintenance model performs alarm noise reduction, alarm deduplication, alarm merging, and alarm summarization on the at least one alarm information, thereby providing the tenant with an alarm analysis result, i.e., the first alarm information, in natural language.
[0136] In the method of FIG. 6, the cloud management platform analyzes the real-time running information set of the M services in real time through the second operation and maintenance model, determines the abnormal running information in the running information set, and then generates at least one alarm information according to the abnormal running information. Meanwhile, the cloud management platform processes the at least one alarm information through the first operation and maintenance model, generates the first alarm information, thereby reducing the number of alarm information provided for the tenant, providing only important and valuable alarm information for the tenant, thereby effectively suppressing the alarm storm and improving the efficiency of the tenant in handling faults.
[0137] FIG. 7 is a schematic structural diagram of a computing device provided by an embodiment of the present application. The computing device 700 in FIG. 7 includes a sending module 710 and a receiving module 720. The computing device 700 in FIG. 7 can be applied to a cloud management platform, for example, the cloud management platform in FIG. 1.
[0138] When the computing device 700 is used to execute the method in FIG. 2, the sending module 710 is configured to provide the first alarm information for the tenant in the case that at least one service of the M services fails and there are multiple alarm information. The receiving module 720 is configured to receive the first input information of the tenant. The receiving module 720 is configured to execute step 220 in FIG. 2. The sending module 710 is further configured to provide the first output information for the tenant in response to the first input information. The sending module 710 is configured to execute steps 210 and 230 in FIG. 2. The multiple alarm information, the first alarm information, the first input information, and the first output information are described in FIG. 2.
[0139] In some embodiments, when the computing device 700 is used to execute the method in FIG. 6, the computing device 700 further includes a processing module (not shown in the figure). The processing module is configured to obtain the first operation and maintenance model, the second operation and maintenance model, and the running information set of the M services; determine the first abnormal information set according to the second operation and maintenance model and the running information set of the M services; determine at least one alarm information according to the first abnormal information set; and determine the first alarm information according to the first operation and maintenance model and the at least one alarm information. The processing module is configured to execute steps 610-640 in FIG. 6. The first operation and maintenance model, the second operation and maintenance model, the running information set, and the first abnormal information set are described in FIG. 2 or FIG. 6.
[0140] The sending module 710 and the receiving module 720 can be implemented by software or by hardware. For example, the implementation of the sending module 710 is described as follows. Similarly, the implementation of the receiving module 720 can refer to the implementation of the sending module 710.
[0141] As an example of a software functional unit, the sending module 710 can include code running on a compute instance. The compute instance can include at least one of a physical host (computing device), a virtual machine, a container. Further, the compute instance can be one or more. For example, the sending module 710 can include code running on multiple hosts / virtual machines / containers. It is noted that the multiple hosts / virtual machines / containers running the code can be distributed in the same region, or in different regions. Further, the multiple hosts / virtual machines / containers running the code can be distributed in the same availability zone (AZ), or in different AZs, each of which includes one data center or multiple data centers in close geographical proximity. Typically, a region can include multiple AZs.
[0142] Similarly, the multiple hosts / virtual machines / containers running the code can be distributed in the same VPC, or in multiple VPCs. Typically, a VPC is set up within a region, and communication between two VPCs in the same region, or between VPCs in different regions, requires a communication gateway in each VPC to enable interconnection between VPCs.
[0143] As an example of a hardware functional unit, the sending module 710 can include at least one computing device, such as a server, etc. Alternatively, the sending module 710 can be a device implemented using an application-specific integrated circuit (ASIC), or a programmable logic device (PLD), etc. The PLD can be implemented using a complex programmable logical device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL), or any combination thereof.
[0144] The multiple computing devices included in the sending module 710 can be distributed in the same region or in different regions. The multiple computing devices included in the sending module 710 can be distributed in the same AZ or in different AZs. Similarly, the multiple computing devices included in the sending module 710 can be distributed in the same VPC or in multiple VPCs. The multiple computing devices can be any combination of servers, ASICs, PLDs, CPLDs, FPGAs, and GALs.
[0145] Therefore, the modules of the examples described in the embodiments of the present application can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether the functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. A person skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the present application.
[0146] It should be noted that: the device provided by the above-mentioned embodiments is in executing the above-mentioned method, only above-mentioned each functional module is divided and carries out example explanation, actually applies, can according to need to complete by different functional module to the functional distribution of above-mentioned, namely divides the internal structure of device into different functional module, to complete above description all or partial function. For example, the sending module 710 can be used to execute any step in the above method, and the receiving module 720 can be used to execute any step in the above method. The steps responsible for the sending module 710 and the receiving module 720 can be specified as needed, and the functions of the above device are realized by the sending module 710 and the receiving module 720 respectively implementing different steps in the above method.
[0147] In addition, the device and method embodiments provided by the above-mentioned embodiments belong to the same concept, and the specific implementation process is described in the method embodiments above, which will not be repeated here.
[0148] The method provided by the embodiments of the present application can be executed by a computing device, which can also be referred to as a computer system. The computer system includes a hardware layer, an operating system layer running on the hardware layer, and an application layer running on the operating system layer. The hardware layer includes hardware such as a processing unit, a memory, and a memory control unit, and the functions and structures of the hardware are described in detail later. The operating system is any one or more computer operating systems that implement business processing through processes, such as a Linux operating system, a Unix operating system, an Android operating system, an iOS operating system, or a windows operating system. The application layer includes application programs such as a browser, an address book, word processing software, and instant messaging software. Optionally, the computer system is a handheld device such as a smartphone or a terminal device such as a personal computer, and the present application is not particularly limited as long as the method provided by the embodiments of the present application can be executed. The execution subject of the method provided by the embodiments of the present application can be a computing device, or a functional module in the computing device that can call and execute a program.
[0149] FIG. 8 is a schematic structural block diagram of a computing device 800 provided by an embodiment of the present application. The computing device 800 can be a server or a computer or other device with computing capability. The computing device 800 shown in FIG. 8 includes at least one processor 810 and a memory 820.
[0150] It should be understood that the present application does not limit the number of processors and memories in the computing device 800.
[0151] The processor 810 executes instructions in the memory 820, so that the computing device 800 implements the method provided by the present application. Alternatively, the processor 810 executes instructions in the memory 820, so that the computing device 800 implements each functional module provided by the present application, thereby implementing the method provided by the present application.
[0152] Optionally, the computing device 800 further includes a communication interface 830. The communication interface 830 uses a transceiving module such as, but not limited to, a network interface card and a transceiver to implement communication between the computing device 800 and other devices or communication networks.
[0153] Optionally, the computing device 800 further includes a system bus 840, wherein the processor 810, the memory 820 and the communication interface 830 are connected with the system bus 840 respectively. The processor 810 can access the memory 820 through the system bus 840, for example, the processor 810 can read and write data in the memory 820 or execute code in the memory 820 through the system bus 840. The system bus 840 is a peripheral component interconnect express (PCI) bus or an extended industry standard architecture (EISA) bus, etc. The system bus 840 is divided into an address bus, a data bus, a control bus, etc. For the convenience of representation, only one thick line is shown in FIG. 8, but it does not mean that there is only one bus or only one type of bus.
[0154] In one possible implementation, the function of the processor 810 is mainly to interpret the instructions (or code) of the computer program and process the data in the computer software. The instructions of the computer program and the data in the computer software can be saved in the memory 820 or the cache of the processor 810.
[0155] Optionally, the processor 810 can be an integrated circuit chip with a processing capability of signals. As an example but not limitation, the processor 810 is a general purpose processor, a digital signal processor (DSP), an ASIC, an FPGA or other programmable logic device, a discrete gate or transistor logic device, a discrete hardware component. The general purpose processor is a microprocessor, etc. For example, the processor 810 is a central processing unit (CPU).
[0156] The memory 820 can provide a running space for a process in the computing device 800, for example, the memory 820 saves the computer program (specifically, the code of the program) used to generate the process. After the computer program is run by the processor to generate the process, the processor allocates a corresponding storage space for the process in the memory 820. Further, the above-mentioned storage space further includes a text segment, an initialized data segment, a bit initialized data segment, a stack segment, a heap segment, etc. The memory 820 saves the data generated during the running of the process in the above-mentioned storage space of the process, for example, intermediate data, process data, etc.
[0157] Optionally, memory (also referred to as the "memory") is used for the temporary storage of data exchanged between the processor 810 and an external memory, such as a hard disk. As long as the computer is running, the processor 810 will call the data needed for operation to the memory for operation, and after the operation is completed, the results will be transmitted.
[0158] By way of example, and not limitation, memory 820 is volatile memory or nonvolatile memory, or can include both volatile and nonvolatile memory. By way of example, and not limitation, nonvolatile memory, such as ROM, can be used for storage of data
[0159] The structure of the computing device 800 listed above is only an example, and the application is not limited thereto. The computing device 800 of the embodiments of the application includes various hardware in the prior art computer system, for example, the computing device 800 also includes other memories in addition to the memory 820, such as disk memories and the like. Those skilled in the art should understand that the computing device 800 can also include other devices necessary for normal operation. Meanwhile, according to specific needs, those skilled in the art should understand that the above computing device 800 can also include hardware devices for implementing other additional functions. In addition, those skilled in the art should understand that the above computing device 800 can also only include devices necessary for the embodiments of the application, and does not have to include all the devices shown in FIG. 8.
[0160] The embodiments of the application also provide a computing device cluster. The computing device cluster includes at least one computing device. The computing device can be a server. In some embodiments, the computing device can also be a desktop computer, a notebook computer, or a terminal device such as a smart phone.
[0161] As shown in FIG. 9, the computing device cluster includes at least one computing device 800. The memory 820 in one or more computing devices 800 in the computing device cluster can store the same instructions for executing the above method.
[0162] In some possible implementation manners, the memory 820 in one or more computing devices 800 in the computing device cluster can also respectively store partial instructions for executing the above method. In other words, the combination of one or more computing devices 800 can collectively execute the instructions of the above method.
[0163] It should be noted that the memories 820 in different computing devices 800 in the computing device cluster can store different instructions, respectively for executing partial functions of the above apparatus. That is, the instructions stored in the memories 820 in different computing devices 800 can implement the functions of one or more modules in the above apparatus.
[0164] In some possible implementation manners, one or more computing devices in the computing device cluster can be connected through a network. The network can be a wide area network or a local area network, etc. FIG. 10 shows a possible implementation manner. As shown in FIG. 10, two computing devices 800A and 800B are connected through a network. Specifically, the communication interfaces in the respective computing devices are connected to the network.
[0165] It should be understood that the functions of the computing device 800A shown in FIG. 10 can also be completed by multiple computing devices 800. Similarly, the functions of the computing device 800B can also be completed by multiple computing devices 800.
[0166] In this embodiment, a computer program product containing instructions is also provided. The computer program product can be a software or program product containing instructions, which can be run on a computing device cluster or stored in any available medium. When it is run by the computing device cluster, it causes the computing device cluster to perform the method provided above, or causes the computing device cluster to realize the functions of the apparatus provided above.
[0167] In this embodiment, a computer readable storage medium is also provided. The computer readable storage medium can be any available medium that the computing device can store or a data storage device such as a data center containing one or more available media. The available medium can be a magnetic medium (for example, a floppy disk, a hard disk, a magnetic tape), an optical medium (for example, a digital video disc (DVD)), or a semiconductor medium (for example, a solid state disk), etc. The computer readable storage medium contains instructions, which, when executed by the computing device cluster, cause the computing device cluster to perform the method provided above.
[0168] Those skilled in the art can understand that the units and algorithm steps of the examples described in combination with the embodiments disclosed herein can be realized by electronic hardware or a combination of computer software and electronic hardware. Whether the functions are realized in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to realize the described functions for each specific application, but such implementation should not be considered beyond the scope of the present application.
[0169] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the above-described system, apparatus and unit can refer to the corresponding processes in the foregoing method embodiments, which will not be described here.
[0170] In several embodiments provided in the present application, it should be understood that the disclosed system, apparatus and method can be implemented in other ways. For example, the apparatus embodiments described above are only schematic. The division of the units is only a logical function division. In actual implementation, another division mode can be used, for example, a plurality of units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between the units or components shown or discussed can be indirect coupling or communication connection through some interface, device or unit, which can be electrical, mechanical or other form.
[0171] The units described as separate components may or may not be physically separate, and the components displayed as units may or may not be physical units, that is, may be located in one place, or may be distributed on multiple network units. Part or all of the units can be selected according to actual needs to achieve the purpose of the embodiment scheme.
[0172] In addition, the functional units in each embodiment of the present application can be integrated in one processing unit, or each unit can be physically present separately, or two or more units can be integrated in one unit.
[0173] The functions, if realized in the form of software functional units and sold or used as independent products, can be stored in a computer readable storage medium. Based on such understanding, the technical solutions of the present application or the part of the present application that essentially contributes to the prior art or the part of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes a plurality of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, and various program code storage media.
[0174] The above is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art can easily think of changes or replacements within the technical scope disclosed in the present application, which should be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. A failure handling method characterized by, The method is applied to a cloud management platform for managing an infrastructure providing cloud services, the infrastructure comprising at least one computing node for running M services, M being a positive integer, the method comprising: in the case where at least one service of the M services fails and there are multiple alarm information, providing a first alarm information for a tenant, the first alarm information being used for indicating that the at least one service fails, the first alarm information being generated by summarizing alarm information triggered according to a same root cause failure in the multiple alarm information according to a first operation and maintenance model, the first operation and maintenance model being obtained by training according to a historical failure database and deployment information of the M services, the multiple alarm information being generated according to a running information set of the M services, the running information set comprising at least one running information; receiving first input information of the tenant, the first input information being used for requesting to obtain a reason for the root cause failure corresponding to the first alarm information; in response to the first input information, providing first output information for the tenant, the first output information being used for indicating the reason for the root cause failure corresponding to the first alarm information, the first output information being generated according to the first operation and maintenance model.
2. The method of claim 1, wherein, The historical failure database comprises at least one of the following: at least one historical operation and maintenance case, a failure mode library or a diagnosis tree; The historical operation and maintenance case comprises at least one of the following: historical failure alarm information, a location of a failure corresponding to the historical failure alarm information, a reason for the failure corresponding to the historical failure alarm information, a processing measure for the failure corresponding to the historical failure alarm information; The failure mode library comprises at least one of the following: identification information of a failure, a location of the failure, a reason for the failure, a processing measure for the failure, a level of the failure, an association relationship between failures; The diagnosis tree comprises a first diagnosis tree and / or a second diagnosis tree, the first diagnosis tree being used for determining a root cause of a failure according to failure alarm information, and the second diagnosis tree being used for determining a processing measure for the failure according to the root cause of the failure.
3. The method according to claim 1 or 2, characterized in that, The deployment information of the M services comprises at least one of the following: information of a computing node cluster in which a computing node running the M services is located, an association relationship between the M services, a service topology of the M services.
4. The method according to any one of claims 1 to 3, characterized in that, The running information set comprises at least one of the following: a data index information set, a call link information set, a log information set; The data index information set comprises at least one data index in a running process of the M services, the call link information set comprises at least one request in a processing process of the M services, and the log information set comprises at least one log information in the running process of the M services.
5. The method according to any one of claims 1 to 4, characterized in that, The method further comprises: determine a first exception information set according to the operation information set of the M services and a second operation and maintenance model, the second operation and maintenance model being obtained by training according to a historical operation information set of the M services, the first exception information set including at least one exception information, each exception information in the at least one exception information being operation information in the operation information set that meets an exception feature; determine the plurality of alarm information according to the first exception information set.
6. The method of claim 5, wherein, The second operation and maintenance model includes at least one of a first sub-model, a second sub-model, or a third sub-model, and the method further includes: train a first initial model according to a historical data index information set of the M services to obtain a first sub-model, the first sub-model being used to determine a first data index information set from the data index information set of the M services, the first data index information set including at least one abnormal first data index, and the first data index information set belonging to the first exception information set; and / or train a second initial model according to a historical call link information set of the M services to obtain a second sub-model, the second sub-model being used to determine a first call link information set from the call link information set of the M services, the first call link information set including at least one abnormal call link information, and the first call link information set belonging to the first exception information set; and / or train a third initial model according to a historical log information set of the M services to obtain a third sub-model, the third sub-model being used to determine a first log information set from the log information set of the M services, the first log information set including at least one abnormal log information, and the first log information set belonging to the first exception information set; The first initial model, the second initial model, and the third initial model are a base model FM.
7. The method according to any one of claims 1 to 6, characterized in that, The method further includes: train an initial large model according to the historical fault database and deployment information of the M services to obtain the first operation and maintenance model, the initial large model being a large language model LLM.
8. The method according to any one of claims 1 to 7, characterized in that, In response to the first input information, provide a first output information for the tenant, including: The first operation and maintenance model searches in the historical fault database and / or the operation information set of the M services according to the first input information to determine a first response information set, the first response information set including at least one response information related to the first input information; The first operation and maintenance model generates the first output information according to the first response information set.
9. The method according to any one of claims 1 to 8, characterized in that, The method further includes: In the case that the first alarm information or a root cause fault corresponding to the first alarm information is not included in the historical fault database, at least one of the following is added to the historical fault database: the first alarm information, a cause of the root cause fault corresponding to the first alarm information, a location of the root cause fault corresponding to the first alarm information, and a processing measure of the root cause fault corresponding to the first alarm information.
10. The method according to any one of claims 1 to 9, characterized in that, The first output information further includes a location of a root cause corresponding to the first alarm information and / or a processing measure of the root cause corresponding to the first alarm information.
11. A computing device, comprising: The device is applied to a cloud management platform, the cloud management platform is used for managing infrastructure providing cloud services, the infrastructure includes at least one computing node, the at least one computing node is used for running M services, M is a positive integer, and the device includes: The sending module is configured to provide first alarm information for a tenant in a case where at least one service of the M services fails and a plurality of alarm information exists, the first alarm information is used for indicating that the at least one service fails, the first alarm information is generated by summarizing alarm information triggered according to a same root cause in the plurality of alarm information according to a first operation and maintenance model, the first operation and maintenance model is obtained by training according to a historical fault database and deployment information of the M services, and the plurality of alarm information is generated according to a running information set of the M services, the running information set includes at least one running information. The receiving module is configured to receive first input information of the tenant, the first input information is used for requesting to obtain a cause of the root cause corresponding to the first alarm information. The sending module is further configured to provide first output information for the tenant in response to the first input information, the first output information is used for indicating the cause of the root cause corresponding to the first alarm information, and the first output information is generated according to the first operation and maintenance model.
12. The apparatus of claim 11, wherein, The historical fault database includes at least one of the following: at least one historical operation and maintenance case, a fault mode library or a diagnosis tree; The historical operation and maintenance case includes at least one of the following: historical fault alarm information, a location of a fault corresponding to the historical fault alarm information, a cause of the fault corresponding to the historical fault alarm information, and a processing measure of the fault corresponding to the historical fault alarm information; The fault mode library includes at least one of the following: identification information of a fault, a location of the fault, a cause of the fault, a processing measure of the fault, a level of the fault, and an association relationship between faults; The diagnosis tree includes a first diagnosis tree and / or a second diagnosis tree, the first diagnosis tree is used for determining a root cause of a fault according to fault alarm information, and the second diagnosis tree is used for determining a processing measure of the fault according to the root cause of the fault.
13. The apparatus of claim 11 or 12, wherein, The deployment information of the M services includes at least one of the following: information of a computing node cluster in which a computing node running the M services is located, an association relationship between the M services, and a service topology of the M services.
14. The apparatus of any one of claims 11 to 13, wherein, The running information set includes at least one of the following: a data index information set, a call link information set, and a log information set; The data index information set includes at least one data index in a running process of the M services, the call link information set includes call link information of at least one request in a processing process of the M services in the running process, and the log information set includes at least one log information in the running process of the M services.
15. The apparatus of any one of claims 11 to 14, wherein, The apparatus further includes a processing module configured to: determine a first set of abnormal information according to the set of running information of the M services and a second operation and maintenance model, the second operation and maintenance model being obtained by training a set of historical running information of the M services, each of the at least one abnormal information in the first set of abnormal information being running information in the set of running information that meets an abnormal characteristic; determine the plurality of alarm information according to the first set of abnormal information.
16. The apparatus of claim 15, wherein, The second operation and maintenance model includes at least one of a first sub-model, a second sub-model, or a third sub-model, and the processing module is further configured to: train a first initial model according to a set of historical data indicator information of the M services to obtain a first sub-model, the first sub-model being configured to determine a first set of data indicator information from the set of data indicator information of the M services, the first set of data indicator information including at least one abnormal first data indicator, the first set of data indicator information belonging to the first set of abnormal information; and / or, train a second initial model according to a set of historical call link information of the M services to obtain a second sub-model, the second sub-model being configured to determine a first set of call link information from the set of call link information of the M services, the first set of call link information including at least one abnormal call link information, the first set of call link information belonging to the first set of abnormal information; and / or, train a third initial model according to a set of historical log information of the M services to obtain a third sub-model, the third sub-model being configured to determine a first set of log information from the set of log information of the M services, the first set of log information including at least one abnormal log information, the first set of log information belonging to the first set of abnormal information; wherein the first initial model, the second initial model, and the third initial model are a base model FM.
17. The apparatus of any one of claims 11 to 16, wherein, The apparatus further includes a processing module configured to: train an initial large model according to the historical fault database and the deployment information of the M services to obtain the first operation and maintenance model, the initial large model being a large language model LLM.
18. The apparatus of any one of claims 11 to 17, wherein, The apparatus further includes a processing module configured to run the first operation and maintenance model, The first operation and maintenance model retrieves in the historical fault database and / or the set of running information of the M services according to the first input information to determine a first set of response information, the first set of response information including at least one response information related to the first input information; The first operation and maintenance model generates the first output information according to the first set of response information.
19. The apparatus of any of claims 11 to 18, wherein, The apparatus further includes a processing module configured to: In a case where the first alarm information is not included in the historical fault database or a root cause fault corresponding to the first alarm information, at least one of the first alarm information, a cause of the root cause fault corresponding to the first alarm information, a location of the root cause fault corresponding to the first alarm information, and a treatment measure of the root cause fault corresponding to the first alarm information is added to the historical fault database.
20. The apparatus of any one of claims 11-19, wherein, The first output information further includes a location of a root cause fault corresponding to the first alarm information and / or a treatment measure of the root cause fault corresponding to the first alarm information.
21. A cluster of computing devices, characterized in that, The at least one computing device includes a processor and a memory; The processor of the at least one computing device is configured to execute instructions stored in the memory of the at least one computing device to cause the cluster of computing devices to perform the method of any one of claims 1-10.
22. A computer program product comprising instructions, wherein: The instructions, when executed by the cluster of computing devices, cause the cluster of computing devices to perform the method of any one of claims 1-10.
23. A computer-readable storage medium, characterized in that, The computer program instructions, when executed by the cluster of computing devices, cause the cluster of computing devices to perform the method of any one of claims 1-10.
Citation Information
Patent Citations
Fault processing method and device based on network alarm association
CN110247792A
Message obtaining method and device
CN110888754A
Method and device for determining root cause fault
CN114520994A
Operation and maintenance fault identification method and device
CN117931589A
Diagnosing a fault incident in a data center
US20110191630A1