Abnormality diagnosis method and device, computer program product and storage medium

By automatically collecting and visualizing full-link resource information in the anomaly diagnosis platform, users can analyze the root cause and submit repair plans, which solves the accuracy problem of unknown anomaly diagnosis and improves the diagnostic capabilities and efficiency of the diagnosis platform.

CN120729698APending Publication Date: 2025-09-30HANGZHOU ALICLOUD FEITIAN INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410362561.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-03-27
Publication Date
2025-09-30

AI Technical Summary

Technical Problem

The existing anomaly diagnosis platform is unable to effectively diagnose unknown anomalies, resulting in complicated offline work order flows and poor repair results.

Method used

After receiving an unknown exception request in the exception diagnosis platform, it automatically collects the full-link resource information of the target cloud computing product, provides a visual interface for users to analyze the root cause, and obtains the root cause and repair plan submitted by the user, enriching the platform knowledge to improve the diagnostic capability.

Benefits of technology

It achieves accurate diagnosis of unknown anomalies, reduces offline work order flow, saves time costs, and improves the diagnostic accuracy and efficiency of the diagnostic platform.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120729698A_ABST
    Figure CN120729698A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides an abnormality diagnosis method and device, a computer program product and a storage medium. According to the embodiment of the invention, an anomaly diagnosis request for any resource in the cloud computing product can be received, if a root cause causing an anomaly cannot be diagnosed, information collection can be automatically carried out according to a preset information dimension by taking each resource contained in the cloud computing product as a collection object, and the collected information is displayed in a service interface oriented to a product provider. A product provider can submit the analyzed root cause through a service interface of the product provider as the root cause causing the resource abnormality. Therefore, unknown problems can be diagnosed through the abnormity diagnosis platform, the diagnosis accuracy can be effectively guaranteed, the product user and the product provider do not need to flow work orders offline any more, and the time cost of the two parties is effectively saved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of cloud computing technology, and in particular to an abnormality diagnosis method, device, computer program product, and storage medium. Background Art

[0002] The anomaly diagnosis platform can be used to diagnose anomalies in cloud computing products. Currently, the anomaly diagnosis platform does not have the ability to diagnose unknown anomalies.

[0003] For such anomalies, product users need to manually submit work orders and data such as logs on the resources where the anomalies occurred in the cloud computing product to the product provider. The product provider will then conduct a single-point analysis of these data on the resources to find the root cause of the anomaly and a repair plan. However, in many cases, the repair plan provided still cannot resolve the anomaly, causing the same anomaly to recur.

[0004] This offline solution path, which is separated from the abnormal diagnosis platform, is too complicated and the repair effect is still poor. Therefore, there is an urgent need to improve the diagnostic capabilities of the abnormal diagnosis platform. Summary of the Invention

[0005] Various aspects of the present application provide an anomaly diagnosis method, device, computer program product, and storage medium to enhance anomaly diagnosis capabilities.

[0006] The present invention provides an abnormality diagnosis method, including:

[0007] In response to receiving an anomaly diagnosis request for a target resource in a target cloud computing product, if the root cause of the anomaly is not diagnosed, collecting information based on preset information dimensions with each resource included in the target cloud computing product as a collection object;

[0008] Displaying the collected information on a first service interface for the first type of users, so that the first type of users can analyze the root cause of the anomaly based on the information;

[0009] The root cause submitted by the first type of user through the first service interface is obtained as the root cause of the target resource exception.

[0010] Furthermore, the collected information is displayed on the first service interface for the first type of users, including:

[0011] Convert the collected information into visual materials;

[0012] The visualization material is displayed on the first service interface to visualize the collected information.

[0013] Furthermore, the preset information dimension includes descriptive information corresponding to abnormal problems that have been diagnosed on each resource included in the cloud computing product, descriptive information corresponding to operation events that occurred on the cloud computing product within a preset diagnostic period, and / or full-link tracking information corresponding to the operation events.

[0014] Furthermore, the method further comprises:

[0015] Obtaining a repair plan formulated for the analyzed root cause, submitted by the first category of users through the first service interface;

[0016] Executing a repair operation according to the repair plan to complete the problem repair of the target resource;

[0017] The repair operation involves one or more resources in the target cloud computing product.

[0018] Furthermore, the method further comprises:

[0019] Obtaining a new abnormal problem defined by the abnormal diagnosis request submitted by the first category of users through the first service interface, and a causal relationship between the monitoring item and the new abnormal problem analyzed based on the information;

[0020] The causal relationship provided by the first category of users is supplemented as a new basis for anomaly detection.

[0021] Furthermore, the method further comprises:

[0022] Obtaining a root cause analysis plan for the new abnormal problem submitted by the first category of users through the first service interface, and supplementing it as a new basis for root cause analysis; and / or,

[0023] Obtain the correlation between the repair solution submitted by the first category of users through the first service interface and the root cause, and supplement it as a new basis for problem repair.

[0024] Furthermore, the method further comprises:

[0025] After completing the supplemental operation, receiving a subsequent abnormality diagnosis request for any resource in the target computing product;

[0026] If it is detected in the anomaly detection step that the monitoring items monitored from the cloud computing product meet the supplemented causal relationship, it is determined that the new anomaly problem has been detected;

[0027] In the root cause analysis phase, a root cause analysis is performed based on the root cause analysis plan formulated for the new abnormal problem to determine the root cause;

[0028] In the problem repair phase, a repair operation is performed according to the repair solution associated with the determined root cause to complete the problem repair for the resource.

[0029] Furthermore, the method further comprises:

[0030] If the target cloud computing product has enabled active diagnosis mode, then without receiving any abnormality diagnosis request for the target cloud computing product, anomaly detection will be actively performed based on the monitoring items obtained from the various resources included in the target cloud computing product, so as to complete the problem repair before the cloud computing product fails.

[0031] Furthermore, the method further comprises:

[0032] If a response result query instruction for the abnormal diagnosis request is received from a second type of user through a second service interface, the root cause and the repair solution are displayed in the second service interface.

[0033] Furthermore, the abnormality diagnosis request is submitted by the first category user through the first service interface, or the abnormality diagnosis request is submitted by the second category user through the second service interface.

[0034] Furthermore, before performing the repair operation, the method further includes:

[0035] If it is detected that the repair plan includes a repair operation that affects resource usage, a query message is displayed in the second service interface for the second type of user to confirm whether to allow the repair;

[0036] If a repair permission instruction is received from the second type of user through the second service interface, the repair operation is performed.

[0037] Furthermore, the method further comprises:

[0038] After receiving the abnormality diagnosis request, if after performing abnormality detection based on the causal relationship between the known monitoring items and the abnormal problem, no detection result can be obtained or the detection result does not point to the known abnormal problem, it is determined that the root cause of the abnormality has not been diagnosed.

[0039] An embodiment of the present application further provides a computing device, including a memory, a processor, and a communication component;

[0040] The memory is used to store one or more computer instructions;

[0041] The processor is coupled to the memory and the communication component, and is configured to execute the one or more computer instructions to perform the aforementioned abnormality diagnosis method.

[0042] An embodiment of the present application further provides a computer-readable storage medium storing a computer program. When the computer program is executed by one or more processors, the one or more processors are caused to execute the aforementioned abnormality diagnosis method.

[0043] An embodiment of the present application further provides a computer program product, comprising a computer program. When the computer program is executed by one or more processors, the one or more processors are caused to execute the aforementioned abnormality diagnosis method.

[0044] In an embodiment of the present application, after receiving an exception diagnosis request for any resource in a cloud computing product, if the root cause of the exception cannot be diagnosed, information can be automatically collected according to the preset information dimension, with each resource contained in the cloud computing product as the collection object, and the collected information can be displayed to the service interface facing the product provider. In this way, the observation data of the entire link resources in the cloud computing product can be displayed to the product provider, and is no longer limited to the single point resource pointed to by the exception diagnosis request, which can enable the product provider to obtain a more comprehensive analysis perspective, thereby more accurately analyzing the root cause of the exception. The product provider can submit the analyzed root cause through its service interface as the root cause of the resource exception. Accordingly, in this embodiment, the diagnostic capability of the exception diagnosis platform is improved, and unknown problems can be diagnosed through the exception diagnosis platform, and the accuracy of the diagnosis can be effectively guaranteed. Product users and product providers no longer need to transfer work orders offline, effectively saving the time cost of both parties. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:

[0046] Figure 1 A flowchart of an abnormality diagnosis method provided by an exemplary embodiment of the present application;

[0047] Figure 2 An optional logic diagram of an abnormality diagnosis method provided by an exemplary embodiment of the present application;

[0048] Figure 3 A schematic structural diagram of an abnormality diagnosis platform provided by an exemplary embodiment of the present application;

[0049] Figure 4 A schematic diagram of interaction logic in an abnormality diagnosis method provided by an exemplary embodiment of the present application;

[0050] Figure 5 A schematic structural diagram of a computing device provided as another exemplary embodiment of the present application. DETAILED DESCRIPTION

[0051] To make the purpose, technical solutions, and advantages of this application more clear, the technical solutions of this application will be clearly and completely described below in conjunction with the specific embodiments of this application and the corresponding drawings. Obviously, the embodiments described are only part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0052] Before describing in detail the technical solutions provided by the embodiments of the present application, several technical concepts involved in the embodiments of the present application are briefly explained as follows.

[0053] A cloud computing product can be understood as a cloud computing service provided by a cloud computing provider, which is used to achieve the slimming down of terminal devices. The program products that originally ran on the terminal devices and mainly relied on the computing power of the terminal devices are now implemented by cloud computing products configured on the server side. The cloud computing products involved in the embodiments of the present application may include cloud desktops or cloud applications, etc. Cloud computing products usually contain multiple resources, and multiple resources communicate and cooperate with each other to realize the product functions of cloud computing products. Here, the resources contained in the cloud computing product can be cloud server instances, virtual machines, gateways, product clients on user terminals, etc. No further examples are given here.

[0054] Cloud desktop, also known as desktop virtualization or cloud computer, is a new model that replaces traditional computers. After adopting cloud desktop, users no longer need to purchase a computer host. The CPU, memory, hard disk and other components contained in the host are all virtualized in the back-end server. A thin client (also known as the host side) can be used on the terminal device to connect to the monitor, keyboard and / or mouse. After the user installs the thin client, he accesses the cloud desktop on the server through a unique communication protocol to realize interactive operations, achieving the same experience as a computer. At the same time, cloud desktop not only supports the replacement of traditional computers, but also supports access on the Internet from other smart devices such as mobile phones and tablets. It is the latest solution for mobile office.

[0055] Cloud applications are the application layer embodiment of cloud computing technology. Compared to locally installed applications, cloud applications offer advantages such as no download or installation required, immediate use, minimal device requirements, and cross-platform functionality. Cloud-based applications are becoming a future trend. Cloud applications stream multimedia content to end devices in real time, enabling smooth user experience by simply displaying the multimedia content. Interactions between users and the end devices are also handled by the cloud application.

[0056] Anomaly diagnosis can be understood as a solution that integrates anomaly detection, root cause analysis, and problem remediation technologies. After an anomaly is discovered, it locates the underlying anomaly, analyzes the root cause, and provides a remediation solution.

[0057] The anomaly diagnosis platform can be understood as an anomaly diagnosis service provided by the cloud computing provider, which is used to implement anomaly diagnosis of cloud computing products.

[0058] During the research process, the inventors discovered that various anomalies may occur in the operation of various resources in cloud computing products. In traditional solutions, the anomaly diagnosis platform does not have the ability to diagnose unknown anomalies. When the anomaly diagnosis platform is unable to complete the diagnosis, the product user needs to submit a work order to the product provider for the resource with an anomaly in the cloud computing product. The product user will usually also provide some logs and other data on the resource to the product provider as a reference. The product provider needs to check this data to analyze the root cause of the anomaly of the resource. This offline work order flow solution requires a lot of time costs for both the product user and the product provider. Moreover, due to the limited data that the product user can provide, the root cause analyzed by the product provider for the resource may not be accurate, which may cause the resource to have repeated anomalies.

[0059] To this end, this embodiment proposes an anomaly diagnosis method to enhance the diagnostic capabilities of the anomaly diagnosis platform, enabling it to diagnose unknown anomalies and ensure the accuracy of the diagnosis, thereby avoiding offline work order transfers between product users and product providers for unknown anomalies, saving time costs for both parties.

[0060] The following describes in detail the technical solutions provided by various embodiments of the present application in conjunction with the accompanying drawings.

[0061] Figure 1 This is a flow chart of an abnormality diagnosis method provided by an exemplary embodiment of the present application. The method can be executed by an abnormality diagnosis platform. The abnormality diagnosis platform can be implemented as software, hardware, or a combination of software and hardware. The abnormality diagnosis platform can be integrated into a computing device. The computing device can be implemented as an independent server or a server cluster, etc. This embodiment does not limit this. Figure 1 , the method may include:

[0062] Step 100: In response to receiving an anomaly diagnosis request for a target resource in a target cloud computing product, if the root cause of the anomaly is not diagnosed, information is collected based on preset information dimensions, with each resource included in the target cloud computing product as a collection object;

[0063] Step 101: Display the collected information on a first service interface for a first type of user, so that the first type of user can analyze the root cause of the anomaly based on the information;

[0064] Step 102: Obtain the root cause submitted by the first type of user through the first service interface as the root cause of the target resource exception.

[0065] In this embodiment, the target cloud computing product can be any cloud computing product provided by the cloud computing provider. It should be understood that the anomaly diagnosis platform in this embodiment can be used to perform anomaly diagnosis on a variety of cloud computing products. In addition, in this embodiment, product providers can be further subdivided for different cloud computing products. In actual applications, a product provider can provide one or more cloud computing products, and the product provider can be understood as product R&D personnel, etc. The product user can be understood as various types of personnel who need to use cloud computing products, including but not limited to customer administrators, product operation and maintenance personnel, on-duty personnel, and product testers, etc. No further examples are given here. For the sake of ease of description, in the following text, the product provider is described as the first type of user and the product user is described as the second type of user.

[0066] As mentioned above, cloud computing products can contain multiple resources. In practical applications, an anomaly diagnosis request typically explicitly targets a specific resource within the cloud computing product, specifically the resource where the anomaly occurred. Therefore, it is understood that the target resource mentioned in this embodiment can refer to any resource within the target cloud computing product where the anomaly occurred.

[0067] refer to Figure 1 In step 100, a request for abnormality diagnosis of a target resource in a target cloud computing product may be received. In this embodiment, there is no restriction on the initiator of the abnormality diagnosis request. Both the first and second category users in this embodiment may initiate an abnormality diagnosis request if they deem it necessary.

[0068] After the abnormality diagnosis request reaches the abnormality diagnosis platform in this embodiment, the abnormality diagnosis platform can perform abnormality diagnosis according to existing knowledge. Figure 2 An optional logic diagram of an abnormality diagnosis method provided by an exemplary embodiment of the present application. Figure 2 , there may be two situations:

[0069] One scenario is that the anomaly diagnosis platform can diagnose the root cause of the anomaly based on existing knowledge. In this case, the root cause diagnosed by the anomaly diagnosis platform can be used as the root cause of the target resource anomaly.

[0070] Another situation is that, as described in step 100, the anomaly diagnosis platform cannot diagnose the root cause of the anomaly based on existing knowledge. In this case, it means that an unknown anomaly may have occurred, and the anomaly diagnosis platform itself does not yet have the ability to diagnose the anomaly. According to the traditional solution mentioned above, in this case, the anomaly diagnosis platform will usually issue a notification that cannot be targeted, thereby triggering the offline work order flow. However, in this embodiment, in this case, the anomaly diagnosis platform will not issue such a notification, but will automatically collect information according to the preset information dimensions, as described in step 100, with the various resources included in the target cloud computing product as the collection objects.

[0071] The preset information dimensions proposed in this embodiment may include, but are not limited to, descriptive information corresponding to abnormal problems that have been diagnosed on each resource included in the cloud computing product, descriptive information corresponding to operation events that occurred on the cloud computing product within a preset diagnostic period, and full-link tracking information corresponding to the operation events.

[0072] Among them, the anomaly diagnosis platform may have already performed anomaly diagnosis on each resource included in the cloud computing product and identified anomaly problems. Based on this, this embodiment proposes that the descriptive information corresponding to the diagnosed anomaly problems on each resource included in the cloud computing product can be used as a preset information dimension. This descriptive information may include, but is not limited to, the resource name, the name of the anomaly problem, the time the anomaly occurred, the time the anomaly ended, the number of anomaly occurrences, the root cause of the anomaly, and the repair solution adopted, etc. Further examples are not provided here. The information collected under this preset information dimension can provide the first type of users with clues about the correlation between anomaly problems and other aspects.

[0073] Among them, there are usually many event monitoring services in the cloud computing scenario, which can monitor the operation events occurring on the cloud computing products. The operation events here can refer to the operations performed by the user, such as shutting down the virtual machine, etc., or they can refer to the operations automatically performed by the resource, for example, the resource itself performs configuration update operations, etc. These event monitoring services can provide descriptive information corresponding to the operation events. In this embodiment, the descriptive information corresponding to the operation events occurring on the cloud computing product may include but is not limited to the name of the operation event, the time when the event occurred, etc., and no further examples are given here. The information collected under this preset information dimension can provide the first type of users with clues about the correlation between the operation event and the time when the anomaly occurred. For example, if the target resource becomes abnormal after a certain operation event occurs, the operation event can be positioned as an event that may cause the anomaly, thereby continuously narrowing the scope of analysis.

[0074] Among them, the full-link tracking platform in the cloud computing scenario can provide a full-link tracking service for operation events, so as to construct full-link tracking information (also called trace information) for each operation event. The full-link tracking information can describe all the call links of all resources involved in an operation event. Therefore, the information collected based on the preset information dimension can be used by the first type of users to focus on and analyze the overall call link of a certain operation event that may cause an anomaly, so as to locate the resources related to the root cause. Then, the located resources are screened for monitoring items that have a causal relationship with the anomaly.

[0075] It's worth emphasizing here that, in this embodiment, the anomaly diagnosis platform collects information based on the resources contained in the target cloud computing product. That is, the anomaly diagnosis platform collects information from each resource contained in the target cloud computing product, according to preset information dimensions. This allows for observation data on all resources in the target cloud computing product, rather than being limited to the target resources. It can be seen that the information collected by the anomaly diagnosis platform in this embodiment provides the first category of users with hierarchical, multi-perspective, and full-link observation data. Based on this information, the first category of users can gradually narrow the scope of their analysis, discover useful clues, and more efficiently and accurately analyze the root cause of the anomaly.

[0076] Continue to refer Figure 1 In step 101, the anomaly diagnosis platform may display the collected information on a first service interface for a first category of users, allowing the first category of users to analyze the root cause of the anomaly based on the information. Here, the first category of users refers to R&D personnel of the product provider, etc.

[0077] Figure 3 This is a schematic diagram of the structure of an abnormality diagnosis platform provided by an exemplary embodiment of the present application. Figure 3 The abnormality diagnosis platform in this embodiment can provide two service interfaces: a first service interface for the first type of users, and a second service interface for the second type of users. As mentioned above, the first type of users may refer to product providers, such as product developers, etc. Accordingly, the first service interface can be understood as an interface from the perspective of the product provider, and the aforementioned product developers, etc. can use the first service interface for human-computer interaction. The second type of users may refer to product users, and accordingly, the second service interface can be understood as an interface from the perspective of the product user, and the aforementioned customer administrators, product operation and maintenance personnel, on-duty personnel, and product testers, etc. can use the second service interface for human-computer interaction.

[0078] As mentioned above, both the first and second category users in this embodiment can initiate an abnormality diagnosis request if they deem it necessary. Figure 3, the first and second types of users can initiate abnormality diagnosis requests through the service interfaces they use. In actual applications, both the first and second service interfaces can provide interface controls for initiating abnormality diagnosis requests to support initiating abnormality diagnosis requests, which will not be explained in detail here.

[0079] Based on this, in step 101, the abnormality diagnosis platform can display the collected information on the first service interface. The first type of user can observe the information collected by the abnormality diagnosis platform through the first service interface. The first type of user can analyze this information to determine the root cause of the abnormality.

[0080] In a preferred implementation, the abnormality diagnosis platform can convert the collected information into visualization materials, and display the visualization materials in the first service interface to visualize the collected information. The visualization materials may include but are not limited to line graphs, bar graphs, pie charts, and scatter plots, and connections can be established between the visualization materials to display the relevant information in a correlated manner. In this way, the abnormality diagnosis platform can convert information that is originally difficult to read into visualization materials, and then display the converted visualization materials in the first service interface to achieve a visual display of the information instead of displaying the original appearance of the information.

[0081] An exemplary visualization effect can be: using a scatter plot with time as the horizontal axis and operation events as the vertical axis to display the distribution of operation events for the target cloud computing product within a preset diagnostic period. Furthermore, corresponding descriptive information can be associated with the scatter plots corresponding to the operation events. Furthermore, clicking on the scatter plot corresponding to an operation event can display the full-link tracking information corresponding to the operation event. It should be understood that the visualization effect in this embodiment is not limited to this, and further examples are not provided here.

[0082] In addition, this embodiment does not limit the analysis method used by the product provider when performing manual analysis. Figure 4 This is a schematic diagram of the interaction logic in an abnormality diagnosis method provided by an exemplary embodiment of the present application. Figure 4 In actual applications, product providers can use various available tools such as log analysis tools, full-link tracking platforms, cloud assistants, virtual network consoles, and AI analysis tools for auxiliary analysis to more efficiently and accurately analyze the root causes of anomalies.

[0083] The key here is that the information collected by the anomaly diagnosis platform involves the full-link resources contained in the target cloud computing product. Therefore, the product provider can obtain observation data of the full-link resources contained in the target cloud computing product. This can provide the product provider with a comprehensive and diverse analysis perspective, thereby helping the product provider to discover clues more efficiently and more accurately determine the root cause of the anomaly.

[0084] During the research process, the inventor discovered that the root cause of an abnormality in a resource in a cloud computing product may not be related to the resource itself, but may be related to other resources in the cloud computing product. Therefore, the method of observing the entire link resources adopted in this embodiment can effectively improve the accuracy and rationality of the analyzed root cause, thereby providing a more comprehensive and reasonable reference for subsequent problem repair links.

[0085] Continue to refer Figure 1 In step 102, the first type of user can submit the analyzed root cause through the first service interface. In this way, the abnormality diagnosis platform can use the root cause submitted by the first type of user as the root cause of the target resource abnormality.

[0086] It is understandable that after the anomaly diagnosis platform receives the anomaly diagnosis request, the aforementioned steps 100-102 are executed without the knowledge of the initiator of the anomaly diagnosis request. From the initiator's perspective, the anomaly diagnosis platform completes the response to the anomaly diagnosis request, and the initiator is unaware of whether the anomaly diagnosis platform triggers the first type of user's manual analysis process.

[0087] Therefore, in this embodiment, from the perspective of the initiator of the abnormality diagnosis request, the abnormality diagnosis platform has the ability to diagnose unknown abnormalities, and it is no longer necessary to conduct offline work order circulation to realize the diagnosis of unknown abnormalities.

[0088] Furthermore, based on the second service interface provided by the anomaly diagnosis platform, if a response query command for an anomaly diagnosis request is received from a second-category user through the second service interface, the root cause and remediation solution will be displayed on the second service interface. In practice, the second-category user is not particularly concerned with internal technical details. They are more concerned with the cause of the currently perceived anomaly and how to resolve it. Therefore, the second service interface can minimize technical details and present the root cause and remediation solution in a more understandable manner, allowing the second-category user to quickly understand the root cause of the anomaly and the adopted remediation solution.

[0089] In summary, in this embodiment, after receiving an exception diagnosis request for any resource in a cloud computing product, if the root cause of the exception cannot be diagnosed, information can be automatically collected according to the preset information dimension, with each resource contained in the cloud computing product as the collection object, and the collected information can be displayed to the service interface facing the product provider. In this way, the observation data of the entire link resources in the cloud computing product can be displayed to the product provider, and is no longer limited to the single point resource pointed to by the exception diagnosis request, which can enable the product provider to obtain a more comprehensive analysis perspective, thereby more accurately analyzing the root cause of the exception. The product provider can submit the analyzed root cause through its service interface as the root cause of the resource exception. Accordingly, in this embodiment, the diagnostic capability of the exception diagnosis platform is improved, and unknown problems can be diagnosed through the exception diagnosis platform, and the accuracy of the diagnosis can be effectively guaranteed. Product users and product providers no longer need to transfer work orders offline, effectively saving the time cost of both parties.

[0090] In the above or following embodiments, after receiving the exception diagnosis request, if the exception diagnosis platform is to diagnose the root cause of the exception, in addition to collecting information and displaying the collected information on the second service interface so that the first type of user can analyze the root cause of the exception, it can also support the first type of user to formulate a repair plan.

[0091] In this regard, the present embodiment proposes that the first type of user can formulate a repair plan based on the analyzed root cause and submit the repair plan through the second service interface. Here, the repair plan can be used to indicate the required repair operation.

[0092] Based on this, in this embodiment, the abnormality diagnosis platform can obtain the repair plan formulated for the analyzed root cause submitted by the first type of user through the first service interface; perform the repair operation according to the repair plan to complete the problem repair for the target resource.

[0093] It's worth noting that, because the first category of users in this embodiment performs analysis based on observation data from all resources in the target cloud computing product, the root cause identified may involve one or more resources in the target cloud computing product, and is not limited to the target resources. Accordingly, the remediation actions specified in the remediation plan for the analyzed root cause by the first category of users may involve one or more resources in the target cloud computing product, and are likewise not limited to the target resources.

[0094] It can be seen that based on the observation data of the full-link resources in the target cloud computing product provided by the anomaly diagnosis platform in this embodiment, the repair plan formulated by the first type of users can be more comprehensive and reasonable, and therefore can better ensure that the anomalies of the target resources are fundamentally resolved.

[0095] Furthermore, during their research, the inventors discovered that some repair operations may affect resource usage. Therefore, this implementation proposes an optimization solution: Before executing a repair operation, it is possible to detect whether the repair plan includes repair operations that affect resource usage. If so, a query message can be displayed on the second service interface for the second-category user to confirm whether to allow the repair. If a permission instruction is submitted by the second-category user through the second service interface, the repair operation is then executed. This effectively ensures that the second-category user has a good experience using the resources of the target cloud computing product.

[0096] At this point, the problem can be repaired for the target resource.

[0097] Further, refer to Figure 3 This embodiment also proposes a new knowledge entry function within the first service interface of the anomaly diagnosis platform. This function allows the root causes and repair solutions submitted by the first category of users to be added to the anomaly diagnosis platform as knowledge, thereby continuously enriching the knowledge within the anomaly diagnosis platform and enhancing its diagnostic capabilities.

[0098] During the research process, the inventors found that the process from the occurrence of an exception to its resolution can be roughly divided into several steps: anomaly discovery step, anomaly detection step, root cause analysis step, and problem repair step.

[0099] Among them, the abnormality discovery link can be understood as the link of perceiving the abnormality. On the one hand, in this embodiment, the abnormality diagnosis platform can perform tracking on each resource in the cloud computing product to observe whether there are abnormal changes in the monitoring items on each resource. If so, the tracking can be used to push an alarm to the second service interface to prompt the second type of user to perceive an abnormality on a certain resource. On the other hand, in this embodiment, the second type of user may also perceive a usage failure in the process of using a certain resource in the cloud computing product. For example, the product client is stuck, etc. This situation can also prompt the second type of user to perceive an abnormality on a certain resource. On the other hand, in this embodiment, the first type of user may find that some resources in the cloud computing product have abnormal operation during the product debugging before the cloud computing product is released. This situation can prompt the first type of user to perceive an abnormality on a certain resource.

[0100] To this end, in this embodiment, no matter which of the above aspects or types of users, when an abnormality is perceived, as described above, an abnormality diagnosis request can be initiated to the abnormality diagnosis platform when it is deemed necessary.

[0101] However, during the anomaly detection phase, while only notifying a resource of an anomaly, the specific anomaly requires further investigation. To this end, we can define anomaly problems to describe similar anomalies. Anomalies not categorized as known anomaly problems are considered unknown anomalies.

[0102] In the anomaly detection link, anomaly detection can be used to determine the abnormal problems that have occurred. It can be seen that if an unknown anomaly is perceived in the anomaly discovery link, the detection results cannot be obtained in the anomaly detection link or the obtained detection results do not point to known abnormal problems. The inventors found in the research process that the occurrence of abnormal problems is usually foreshadowed by some monitoring items. Therefore, for known abnormal problems, monitoring items that foreshadow them can be observed, so that when a specified abnormal change is observed in such monitoring items (for example, the value on a certain monitoring item exceeds a preset threshold), it is determined that the abnormal problem is about to occur. To this end, it is proposed in this embodiment that a causal relationship between monitoring items and abnormal problems can be constructed and entered into the abnormal diagnosis platform as knowledge to support the abnormal detection capability of the anomaly detection platform.

[0103] To this end, this embodiment proposes that: the abnormal diagnosis platform can obtain the new abnormal problems defined by the aforementioned abnormal diagnosis request submitted by the first type of users through the first service interface, as well as the causal relationship between the monitoring items and the new abnormal problems analyzed based on the information; and supplement the causal relationship provided by the first type of users as a new basis for abnormal detection.

[0104] During the research process, the inventors discovered that the abnormality diagnosis platform is unable to diagnose the root cause of the abnormality. Usually, as mentioned above, in the abnormality detection link, after performing abnormality detection based on the causal relationship between known monitoring items and abnormal problems, no detection results can be obtained or the detection results do not point to known abnormal problems. Therefore, the first type of users can define new abnormal problems so that unknown abnormalities can be classified. Moreover, the first type of users can also determine the monitoring items that foreshadow the new abnormal problem based on the information displayed in the second service interface, thereby constructing a causal relationship between these monitoring items and the new abnormal problem. This causal relationship can describe what changes in these monitoring items will cause the new abnormal problem. On this basis, the first type of users can submit the constructed causal relationship through the first service interface, and the abnormality diagnosis platform can supplement the causal relationship as a new basis for abnormal diagnosis.

[0105] Optionally, in this embodiment, a machine learning model can be used in the anomaly diagnosis platform to be responsible for anomaly detection. Based on this, the causal relationship between the constructed monitoring items and the new abnormal problem can be input into the machine learning model so that the machine learning model can learn the causal relationship, thereby having the ability to detect the new abnormal problem in the subsequent reasoning process.

[0106] In this way, in this embodiment, the causal relationship between monitoring items and abnormal problems in the anomaly diagnosis platform can be continuously enriched, thereby continuously enriching the abnormal problems supported by the anomaly diagnosis platform. Moreover, for new abnormal problems, only one manual analysis of the first type of users is required to be triggered, and the anomaly detection capability for new abnormal problems can be accumulated in the anomaly detection platform. Subsequently, the anomaly detection platform can automatically detect the new abnormal problem, eliminating the need to trigger the first type of users for manual analysis.

[0107] During root cause analysis, the root cause of an anomaly can be diagnosed. Similarly, root cause analysis can only identify root causes for known anomalies. If no detection results are obtained during anomaly detection, or if the detection results only point to a known anomaly, the root cause cannot be diagnosed.

[0108] To this end, this embodiment proposes that the anomaly diagnosis platform can obtain root cause analysis plans for new anomaly issues submitted by the first category of users through the first service interface, supplementing them as new evidence for root cause analysis. The root cause analysis plan can indicate the analysis logic to be executed on each monitoring item that has a causal relationship with the new anomaly issue. By analyzing each monitoring item that has a causal relationship with the new anomaly issue according to this analysis logic, the root cause can be determined. Because root cause analysis plans are typically complex and diverse, specific examples of root cause analysis plans are not provided here.

[0109] In this way, in this embodiment, the root cause analysis solutions in the abnormality diagnosis platform can be continuously enriched, thereby continuously enriching the root cause analysis capability of the abnormality diagnosis platform.

[0110] In the problem repair phase, a repair operation can be performed to repair the problem. In this embodiment, it is proposed that the abnormality diagnosis platform can obtain the correlation between the repair solution submitted by the first type of user through the first service interface and the root cause, and supplement it as a new basis for problem repair.

[0111] In this way, in this embodiment, the repair solutions in the abnormality diagnosis platform can be continuously enriched, thereby continuously enriching the problem repair capability of the abnormality diagnosis platform.

[0112] refer to Figure 3 In actual applications, the anomaly diagnosis platform can provide an anomaly diagnosis capability management module, a root cause analysis capability management module, and a problem repair capability management module. Based on these modules, it can manage various aspects of new knowledge provided by the first type of users through the new knowledge entry function in the first service interface, thereby continuously enriching the capabilities of the anomaly diagnosis platform in these aspects.

[0113] In summary, by supplementing the new abnormal problems defined by the first type of users in the manual analysis process, the constructed causal relationships, the analyzed root causes, and the developed repair plans as new knowledge for the abnormal diagnosis platform, the abnormal diagnosis platform can be equipped with the ability to diagnose the new abnormal problems.

[0114] On this basis, in this embodiment, the anomaly diagnosis platform can continue to receive subsequent anomaly diagnosis requests for any resource in the target computing product. If it is detected in the anomaly detection link that the monitoring items monitored from the cloud computing product meet the supplemented causal relationship, it is determined that a new anomaly problem is detected; in the root cause analysis link, a root cause analysis is performed based on the root cause analysis plan formulated for the new anomaly problem to determine the root cause; in the problem repair link, the repair operation is executed according to the repair plan associated with the determined root cause to complete the problem repair for the resource. Similarly, here, before executing the repair operation, it can also be detected whether the repair plan contains repair operations that affect resource usage; if it does, the query information is displayed in the second service interface for the second type of user to confirm whether to allow the repair; if the second type of user submits a repair permission instruction through the second service interface, the repair operation is executed, thereby ensuring the second type of user's experience of using the resources in the target cloud computing product.

[0115] Thus, in this embodiment, the solution path for unknown anomalies generally includes: initiating an anomaly diagnosis request, providing visualized observation data of the entire resource chain to the first-class user, manual analysis by the first-class user, completing the problem repair, entering new knowledge into the anomaly diagnosis platform, and enriching the diagnostic capabilities of the anomaly diagnosis platform. For known anomalies, the solution path generally includes: detecting the known anomaly problem, determining the root cause, determining a repair solution, and completing the problem repair. Therefore, through the positive cycle of these two solution paths, the diagnostic capabilities of the anomaly diagnosis platform can be continuously enriched.

[0116] It can be seen that in this embodiment, after completing the aforementioned supplementary operations, the exception diagnosis platform has the ability to diagnose new exception problems. Therefore, after receiving subsequent exception diagnosis requests, the exception diagnosis platform can detect new exception problems without obstacles, diagnose the root cause and provide repair solutions, and realize fully automated diagnosis, without the need to trigger manual analysis of the first type of users, and there is no need to trigger offline work order flow.

[0117] Furthermore, during their research, the inventors discovered that, for various reasons, some abnormalities may not be suitable for monitoring items through tracking. Therefore, although these abnormalities are known, the monitoring items with which they are causally related cannot be observed in a timely manner, and abnormal changes in the related monitoring items cannot be perceived. This results in the inability to initiate abnormal diagnosis requests corresponding to these abnormalities in a timely manner. If these abnormalities are not repaired in a timely manner, they may cause user failures, affecting the product experience of the second category of users.

[0118] In this regard, this embodiment also proposes that the anomaly diagnosis platform can support cloud computing products in active diagnosis mode. Based on this, if the target cloud computing product has active diagnosis mode enabled, the platform will proactively perform anomaly detection based on monitoring items obtained from various resources included in the target cloud computing product, even if no request for anomaly detection is received for the target cloud computing product. This allows the platform to complete problem repair before the cloud computing product fails.

[0119] Based on this, when anomalies cannot be detected in time, the anomaly diagnosis platform can proactively detect anomalies in cloud computing products, thereby promptly detecting hidden anomalies and automatically performing root cause analysis and problem repair on the hidden anomalies to avoid hidden anomalies causing usage failures.

[0120] In summary, the anomaly detection solution provided in this embodiment can achieve at least the following technical effects:

[0121] 1) It can construct the causal relationship between monitoring items and abnormal problems, develop root cause analysis plans and repair plans for the newly discovered abnormal problems, and supplement these analysis results as new knowledge for the abnormal diagnosis platform, thereby continuously enriching the abnormal diagnosis platform's diagnostic capabilities for new abnormal problems.

[0122] 2) The anomaly diagnosis platform can support the service needs of different types of users. From the R&D perspective, it can provide rich and comprehensive clues and display key information in a more convenient way such as visualization, so that R&D personnel can quickly locate abnormal problems, quickly determine the root cause and develop repair plans; from the perspective of various personnel using cloud computing products, technical details can be hidden as much as possible, focusing on the automated solution of abnormal problems, thereby meeting such service needs.

[0123] 3) Each anomaly only requires one manual analysis. By establishing a forward-looking problem analysis and resolution loop, once R&D identifies the root cause of a new anomaly, they simply need to enter the analysis process and solution into the anomaly diagnosis platform. This will then enable the platform to analyze the root cause of the anomaly, eliminating the need for reanalysis the next time the same anomaly is encountered.

[0124] It should be noted that some of the processes described in the above embodiments and the accompanying drawings include multiple operations that appear in a specific order, but it should be clearly understood that these operations may not be executed in the order in which they appear in this article or may be executed in parallel. The serial numbers of the operations, such as 101, 102, etc., are only used to distinguish between different operations, and the serial numbers themselves do not represent any order of execution. In addition, these processes may include more or fewer operations, and these operations may be executed in sequence or in parallel. It should be noted that the descriptions of "first", "second", etc. in this article are used to distinguish different users or service interfaces, etc., and do not represent the order of precedence, nor do they limit "first" and "second" to be different types.

[0125] Figure 5 This is a schematic diagram of a computing device provided by another exemplary embodiment of the present application. Figure 5 As shown, the computing device includes a memory 50 , a processor 51 and a communication component 52 .

[0126] The processor 51 is coupled to the memory 50 and the communication component 52, and is configured to execute the computer program in the memory 50 to:

[0127] In response to receiving an anomaly diagnosis request for a target resource in a target cloud computing product, if the root cause of the anomaly is not diagnosed, collecting information based on preset information dimensions with each resource included in the target cloud computing product as a collection object;

[0128] Displaying the collected information on a first service interface for the first type of users, so that the first type of users can analyze the root cause of the anomaly based on the information;

[0129] The root cause submitted by the first type of user through the first service interface is obtained as the root cause of the target resource exception.

[0130] In an optional embodiment, when displaying the collected information on the first service interface for the first type of users, the processor 51 may be specifically configured to:

[0131] Convert the collected information into visual materials;

[0132] The visualization material is displayed on the first service interface to visualize the collected information.

[0133] Optionally, the preset information dimension includes descriptive information corresponding to abnormal problems that have been diagnosed on each resource included in the cloud computing product, descriptive information corresponding to operation events that occurred on the cloud computing product within a preset diagnostic period, and / or full-link tracking information corresponding to the operation events.

[0134] In an optional embodiment, the processor 51 may also be configured to:

[0135] Obtaining a repair plan formulated for the analyzed root cause, submitted by the first category of users through the first service interface;

[0136] Executing a repair operation according to the repair plan to complete the problem repair of the target resource;

[0137] The repair operation involves one or more resources in the target cloud computing product.

[0138] In an optional embodiment, the processor 51 may also be configured to:

[0139] Obtaining a new abnormal problem defined by the abnormal diagnosis request submitted by the first category of users through the first service interface, and a causal relationship between the monitoring item and the new abnormal problem analyzed based on the information;

[0140] The causal relationship provided by the first category of users is supplemented as a new basis for anomaly detection.

[0141] In an optional embodiment, the processor 51 may also be configured to:

[0142] Obtaining a root cause analysis plan for the new abnormal problem submitted by the first category of users through the first service interface, and supplementing it as a new basis for root cause analysis; and / or,

[0143] Obtain the correlation between the repair solution submitted by the first category of users through the first service interface and the root cause, and supplement it as a new basis for problem repair.

[0144] In an optional embodiment, the processor 51 may also be configured to:

[0145] After completing the supplemental operation, receiving a subsequent abnormality diagnosis request for any resource in the target computing product;

[0146] If it is detected in the anomaly detection step that the monitoring items monitored from the cloud computing product meet the supplemented causal relationship, it is determined that the new anomaly problem has been detected;

[0147] In the root cause analysis phase, a root cause analysis is performed based on the root cause analysis plan formulated for the new abnormal problem to determine the root cause;

[0148] In the problem repair phase, a repair operation is performed according to the repair solution associated with the determined root cause to complete the problem repair for the resource.

[0149] In an optional embodiment, the processor 51 may also be configured to:

[0150] If the target cloud computing product has enabled active diagnosis mode, then without receiving any abnormality diagnosis request for the target cloud computing product, anomaly detection will be actively performed based on the monitoring items obtained from the various resources included in the target cloud computing product, so as to complete the problem repair before the cloud computing product fails.

[0151] In an optional embodiment, the processor 51 may also be configured to:

[0152] If a response result query instruction for the abnormal diagnosis request is received from a second type of user through a second service interface, the root cause and the repair solution are displayed in the second service interface.

[0153] In an optional embodiment, the abnormality diagnosis request is submitted by the first category of users through the first service interface, or the abnormality diagnosis request is submitted by the second category of users through the second service interface.

[0154] In an optional embodiment, before performing the repair operation, the processor 51 may further be configured to:

[0155] If it is detected that the repair plan includes a repair operation that affects resource usage, a query message is displayed in the second service interface for the second type of user to confirm whether to allow the repair;

[0156] If a repair permission instruction is received from the second type of user through the second service interface, the repair operation is performed.

[0157] In an optional embodiment, the processor 51 may also be configured to:

[0158] After receiving the abnormality diagnosis request, if after performing abnormality detection based on the causal relationship between the known monitoring items and the abnormal problem, no detection result can be obtained or the detection result does not point to the known abnormal problem, it is determined that the root cause of the abnormality has not been diagnosed.

[0159] Further, if Figure 5 As shown, the computing device also includes: a power supply component 53 and other components. Figure 5 Only some components are shown schematically, and it does not mean that the computing device only includes Figure 5 Components shown.

[0160] It is worth noting that the technical details in the above-mentioned embodiments of the computing device can be referred to the relevant description in the aforementioned method embodiment. In order to save space, they will not be repeated here, but this should not cause any loss of the scope of protection of this application.

[0161] Accordingly, an embodiment of the present application further provides a computer-readable storage medium storing a computer program, which can implement the steps in the above method embodiment when the computer program is executed.

[0162] Accordingly, an embodiment of the present application further provides a computer program product, which can implement the steps in the above method embodiment when the computer program contained therein is executed.

[0163] above Figure 5 The memory in the computing platform is used to store computer programs and can be configured to store various other data to support operations on the computing platform. Examples of such data include instructions for any application or method operating on the computing platform, contact data, phone book data, messages, pictures, videos, etc. The memory can be implemented by any type of volatile or non-volatile storage device or a combination of them, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk or optical disk.

[0164] above Figure 5 The communication component in is configured to facilitate wired or wireless communication between the device where the communication component is located and other devices. The device where the communication component is located can access a wireless network based on a communication standard, such as WiFi, 2G, 3G, 4G / LTE, 5G and other mobile communication networks, or a combination thereof. In an exemplary embodiment, the communication component receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component also includes a near field communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on radio frequency identification (RFID) technology, infrared data association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology and other technologies.

[0165] above Figure 5 The power supply component in a device provides power to various components of the device in which the power supply component is located. The power supply component may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power to the device in which the power supply component is located.

[0166] Those skilled in the art will appreciate that the embodiments of the present application can be provided as methods, systems, or computer program products. Therefore, the present application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the present application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.

[0167] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the steps in the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0168] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.

[0169] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.

[0170] It should also be noted that the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, commodity, or apparatus that includes a series of elements includes not only those elements but also other elements not explicitly listed, or includes elements inherent to such process, method, commodity, or apparatus. In the absence of further limitations, an element defined by the phrase "comprises a ..." does not exclude the presence of other identical elements in the process, method, commodity, or apparatus that includes the element.

[0171] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of relevant countries and regions, and provide corresponding operation entrances for users to choose to authorize or refuse.

[0172] The foregoing is merely an embodiment of the present application and is not intended to limit the present application. For those skilled in the art, the present application may have various modifications and variations. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application should all be included in the protection scope of the present application.

Claims

1. A method for abnormality diagnosis, characterized in that: include: In response to receiving an anomaly diagnosis request for a target resource in a target cloud computing product, if the root cause of the anomaly is not diagnosed, collecting information based on preset information dimensions with each resource included in the target cloud computing product as a collection object; Displaying the collected information on a first service interface for the first type of users, so that the first type of users can analyze the root cause of the anomaly based on the information; The root cause submitted by the first type of user through the first service interface is obtained as the root cause of the target resource exception.

2. The method according to claim 1, characterized in that The collected information is displayed on the first service interface for the first type of users, including: Convert the collected information into visual materials; The visualization material is displayed on the first service interface to visualize the collected information.

3. The method according to claim 1, characterized in that The preset information dimensions include descriptive information corresponding to abnormal problems that have been diagnosed on each resource included in the cloud computing product, descriptive information corresponding to operation events that occurred on the cloud computing product within a preset diagnostic period, and / or full-link tracking information corresponding to the operation events.

4. The method according to claim 1, wherein Also includes: Obtaining a repair plan formulated for the analyzed root cause, submitted by the first category of users through the first service interface; Executing a repair operation according to the repair plan to complete the problem repair of the target resource; The repair operation involves one or more resources in the target cloud computing product.

5. The method according to claim 4, characterized in that Also includes: Obtaining a new abnormal problem defined by the abnormal diagnosis request submitted by the first category of users through the first service interface, and a causal relationship between the monitoring item and the new abnormal problem analyzed based on the information; The causal relationship provided by the first category of users is supplemented as a new basis for anomaly detection.

6. The method according to claim 5, characterized in that Also includes: Obtaining a root cause analysis plan for the new abnormal problem submitted by the first category of users through the first service interface, and supplementing it as a new basis for root cause analysis; and / or, Obtain the correlation between the repair solution submitted by the first category of users through the first service interface and the root cause, and supplement it as a new basis for problem repair.

7. The method according to claim 6, characterized in that Also includes: After completing the supplemental operation, receiving a subsequent abnormality diagnosis request for any resource in the target computing product; If it is detected in the anomaly detection step that the monitoring items monitored from the cloud computing product meet the supplemented causal relationship, it is determined that the new anomaly problem has been detected; In the root cause analysis phase, a root cause analysis is performed based on the root cause analysis plan formulated for the new abnormal problem to determine the root cause; In the problem repair phase, a repair operation is performed according to the repair solution associated with the determined root cause to complete the problem repair for the resource.

8. The method according to claim 6, characterized in that Also includes: If the target cloud computing product has enabled active diagnosis mode, then without receiving any abnormality diagnosis request for the target cloud computing product, anomaly detection will be actively performed based on the monitoring items obtained from the various resources included in the target cloud computing product, so as to complete the problem repair before the cloud computing product fails.

9. The method according to claim 4, characterized in that Also includes: If a response result query instruction for the abnormal diagnosis request is received from a second type of user through a second service interface, the root cause and the repair solution are displayed in the second service interface.

10. The method according to claim 9, characterized in that The abnormality diagnosis request is submitted by the first category user through the first service interface, or the abnormality diagnosis request is submitted by the second category user through the second service interface.

11. The method according to claim 4, characterized in that Before performing the repair operation, also include: If it is detected that the repair plan includes a repair operation that affects resource usage, a query message is displayed in the second service interface for the second type of user, so that the second type of user can confirm whether to allow the repair; If a repair permission instruction is received from the second type of user through the second service interface, the repair operation is performed.

12. The method according to claim 1, characterized in that Also includes: After receiving the abnormality diagnosis request, if after performing abnormality detection based on the causal relationship between the known monitoring items and the abnormal problem, no detection result can be obtained or the detection result does not point to the known abnormal problem, it is determined that the root cause of the abnormality has not been diagnosed.

13. A computing device, characterized in that including memory, processor, and communication components; The memory is used to store one or more computer instructions; The processor is coupled to the memory and the communication component, and is configured to execute the one or more computer instructions to execute the abnormality diagnosis method according to any one of claims 1 to 12.

14. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by one or more processors, the one or more processors are caused to execute the abnormality diagnosis method according to any one of claims 1 to 12.

15. A computer program product, characterized in that The invention comprises a computer program, which, when executed by one or more processors, causes the one or more processors to execute the abnormality diagnosis method according to any one of claims 1 to 12.