Fault diagnosis method and device

By applying fault diagnosis methods on the CI/CD pipeline, using SOP nodes and agents to generate diagnostic results, the difficulty of fault location and diagnosis of CI/CD pipeline is solved, and the degree of automation and efficiency are improved.

CN119961033AActive Publication Date: 2025-05-09HUAWEI CLOUD COMPUTING TECHNOLOGIES CO LTD

Patent Information

Application Number
CN202411833490.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-12
Publication Date
2025-05-09
Estimated Expiration
2044-12-12

AI Technical Summary

Technical Problem

It is difficult to locate and diagnose faults in CI/CD pipelines, mainly due to the wide variety of faults, complex logs, complex distributed architecture and difficulty in team collaboration.

Method used

The diagnostic results are determined by obtaining the category and problem description of CI/CD pipeline failures, as well as multiple standard operating instructions SOP nodes. If the first diagnostic result cannot solve the fault, use the RAG agent and the LLM agent to generate the second diagnostic result to be used as a solution to the fault.

Benefits of technology

It improves the automation and scalability of CI/CD pipeline fault diagnosis, and can quickly obtain current fault solutions based on solutions of historical faults, reducing the possibility of manual intervention.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119961033A_ABST
    Figure CN119961033A_ABST
Patent Text Reader

Abstract

The invention provides a fault diagnosis method, the method is applied to a fault diagnosis service of a CI / CD pipeline, and the method comprises the steps: obtaining the type of a fault of the CI / CD pipeline, a first problem description of the fault, and a plurality of SOP nodes of the fault diagnosis service, the SOP nodes being used for obtaining a diagnosis result according to the problem description; and determining at least one first diagnosis result according to the first problem description and at least one first SOP node corresponding to the category of the fault in the plurality of SOP nodes. Outputting the first diagnosis result under the condition that the at least one first diagnosis result can be used as a fault solution; or, under the condition that the first diagnosis result cannot be used as a fault solution, a second diagnosis result is obtained according to at least one first diagnosis result and an RAG agent program and / or an LLM agent program of the fault diagnosis service, and the second diagnosis result is used as the fault solution; and outputting the second diagnosis result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer technology, and more specifically, to a fault diagnosis method and device. Background Art

[0002] Continuous integration (CI) and continuous delivery / deployment (CD), often referred to as CI / CD, are a practice in software development. Specifically, continuous integration emphasizes that developers frequently submit code changes to the code repository and automatically build and test to ensure the stability and reliability of the code; continuous delivery further deploys these tested codes to the test environment and is ready to be released to the production environment at any time, but this step does not automatically deploy the code to the production environment; continuous deployment is an extension of continuous delivery, which automatically releases the code to the production environment to achieve rapid iteration and delivery. CI / CD significantly accelerates the software development life cycle, improves development efficiency and software quality through automated processes, and is an indispensable part of modern agile software development.

[0003] The diversity and complexity of the tools involved in the CI / CD pipeline increase the probability of failure and the difficulty of troubleshooting and repairing. The difficulty of fault location and diagnosis lies in the wide variety of faults, complex logs, complex distributed architecture, and difficulty in team collaboration. In order to improve the efficiency of fault location and diagnosis of the CI / CD pipeline, it is necessary to optimize the automatic diagnosis service supporting the CI / CD pipeline. Summary of the invention

[0004] The present application provides a fault diagnosis method and device to improve the efficiency of fault diagnosis for a continuous integration and continuous delivery / deployment (CI / CD) pipeline.

[0005] In a first aspect, a fault diagnosis method is provided, which is applied to a fault diagnosis service of a CI / CD pipeline, and the method includes: obtaining the category of the fault of the CI / CD pipeline and the first problem description of the fault, and multiple standard operation instruction SOP nodes of the fault diagnosis service, wherein the SOP node is used to obtain a diagnosis result according to the problem description; according to the first problem description and at least one first SOP node corresponding to the category of the fault in the multiple SOP nodes, at least one first diagnosis result is determined. In the case where at least one first diagnosis result can be used as a solution to the fault, the first diagnosis result is output; or, in the case where the first diagnosis result cannot be used as a solution to the fault, a second diagnosis result is obtained according to the at least one first diagnosis result and the retrieval enhancement generation RAG agent and / or large language model LLM agent of the fault diagnosis service, and the second diagnosis result is used as a solution to the fault; and the second diagnosis result is output.

[0006] The embodiment of the present application determines the diagnosis result according to the SOP node of the fault diagnosis service of the CI / CD pipeline, which can improve the automation and scalability of the fault diagnosis method, so as to obtain the solution to the current fault according to the solution to the corresponding historical fault. When encountering a fault that cannot be solved by the first diagnosis result, the second diagnosis result can also be determined according to the RAG agent program to be used as a solution to the fault, thereby improving the efficiency of fault diagnosis for the CI / CD pipeline.

[0007] In certain implementations of the first aspect, at least one first SOP node forms a SOP tree, and at least one first diagnostic result is determined based on the first problem description and at least one first SOP node corresponding to the category of the fault among multiple SOP nodes, including: obtaining a third diagnostic result based on the second SOP node in the SOP tree, wherein the third diagnostic result is used to determine the first diagnostic result, or the third diagnostic result is used as the first diagnostic result, and the second SOP node is a SOP node among the at least one first SOP node; determining the first diagnostic result and whether the first diagnostic result can be used as a solution to the fault based on the first condition, the second condition, the third condition and the third diagnostic result, wherein the first condition is that the third diagnostic result can be used as a solution to the fault, the second condition is that the second SOP node is a leaf node of the SOP tree, and the third condition is that the third diagnostic result can be used as a problem description of a child node of the input second SOP node.

[0008] In such an implementation, the structure of the SOP tree can guide the fault diagnosis service to execute the path of the SOP node, enabling the fault diagnosis service to determine the first diagnostic result and whether the first diagnostic result can be used as a solution to the fault, thereby reducing the possibility of manual intervention in the fault diagnosis process, thereby improving the efficiency of fault diagnosis for the CI / CD pipeline.

[0009] In certain implementations of the first aspect, when the first condition and the second condition are not met: when the third condition is not met, the third diagnostic result is used as the first diagnostic result, wherein the first diagnostic result cannot be used as a solution to the fault; or, when the third condition is met, a third diagnostic result is obtained, including: obtaining the third diagnostic result based on the sub-nodes of the second SOP node and the second problem description, wherein the diagnostic result output by the second SOP node is used as the second problem description, and after determining the second problem description, the sub-nodes of the second SOP node are used as the second SOP node.

[0010] In such an implementation, the SOP tree can break down a broad problem description into multiple problem descriptions and diagnostic results with logical relationships and dependencies. The SOP nodes preset by the fault diagnosis service can correspond to clear and specific functions to improve the efficiency of fault diagnosis for the CI / CD pipeline.

[0011] In certain implementations of the first aspect, when the first condition is not satisfied and the second condition is satisfied, the third diagnosis result is used as the first diagnosis result, wherein the first diagnosis result cannot be used as a solution to the fault.

[0012] In such an implementation, when the existing knowledge of the SOP tree cannot solve the fault, the first diagnostic result can be used as input to the agent program of the fault diagnosis service to obtain a second diagnostic result. The fault diagnosis service can dynamically learn the solution to the fault, thereby improving the efficiency of fault diagnosis for the CI / CD pipeline in subsequent use.

[0013] In certain implementations of the first aspect, when the first condition is met, the third diagnostic result is used as the first diagnostic result, wherein the first diagnostic result can be used as a solution to the fault.

[0014] In certain implementations of the first aspect, the method further includes: generating at least one third SOP node according to the fault category and the first problem description, wherein the third SOP node corresponds to the fault category and is used to obtain a second diagnostic result according to the first problem description.

[0015] In such an implementation, the RAG agent can dynamically learn solutions to faults based on existing knowledge and unlearned faults to improve the efficiency of fault diagnosis for the CI / CD pipeline in subsequent use.

[0016] In certain implementations of the first aspect, the SOP node includes metadata of at least one knowledge unit corresponding to the same category as the SOP node, wherein the knowledge unit includes problem descriptions and solutions of historical failures of the CI / CD pipeline, and the SOP node is obtained based on the at least one knowledge unit.

[0017] In such an implementation, the SOP node includes metadata of the knowledge unit, which can more accurately and quickly determine the association between the SOP node and the historical faults, so as to improve the automation and scalability of the fault diagnosis method.

[0018] In certain implementations of the first aspect, the SOP node is used to obtain a diagnosis result based on a problem description, and a problem description and a solution of at least one knowledge unit corresponding to the SOP node.

[0019] In such an implementation, the knowledge units correspond to common faults in the CI / CD pipeline and difficult faults whose solutions mainly rely on human experience. Obtaining diagnostic results based on the knowledge units can further improve the automation and scalability of the fault diagnosis method.

[0020] In certain implementations of the first aspect, the SOP node is used to obtain a diagnostic result based on a problem description and an agent program corresponding to the SOP node, wherein the agent program includes at least one of an application programming interface (API) agent program, an LLM agent program, and a RAG agent program.

[0021] In such an implementation, the API agent can make the fault diagnosis steps repeatable, traceable, and track progress to improve the automation and scalability of fault diagnosis. The LLM agent can effectively process natural language data such as logs, code snippets, and problem descriptions to enable fault diagnosis services to work collaboratively with other services. The RAG agent can effectively integrate human experience and external databases to try to solve unlearned faults, thereby improving the automation and scalability of fault diagnosis methods.

[0022] In a second aspect, a fault diagnosis device is provided, which is applied to fault diagnosis services of continuous integration and continuous delivery / deployment CI / CD pipelines, and the device includes: an acquisition module, used to obtain the category of the fault of the CI / CD pipeline and a first problem description of the fault, and multiple standard operation instruction SOP nodes of the fault diagnosis service, wherein the SOP node is used to obtain a diagnosis result according to the problem description; a processing module, used to determine at least one first diagnosis result according to the first problem description and at least one first SOP node corresponding to the category of the fault in the multiple SOP nodes; when at least one first diagnosis result can be used as a solution to the fault, output the first diagnosis result; or, when the first diagnosis result cannot be used as a solution to the fault, generate a RAG agent and / or a large language model LLM agent according to the at least one first diagnosis result and the retrieval enhancement of the fault diagnosis service to obtain a second diagnosis result, and the second diagnosis result is used as a solution to the fault; and output the second diagnosis result.

[0023] In some implementations of the second aspect, the at least one first SOP node forms a SOP tree, and the processing module is further configured to:

[0024] A third diagnostic result is obtained according to the second SOP node in the SOP tree, wherein the third diagnostic result is used to determine the first diagnostic result, or the third diagnostic result is used as the first diagnostic result, and the second SOP node is a SOP node in at least one first SOP node; according to the first condition, the second condition, the third condition and the third diagnostic result, the first diagnostic result and whether the first diagnostic result can be used as a solution to the fault are determined, wherein the first condition is that the third diagnostic result can be used as a solution to the fault, the second condition is that the second SOP node is a leaf node of the SOP tree, and the third condition is that the third diagnostic result can be used as a problem description of a child node of the input second SOP node.

[0025] In certain implementations of the second aspect, when the first condition and the second condition are not met: when the third condition is not met, the third diagnostic result is used as the first diagnostic result, wherein the first diagnostic result cannot be used as a solution to the fault; or, when the third condition is met, the processing module is also used to: obtain the third diagnostic result based on the sub-nodes of the second SOP node and the second problem description, wherein the diagnostic result output by the second SOP node is used as the second problem description, and after determining the second problem description, the sub-nodes of the second SOP node are used as the second SOP node.

[0026] In certain implementations of the second aspect, when the first condition is not satisfied and the second condition is satisfied, the third diagnosis result is used as the first diagnosis result, wherein the first diagnosis result cannot be used as a solution to the fault.

[0027] In certain implementations of the second aspect, when the first condition is met, the third diagnosis result is used as the first diagnosis result, wherein the first diagnosis result can be used as a solution to the fault.

[0028] In certain implementations of the second aspect, the processing module is further used to generate at least one third SOP node based on the fault category and the first problem description, wherein the third SOP node corresponds to the fault category and is used to obtain a second diagnostic result based on the first problem description.

[0029] In certain implementations of the second aspect, the SOP node includes metadata of at least one knowledge unit corresponding to the same category as the SOP node, wherein the knowledge unit includes problem descriptions and solutions of historical failures of the CI / CD pipeline, and the SOP node is obtained based on the at least one knowledge unit.

[0030] In certain implementations of the second aspect, the SOP node is used to obtain a diagnosis result based on a problem description, and a problem description and a solution of at least one knowledge unit corresponding to the SOP node.

[0031] In certain implementations of the second aspect, the SOP node is used to obtain a diagnostic result based on a problem description and an agent program corresponding to the SOP node, wherein the agent program includes at least one of an application programming interface API agent program, an LLM agent program, and a RAG agent program.

[0032] In a third aspect, a computing device cluster is provided, the computing device cluster comprising at least one computing device, each computing device comprising a processor and a memory; the processor of the at least one computing device is used to execute instructions stored in the memory of the at least one computing device, so that the computing device cluster executes the method described in any implementation manner of the first aspect.

[0033] In a fourth aspect, a computer program product comprising instructions is provided. When the instructions are executed by a computing device cluster, the computing device cluster executes the method as described in any implementation manner of the first aspect.

[0034] In a fifth aspect, a computer-readable storage medium is provided, comprising computer program instructions. When the computer program instructions are executed by a computing device cluster, the computing device cluster executes the method as described in any implementation manner of the first aspect. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] Figure 1 It is a flowchart of a fault diagnosis method provided in an embodiment of the present application.

[0036] Figure 2 It is a schematic block diagram of the architecture of a fault diagnosis service provided in an embodiment of the present application.

[0037] Figure 3 It is a flowchart of a fault diagnosis method provided in an embodiment of the present application.

[0038] Figure 4 It is a schematic block diagram of a fault diagnosis device provided in an embodiment of the present application.

[0039] Figure 5 It is a schematic block diagram of a computing device provided in an embodiment of the present application.

[0040] Figure 6 It is a schematic block diagram of a computing device cluster provided in an embodiment of the present application.

[0041] Figure 7 It is a schematic block diagram of a computing device cluster provided in an embodiment of the present application. DETAILED DESCRIPTION

[0042] First, technical terms related to this application are explained.

[0043] Continuous integration (CI) and continuous delivery / deployment (CD), often referred to as CI / CD, are a practice in software development. Specifically, continuous integration emphasizes that developers frequently submit code changes to the code repository and automatically build and test to ensure the stability and reliability of the code; continuous delivery further deploys these tested codes to the test environment and is ready to be released to the production environment at any time, but this step does not automatically deploy the code to the production environment; continuous deployment is an extension of continuous delivery, which automatically releases the code to the production environment to achieve rapid iteration and delivery.

[0044] In the CI / CD process, standard operating procedures (SOPs) define a series of standard steps and specifications for executing continuous integration and continuous delivery / deployment. SOP nodes refer to each node in these standard operating instructions, representing a specific task or operation link, such as code submission, automated building, testing, deployment, etc. The SOP tree is a collection of SOP nodes organized in a tree structure. It intuitively shows the logical relationship and execution order between the nodes in the CI / CD process, ensuring the coherence and accuracy of the entire process.

[0045] In the SOP tree, the parent-child node relationship is the most basic and core relationship. The parent node represents a larger task or operation, while the child node represents the smaller and more specific tasks or operations that the task or operation is decomposed into. For example, in the CI / CD process, a parent node may represent an overall build task, while its child nodes may represent specific build steps such as code pulling, compiling, and packaging. Brother nodes refer to two or more nodes with the same parent node. Brother nodes are each responsible for different tasks or operations, and these tasks or operations are performed within the framework of the larger task or operation represented by their common parent node. For example, in the CI / CD process, automated building and automated testing may be two parallel brother nodes. In the SOP tree, a node without a parent node is called a root node, which is usually used as the entrance to the SOP process. A node without a child node is called a leaf node. After completing the task corresponding to the leaf node, the SOP process ends. Other nodes in the SOP tree are usually called internal nodes.

[0046] In the CI / CD process, the agent is used as a bridge between the SOP node and the actual operation execution. Specifically, the agent is assigned to execute tasks in one or more SOP nodes. When the CI / CD process advances to a certain SOP node, the corresponding agent will be triggered and start to execute the specific operations defined in the node. In addition, the agent can also make adaptive adjustments based on the results of environmental perception and user interaction to improve the flexibility and robustness of task execution.

[0047] Retrieval-augmented generation (RAG) is a model architecture that combines the advantages of information retrieval and text generation. It retrieves documents or paragraphs related to the input from an external knowledge base through a retrieval system before generating text, and then inputs the retrieved information into the generation model as context, thereby enhancing the accuracy and relevance of the generated content. RAG technology is mainly divided into a retrieval stage and a generation stage. In the retrieval stage, the model uses a retrieval system (such as vector-based retrieval technology) to retrieve documents or paragraphs related to the input query from a pre-established knowledge base. These knowledge bases can be any form of document collection, such as Wikipedia, professional databases, academic papers, etc. In the generation stage, a generative model, such as a pre-trained large language model (LLM), uses the retrieved information as context input and combines it with the original query to generate the final text content.

[0048] Retrieval-augmented generation agent (RAG agent) is an intelligent entity that combines RAG technology with the functions and characteristics of an agent. RAG agent can perceive the environment, process reasoning, make decisions and perform tasks. It can retrieve information from a large knowledge base and generate text based on it, thereby achieving more accurate and diverse text content creation. It not only has powerful language understanding and generation capabilities, but also can enhance the accuracy of generated content based on real-time retrieved information, making text generation more intelligent and personalized.

[0049] The automatic diagnosis service is an intelligent tool that can automatically monitor, analyze and diagnose potential problems in the CI / CD process. By collecting and analyzing log information, performance indicators and other relevant data in real time, the automatic diagnosis service can quickly identify faults in the code, configuration problems or performance bottlenecks, and provide detailed diagnostic reports and solution suggestions.

[0050] The question-knowledge base, also known as the question-answer database, is an information organization method that uses key-value pairs to store information. Each question is used as a key, and the related answers, solutions, or fault information are stored as values, forming multiple question-knowledge pairs, which are also called knowledge units. In the CI / CD pipeline, the question-knowledge base usually brings together known problems and corresponding solutions from multiple sources such as historical faults, code analysis, log records, and operation and maintenance team experience.

[0051] The diversity and complexity of the tools involved in the CI / CD pipeline increase the probability of failure and the difficulty of troubleshooting and repairing. The difficulty of fault location and diagnosis lies in the wide variety of faults, complex logs, complex distributed architecture, and difficulty in team collaboration. In order to improve the efficiency of fault location and diagnosis of the CI / CD pipeline, it is necessary to optimize the automatic diagnosis service supporting the CI / CD pipeline.

[0052] In view of this, the present application provides a fault diagnosis method, which is applied to the fault diagnosis service of continuous integration and continuous delivery / deployment CI / CD pipeline, and the method includes: obtaining the category of the fault of the CI / CD pipeline and the first problem description of the fault, and multiple standard operation instruction SOP nodes of the fault diagnosis service, wherein the SOP node is used to obtain a diagnosis result according to the problem description; according to the first problem description and at least one first SOP node corresponding to the category of the fault in the multiple SOP nodes, at least one first diagnosis result is determined. In the case where at least one first diagnosis result can be used as a solution to the fault, the first diagnosis result is output; or, in the case where the first diagnosis result cannot be used as a solution to the fault, a second diagnosis result is obtained according to the at least one first diagnosis result and the retrieval enhancement generation RAG agent and / or large language model LLM agent of the fault diagnosis service, and the second diagnosis result is used as a solution to the fault; the second diagnosis result is output.

[0053] The embodiment of the present application determines the diagnosis result according to the SOP node of the fault diagnosis service of the CI / CD pipeline, which can improve the automation and scalability of the fault diagnosis method, so as to obtain the solution to the current fault according to the solution to the corresponding historical fault. When encountering a fault that cannot be solved by the first diagnosis result, the second diagnosis result can also be determined according to the RAG agent program to be used as a solution to the fault, thereby improving the efficiency of fault diagnosis for the CI / CD pipeline.

[0054] Combine the following Figure 1 The flowchart of the fault diagnosis method shown in FIG. 1 illustrates the specific process of applying the method to the fault diagnosis service of the CI / CD pipeline.

[0055] S110, obtaining the category of the fault of the CI / CD pipeline and the first problem description of the fault, as well as multiple SOP nodes of the fault diagnosis service, wherein the SOP nodes are used to obtain a diagnosis result according to the problem description.

[0056] First, the method of obtaining the category of CI / CD pipeline faults and the description of the fault problem is described. In some embodiments, when a user executes a CI / CD pipeline and the CI / CD pipeline is interrupted due to a fault, the fault diagnosis service can directly obtain the fault log of the CI / CD pipeline. Of course, the user can also use natural language to describe the cause, type, possible solutions, etc. of the fault, and input the corresponding description through the large language model of the fault diagnosis service, so that the fault diagnosis service can obtain more fault information.

[0057] In a possible implementation, after obtaining the fault log of the CI / CD pipeline, the fault log can be sorted into the fault category and fault problem description according to the log parsing service in the fault diagnosis service. First, log screening is performed to eliminate irrelevant logs and only retain log entries directly related to the fault; then log analysis is performed to determine the approximate category of the fault and the core fault information using pattern matching and log classification technology; finally, context extraction is performed to collect context metadata related to this task through the metadata of the CI / CD process and the execution machine, such as the identifier of the build machine, code version, build configuration, execution machine file directory, environment configuration and other metadata. The final problem description usually includes accurate fault logs and context metadata, and can also include the user's natural language description. The category of the fault can be represented by the relevant metadata. It should be understood that the classification method of fault information includes classification based on at least one combination of the aforementioned multiple metadata, and the present application does not limit the classification method.

[0058] Before the fault diagnosis service is applied to the CI / CD pipeline, it is necessary to set up a SOP knowledge base for the fault diagnosis service. The following describes a method for obtaining multiple preset SOPs.

[0059] S210, obtaining multiple knowledge units according to the problem description and solution of the historical failure of the CI / CD pipeline.

[0060] In some embodiments, the development team builds a question-and-answer database based on experience. First, data such as problem descriptions (such as fault logs) and solutions (such as optimization strategies) of historical faults closely related to the CI / CD pipeline are collected and stored in the database; secondly, these data are sorted by category; finally, the correspondence between the problem description and the solution of each fault is clarified to form multiple knowledge units.

[0061] In one possible implementation, a knowledge unit includes multiple metadata, and the metadata is used to determine the classification of the knowledge unit. For example, metadata related to the problem description may include a problem identifier, a problem title that is a brief description or summary of the problem, a fault log, stages of the CI / CD testing process such as build, test, and deployment, and problem types such as build failure, deployment failure, and dependency conflict, as well as the fault code of the problem. Metadata related to the solution (i.e., knowledge) may include an answer identifier, and sources of answers such as expert provision, historical records, and external documents. Other metadata may include tags consisting of keywords or phrases related to the problem or knowledge, as well as identifiers of associated problems, etc.

[0062] It should be understood that the classification method of knowledge units includes a combination of at least one of multiple categories such as the stage of the CI / CD testing process, the type of problem, the fault code of the problem, the source of the answer, the question label, etc. This application does not limit the selection of categories in the clustering algorithm.

[0063] In one possible implementation, a clustering algorithm is used to group multiple knowledge units in a question-answering database into multiple categories. Specifically, a clustering algorithm groups data points in a data set according to a certain standard (such as distance or similarity) so that data points in the same cluster are similar to each other, while data points in different clusters are quite different. This grouping process does not require pre-defined category labels, but allows the algorithm to automatically discover patterns and structures in the data. Common clustering algorithms include K-means, hierarchical clustering, and density-based spatial clustering applications with noise (DBSCAN).

[0064] S220, generating a plurality of SOP nodes according to a plurality of knowledge units and a preset generation rule, wherein the SOP nodes include metadata of at least one knowledge unit corresponding to the same category as the SOP nodes.

[0065] In some embodiments, the SOP node includes metadata, for example, the metadata may include a node identifier, a node name, a node type such as build, test, and deploy, trigger conditions such as code submission, a specific time, and completion of dependent tasks, input parameters such as code repository address, build parameters, and environment variables, operation instructions such as querying static check configuration and obtaining build configuration, expected outputs such as build products and test reports, execution environments such as specific versions of compilers and test frameworks, log output configurations, and dependencies, etc. The metadata of the SOP node also includes metadata of at least one corresponding knowledge unit, so that the SOP node includes as complete and comprehensive information as possible to determine the correspondence between the knowledge unit and the SOP node.

[0066] In a possible implementation, the metadata of the knowledge unit can be preprocessed using algorithms such as word segmentation and stem extraction, and then multiple SOP nodes can be obtained based on predefined generation rules such as text similarity or keyword matching. Among them, specific processes such as word segmentation, stem extraction, and text similarity matching can be implemented through a large language model (LLM). Alternatively, a suitable prompt template is set for the LLM so that the LLM can generate SOP nodes that conform to a certain format based on the metadata of the knowledge unit. Finally, the generated multiple SOP nodes may require further manual review and modification.

[0067] In such an implementation, the SOP node includes metadata of the knowledge unit, which can more accurately and quickly determine the association between the SOP node and the historical faults, so as to improve the automation and scalability of the fault diagnosis method.

[0068] S230, obtaining a SOP knowledge base according to a plurality of SOP nodes, and the SOP nodes in the SOP knowledge base form a SOP tree.

[0069] In some embodiments, multiple SOP trees are obtained according to multiple SOP nodes and multiple knowledge units. The metadata of multiple SOP nodes is read, including node ID, name, type, execution condition, input parameter, etc., and the detailed content of the knowledge unit they include, to understand the function and role of each node. Then, according to the metadata of some types of SOP nodes, such as their functions or positions in the CI / CD process, they are divided into different categories. For example, nodes can be divided into categories such as construction, testing, and deployment. Secondly, the dependency between SOP nodes is analyzed. Usually, the execution of a SOP node may depend on the completion of another node, and this dependency determines the parent-child relationship in the SOP tree. In addition, other metadata such as trigger conditions, construction parameters, and execution environment also indirectly include information on the dependency between SOP nodes. Finally, based on the dependency, the parent node and child node of each SOP node are determined. The parent node is the node that triggers the execution of the child node, and the child node is the node that is executed after the parent node is executed. The metadata of the knowledge unit is preprocessed, and the mode of constructing the SOP tree according to the preset matching rules is similar to the aforementioned embodiment, and will not be repeated here. Obviously, according to different classification methods, this embodiment will generate multiple SOP trees organized in different ways.

[0070] In a possible implementation, the SOP node and the corresponding knowledge unit may not be stored in the same database. Figure 2The schematic block diagram of the error diagnosis service shown in the figure, SOP nodes form multiple SOP trees according to categories and are stored in the SOP knowledge base, and the knowledge units are stored in the question and answer database. Specifically, SOP node #101 corresponds to 2 knowledge units #101, and SOP node #203 corresponds to 2 knowledge units #203. It can be considered that SOP node #101 includes 2 knowledge units #101, and SOP node #203 includes 2 knowledge units #203.

[0071] S120 , determining at least one first diagnosis result according to the first problem description and at least one first SOP node corresponding to the category of the fault among the plurality of SOP nodes.

[0072] In some embodiments, the representation of the diagnosis result is similar to the representation of the problem description. The diagnosis result includes the diagnosis log and context metadata, and may also include the natural language output by the LLM. In addition, the diagnosis result may also include a category determined based on the context metadata. Therefore, in the SOP tree, the diagnosis result output by an SOP node can be used as the problem description input by its child node.

[0073] In some embodiments, at least one first SOP node is formed as Figure 2 The SOP tree shown can obtain a third diagnostic result according to the second SOP node in the SOP tree, wherein the third diagnostic result is used to determine the first diagnostic result, or the third diagnostic result is used as the first diagnostic result, and the second SOP node is a SOP node in at least one first SOP node. Then, according to the first condition, the second condition, the third condition and the third diagnostic result, determine the first diagnostic result and whether the first diagnostic result can be used as a solution to the fault, wherein the first condition is that the third diagnostic result can be used as a solution to the fault, the second condition is that the second SOP node is a leaf node of the SOP tree, and the third condition is that the third diagnostic result can be used as a problem description of a child node of the second SOP node.

[0074] In such an implementation, the structure of the SOP tree can guide the fault diagnosis service to execute the path of the SOP node, enabling the fault diagnosis service to determine the first diagnostic result and whether the first diagnostic result can be used as a solution to the fault, thereby reducing the possibility of manual intervention in the fault diagnosis process, thereby improving the efficiency of fault diagnosis for the CI / CD pipeline.

[0075] Specifically, the SOP node can perform various tasks through the agent corresponding to the SOP node in the fault diagnosis service, wherein the agent corresponds to multiple categories and can perform tasks in multiple ways, for example, the application programming interface (API) agent can query open source component information, obtain access control template details, retrieve and build machine environment, etc. These tasks are completed by calling the corresponding API; LLMagent can call the API of a large language model (such as ChatGPT, Qwen or a self-deployed large language model) to generate the required natural language and code snippets through preset or dynamically spliced ​​prompts; RAGagent can retrieve and obtain relevant information from the knowledge base, and combine the text generation capabilities of the aforementioned LLM to support the execution of tasks. The knowledge base may include a knowledge base from the public Internet, or a question-and-answer database provided in this application. It should be understood that the above description is only an example, and this application does not limit the way in which the agent executes SOP.

[0076] In such an implementation, the API agent can make the fault diagnosis steps repeatable, traceable, and track progress to improve the automation and scalability of fault diagnosis. The LLM agent can effectively process natural language data such as logs, code snippets, and problem descriptions to enable fault diagnosis services to work collaboratively with other services. The RAG agent can effectively integrate human experience and external databases to try to solve unlearned faults, thereby improving the automation and scalability of fault diagnosis methods.

[0077] According to the following Figure 3 The flowchart diagram shown exemplifies a method for determining at least one first diagnosis result and whether the first diagnosis result can be used as a solution to a fault.

[0078] In some embodiments, when the first condition and the second condition are not satisfied and the third condition is satisfied, a third diagnostic result is obtained according to the subnodes of the second SOP node and the second problem description, wherein the diagnostic result output by the second SOP node is used as the second problem description, and after determining the second problem description, the subnodes of the second SOP node are used as the second SOP node. Figure 2In the SOP tree corresponding to category #1 shown, the initial value of the second SOP node is the root node (i.e., SOP node #101), and the category corresponding to the root node is "build failure", and the specific function is to determine the specific reason for the build failure. Assuming that the current CI / CD pipeline fails in the build phase because the API agent retrieves that version 5.6 of one of the open source software packages A used has a high-risk vulnerability, resulting in the software package being intercepted (referred to as the vulnerability interception scenario), the diagnostic result output by SOP node #101 is similar to "(vulnerability interception, package A, 5.6, including logs of vulnerability information)". Obviously, this diagnostic result cannot be directly used as a solution to the vulnerability interception problem, that is, the diagnostic result will not be used as the first diagnostic result, that is, the first condition is not met. In addition, the root node is not a leaf node, that is, the second condition is not met.

[0079] Next, it is determined whether the diagnostic result output by the second SOP node can be used as the problem description of its child node. For example, the specific function of a child node #102 of SOP node #101 is to determine the version number compatible with the specific version of a software package according to the specific version of the software package, for example, according to version 5.6 of software package A, it is determined that version 5.5 and version 5.4 are compatible versions, that is, the diagnostic result output by SOP node #101 is used as the second problem description, and SOP node #102 is used as the new second SOP node.

[0080] In such an implementation, the SOP tree can break down a broad problem description into multiple problem descriptions and diagnostic results with logical relationships and dependencies. The SOP nodes preset by the fault diagnosis service can correspond to clear and specific functions to improve the efficiency of fault diagnosis for the CI / CD pipeline.

[0081] In a possible implementation, the SOP node is used to obtain a diagnostic result based on a problem description, and a problem description and a solution of at least one knowledge unit corresponding to the SOP node. Specifically, in order to determine whether the diagnostic result output by the second SOP node can be used as a problem description of a child node, the text similarity between the log of the diagnostic result and the log of the historical fault in the knowledge unit of the child node can be calculated using an LLM agent. Subsequently, the diagnostic result can be adjusted or a more reliable diagnostic result can be regenerated based on the text similarity, so that the diagnostic result output by the second SOP node can be used as a problem description of a child node as much as possible. Of course, the problem description of the historical fault and the current diagnostic result also include other information such as context metadata, which can also improve the reliability of the diagnostic result. The present application does not limit the specific method for obtaining the diagnostic result.

[0082] In such an implementation, the knowledge units correspond to common faults in the CI / CD pipeline and difficult faults whose solutions mainly rely on human experience. Obtaining diagnostic results based on the knowledge units can further improve the automation and scalability of the fault diagnosis method.

[0083] In a possible implementation, when the first condition is met, the third diagnostic result is used as the first diagnostic result, wherein the first diagnostic result can be used as a solution to the fault. For example, the APIagent finds software package A with version 5.5 in the software package repository, and SOP node #102 can output the diagnostic result "Change the version of software package A to 5.5 and rebuild" through the LLM agent. In the case that there are no high-risk vulnerabilities in versions 5.5 and 5.4 of software package A, "Change the version of software package A to 5.5 and rebuild" is used as the first diagnostic result and can be used as a solution to the fault.

[0084] In a possible implementation, when the first condition, the second condition and the third condition are not met, the third diagnostic result is used as the first diagnostic result, wherein the first diagnostic result cannot be used as a solution to the fault. For example, when versions 5.5 and 5.4 of software package A also have vulnerabilities similar to version 5.6, the diagnostic result obtained by SOP node #102 is "version 5.5 of software package A has similar vulnerabilities". This diagnostic result is only an explanation for the fault, and it is obviously not used as a solution to the fault, that is, it does not meet the first condition. In addition, SOP node #103 is not a leaf node and does not meet the second condition. Moreover, the specific function of the child node #103 of SOP node #102 is set under the assumption that the diagnostic result of SOP node #102 is "change the version of software package A to 5.5 and rebuild", such as "marking and trying to repair the programming interface different from the software package after the version is changed and the software package originally used", and it is obvious that "version 5.5 of software package A has similar vulnerabilities" cannot be used as a problem description of input SOP node #103. In summary, the SOP tree cannot further resolve the fault, and “version 5.5 of software package A has a similar vulnerability” is used as the first diagnostic result and cannot be used as a solution to the fault.

[0085] In some embodiments, the log tool of software package C will output a fault log to indicate that software package C triggered a fault during the entire software building process, but the log does not include the type and context information of the fault. Specifically, a suitable prompt word template can be set for the LLM agent so that the LLM agent outputs a natural language description to indicate that the fault is related to software package C, but the problem description of the fault is incomplete. The natural language description is used as the first diagnostic result and cannot be used as a solution to the fault. In addition, the LLM agent can also obtain common faults that may occur in software package C based on the question and answer database for further reference by developers.

[0086] In a possible implementation, when the first condition is not met and the second condition is met, the third diagnosis result is used as the first diagnosis result, wherein the first diagnosis result cannot be used as a solution to the fault. Figure 2 The SOP tree corresponding to category #1 shown does not include SOP node #103, so SOP node #102 is a leaf node, satisfying the second condition, and "version 5.5 of software package A has a similar vulnerability" is used as the first diagnostic result and cannot be used as a solution to the fault.

[0087] In such an implementation, when the existing knowledge of the SOP tree cannot solve the fault, the first diagnostic result can be used as input to the agent program of the fault diagnosis service to obtain a second diagnostic result. The fault diagnosis service can dynamically learn the solution to the fault, thereby improving the efficiency of fault diagnosis for the CI / CD pipeline in subsequent use.

[0088] S130 , in a case where at least one first diagnosis result can be used as a solution to the fault, output a first diagnosis result.

[0089] In some embodiments of the aforementioned vulnerability interception scenario, "change the version of software package A to 5.5 and rebuild" is used as the first diagnostic result, and the first diagnostic result can be used as a solution to the fault. After the fault diagnosis service outputs the first diagnostic result, the automatic build service of the CI / CD pipeline can complete the solution that can be automatically executed without human intervention, or wait for the developer to complete the steps that require human intervention and then perform appropriate post-processing. For example, in the aforementioned vulnerability interception scenario, when the API agent corresponding to SOP node #102 finds software package A with version 5.5 in the software package repository, and this version of software package A has no vulnerabilities, the automatic build service can call the API agent to change the version of software package A in the build instruction to 5.5 and re-execute the build task.

[0090] S140, when the first diagnostic result cannot be used as a solution to the fault, obtain a second diagnostic result based on at least one first diagnostic result and the RAG agent and / or LLM agent of the fault diagnosis service, and use the second diagnostic result as a solution to the fault; output the second diagnostic result.

[0091] In some embodiments of the aforementioned vulnerability interception scenario, "version 5.5 of software package A has a similar vulnerability" is used as the first diagnostic result and cannot be used as a solution to the fault. The RAG agent can search for software packages with similar functions to software package A in the knowledge base according to the functions of the code to be built, and obtain the second diagnostic result "Switch to software package B with the same functions and similar code interface as software package A", and generate code with the same functions using software package B according to the functions of the existing code using software package A. It should be understood that the generated code snippet and the natural language description of the code can also be regarded as part of the second diagnostic result.

[0092] In some embodiments, at least one third SOP node is generated according to the category of the fault and the first problem description, wherein the third SOP node corresponds to the category of the fault and is used to obtain the second diagnostic result according to the first problem description. The specific method of generating the third SOP node and the method of storing the third SOP node in the SOP knowledge base are similar to the methods described in S210, S220 and S230, and are not repeated here.

[0093] In such an implementation, the RAG agent can dynamically learn solutions to faults based on existing knowledge and unlearned faults to improve the efficiency of fault diagnosis for the CI / CD pipeline in subsequent use.

[0094] The present application also provides a fault diagnosis device, such as Figure 4 As shown, including:

[0095] The acquisition module is used to obtain the category of the fault of the CI / CD pipeline and the first problem description of the fault, as well as multiple standard operation instruction SOP nodes of the fault diagnosis service, wherein the SOP node is used to obtain the diagnosis result according to the problem description.

[0096] A processing module is used to: determine at least one first diagnostic result according to a first problem description and at least one first SOP node corresponding to a category of a fault among a plurality of SOP nodes. If at least one first diagnostic result can be used as a solution to the fault, the first diagnostic result is output; or, if the first diagnostic result cannot be used as a solution to the fault, a second diagnostic result is obtained according to at least one first diagnostic result and a retrieval enhancement RAG agent and / or a large language model LLM agent of a fault diagnosis service, and the second diagnostic result is used as a solution to the fault; and the second diagnostic result is output.

[0097] The processing module and the acquisition module can be implemented by software or hardware. For example, the implementation of the processing module is described below by taking the processing module as an example. Similarly, the implementation of the acquisition module can refer to the implementation of the processing module.

[0098] As an example of a software functional unit, a processing module may include code running on a computing instance. Among them, the computing instance may include at least one of a physical host (computing device), a virtual machine, and a container. Furthermore, the above-mentioned computing instance may be one or more. For example, the processing module may include code running on multiple hosts / virtual machines / containers. It should be noted that the multiple hosts / virtual machines / containers used to run the code may be distributed in the same region or in different regions. Furthermore, the multiple hosts / virtual machines / containers used to run the code may be distributed in the same availability zone (AZ) or in different AZs, each AZ including one data center or multiple data centers with similar geographical locations. Among them, usually a region may include multiple AZs.

[0099] Similarly, multiple hosts / virtual machines / containers used to run the code can be distributed in the same virtual private cloud (VPC) or in multiple VPCs. Usually, a VPC is set up in a region. For cross-region communication between two VPCs in the same region and between VPCs in different regions, a communication gateway needs to be set up in each VPC to achieve interconnection between VPCs through the communication gateway.

[0100] As an example of a hardware functional unit, the processing module may include at least one computing device, such as a server, etc. Alternatively, the processing module may also be a device implemented using an application-specific integrated circuit (ASIC) or a programmable logic device (PLD). The PLD may be a complex programmable logical device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL) or any combination thereof.

[0101] The multiple computing devices included in the processing module can be distributed in the same region or in different regions. The multiple computing devices included in the processing module can be distributed in the same AZ or in different AZs. Similarly, the multiple computing devices included in the processing module can be distributed in the same VPC or in multiple VPCs. The multiple computing devices can be any combination of computing devices such as servers, ASICs, PLDs, CPLDs, FPGAs, and GALs.

[0102] The present application also provides a computing device 1200. Figure 5 As shown, the computing device 1200 includes: a bus 1202, a processor 1204, a memory 1206, and a communication interface 1208. The processor 1204, the memory 1206, and the communication interface 1208 communicate through the bus 1202. The computing device 1200 can be a server or a terminal device. It should be understood that the present application does not limit the number of processors and memories in the computing device 1200.

[0103] The bus 1202 may be a peripheral component interconnect (PCI) bus or an extended industry standard architecture (EISA) bus. The bus may be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 5 The bus 1202 may include a path for transmitting information between various components of the computing device 1200 (eg, the memory 1206, the processor 1204, and the communication interface 1208).

[0104] The processor 1204 may include any one or more of a central processing unit (CPU), a graphics processing unit (GPU), a microprocessor (MP), or a digital signal processor (DSP).

[0105] The memory 1206 may include a volatile memory, such as a random access memory (RAM). The processor 1204 may also include a non-volatile memory, such as a read-only memory (ROM), a flash memory, a hard disk drive (HDD), or a solid state drive (SSD).

[0106] The memory 1206 stores executable program codes, and the processor 1204 executes the executable program codes to respectively implement the functions of the aforementioned processing module and acquisition module, thereby implementing the fault diagnosis method. That is, the memory 1206 stores instructions for executing the fault diagnosis method.

[0107] The communication interface 1208 uses a transceiver module such as, but not limited to, a network interface card or a transceiver to implement communication between the computing device 1200 and other devices or communication networks.

[0108] The embodiment of the present application also provides a computing device cluster. The computing device cluster includes at least one computing device. The computing device can be a server, such as a central server, an edge server, or a local server in a local data center. In some embodiments, the computing device can also be a terminal device such as a desktop computer, a laptop computer, or a smart phone.

[0109] like Figure 6 As shown, the computing device cluster includes at least one computing device 1200. The memory 1206 in one or more computing devices 1200 in the computing device cluster may store the same instructions for executing the fault diagnosis method.

[0110] In some possible implementations, the memory 1206 of one or more computing devices 1200 in the computing device cluster may also store partial instructions for executing the fault diagnosis method. In other words, the combination of one or more computing devices 1200 may jointly execute instructions for executing the fault diagnosis method.

[0111] It should be noted that the memory 1206 in different computing devices 1200 in the computing device cluster can store different instructions, which are respectively used to execute part of the functions of the fault diagnosis device. That is, the instructions stored in the memory 1206 in different computing devices 1200 can implement the functions of one or more modules in the processing module and the acquisition module.

[0112] In some possible implementations, one or more computing devices in the computing device cluster may be connected via a network, which may be a wide area network or a local area network. Figure 7 A possible implementation is shown. Figure 7 As shown, two computing devices 1200A and 1200B are connected via a network. Specifically, they are connected to the network via a communication interface in each computing device. In this type of possible implementation, the memory 1206 in one or more computing devices 1200 in the computing device cluster may store the same instructions for executing the fault diagnosis method.

[0113] It should be understood that Figure 7 The functions of the computing device 1200A shown in FIG. 1200A may also be completed by multiple computing devices 1200. Similarly, the functions of the computing device 1200B may also be completed by multiple computing devices 1200.

[0114] The embodiment of the present application also provides a computer program product including instructions. The computer program product may be software or a program product including instructions that can be run on a computing device or stored in any available medium. When the computer program product is run on at least one computing device, the at least one computing device executes a fault diagnosis method.

[0115] The embodiment of the present application also provides a computer-readable storage medium. The computer-readable storage medium can be any available medium that can be stored by a computing device or a data storage device such as a data center containing one or more available media. The available medium can be a magnetic medium (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state hard disk). The computer-readable storage medium includes instructions that instruct the computing device to execute the fault diagnosis method.

[0116] The above is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any changes or substitutions that can be easily thought of by a person skilled in the art within the technical scope disclosed in the present application should be included in the protection scope of the present application. Therefore, the protection scope of the present application shall be based on the protection scope of the claims.

Claims

1. A fault diagnosis method, characterized in that: The method is applied to the fault diagnosis service of continuous integration and continuous delivery / deployment CI / CD pipeline, and the method includes: Obtaining a type of the fault of the CI / CD pipeline and a first problem description of the fault, and a plurality of standard operating instruction SOP nodes of the fault diagnosis service, wherein the SOP nodes are used to obtain a diagnosis result according to the problem description; Determine at least one first diagnosis result according to the first problem description and at least one first SOP node corresponding to the category of the fault among the plurality of SOP nodes; If at least one of the first diagnosis results can be used as a solution to the fault, output the first diagnosis result; or, In the case that the first diagnostic result cannot be used as a solution to the fault, a RAG agent and / or a large language model LLM agent is generated based on the at least one first diagnostic result and the retrieval enhancement of the fault diagnosis service to obtain a second diagnostic result, and the second diagnostic result is used as a solution to the fault; and the second diagnostic result is output.

2. The method according to claim 1, characterized in that The at least one first SOP node forms a SOP tree, and the determining at least one first diagnosis result according to the first problem description and at least one first SOP node corresponding to the category of the fault among the plurality of SOP nodes comprises: Obtaining a third diagnostic result according to a second SOP node in the SOP tree, wherein the third diagnostic result is used to determine the first diagnostic result, or the third diagnostic result is used as the first diagnostic result, and the second SOP node is a SOP node in the at least one first SOP node; According to the first condition, the second condition, the third condition and the third diagnostic result, determine the first diagnostic result and whether the first diagnostic result can be used as a solution to the fault, wherein the first condition is that the third diagnostic result can be used as a solution to the fault, the second condition is that the second SOP node is a leaf node of the SOP tree, and the third condition is that the third diagnostic result can be used as a problem description of a child node of the second SOP node.

3. The method according to claim 2, characterized in that If the first and second conditions are not met: In the case where the third condition is not satisfied, the third diagnosis result is used as the first diagnosis result, wherein the first diagnosis result cannot be used as a solution to the fault; or, When the third condition is met, obtaining the third diagnostic result includes: obtaining the third diagnostic result according to the child nodes of the second SOP node and the second problem description, wherein the diagnostic result output by the second SOP node is used as the second problem description, and after determining the second problem description, the child nodes of the second SOP node are used as the second SOP node.

4. The method according to claim 2 or 3, characterized in that: In a case where the first condition is not satisfied and the second condition is satisfied, the third diagnosis result is used as the first diagnosis result, wherein the first diagnosis result cannot be used as a solution to the fault.

5. The method according to any one of claims 2 to 4, characterized in that In a case where the first condition is satisfied, the third diagnosis result is used as the first diagnosis result, wherein the first diagnosis result can be used as a solution to the fault.

6. The method according to any one of claims 1 to 5, characterized in that The method further comprises: At least one third SOP node is generated according to the category of the fault and the first problem description, wherein the third SOP node corresponds to the category of the fault and is used to obtain the second diagnosis result according to the first problem description.

7. The method according to any one of claims 1 to 6, characterized in that The SOP node includes metadata of at least one knowledge unit corresponding to the same category as the SOP node, wherein the knowledge unit includes problem descriptions and solutions of historical failures of the CI / CD pipeline, and the SOP node is obtained based on the at least one knowledge unit.

8. The method according to claim 7, characterized in that The SOP node is used to obtain the diagnosis result according to the problem description, and the problem description and solution of at least one knowledge unit corresponding to the SOP node.

9. The method according to any one of claims 1 to 8, characterized in that The SOP node is used to obtain the diagnosis result according to the problem description and the agent program corresponding to the SOP node, wherein the agent program includes at least one of an application programming interface API agent program, an LLM agent program and a RAG agent program.

10. A fault diagnosis device, characterized in that: The device is applied to the fault diagnosis service of continuous integration and continuous delivery / deployment CI / CD pipeline, and the device includes: An acquisition module, used to acquire the category of the fault of the CI / CD pipeline and the first problem description of the fault, and a plurality of standard operation instruction SOP nodes of the fault diagnosis service, wherein the SOP nodes are used to obtain a diagnosis result according to the problem description; a processing module, configured to determine at least one first diagnosis result according to the first problem description and at least one first SOP node corresponding to the category of the fault among the plurality of SOP nodes; and output the first diagnosis result if at least one first diagnosis result can be used as a solution to the fault; or In the case that the first diagnostic result cannot be used as a solution to the fault, a RAG agent and / or a large language model LLM agent is generated based on the at least one first diagnostic result and the retrieval enhancement of the fault diagnosis service to obtain a second diagnostic result, and the second diagnostic result is used as a solution to the fault; and the second diagnostic result is output.

11. The device according to claim 10, characterized in that The at least one first SOP node forms a SOP tree, and the processing module is further used for: Obtaining a third diagnostic result according to a second SOP node in the SOP tree, wherein the third diagnostic result is used to determine the first diagnostic result, or the third diagnostic result is used as the first diagnostic result, and the second SOP node is a SOP node in the at least one first SOP node; According to the first condition, the second condition, the third condition and the third diagnostic result, determine the first diagnostic result and whether the first diagnostic result can be used as a solution to the fault, wherein the first condition is that the third diagnostic result can be used as a solution to the fault, the second condition is that the second SOP node is a leaf node of the SOP tree, and the third condition is that the third diagnostic result can be used as a problem description of a child node of the second SOP node.

12. The device according to claim 11, characterized in that If the first and second conditions are not met: In the case where the third condition is not satisfied, the third diagnosis result is used as the first diagnosis result, wherein the first diagnosis result cannot be used as a solution to the fault; or, When the third condition is met, the processing module is also used to: obtain the third diagnostic result based on the child nodes of the second SOP node and the second problem description, wherein the diagnostic result output by the second SOP node is used as the second problem description, and after determining the second problem description, the child nodes of the second SOP node are used as the second SOP node.

13. The device according to claim 11 or 12, characterized in that In a case where the first condition is not satisfied and the second condition is satisfied, the third diagnosis result is used as the first diagnosis result, wherein the first diagnosis result cannot be used as a solution to the fault.

14. The device according to any one of claims 11 to 13, characterized in that In a case where the first condition is satisfied, the third diagnosis result is used as the first diagnosis result, wherein the first diagnosis result can be used as a solution to the fault.

15. The device according to any one of claims 10 to 14, characterized in that The processing module is also used for: At least one third SOP node is generated according to the category of the fault and the first problem description, wherein the third SOP node corresponds to the category of the fault and is used to obtain the second diagnosis result according to the first problem description.

16. The device according to any one of claims 10 to 15, characterized in that The SOP node includes metadata of at least one knowledge unit corresponding to the same category as the SOP node, wherein the knowledge unit includes problem descriptions and solutions of historical failures of the CI / CD pipeline, and the SOP node is obtained based on the at least one knowledge unit.

17. The device according to claim 16, characterized in that The SOP node is used to obtain the diagnosis result according to the problem description, and the problem description and solution of at least one knowledge unit corresponding to the SOP node.

18. The device according to any one of claims 10 to 17, characterized in that The SOP node is used to obtain the diagnosis result according to the problem description and the agent program corresponding to the SOP node, wherein the agent program includes at least one of an application programming interface API agent program, an LLM agent program and a RAG agent program.

19. A computing device cluster, characterized in that: comprising at least one computing device, each computing device comprising a processor and a memory; The processor of the at least one computing device is configured to execute instructions stored in the memory of the at least one computing device, so that the computing device cluster executes the method according to any one of claims 1 to 9.

20. A computer program product comprising instructions, characterized in that When the instructions are executed by a computing device cluster, the computing device cluster is caused to perform the method according to any one of claims 1 to 9.

21. A computer-readable storage medium, characterized in that: The method comprises computer program instructions, and when the computer program instructions are executed by a computing device cluster, the computing device cluster performs the method according to any one of claims 1 to 9.

Citation Information

Patent Citations

  • Server fault diagnosis method and device, equipment and readable storage medium

    CN113886120A

  • GitLab assembly line publishing method, system, equipment and medium

    CN117742797A

  • Application functionality testing, resiliency testing, chaos testing, and performance testing in a single platform

    US11847046B1

Cited By

  • Intelligent problem checking method and device, computer equipment and storage medium

    CN120469847A

  • Intelligent problem troubleshooting method, device, computer equipment and storage medium

    CN120469847B

  • Fault diagnosis method and apparatus

    WO2026123725A1