Fault handling method and apparatus for continuous integration and continuous delivery pipeline

CN122838153APending Publication Date: 2026-09-29CHONGQING CHANGAN AUTOMOBILE CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611043471.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-14
Publication Date
2026-09-29

AI Technical Summary

Technical Problem

[0004]本申请提供了一种持续集成与持续交付流水线的故障处理方法及装置,以解决现有的CI/CD流水线的故障处理效率较低的问题

Benefits of technology

(1)本申请通过预先构建得到技能库,并在持续集成与持续交付流水线发生故障时,将预设大模型输出的第二分析结果与技能库中的各技能单元进行匹配,从而自动确定出当前故障对应的修复策略,并对当前故障进行自动修复,因而无需依赖人工来进行故障分析和故障修复,提高了CI/CD流水线的故障处理效率。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122838153A_ABST
    Figure CN122838153A_ABST
Patent Text Reader

Abstract

The application relates to a fault processing method and device of a continuous integration and continuous delivery pipeline. The method comprises the following steps: acquiring a historical fault log and a historical repair strategy corresponding to the continuous integration and continuous delivery pipeline when a historical fault occurs, inputting the historical fault log into a preset large model for analysis to obtain a first analysis result; correcting a first repair strategy based on the historical repair strategy, and verifying and packaging the corrected first repair strategy to obtain a skill library; in the case that a current fault of the pipeline is detected, acquiring a current fault log, inputting the current fault log into the preset large model for analysis to obtain a second analysis result; determining whether a target skill unit matched with the second analysis result exists in the skill library; in the case that the target skill unit exists in the skill library, executing a repair strategy corresponding to the target skill unit to repair the current fault. Thus, the CI / CD pipeline fault processing efficiency can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of software development technology, specifically to a fault handling method and apparatus for a continuous integration and continuous delivery pipeline. Background Technology

[0002] Continuous Integration and Continuous Delivery (CI / CD) pipelines are a collective term for the entire automated pipeline of software development, from code submission, compilation and building, automated testing to deployment and release. With the rapid iteration of business requirements, the continuous compression of product development cycles, and the increasing frequency of application releases, the use of CI / CD pipelines is becoming more and more frequent.

[0003] However, existing CI / CD pipelines typically rely on manual fault analysis and repair after a failure occurs, resulting in low fault handling efficiency. Therefore, improving the fault handling efficiency of CI / CD pipelines has become an urgent technical problem to be solved. Summary of the Invention

[0004] This application provides a fault handling method and apparatus for continuous integration and continuous delivery pipelines to solve the problem of low fault handling efficiency in existing CI / CD pipelines.

[0005] In a first aspect, this application provides a fault handling method for a continuous integration and continuous delivery pipeline, the method comprising: Obtain historical fault logs and historical repair strategies corresponding to historical faults in the continuous integration and continuous delivery pipeline, and input the historical fault logs into a preset large model for analysis to obtain a first analysis result, wherein the first analysis result includes a first repair strategy; The first repair strategy is modified based on the historical repair strategy, and the modified first repair strategy is verified and encapsulated to obtain a skill library. The skill library includes multiple skill units and metadata of the multiple skill units, and each skill unit includes a repair strategy for the corresponding fault. If a fault is detected in the pipeline, the current fault log is obtained and input into the preset large model for analysis to obtain a second analysis result. Determine whether a target skill unit exists in the skill base that matches the second analysis result; If the target skill unit is found to exist in the skill library, the repair strategy corresponding to the target skill unit is executed to repair the current fault.

[0006] Optionally, the second analysis result includes current fault description information, the current execution environment of the pipeline, and the current execution stage; the metadata includes historical fault description information, rule requirements, and skill status corresponding to each skill unit, wherein the rule requirements include at least execution environment requirements and execution stage requirements; Determining whether a target skill unit in the skill base matches the second analysis result includes: Calculate the text similarity score between the current fault description information and the historical fault description information corresponding to each skill unit in the skill library, and determine the skill unit whose text similarity score is greater than a first preset threshold as the initial skill unit; Based on the rule requirements and skill status corresponding to each initial skill unit, each initial skill unit is filtered to obtain candidate skill units. The execution environment of the candidate skill unit matches the current execution environment of the pipeline, the execution stage of the candidate skill unit matches the current execution stage of the pipeline, and the candidate skill unit is in a skill-enabled state. Calculate the comprehensive score of each candidate skill unit, and determine whether the highest comprehensive score is greater than a second preset threshold. If the highest score of the comprehensive score is greater than the second preset threshold, it is determined that the target skill unit exists in the skill library, wherein the target skill unit is the candidate skill unit corresponding to the highest score of the comprehensive score; If the highest score of the comprehensive score is less than or equal to the second preset threshold, it is determined that the target skill unit does not exist in the skill library.

[0007] Optionally, calculating the overall score for each of the candidate skill units includes: Obtain the degree of matching between the running scenario of each candidate skill unit and the current running scenario of the pipeline, and determine the running scenario matching score corresponding to each candidate skill unit based on the degree of matching. Obtain the historical self-healing success rate of each candidate skill unit, and determine the historical self-healing success rate score corresponding to each candidate skill unit based on the historical self-healing success rate. The text similarity score, the running scenario matching score, and the historical self-healing success rate score corresponding to each candidate skill unit are weighted and summed to obtain the comprehensive score corresponding to each candidate skill unit.

[0008] Optionally, the second analysis result may also include a second repair strategy; After determining whether a target skill unit matching the second analysis result exists in the skill base, the method further includes: If it is determined that the target skill unit does not exist in the skill library, the current fault description information and the second repair strategy are synchronized to the manual review node to review and verify the second repair strategy, and the verified second repair strategy is encapsulated into a new skill unit. The new skill unit is stored in the skill library.

[0009] Optionally, obtaining historical fault logs and historical repair strategies corresponding to historical faults in the continuous integration and continuous delivery pipeline includes: The preset data entry fields in the historical logs of the pipeline are detected, and the historical fault logs are extracted from the historical logs of the pipeline based on the preset data entry fields. Extract the historical repair strategies corresponding to each historical fault from the historical repair records associated with the pipeline.

[0010] Optionally, the step of inputting the historical fault logs into a preset large model for analysis to obtain a first analysis result includes: The historical fault logs are categorized and archived, and then the categorized and archived historical fault logs are cleaned. Feature extraction is performed on the cleaned historical fault logs to obtain key feature data for each historical fault; The key feature data of each historical fault are used as fault sample data, and each fault sample data and preset prompt words are sequentially input into the preset large model for analysis to obtain the first analysis result. The preset prompt words are used to guide the preset large model to analyze each fault sample data according to the preset prompt words. The preset prompt words include role setting, fault root cause analysis rules, fault classification rules, repair strategy generation rules and output format template.

[0011] Optionally, the step of modifying the first repair strategy based on the historical repair strategy, and verifying and encapsulating the modified first repair strategy to obtain a skill library includes: Based on the historical repair strategy, the first repair strategy is modified from preset dimensions, wherein the preset dimensions include at least one of the following: the accuracy of root cause analysis, the completeness of the repair strategy, the executability of the repair strategy, and the completeness of scenario constraints. The initial execution script in the revised first repair strategy is parsed and validated, and the repair logic in the revised first repair strategy is parsed and validated in sequence. After the verification is passed, the metadata and target execution scripts corresponding to each historical fault are generated. The metadata and target execution scripts corresponding to each historical fault are encapsulated to obtain the skill units corresponding to each historical fault. The skill library is constructed based on the skill units corresponding to each historical fault.

[0012] Secondly, this application also provides a fault handling apparatus for a continuous integration and continuous delivery pipeline, the apparatus comprising: The first acquisition and analysis module is used to acquire historical fault logs and historical repair strategies corresponding to historical faults in the continuous integration and continuous delivery pipeline, and input the historical fault logs into a preset large model for analysis to obtain a first analysis result, wherein the first analysis result includes a first repair strategy. The verification and encapsulation module is used to modify the first repair strategy based on the historical repair strategy, and to verify and encapsulate the modified first repair strategy to obtain a skill library, wherein the skill library includes multiple skill units and metadata of the multiple skill units, and each skill unit includes a repair strategy for the corresponding fault. The second acquisition and analysis module is used to acquire the current fault log when a fault is detected in the pipeline, and input the current fault log into the preset large model for analysis to obtain the second analysis result; The determination module is used to determine whether there is a target skill unit in the skill library that matches the second analysis result; The repair module is used to execute the repair strategy corresponding to the target skill unit when it is determined that the target skill unit exists in the skill library, so as to repair the current fault.

[0013] Thirdly, this application also provides an electronic device, including a processor, a communication interface, a memory, and a communication bus, wherein the processor, the communication interface, and the memory communicate with each other through the communication bus; Memory, used to store computer programs; The processor, when executing a program stored in memory, implements the fault handling method for the continuous integration and continuous delivery pipeline described in the first aspect.

[0014] Fourthly, this application also provides a computer-readable storage medium storing computer-executable instructions for performing the fault handling method for the continuous integration and continuous delivery pipeline described in the first aspect.

[0015] The beneficial effects of this application are: (1) This application obtains a skill library in advance, and when a failure occurs in the continuous integration and continuous delivery pipeline, it matches the second analysis result output by the preset large model with each skill unit in the skill library, thereby automatically determining the repair strategy corresponding to the current failure and automatically repairing the current failure. Therefore, it does not need to rely on manual fault analysis and fault repair, thus improving the fault handling efficiency of the CI / CD pipeline.

[0016] (2) This application modifies the first repair strategy output by the preset large model through the historical repair strategy to obtain the skill library, and determines the fault repair strategy based on the skill library. This can avoid the defects of the preset large model inference being unstable, prone to hallucination, uncontrollable self-healing execution risk, and insufficient reuse of historical experience, thereby improving the reliability of fault handling in the CI / CD pipeline. Attached Figure Description

[0017] Figure 1 A flowchart illustrating a fault handling method for a continuous integration and continuous delivery pipeline provided in this application embodiment; Figure 2 A schematic diagram illustrating the process of generating a skill library, provided for an embodiment of this application; Figure 3 A flowchart illustrating another fault handling method for a continuous integration and continuous delivery pipeline provided in this application embodiment; Figure 4 A schematic diagram of a fault handling device for a continuous integration and continuous delivery pipeline provided in this application embodiment; Figure 5 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation

[0018] The embodiments of this application will be described below with reference to the accompanying drawings and preferred embodiments. Those skilled in the art can easily understand other advantages and effects of this application from the content disclosed in this specification. This application can also be implemented or applied through other different specific embodiments, and various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of this application. It should be understood that the preferred embodiments are only for illustrating this application and are not intended to limit the scope of protection of this application.

[0019] It should be noted that the illustrations provided in the following embodiments are only schematic representations of the basic concept of this application. Therefore, the drawings only show the components related to this application and are not drawn according to the actual number, shape and size of the components in the actual implementation. In the actual implementation, the form, quantity and proportion of each component can be arbitrarily changed, and the layout of the components may also be more complex.

[0020] For ease of description, spatial relative terms may be used in the text to describe the relative position or movement of one element or feature relative to another element or feature, as shown in the figure. These relative terms include, for example, "inside," "outside," "middle," "outer," "below," "below," "above," "front," "back," etc. Such spatial relative terms are intended to include different orientations of the device in use or operation, other than those depicted in the figure. For example, if the device in the figure undergoes a positional flip, orientation change, or change of motion, these directional indications will change accordingly. For instance, an element described as "below other elements or features" or "below other elements or features" will subsequently be oriented "above other elements or features" or "above other elements or features." Therefore, the example term "below" can include both upper and lower orientations. The device may be otherwise oriented (rotated 90 degrees or in other directions), and the spatial relative descriptors used in the text will be interpreted accordingly.

[0021] To address the issue of low fault handling efficiency in existing CI / CD pipelines, this application provides a fault handling method and apparatus for continuous integration and continuous delivery pipelines, which can improve the fault handling efficiency of CI / CD pipelines.

[0022] See Figure 1 , Figure 1 This is a flowchart illustrating a fault handling method for a continuous integration and continuous delivery pipeline provided in an embodiment of this application. Figure 1 As shown, the fault handling method for this continuous integration and continuous delivery pipeline may include the following steps: Step S102: Obtain historical fault logs and historical repair strategies corresponding to historical faults in the continuous integration and continuous delivery pipeline, and input the historical fault logs into a preset large model for analysis to obtain the first analysis result, wherein the first analysis result includes the first repair strategy.

[0023] Specifically, the aforementioned continuous integration and continuous delivery pipeline is a collective term for the entire automated pipeline of software development, from code submission, compilation and building, automated testing to deployment and launch. The aforementioned historical fault logs and historical remediation strategies correspond one-to-one with historical faults occurring in the continuous integration and continuous delivery pipeline; that is, one historical fault corresponds to one historical fault log and one historical remediation strategy. The aforementioned historical fault logs refer to the log context content recorded when the corresponding historical fault occurred. The aforementioned historical remediation strategies refer to fault remediation strategies derived from human experience for the corresponding historical faults. The aforementioned preset large model can be any pre-trained large model, such as a Generative Pre-trained Transformer (GPT) model, a DeepSeek model, etc. The aforementioned first analysis result refers to the analysis result obtained after analyzing the historical fault logs using the preset large model. It may include the first remediation strategy, and of course, it may also include other information such as fault root cause analysis and fault type. This application embodiment does not specifically limit this.

[0024] Step S104: Modify the first repair strategy based on the historical repair strategy, and verify and encapsulate the modified first repair strategy to obtain a skill library. The skill library includes multiple skill units and metadata of multiple skill units, and each skill unit includes a repair strategy for the corresponding fault.

[0025] Because the pre-set large model may have defects such as unstable reasoning, susceptibility to hallucinations, uncontrollable risks in self-healing execution, and insufficient reuse of historical experience, the first repair strategy output by the pre-set large model can be modified based on historical repair strategies to obtain a modified first repair strategy. Then, the modified first repair strategy can be verified and encapsulated to obtain a skill library, which facilitates the rapid determination of a repair strategy based on this skill library in the event of subsequent failures.

[0026] Step S106: If a fault is detected in the pipeline, obtain the current fault log and input the current fault log into the preset large model for analysis to obtain the second analysis result.

[0027] Specifically, the aforementioned current fault log refers to the log context content that records the current fault. The aforementioned second analysis result refers to the analysis result obtained after analyzing the current fault log using a preset large model. It may include a second repair strategy, and of course, it may also include other information such as fault root cause analysis and fault type. This application embodiment does not specifically limit it.

[0028] Step S108: Determine whether there is a target skill unit in the skill library that matches the second analysis result.

[0029] Specifically, after determining the second analysis result, it can be determined whether there is a target skill unit in the skill library that matches the second analysis result. If there is a target skill unit in the skill library that matches the second analysis result, it means that the skill library has stored the repair strategy corresponding to the fault, and step S110 can be executed; if there is no target skill unit in the skill library that matches the second analysis result, it means that the skill library has not stored the repair strategy corresponding to the fault. At this time, the second analysis result can be manually reviewed and verified to obtain the repair strategy corresponding to the fault.

[0030] Step S110: If the target skill unit is found to exist in the skill library, execute the repair strategy corresponding to the target skill unit to repair the current fault.

[0031] Specifically, if a target skill unit that matches the second analysis result exists in the skill library, it means that the skill library has stored the repair strategy corresponding to the fault. At this time, the repair strategy corresponding to the target skill unit can be directly executed to repair the current fault.

[0032] In this way, a skill library can be pre-built, and when a failure occurs in the continuous integration and continuous delivery pipeline, the second analysis result output by the pre-built large model is matched with each skill unit in the skill library. This automatically determines the corresponding repair strategy for the current failure and automatically repairs it, thus eliminating the need for manual failure analysis and repair, and improving the failure handling efficiency of the CI / CD pipeline. Furthermore, by modifying the first repair strategy output by the pre-built large model using historical repair strategies to obtain the skill library, and relying on the skill library to determine the failure repair strategy, the shortcomings of the pre-built large model, such as unstable inference, susceptibility to illusions, uncontrollable risks in self-healing execution, and insufficient reuse of historical experience, can be avoided, thereby improving the reliability of failure handling in the CI / CD pipeline.

[0033] In an optional embodiment, the second analysis result includes current fault description information, the current execution environment of the pipeline, and the current execution stage; the metadata includes historical fault description information, rule requirements, and skill status corresponding to each skill unit, wherein the rule requirements include at least execution environment requirements and execution stage requirements; Step S108 above, determining whether there is a target skill unit in the skill base that matches the second analysis result, includes: Calculate the text similarity score between the current fault description information and the historical fault description information corresponding to each skill unit in the skill library, and determine the skill unit with a text similarity score greater than the first preset threshold as the initial skill unit; Based on the rule requirements and skill status corresponding to each initial skill unit, each initial skill unit is filtered to obtain candidate skill units. The execution environment of the candidate skill unit matches the current execution environment of the pipeline, the execution stage of the candidate skill unit matches the current execution stage of the pipeline, and the candidate skill unit is in the skill enabled state. Calculate the overall score for each candidate skill unit and determine whether the highest overall score is greater than the second preset threshold. If the highest score in the overall score is greater than the second preset threshold, it is determined that there is a target skill unit in the skill library, where the target skill unit is the candidate skill unit corresponding to the highest score in the overall score; If the highest score in the overall score is less than or equal to the second preset threshold, it is determined that there is no target skill unit in the skill library.

[0034] Specifically, the second analysis result mentioned above may include current fault description information, the current execution environment of the pipeline, and the current execution stage. The current fault description information can be understood as descriptive information related to fault location and root cause analysis of the current fault. The current execution environment of the pipeline can be understood as information such as the current hardware environment, network status, and software version. The current execution stage of the pipeline can be the code submission stage, compilation and build stage, automated testing stage, or deployment stage, etc.

[0035] When determining whether a target skill unit in the skill base matches the second analysis result, a preset similarity calculation formula can be used to calculate the text similarity score between the current fault description information in the second analysis result and the historical fault description information corresponding to each skill unit in the skill base. Skill units with text similarity scores greater than a first preset threshold are then identified as initial skill units. The preset similarity calculation formula can employ cosine similarity, Euclidean distance, or other similar methods. The first preset threshold can be set according to actual needs, such as 80% or 90%.

[0036] Next, based on the rule requirements (such as execution environment requirements, execution stage requirements, etc.) and skill status (such as skill enabled status, skill disabled status, etc.) corresponding to each initial skill unit, candidate skill units can be selected from each initial skill unit whose execution environment matches the current execution environment of the pipeline, whose execution stage matches the current execution stage of the pipeline, and whose skill status is skill enabled.

[0037] Next, the comprehensive score of each candidate skill unit can be calculated and sorted. Then, it is determined whether the highest comprehensive score is greater than a second preset threshold. This second preset threshold can be set according to actual needs, and this embodiment does not impose a specific limitation. If the highest comprehensive score is greater than the second preset threshold, it can be determined that the target skill unit exists in the skill library; if the highest comprehensive score is less than or equal to the second preset threshold, it can be determined that the target skill unit does not exist in the skill library.

[0038] In this way, based on the current fault description information and the historical fault description information corresponding to each skill unit in the skill library, the initial skill units with a high degree of similarity to the current fault description information can be identified. Then, each initial skill unit is filtered according to the rule requirements and skill status to obtain each candidate skill unit. Next, based on the comprehensive score of each candidate skill unit, it is determined whether there is a target skill unit in the skill library that matches the second analysis result, thereby providing an accurate and reliable basis for the subsequent repair of the current fault.

[0039] In an optional embodiment, the above steps, including calculating the comprehensive score of each candidate skill unit, include: Obtain the degree of matching between the operating scenario of each candidate skill unit and the current operating scenario of the pipeline, and determine the operating scenario matching score for each candidate skill unit based on the degree of matching. Obtain the historical self-healing success rate of each candidate skill unit, and determine the historical self-healing success rate score corresponding to each candidate skill unit based on the historical self-healing success rate. The text similarity score, running scenario matching score, and historical self-healing success rate score corresponding to each candidate skill unit are weighted and summed to obtain the comprehensive score corresponding to each candidate skill unit.

[0040] Specifically, when calculating the comprehensive score of each candidate skill unit, the degree of matching between the operating scenario of each candidate skill unit and the current operating scenario of the pipeline can be obtained, and the operating scenario matching score corresponding to each candidate skill unit can be determined based on the matching degree. Simultaneously, the historical self-healing success rate of each candidate skill unit can be obtained, and the historical self-healing success rate score corresponding to each candidate skill unit can be determined based on the historical self-healing success rate. Then, the text similarity score, operating scenario matching score, and historical self-healing success rate score corresponding to each candidate skill unit can be weighted and summed to obtain the comprehensive score corresponding to each candidate skill unit.

[0041] In this way, the comprehensive score of each candidate skill unit can be calculated from multiple dimensions such as text similarity, running scenario matching degree and historical self-healing success rate. Based on the comprehensive score, the most accurate and reliable repair strategy can be matched from the skill library, which will facilitate subsequent fault repair based on the repair strategy.

[0042] In an optional embodiment, the second analysis result further includes a second repair strategy; after step S108, determining whether there is a target skill unit in the skill base that matches the second analysis result, the method further includes: If it is determined that the target skill unit does not exist in the skill library, the current fault description information and the second repair strategy are synchronized to the manual review node to review and verify the second repair strategy, and the verified second repair strategy is encapsulated into a new skill unit. Store new skill units in the skill library.

[0043] Specifically, when it is determined that the target skill unit does not exist in the skill library, the current fault description information and the second repair strategy can be synchronized to the manual review node. This allows operations or development personnel to review and verify the second analysis results output by the preset large model, based on the actual troubleshooting results and the repair strategy. If the current fault is determined to be a newly added typical fault with reusability, the verified second repair strategy can be packaged into a new skill unit and added to the skill library, completing the continuous iteration, enrichment, and optimization of the skill library. If the current fault is determined to be an intermittent fault, without the need for reuse, or only requiring temporary handling, the verified second repair strategy can be directly fed back to the pipeline platform and the triggering person in a standardized result format.

[0044] This enables a self-healing closed loop throughout the entire process, including fault analysis, intelligent matching, automatic execution, manual review, and skill accumulation, continuously improving the self-service fault repair rate of the pipeline, shortening fault recovery time, and ensuring the efficient and stable operation of the CI / CD pipeline.

[0045] In an optional embodiment, step S102, obtaining historical fault logs and historical remediation strategies corresponding to historical faults in the continuous integration and continuous delivery pipeline, includes: The preset data entry fields in the historical logs of the pipeline are detected, and historical fault logs are extracted from the historical logs of the pipeline based on the preset data entry fields. Extract the historical repair strategies corresponding to each historical fault from the historical repair records related to the production line.

[0046] Specifically, when obtaining historical fault logs and historical remediation strategies corresponding to historical failures in continuous integration and continuous delivery pipelines, preset instrumentation fields (such as "stageinfo_start", "stageinfo_end", etc.) within the pipeline's historical logs can be detected. Based on these preset instrumentation fields, historical fault logs, such as key error segments and stage context information, can be extracted from the pipeline's historical logs. This avoids the problem of exceeding the token limit of the preset large model by inputting all raw logs.

[0047] At the same time, historical repair strategies corresponding to each historical fault can be extracted from historical repair records related to the pipeline (such as Git commit records, change records, and internal operation and maintenance manuals).

[0048] In this way, we can obtain historical fault logs and historical repair strategies corresponding to historical failures in the continuous integration and continuous delivery pipeline. This makes it convenient to input the historical fault logs into the preset large model for analysis, obtain the first analysis result, and modify the first repair strategy based on the historical repair strategy. The modified first repair strategy is then verified and encapsulated to obtain a skill library.

[0049] In an optional embodiment, step S102 above, inputting historical fault logs into a preset large model for analysis to obtain a first analysis result, includes: Historical fault logs are categorized and archived, and then cleaned after categorization and archiving. Feature extraction is performed on the cleaned historical fault logs to obtain key feature data for each historical fault; The key feature data of each historical fault are used as fault sample data. Each fault sample data and preset prompt words are sequentially input into the preset large model for analysis to obtain the first analysis result. The preset prompt words are used to guide the preset large model to analyze each fault sample data according to the preset prompt words. The preset prompt words include role settings, fault root cause analysis rules, fault classification rules, repair strategy generation rules and output format templates.

[0050] Specifically, historical fault logs can be categorized and archived according to dimensions such as project name, fault type, and build tools. After categorization and archiving, the historical fault logs are cleaned, including log deduplication, sensitive information anonymization, and redundant information filtering. Next, feature extraction can be performed on the cleaned historical fault logs to obtain key feature data for each historical fault, such as dependency versions, network policies, permission configurations, and environment variables. This key feature data is then formatted into structured data, ultimately forming high-quality fault sample data that can be directly input into a pre-set large-scale model for analysis.

[0051] Next, the fault sample data and preset prompts can be sequentially input into the preset large model for analysis. The preset prompts (Prompt) precisely guide the behavior of the preset large model, enabling it to complete the following steps in sequence: fault root cause analysis, fault classification (such as environmental anomalies, configuration errors, code defects, dependency conflicts, network problems, etc.), and remediation strategies (such as directly executable remediation commands, configuration modification snippets, and Gradle / Maven build script corrections, etc.). The final output is a standardized remediation strategy with a unified structure.

[0052] As an optional implementation, the preset prompt can be composed of role settings, root cause analysis rules, fault classification rules, repair strategy generation rules, and output format templates. For example, the preset prompt can be as follows: "You are a CI / CD pipeline operations expert, and you are only performing fault analysis on build logs. Strictly follow the order below to execute the tasks, and do not disrupt the steps:" Step 1: Read the fault sample data, accurately locate the error stage, exception stack, and error file, and complete the root cause analysis of the fault. Step 2: Categorize the faults according to the following categories: environmental anomalies, configuration errors, code defects, dependency conflicts, network problems, etc., determine the fault type, and cluster the current fault with similar historical faults. Step 3: Generate ready-to-run fix commands, configuration modification snippets, and Gradle or Maven build fix scripts.

[0053] The output must strictly include the following fixed fields; no new or deleted entries may be added: [Error Location]: Specify the location of the error, the construction stage, and the original error message; [Root Cause Analysis]: Extracting the essence of the problem; [Detailed Repair Operation Steps]: Write out the operation process step by step; [Effect Verification Method]: Provide execution instructions to verify whether the repair has taken effect; [Potential Risk Warning]: This section explains the compatibility issues that this modification may cause and provides a rollback solution.

[0054] The output format should be standardized, the solution should be feasible, and vague and general textual descriptions should be avoided. In this way, preset prompts can be used to generate a first analysis result that includes error location, root cause analysis, specific repair operation steps, effect verification methods, and potential risk warnings. This facilitates the subsequent modification of the first repair strategy based on historical repair strategies, and the verification and encapsulation of the modified first repair strategy to obtain a skill library.

[0055] In an optional embodiment, step S104 above, which modifies the first repair strategy based on historical repair strategies and verifies and encapsulates the modified first repair strategy to obtain a skill library, includes: Based on historical repair strategies, the first repair strategy is modified from preset dimensions, wherein the preset dimensions include at least one of the following: accuracy of root cause analysis, completeness of repair strategy, executability of repair strategy and completeness of scenario constraints. The initial execution script in the revised first repair strategy is parsed and validated, and the repair logic in the revised first repair strategy is parsed and validated in sequence. After the verification is passed, the metadata and target execution scripts corresponding to each historical fault are generated. The metadata and target execution scripts corresponding to each historical fault are encapsulated to obtain the skill units corresponding to each historical fault. A skill library is constructed based on the skill units corresponding to each historical fault.

[0056] Specifically, the first remediation strategy can be modified based on historical remediation strategies, considering preset dimensions. These preset dimensions include one or more of the following: accuracy of root cause analysis, completeness of the remediation strategy, executability of the remediation strategy, and completeness of scenario constraints. The accuracy of root cause analysis refers to comparing the root cause analysis obtained from the model with the manually determined actual root causes of the fault. If they are inconsistent, the root cause analysis obtained from the model is modified based on the manually determined actual root causes. The completeness of the remediation strategy refers to the key operational steps covered by the model's output remediation operations. If it is determined that the model's output remediation operations are missing key operational steps, the model's output remediation operations are modified by comparing them with the operational steps in historical remediation strategies. The executability of the remediation strategy refers to running the model's output remediation script and configuration snippets in a test pipeline environment to verify the legality of the command syntax and the fault remediation effect. If the command syntax is invalid or the fault remediation effect is poor, the remediation script from historical remediation strategies is used to correct it. The completeness of scenario constraints refers to verifying whether the remediation strategy differentiates between JDK versions, build tool versions, intranet network environments, and other usage prerequisites, thereby assessing whether the scenario constraints are complete.

[0057] Next, the skill generator can be activated to complete standardized encapsulation according to engineering specifications. Specifically, the skill generator first uses its built-in syntax parsing module to perform syntax validation, format standardization, and cross-environment adaptation on the initial execution script in the corrected first repair strategy. It completes environment declarations, permission assignments, and exception handling logic (such as try-catch, timeout retries, and failure rollback) to ensure the initial execution script can be executed seamlessly on different Jenkins nodes and environments. Then, the rule engine performs time-series parsing and validation on the repair logic in the corrected first repair strategy, breaking down the complex process into atomic operation units and clarifying the inputs, outputs, and dependencies of each unit. After successful validation, metadata (such as unique identifiers, fault tags, and rule requirements) and executable script files corresponding to each historical fault are generated. Then, the metadata and target execution scripts corresponding to each historical fault are encapsulated to obtain the skill units corresponding to each historical fault. Based on these skill units, a skill library is built.

[0058] In this way, a skill library can be built in advance, which makes it easier to match the corresponding skill units for each fault, i.e., the repair strategy, based on the skill library.

[0059] As an optional implementation, the generation process of this skill library is as follows: Figure 2 As shown, it may include the following steps: Step S202: Obtain the historical fault logs corresponding to the historical faults that occurred in the continuous integration and continuous delivery pipeline.

[0060] Step S204: Analyze the historical fault logs using a preset large model to obtain the first analysis result.

[0061] Step S206: Obtain historical repair strategies.

[0062] Step S208: Cross-compare the historical repair strategy with the first repair strategy.

[0063] Step S210: Based on the comparison results, generate a skill library.

[0064] In this way, proven and effective historical repair strategies can be transformed into reliable, standardized, and reusable knowledge and executable repair steps, effectively reducing the loss of expert experience caused by changes in business personnel and pipeline maintenance personnel. At the same time, after centralized analysis and structured governance, massive historical fault logs will form unique data assets for the enterprise, which can serve as authoritative reference and static decision support in subsequent skill matching and fault diagnosis processes, improving the accuracy and credibility of solution matching.

[0065] As an optional implementation, the fault handling process of the continuous integration and continuous delivery pipeline is as follows: Figure 3As shown, it may include the following steps: Step S302: Obtain the current fault log of the pipeline.

[0066] Step S304: Analyze the current fault log using a preset large model to obtain the second analysis result.

[0067] Step S306: Determine whether there is a target skill unit in the skill library that matches the second analysis result.

[0068] If the target skill unit exists in the skill library, proceed to step S308; if the target skill unit does not exist in the skill library, proceed to step S310.

[0069] Step S308: Execute the repair strategy corresponding to the target skill unit to repair the current fault.

[0070] Step S310: Synchronize the current fault description information and the second repair strategy to the manual review node to review and verify the second repair strategy, and encapsulate the verified second repair strategy into a new skill unit. Step S312: Store the new skill unit in the skill library.

[0071] Step S314: Execute the verified second repair strategy to repair the current fault.

[0072] In this way, when the pipeline fails, the current fault log of the pipeline can be obtained and sent to a pre-set large model for intelligent analysis. The model can accurately locate the root cause of the failure and generate a one-click repair solution that can be implemented. Then, the diagnostic results and the repair solution are matched with multiple strategies in the skill library to verify their effectiveness. The final verification conclusion and repair steps are then fed back to the pipeline triggering personnel. The triggering personnel can complete the correction themselves according to the standardized solution in the feedback, or hand it over to the pipeline operation and maintenance personnel for assistance, so as to achieve rapid fault handling and closed loop.

[0073] See Figure 4 , Figure 4 This is a schematic diagram of a fault handling device for a continuous integration and continuous delivery pipeline provided in an embodiment of this application. Figure 4 As shown, the fault handling apparatus 400 for the continuous integration and continuous delivery pipeline includes: The first acquisition and analysis module 402 is used to acquire historical fault logs and historical repair strategies corresponding to historical faults in the continuous integration and continuous delivery pipeline, and input the historical fault logs into a preset large model for analysis to obtain the first analysis result, wherein the first analysis result includes the first repair strategy. The verification and encapsulation module 404 is used to modify the first repair strategy based on the historical repair strategy, and to verify and encapsulate the modified first repair strategy to obtain a skill library. The skill library includes multiple skill units and metadata of multiple skill units, and each skill unit includes a repair strategy for the corresponding fault. The second acquisition and analysis module 406 is used to acquire the current fault log when a fault is detected in the pipeline, and input the current fault log into a preset large model for analysis to obtain the second analysis result; Module 408 is used to determine whether there is a target skill unit in the skill base that matches the second analysis result; Repair module 410 is used to execute the repair strategy corresponding to the target skill unit when it is determined that the target skill unit exists in the skill library, so as to repair the current fault.

[0074] Furthermore, the second analysis results include current fault description information, the current execution environment of the pipeline, and the current execution stage; metadata includes historical fault description information, rule requirements, and skill status corresponding to each skill unit, wherein the rule requirements include at least execution environment requirements and execution stage requirements; the determination module 408 includes: The calculation submodule is used to calculate the text similarity score between the current fault description information and the historical fault description information corresponding to each skill unit in the skill library, and to determine the skill unit with the text similarity score greater than the first preset threshold as the initial skill unit. The filtering submodule is used to filter each initial skill unit based on the rule requirements and skill status corresponding to each initial skill unit to obtain candidate skill units. The execution environment of the candidate skill unit matches the current execution environment of the pipeline, the execution stage of the candidate skill unit matches the current execution stage of the pipeline, and the candidate skill unit is in the skill enabled state. The judgment submodule is used to calculate the comprehensive score of each candidate skill unit and determine whether the highest comprehensive score is greater than the second preset threshold. The first determining submodule is used to determine that a target skill unit exists in the skill library when the highest score of the comprehensive score is greater than the second preset threshold. The target skill unit is the candidate skill unit corresponding to the highest score of the comprehensive score. The second determination submodule is used to determine that the target skill unit does not exist in the skill library when the highest score of the comprehensive score is less than or equal to the second preset threshold.

[0075] Furthermore, the judgment submodule includes: The first determining unit is used to obtain the degree of matching between the operating scenario of each candidate skill unit and the current operating scenario of the pipeline, and to determine the operating scenario matching score of each candidate skill unit based on the degree of matching. The second determining unit is used to obtain the historical self-healing success rate of each candidate skill unit, and to determine the historical self-healing success rate score corresponding to each candidate skill unit based on the historical self-healing success rate. The calculation unit is used to perform weighted summation of the text similarity score, running scenario matching score and historical self-healing success rate score corresponding to each candidate skill unit to obtain the comprehensive score corresponding to each candidate skill unit.

[0076] Furthermore, the second analysis results also include a second remediation strategy; the fault handling apparatus 400 for the continuous integration and continuous delivery pipeline also includes: The synchronization module is used to synchronize the current fault description information and the second repair strategy to the manual review node when it is determined that the target skill unit does not exist in the skill library, so as to review and verify the second repair strategy and encapsulate the verified second repair strategy into a new skill unit. Store new skill units in the skill library.

[0077] Furthermore, the first acquisition and analysis module 402 includes: The first extraction submodule is used to detect preset tracking fields in the historical logs of the pipeline, and extract historical fault logs from the historical logs of the pipeline based on the preset tracking fields. The second extraction submodule is used to extract the historical repair strategies corresponding to each historical fault from the historical repair records related to the pipeline.

[0078] Furthermore, the first acquisition and analysis module 402 also includes: The classification and archiving submodule is used to classify and archive historical fault logs, and to clean the classified and archived historical fault logs. The feature extraction submodule is used to extract features from the cleaned historical fault logs to obtain key feature data for each historical fault. The analysis submodule is used to take the key feature data of each historical fault as the fault sample data, and input each fault sample data and preset prompt words into the preset large model for analysis to obtain the first analysis result. The preset prompt words are used to guide the preset large model to analyze each fault sample data according to the preset prompt words. The preset prompt words include role settings, fault root cause analysis rules, fault classification rules, repair strategy generation rules and output format templates.

[0079] Furthermore, the verification and encapsulation module 404 includes: The correction submodule is used to correct the first repair strategy based on the historical repair strategy from preset dimensions, wherein the preset dimensions include at least one of the following: the accuracy of root cause analysis, the completeness of the repair strategy, the executability of the repair strategy, and the completeness of scenario constraints. The verification submodule is used to perform syntax parsing and verification on the initial execution script in the modified first repair strategy, and to perform timing parsing and verification on the repair logic in the modified first repair strategy. The generation submodule is used to generate metadata and target execution scripts for each historical fault after the verification is passed. The encapsulation submodule is used to encapsulate the metadata and target execution scripts corresponding to each historical fault to obtain the skill unit corresponding to each historical fault. A submodule is constructed to build a skill library based on the skill units corresponding to each historical fault.

[0080] It should be noted that the fault handling device 400 for the continuous integration and continuous delivery pipeline can implement the fault handling method for the continuous integration and continuous delivery pipeline provided in any of the aforementioned method embodiments, and can achieve the same technical effect, which will not be elaborated here.

[0081] like Figure 5 As shown, this application embodiment also provides an electronic device, including a processor 511, a communication interface 512, a memory 513 and a communication bus 514, wherein the processor 511, the communication interface 512 and the memory 513 communicate with each other through the communication bus 514. Memory 513 is used to store computer programs; In one embodiment of this application, when the processor 511 executes the program stored in the memory 513, it implements the fault handling method for the continuous integration and continuous delivery pipeline provided in any of the foregoing method embodiments.

[0082] In addition, embodiments of this application also provide a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the fault handling method for the continuous integration and continuous delivery pipeline provided in any of the foregoing method embodiments.

[0083] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.

[0084] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented using software plus a general-purpose hardware platform, or of course, using hardware. Based on this understanding, the above technical solutions, in essence or the parts that contribute to the related technology, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.

[0085] It should be understood that the terminology used herein is for the purpose of describing particular exemplary embodiments only and is not intended to be limiting. Unless the context clearly indicates otherwise, the singular forms “a,” “an,” and “described” as used herein may also include the plural forms. The terms “comprising,” “including,” “containing,” and “having” are inclusive and therefore indicate the presence of the stated features, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, elements, components, and / or combinations thereof. The method steps, processes, and operations described herein are not construed as requiring them to be performed in a particular order described or illustrated unless the order of performance is explicitly indicated. It should also be understood that additional or alternative steps may be used.

[0086] The above description is merely a specific embodiment of this application, enabling those skilled in the art to understand or implement this application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of this application. Therefore, this application is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features claimed herein.

Claims

1. A fault handling method for a continuous integration and continuous delivery pipeline, characterized in that, The method includes; Obtain historical fault logs and historical repair strategies corresponding to historical faults in the continuous integration and continuous delivery pipeline, and input the historical fault logs into a preset large model for analysis to obtain a first analysis result, wherein the first analysis result includes a first repair strategy; The first repair strategy is modified based on the historical repair strategy, and the modified first repair strategy is verified and encapsulated to obtain a skill library. The skill library includes multiple skill units and metadata of the multiple skill units, and each skill unit includes a repair strategy for the corresponding fault. If a fault is detected in the pipeline, the current fault log is obtained and input into the preset large model for analysis to obtain a second analysis result. Determine whether a target skill unit exists in the skill base that matches the second analysis result; If the target skill unit is found to exist in the skill library, the repair strategy corresponding to the target skill unit is executed to repair the current fault.

2. The method according to claim 1, characterized in that, The second analysis result includes current fault description information, the current execution environment of the pipeline, and the current execution stage; the metadata includes historical fault description information, rule requirements, and skill status corresponding to each skill unit, wherein the rule requirements include at least execution environment requirements and execution stage requirements; Determining whether a target skill unit in the skill base matches the second analysis result includes: Calculate the text similarity score between the current fault description information and the historical fault description information corresponding to each skill unit in the skill library, and determine the skill unit whose text similarity score is greater than a first preset threshold as the initial skill unit; Based on the rule requirements and skill status corresponding to each initial skill unit, each initial skill unit is filtered to obtain candidate skill units. The execution environment of the candidate skill unit matches the current execution environment of the pipeline, the execution stage of the candidate skill unit matches the current execution stage of the pipeline, and the candidate skill unit is in a skill-enabled state. Calculate the comprehensive score of each candidate skill unit, and determine whether the highest comprehensive score is greater than a second preset threshold. If the highest score of the comprehensive score is greater than the second preset threshold, it is determined that the target skill unit exists in the skill library, wherein the target skill unit is the candidate skill unit corresponding to the highest score of the comprehensive score; If the highest score of the comprehensive score is less than or equal to the second preset threshold, it is determined that the target skill unit does not exist in the skill library.

3. The method according to claim 2, characterized in that, The calculation of the overall score for each of the candidate skill units includes: Obtain the degree of matching between the running scenario of each candidate skill unit and the current running scenario of the pipeline, and determine the running scenario matching score corresponding to each candidate skill unit based on the degree of matching. Obtain the historical self-healing success rate of each candidate skill unit, and determine the historical self-healing success rate score corresponding to each candidate skill unit based on the historical self-healing success rate. The text similarity score, the running scenario matching score, and the historical self-healing success rate score corresponding to each candidate skill unit are weighted and summed to obtain the comprehensive score corresponding to each candidate skill unit.

4. The method according to claim 1, characterized in that, The second analysis results also include a second repair strategy; After determining whether a target skill unit matching the second analysis result exists in the skill base, the method further includes: If it is determined that the target skill unit does not exist in the skill library, the current fault description information and the second repair strategy are synchronized to the manual review node to review and verify the second repair strategy, and the verified second repair strategy is encapsulated into a new skill unit. The new skill unit is stored in the skill library.

5. The method according to claim 1, characterized in that, The acquisition of historical fault logs and historical remediation strategies corresponding to historical faults in the continuous integration and continuous delivery pipeline includes: The preset data entry fields in the historical logs of the pipeline are detected, and the historical fault logs are extracted from the historical logs of the pipeline based on the preset data entry fields. Extract the historical repair strategies corresponding to each historical fault from the historical repair records associated with the pipeline.

6. The method according to claim 1, characterized in that, The step of inputting the historical fault logs into a preset large model for analysis to obtain a first analysis result includes: The historical fault logs are classified and archived, and then the classified and archived historical fault logs are cleaned. Feature extraction is performed on the cleaned historical fault logs to obtain key feature data for each historical fault; The key feature data of each historical fault are used as fault sample data, and each fault sample data and preset prompt words are sequentially input into the preset large model for analysis to obtain the first analysis result. The preset prompt words are used to guide the preset large model to analyze each fault sample data according to the preset prompt words. The preset prompt words include role setting, fault root cause analysis rules, fault classification rules, repair strategy generation rules and output format template.

7. The method according to claim 1, characterized in that, The first repair strategy is modified based on the historical repair strategy, and the modified first repair strategy is verified and encapsulated to obtain a skill library, including: Based on the historical repair strategy, the first repair strategy is modified from preset dimensions, wherein the preset dimensions include at least one of the following: the accuracy of root cause analysis, the completeness of the repair strategy, the executability of the repair strategy, and the completeness of scenario constraints. The initial execution script in the revised first repair strategy is parsed and validated, and the repair logic in the revised first repair strategy is parsed and validated in sequence. After the verification is passed, the metadata and target execution scripts corresponding to each historical fault are generated. The metadata and target execution scripts corresponding to each historical fault are encapsulated to obtain the skill units corresponding to each historical fault. The skill library is constructed based on the skill units corresponding to each historical fault.

8. A fault handling device for a continuous integration and continuous delivery pipeline, characterized in that, The device includes; The first acquisition and analysis module is used to acquire historical fault logs and historical repair strategies corresponding to historical faults in the continuous integration and continuous delivery pipeline, and input the historical fault logs into a preset large model for analysis to obtain a first analysis result, wherein the first analysis result includes a first repair strategy. The verification and encapsulation module is used to modify the first repair strategy based on the historical repair strategy, and to verify and encapsulate the modified first repair strategy to obtain a skill library, wherein the skill library includes multiple skill units and metadata of the multiple skill units, and each skill unit includes a repair strategy for the corresponding fault. The second acquisition and analysis module is used to acquire the current fault log when a fault is detected in the pipeline, and input the current fault log into the preset large model for analysis to obtain the second analysis result; The determination module is used to determine whether there is a target skill unit in the skill library that matches the second analysis result; The repair module is used to execute the repair strategy corresponding to the target skill unit when it is determined that the target skill unit exists in the skill library, so as to repair the current fault.

9. An electronic device, characterized in that, It includes a processor, a communication interface, a memory, and a communication bus, wherein the processor, the communication interface, and the memory communicate with each other through the communication bus; Memory, used to store computer programs; A processor, when executing a program stored in memory, implements the fault handling method for a continuous integration and continuous delivery pipeline as described in any one of claims 1-7.

10. A computer-readable storage medium, characterized in that, The system stores computer-executable instructions for performing the fault handling method for the continuous integration and continuous delivery pipeline as described in any one of claims 1-7.