Continuous integration unstable build detection method

CN122816637APending Publication Date: 2026-09-25SUZHOU UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610840843.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-11
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

[0005]针对现有技术中存在的不足,本发明提供了一种持续集成不稳定构建检测方法,用以解决GitHub Actions等开源环境下的不稳定构建检测问题

Benefits of technology

[0030](1)本发明提供了一种持续集成不稳定构建检测方法,仅依赖可公开获取的执行日志与流水线元数据,能够充分发挥语义理解与逻辑建模的协同效应,有效弥补单一特征源在处理复杂或新型非确定性失败时的信息缺失,从而在GitHub Actions等动态云环境中实现稳健、高效的不稳定构建检测;由此有效打破开发者过度依赖手动重跑来判别失败性质的被动现状,大幅减少不必要的重复执行操作,缩短开发团队的反馈延迟,显著节省昂贵的云端计算资源。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122816637A_ABST
    Figure CN122816637A_ABST
Patent Text Reader

Abstract

The application provides a continuous integration unstable build detection method, one aspect is based on execution log to retrieve similar samples, and outputs semantic feature confidence through a logistic regression model; another aspect is to extract structured context features of a target CI task, and outputs structural feature confidence through a classifier; finally, the two confidences are weighted and fused to obtain a comprehensive prediction probability to determine whether it is an unstable build. The application only relies on publicly available execution logs and pipeline metadata, fully utilizes the synergistic effect of semantic understanding and logical modeling, and effectively makes up for the information loss of a single feature source when dealing with complex or new non-deterministic failures.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of software engineering technology, specifically relating to a method for detecting unstable continuous integration builds. Background Technology

[0002] Continuous Integration (CI) is a crucial cornerstone of modern collaborative software development. It identifies potential defects earlier and accelerates the development cycle by automatically building code changes before integrating code into a central repository. Platforms such as GitHub Actions have become the most widely used and leading CI services in the industry due to their seamless integration with the open-source community and their generous free quotas.

[0003] While CI provides developers with crucial build feedback, build results are not always reliable. In practice, CI builds often fail intermittently due to unpredictable factors such as network instability and the instability of test cases themselves. When a build fails, developers typically need to pause their work to diagnose potential problems. To address this, many CI platforms (including GitHub Actions) offer rerun mechanisms, allowing the CI build to be re-executed. If the build result changes after a rerun without modifying any source code, this is called a flaky build. Flaky builds not only challenge the fundamental assumption that "build failure means code defects" and erode developer trust in CI systems, but also lead to significant additional troubleshooting time and wasted computing resources.

[0004] Currently, automated detection technologies for unstable build failures still have significant limitations. On the one hand, existing research on CI rerun characteristics mainly focuses on other platforms such as Travis CI or the OpenStack community, and does not specifically target the mainstream environment of GitHub Actions. On the other hand, most existing unstable build detection methods are primarily geared towards specific industrial-grade internal CI systems, and they often rely on features that are difficult to obtain in open-source environments. Therefore, there is an urgent need in this field for an automated detection method that can run in a typical open-source environment to help developers efficiently identify unstable build failures. Summary of the Invention

[0005] To address the shortcomings of existing technologies, this invention provides a method for detecting unstable builds in continuous integration, which solves the problem of detecting unstable builds in open-source environments such as GitHub Actions.

[0006] The present invention achieves the above-mentioned technical objectives through the following technical means.

[0007] A method for detecting unstable continuous integration builds:

[0008] Based on the execution logs, the K samples with the highest similarity to the target CI task are retrieved from the sample set; based on the labels and similarity of the retrieved samples, a logistic regression model is used to output the semantic feature confidence score. ;

[0009] Structured context features for the target CI task are extracted from the developer, code change, and project / context dimensions, respectively, and the confidence scores of these structured features are output through a classifier. ;

[0010] Will and Weighted fusion yields a comprehensive prediction probability. ;based on Determine whether the target CI task is an unstable build.

[0011] Furthermore, sentence vectors are obtained from the execution logs, and similar samples are retrieved by calculating the cosine similarity between the sentence vectors of the target CI task and each sample.

[0012] Furthermore, before converting to sentence vectors, the execution log is preprocessed, including filtering procedural outputs and masking dynamic variables.

[0013] Furthermore, the developer dimension extraction includes:

[0014] The number of code commits made by the committer in the current repository;

[0015] Does the submitter own a sufficient number of repositories that meet the requirements?

[0016] The submitter's Bayesian trust score is calculated as follows:

[0017]

[0018] In the formula, For Bayesian trust scores, The number of successful builds in the submitter's past history. This represents the total number of builds made by the submitter. Based on the past success rate of building the current repository, This is a smooth value.

[0019] Furthermore, the Bayesian trust score is divided into those from the past month and those from the past three months.

[0020] Furthermore, in terms of code change dimensions, the scale of changes and code entropy at the file and line levels are statistically analyzed, and the number of structural changes in import declarations, class and method signatures, and method bodies is extracted by comparing with the abstract syntax tree.

[0021] Code entropy is calculated as follows:

[0022]

[0023]

[0024] In the formula, For code entropy, For the first The number of modified files with different file extensions. This represents the total number of modified files. For the first The percentage of modified files with different file extensions.

[0025] The project and context dimensions extract information including: whether the environment configuration has changed, concurrent task load, and historical failure rate.

[0026] Furthermore, in the code change dimension, the change scale includes: the total number of changed lines in the source code files and the cumulative number of newly added lines of code in all files during this change.

[0027] Furthermore, in terms of project and context dimensions, the following are extracted: total number of lines of source code in the repository before this build, test code density ratio, total number of lines of test code in the repository before this build, number of test cases that passed in the target CI task, total number of test cases actually executed in the target CI task, and build failure rate of the current repository in the past 3 months.

[0028] Furthermore, the sample set is constructed based on previously failed CI tasks. By rerunning the CI task samples, if the rerun is successful, it is marked as an unstable construction; if the rerun fails multiple times, it is marked as a deterministic failure.

[0029] The beneficial effects of this invention are as follows:

[0030] (1) This invention provides a method for detecting unstable builds in continuous integration. It relies only on publicly available execution logs and pipeline metadata, which can give full play to the synergistic effect of semantic understanding and logical modeling. It can effectively make up for the lack of information when dealing with complex or novel nondeterministic failures by a single feature source, thereby achieving robust and efficient unstable build detection in dynamic cloud environments such as GitHub Actions. This effectively breaks the passive status quo where developers rely too much on manual reruns to determine the nature of failures, greatly reduces unnecessary repeated execution operations, shortens the feedback delay of the development team, and significantly saves expensive cloud computing resources.

[0031] (2) This invention combines an automated rerun mechanism to solve the problems of poor labeling quality and many false negative samples in existing datasets, and provides the industry with a highly reliable nondeterministic construction detection and training framework. Attached Figure Description

[0032] Figure 1 This is a schematic diagram of the framework of the detection method of the present invention;

[0033] Figure 2 This is a horizontal comparison of the detection results of the present invention and existing methods;

[0034] Figure 3 This is a comparison of the mutual information importance of multiple features extracted in this invention. Detailed Implementation

[0035] The embodiments of the present invention are described in detail below. The embodiments described below with reference to the accompanying drawings are exemplary and intended to explain the present invention, and should not be construed as limiting the present invention.

[0036] Reference Figure 1 As shown, the continuous integration unstable build detection method of the present invention includes the following parts:

[0037] I. Sample Set

[0038] A sample set is constructed by collecting samples from past failed CI tasks, including:

[0039] 1. Tags

[0040] For previously failed CI tasks, a rerun is performed using a rerun tool (such as CI-Rerunner). If the task status changes from "failed" to "successful", the task sample is marked as "unstable build". Conversely, if the task status remains "failed" after multiple reruns (up to 10 times in this embodiment), the task sample is marked as "deterministic failure".

[0041] 2. Semantic Vectors

[0042] We use the extraction method described below, "Semantic Log Features," to extract high-dimensional sentence vectors from the CI task.

[0043] Using the methods described above, a sample set is constructed based on the sentence vectors and labels of each CI task sample.

[0044] II. Feature Extraction

[0045] Feature extraction for the target CI task (i.e., the failed CI task to be detected as "unstable construction") is carried out in two parallel ways: "semantic log features" and "structured context features".

[0046] 1. Semantic log features

[0047] 1.1 Preprocess the raw execution logs of CI tasks, including:

[0048] 1) Filter redundant procedural outputs using regular expressions; the procedural outputs refer to the input logs generated during the construction and execution process, including but not limited to: thought chain text during large language model inference, unformatted temporary results, or printed outputs of intermediate variables.

[0049] 2) Use placeholders to mask the following dynamic variables, including: timestamp, memory address, URL (Uniform Resource Locator), hash value, port number, IP address, and file path.

[0050] 1.2. Using a sentence vector model (such as all-MiniLM-L6-v2), the preprocessed execution logs are converted into high-dimensional vectors.

[0051] 1.3 Based on high-dimensional vectors, cosine similarity is calculated to retrieve the K most similar samples (i.e., those with the highest cosine similarity) to the target CI task from the sample set. The labels of these K samples and their similarity to the target CI task are then input into a logistic regression model (which can be trained and optimized using the aforementioned sample set) to predict the probability that the target CI task belongs to an "unstable construction." The corresponding probability value is used as the "semantic feature confidence score" of the target CI task. .

[0052] 2. Structured Context Features

[0053] Deep features are extracted from the following three dimensions using relevant analysis tools (such as CI-Miner):

[0054] 2.1 Developer Perspective

[0055] Used to quantify the submitter's historical experience and the reliability of code submissions, effectively eliminating uncertainties caused by team personnel changes or novice operations; including the submitter's experience and Bayesian trust score for the target CI task, where:

[0056] Experience is quantified by two metrics: first, the number of code commits made by the committer in the current repository (i.e., the storage space where the code corresponding to the target CI task is located); and second, whether the committer has enough repositories that meet the specified requirements (specifically, 3 or more repositories with a workflow execution count greater than 1000 in this embodiment). This metric is a Boolean value, i.e., "yes" or "no".

[0057] Bayesian trust score, which estimates the developer's success rate using Bayesian smoothing, is calculated as follows:

[0058]

[0059] In the formula, For Bayesian trust scores, The number of successful builds in the submitter's past history. This represents the total number of builds made by the submitter. Based on the past success rate of building the current repository, For smoothing values, in this embodiment A value of 5 is used to account for situations where the submitter has a limited number of past samples. The fewer past samples a submitter has, the closer the result is to the repository average; conversely, the more past samples a submitter has, the closer the result is to the submitter's actual build success rate. In specific calculations, this can be further divided into calculating the submitter's Bayesian trust score for the most recent month (i.e., ...). and Statistical analysis of the Bayesian trust scores for the most recent month and the most recent three months (i.e., ... and Statistics for the past 3 months.

[0060] In addition to the three features mentioned above (number of submissions, whether the repository meets the specified requirements, and Bayesian trust score), the following features can also be extracted as supplementary features:

[0061] The total number of developers involved in the target CI task, the number of active developers in the past three months, the cumulative number of cross-project contributions by the submitter, whether the target CI task is the submitter's first triggered build, whether the submitter is a core member of the development team, and whether the submitter is the original code author.

[0062] 2.2, Code Change Dimension

[0063] This is used to accurately characterize the complexity and structural differences of the code modifications that triggered the current build from both macro-level version control and micro-level abstract syntax tree perspectives, including:

[0064] At the macro level, the extracted features include:

[0065] 1) The total number of Git commits included in the target CI task;

[0066] 2) Number of commits for non-code files (such as configuration files, documentation, etc.);

[0067] 3) Scale of changes, including the absolute scale of changes in the number of lines added, deleted, or modified in source code files and test code files, as well as the number of files added, deleted, or modified. Specifically:

[0068] 3.1) Total number of lines of change in the source code file (including additions, deletions, and modifications);

[0069] 3.2) Total number of lines of change in the test code file (including additions, deletions, and modifications);

[0070] 3.3) The number of new test cases added to the abstract syntax tree;

[0071] 3.4) The number of test cases deleted from the abstract syntax tree;

[0072] 3.5) The total number of lines of code added to all files during this change;

[0073] 3.6) The cumulative number of lines of code deleted across all files during this change;

[0074] 3.7) Number of new files added in this change;

[0075] 3.8) Number of files deleted in this change;

[0076] 3.9) Number of documents whose content was modified during this change;

[0077] 3.10) The number of file extensions (i.e., file names) types involved in this change (after deduplication);

[0078] 4) Submission intent classification tags for target CI tasks (e.g., fix, feature, refactor, etc.);

[0079] 5) Does this code change span multiple independent business modules? This characteristic is a discrete Boolean value;

[0080] 6) In addition to the feature values ​​that can be directly obtained above, to further quantify the dispersion of changes, this invention uses the Shannon entropy formula to calculate "code entropy" to characterize the distribution of changed lines of code in various files:

[0081]

[0082]

[0083] In the formula, For code entropy, For the first The number of modified files with different file extensions (e.g., how many files with the extension ".py" have been modified). This represents the total number of modified files. For the first The percentage of modified files with different file extensions.

[0084] At the micro level, by comparing the abstract syntax trees before and after modification, the total number of difference nodes between the source code and test code is precisely extracted. At the dependency declaration level, the number of added, deleted, and replaced import dependency nodes is counted. At the class and method level, it is distinguished into structural changes and internal logic changes, and the number of added or deleted classes and methods, the number of classes and methods with structural changes in external signatures or inheritance relationships, and the number of classes and methods with unchanged external signatures but substantial code modifications within their method bodies are counted. Specifically, this includes:

[0085] 7) The total size of the difference nodes before and after the source code abstract syntax tree modification;

[0086] 8) Test the total size of the difference nodes before and after the modification of the abstract syntax tree of the code;

[0087] 9) The number of newly added class declaration nodes in the abstract syntax tree;

[0088] 10) The number of class declaration nodes deleted from the abstract syntax tree;

[0089] 11) The number of classes whose class names or inheritance relationships have undergone structural changes in the abstract syntax tree;

[0090] 12) The number of classes in the abstract syntax tree whose external signatures remain unchanged but whose internal code has undergone substantial modifications;

[0091] 13) The number of newly added method declaration nodes in the abstract syntax tree;

[0092] 14) The number of method declaration nodes deleted from the abstract syntax tree;

[0093] 15) The number of methods whose method signatures (such as input parameters, return values, etc.) in the abstract syntax tree undergo structural changes;

[0094] 16) The number of methods in the abstract syntax tree whose method signatures remain unchanged but whose method bodies are modified;

[0095] 17) The number of newly added member variables (fields) within the class in the abstract syntax tree;

[0096] 18) The number of member variables (fields) deleted within a class in the abstract syntax tree;

[0097] 19) The number of member variables whose global variable declarations within a class have changed in the abstract syntax tree;

[0098] 20) The number of new import dependency nodes added to the file header;

[0099] 21) The number of import dependency nodes deleted from the file header;

[0100] 22) The number of import dependency nodes whose headers have been changed or replaced.

[0101] 2.3 Project and Context Dimension

[0102] This is used to capture the underlying environmental state, dynamic task load, and historical failure baseline during pipeline operation, converting nondeterministic environmental events into numerical variables that can be directly calculated by the classification model, including changes in monitoring environment configuration, concurrent task load, and historical failure rate. Among these:

[0103] Regarding the project baseline, we extract the historical build failure rate of the current repository over the past month and three months, the total number of lines of code before the build, the test code density ratio, the total number of external dependent components introduced, the number of comments associated with pull requests, and the description text complexity score.

[0104] In terms of environment configuration, key environment change behaviors are strictly quantified into discrete boolean values ​​to specifically detect and identify whether container orchestration files such as Dockerfile or docker-compose have been modified, whether the executor base image has been changed, whether custom Action scripts have been modified, and whether unverified third-party Action components have been used.

[0105] Regarding pipeline status and execution results, the size of the CI / CD workflow configuration file, the frequency of recent configuration modifications, and the number of concurrent tasks in the current system are extracted. At the same time, the cumulative time of the current build task, the total number of test cases that were executed and passed or failed are fully recorded. The build step type where the first exception occurred, the parsed specific failure reason classification label, and the cumulative number of alarm messages appearing in the log are also extracted from the build log. Discrete Boolean values ​​are used to clearly identify whether the current build is a retry behavior of the same user.

[0106] The extracted features are as follows:

[0107] 1) The build failure rate of the current repository over the past month;

[0108] 2) The build failure rate of the current repository over the past 3 months;

[0109] 3) The total number of lines of source code in the repository before this build;

[0110] 4) The total number of lines of test code in the repository before this build;

[0111] 5) Test code density ratio (i.e., the number of test code lines per thousand lines of source code);

[0112] 6) The total number of source code files in the current repository;

[0113] 7) The total number of document files in the current repository;

[0114] 8) The total number of configuration files in the current repository;

[0115] 9) Total number of other location-type files in the current repository;

[0116] 10) Total number of external dependency packages / components introduced;

[0117] 11) The number of lines of code changed in the dependency management configuration file;

[0118] 12) Size (in bytes) of the CI / CD workflow configuration file;

[0119] 13) The cumulative number of times the CI / CD configuration file has been modified recently (within the last month);

[0120] 14) The number of tasks executing in parallel in the system during the current pipeline operation;

[0121] 15) A discrete boolean value indicating whether the docker-compose orchestration file was modified in this change;

[0122] 16) A discrete boolean value indicating whether the Dockerfile environment configuration file has been modified in this change;

[0123] 17) Discrete Boolean value, indicating whether an unverified third-party Action component was used;

[0124] 18) Discrete Boolean value indicating whether a custom Action script or workflow file has been modified;

[0125] 19) A discrete Boolean value indicating whether the Action caching mechanism is enabled for the current build task;

[0126] 20) Discrete Boolean value, indicating whether there is cross-job build artifact sharing behavior;

[0127] 21) Discrete Boolean value, indicating whether the base environment image of the current executor (Runner) has changed;

[0128] 22) The hosting type of the executor machine (e.g., GitHub-hosted, Self-hosted);

[0129] 23) The type of operating system the executor runs on (e.g., Linux, Windows, macOS);

[0130] 24) A discrete Boolean value indicating whether external GitHub network resources were called during the build process;

[0131] 25) The cumulative time (in seconds) of the target CI task during execution;

[0132] 26) The total number of test cases actually executed in the target CI task;

[0133] 27) The number of test cases that passed execution in the target CI task;

[0134] 28) The number of test cases that failed to execute in the target CI task;

[0135] 29) The type of build step where the first exception occurred (e.g., environment initialization, dependency installation, script execution);

[0136] 30) Parse and extract specific failure reason type tags from the build logs;

[0137] 31) Calculate the cumulative number of warning messages appearing in the logs;

[0138] 32) The cumulative number of existing discussion comments under the associated pull request (PR);

[0139] 33) The complexity metric score of the associated PR description text;

[0140] 34) A discrete Boolean value that indicates whether the issue number is explicitly referenced in the commit message;

[0141] 35) Discrete Boolean value, indicating whether the target branch being constructed is the main branch;

[0142] 36) The final execution status of the previous historical build task in the same workflow;

[0143] 37) Discrete Boolean value, indicating whether the current file change affects the file that caused the previous build failure;

[0144] 38) Discrete Boolean value, indicating whether this change has affected a historically frequently erroring hot file;

[0145] 39) Event types that trigger the current build pipeline (such as push, pull_request, schedule);

[0146] 40) Construct the specific hour of the day corresponding to the trigger (value range 0-23);

[0147] 41) Construct the specific week number (Monday to Sunday) corresponding to the week when the trigger is executed;

[0148] 42) Discrete Boolean value, indicating whether the current build task is a retry behavior for the same user.

[0149] 2.4 Classification prediction based on structured context features

[0150] To optimize the feature space, the extracted structured context features are sorted by relevance and dimensionality reduced based on the mutual information criterion to remove redundant noise.

[0151] The optimized features are then input into a supervised classifier (such as Support Vector Machine (SVM) or Multilayer Perceptron (MLP), which can be trained and optimized using the aforementioned sample set) to obtain the probability that the target CI task belongs to "unstable construction". The corresponding probability value is used as the "structural feature confidence" of the target CI task. .

[0152] III. Fusion Output

[0153] Semantic feature confidence and structural feature confidence We perform weighted fusion to obtain the comprehensive prediction probability:

[0154]

[0155] In the formula, To comprehensively predict probabilities, These are the weighting coefficients.

[0156] when If the set decision threshold is exceeded, the target CI task is determined to be an unstable construction.

[0157] By leveraging the synergistic effect of semantic understanding and logical modeling, the above approach can effectively compensate for the lack of information when dealing with complex or novel nondeterministic failures from a single feature source, thereby enabling robust and efficient unstable build detection in dynamic cloud environments such as GitHub Actions.

[0158] IV. Testing and Verification

[0159] Forward chaining validation was performed on an open-source GitHub repository dataset, where the proposed multi-dimensional feature fusion model achieved a breakthrough improvement in overall detection accuracy. Figure 2 As shown in the comprehensive detection performance comparison chart, under the same test environment, the optimal fusion model of this invention, which uses Support Vector Machine (SVM) as the underlying classifier, achieves an average F1-score of 0.803, which is a relative improvement of up to 20.3% compared to the existing baseline method's 0.667. Especially in some highly complex open-source projects (such as Eclipse / XText), the advantages of this invention are more significant, thereby greatly reducing the false negatives and false positives caused by nondeterministic construction.

[0160] This invention's detection method balances low false positive rate and high recall rate, exhibiting strong adaptability to industrial scenarios. In actual CI / CD pipelines, when the focus is on reducing false positives to minimize developer interference, this invention, employing a Multilayer Perceptron (MLP), achieves a maximum average precision of 0.766 and an AUC score of 0.852. Conversely, when the scenario requires capturing unstable builds as comprehensively as possible, this invention, by employing a Random Forest (RF), achieves an average recall rate as high as 0.915.

[0161] This invention successfully overcomes the limitations of single error log features, quantifying and verifying the core value of structured context information. For example... Figure 3 As shown, this invention ranks the mutual information importance of multidimensional features extracted by CI-Miner in a typical large-scale open-source project application scenario. The results clearly show that developer-dimensional features (especially the recent Bayesian trust score of the committer) and initial code size features have a very strong statistical correlation with the occurrence of unstable builds, and their importance significantly surpasses traditional indicators such as the amount of code changes. The verification results of this typical scenario not only confirm that nondeterministic errors are closely related to code change logic and committer behavior, but also fundamentally prove the necessity and advancement of the "dual-path extraction and multimodal fusion" architecture of this invention in capturing deep-seated nondeterministic causes.

[0162] This invention is not limited to the above-described embodiments. Any obvious improvements, substitutions, or modifications that can be made by those skilled in the art without departing from the essence of this invention are within the scope of protection of this invention.

Claims

1. A method for detecting unstable continuous integration builds, characterized in that: Based on the execution log, the K samples with the highest similarity to the target CI task are retrieved from the sample set; Based on the labels and similarity of the retrieved samples, a logistic regression model is used to output the semantic feature confidence score. ; Feature extraction is performed on the target CI task from the perspectives of developer, code change, and project and context, and the classifier outputs the structural feature confidence scores. ; Will and Weighted fusion yields a comprehensive prediction probability. ;based on Determine whether the target CI task is an unstable build.

2. The continuous integration unstable build detection method according to claim 1, characterized in that: Sentence vectors are obtained from the execution logs. Similar samples are retrieved by calculating the cosine similarity between the sentence vectors of the target CI task and each sample.

3. The continuous integration unstable build detection method according to claim 2, characterized in that: Before converting to sentence vectors, the execution log is preprocessed, including filtering procedural outputs and masking dynamic variables.

4. The continuous integration unstable build detection method according to claim 1, characterized in that: Developer-level extraction includes: The number of code commits made by the committer in the current repository; Does the submitter own a sufficient number of repositories that meet the requirements? The submitter's Bayesian trust score is calculated as follows: In the formula, For Bayesian trust scores, The number of successful builds in the submitter's past history. This represents the total number of builds made by the submitter. Based on the past success rate of building the current repository, This is a smooth value.

5. The continuous integration unstable build detection method according to claim 4, characterized in that: The Bayesian trust scores are divided into those from the past month and those from the past three months.

6. The continuous integration unstable build detection method according to claim 1, characterized in that: The code change dimension is used to count the scale of changes and code entropy at the file and line levels, and to extract the number of structural changes in import declarations, class and method signatures, and method bodies by comparing with the abstract syntax tree. Code entropy is calculated as follows: In the formula, For code entropy, For the first The number of modified files with different file extensions. This represents the total number of modified files. For the first The percentage of modified files with different file extensions.

7. The continuous integration unstable build detection method according to claim 1, characterized in that: The project and context dimensions extract information including: whether the environment configuration has changed, concurrent task load, and historical failure rate.

8. The continuous integration unstable build detection method according to claim 6, characterized in that: In the code change dimension, the change scale includes: the total number of lines of change in the source code files and the cumulative number of new lines of code added to all files in this change.

9. The continuous integration unstable build detection method according to claim 1, characterized in that: The project and context dimensions extract the following: total number of lines of source code in the repository before this build, test code density ratio, total number of lines of test code in the repository before this build, number of test cases that passed in the target CI task, total number of test cases actually executed in the target CI task, and build failure rate of the current repository in the past 3 months.

10. The continuous integration unstable build detection method according to claim 1, characterized in that: The sample set is constructed based on previously failed CI tasks. By rerunning the CI task samples, if the rerun is successful, it is marked as an unstable construction; if the rerun fails multiple times, it is marked as a deterministic failure.