An artificial intelligence software development-based management method and system
Patent Information
- Application Number
- CN202610938670.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-26
- Publication Date
- 2026-08-28
AI Technical Summary
[0009] The beneficial effects of the technical solution provided by this invention include at least the following:
Smart Images

Figure CN122653674A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of artificial intelligence-assisted software engineering technology, and in particular to a management method and system based on artificial intelligence software development. Background Technology
[0002] As software development systems grow increasingly large in scale, microservices and distributed architectures are gradually becoming mainstream. However, data from different stages of software development is often stored in isolation, forming data silos. When defects occur, it is difficult to quickly trace the source and pinpoint the affected links. Furthermore, existing defect detection technologies are mostly limited to a single stage and lack the ability to analyze causal relationships across stages, causing many problems to only be exposed after deployment, resulting in high repair costs.
[0003] Existing methods for detecting code defects mostly rely on machine learning models for risk scoring, but the output is mostly probabilistic warnings, lacking interpretable causal paths. Developers find it difficult to understand the logic behind the warnings, resulting in a weak willingness to adopt them. At the same time, code review tools are often based on static rules and cannot provide targeted suggestions in conjunction with real-time development context, which limits their practicality. Especially in complex architectures, a single code change may trigger a chain of failures, and traditional methods are unable to quantify and assess the scope of their impact.
[0004] While existing anomaly monitoring tools can capture exceptions during code runtime, they cannot simulate fault propagation in advance during the development phase. Assessment relies on human experience, leading to delayed risk control and often underestimating the impact of faults.
[0005] Furthermore, existing quality tools are mostly based on fixed rules or periodically trained models, which cannot be dynamically optimized as projects evolve and team behaviors change. After long-term use, their accuracy and adaptability decline, making it difficult to continuously support rapid iterative software development processes. Therefore, there is an urgent need for a systematic solution that can integrate data from the entire development cycle and achieve intelligent early warning and real-time guidance.
[0006] Invention The purpose of this invention is to propose a management method and system based on artificial intelligence software development to solve the problems in the prior art.
[0007] To achieve the above objectives, the present invention adopts the following technical solution: a management method based on artificial intelligence software development, comprising the following steps: S1. Throughout the entire software development lifecycle, multi-dimensional data is collected synchronously and a software development knowledge graph is constructed. The multi-dimensional data includes code modification records, test execution logs, online performance indicators, and fault events. S2, based on the software development knowledge graph, a cross-stage defect association mining model is used to automatically identify and extract defect association patterns, and store the defect association patterns and their quantitative confidence in a dynamic rule base, including code modification records, the execution results of test cases associated with the code modification records, and the cross-stage causal chain between the three of the operation failures caused in actual operation. S3: When a developer performs any operation such as writing code, submitting code, or executing tests, the current operation context and its associated unique global identifier are captured in real time, matched with the defect association pattern, and intelligent guidance information is automatically generated based on the matching result. S4, record the developer's adoption behavior of the intelligent guidance information, and track the verification effect of the corresponding code changes in the subsequent testing and running stages based on the unique global identifier, and perform feedback optimization on the cross-stage defect association mining model.
[0008] A management system based on artificial intelligence software development for this method, the system comprising: The data acquisition and knowledge graph construction module is used to synchronously collect multi-dimensional data and construct a software development knowledge graph throughout the entire software development lifecycle. The multi-dimensional data includes code modification records, test execution logs, online performance indicators, and fault events. The cross-stage defect association mining module is used to automatically identify and extract defect association patterns based on the software development knowledge graph through a cross-stage defect association mining model, and store the defect association patterns and their quantitative confidence in a dynamic rule base, including code modification records, execution results of test cases associated with the code modification records, and cross-stage causal chains between the three factors: operation failures caused in actual operation. The intelligent guidance generation and real-time push module is used to capture the current operation context and its associated unique global identifier in real time when the developer performs any operation such as code writing, code submission or test execution, match it with the defect association pattern, and automatically generate intelligent guidance information based on the matching result; The feedback-driven model optimization module is used to record the developer's adoption behavior of the intelligent guidance information, and to track the verification effect of the corresponding code changes in the subsequent testing and running phases based on the unique global identifier, so as to perform feedback optimization on the cross-stage defect association mining model.
[0009] The beneficial effects of the technical solution provided by this invention include at least the following: This invention constructs a knowledge graph covering the entire software development lifecycle and connects data from each stage based on a unique global identifier, enabling end-to-end visual tracing of code defects. This significantly improves the efficiency of root cause localization, reduces maintenance costs, and provides a structured data foundation for subsequent intelligent analysis.
[0010] This invention integrates temporal causal graphs and counterfactual reasoning techniques to extract interpretable defect association rules from historical data. This makes the defect association pattern information have causal logic, enhances developers' trust, and improves the accuracy of early warnings and subsequent intelligent guidance.
[0011] The system of this invention captures the operational context of developers in real time when they are performing development tasks, and dynamically generates intelligent guidance information by intelligently matching historical defect association patterns. This guidance information is used to alert risks, recommend specific modification solutions, and supplement test cases, thereby achieving real-time intervention, intercepting defects in the development stage, and reducing the possibility of online operation failures.
[0012] The system of this invention continuously collects developers' adoption behavior and subsequent verification effects of intelligent guidance information through a dual-channel feedback mechanism, which drives the co-evolution of models and strategies. Combined with the quantitative assessment of the impact on project modules, the system can effectively adapt to project changes and ultimately form a closed-loop adaptive intelligent quality control system, continuously improving software delivery quality and development efficiency. Attached Figure Description
[0013] To more clearly illustrate the technical solutions and advantages in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0014] Figure 1 A flowchart of the method provided in an embodiment of the present invention; Figure 2 This is a system structure diagram provided for an embodiment of the present invention. Detailed Implementation
[0015] To further illustrate the technical means and effects adopted by the present invention to achieve its intended purpose, the following, in conjunction with the accompanying drawings and preferred embodiments, details the specific implementation, structure, features, and effects of a management method and system based on artificial intelligence software development proposed according to the present invention. In the following description, different "one embodiment" or "another embodiment" do not necessarily refer to the same embodiment. Furthermore, specific features, structures, or characteristics in one or more embodiments can be combined in any suitable form.
[0016] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains.
[0017] The following examples are for illustrative purposes only and are not intended to limit the scope of the invention.
[0018] The following description, in conjunction with the accompanying drawings, details a specific solution for a management method and system based on artificial intelligence software development provided by this invention.
[0019] Please see Figure 1 The diagram illustrates a method flowchart of a management method based on artificial intelligence software development according to an embodiment of the present invention. The method includes the following steps: S1 synchronously collects multi-dimensional data and constructs a software development knowledge graph throughout the entire software development lifecycle. The multi-dimensional data includes code modification records, test execution logs, online performance indicators, and fault events. S2, based on software development knowledge graph, automatically identifies and extracts defect association patterns through a cross-stage defect association mining model, and stores the defect association patterns and their quantitative confidence in a dynamic rule base, including code modification records, the execution results of test cases associated with code modification records, and the cross-stage causal chain between the three factors of operational failures caused in actual operation. S3 captures the current operation context and its associated unique global identifier in real time when the developer performs any operation such as writing code, submitting code, or executing tests. It matches this context with defect association patterns and automatically generates intelligent guidance information based on the matching results. S4 records developers' adoption of intelligent guidance information and tracks the verification effect of corresponding code changes in subsequent testing and operation phases based on a unique global identifier, providing feedback optimization for the cross-phase defect association mining model.
[0020] As one embodiment of the present invention, the steps of synchronously collecting multi-dimensional data and constructing a software development knowledge graph throughout the entire software development lifecycle include: Deploy lightweight data acquisition probes in the development environment, continuous integration pipeline, and runtime environment, and create and bind a unique global identifier for every code commit record throughout the software development lifecycle. During continuous integration, test execution logs and corresponding code commit records are captured in real time and associated with each other based on a unique global identifier. Test execution logs include test cases and test execution results. Collect online performance metrics and fault events using data acquisition probes in the operating environment; Using code commit records, test cases, test execution results, online performance metrics, and fault events as nodes, and the relationships between nodes as edges, a software development knowledge graph is constructed and stored in a graph database.
[0021] It should be noted that during the development phase, when a developer performs a code commit operation, the system automatically generates a unique global identifier for that commit (for example, based on a combination of the project identifier and a hash value), and binds this global identifier to the code commit record as the basis for tracing the code change throughout its entire lifecycle.
[0022] Lightweight data acquisition probes are deployed in the development environment, continuous integration (CI) pipeline, and production environment. Each probe collects corresponding data according to a predefined data pattern and must record its associated global identifier, including: Development environment probe: Captures code commit events through an integrated development environment plugin and attaches a global identifier; CI probes: During automated test execution, they capture information such as test case identifiers, execution results, and execution time, and associate them with corresponding global identifiers by constructing trigger information. Runtime environment probe: Collects service performance metrics and system failure events through the application performance monitoring agent, and traces back the runtime data to the corresponding global identifier based on the mapping relationship between deployment version and code commit.
[0023] The knowledge graph is constructed and stored using the aforementioned collected entities (code commit records, test cases, test execution results, online performance metrics, and fault events) as nodes, and the temporal, causal, and attribution relationships between entities as edges. The edge types of this graph include, but are not limited to: [commit_trigger_test], [test_association_performance metrics], [performance anomaly_caused_fault], etc. This graph is stored in a graph database, supporting cross-stage subgraph queries and path discovery based on global identifiers, providing a structured and traceable data foundation for subsequent defect association mining.
[0024] This step uses global identifiers to establish end-to-end data association from code submission to online failures, constructing a unified knowledge graph covering the development, testing, and operation phases. This integrates previously isolated phase data into a networked structure with causal and temporal relationships, providing necessary data support for subsequent cross-phase defect propagation analysis.
[0025] As one embodiment of the present invention, the step of automatically identifying and extracting defect association patterns based on a software development knowledge graph and through a cross-stage defect association mining model includes: Based on the software development knowledge graph, and based on the code commit records identified by unique global identifiers, we extract the corresponding cross-stage knowledge subgraphs. For the edges representing the relationships between events in the cross-stage knowledge subgraph, weights are assigned based on the temporal and causal correlation strength they reflect. The cross-stage knowledge subgraph is then reconstructed into a weighted cross-stage temporal causal graph. The weight assignment includes assigning temporal weights to edges representing the order of occurrence of forward temporal events and assigning causal inference weights to edges representing the reverse root cause tracing relationship.
[0026] It should be noted that, in this embodiment, based on the software development knowledge graph, the unique global identifier corresponding to any code submission record is used as the query starting point. A breadth-first traversal is performed in the graph database to extract cross-stage knowledge subgraphs covering the following key nodes: (1) The code commit history and its associated modified files; (2) All test cases triggered by code commit records and their corresponding execution results; (3) A sequence of online runtime performance metrics associated with the deployment version of the code commit record; (4) The record of the fault event that ultimately occurred (if it exists).
[0027] The cross-stage knowledge subgraph is structurally reconstructed and weighted to form a cross-stage temporal causal graph, including: Temporal weight assignment: For edges representing the forward temporal order of events... (e.g., "Submit → Test Execution", "Test Execution → Performance Collection"), its time sequence weights The calculation formula is as follows:
[0028] In the formula, The time difference between the occurrence time of node B and the occurrence time of node A (unit: h). The decay rate parameter can be calibrated based on the average interval of the event chain in historical data, for example, by taking... .
[0029] Causal inference weight assignment: For edges representing root cause tracing relationships... (e.g., "fault event → performance anomaly", "test failure → code submission"), its causal inference weight The conditional probability based on historical subgraph statistics is used for assignment, as shown in the following formula:
[0030] In the formula, probability This is estimated by statistically analyzing all subgraphs containing nodes of the same type of result in the historical knowledge graph and calculating the frequency of nodes containing the same type of cause. Laplace smoothing can be used to avoid the zero probability problem. For cases where there is both temporal order and causal tracing between the same pair of nodes, they are split into two directed edges to carry temporal weights and causal inference weights respectively.
[0031] In one embodiment of the present invention, after reconstructing the cross-stage knowledge subgraph into a weighted cross-stage temporal causal graph, the following steps are performed: On the cross-stage temporal cause-effect graph, a set of cross-stage meta-path templates representing typical defect propagation paths are predefined; The graph traversal algorithm automatically matches and filters out paths that meet the meta-path template and have a frequency higher than the first preset threshold in the cross-stage temporal cause-effect graph, and uses them as defect propagation path instances. Based on temporal convolutional networks and causal discovery algorithms, a machine learning model is constructed as a cross-stage defect association mining model; Using cross-stage temporal causal graphs and defect propagation path instances as training samples, a cross-stage defect association mining model is trained to extract potential causal rules among code changes, test status, and runtime failures.
[0032] It should be noted that, in this embodiment, based on knowledge from the software engineering domain, a set of cross-stage meta-path templates representing typical defect propagation patterns are predefined. Each template is a sequence of node and edge types, describing the abstract path of a defect penetrating from the development stage to the runtime stage. For example, a typical meta-path template can be defined in the following form: [Code Submission] → (Trigger) → [Test Failure] → (Association) → [Performance Metric Abnormality] → (Cause) → [Runtime Failure]; Metapath templates are stored in the metapath template library, can be initialized according to different project types, and can be dynamically expanded in subsequent use.
[0033] On a cross-stage temporal cause-effect graph, perform the following steps to discover specific instances of defect propagation: (1) Use graph traversal algorithms (such as depth-first path matching algorithms) to find all paths in the cross-stage temporal causal graph that are completely matched with any meta-path template in terms of node and edge types; (2) Count the frequency of each matching path in the historical graph, filter out the paths with a frequency higher than the first preset threshold (e.g., the number of occurrences in the statistical period is greater than 5 times), and mark them as defect propagation path instances. Each instance not only contains a node sequence, but also inherits the temporal weight and causal inference weight of each edge on its path.
[0034] Construct a cross-stage defect association mining model, which consists of two core components: (1) Temporal feature extractor: It is constructed based on a temporal convolutional network. Its input is the sequence of node attributes and edge weights extracted from instances along the defect propagation path, and its output is a high-dimensional temporal feature representation of the path. (2) Causal rule decoder: It is built based on causal discovery algorithms (such as structural equation modeling or fraction-based causal learning algorithms), and takes the aforementioned time series features as input to learn and output the potential causal dependencies between nodes and their strength. Each defect propagation path instance and its local cross-stage temporal causal graph substructure together constitute a set of training samples, where positive samples are verified defect propagation instances, and negative samples can be obtained by generating non-causal paths through random walks or by sampling from fault-free submission graphs. Based on the training samples, the model is trained to maximize the accuracy of identifying real defect propagation paths and to calibrate its output causal strength score with the existing causal inference weights in the path instances.
[0035] In one embodiment of the present invention, after the cross-stage defect correlation mining model is trained, the following steps are performed: Extract defect association patterns in IF-THEN form from potential causal rules; For each defect association pattern, the counterfactual reasoning method is used to calculate its causal strength score, which is used as the confidence level of the association represented by the potential causal rule. For each defect association pattern, generate a rule feature vector corresponding to its semantics and structure, and store it in a dynamic rule base.
[0036] It should be noted that in this embodiment, the latent causal rules learned by the trained cross-stage defect association mining model are extracted. To improve the interpretability and operability of the rules, each rule is converted into a defect association pattern in the form of "IF-THEN". The general structure of this pattern is as follows: IF <precondition set> THEN <expected consequences set> WITH <contextual constraints>; The precondition set describes the combination of code and test status that triggers the rule, such as: "Code modification involves the interface file of module X" AND "Associated integration test case Y failed to execute"; The expected consequences set is used to predict operational issues that may arise if no intervention is taken, such as: "may cause the API latency P95 metric for service Z to increase by more than 20%" OR "there is a high probability of triggering a specific anomaly type F"; Context constraints are used to limit the environment or architectural conditions under which a rule takes effect, such as: "Only true in the call chain of microservice A".
[0037] To assess the reliability of each "IF-THEN" pattern, a counterfactual reasoning method is used to calculate its causal strength score. The specific steps are as follows: (1) For a defect association pattern, if its structure is: IF C THEN D, then locate all original instances of the precondition C that satisfy the defect association pattern in the historical knowledge graph. (2) By using graph structure intervention technology, in the local subgraph corresponding to each instance above, the premise C is simulated to be removed or negated to obtain ¬C (for example, "test failed" is changed to "test passed"), while keeping other context nodes unchanged, to generate a counterfactual graph instance; (3) Using the trained cross-stage defect association mining model, reasoning is performed on the original instance and the counterfactual instance respectively to obtain the predicted probability difference of consequence D, as shown in the following formula: Causality strength score = P(D|C) - P(D|¬C) The average of the predicted probability differences calculated for all instances is taken as the final quantitative confidence level of the defect association pattern.
[0038] To support long-term rule management and efficient retrieval, two types of feature vectors are generated for each defect association pattern: (1) Rule Archive Vector: This vector serves as the complete digital archive of the rule, integrating the following three types of features: Structural features: Encode the topological structure of the “IF-THEN” defect association pattern in the cross-stage temporal cause-effect graph; Full semantic features: Using a natural language processing model, all textual descriptions (including premises and consequences) in the defect association pattern are encoded to generate semantic embedding vectors; Statistical characteristics: including the frequency of occurrence of the defect association pattern in historical data, and the calculated causal strength score; (2) Rule matching key vector: focuses only on the semantic and contextual information of the "IF" part (i.e., precondition) in the encoded defect association pattern, specifically including: Premise semantic encoding: Encodes key text (such as module name and interface name) in premise C using a semantic model shared with the dynamic context encoder; Prerequisite structure encoding: Extract the code change features involved in the prerequisite C (such as modifying file type or function signature pattern); Module context encoding: Extracts information about the project module or service associated with the defect association pattern; The rule archive vector and rule matching key vector, along with the text in "IF-THEN" format and its causal strength score, are stored together in the dynamic rule base. The rule matching key vector will be used specifically for subsequent real-time similarity retrieval.
[0039] In one embodiment of the present invention, the step of capturing the current operation context and its associated unique global identifier in real time when the developer performs any of the operations of writing code, submitting code, or executing tests, matching it with defect association patterns, and automatically generating intelligent guidance information based on the matching results includes: By using plugins in an integrated development environment or code management platform, the developer's current operation events can be captured in real time, and a unique global identifier bound to the current operation can be obtained. Based on a unique global identifier, the context feature set associated with the current operation is extracted in real time and encoded into a dynamic context vector. The context feature set includes at least the code syntax structure, historical fragments of modified code, and module information of the corresponding project. The dynamic context vector is matched with the pre-stored rule feature vectors in the dynamic rule base based on similarity. This matching process includes the following two levels of retrieval: First-level module retrieval: Based on project module information, a first subset of rules related to the current development module or change type is selected from the dynamic rule base; Secondary semantic retrieval: In the first rule subset, based on the semantic similarity of code syntax structure and change patterns, match to locate the defect association pattern most relevant to the current operation context.
[0040] It should be noted that in this embodiment, a client plugin deployed in an integrated development environment or code management platform monitors key developer operation events in real time, including but not limited to: file saving, code completion, executing local tests, and initiating code commits (Commit / Push). When a key operation event is captured, the plugin synchronously obtains a unique global identifier bound to the current code change. Using this global identifier as an index, it extracts a context feature set from the local cache or project knowledge graph in real time, including: Code syntax structure features: the types of key nodes in the abstract syntax tree of the currently edited or committed code, function / method signatures, and introduced dependencies, etc. Characteristics of historical code changes: Recent code change history associated with this code change (such as the same files or modules modified in the last 5 commits); Project module topology characteristics: The location of the files involved in the current code change within the project module topology, their associated services, and their upstream and downstream call relationships; The above multi-dimensional features are input into a pre-trained lightweight encoder (such as a Transformer-based model), which outputs a fixed-dimensional dynamic context vector that encapsulates the semantics of the current operation and its structural significance in the project.
[0041] The two-level vector retrieval matching mechanism includes: First-level retrieval: Based on the topological features of the project modules in the dynamic context vector, calculate the cosine similarity between the features of the module and the feature vector of each rule in the dynamic rule base, and filter out the rules with a cosine similarity higher than the threshold α (e.g., α=0.7) to form the first rule subset. This step utilizes the principle of modular locality to quickly eliminate a large number of irrelevant rules, thereby reducing the search scope. Secondary retrieval: Within the first rule subset, more refined semantic matching is performed. The similarity between the overall semantic representation of the dynamic context vector and the rule feature vector of each rule in the first rule subset is calculated (e.g., using inner product or Euclidean distance). The K defect association patterns with the highest similarity (e.g., K=3) are selected as the matching results most relevant to the current operation context.
[0042] As one embodiment of the present invention, for a successfully matched defect association pattern, a potential fault propagation path is inferred based on the historical fault type and location associated with the cross-stage causal chain it represents, and the position of the current operating context in the project service architecture dependency graph. Assess the scope and severity of the impact of potential fault propagation paths on related services or downstream modules; Based on the causal strength score and impact severity level of the comprehensive defect correlation pattern, multiple candidate intelligent guidance messages are automatically generated, including risk warnings, code modification suggestions, supplementary test cases, or review point prompts; Based on the historical adoption preference data of individual developers, multiple candidate intelligent guidance messages are sorted, and the sorted results are pushed to the developers.
[0043] It should be noted that after matching a relevant defect association pattern, the system does not directly use the historical path of that defect association pattern, but instead performs the following dynamic deduction based on the latest status of the current project: Using the cross-stage causal chain represented by the matched defect association pattern as a template, and based on the mapping rules between node type and semantics, the generalized nodes in the cross-stage causal chain are replaced with specific entities in the current operation context, as shown in the following example: The generalized node 'interface file of module X' in the cross-stage causal chain is matched with the module identifier and interface type definition of the code file in the current context and replaced with the specific entity 'AuthController class of user_service'. Replace the generalized node 'database performance metrics' in the cross-stage causal chain with the specific entity 'QPS metrics of user_db' based on the data source mapping relationship configured in the project. Load the service architecture dependency graph of the current project. This graph represents the calls and dependencies between microservices, libraries, and modules. Locate the instantiated cross-stage causal chain in this graph.
[0044] Propagation simulation: Starting from the consequence node of the cross-stage causal chain (i.e. the predicted fault starting point), a propagation simulation with a finite step length is performed along the call chain on the service architecture dependency graph to generate one or more specific potential fault propagation paths and mark the affected related services and downstream modules. An impact assessment is conducted for each deduced potential failure propagation path, including: The number of affected related services and downstream modules and their hierarchical depth in the business architecture are statistically analyzed, and an impact scope index is calculated by weighted summation. The formula is as follows:
[0045] In the formula, , and These are the number of core services affected, the number of regular services affected, and the depth of the affected service level. , and These are the weighting coefficients for the three factors, which can be preset according to business importance, for example, by taking different values for each. , , .
[0046] Based on historical operational data such as average recovery time, customer impact, and financial loss level for similar faults, a preset severity level is assigned to the predicted fault type. Combined with an impact range index, and through a preset evaluation matrix, the overall severity level of the fault propagation path is ultimately determined. The severity level is preset based on historical fault data, for example: P0 level (fatal): Causes system unavailability or core data errors; Level P1 (Severe): Causes core functionality to be downgraded; Level P2 (General): Causes assistive function abnormalities; Level P3 (Minor): Causes minor user experience issues.
[0047] The evaluation matrix is a two-dimensional lookup table that takes the impact range index and the fault type severity level as input and outputs a comprehensive severity level. An example of an evaluation matrix is shown below:
[0048] The system automatically generates various types of candidate guidance information based on the following logic: Input: Matched defect association patterns and their causal strength scores, deduced potential fault propagation paths and their overall severity levels; Output: Directly inform developers of potential defects, predicted consequences, scope of impact, and severity level; If the causal strength score is high and the pattern contains a referable remedial solution, specific code modification suggestions will be generated. If the historical test coverage is insufficient, it is recommended to supplement with targeted test cases or scenarios; For changes with a wide impact, generate a checklist that should be the focus of code review; Subsequently, the generated candidate information is initially ranked by importance based on the product of (causal strength score × overall severity level). To improve the developer adoption rate, this invention introduces a personalized ranking mechanism, including: The system maintains a lightweight preference model for each developer, recording the developer's historical adoption (click, view, execute) and ignoring behaviors of guidance information of different types and severity. The candidate list initially sorted in the previous step is input into the preference model for re-sorting. For example, for developers who are used to looking at code suggestions first, the ranking of such suggestions is improved. Finally, the top-N most relevant intelligent guidance information (e.g., N=2) after reordering will be pushed to the developers through channels such as IDE notifications, emails, or chatbots.
[0049] As one embodiment of the present invention, the steps of recording the developer's adoption behavior of intelligent guidance information, tracking the verification effect of corresponding code changes in subsequent testing and running phases based on a unique global identifier, and optimizing the cross-phase defect correlation mining model include: Based on a unique global identifier, we collect adoption behavior data and verification effect data associated with the project to form a local feedback dataset; For scenarios where defects actually occur after developers ignore intelligent guidance information, a time decay weight is applied to the data corresponding to the scenario based on the interval between the time of defect occurrence and the time of intelligent guidance information push, and a weighted synthetic negative sample is constructed. Establish a parallel operation pattern discovery channel and a guidance strategy channel, and perform dual-channel parallel optimization, including: Pattern discovery channel: Optimize the model parameters of the cross-stage defect association mining model based on adoption behavior data and weighted synthetic negative samples; Guidance Strategy Channel: By using reinforcement learning-based optimization algorithms and verification data as reward signals, the intelligent guidance information generation and ranking strategy is optimized.
[0050] It should be noted that in this embodiment, the unique global identifier bound to the code commit record is used as the main thread to continuously track and collect the following two types of feedback data: Adoption behavior data: Records developers' real-time reactions to the push notification's intelligent guidance information, including actions such as viewing, adopting and performing modifications or ignoring, as well as the timestamps of the actions; Verification effect data: During the subsequent CI testing and online operation phase, monitor the code changes associated with this global identifier. If it causes new test failures or operational malfunctions, it will be recorded as a negative verification effect data and associated with the specific defect type. If it does not cause any problems within the preset observation period (such as a complete CI / CD pipeline and the subsequent 72 hours of online operation), it will be recorded as a positive verification effect data.
[0051] In this embodiment, based on the above feedback data, two independent optimization channels are established for parallel learning: (1) Pattern Discovery Channel: The optimization goal is to improve the accuracy of the cross-stage defect association mining model in identifying real defect propagation patterns. The construction of training samples includes: Positive samples: Data pairs in the adoption behavior data that are marked as adopted and modified, and whose associated validation effect data are positive; Negative samples: Data pairs in the adoption behavior data that are marked as ignored and whose associated validation effect data are negative; Weighted processing of synthetic negative samples: For the above strong negative samples, in order to more precisely measure their importance, a time decay weight is applied based on the time interval between the time of occurrence of the associated defect and the time of push of guidance information. The larger the weight, the more direct the causal relationship between the ignored behavior and the subsequent defect.
[0052] Optimization method: Use a training dataset consisting of positive samples and weighted negative samples to incrementally train or fine-tune the parameters of the cross-stage defect association mining model, so as to strengthen the model's memory of the verified defect association patterns and correct its confidence assessment of the effective warnings that are ignored.
[0053] (2) Guidance Strategy Channel: The optimization goal is to optimize the generation and sorting strategy of intelligent guidance information to improve the adoption rate and problem prevention success rate of developers. An optimization algorithm based on reinforcement learning is adopted. The generation and push process of intelligent guidance is modeled as a sequential decision problem, with the reward calculated using verification data as the core. If the guidance is adopted and the associated verification data is positive, a high positive reward will be given. If guidance is ignored but the correlation verification effect data is negative, a high negative reward (penalty) will be given. If guidance is adopted but the associated verification effect is negative, or if guidance is ignored but the associated verification effect is positive, a neutral reward close to zero is given. Using the current development context as the state, the type and order of the generated guidance information as the action, and the aforementioned rewards as the learning signal, the policy network of the reinforcement learning algorithm is iteratively updated to ultimately learn what kind of guidance should be provided in what context and how to order it to maximize long-term rewards.
[0054] As one embodiment of the present invention, the process of establishing a parallel-operating pattern discovery channel and a guidance strategy channel, and performing dual-channel parallel optimization, further includes the following steps: Structural interventions are performed on the graph structure features of cross-stage temporal causal graphs to generate a set of comparative samples that differ from the original graph structure; The original sample and the comparison sample are respectively input into the cross-stage defect association mining model to obtain the corresponding prediction results. By calculating the consistency measure between the two sets of prediction results, a causal robustness assessment signal is generated. The causal robustness assessment signal is used as a joint optimization objective, and is also used in the optimization process of the model parameters of the cross-stage defect association mining model in the pattern discovery channel, as well as the optimization process of the intelligent guidance information generation and ranking strategy in the guidance strategy channel.
[0055] It should be noted that in this embodiment, in each training iteration of the pattern discovery channel, a cross-stage temporal causal graph substructure is selected from the samples of the training batch as the original sample, and random edge perturbation is performed on the graph structure of the original sample: non-critical edges (such as edges with low temporal weights) in the original graph are randomly discarded with a certain probability, or virtual edges that do not change the causal main chain are added, so as to generate a set of comparison samples that have controllable differences in structure from the original graph, but still point to the same type of defect propagation pattern in semantics.
[0056] The original sample and the generated comparison sample are input into the current cross-stage defect association mining model in pairs. The probability distribution difference between the prediction results of the original sample and the prediction results of the comparison sample is calculated using an algorithm based on Jensen-Shannon divergence or mean square error. The negative value of the above difference value is used as the causal robustness assessment signal.
[0057] The causal robustness assessment signal L is used as an additional optimization objective and is injected into the optimization process of both the pattern discovery channel and the guidance strategy channel, including: In the pattern discovery channel, a robustness regularization term (-L) is added to the original loss function to encourage the cross-stage defect association mining model to maximize its robustness to reasonable perturbations of the graph structure while minimizing prediction error. In the guidance strategy channel: a robustness reward proportional to the causal robustness evaluation signal L is added to the reward function of the reinforcement learning algorithm, so that the decision-making in this channel not only pursues the adoption of short-term guidance or the avoidance of problems, but also tends to select rules based on highly robust causal inference to generate guidance, thereby improving the reliability of long-term decisions and the credibility of the system.
[0058] Please see Figure 2 The diagram illustrates a system architecture of a management system based on artificial intelligence software development according to an embodiment of the present invention. The system includes: The data acquisition and knowledge graph construction module is used to synchronously collect multi-dimensional data and construct a software development knowledge graph throughout the entire software development lifecycle. The multi-dimensional data includes code modification records, test execution logs, online performance indicators, and fault events. The cross-stage defect association mining module is used to automatically identify and extract defect association patterns based on the software development knowledge graph through a cross-stage defect association mining model, and store the defect association patterns and their quantitative confidence in a dynamic rule base, including the cross-stage causal chain between code modification records, the execution results of test cases associated with code modification records, and the operational failures caused in actual operation. The intelligent guidance generation and real-time push module is used to capture the current operation context and its associated unique global identifier in real time when the developer performs any operation such as code writing, code submission or test execution, match it with defect association patterns, and automatically generate intelligent guidance information based on the matching results; The feedback-driven model optimization module records developers' adoption of intelligent guidance information and tracks the verification effect of corresponding code changes in subsequent testing and operation phases based on a unique global identifier, thereby optimizing the cross-phase defect correlation mining model.
[0059] The above-described embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application, and should all be included within the protection scope of this application.
Claims
1. A management method based on artificial intelligence software development, characterized in that, The method includes: S1. Throughout the entire software development lifecycle, multi-dimensional data is collected synchronously and a software development knowledge graph is constructed. The multi-dimensional data includes code modification records, test execution logs, online performance indicators, and fault events. S2, based on the software development knowledge graph, a cross-stage defect association mining model is used to automatically identify and extract defect association patterns, and store the defect association patterns and their quantitative confidence in a dynamic rule base, including code modification records, the execution results of test cases associated with the code modification records, and the cross-stage causal chain between the three of the operation failures caused in actual operation. S3: When a developer performs any operation such as writing code, submitting code, or executing tests, the current operation context and its associated unique global identifier are captured in real time, matched with the defect association pattern, and intelligent guidance information is automatically generated based on the matching result. S4, record the developer's adoption behavior of the intelligent guidance information, and track the verification effect of the corresponding code changes in the subsequent testing and running stages based on the unique global identifier, and perform feedback optimization on the cross-stage defect association mining model.
2. The management method based on artificial intelligence software development according to claim 1, characterized in that: The steps for synchronously collecting multi-dimensional data and constructing a software development knowledge graph throughout the entire software development lifecycle include: Deploy lightweight data acquisition probes in the development environment, continuous integration pipeline, and runtime environment, and create and bind a unique global identifier for every code commit record throughout the software development lifecycle. During continuous integration, test execution logs and corresponding code commit records are captured in real time, and the two are associated based on the unique global identifier. The test execution logs include test cases and test execution results. Collect online performance metrics and fault events using data acquisition probes in the operating environment; Using code commit records, test cases, test execution results, online performance metrics, and fault events as nodes, and the relationships between these nodes as edges, a software development knowledge graph is constructed and stored in a graph database.
3. The management method based on artificial intelligence software development according to claims 1 and 2, characterized in that: The steps for automatically identifying and extracting defect association patterns based on the software development knowledge graph and through a cross-stage defect association mining model include: Based on the software development knowledge graph, and based on the code submission records identified by the unique global identifier, extract the corresponding cross-stage knowledge subgraphs; For the edges representing the relationships between events in the cross-stage knowledge subgraph, weights are assigned based on the temporal and causal correlation strength they reflect, and the cross-stage knowledge subgraph is reconstructed into a weighted cross-stage temporal causal graph. The weight assignment includes assigning temporal weights to edges representing the order of occurrence of forward temporal events and assigning causal inference weights to edges representing the reverse root cause tracing relationship.
4. The management method based on artificial intelligence software development according to claim 3, characterized in that: After reconstructing the cross-stage knowledge subgraph into a weighted cross-stage temporal causal graph, the following steps are performed: On the cross-stage temporal cause-effect graph, a set of cross-stage meta-path templates representing typical defect propagation paths are predefined; Using a graph traversal algorithm, paths that conform to the meta-path template and have a frequency higher than a first preset threshold are automatically matched and filtered in the cross-stage temporal cause-effect graph, and are used as defect propagation path instances. Based on temporal convolutional networks and causal discovery algorithms, a machine learning model is constructed as a cross-stage defect association mining model; Using the cross-stage temporal causal graph and the defect propagation path instances as training samples, the cross-stage defect association mining model is trained to extract potential causal rules among code changes, test status, and runtime failures.
5. The management method based on artificial intelligence software development according to claim 4, characterized in that: After the cross-stage defect correlation mining model is trained, the following steps are performed: Extract the defect association patterns in IF-THEN form from the potential causal rules; For each of the aforementioned defect association patterns, a counterfactual reasoning method is used to calculate its causal strength score, which serves as the confidence level of the association represented by the potential causal rule. For each defect association pattern, generate a rule feature vector corresponding to its semantics and structure, and store it in a dynamic rule base.
6. The management method based on artificial intelligence software development according to claim 1, characterized in that: The step of capturing the current operation context and its associated unique global identifier in real time when the developer performs any operation such as code writing, code submission, or test execution, matching it with the defect association pattern, and automatically generating intelligent guidance information based on the matching result includes: By using plugins in an integrated development environment or code management platform, the developer's current operation events can be captured in real time, and the unique global identifier bound to the current operation can be obtained. Based on the unique global identifier, the context feature set associated with the current operation is extracted in real time and encoded into a dynamic context vector. The context feature set includes at least the code syntax structure, historical fragments of modified code, and module information of the corresponding project. The dynamic context vector is matched with the pre-stored rule feature vectors in the dynamic rule base based on similarity. This matching process includes the following two-level retrieval: First-level module retrieval: Based on project module information, a first subset of rules related to the current development module or change type is selected from the dynamic rule base; Secondary semantic retrieval: In the first rule subset, matching is performed based on the semantic similarity of code syntax structure and change patterns to locate the defect association pattern most relevant to the current operation context.
7. The management method based on artificial intelligence software development according to claim 6, characterized in that: For successfully matched defect association patterns, the potential fault propagation path is inferred based on the historical fault types and locations associated with the cross-stage causal chains they represent, as well as the position of the current operating context in the project service architecture dependency graph. Assess the scope and severity of the impact of the potential fault propagation path on related services or downstream modules; By combining the causal strength score of the defect association pattern and the severity level of the impact, multiple candidate intelligent guidance messages are automatically generated, including risk warnings, code modification suggestions, supplementary test cases, or review point prompts. Based on the historical adoption preference data of individual developers, the multiple candidate intelligent guidance information are sorted, and the sorted results are pushed to the developers.
8. The management method based on artificial intelligence software development according to claim 1, characterized in that: The steps of recording developers' adoption of the intelligent guidance information, tracking the verification effect of corresponding code changes in subsequent testing and operation phases based on the unique global identifier, and optimizing the cross-phase defect association mining model include: Based on the unique global identifier, adoption behavior data and verification effect data associated with the project are collected to form a local feedback dataset; For scenarios where defects actually occur after developers ignore intelligent guidance information, a time decay weight is applied to the data corresponding to the scenario based on the interval between the time of defect occurrence and the time of intelligent guidance information push, and a weighted synthetic negative sample is constructed. Establish a parallel operation pattern discovery channel and a guidance strategy channel, and perform dual-channel parallel optimization, including: Pattern discovery channel: Based on adoption behavior data and weighted synthetic negative samples, optimize the model parameters of the cross-stage defect association mining model; Guidance Strategy Channel: Using a reinforcement learning-based optimization algorithm, the verification effect data is used as a reward signal to optimize the intelligent guidance information generation and ranking strategy.
9. The management method based on artificial intelligence software development according to claim 8, characterized in that: The process of establishing a parallel operation mode discovery channel and a guidance strategy channel for dual-channel parallel optimization also includes the following steps: Structural intervention is performed on the graph structure features of the cross-stage temporal causal graph to generate a set of comparative samples that differ from the original graph structure; The original sample and the comparison sample are respectively input into the cross-stage defect association mining model to obtain the corresponding prediction results. By calculating the consistency measure between the two sets of prediction results, a causal robustness assessment signal is generated. The causal robustness assessment signal is used as a joint optimization objective, and is also used in the optimization process of the model parameters of the cross-stage defect association mining model in the pattern discovery channel, as well as the optimization process of the intelligent guidance information generation and ranking strategy in the guidance strategy channel.
10. A management system based on artificial intelligence software development, characterized in that: The system includes: The data acquisition and knowledge graph construction module is used to synchronously collect multi-dimensional data and construct a software development knowledge graph throughout the entire software development lifecycle. The multi-dimensional data includes code modification records, test execution logs, online performance indicators, and fault events. The cross-stage defect association mining module is used to automatically identify and extract defect association patterns based on the software development knowledge graph through a cross-stage defect association mining model, and store the defect association patterns and their quantitative confidence in a dynamic rule base, including code modification records, execution results of test cases associated with the code modification records, and cross-stage causal chains between the three factors: operation failures caused in actual operation. The intelligent guidance generation and real-time push module is used to capture the current operation context and its associated unique global identifier in real time when the developer performs any operation such as code writing, code submission or test execution, match it with the defect association pattern, and automatically generate intelligent guidance information based on the matching result; The feedback-driven model optimization module is used to record the developer's adoption behavior of the intelligent guidance information, and to track the verification effect of the corresponding code changes in the subsequent testing and running phases based on the unique global identifier, so as to perform feedback optimization on the cross-stage defect association mining model.