Multi-agent based software defect repair method and apparatus
Patent Information
- Application Number
- CN202611101901.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-23
- Publication Date
- 2026-09-29
AI Technical Summary
传统路由方式主要关注修复成功率,缺少对调用成本的显式建模,容易导致高成本派发;或采用简单后处理策略进行成本筛选,难以在成功率与成本之间实现稳定平衡
[0044]上述方案至少具有以下的有益效果:能够在保持高修复命中率的同时降低平均修复成本;采用issue-agent粒度成本建模,相较全局粗粒度成本具有更高决策精度;支持与多目标决策组合部署,满足实际预算约束场景需求。
Smart Images

Figure CN122837899A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of artificial intelligence, and in particular to a method and apparatus for repairing software defects based on multiple agents. Background Technology
[0002] Existing software defect repair systems typically configure multiple large-scale intelligent agents with varying capabilities and costs. Traditional routing methods primarily focus on repair success rates, lacking explicit modeling of call costs, which can easily lead to high-cost dispatches; or they employ simple post-processing strategies for cost filtering, making it difficult to achieve a stable balance between success rate and cost. Summary of the Invention
[0003] The following is an overview of the topics described in detail in this article.
[0004] The purpose of this application is to at least partially solve one of the technical problems existing in the related technologies. The embodiments of this application provide a software defect repair method and apparatus based on multi-agent systems.
[0005] An embodiment of the first aspect of this application provides a software defect repair method based on multi-agent systems, comprising:
[0006] Acquire a data set consisting of the software defect task (issue), the intelligent agent, the execution trajectory of the agent in fixing the issue, and the cost;
[0007] A heterogeneous graph-based routing model is constructed based on the dataset. The heterogeneous graph uses issue and agent as nodes, and a supervision edge is generated between the issue node and the agent node. Data on the agent fixing the issue is written into the supervision edge.
[0008] The routing model is trained, and the matching score between issue and agent is calculated based on the heterogeneous graph according to the first constraint and the second constraint. The loss function of the routing model is calculated based on the matching score. The parameters of the routing model are updated based on the loss function to obtain the target routing model. The first constraint is used to constrain the score of the repaired sample to be higher than the score of the unrepaired sample. The second constraint is used to constrain the score of the low-cost repaired sample corresponding to the same issue to be higher than the score of the high-cost repaired sample.
[0009] Input the issue to be routed and the candidate agent into the target routing model to obtain the optimal routing result between the issue to be routed and the candidate agent.
[0010] According to certain embodiments of the first aspect of this application, constructing a heterogeneous graph-based routing model based on the dataset includes:
[0011] The dataset is sequentially processed by field mapping, time alignment, dirty data removal, and primary key unification to obtain standardized intermediate tables, which include issue intermediate tables, event intermediate tables, trajectory intermediate tables, and cost intermediate tables.
[0012] According to certain embodiments of the first aspect of this application, constructing a heterogeneous graph-based routing model based on the dataset includes:
[0013] Based on the standardized intermediate table, issue text vectorization, agent capability profile encoding, action feature extraction, and code node feature extraction are performed to obtain a node type feature matrix, which includes an issue feature matrix, an agent feature matrix, an action feature matrix, and a code feature matrix.
[0014] According to certain embodiments of the first aspect of this application, constructing a heterogeneous graph-based routing model based on the dataset includes:
[0015] Based on the node type feature matrix and relation mapping table, a node set and an edge set are constructed to obtain a heterogeneous graph sample.
[0016] According to certain embodiments of the first aspect of this application, constructing a heterogeneous graph-based routing model based on the dataset includes:
[0017] When there are attempt records between an issue and an agent, an initial supervisory edge is created between the issue node and the agent node in the heterogeneous graph;
[0018] The initial monitoring edge is associated with the execution trajectory of the issue and the agent;
[0019] Write the repair label into the initial supervision edge; the repair label is used to record whether the agent successfully repaired the issue.
[0020] The cost is written into the initial supervision edge, and the cost includes the lexical cost, tool call cost, and retry cost for each round of invocation;
[0021] The set of supervised edges is obtained based on the modified initial supervised edges.
[0022] According to certain embodiments of the first aspect of this application, during the training of the routing model, supervision edges with missing trajectories or costs are removed, training samples are constructed by grouping supervision edges according to issue, and loss functions are not calculated for issues with no positive or negative samples.
[0023] According to certain embodiments of the first aspect of this application, the loss function is expressed as: L = L_A + λ_B * L_B;
[0024] In the formula, L is the total loss function, L_A is the first sub-loss function corresponding to the first constraint, L_B is the second sub-loss function corresponding to the second constraint, and λ_B is the weight hyperparameter corresponding to the preset second sub-loss function;
[0025] Where L_A = mean_{p∈P,n∈N}softplus(s_n-s_p);
[0026] When c_i <c_j,L_B=mean(softplus(s_j-s_i));
[0027] softplus(x)=log(1+exp(x)),x=s_n-s_p;
[0028] mean is the arithmetic mean, P is the set of positive samples of the issue, N is the set of negative samples of the issue, s_p is the matching score of positive sample p, s_n is the matching score of negative sample n, softplus is the loss contribution, x is the difference in matching scores, c_i is the cost of the i-th issue sample, c_j is the cost of the j-th issue sample, s_i is the matching score of the i-th issue sample, and s_j is the matching score of the j-th issue sample.
[0029] According to certain embodiments of the first aspect of this application, the loss function is expressed as: L = L_A + λ_B * L_B + λ_C * L_C;
[0030] In the formula, L is the total loss function, L_A is the first sub-loss function corresponding to the first constraint, L_B is the second sub-loss function corresponding to the second constraint, L_C is the third sub-loss function for suppressing low-cost negative samples, λ_B is the weight hyperparameter corresponding to the preset second sub-loss function, and λ_C is the weight hyperparameter corresponding to the preset third sub-loss function.
[0031] Where L_A = mean_{p∈P,n∈N}softplus(s_n-s_p);
[0032] When c_i <c_j,L_B=mean(softplus(s_j-s_i));
[0033] When c_n <c_p,L_C=mean(softplus(s_n−s_p));
[0034] softplus(x)=log(1+exp(x)),x=s_n−s_p;
[0035] mean is the arithmetic mean, P is the set of positive samples of the issue, N is the set of negative samples of the issue, s_p is the matching score of positive sample p, s_n is the matching score of negative sample n, softplus is the loss contribution, x is the difference in matching scores, c_i is the cost of the i-th issue sample, c_j is the cost of the j-th issue sample, s_i is the matching score of the i-th issue sample, s_j is the matching score of the j-th issue sample, c_p is the cost of negative sample p, and c_n is the cost of negative sample n.
[0036] According to certain embodiments of the first aspect of this application, the issue to be routed and the candidate agent are input into the target routing model to obtain the optimal routing result between the issue to be routed and the candidate agent, including:
[0037] From the set of candidate agents whose routing scores satisfy score>=best_score-margin, select the candidate agent with the lowest cost as the optimal routing result;
[0038] Wherein, best_score is the optimal routing score of the candidate agent, margin is a preset floating parameter, and margin is determined by the utility function of the validation set, the utility function being a weighted combination of hit rate and cost metrics.
[0039] An embodiment of the second aspect of this application provides a software defect repair device based on multiple agents, comprising:
[0040] The data input module is used to acquire a data set consisting of the software defect task to be repaired (issue), the intelligent agent, the execution trajectory of the agent in repairing the issue, and the cost.
[0041] The model building module is used to build a routing model based on a heterogeneous graph according to the dataset. The heterogeneous graph has issue and agent as nodes, and a supervision edge is generated between the issue node and the agent node. The data of the agent fixing the issue is written into the supervision edge.
[0042] The model training module is used to train the routing model, calculate the matching score between issue and agent based on the heterogeneous graph according to the first constraint and the second constraint, calculate the loss function of the routing model according to the matching score, and update the parameters of the routing model according to the loss function to obtain the target routing model. The first constraint is used to constrain the score of the repaired sample to be higher than the score of the unrepaired sample, and the second constraint is used to constrain the score of the low-cost repaired sample corresponding to the same issue to be higher than the score of the high-cost repaired sample.
[0043] The routing decision module is used to input the issue to be routed and the candidate agent into the target routing model to obtain the optimal routing result between the issue to be routed and the candidate agent.
[0044] The above-mentioned approach has at least the following beneficial effects: it can reduce the average repair cost while maintaining a high repair hit rate; it adopts issue-agent granular cost modeling, which has higher decision accuracy compared to global coarse-grained cost; and it supports deployment in combination with multi-objective decision-making to meet the needs of actual budget-constrained scenarios. Attached Figure Description
[0045] The accompanying drawings are used to provide a further understanding of the technical solutions of this application and constitute a part of the specification. They are used together with the embodiments of this application to explain the technical solutions of this application and do not constitute a limitation on the technical solutions of this application.
[0046] Figure 1 This is a flowchart illustrating the steps of a multi-agent-based software defect repair method.
[0047] Figure 2 This is a step diagram of the supervision edge annotation process;
[0048] Figure 3 This is a step-by-step diagram of the model training process;
[0049] Figure 4 It is a flowchart of the reasoning and decision-making process;
[0050] Figure 5 This is a schematic diagram of a six-dimensional profile of an agent's capabilities;
[0051] Figure 6 This is a schematic diagram of a heterogeneous graph;
[0052] Figure 7 This is a comparison chart of test results from multiple test sets. Detailed Implementation
[0053] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.
[0054] It should be noted that although functional modules are divided in the device schematic diagram and a logical order is shown in the flowchart, in some cases, the steps shown or described may be performed in a different order than the module division in the device or the order in the flowchart. The terms "first," "second," etc., in the specification, claims, or the aforementioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence.
[0055] The embodiments of this application provide a software defect repair method and apparatus based on multi-agent technology, which are applied to agent selection and cost optimization in automatic software defect repair tasks (such as SWE-Bench scenarios), and can reduce the average repair cost while maintaining a high repair hit rate.
[0056] The embodiments of this application will be further described below with reference to the accompanying drawings.
[0057] The software defect repair device based on multi-agent includes: a data input module, a model building module, a model training module, and a routing decision module.
[0058] The data input module is used to obtain a data set consisting of the software defect task to be repaired (issue), the agent, the execution trajectory of the agent in repairing the issue, and the cost.
[0059] The model building module is used to build a heterogeneous graph-based routing model based on the dataset. The heterogeneous graph uses issue and agent as nodes, and a supervision edge is generated between the issue node and the agent node. The data of the agent fixing the issue is written into the supervision edge.
[0060] The model training module is used to train the routing model. Based on the first and second constraints, it calculates the matching score between the issue and the agent according to the heterogeneous graph. It calculates the loss function of the routing model based on the matching score and updates the parameters of the routing model according to the loss function to obtain the target routing model. The first constraint is used to constrain the score of the repaired sample to be higher than the score of the unrepaired sample. The second constraint is used to constrain the score of the low-cost repaired sample corresponding to the same issue to be higher than the score of the high-cost repaired sample.
[0061] The routing decision module is used to input the issue to be routed and the candidate agent into the target routing model to obtain the optimal routing result between the issue to be routed and the candidate agent.
[0062] In addition, the software defect repair device can be further subdivided into the following modules according to function: data access module M1, feature engineering module M2, heterogeneous graph construction module M3, supervised edge annotation module M4, training sample arrangement module M5, model training module M6, routing inference module M7, multi-objective decision-making module M8, and online feedback module M9. The modules communicate with each other through structured interfaces.
[0063] Therefore, the multi-agent-based software defect repair device executes the following multi-agent-based software defect repair method.
[0064] Reference Figure 1A software defect repair method based on multi-agent systems includes the following steps:
[0065] Step S100: Obtain a data set consisting of the software defect task (issue), the intelligent agent, the execution trajectory of the agent in fixing the issue, and the cost;
[0066] Step S200: Construct a routing model based on a heterogeneous graph based on the dataset. The heterogeneous graph has issue and agent as nodes. A supervision edge is generated between the issue node and the agent node. The data of the agent fixing the issue is written into the supervision edge.
[0067] Step S300: Train the routing model. Based on the first and second constraints, calculate the matching score between the issue and the agent according to the heterogeneous graph. Calculate the loss function of the routing model based on the matching score. Update the parameters of the routing model according to the loss function to obtain the target routing model. The first constraint is used to constrain the score of the repaired sample to be higher than the score of the unrepaired sample. The second constraint is used to constrain the score of the low-cost repaired sample corresponding to the same issue to be higher than the score of the high-cost repaired sample.
[0068] Step S400: Input the issue to be routed and the candidate agent into the target routing model to obtain the optimal routing result between the issue to be routed and the candidate agent.
[0069] For step S100, obtain the software defect task issue, the intelligent agent, the execution trajectory and cost of the agent repairing the issue, and form a data set by combining the software defect task issue, the intelligent agent, the execution trajectory and cost of the agent repairing the issue.
[0070] Additionally, the dataset may include warehouse identifiers.
[0071] It should be noted that a software defect issue is an instance of a software problem to be fixed, which typically includes a problem description, the code repository to which it belongs, and testing or acceptance information. The original issue may also include a defect description, steps to reproduce the problem, and related test information.
[0072] Repository identifier, a field that uniquely identifies the source code repository to which an issue belongs, such as Git remote address, owner / repo name, or internal repository code; used to associate issues, code nodes, and historical traces.
[0073] The intelligent agent also includes agent history records, which record each agent's past task performance, capability tags, etc.
[0074] Execution trail: A structured record of the process generated by an agent when handling a specific issue. It may include at least the tools / commands invoked, file fragments edited, intermediate inference summaries, and number of retries; it is output by the runtime environment or log collection module.
[0075] Costs are recorded in cost logs and billing records aligned with execution timelines. These records include token usage, tool call counts, and additional overhead from retries, and are aggregated into edge_cost_usd at the issue-agent granularity.
[0076] For step S200, firstly, the data access module M1 performs field mapping, time alignment, dirty data removal, and primary key unification on the data set in sequence to obtain standardized intermediate tables. The standardized intermediate tables include issue intermediate tables, event intermediate tables, trajectory intermediate tables, and cost intermediate tables.
[0077] The input to the data access module M1 is a dataset; the output is a standardized intermediate table consisting of issue_table, agent_table, trace_table, and cost_table.
[0078] The minimum field convention for the intermediate table `issue_table` is: `issue_id`, `repo_id` (repository identifier), `title`, `body`, `created_at`, etc.
[0079] The minimum field convention for the agent_table intermediate table in Agnient is: agent_id, model_name, capability_tags, etc.
[0080] The minimum field convention for the intermediate trace table is: trace_id, issue_id, agent_id, step_index, action_type, payload, and timestamp.
[0081] The minimum field convention for the intermediate cost table, cost_table, is: cost_id, trace_id (or issue_id + agent_id), cost_type, amount_usd, and timestamp.
[0082] Through field mapping, time alignment, dirty data removal, and unified primary key processing, each standardized intermediate table has a unique primary key, clear foreign key relationships, and consistent field types and naming within the system.
[0083] Then, the feature engineering module M2 performs issue text vectorization, agent capability profile encoding, action feature extraction, and code node feature extraction based on the standardized intermediate table to obtain the node type feature matrix, which includes the issue feature matrix, agent feature matrix, action feature matrix, and code feature matrix.
[0084] The input to the feature engineering module M2 is the standardized intermediate table output by M1; the output is the node type feature matrix X_issue, X_agent, X_action, and X_code.
[0085] Issue text vectorization involves vectorizing issue data into sentence vectors, TF-IDF, or pre-trained language model embeddings to obtain issue features, which are then further used to form an issue feature matrix.
[0086] The agent capability profile encoding involves encoding the agent's capabilities into historical success rate, proficient languages, average cost, etc., to obtain agent capability features, which are then further formed into an agent feature matrix.
[0087] Specifically, refer to Figure 5 The problem text, gold standard patch, evaluation results (resolved tags), prediction patches (preds), and execution trajectories (trajs) are aligned and integrated. Based on the integrated data, six-dimensional capability feature engineering is performed to obtain C1 localization capability (Ioc_recall_mean), C2 edit hit rate (edit_gold_overlap_mean), C3 search efficiency (n_steps_median), C4 cost control (cost_usd_mean), C5 patch range control (over_edit_ratio_median), and C6 correctness (resolved_rate). Based on the capability profile generated from the six-dimensional feature output, a priori agent six-dimensional capability profile table is generated.
[0088] By extracting action features, we can obtain the frequency of action types, tool call distribution, etc., and further form an action feature matrix.
[0089] By extracting code features from code nodes, statistical and semantic features are extracted from the source code fragments or change diffs associated with code nodes, such as file path hashes, number of changed lines, identifiers of involved functions / classes, code embedding vectors, etc., and then a code feature matrix is formed.
[0090] Next, the heterogeneous graph construction module M3 constructs a set of nodes and a set of edges based on the node type feature matrix and the relation mapping table to obtain a heterogeneous graph sample.
[0091] The input to the heterogeneous graph construction module M3 is the node type feature matrix and relation mapping table output by M2, and the output is the heterogeneous graph sample G=(V,E), including edge index and node features.
[0092] Among them, reference Figure 6 Based on the node type feature matrix, issue nodes, agent nodes, action nodes, and code nodes are formed. Node features include structural features, capability features, and prior capabilities.
[0093] The edges in the edge set are the edges that connect the nodes. The fields of each edge must include at least attempted, performed, belongs_to, and resolved_by.
[0094] Finally, the supervised edge annotation step is performed through the supervised edge annotation module M4. Edge features include relation type, supervision signals (resolved, cost, etc.), and process information (steps, errors, etc.).
[0095] The input to the supervised edge annotation module M4 is: heterogeneous graph G, repair labels and cost logs, and the output is: supervised edge set E_sup.
[0096] For each issue-agent attempt (where the agent attempts to perform automated defect remediation tasks on the issue), a remediation label "resolved" is generated based on the acceptance results. In benchmarks such as SWE-Bench, this label typically comes from whether automated tests pass or fail; in engineering systems, it can also come from CI pipelines, regression tests, or manual verification. It is uniformly mapped to resolved=1 or 0, where 1 corresponds to successful remediation and 0 corresponds to unresolved remediation.
[0097] Reference Figure 2 The supervision of edge annotation steps includes:
[0098] Step S211: When there are attempt records between the issue and the agent, create an initial supervision edge between the issue node and the agent node in the heterogeneous graph;
[0099] Step S212: Associate the initial supervision edge with the execution trajectory of the issue and the agent;
[0100] Step S213: Write the repair label into the initial supervision edge; the repair label is used to record whether the agent successfully repaired the issue;
[0101] Step S214: Write the cost into the initial supervision edge; the cost includes the lexical cost of each round of invocation, the tool invocation cost, and the retry cost;
[0102] Step S215: Obtain the set of supervision edges based on the modified initial supervision edges.
[0103] Specifically, the process iterates through the history or sampled candidate agent list for each issue, locating the corresponding issue node and agent node in the heterogeneous graph. If a complete attempt record exists for this issue-agent combination, a supervision edge is created or updated between the two nodes, and associated with the trace_id of this attempt. The resolved value is written based on the source of the repair tag: 1 for successful repair, 0 otherwise. Billing entries related to this attempt are filtered from the cost log, aggregated into a total cost edge_cost_usd according to the cost aggregation rules, and written to the supervision edge.
[0104] The fields of the supervised edge include: issue_id, agent_id, resolved, edge_cost_usd (total cost of issue-agent task), and split (training / validation / testing partition label).
[0105] For cost aggregation rules, when a single issue-agent task involves multiple rounds of calls, the cost aggregation function is defined as follows:
[0106] edge_cost_usd=
[0107] Σ(token_cost_i+tool_call_cost_i+retry_cost_i).
[0108] Therefore, the cost granularity is not the global average price of the agent, but the actual cost of the issue-agent.
[0109] Understandably, if simplification is needed, only token-based billing can be used, but consistency should be maintained within the same system.
[0110] Step S300: First, the training sample arrangement module M5 divides the training mask (train_mask), verification mask (val_mask), and test mask (test_mask) according to the issue dimension to ensure that the same issue does not cross sets.
[0111] The input to the training sample orchestration module M5 is the set of supervised edges E_sup, and the output is train_mask, val_mask, test_mask, and an index constructed by pair according to issue.
[0112] Then, the model training module M6 performs heterogeneous message passing, edge scoring, loss calculation, backpropagation, and early stopping to obtain the trained target routing model, which can then be deployed as an inference model.
[0113] The inputs to the model training module M6 are: G, E_sup, mask and hyperparameters, and the outputs are: the trained model parameters θ* and training logs.
[0114] Reference Figure 3 Specifically, the model training steps include:
[0115] Step S311: Divide the training mask, verification mask, and test mask according to the issue dimension;
[0116] Step S312: Initialize the parameters of the routing network of the heterogeneous graph;
[0117] Step S313: Forward propagation to obtain the matching score for each supervised edge;
[0118] Step S314: Construct type A pair, type B pair, and type C pair according to the candidate set within the issue;
[0119] Step S315: Calculate the loss function based on the matching score, perform backpropagation based on the loss function, and update the parameters of the routing network;
[0120] Step S316: Calculate Hit@1, MRR, and cost indicators in the validation set and select the best option by early stopping.
[0121] Step S317: Solidify the optimal parameters and output the inference model.
[0122] When training the routing model, set the basic training parameters.
[0123] Learning rate lr: 1e -4 up to 5e -3 ;
[0124] Weight decay wd: 1e -6 up to 1e -3 ;
[0125] Hidden layer dimension h: 32 to 256;
[0126] Training epochs: 20 to 200.
[0127] Let the positive sample set of an issue be P (resolved=1), the negative sample set be N (resolved=0), the matching score be s_i, and the cost be c_i. Define the softplus function: softplus(x) = log(1 + exp(x)), where x is the difference between two candidate matching scores; the output of softplus(x) is a non-negative scalar, which takes a large value when the sorting constraint is violated, and serves as the contribution of the pairwise loss for this item.
[0128] Class A constraint implements repair priority, which gives:
[0129] L_A=mean_{p∈P,n∈N}softplus(s_n-s_p).
[0130] Class B constraint implements low cost priority within positive samples, which gives:
[0131] For any p_i,p_j∈P, if c_i< c_j, construct:
[0132] L_B=mean(softplus(s_j-s_i)).
[0133] Class C constraint implements low-cost negative sample suppression, which gives:
[0134] For any n∈N,p∈P, if c_n<c_p, construct:
[0135] L_C=mean(softplus(s_n−s_p)).
[0136] In a specific embodiment, the total loss function is expressed as: L=L_A+λ_B*L_B;
[0137] In another specific embodiment, the total loss function is expressed as: L=L_A+λ_B*L_B+λ_C*L_C;
[0138] In the formula, L is the total loss function, L_A is the first sub-loss function corresponding to the first constraint, L_B is the second sub-loss function corresponding to the second constraint, and λ_B is a preset weight hyperparameter corresponding to the second sub-loss function;
[0139] wherein, L_A=mean_{p∈P,n∈N}softplus(s_n-s_p);
[0140] when c_i<c_j, L_B=mean(softplus(s_j-s_i));
[0141] when c_n<c_p, L_C=mean(softplus(s_n−s_p));
[0142] softplus(x)=log(1+exp(x)),x=s_n-s_p;
[0143] mean is the arithmetic mean, P is the set of positive samples of the issue, N is the set of negative samples of the issue, s_p is the matching score of positive sample p, s_n is the matching score of negative sample n, softplus is the loss contribution, x is the difference in matching scores, c_i is the cost of the i-th issue sample, c_j is the cost of the j-th issue sample, s_i is the matching score of the i-th issue sample, and s_j is the matching score of the j-th issue sample.
[0144] The recommended range for λ_B is 0.01 to 0.2; the recommended range for λ_C is 0 to 0.1. λ_B and λ_C are non-negative weight hyperparameters used to balance the repair priority and cost priority during the training phase. L serves as the optimization objective for a single iteration of M6, and θ is updated via backpropagation. When λ_B = λ_C = 0, the model degenerates into a per-issue ranking model containing only repair priority constraints.
[0145] By constructing constraints based on pairs within an issue, noise introduced by cross-issue comparisons is avoided.
[0146] When tuning the routing model, the principle is to minimize the average cost while ensuring that the decrease in Hit@1 in the validation set does not exceed the threshold δ (e.g., 0 to 1%).
[0147] During the training of the routing model, anomaly and missing value handling is performed: supervision edges with missing trajectories or costs are removed, training samples are constructed by grouping supervision edges according to issue, and the loss function is not calculated for issues with no positive or negative samples.
[0148] Specifically, if the cost is missing, the data access module M1 provides a fallback strategy, the supervision edge labeling module M4 uses a preset default cost when writing the edge, or marks the supervision edge as invalid; the training sample orchestration module M5 removes invalid edges when generating the training set.
[0149] If an issue has no positive or negative samples during training, the training sample arrangement module M5 will not include the issue in the pair index, and the model training module M6 will not calculate the A, B, and C loss terms for the issue.
[0150] If two repaired samples have the same cost, the model training module M6 will not generate the cost order pair when constructing the B-class pair, or will skip the pair according to the predetermined stable sorting rules.
[0151] For example, offline training can be used to complete data cleaning, graph construction, model training and threshold tuning, and output a versioned model.
[0152] Through the model training process, within the candidate agent set of the same issue, both "repairability priority ranking" and "low cost priority ranking among repairable samples" are simultaneously constrained. These two types of constraints are embedded into the parameter learning during the training phase by the model training module M6 through a joint loss function.
[0153] For step S400, the inference model is deployed online. The issue to be routed and the candidate agent are input into the target routing model to obtain the optimal routing result between the issue to be routed and the candidate agent.
[0154] Specifically, the routing inference module M7 inputs the new issue to be routed and the candidate agent into the target routing model, calculates the score for each (issue, agent) combination, sorts the scores, and outputs Top-1 and Top-K lists, which include scores, cost estimates, confidence levels, etc.
[0155] The multi-objective decision module M8 selects the candidate agent with the lowest cost from the set of candidate agents whose routing scores satisfy score>=best_score-margin, based on the Top-1 and Top-K lists and cost prediction values, as the optimal routing result.
[0156] Where, best_score is the optimal route score of the candidate agent, and margin is a preset floating parameter. Margin is determined by the utility function of the validation set, which is a weighted combination of hit rate and cost metrics. For example, the value range of margin is 0 to 0.2; it is obtained by maximizing U = Hit@1-α*AvgCost on the validation set.
[0157] For example, refer to Figure 4 The reasoning and decision-making process includes:
[0158] Step S411: Receive the issue to be routed and the list of candidate agents;
[0159] Step S412: Construct the edge features of issue and agent and input them into the inference model to obtain scores;
[0160] Step S413: Generate the sequence ranked_agents based on the scores;
[0161] Step S414: The first character of the output sequence is Top-1;
[0162] Step S415: Enable multi-objective decision-making and select the lowest cost from the candidate agents that satisfy score>= best_score-margin as the optimal routing result;
[0163] Step S416 returns the optimal agent, confidence level, and explanatory fields; the explanatory fields include score / cost / candidate ranking, etc.
[0164] Finally, new samples are collected daily / weekly for incremental retraining. Through the online feedback module M9, the sample library is updated, cost distribution is refreshed, and periodic retraining is triggered based on actual execution results and real costs, outputting new version data and models. By updating cost statistics and model parameters, long-term adaptability is ensured.
[0165] The results were evaluated and analyzed. The main metrics included Hit@1 (Top 1 fix rate), and the auxiliary metrics included MRR, Accuracy, ROC-AUC, and PR-AUC. The cost metrics included Avg Cost / Issue and Cost / Resolved. The analysis dimensions included: grouping by repository / language / complexity, cost-effectiveness trade-off analysis, and capability profile contribution analysis.
[0166] To apply this multi-agent-based software defect repair method in an implementation environment, the following settings are used:
[0167] Task scenario: Software defect repair routing (SWE-Bench style issue-agent routing task);
[0168] Number of candidate agents: 7;
[0169] Monitoring edge size: 3500 samples (issue, agent);
[0170] Division method: Divided by issue, train=325, val=75, test=100;
[0171] Final experimental random seed: seed=17;
[0172] Cost field: issue-agent granularity actual cost edge_cost_usd.
[0173] In Example 1, a basic cost-aware route is configured, and only Class A and Class B losses are used during the training phase. The goal is to reduce the average cost without reducing Hit@1.
[0174] The steps are as follows:
[0175] Read issue, agent, behavior trajectory, and cost data to construct a heterogeneous graph;
[0176] Generate supervised edges and write them to resolved and edge_cost_usd;
[0177] Construct sorted pairs of type A (resolved > unresolved) and type B (cheap resolved > expensive resolved) within the same issue;
[0178] The model is trained using L = L_A + λ_B * L_B + λ_C * L_C, where in this embodiment λ_C = 0;
[0179] Select the optimal parameters on the validation set and solidify the model;
[0180] Perform Top-1 routing and collect metrics on the test set.
[0181] The key training configuration includes: the ranking loss is cost_aware_ranking;
[0182] With cost_resolved_pair_weight = 0.06 and cheap_unres_weight = 0.0, the optimal validation set MRR is 0.8378.
[0183] The following test results were obtained: Hit@1 / Top1 repair rate: 0.87; MRR: 0.88;
[0184] Avg Cost / Issue: 0.2657; Cost / Resolved: 0.3054; Accuracy: 0.65; ROC-AUC: 0.6002; PR-AUC: 0.8703.
[0185] In Example 2, an enhanced robust routing is configured, and Class A loss, Class B loss, and Class C loss are used during the training phase. The goal is to reduce the probability of "low-cost but unrepairable" candidates being mistakenly ranked at the top.
[0186] In Example 3, a multi-target deployment version of the route is configured. During the training phase, Class A and Class B losses are used, and during the inference phase, margin is enabled for secondary screening based on the lowest cost. The objective is to stably output repairable candidates with lower costs under online budget constraints. Specifically, margin=0.02 (verified using seed17).
[0187] The following test results were obtained: Multi-Objective Router: Hit@1=0.86, Avg Cost / Issue=0.2530; Hybrid Router: Hit@1=0.86, Avg Cost / Issue=0.2986.
[0188] With the hit rate remaining constant, the multi-objective strategy further reduces the average cost, demonstrating the engineering practicality of this implementation method.
[0189] This multi-agent-based software defect repair method is compared with the baseline scheme.
[0190] The baseline is defined as follows:
[0191] Random Expected: The theoretical expectation of uniformly randomly selected candidate agents;
[0192] Leaderboard Single Best: Fixed single strong baseline agent;
[0193] Test-set Single Best: Dynamically optimal single agent for the test set (for diagnostic purposes only, not for deployment).
[0194] The comparison results include:
[0195] Random Expected: Hit@1=0.8086, Avg Cost / Issue=0.5574;
[0196] Leaderboard Single Best: Hit@1=0.82, Avg Cost / Issue=0.8740;
[0197] Test-set Single Best (acoder): Hit@1=0.83, Avg Cost / Issue=0.8116;
[0198] This method DKG-GNN: Hit@1=0.87, Avg Cost / Issue=0.2657.
[0199] Compared to Random Expected, this method improves Hit@1 by approximately +6.14 percentage points and significantly reduces average cost; compared to Test-set Single Best, this method improves Hit@1 by approximately +4.00 percentage points while also being more cost-effective. These results demonstrate that this method achieves significant technical effectiveness in achieving the combined goals of success rate and cost.
[0200] Through ablation analysis, after removing the `triggered_error` relation: Hit@1=0.81; using the modified relation variant: Hit@1=0.80; final configuration (No Modified + Triggered_Error): Hit@1=0.87. This illustrates that graph relation configuration and cost-aware training work together to improve performance.
[0201] The DKG-GNN (Final) method was compared with other methods, such as Random Expected, Random Sampled, Leaderboard Single Best (liveswe_claude45), and Test-set Single Best (acoder), and the core comparison results are shown in Table 1.
[0202] Table 1. Core Comparison Results
[0203]
[0204] The advantages of this multi-agent-based software defect repair method are demonstrated on the SWE-Bench style defect repair routing task:
[0205] Reference Figure 7 On the seed=17 test set, the Hit@1 score of our proposed DKG-GNN was 87.0%, an improvement of approximately 6.1 percentage points compared to Random Expected (80.86%), and an improvement of 5.0 percentage points compared to the fixed strong baseline Leaderboard SingleBest (82.0%). It also exceeded the upper bound of the Test-set Single Best diagnostic score (83.0%) by approximately 4.0 percentage points. This demonstrates the effectiveness of supervised edge combined with dual-constraint training.
[0206] Under the same partitioning, Avg Cost / Issue is $0.266, a decrease of approximately 52.2% compared to Random Expected ($0.557) and approximately 69.6% compared to Leaderboard Single Best ($0.874), achieving a joint optimization that prioritizes success rate while considering cost, even with a higher hit rate; the Multi-Objective deployment version further reduces it to $0.253 (Hit@1=86.0%). This indicates that cost-aware ranking leads to improved economics.
[0207] Issue-agent granular cost is superior to global average cost routing. The edge_cost_usd of the supervised edge is bound to the real cost of a single attempt. During the training phase, Class B constraints re-rank the repairable samples within the same issue. It shows that Top-1 is neither randomly assigned nor fixed to the most expensive / strongest single agent, but rather selects a combination of high hit rate and low cost from the candidates.
[0208] The average hit@1 rate for 20-seed was 80.85% (81.20% for Hybrid), which is consistently higher than the random expectation (77.84%) and the baseline mean for single agent (79.70%). This indicates that the per-issue ranking constraint can be repeatedly effective under different train / val / test partitions and is robust across partitions.
[0209] Causal verification (ablation) of graph structure and constraints, with seed=17, shows that the final configuration (No Modified + Triggered_Error) hit@1=87.0%; restoring Modified edges reduces it to 80.0%; removing Triggered_Error reduces it to 81.0% (although the cost is extremely low, it sacrifices hit rate). This demonstrates that the design of heterogeneous graph relationships and A / B class ordering constraints work together to improve routing quality, rather than simply being a superposition of algorithms.
[0210] In summary, this method can reduce the average dispatch cost while maintaining a high repair hit rate; it adopts issue-agent granular cost modeling, which has higher decision accuracy compared to global coarse-grained cost; and it supports deployment in combination with multi-objective decision-making to meet the needs of actual budget-constrained scenarios.
[0211] The above is a detailed description of the preferred embodiments of this application, but this application is not limited to the embodiments. Those skilled in the art can make various equivalent modifications or substitutions without departing from the spirit of this application, and these equivalent modifications or substitutions are all included within the scope defined by the claims of this application.
Claims
1. A software defect repair method based on multi-agent systems, characterized in that, include: Acquire a data set consisting of the software defect task (issue), the intelligent agent, the execution trajectory of the agent in fixing the issue, and the cost; A heterogeneous graph-based routing model is constructed based on the dataset. The heterogeneous graph uses issue and agent as nodes, and a supervision edge is generated between the issue node and the agent node. Data on the agent fixing the issue is written into the supervision edge. The routing model is trained, and the matching score between issue and agent is calculated based on the heterogeneous graph according to the first constraint and the second constraint. The loss function of the routing model is calculated based on the matching score. The parameters of the routing model are updated based on the loss function to obtain the target routing model. The first constraint is used to constrain the score of the repaired sample to be higher than the score of the unrepaired sample. The second constraint is used to constrain the score of the low-cost repaired sample corresponding to the same issue to be higher than the score of the high-cost repaired sample. Input the issue to be routed and the candidate agent into the target routing model to obtain the optimal routing result between the issue to be routed and the candidate agent.
2. The software defect repair method based on multi-agent technology according to claim 1, characterized in that, The step of constructing a heterogeneous graph-based routing model based on the dataset includes: The dataset is sequentially processed by field mapping, time alignment, dirty data removal, and primary key unification to obtain standardized intermediate tables, which include issue intermediate tables, event intermediate tables, trajectory intermediate tables, and cost intermediate tables.
3. The software defect repair method based on multi-agent technology according to claim 2, characterized in that, The step of constructing a heterogeneous graph-based routing model based on the dataset includes: Based on the standardized intermediate table, issue text vectorization, agent capability profile encoding, action feature extraction, and code node feature extraction are performed to obtain a node type feature matrix, which includes an issue feature matrix, an agent feature matrix, an action feature matrix, and a code feature matrix.
4. The software defect repair method based on multi-agent technology according to claim 3, characterized in that, The step of constructing a heterogeneous graph-based routing model based on the dataset includes: Based on the node type feature matrix and relation mapping table, a node set and an edge set are constructed to obtain a heterogeneous graph sample.
5. The software defect repair method based on multi-agent technology according to claim 4, characterized in that, The step of constructing a heterogeneous graph-based routing model based on the dataset includes: When there are attempt records between an issue and an agent, an initial supervision edge is created between the issue node and the agent node in the heterogeneous graph; The initial monitoring edge is associated with the execution trajectory of the issue and the agent; Write the repair label into the initial supervision edge; the repair label is used to record whether the agent successfully repaired the issue. The cost is written into the initial supervision edge, and the cost includes the lexical cost, tool call cost, and retry cost for each round of invocation; The set of supervised edges is obtained based on the modified initial supervised edges.
6. The software defect repair method based on multi-agent technology according to claim 5, characterized in that, During the training of the routing model, supervision edges with missing trajectories or costs are removed, and training samples are constructed by grouping supervision edges according to issue. The loss function is not calculated for issues with no positive or negative samples.
7. The software defect repair method based on multi-agent technology according to claim 1, characterized in that, The loss function is expressed as: L = L_A + λ_B * L_B; In the formula, L is the total loss function, L_A is the first sub-loss function corresponding to the first constraint, L_B is the second sub-loss function corresponding to the second constraint, and λ_B is the weight hyperparameter corresponding to the preset second sub-loss function; Where L_A = mean_{p∈P,n∈N}softplus(s_n-s_p); When c_i <c_j,L_B=mean(softplus(s_j-s_i)); softplus(x)=log(1+exp(x)),x=s_n-s_p; mean is the arithmetic mean, P is the set of positive samples of the issue, N is the set of negative samples of the issue, s_p is the matching score of positive sample p, s_n is the matching score of negative sample n, softplus is the loss contribution, x is the difference in matching scores, c_i is the cost of the i-th issue sample, c_j is the cost of the j-th issue sample, s_i is the matching score of the i-th issue sample, and s_j is the matching score of the j-th issue sample.
8. The software defect repair method based on multi-agent technology according to claim 1, characterized in that, The loss function is expressed as: L = L_A + λ_B * L_B + λ_C * L_C; In the formula, L is the total loss function, L_A is the first sub-loss function corresponding to the first constraint, L_B is the second sub-loss function corresponding to the second constraint, L_C is the third sub-loss function for suppressing low-cost negative samples, λ_B is the weight hyperparameter corresponding to the preset second sub-loss function, and λ_C is the weight hyperparameter corresponding to the preset third sub-loss function. Where L_A = mean_{p∈P,n∈N}softplus(s_n-s_p); When c_i <c_j,L_B=mean(softplus(s_j-s_i)); When c_n <c_p,L_C=mean(softplus(s_n−s_p)); softplus(x)=log(1+exp(x)),x=s_n−s_p; mean is the arithmetic mean, P is the set of positive samples of the issue, N is the set of negative samples of the issue, s_p is the matching score of positive sample p, s_n is the matching score of negative sample n, softplus is the loss contribution, x is the difference in matching scores, c_i is the cost of the i-th issue sample, c_j is the cost of the j-th issue sample, s_i is the matching score of the i-th issue sample, s_j is the matching score of the j-th issue sample, c_p is the cost of negative sample p, and c_n is the cost of negative sample n.
9. The software defect repair method based on multi-agent technology according to claim 1, characterized in that, Input the issue to be routed and the candidate agent into the target routing model to obtain the optimal routing result between the issue to be routed and the candidate agent, including: From the set of candidate agents whose routing scores satisfy score>=best_score-margin, select the candidate agent with the lowest cost as the optimal routing result; Wherein, best_score is the optimal routing score of the candidate agent, margin is a preset floating parameter, and margin is determined by the utility function of the validation set, the utility function being a weighted combination of hit rate and cost metrics.
10. A software defect repair device based on multi-agent systems, characterized in that, include: The data input module is used to acquire a data set consisting of the software defect task to be repaired (issue), the intelligent agent, the execution trajectory of the agent in repairing the issue, and the cost. The model building module is used to build a routing model based on a heterogeneous graph according to the dataset. The heterogeneous graph has issue and agent as nodes, and a supervision edge is generated between the issue node and the agent node. The data of the agent fixing the issue is written into the supervision edge. The model training module is used to train the routing model, calculate the matching score between issue and agent based on the heterogeneous graph according to the first constraint and the second constraint, calculate the loss function of the routing model according to the matching score, and update the parameters of the routing model according to the loss function to obtain the target routing model. The first constraint is used to constrain the score of the repaired sample to be higher than the score of the unrepaired sample, and the second constraint is used to constrain the score of the low-cost repaired sample corresponding to the same issue to be higher than the score of the high-cost repaired sample. The routing decision module is used to input the issue to be routed and the candidate agent into the target routing model to obtain the optimal routing result between the issue to be routed and the candidate agent.