Task context-based ai code dynamic review intensity and routing method and system
Patent Information
- Application Number
- CN202610826272.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-09
- Publication Date
- 2026-09-01
AI Technical Summary
[0003]现有技术已经公开了利用人工智能生成代码审查意见、预测合并请求风险、推荐代码审查人、对代码差异进行结构化分析以及基于历史样本执行自动代码审查等方案,上述方案能够提高审查效率,但多数仍围绕代码片段或合并请求文本如何被模型审查、变更是否可能引入缺陷或谁适合参与审查展开,尚未充分解决AI研发执行平台中如何基于任务执行全链路证据动态确定审查强度,并把不同专项风险路由至不同审查角色的治理问题;因此,现有技术存在以下缺陷:
1、本发明不只依赖代码差异或合并请求文本确定审查强度,而是融合AI执行记录、验证不确定性、多模型仲裁结果和影响图谱等多源证据生成审查证据包,在此基础上动态确定审查强度等级并生成专项风险标签,证据维度更完整,能够识别仅凭代码差异无法发现的风险,从而降低高风险变更被低估或漏审的概率;且低风险变更可通过AI基础检查在数秒内自动通过,无需人工介入;高风险变更自动进入专项人工审查或多方会审,可减少人工审查资源浪费,使有限的专家审查资源集中于高风险变更。
Smart Images

Figure CN122673076A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of intelligent software engineering technology, specifically to a method and system for dynamic review intensity and routing of AI code based on task context. Background Technology
[0002] With the development of AI code generation, AI R&D execution platforms, and code hosting platforms, the number of merge requests, the speed of changes, and the complexity of reviews have increased significantly. In AI-generated code and AI-assisted development scenarios, the core goal of code review is not to replace manual review, but to dynamically determine the review intensity and path based on multi-dimensional evidence, so that low-risk changes can quickly pass basic checks, while high-risk changes automatically enter into specialized manual reviews or multi-party reviews.
[0003] Existing technologies have disclosed solutions for using artificial intelligence to generate code review opinions, predict merge request risks, recommend code reviewers, perform structured analysis of code differences, and perform automated code reviews based on historical samples. While these solutions can improve review efficiency, most still focus on how code snippets or merge request texts are reviewed by the model, whether changes might introduce defects, and who is suitable to participate in the review. They have not yet fully addressed the governance issues of how to dynamically determine the review intensity based on evidence from the entire task execution chain in an AI R&D execution platform, and how to route different specific risks to different review roles. Therefore, existing technologies have the following shortcomings: 1) There is a lack of a review intensity determination mechanism for evidence across the entire AI R&D task chain. Existing solutions mostly rely on merging request texts, code differences, historical review samples, or document risk characteristics for judgment. They rarely model AI execution records, deviations from task objectives, verification uncertainties, scope of impact, historical defects, and organizational strategies into a computable review evidence package, resulting in a lack of comprehensiveness in the review intensity determination.
[0004] 2) Lack of linkage between review intensity and special review routing. Existing solutions can generate review opinions, recommend reviewers, or predict defect risks, but they usually do not automatically generate multi-role review paths based on special risk tags such as security, database, architecture, testing, performance, configuration, and compliance. This results in changes to different types of risks not being accurately routed to the corresponding special review roles.
[0005] 3) The lack of an organizational governance loop to address the unique risks of AI-generated code. AI-generated code may have risks such as passing tests on the surface but deviating from the task objectives, weakening verification, omitting boundary conditions, or conflicting with the results of multiple models. Ordinary code review models cannot identify these risks by relying solely on code differences, resulting in insufficient review coverage of AI-generated code.
[0006] 4) Lack of a mechanism for aggregating the execution status and conclusions of review routes. Existing automated code reviews typically output a single review suggestion or score, lacking unified tracking and summarization of multiple review statuses such as AI pre-review, module leader review, project leader review, and multi-party review, resulting in an untraceable review process and incomplete review conclusions.
[0007] 5) The lack of feedback optimization of review conclusions on strategies and model capability profiles means that existing solutions rarely write back manual rejection reasons, false positives / missed positives, supplementary verification results, post-launch defects and production accidents to review intensity thresholds, special risk label rules, role routing rules and model capability profiles, which makes it impossible for the review system to continuously learn and optimize from historical review results. Summary of the Invention
[0008] The purpose of this invention is to provide a method and system for dynamic review intensity and routing of AI code based on task context. The review intensity is determined more comprehensively and the review resources are allocated more reasonably. It can identify AI-specific risks, the review process is traceable, a closed-loop optimization can be formed, and the review accuracy can be continuously improved, thus solving the problems mentioned in the background art.
[0009] To achieve the above objectives, the present invention provides the following technical solution: The AI code dynamic review intensity and routing method based on task context includes the following steps: S1. Receive code change or merge request: Receive a code change from the code platform, AI R&D execution platform or continuous integration system; S2. Collect review evidence: Collect task context, code differences, change files, test results, verification evidence, historical defects, task contract records, and impact graph information; S3. Generate a review evidence package: Unify and standardize evidence from different sources to generate a structured review evidence package specific to this change. S4. Calculate the review intensity score: The review intensity score is calculated based on the scope of impact of the change, business sensitivity, verification uncertainty, historical failure weight, AI-generated confidence risk, access to sensitive resources, and the weight of special risk labels. S5. Determine the review level: Map the review intensity score to review levels from S0 to S4; S6. Generate review routes: Determine the AI review, manual review, and special review paths to be performed based on the review level and specific risk labels; S7. Perform review and summarize conclusions: Summarize the results of AI review, special inspection and manual confirmation, and output conclusions such as pass, supplementary verification, re-review after modification, block and merge or manual takeover. S8. Review Conclusion Feedback: Write the review process and conclusions into the audit log, and write back the reasons for false reports, omissions, rejections and subsequent defects to the strategy library.
[0010] Preferably, the review evidence package is used to drive review intensity calculation and review route generation, and the review evidence package includes at least task context, change information, execution evidence, verification evidence, impact graph evidence, knowledge credibility evidence, personnel responsibility evidence, and model arbitration evidence.
[0011] Preferably, the model arbitration evidence comes from the multi-model arbitration module. When multiple AI models are called to generate review opinions for the same code change, the multi-model arbitration module determines whether there is an evidence conflict according to consistency judgment, conflict judgment, conflict degree quantification and escalation processing.
[0012] Preferably, the review intensity score Obtain it using the following formula: ; in, The review intensity score ranges from 0 to 10. This refers to the scope of impact of the change, normalized to a risk score of 0-10, and is used to measure the number of modules, files, and downstream dependencies involved in the code change; Business sensitivity refers to a risk score normalized to 0-10, used to measure the criticality of the business path involved in a change. This refers to the uncertainty in verification, which is normalized to a risk score of 0-10. It is used to measure factors including insufficient test coverage, number of automatic repair retries, and low knowledge credibility level. Historical fault weights, normalized to a risk score of 0-10, are used to measure the frequency and severity of defects in a module or file throughout history. This refers to the confidence risk generated by AI, normalized to a risk score of 0-10, used to measure the risks of path default, multiple rounds of repair, and low confidence levels during AI execution. Sensitive resource access is normalized to a risk score of 0-10 and used to measure whether changes involve sensitive resources such as production configurations, keys, personal information, and non-rollbackable operations. The specific risk label weight is normalized to a risk score of 0-10 and is used to measure the number and severity of specific risk labels that have been hit. This refers to the weighting coefficient for the scope of the change's impact; the default value is 0.20. This refers to the weighting coefficient for business sensitivity, with a default value of 0.20. This refers to the weighting coefficient for verifying uncertainty; the default value is 0.15. This refers to the weighting coefficient for historical faults, with a default value of 0.15. This refers to the weighting coefficient for the confidence risk generated by AI, with a default value of 0.10; This refers to the weighting coefficient for sensitive resource access, with a default value of 0.10. This refers to the weighting coefficient of the specific risk label, with a default value of 0.10.
[0013] Preferably, when the scope of impact of the change, business sensitivity, verification uncertainty, historical fault weight, AI-generated confidence risk, sensitive resource reach, and special risk label weight are mapped from the original data to a risk score of 0-10, the following operations are performed: The scope of impact of the change is determined by the number of modules involved M and the number of downstream dependent services D. The original value is calculated as M×2 + D×3, and the normalized risk score is min(original value, 10). Business sensitivity is mapped based on the preset sensitivity level of the business paths involved in the change, and the highest value is taken when the change involves multiple business paths; To verify the uncertainty, we take the number of test coverage gaps G, the number of automatic repair retries R, and the knowledge credibility level penalty K, and calculate the risk score = min(G×2 + R×1.5 + K, 10); Historical fault weights are calculated by taking the number of defects N and the highest severity level S of the module in the past six months, and the risk score is calculated as min(N×S, 10). AI generates confidence risk. When it is not generated by AI, the risk score is 0. When it is generated by AI, the number of path defaults V and the low confidence level marker C are taken. The risk score is calculated as min(V×3 + C×4 + (repair rounds-1)×2, 10). Sensitive resource access is incremented based on the type of sensitive resource accessed: production configuration adds 4, key or signature adds 4, personal information adds 3, and non-rollbackable operation adds 5. The smaller value between the incremented value and 10 is taken. The weight of the special risk label is the sum of the number of special risk labels that have been hit and the preset severity coefficient of each label. The risk score is calculated as min(the sum of the severity coefficients of each label, 10).
[0014] Preferably, the review intensity score is mapped to review levels from S0 to S4, including: When review intensity rating If the level is below 4.0, and the changes only involve documents, comments, or low-risk styles and have no business module connection to the display of the map, then the review level is determined to be S0. When review intensity rating If the score is below 4.0, and the code involves non-core logic but has a small impact, the review level is determined to be S1. When review intensity rating If the score is below 7.0 but not below 4.0, the review level is determined to be S2. When review intensity rating If the score is not lower than 7.0, the review level is determined to be S3; If any of the following are detected: changes to production environment configuration, core permission logic, non-rollbackable database migration, changes to billing and settlement logic, changes to personal information processing logic, changes to cross-service communication protocols, or highly conflicting AI multi-model outputs, or if multiple high-risk special tags are detected, the review level will be determined as S4.
[0015] Preferably, the specific risk label is a structured routing basis automatically generated or modified based on the review evidence package, wherein the specific risk label is generated according to the following rules: When code changes involve authentication, permission checks, keys, signatures, sensitive data processing, login status, tenant isolation, or security configuration, generate a security audit label. When code changes involve table structure, indexes, migration scripts, field state transitions, transaction boundaries, query performance, or data consistency, a database audit label is generated. When code changes affect cross-service interfaces, common components, domain models, dependency upgrades, messaging protocols, or architectural boundaries, generate architectural review tags. When test coverage is insufficient, historical defects are frequent, verification results conflict, or there are gaps in the observation of the spectrum, a test review or supplementary verification label is generated. When code changes involve production configuration, pricing, settlement, personal information, or non-rollback operations, a compliance review label is generated. When a code change simultaneously triggers multiple high-risk special tag combinations, or triggers any mandatory blocking condition, a multi-party review tag is generated. When code changes involve hot paths, changes in algorithm complexity, changes in expected response time or throughput, adjustments to connection pool or cache parameters, or significant changes in resource consumption, a performance audit label is generated. When code changes involve environment variables, feature switches, rate limiting or circuit breaker parameters, log levels, middleware configurations, or deployment descriptors, generate configuration review tags.
[0016] Preferably, after the review route is generated, one or more review nodes are created according to the review route scheme. Each node corresponds to AI pre-review, module leader review, project leader review, supplementary verification or multi-party review. The review nodes have states such as pending, under review, passed, requiring modification, supplementary verification, under review, blocked merging, manual takeover and timeout upgrade. The review conclusion is aggregated after all necessary nodes are completed. If any forced node outputs blocked merging or manual takeover, the overall conclusion cannot be automatically passed.
[0017] Preferably, each review node is configured with a timeout limit. When a review node in the pending state exceeds the time limit and is not claimed, an upgrade notification is automatically sent to the node's superior and the node status is changed to timeout upgrade. When a review node in the review process exceeds the time limit and does not output a conclusion, a timeout upgrade is also triggered. After the timeout upgrade, if it is still not completed within the extra grace period, the node is automatically changed to manual takeover status, whereby a higher-level manager intervenes or organizes a multi-party review.
[0018] According to another aspect of the present invention, a task-context-based AI code dynamic review intensity and routing system is provided for implementing the task-context-based AI code dynamic review intensity and routing method as described above, comprising: The change input module is used to receive merge requests, branch changes, code differences, AI execution results, or R&D task outputs generated by the code platform. The evidence collection module is used to collect task context, code change information, task contract execution records, automatic repair verification results, change impact graph, knowledge trust level, tool access records, and historical defect records. The risk evidence fusion module is used to convert evidence from multiple sources into a unified change risk profile and generate a risk feature vector. The review intensity determination module is used to determine the review intensity level based on risk feature vectors, team strategies, and dynamic thresholds; The specialized risk label generation module is used to generate specialized risk labels for security, database, architecture, testing, performance, configuration, and compliance based on the changed object, resource type, impact graph, historical defects, verification gaps, and team strategies. The review routing module is used to generate review routing schemes for AI pre-review, module leader review, security review, database review, architecture review, test review, or multi-party review based on review intensity level, specific risk labels, role and responsibility table, and mandatory blocking conditions. The review routing execution module is used to track the status of each review node, including pending, passed, required modification, supplementary verification, under review, blocked merging, and manual takeover, and to summarize the review conclusions from multiple paths. The AI pre-review module is used to conduct an initial review of code quality, potential defects, security issues, testing gaps, boundary conditions, and consistency with task objectives. The human expert review module is used to assign high-risk or specific risk changes to the corresponding roles, including module owners, security owners, database owners, architecture owners, and test owners. The audit and tracking module is used to record evidence sources, scoring processes, review intensity, routing results, review conclusions, and manual confirmation information. The review conclusion feedback module is used to write back the review results, false alarm or omission markers, rejection reasons, approval reasons, supplementary verification results, and subsequent defect results to the review strategy library, special risk label rules, role routing rules, and model capability profile.
[0019] Compared with the prior art, the beneficial effects of the present invention are: 1. This invention does not rely solely on code differences or merge request text to determine the intensity of review. Instead, it integrates multi-source evidence, such as AI execution records, verification uncertainties, multi-model arbitration results, and impact maps, to generate a review evidence package. Based on this, it dynamically determines the review intensity level and generates specific risk labels. The evidence dimensions are more complete, enabling the identification of risks that cannot be discovered by code differences alone, thereby reducing the probability of high-risk changes being underestimated or missed. Furthermore, low-risk changes can be automatically approved within seconds through AI basic checks without human intervention. High-risk changes automatically enter specialized manual review or multi-party review, which can reduce the waste of manual review resources and concentrate limited expert review resources on high-risk changes.
[0020] 2. This invention incorporates task contract breach records, AI confidence levels, multi-model arbitration results, and verification uncertainties into the review intensity assessment. It can identify risks unique to AI-generated code, supplementing the review basis from the AI execution process dimension. Based on specific risk labels, it automatically determines whether special reviews such as security, database, architecture, testing, performance, configuration, and compliance are required. By tracking the status of each review node and aggregating multiple review conclusions through a routing execution state machine, it achieves structured tracking of the review process and complete aggregation of conclusions. This reduces the problems of misrouting of different types of risks or the irreversibility of the review process. Review conclusions, false positives / false negatives, rejection reasons, and subsequent defects can be written back to the strategy library and model capability profile, enabling subsequent review intensity assessment, label rules, and role routing to be gradually optimized. It can continuously learn from historical review experience, reducing the false positive and false negative rates in subsequent reviews. Attached Figure Description
[0021] Figure 1 This is a block diagram of the AI code dynamic review strength and routing system of the present invention; Figure 2 This is a flowchart of the AI code dynamic review strength and routing method of the present invention; Figure 3 This is a structural diagram of the examination evidence package and risk feature vector of the present invention. Detailed Implementation
[0022] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0023] To address the existing issues of lacking a mechanism for determining the intensity of review evidence across the entire AI R&D task chain, lacking linkage between review intensity and specific review routes, lacking a closed-loop organizational governance system to address the unique risks of AI-generated code, lacking a mechanism for aggregating the execution status and conclusions of review routes, and lacking a mechanism for feedback optimization of review conclusions on strategies and model capability profiles, please refer to [link to relevant documentation]. Figures 1-3 This embodiment provides the following technical solution: Example 1 Application in high-risk special review scenarios: The AI R&D execution platform receives a merge request automatically generated by AI, which involves changes to the order status processing logic, an external interface field, and two unit test files.
[0024] Among them, the AI code dynamic review intensity and routing method based on task context includes: After receiving the merge request, the system collects code differences, task descriptions, AI execution logs, test results, and change impact graphs. The impact graph shows that the change affects the order status interface and downstream notification service. The knowledge base shows that the module has experienced two status transition defects in the past three months. The knowledge credibility evidence shows that the module's interface specification is K2 level (confirmed knowledge) and the order status transition constraint is K3 level (mandatory constraint knowledge). The task contract record shows that no path violation occurred during this AI execution, but the automatic repair rounds reached 2, indicating that the AI has some uncertainty about this issue; The system calculates the review intensity score. The score is 7.8 (where the scope of change impact scores higher due to its involvement of the core order link and downstream notification services, business sensitivity scores higher due to the order status being a critical business path, verification uncertainty scores higher due to the number of automatic repair rounds reaching 2, historical fault weight scores higher due to the module having two status transition defects within three months, and special risk label weight scores higher due to the simultaneous occurrence of the cross-service interface and historical defect high incidence labels), which is higher than the high-risk threshold of 7.0 and hits the two special risk labels of cross-service interface and historical defect high incidence. The system determines the review level as S3 and generates the review route: AI deep pre-review, module owner review, architecture-specific review, and test owner confirmation of interface regression coverage; After the review was completed, the module manager requested that the boundary conditions be modified, and the test manager requested that a state rollback test case be added. The system recorded the reasons for rejection and rewrote the review policy.
[0025] Example 2 The difference from Example 1 lies in the application scenario. This example is applied to a low-risk, fast-through scenario: the AI R&D execution platform receives a merge request generated with AI assistance. The changes only involve updating the parameter descriptions of three interfaces in the interface documentation and correcting a spelling error in a code comment, modifying a total of two files and four lines of text.
[0026] Among them, the AI code dynamic review intensity and routing method based on task context includes: After receiving the merge request, the system collects code differences, task descriptions, AI execution logs, and change impact graphs. Code difference analysis shows that this change does not involve any executable code, interface definitions, database operations or configuration files, only the documentation and code comments have been modified; The impact graph shows that this change is not related to any business modules, data tables, or downstream services; the knowledge base shows that the module has no related defect records in the past six months. The task contract record shows that the AI execution was completed in one go, with no path violation or automatic repair retry, indicating a high level of confidence in the AI. The system calculates the review intensity score. The score is 1.6 (the scope of the change is extremely low because it only involves documents and comments; the business sensitivity is 0 because there is no business logic connection; the verification uncertainty is extremely low because it was completed in one go and there was no retry; the historical fault weight is 0 because there are no recent defects in this module; the AI-generated confidence risk is extremely low because it has a high confidence level; the sensitive resource reach is 0; and the special risk label weight is 0 because it did not hit any special risk label), which is lower than the low risk threshold of 4.0. In addition, the change type is pure document / comment and the impact graph has no business module connection. The system sets the review level to S0 (basic review) and generates a review route: only performs basic AI checks (syntax, format, link validity), without requiring manual review nodes; Once the AI basic check is passed, the system automatically aggregates the review conclusion and concludes that the request is passed. The merge request is automatically merged without human intervention. The entire review process, from receipt to approval, takes approximately 12 seconds. The system will record this review in the audit ledger and mark it as S0 automatic approval, which will facilitate subsequent statistics on the automatic approval rate of low-risk changes and potential omissions in the audit.
[0027] Example 3 The difference from Implementation Example 1 lies in the application scenario. This implementation example is applied to a high-risk multi-party review scenario: the AI R&D execution platform receives a merge request automatically generated by AI, which changes the amount calculation logic of the user payment module, a database migration script (changing the order amount field from an integer type to a high-precision floating-point type, and this migration cannot be rolled back), and the interface signature logic of the payment gateway.
[0028] Among them, the AI code dynamic review intensity and routing method based on task context includes: After receiving the merge request, the system collects code differences, task descriptions, AI execution logs, test results, change impact graphs, and model arbitration evidence. Code difference analysis shows that this change involves the core logic of payment amount calculation, database migration script and interface signature algorithm. The impact map shows that the change affects the payment module, settlement module and financial reconciliation service. Trusted knowledge evidence shows that the payment amount accuracy constraint is K3 level (mandatory constraint knowledge, derived from financial compliance requirements) and the interface signature specification is K2 level. The system invoked two AI models to pre-screen the change: Model A determined that the change's logic was correct and the risk was controllable (confidence level 0.82); Model B determined that the change posed a risk of precision overflow and signature compatibility issues (confidence level 0.75). The two AI models reached opposite conclusions (Model A indicated approval, Model B indicated risk / blockage), and both models' confidence levels exceeded the minimum confidence threshold of 0.6, indicating a significant difference in confidence levels. = |0.82 - 0.75| = 0.07 < 0.3, indicating a high degree of conflict; The system calculates the review intensity score. The score is 9.2 (high score for scope of impact of change due to involvement of core payment links and three downstream services; extremely high score for business sensitivity due to involvement of tariff settlement; high score for sensitive resource access due to involvement of non-rollback database migration; special risk label weight due to simultaneous occurrence of security, database, and compliance labels), which is higher than the high risk threshold of 7.0. It also meets the mandatory S4 blocking conditions: non-rollback database migration, tariff settlement logic change, and high conflict of AI multi-model output. The system sets the review level to S4 (multi-party review / prohibited automatic merging) and generates the following review route: AI deep pre-review (completed but with conflicting conclusions, marked as manual takeover), security manager review (signature logic), database manager review (non-rollbackable migration), compliance manager review (amount accuracy compliance), payment module manager review, and multi-party review node. The review routing scheme allows security review and database review to be executed in parallel, compliance review to be executed serially after both are passed, and multi-party review to be executed after all specific nodes are completed. The system disables the automatic merging of this request and sends a notification to all responsible persons at the review nodes, with the review node timeout limit configured to 48 hours; After review, the security manager confirmed the signature compatibility issue was genuine and requested modifications; the database manager confirmed that the migration script needed to include a rollback contingency plan document (although automatic rollback was not possible, a manual recovery plan was required); the compliance manager confirmed that the change in monetary precision complied with financial requirements but additional precision boundary testing was needed. After the developer submits the changes, the system will change the status of the relevant node to "under review". After the review is passed, the multi-party review node will be jointly confirmed by the person in charge of the payment module, the person in charge of security, and the person in charge of compliance, and finally output the conclusion of approval.
[0029] The system records the entire review process in the audit ledger, including model arbitration conflict records, review opinions at each stage, and reasons for final approval. It also marks the accuracy overflow risk proposed by Model B as a valid discovery confirmed by manual review and writes it back into the model capability profile to enhance the credibility weight of Model B in the payment field.
[0030] It should be noted that when the system calls multiple AI models (at least two) to generate review opinions for the same code change, the model arbitration module determines whether there is a conflict of evidence according to the following logic: 1) Consistency determination: Compare the review conclusion directions (approval / risk / blockage) of each model for the same change. If all models have consistent conclusion directions, the detailed opinion of the model with the highest confidence level shall be taken as the AI pre-review result. 2) Conflict determination: If at least two models give opposite conclusions (e.g., one model determines that it passes and the other determines that it is a high-risk block), and their respective confidence levels are higher than the preset minimum confidence threshold (e.g., 0.6), then it is determined that the model outputs conflict. 3) Quantification of Conflict Level: Defining Confidence Difference = |max(confidence) - min(confidence)|, when model outputs conflict and confidence levels differ. When the value is below the preset conflict threshold (e.g., 0.3, meaning both parties have high confidence but the conclusions are opposite), it is judged as a high conflict and the S4 forced blocking condition is triggered. 4) Upgrade processing: In the event of a high degree of conflict, the system will write the review opinions, confidence levels and conflict determination results of each model into the model arbitration evidence, generate a manual takeover mark, and make a final decision through manual review.
[0031] It should be noted that the aforementioned review intensity score Obtain it using the following formula: ; in, The review intensity score ranges from 0 to 10. This refers to the scope of impact of the change, normalized to a risk score of 0-10, and is used to measure the number of modules, files, and downstream dependencies involved in the code change; Business sensitivity refers to a risk score normalized to 0-10, used to measure the criticality of the business path involved in a change. This refers to the uncertainty in verification, which is normalized to a risk score of 0-10. It is used to measure factors including insufficient test coverage, number of automatic repair retries, and low knowledge credibility level. Historical fault weights, normalized to a risk score of 0-10, are used to measure the frequency and severity of defects in a module or file throughout history. This refers to the confidence risk generated by AI, normalized to a risk score of 0-10, used to measure the risks of path default, multiple rounds of repair, and low confidence levels during AI execution. Sensitive resource access is normalized to a risk score of 0-10 and used to measure whether changes involve sensitive resources such as production configurations, keys, personal information, and non-rollbackable operations. The specific risk label weight is normalized to a risk score of 0-10 and is used to measure the number and severity of specific risk labels that have been hit. This refers to the weighting coefficient for the scope of the change's impact; the default value is 0.20. This refers to the weighting coefficient for business sensitivity, with a default value of 0.20. This refers to the weighting coefficient for verifying uncertainty; the default value is 0.15. This refers to the weighting coefficient for historical faults, with a default value of 0.15. This refers to the weighting coefficient for the confidence risk generated by AI, with a default value of 0.10; This refers to the weighting coefficient for sensitive resource access, with a default value of 0.10. This refers to the weighting coefficient of the specific risk label, with a default value of 0.10.
[0032] It should be noted that when mapping the scope of impact of the changes, business sensitivity, verification uncertainty, historical fault weight, AI-generated confidence risk, sensitive resource access, and special risk label weights from the original data to a risk score of 0-10, the following operations are performed: The scope of impact of the change is determined by the number of modules involved (M) and the number of downstream dependent services (D). The original value is calculated as M×2 + D×3, and the normalized risk score is min(original value, 10). For example, when the change involves 2 modules and 1 downstream service, the risk score is min(2×2+1×3, 10) = 7. Business sensitivity is mapped according to the preset sensitivity level of the business path involved in the change. For example, the payment / settlement path is mapped to 9, the order core path is mapped to 7, the internal management tool is mapped to 2, and the pure document is mapped to 0. When the change involves multiple business paths, the highest value is taken. To verify the uncertainty, we take the number of test coverage gaps G, the number of automatic repair retries R, and the knowledge credibility level penalty K (+2 for each K0 item and +1 for each K1 item). We calculate the risk score = min(G×2 + R×1.5 + K, 10). For example, when there is 1 test coverage gap, 2 automatic repair retries, and 1 K0 knowledge reference, the risk score = min(1×2+2×1.5+2, 10) = 7. The historical fault weight is calculated by taking the number of defects N and the highest severity level S (fatal = 3, severe = 2, moderate = 1) in the module within the past six months, and the risk score is calculated as min(N×S, 10); for example, if there are 2 severe defects in the past six months, the risk score is min(2×2, 10) = 4. AI generates confidence risk. When not generated by AI, the risk score is 0. When generated by AI, the number of path defaults V and the low confidence level C (0 or 1) are taken, and the risk score is calculated as min(V×3 + C×4 + (repair rounds-1)×2, 10). For example, if there is no path default but 2 rounds of repair are performed, the risk score is min(0+0+(2-1)×2, 10) = 2. Sensitive resource access is calculated by accumulating the risk score based on the type of sensitive resource accessed: production configuration +4, key / signature +4, personal information +3, and non-rollbackable operations +5. The smaller of the accumulated value and 10 is taken. For example, when a non-rollbackable database migration is involved, the risk score is 5. The weight of the special risk label is the sum of the number of special risk labels that have been hit and the preset severity coefficient of each label. The risk score is calculated as min(sum of severity coefficients of each label, 10). For example, when the security label (coefficient 4) and the database label (coefficient 3) are hit, the risk score is min(4+3, 10) = 7.
[0033] It should be noted that the knowledge credibility levels are defined (K0 to K3) as follows: K0 (Unverified Knowledge): Knowledge items with unclear sources or not confirmed by the team, such as coding rules obtained from the Internet but not verified by this project, or design documents that are more than one year old and have not been reviewed again. K0 level knowledge items will have an additional verification uncertainty score in the review scoring.
[0034] K1 (Experience-based knowledge): Knowledge items derived from team members' personal experience or informal communication records that have not yet been formally documented or reviewed and confirmed, such as verbally agreed interface specifications or architectural decisions that have not been formally archived. K1 level knowledge items can be used as a review reference but not as a basis for judgment.
[0035] K2 (Confirmed Knowledge): Knowledge items that have been reviewed by the team or formally archived, such as reviewed design documents, formally released interface specifications, archived architecture decision records, and approved coding standards. K2 level knowledge items can be directly used as the basis for review and judgment.
[0036] K3 (Mandatory Constraint Knowledge): These are derived from organizational-level mandatory standards, regulatory requirements, security compliance strategies, or core constraints that have been validated in production, such as security compliance strategies, data protection regulatory requirements, and core architectural constraints that have been validated in production multiple times. Violations of K3-level knowledge items directly trigger an increase in review intensity or generate a special risk label.
[0037] It should be noted that the censorship intensity score is mapped to censorship levels from S0 to S4, including: When review intensity rating If the level is below 4.0, and the changes only involve documents, comments, or low-risk styles and have no business module connection to the display of the map, then the review level is determined to be S0. When review intensity rating If the score is below 4.0, and the code involves non-core logic but has a small impact, the review level is determined to be S1. When review intensity rating If the score is below 7.0 but not below 4.0, the review level is determined to be S2. When review intensity rating If the score is not lower than 7.0, the review level is determined to be S3; If any of the following are detected: changes to production environment configuration, core permission logic, non-rollbackable database migration, changes to billing and settlement logic, changes to personal information processing logic, changes to cross-service communication protocols, or highly conflicting AI multi-model outputs, or if multiple high-risk special tags are detected, the review level will be determined as S4.
[0038] Specifically, the review level threshold mapping rules are shown in Table 1: Table 1: Review Level Threshold Mapping Rules Specifically, the review levels are defined as shown in Table 2: Table 2: Review Levels It should be noted that the specific risk labels are structured routing criteria automatically generated or modified based on the review evidence package, wherein the specific risk labels are generated according to the following rules: When code changes involve authentication, permission checks, keys, signatures, sensitive data processing, login status, tenant isolation, or security configuration, generate a security audit label. When code changes involve table structure, indexes, migration scripts, field state transitions, transaction boundaries, query performance, or data consistency, a database audit label is generated. When code changes affect cross-service interfaces, common components, domain models, dependency upgrades, messaging protocols, or architectural boundaries, generate architectural review tags. When test coverage is insufficient, historical defects are frequent, verification results conflict, or there are gaps in the observation of the spectrum, a test review or supplementary verification label is generated. When code changes involve production configuration, pricing, settlement, personal information, or non-rollback operations, a compliance review label is generated. When a code change simultaneously triggers multiple high-risk special tag combinations, or triggers any mandatory blocking condition, a multi-party review tag is generated. When code changes involve hot paths, changes in algorithm complexity, changes in expected response time or throughput, adjustments to connection pool or cache parameters, or significant changes in resource consumption, a performance audit label is generated. When code changes involve environment variables, feature switches, rate limiting or circuit breaker parameters, log levels, middleware configurations or deployment descriptors, generate configuration review tags; The aforementioned tags and review levels together determine the routing. For example, for the same S2 change, if the database tag is hit, then at least the database owner's confirmation is required; if both the security tag and the non-rollback tag are hit, it can be upgraded to S4 multi-party review or automatic merging can be prohibited.
[0039] Specifically, the specific risk labels and minimum routing requirements are shown in Table 3: Table 3: Specific Risk Labels and Minimum Routing Requirements It should be noted that review routes are generated based on review levels and specific risk tags. For example, when changes affect database table structure, indexes, migration scripts, or query performance risks, a database-specific review is generated; when authentication, signature, key, permission judgment, or sensitive data processing is affected, a security-specific review is generated; when cross-service interfaces, common components, domain models, or architectural boundaries are affected, an architecture-specific review is generated; and when the impact graph shows high test coverage gaps or a high incidence of historical defects, a test-specific review is generated.
[0040] It should be noted that after the review route is generated, one or more review nodes are created according to the review route scheme. Each node corresponds to AI pre-review, module leader review, project leader review, supplementary verification or multi-party review. The review nodes have states such as pending, under review, passed, requiring modification, supplementary verification, under review, blocked merging, manual takeover and timeout upgrade. The review conclusion is aggregated after all necessary nodes are completed. If any forced node outputs blocked merging or manual takeover, the overall conclusion cannot be automatically passed.
[0041] It should be noted that each review node is configured with a timeout limit. When a review node in the pending state exceeds the time limit without being claimed, an escalation notification is automatically sent to the node's superior and the node status is changed to timeout escalation. When a review node in the review process exceeds the time limit without outputting a conclusion, a timeout escalation is also triggered. After timeout escalation, if the review is still not completed within an additional grace period, the node is automatically changed to manual takeover status, whereby a higher-level manager intervenes or organizes a multi-party review.
[0042] Specifically, the review conclusions include at least approval, modification requests, supplementary verification, special review, blocking and merging, and manual takeover. After the review conclusions are generated, the system will write the review evidence package, scoring results, special risk labels, routing results, status of each review node, AI review opinions, manual review opinions, false positive / false negative flags, and final processing results into the audit ledger. If defects appear in the subsequent online or testing environment, the defects will be associated with the original review evidence package and review routing, and the review intensity threshold, special risk label rules, role routing rules, and model capability profiles for similar changes will be corrected.
[0043] In summary, this invention does not rely solely on code differences or merge request text to determine review intensity. Instead, it integrates multi-source evidence, including AI execution records, verification uncertainties, multi-model arbitration results, and impact maps, to generate a review evidence package. Based on this, it dynamically determines the review intensity level and generates specific risk labels. This provides a more complete evidence dimension and can identify risks that cannot be detected by code differences alone (such as AI execution uncertainties and deviations from task objectives), thereby reducing the probability of high-risk changes being underestimated or missed. Furthermore, low-risk changes (such as document modifications and comment corrections) can be automatically approved within seconds through basic AI checks without human intervention. High-risk changes automatically enter specialized manual review or multi-party review, reducing the waste of manual review resources and allowing limited expert review resources to be concentrated on high-risk changes. The invention also integrates task contract breach records, AI confidence levels, multi-model arbitration results, and verification uncertainties... Qualitative inclusion in the review intensity assessment can identify risks unique to AI-generated code (such as superficially passing tests but deviating from the task objective, weakened verification, and conflicting conclusions from multiple models). It supplements the review basis from the perspective of AI execution process, automatically determining whether special reviews are needed for security, database, architecture, testing, performance, configuration, compliance, etc., based on specific risk labels. It tracks the status of each review node and aggregates multiple review conclusions through a routing execution state machine, achieving structured tracking of the review process and complete aggregation of conclusions. This reduces the problems of misrouting of different types of risks or the irreversibility of the review process. Review conclusions, false positives / false negatives, rejection reasons, and subsequent defects can be written back to the strategy library and model capability profile, enabling subsequent review intensity assessment, labeling rules, and role routing to be gradually optimized. It can continuously learn from historical review experience and reduce the false positive and false negative rates in subsequent reviews.
[0044] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus.
[0045] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.
Claims
1. A method for dynamic review intensity and routing of AI code based on task context, characterized in that, Includes the following steps: S1. Receive code change or merge request: Receive a code change from the code platform, AI R&D execution platform or continuous integration system; S2. Collect review evidence: Collect task context, code differences, change files, test results, verification evidence, historical defects, task contract records, and impact graph information; S3. Generate a review evidence package: Unify and standardize evidence from different sources to generate a structured review evidence package specific to this change. S4. Calculate the review intensity score: The review intensity score is calculated based on the scope of impact of the change, business sensitivity, verification uncertainty, historical failure weight, AI-generated confidence risk, access to sensitive resources, and the weight of special risk labels. S5. Determine the review level: Map the review intensity score to review levels from S0 to S4; S6. Generate review routes: Determine the AI review, manual review, and special review paths to be performed based on the review level and specific risk labels; S7. Perform review and summarize conclusions: Summarize the results of AI review, special inspection and manual confirmation, and output conclusions such as pass, supplementary verification, re-review after modification, block and merge or manual takeover. S8. Review Conclusion Feedback: Write the review process and conclusions into the audit log, and write back the reasons for false reports, omissions, rejections and subsequent defects to the strategy library.
2. The AI code dynamic review strength and routing method based on task context according to claim 1, characterized in that, The review evidence package is used to drive the review intensity calculation and review route generation. The review evidence package includes at least task context, change information, execution evidence, verification evidence, impact graph evidence, knowledge credibility evidence, personnel responsibility evidence, and model arbitration evidence.
3. The AI code dynamic review strength and routing method based on task context as described in claim 2, characterized in that, The model arbitration evidence comes from the multi-model arbitration module. When multiple AI models are called to generate review opinions for the same code change, the multi-model arbitration module determines whether there is an evidence conflict according to consistency judgment, conflict judgment, conflict degree quantification and escalation processing.
4. The AI code dynamic review strength and routing method based on task context according to claim 3, characterized in that, The review intensity score Obtain it using the following formula: ; in, The review intensity score ranges from 0 to 10. This refers to the scope of impact of the change, normalized to a risk score of 0-10, and is used to measure the number of modules, files, and downstream dependencies involved in the code change; Business sensitivity refers to a risk score normalized to 0-10, used to measure the criticality of the business path involved in a change. This refers to the uncertainty in verification, which is normalized to a risk score of 0-10. It is used to measure factors including insufficient test coverage, number of automatic repair retries, and low knowledge credibility level. Historical fault weights, normalized to a risk score of 0-10, are used to measure the frequency and severity of defects in a module or file throughout history. This refers to the confidence risk generated by AI, normalized to a risk score of 0-10, used to measure the risks of path default, multiple rounds of repair, and low confidence levels during AI execution. Sensitive resource access is normalized to a risk score of 0-10 and used to measure whether changes involve sensitive resources such as production configurations, keys, personal information, and non-rollbackable operations. The specific risk label weight is normalized to a risk score of 0-10 and is used to measure the number and severity of specific risk labels that have been hit. This refers to the weighting coefficient for the scope of the change's impact; the default value is 0.
20. This refers to the weighting coefficient for business sensitivity, with a default value of 0.
20. This refers to the weighting coefficient for verifying uncertainty; the default value is 0.
15. This refers to the weighting coefficient for historical faults, with a default value of 0.
15. This refers to the weighting coefficient for the confidence risk generated by AI, with a default value of 0.10; This refers to the weighting coefficient for sensitive resource access, with a default value of 0.
10. This refers to the weighting coefficient of the specific risk label, with a default value of 0.
10.
5. The AI code dynamic review strength and routing method based on task context according to claim 4, characterized in that, When the scope of impact of the change, business sensitivity, verification uncertainty, historical failure weight, AI-generated confidence risk, sensitive resource accessibility, and special risk label weight are mapped from the original data to a risk score of 0-10, the following operations are performed: The scope of impact of the change is determined by the number of modules involved M and the number of downstream dependent services D. The original value is calculated as M×2 + D×3, and the normalized risk score is min(original value, 10). Business sensitivity is mapped based on the preset sensitivity level of the business paths involved in the change, and the highest value is taken when the change involves multiple business paths; To verify the uncertainty, we take the number of test coverage gaps G, the number of automatic repair retries R, and the knowledge credibility level penalty K, and calculate the risk score = min(G×2 + R×1.5 + K, 10); Historical fault weights are calculated by taking the number of defects N and the highest severity level S of the module in the past six months, and the risk score is calculated as min(N×S, 10). AI generates confidence risk. When it is not generated by AI, the risk score is 0. When it is generated by AI, the number of path defaults V and the low confidence level marker C are taken. The risk score is calculated as min(V×3 + C×4 + (repair rounds-1)×2, 10). Sensitive resource access is incremented based on the type of sensitive resource accessed: production configuration adds 4, key or signature adds 4, personal information adds 3, and non-rollbackable operation adds 5. The smaller value between the incremented value and 10 is taken. The weight of the special risk label is the sum of the number of special risk labels that have been hit and the preset severity coefficient of each label. The risk score is calculated as min(the sum of the severity coefficients of each label, 10).
6. The AI code dynamic review strength and routing method based on task context according to claim 5, characterized in that, The censorship intensity score is mapped to censorship levels from S0 to S4, including: When review intensity rating If the level is below 4.0, and the changes only involve documents, comments, or low-risk styles and have no business module connection to the display of the map, then the review level is determined to be S0. When review intensity rating If the score is below 4.0, and the code involves non-core logic but has a small impact, the review level is determined to be S1. When review intensity rating If the score is below 7.0 but not below 4.0, the review level is determined to be S2. When review intensity rating If the score is not lower than 7.0, the review level is determined to be S3; If any of the following are detected: changes to production environment configuration, core permission logic, non-rollbackable database migration, changes to billing and settlement logic, changes to personal information processing logic, changes to cross-service communication protocols, or highly conflicting AI multi-model outputs, or if multiple high-risk special tags are detected, the review level will be determined as S4.
7. The AI code dynamic review strength and routing method based on task context according to claim 6, characterized in that, The specific risk labels are structured routing criteria automatically generated or modified based on the review evidence package, wherein the specific risk labels are generated according to the following rules: When code changes involve authentication, permission checks, keys, signatures, sensitive data processing, login status, tenant isolation, or security configuration, generate a security audit label. When code changes involve table structure, indexes, migration scripts, field state transitions, transaction boundaries, query performance, or data consistency, a database audit label is generated. When code changes affect cross-service interfaces, common components, domain models, dependency upgrades, messaging protocols, or architectural boundaries, generate architectural review tags. When test coverage is insufficient, historical defects are frequent, verification results conflict, or there are gaps in the observation of the spectrum, a test review or supplementary verification label is generated. When code changes involve production configuration, pricing, settlement, personal information, or non-rollback operations, a compliance review label is generated. When a code change simultaneously triggers multiple high-risk special tag combinations, or triggers any mandatory blocking condition, a multi-party review tag is generated. When code changes involve hot paths, changes in algorithm complexity, changes in expected response time or throughput, adjustments to connection pool or cache parameters, or significant changes in resource consumption, a performance audit label is generated. When code changes involve environment variables, feature switches, rate limiting or circuit breaker parameters, log levels, middleware configurations, or deployment descriptors, generate configuration review tags.
8. The AI code dynamic review strength and routing method based on task context according to claim 7, characterized in that, After the review route is generated, one or more review nodes are created according to the review route scheme. Each node corresponds to AI pre-review, module leader review, project leader review, supplementary verification or multi-party review. The review nodes have states such as pending, under review, passed, requiring modification, supplementary verification, under review, blocked merging, manual takeover and timeout upgrade. The review conclusion is aggregated after all necessary nodes are completed. If any forced node outputs blocked merging or manual takeover, the overall conclusion cannot be automatically passed.
9. The AI code dynamic review strength and routing method based on task context according to claim 8, characterized in that, Each review node is configured with a timeout limit. When a review node in the pending state exceeds the time limit without being claimed, an escalation notification is automatically sent to the node's superior and the node status is changed to timeout escalation. When a review node in the review process exceeds the time limit without outputting a conclusion, a timeout escalation is also triggered. After timeout escalation, if the review is still not completed within an additional grace period, the node is automatically changed to manual takeover status, where a higher-level manager intervenes or organizes a multi-party review.
10. A task-context-based AI code dynamic review intensity and routing system, used to implement the task-context-based AI code dynamic review intensity and routing method as described in claim 9, characterized in that, include: The change input module is used to receive merge requests, branch changes, code differences, AI execution results, or R&D task outputs generated by the code platform. The evidence collection module is used to collect task context, code change information, task contract execution records, automatic repair verification results, change impact graph, knowledge trust level, tool access records, and historical defect records. The risk evidence fusion module is used to convert evidence from multiple sources into a unified change risk profile and generate a risk feature vector. The review intensity determination module is used to determine the review intensity level based on risk feature vectors, team strategies, and dynamic thresholds; The specialized risk label generation module is used to generate specialized risk labels for security, database, architecture, testing, performance, configuration, and compliance based on the changed object, resource type, impact graph, historical defects, verification gaps, and team strategies. The review routing module is used to generate review routing schemes for AI pre-review, module leader review, security review, database review, architecture review, test review, or multi-party review based on review intensity level, specific risk labels, role and responsibility table, and mandatory blocking conditions. The review routing execution module is used to track the status of each review node, including pending, passed, required modification, supplementary verification, under review, blocked merging, and manual takeover, and to summarize the review conclusions from multiple paths. The AI pre-review module is used to conduct an initial review of code quality, potential defects, security issues, testing gaps, boundary conditions, and consistency with task objectives. The human expert review module is used to assign high-risk or specific risk changes to the corresponding roles, including module owners, security owners, database owners, architecture owners, and test owners. The audit and tracking module is used to record evidence sources, scoring processes, review intensity, routing results, review conclusions, and manual confirmation information. The review conclusion feedback module is used to write back the review results, false alarm or omission markers, rejection reasons, approval reasons, supplementary verification results, and subsequent defect results to the review strategy library, special risk label rules, role routing rules, and model capability profile.