Open source code security detection system based on multi-llm collaborative regularization
Patent Information
- Application Number
- CN202610723312.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-25
- Publication Date
- 2026-09-25
AI Technical Summary
[0004]有鉴于此,本申请提供了一种基于多LLM协同正则化的开源代码安全检测系统,主要目的在于解决规则化工具僵化且滞后,成分分析工具无法触及业务逻辑,而单一的LLM模型又不可靠的问题
本申请提供的一种基于多LLM协同正则化的开源代码安全检测系统,本申请包括代码预处理模块、多LLM智能体分析模块、协同正则化模块和报告生成模块;代码预处理模块,用于接收待检测源代码,对待检测源代码进行解析切片处理,得到多个代码片段;多LLM智能体分析模块,用于将多个代码片段输入至多LLM智能体进行数据分析,得到目标漏洞候选集合;协同正则化模块,用于对目标漏洞候选集合进行结果聚类,对结果聚类后的目标漏洞候选集合进行证据加权与交叉验证,生成结构化加权证据包,将结构化加权证据包输入至仲裁者LLM智能体进行裁决,生成结构化裁决信息;报告生成模块,用于基于目标漏洞候选集合和结构化裁决信息生成待检测源代码的代码安全审计报告。通过引入多个专业化LLM智能体并行分析,模拟安全团队中渗透测试、代码审计、供应链安全等不同专家分工协作的场景,克服单一模型存在知识盲区的问题,能够从漏洞模式、数据流、业务逻辑、第三方依赖等多个维度对代码进行审查,显著提高漏洞的检出率。而且,协同正则化模块通过包含结果聚类、证据加权与交叉验证、仲裁者裁决的三级处理机制,系统性应对AI模型输出的不稳定性,极大地降低误报率,同时提升结果的准确性。
Smart Images

Figure CN122818359A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of network security technology, and in particular to an open-source code security detection system based on multi-LLM collaborative regularization. Background Technology
[0002] With the increasing complexity of software systems and the widespread application of open-source code, code security testing has become a crucial part of the software development lifecycle.
[0003] Among related technologies, automated detection methods mainly rely on the following three types of technologies: Traditional static application security testing tools scan source code based on predefined rules, pattern matching, taint analysis, or symbolic execution. However, these tools suffer from high false positive and false negative rates. Their rule bases are insufficient to cover complex business logic and new vulnerability patterns, and their understanding of the code's true intent and deep logical connections is weak, resulting in insufficient accuracy of analysis results. Software composition analysis tools scan project dependency lists and compare them with known vulnerability databases to identify known risks in third-party libraries. However, this technology cannot inherently discover unknown security vulnerabilities introduced by the code itself, nor can it detect improper use of security APIs by developers, thus its detection scope is fundamentally limited. Code security detection based on a single LLM model is also problematic. However, the output of a single LLM model is unstable during detection, and there may be knowledge blind spots for certain types of vulnerabilities, resulting in low detection effectiveness. When analyzing large and complex codebases, it is difficult to ensure the consistency of judgment standards and logic across different modules. Summary of the Invention
[0004] In view of this, this application provides an open-source code security detection system based on multi-LLM collaborative regularization. The main purpose is to solve the problems that rule-making tools are rigid and lagging, component analysis tools cannot reach business logic, and single LLM models are unreliable.
[0005] According to the first aspect of this application, an open-source code security detection system based on multi-LLM collaborative regularization is provided. The system includes a code preprocessing module, a multi-LLM agent analysis module, a collaborative regularization module, and a report generation module. The code preprocessing module is used to receive the source code to be detected, and to perform parsing and slicing processing on the source code to be detected to obtain multiple code fragments; The multi-LLM agent analysis module is used to input the multiple code snippets into the multi-LLM agent for data analysis to obtain a target vulnerability candidate set; The collaborative regularization module is used to perform result clustering on the target vulnerability candidate set, perform evidence weighting and cross-validation on the clustered target vulnerability candidate set, generate a structured weighted evidence package, input the structured weighted evidence package into the arbitrator LLM agent for adjudication, and generate structured adjudication information. The report generation module is used to generate a code security audit report of the source code to be detected based on the target vulnerability candidate set and the structured adjudication information.
[0006] By employing the above-described technical solution, the technical solution provided in this application has at least the following advantages: This application provides an open-source code security detection system based on multi-LLM collaborative regularization. The system includes a code preprocessing module, a multi-LLM agent analysis module, a collaborative regularization module, and a report generation module. The code preprocessing module receives the source code to be detected and performs parsing and segmentation to obtain multiple code fragments. The multi-LLM agent analysis module inputs the multiple code fragments into a multi-LLM agent for data analysis to obtain a target vulnerability candidate set. The collaborative regularization module performs result clustering on the target vulnerability candidate set, performs evidence weighting and cross-validation on the clustered target vulnerability candidate set, generates a structured weighted evidence package, inputs the structured weighted evidence package into an arbitrator LLM agent for adjudication, and generates structured adjudication information. The report generation module generates a code security audit report for the source code to be detected based on the target vulnerability candidate set and the structured adjudication information. By introducing multiple specialized LLM agents for parallel analysis, this approach simulates scenarios where different experts in a security team collaborate on penetration testing, code auditing, and supply chain security. This overcomes the knowledge blind spots inherent in single models and enables code review from multiple dimensions, including vulnerability patterns, data flow, business logic, and third-party dependencies, significantly improving vulnerability detection rates. Furthermore, the collaborative regularization module employs a three-tiered processing mechanism—including result clustering, evidence weighting and cross-validation, and arbitrator adjudication—to systematically address the instability of AI model outputs, greatly reducing false positives while simultaneously improving accuracy.
[0007] The above description is only an overview of the technical solution of this application. In order to better understand the technical means of this application and to implement it in accordance with the contents of the specification, and to make the above and other objects, features and advantages of this application more obvious and understandable, the following are specific embodiments of this application. Attached Figure Description
[0008] Various other advantages and benefits will become apparent to those skilled in the art upon reading the following detailed description of preferred embodiments. The accompanying drawings are for illustrative purposes only and are not intended to limit the scope of this application. Furthermore, the same reference numerals denote the same parts throughout the drawings. In the drawings: Figure 1 This illustration shows a schematic diagram of the architecture of an open-source code security detection system based on multi-LLM collaborative regularization provided in an embodiment of this application; Figure 2 A schematic diagram of the structure of a computer device provided in an embodiment of this application is shown. Detailed Implementation
[0009] In the description of this application, it should be understood that the terms "center", "longitudinal", "lateral", "length", "width", "thickness", "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", "clockwise", "counterclockwise", etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings. They are used only for the convenience of describing this application and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, they should not be construed as limitations on this application.
[0010] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature. In the description of this application, "multiple" means two or more, unless otherwise explicitly specified.
[0011] In this application, unless otherwise expressly specified and limited, the terms "installation," "connection," "linking," and "fixing," etc., should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral connection; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; and they can refer to the internal connection between two components. Those skilled in the art can understand the specific meaning of the above terms in this application according to the specific circumstances.
[0012] Exemplary embodiments of the present application will now be described in more detail with reference to the accompanying drawings. While exemplary embodiments of the present application are shown in the drawings, it should be understood that the present application may be implemented in various forms and should not be limited to the embodiments set forth herein. Rather, these embodiments are provided so that this application will be thorough and complete, and will fully convey the scope of the present application to those skilled in the art.
[0013] The existing technology is as follows: 1. Traditional static code analysis tools (SAST) rely on rules, pattern matching, taint analysis, symbolic execution, etc., with representative tools such as SonarQube and Checkmarx. Main drawbacks: slow rule base updates, inability to cover complex business logic and new vulnerabilities; difficulty in identifying and understanding the true intent and deep logical connections of the code, leading to misjudgments; and the need for security experts to continuously write and update detection rules.
[0014] 2. Software Component Analysis (SCA) tools scan a project's dependencies and compare them against known vulnerability databases such as NVD / CVE. Representative tools include Snyk and Dependabot. Main drawbacks: They cannot discover unknown vulnerabilities (0-day vulnerabilities), only publicly disclosed vulnerabilities. They cannot handle unknown security risks introduced by the code itself; they cannot identify improper use, even if dependent libraries are secure, they cannot detect insecure API calls made by developers.
[0015] 3. Code security inspection based on a single LLM uses a single large language model, such as GPT-4 or Claude, to directly audit the code. Main drawbacks: It may "imagine" non-existent vulnerabilities or give contradictory judgments on the same issue; there are knowledge gaps, as a single model may perform poorly on certain types of vulnerabilities (such as misuse of complex encryption algorithms, race conditions, etc.); when analyzing large and complex codebases, it is difficult to ensure consistency in judgment standards and logic across different modules; it lacks depth protection, as the judgment of a single model is a "single point," and once an error occurs, there is no error correction mechanism.
[0016] Therefore, existing technologies have limitations such as rigid rules, the ability to detect only known problems, and reliance on unreliable single AI models.
[0017] To address this issue, this application proposes an open-source code security detection system based on multi-LLM collaborative regularization. It constructs a heterogeneous multi-LLM agent ensemble, deploying various types of LLMs and training and configuring them with specific roles or specializations. Each agent focuses on detecting specific types of vulnerabilities (such as misuse of encryption algorithms, race conditions, etc.) or code analysis dimensions (such as syntax logic, business processes, etc.), forming a clearly defined detection network. Simultaneously, a collaborative regularization module is designed as the system's "arbitrator" and "coordinator." After receiving the outputs from multiple LLM agents, it integrates and adjudicates the results through cross-validation, confidence weighting, and logical consistency verification mechanisms, ultimately generating a high-confidence final report. Multi-model cross-validation effectively filters out illusions and misjudgments from individual models, improving detection accuracy. LLM agents with different specializations can complement each other's knowledge gaps, forming a more comprehensive vulnerability detection capability. The regularization mechanism ensures the consistency and reliability of the output results, avoiding randomness and enhancing result stability. Combined with LLMs possessing logical analysis and pattern matching capabilities, it can discover complex vulnerabilities requiring a deep understanding of the code context. The system relies on the computing power of servers to provide services to users. Servers can be standalone servers or servers that provide basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDN), and big data and artificial intelligence platforms.
[0018] This application provides an open-source code security detection system based on multi-LLM collaborative regularization, such as... Figure 1 As shown, the system includes a code preprocessing module 1, a multi-LLM agent analysis module 2, a collaborative regularization module 3, and a report generation module 4.
[0019] Code preprocessing module 1 receives the source code to be inspected, parses and slices it to obtain multiple code fragments. Code preprocessing module 1 transforms source code inputs of various forms, such as Git repositories, compressed packages, and code fragments, into structured analysis units that can be deeply understood by LLM.
[0020] The multi-LLM agent analysis module 2 is used to input multiple code snippets into multiple LLM agents for data analysis to obtain a target vulnerability candidate set. Among them, the first LLM agent 201 is used to match known vulnerability characteristics, the second LLM agent 202 is used to trace the complete propagation path of external input in the program and check whether it is adequately cleaned up before flowing into dangerous operations, the third LLM agent 203 is used to analyze the business logic, access control and exception handling path of the code to find deep defects such as permission bypass and condition race. The fourth LLM agent 204 is used to check whether the code uses deprecated or insecure third-party APIs.
[0021] The Collaborative Regularization Module 3 clusters the target vulnerability candidate set, performs evidence weighting and cross-validation on the clustered target vulnerability candidate set, generates a structured weighted evidence package, and inputs the structured weighted evidence package into the arbitrator LLM agent for adjudication, generating structured adjudication information. The Collaborative Regularization Module 3 does not simply summarize the results, but clusters, cross-validates, and weights evidence from multiple parties' data. Finally, a higher-level arbitrator LLM performs logical consistency verification and final adjudication based on all evidence and code context. This overcomes the instability and knowledge blind spots of a single AI model, significantly improving detection accuracy and reducing false positive rates while maintaining high automation efficiency, and outputting a complete chain of evidence and adjudication reasons.
[0022] Report generation module 4 is used to generate a code security audit report of the source code to be tested based on the target vulnerability candidate set and structured adjudication information.
[0023] Specifically, the code preprocessing module 1 is used to receive the source code to be detected, wherein the source code to be detected is any one of the following: source code repository address, source code compressed package, or source code fragment.
[0024] The source code to be checked undergoes preprocessing. If the source code is a repository address, a version control tool is used to download all files from the repository. Non-code files are then filtered out according to preset filtering rules. For example, Git or SVN can be used to download the complete repository, and non-code files such as binary files, logs, documents, and images can be excluded based on preset rules.
[0025] If the source code to be detected is a source code compressed package, the source code compressed package is decompressed, and non-code files in the decompressed source code compressed package are filtered out according to the preset filtering rules; the code preprocessing module 1 supports common formats such as zip, .tar.gz, .rar, etc., and can decompress them to a temporary directory, while checking the integrity of the compressed package, such as whether it is damaged or contains malicious files.
[0026] If the source code to be detected is a source code fragment, then the source code fragment is subjected to format standardization operations, which include unifying indentation, removing extra blank lines, and preserving comments.
[0027] The preprocessed source code to be tested is parsed into an abstract syntax tree and / or control flow graph using an open-source parser library, such as the Clang toolchain for C / C++ code. Specifically, the parser takes the plain text of the source code as input, performs lexical analysis, syntax analysis, and semantic analysis, and finally outputs an abstract syntax tree that reflects the syntactic structure of the code, such as function definitions, variable declarations, loop statements, and nesting relationships. Based on the abstract syntax tree, the parser can further generate a control flow graph, representing the execution order and conditional jump relationships between code blocks within functions, revealing all possible data flow paths of the program.
[0028] The abstract syntax tree and / or control flow graph are cut according to preset partitioning rules to obtain multiple code snippets. The preset partitioning rules refer to partitioning by function or class, based on class keywords and function names, thereby reducing the complexity of LLM processing and improving analysis efficiency and accuracy.
[0029] Specifically, the multi-LLM agent analysis module 2 includes a first LLM agent 201, a second LLM agent 202, a third LLM agent 203, a fourth LLM agent 204, and a vulnerability candidate result integration unit 205. The multi-LLM agent analysis module 2 does not rely on a single model, but instead incorporates four LLM agents that are professionally fine-tuned for different security analysis tasks. These agents function like a team of security experts, simultaneously and in parallel performing deep scans of the same batch of code snippets, but each focusing on completely different core dimensions.
[0030] The first LLM agent 201 is used to perform vulnerability pattern detection on multiple code snippets. It can quickly scan each code snippet, compare it with the built-in vulnerability feature library, identify code patterns with typical vulnerability characteristics, and obtain the first vulnerability candidate record corresponding to each code snippet.
[0031] The second LLM agent 202 is used to perform data flow tracing on multiple code snippets. Starting from identifying all external input points (i.e. pollution sources), it traces the transmission and operation of these data between variables line by line, checking whether they eventually flow into sensitive convergence points without being adequately "sterilized", thereby verifying potential injection vulnerabilities and obtaining the second vulnerability candidate record corresponding to each code snippet.
[0032] The third LLM agent 203 is used to review the logic defects of multiple code snippets, analyze conditional branches and exception handling, and look for logical defects such as permission verification bypass paths and unhandled boundary conditions, so as to obtain the third vulnerability candidate record corresponding to each code snippet.
[0033] The fourth LLM agent 204 is used to perform third-party library risk assessment on multiple code snippets. Combining the project's third-party library dependency information, it checks whether the calls to these library APIs in the code use APIs that are known to be deprecated, have vulnerabilities, or are insecure, and obtains the fourth vulnerability candidate record corresponding to each code snippet.
[0034] The vulnerability candidate result integration unit 205 is used to integrate the first vulnerability candidate record, the second vulnerability candidate record, the third vulnerability candidate record, and the fourth vulnerability candidate record corresponding to each code segment into a target vulnerability candidate set. The vulnerability candidate record contains four key pieces of information: vulnerability type, precise code location, confidence score, and specific analysis basis. The analysis basis is determined by the vulnerability type. For example, the analysis basis for SQL injection is "the pollution path from user input to SQL query", and the analysis basis for null pointer reference is "the call chain of unchecked null variables", etc.
[0035] Specifically, the first LLM agent 201 is used to input multiple code snippets into the first LLM agent 201. For each code snippet, the first LLM agent performs a semantic comparison between the code snippet and a known vulnerability pattern library, identifies at least one suspected vulnerability pattern in the code snippet that matches preset vulnerability features, and generates a vulnerability candidate analysis record for each suspected vulnerability pattern. The preset vulnerability features include string concatenation query features for identifying SQL injection vulnerabilities, direct output user data features for identifying cross-site scripting vulnerabilities, and system command execution features for identifying command injection vulnerabilities. The vulnerability candidate analysis record for the suspected vulnerability pattern is used to describe the specific pattern information of the suspected vulnerability pattern.
[0036] Determine the line number of each suspected vulnerability pattern and the first vulnerability type identifier, wherein the first vulnerability type identifier is any one of SQL injection vulnerability, cross-site scripting vulnerability, and command injection vulnerability.
[0037] Calculate the matching degree of each suspected vulnerability pattern. The product of the matching degree of each suspected vulnerability pattern and the confidence degree of the known vulnerability corresponding to each suspected vulnerability pattern is used as the confidence score of each suspected vulnerability pattern. The matching degree is the ratio of the number of keywords in the code snippet that belong to the suspected vulnerability pattern to the total number of keywords in the code snippet. The confidence degree of the known vulnerability is an adjustable coefficient with an adjustment range of 50%-100%.
[0038] The code line number of each suspected vulnerability pattern, the first vulnerability type identifier of each suspected vulnerability pattern, the confidence score of each suspected vulnerability pattern, and the vulnerability candidate analysis record of each suspected vulnerability pattern are used to generate the first vulnerability candidate record corresponding to the code fragment.
[0039] The first LLM agent 201 automatically receives code snippets and performs semantic comparisons against a built-in vulnerability pattern library containing typical characteristics of SQL injection, XSS, and command injection. Once a matching suspected vulnerability pattern is found, it precisely records the line number and type of the code and calculates a confidence score that quantifies the risk, determined by the code matching degree and an adjustable baseline confidence level. Finally, all this information is integrated to generate a structured first vulnerability candidate record. This enables automated and intelligent high-speed initial screening, quickly identifying vulnerability candidates in code that match known vulnerability characteristics, much like an experienced expert, and performing preliminary risk classification through confidence scoring, thereby greatly improving the coverage and efficiency of vulnerability detection.
[0040] Specifically, the second LLM agent 202 is used to analyze the flow path of data in the code, such as tracing parameters from a user's HTTP request to see if they enter the database query statement without being sanitized. The second LLM agent 202 is specifically designed to verify and discover data contamination vulnerabilities. Specifically, multiple code snippets are input into the second LLM agent 202. For each code snippet, the second LLM agent identifies external input points within the code snippet as taint source variables, resulting in multiple taint source variables. These external input points include file content, database query results, etc.
[0041] For each pollution source variable, the assignment, transmission and operation process of the pollution source variable is semantically analyzed to construct the data flow graph corresponding to the pollution source variable. The data flow graph is used to represent the multiple data propagation paths of data starting from the pollution source variable in the code logic.
[0042] By performing semantic analysis on each pollution source variable, multiple data flow diagrams are obtained, and multiple data propagation paths are extracted from these diagrams.
[0043] For each data propagation path, it is checked whether the data propagation path has undergone security processing operations to obtain the status assessment result of the data propagation path. The security processing operations include encoding operations, verification operations, and filtering operations.
[0044] Multiple sensitive convergence points were identified in the code snippet. These sensitive convergence points can be any one of the following: database query function, command execution function, or file operation function.
[0045] Extract target data propagation paths from multiple data propagation paths whose state assessment results indicate that they have not undergone effective security processing and whose endpoints belong to multiple sensitive convergence points, thereby obtaining at least one target data propagation path. Generate a vulnerability candidate analysis record for each target data propagation path, wherein the vulnerability candidate analysis record of the target data propagation path is used to describe the complete contamination path of the target data propagation path.
[0046] For each target data propagation path, the line number of the target data propagation path is determined based on the path start and path end, and the second vulnerability type identifier of the target data propagation path is determined. The second vulnerability type identifier is any one of SQL injection vulnerability, command injection vulnerability, and file operation vulnerability.
[0047] The clarity and unambiguity of the target data propagation path in the code logic are quantitatively evaluated to obtain the propagation path determinism evaluation value. The product of the propagation path determinism evaluation value and the operation type confidence score corresponding to the target data propagation path is used as the confidence score of the target data propagation path. The operation type confidence score is determined according to the operation type corresponding to the path endpoint of the target data propagation path. The operation type confidence score is an adjustable coefficient with an adjustment range of 50%-100%.
[0048] The code line number of each target data propagation path, the second vulnerability type identifier of each target data propagation path, the confidence score of each target data propagation path, and the vulnerability candidate analysis record of each target data propagation path are used to generate the second vulnerability candidate record corresponding to the code fragment.
[0049] The second LLM agent, 202, goes beyond surface features, deeply tracing the complete flow of external input data within the program. It not only discovers the entire attack path from the source of contamination to the dangerous operation, overcoming the limitation of simple pattern matching tools that can only identify features but not verify reachability, but also quantifies the risks to assign clear reliability levels, significantly reducing false positives. This allows the system to accurately identify hidden security vulnerabilities that require complex data flows to trigger, such as SQL injection and command injection, providing a depth of analysis capabilities far exceeding static pattern matching.
[0050] Specifically, the third LLM agent 203 does not focus on specific vulnerability patterns, but rather examines the business logic and control flow of the code for defects. For example, is it possible for user A to access user B's data through a certain logical branch? Specifically, multiple code snippets are input into the third LLM agent 203. For each code snippet, the conditional branch structure is parsed to identify all possible execution paths, constructing a conditional branch graph. This graph contains multiple possible execution paths, and the conditional branch structures include if / else, switch / case, etc.
[0051] At least one possible execution path containing sensitive operations is identified from multiple possible execution paths. Among the at least one possible execution path containing sensitive operations, a possible execution path with privilege bypass risk is identified as a privilege bypass path, thus obtaining at least one privilege bypass path. A vulnerability candidate analysis record is generated for each privilege bypass path. Sensitive operations include modifying the database, accessing files, and executing privileged commands. Privilege bypass risks include bypassing through abnormal branches, bypassing through default branches, and bypassing through condition combinations. The vulnerability candidate analysis record of the privilege bypass path is used to describe the specific logical defects of the privilege bypass path.
[0052] The first ratio is the ratio of the number of paths with at least one permission bypass path to the number of paths with multiple possible execution paths.
[0053] For each privilege bypass path, determine the line number of the privilege bypass path and the third vulnerability type identifier. The product of the confidence score of the sensitive operation type corresponding to the privilege bypass path and the first proportion is used as the confidence score of the privilege bypass path. The confidence score of the sensitive operation type corresponding to the privilege bypass path is determined according to the sensitive operation type corresponding to the privilege bypass path. The confidence score of the sensitive operation type is an adjustable coefficient with an adjustment range of 50%-100%.
[0054] Perform combined condition coverage testing on the conditional branch graph to generate multiple combinations of input boundary values and obtain multiple test outliers. Combine the multiple test outliers and multiple combinations of input boundary values as a test case set. Among them, the multiple test outliers include null, empty value, extreme value, and illegal type value.
[0055] The code snippets are tested using a test case set to obtain a test result set. Based on the test result set, test cases belonging to unhandled boundary cases are extracted from the multiple test cases included in the test case set as unhandled boundary defects, resulting in multiple unhandled boundary defects. A vulnerability candidate analysis record is generated for each unhandled boundary defect. The vulnerability candidate analysis record of the unhandled boundary defect is used to describe the specific logical defect of the unhandled boundary defect.
[0056] The ratio of the total number of untreated boundary defects to the total number of test cases is used as the second proportion.
[0057] For each unprocessed boundary defect, the line number of the unprocessed boundary defect and the fourth vulnerability type identifier are determined. The confidence score of the sensitive operation type corresponding to the unprocessed boundary defect is the product of the second proportion. The confidence score of the sensitive operation type corresponding to the unprocessed boundary defect is determined based on the sensitive operation type corresponding to the unprocessed boundary defect. The confidence score of the sensitive operation type is an adjustable coefficient with an adjustment range of 50%-100%.
[0058] The code line number of each privilege bypass path, the third vulnerability type identifier of each privilege bypass path, the confidence score of each privilege bypass path, the vulnerability candidate analysis record of each privilege bypass path, the code line number of each unprocessed boundary defect, the fourth vulnerability type identifier of each unprocessed boundary defect, the confidence score of each unprocessed boundary defect, and the vulnerability candidate analysis record of each unprocessed boundary defect are used to generate the third vulnerability candidate record corresponding to the code fragment.
[0059] The third LLM agent, 203, transcends surface-level code analysis, focusing on uncovering deep-seated business logic flaws such as access control and exception handling. By constructing a complete execution path graph, it can systematically discover various hidden privilege bypass vulnerabilities—blind spots that traditional scanning tools struggle to cover. Through automated boundary value combination and outlier attack testing, it can effectively assess code robustness, identifying potential crash points and undefined behaviors. More importantly, through scientific ratio calculations (such as path ratios and anomaly trigger ratios) and adjustable confidence models, it quantifies the discovered logical risks, thus achieving a leap from vulnerability detection to risk quantification assessment, significantly enhancing the depth of auditing and its decision support value.
[0060] Specifically, the fourth LLM agent 204 is used to input multiple code snippets into the fourth LLM agent 204, and for each code snippet, multiple calls to the third-party library application programming interface are identified in the code snippet.
[0061] For each call, the call is matched against a risky application programming interface (API) knowledge base, which contains multiple risky APIs that are obsolete, insecure, or have common vulnerabilities.
[0062] If a match is successful, the call is treated as a risky call. The line number of the risky call, the name of the risky application programming interface, and the fifth vulnerability type identifier are determined. A vulnerability candidate analysis record is generated for each risky call, and the confidence score of the corresponding obsolete and insecure API is obtained as the confidence score value of the risky call. The vulnerability candidate analysis record of the risky call is used to describe the specific risky API information of the risky call. The confidence score of the obsolete and insecure API is an adjustable coefficient with an adjustment range of 50%-100%.
[0063] The fourth vulnerability candidate record corresponding to each code snippet is generated using the line number of each risky call, the name of the risky application programming interface for each risky call, the fifth vulnerability type identifier for each risky call, the confidence score for each risky call, and the vulnerability candidate analysis record for each risky call.
[0064] The fourth LLM agent, 204, is specifically responsible for reviewing the security of third-party dependency libraries used in the code. Its workflow involves automatically scanning code snippets for all calls to third-party library APIs, combining project dependency information, and precisely matching these calls against a continuously updated risk API knowledge base containing records of known obsolete, insecure, or vulnerable APIs. Once a risky call is detected, its precise line number, specific obsolete API name, and vulnerability type are automatically recorded, generating a vulnerability candidate record with detailed analysis. The fourth LLM agent, 204, can quickly locate hidden risk points using "outdated" or "vulnerable" APIs from massive amounts of code, solving the pain point of manually tracking the security status of third-party libraries. Simultaneously, its unique adjustable confidence score mechanism (e.g., setting the confidence level of high-risk APIs to 100% and those that are merely obsolete to 60%) allows users to flexibly balance the stringency and false positive rate of detection according to their actual risk preferences, thereby significantly improving the accuracy and operability of detecting third-party dependency vulnerabilities while ensuring audit efficiency.
[0065] Specifically, the collaborative regularization module 3 is used to cluster the target vulnerability candidate set. Vulnerability candidates pointing to the same code location and with similar vulnerability types are added to the same cluster group, resulting in multiple vulnerability candidate groups. Here, a vulnerability candidate is defined as any one of the following: suspected vulnerability pattern, target data propagation path, privilege circumvention path, unhandled boundary defect, or risky call. The same code location refers to the same file and line number. It should be noted that if different types of vulnerability candidates exist at the same code location, they are grouped into a separate category for subsequent conflict resolution.
[0066] For each vulnerability candidate group, the vulnerability candidate analysis record of each vulnerability candidate in the vulnerability candidate group is extracted. According to the type of the vulnerability candidate group, the vulnerability candidate analysis records of multiple vulnerability candidates are subjected to consistency verification and / or contradiction ruling. The confidence scores of multiple vulnerability candidates are adjusted based on the results of consistency verification and / or contradiction ruling.
[0067] For each vulnerability candidate in the vulnerability candidate group, obtain the current confidence score of the vulnerability candidate. Based on the agent-vulnerability type weight mapping table, determine the agent-vulnerability type weight corresponding to the vulnerability candidate using the agent identifier and vulnerability type identifier of the vulnerability candidate. The product of the current confidence score of the vulnerability candidate and the agent-vulnerability type weight corresponding to the vulnerability candidate is used as the weighted evidence score of the vulnerability candidate.
[0068] The total weighted evidence score of a vulnerability candidate group is calculated using the weighted evidence scores of multiple vulnerability candidates. In a consistency verification scenario, the total weighted evidence score of the vulnerability candidate group equals the sum of the weighted evidence scores of all LLM agents; in a conflict resolution scenario, it equals the weighted evidence score of the LLM agent with the highest weight. Specifically, when vulnerability candidate groups belong to the same code location and have similar vulnerability types, the sum of the weighted evidence scores of multiple vulnerability candidates is used as the total weighted evidence score of the vulnerability candidate group. When vulnerability candidate groups belong to the same code location but have different vulnerability types, or when they belong to the same code location and have similar vulnerability types but conflicting vulnerability candidate analysis records, the LLM agent with the highest weight is determined from among the multiple LLM agents according to the agent-vulnerability type weight mapping table, and the sum of the weighted evidence scores of at least one vulnerability candidate belonging to the LLM agent is used as the total weighted evidence score of the vulnerability candidate group.
[0069] By performing consistency verification and / or conflict resolution on each vulnerability candidate group, we obtain the total weighted evidence score for each vulnerability candidate group, as well as the weighted evidence scores of multiple vulnerability candidates within each vulnerability candidate group.
[0070] The structured weighted evidence package is generated from multiple code snippets, the target vulnerability candidate set, the weighted evidence scores of multiple vulnerability candidates in each vulnerability candidate group, and the total weighted evidence score of each vulnerability candidate group.
[0071] The structured weighted evidence package is input into the LLM arbitrator agent for adjudication, generating structured adjudication information. The structured adjudication information includes the final judgment result, comprehensive confidence level, and adjudication reason for each vulnerability candidate. The final judgment result is any one of the following: confirmed as a vulnerability, marked as a suspected vulnerability, or determined as a false alarm.
[0072] The Arbitrator LLM agent is a large language model execution with stronger reasoning capabilities. Instead of performing basic code scanning, it acts like a judge, making logical decisions on a structured, weighted evidence package.
[0073] The structured weighted evidence package includes core code context, multiple testimonies, weighted influencing factors, and descriptions of contradictions; Core code context: This contains the precise code snippet where the potential vulnerability is located, as well as the upstream (definition, assignment) and downstream (use, call) code necessary to understand that snippet, and is obtained from code preprocessing module 1.
[0074] Multiple testimonies: Present the independent analysis results of the code at this location by all relevant LLM agents (first LLM agent 201, second LLM agent 202, third LLM agent 203, and fourth LLM agent 204) in the form of a structured list, including at least: agent identifier, vulnerability candidate analysis record, and confidence score value.
[0075] Weighting Influence Factor: The weight coefficient of each LLM agent under this vulnerability type, calculated according to the dynamic weighting model.
[0076] Conflict Description: The system-generated natural language text clearly points out the core logical conflict within the current evidence package. For example, "LLM agent A reports an SQL injection here, but analysis by LLM agent C shows that the conditions required to trigger this code path are unattainable in the actual business logic."
[0077] The specific execution process of the arbitrator LLM agent is as follows: Phase 1: The structured weighted evidence package is transformed into structured instruction inputs and given to the arbitrator LLM agent. These instructions not only contain code but also mandate that the arbitrator LLM agent employ "thought chain" reasoning, explicitly requiring it to "reason step by step" and demonstrate its thought process from evidence to conclusion. This is to prevent the model from "jumping" to intuitive conclusions and ensure the interpretability and logical rigor of the ruling. For example, LLM agent A (weight 0.8) believes there is SQL injection in line 45, with the evidence being string concatenation; LLM agent B (weight 0.9) believes the variable has been cleaned in line 20. Based on the code context, determine whether the cleansing logic of LLM agent B is valid and provide a final judgment.
[0078] Phase Two: The arbitrator LLM agent evaluates whether the vulnerability candidate analysis records provided by each LLM agent are logically valid. For example, it checks whether the "cleaning function" claimed by LLM agent B exists and is valid in the actual code. If a high-weighted LLM agent provides conclusive evidence of cleanup, the arbitrator LLM agent will overturn the alert from a low-weighted LLM agent.
[0079] Phase 3: The arbitrator LLM agent outputs the final ruling, including the final judgment result (confirmed as a vulnerability / marked as a suspected vulnerability / judged as a false positive), the overall confidence level, and the reasoning for the ruling. The overall confidence level is a final probability value calculated based on the weighted scores of each LLM agent and the arbitrator LLM agent's own judgment logic, ranging from 0% to 100%. The reasoning for the ruling is a natural language description explaining why it was judged as a vulnerability or a false positive. For example, "Although there is string concatenation on line 45 of the code (adopting the evidence from LLM agent A), its input parameter is forcibly converted to an integer by the intval() function on line 20 (verifying and adopting the evidence from LLM agent B), therefore there is no SQL injection risk at this point, and it is judged as a false positive."
[0080] The Collaborative Regularization Module 3, through its multi-expert parallel consultation and final arbitration architecture, systematically addresses the inherent deficiencies of traditional single tools in terms of accuracy, coverage, and reliability. Like an intelligent "expert review panel," Module 3 not only automatically integrates independent reports from experts across different dimensions, enhancing the certainty of identifying genuine vulnerabilities through "evidence corroboration," but also resolves conflicts of opinion among experts through a "conflicting adjudication" mechanism, and relies on a higher-level "arbitrator LLM" for final logical judgment. Ultimately, it outputs an authoritative audit report with high detection rate, low false alarm rate, and interpretable results, achieving a qualitative leap from "single-point alarm" to "credible collective decision-making."
[0081] Specifically, the collaborative regularization module 3 is used to detect the type of each vulnerability candidate group. If vulnerability candidate groups belong to the same code location and have similar vulnerability types, then the vulnerability candidate analysis record for each vulnerability candidate in the vulnerability candidate group is extracted. Logical verification is performed on the vulnerability candidate analysis records of multiple vulnerability candidates to enhance the credibility of the vulnerability. Based on the logical verification processing rules, the confidence scores of multiple vulnerability candidates are adjusted. Specifically, each clustering result is traversed, and vulnerability detection evidence from all LLM agents in that clustering result is extracted to determine whether the evidence has a complementary or corroborating relationship. Specifically, it is checked whether the evidence from different LLM agents describes different links in the vulnerability cause or exploitation chain and can be logically linked. For example, for the same code location, if the evidence from "Pattern Hunter A" is "discovered SQL statement concatenation pattern," while the evidence from "Data Flow Tracker B" is "confirmed that external user input can flow into this concatenation point without disinfection," then the two pieces of evidence respectively indicate "vulnerability characteristics" and "attack path feasibility," forming a complete evidence chain from static characteristics to dynamic data flow, constituting strong complementary corroboration; it is also checked whether multiple agents independently reach the same conclusion and whether their core basis is consistent. For example, if "Pattern Hunter A", "Logic Examiner C" and another analyzer all report the existence of a "null pointer reference" at the same location, and their analysis basis all points to "variable X was not checked for null after initialization but before being called on line N", then this constitutes multi-source independent verification, which significantly enhances the certainty of the existence of the defect.
[0082] The system dynamically increases the overall confidence level of the vulnerability represented by the cluster based on the type of corroboration relationship and the number and weight of the participating agents. First, it determines whether there are two or more LLM agents within a cluster whose analytical evidence logically forms a complementary or strongly correlated corroboration relationship. For example, LLM agent A proves the existence of a vulnerability pattern, and LLM agent B proves the existence of a triggerable data flow. If the evidence from two LLM agents corroborates each other, the confidence score of the vulnerability candidate is increased by 50%. If the evidence from more than two LLM agents corroborates each other, the evidence is more sufficient, and the confidence score of the vulnerability candidate is increased by 80%. If only one LLM agent provides vulnerability evidence, and no other LLM agents' reports corroborate it, the confidence score of the vulnerability candidate remains unchanged. Simultaneously, the system automatically adds a "single-source evidence" status marker and adjusts its priority in subsequent processes to "requires further verification," which usually means it will enter the "contradictory decision" process or be recommended to security experts for manual review.
[0083] If vulnerability candidate groups belong to the same code location but have different vulnerability types, then the vulnerability candidate analysis record for each vulnerability candidate in the group is extracted. Logical coexistence verification is performed on the vulnerability candidate analysis records of multiple vulnerability candidates to determine whether different vulnerability types can coexist. For example, SQL injection and XSS vulnerabilities cannot coexist in the same line of code, as this is a complete contradiction; while null pointer references and logical errors can coexist, which is a partial contradiction. Then, based on the logical coexistence verification processing rules, the confidence scores of multiple vulnerability candidates are adjusted.
[0084] Specifically, in cases of complete contradiction, if different vulnerability types are logically mutually exclusive in attack vectors, data flows, or triggering conditions, they are determined to be completely contradictory. In this case, the system invokes a dynamic weighting model to query the historical authority weights of agents reporting these contradictory types for that type. The adjudication strategy is as follows: evidence from high-weight agents is adopted, and the confidence scores of vulnerability candidates corresponding to low-weight agents are reduced to below a confidence threshold and marked as "suspected false positive." The confidence threshold is a configurable value; vulnerability candidates below this threshold will be automatically downgraded or filtered by the system. This mechanism effectively controls the spread of false positives. If the weight values of both parties are close, no automatic adjudication is performed. Instead, the issue, the complete chain of evidence, and the weight comparison are packaged and marked as "requires manual verification." It should be noted that when marking "requires manual verification," the system provides not a simple conclusion, but a structured evidence package containing conflict point location, original evidence from each party, weight comparison, and code context, greatly improving the efficiency of manual review.
[0085] For some contradictory / coexistent situations, if the vulnerability type may be caused by different flaws in the same code segment, it is determined to be coexistent. The system retains all reasonable candidates, but finely adjusts their confidence scores: for risk points corresponding to mutually supporting aspects of the evidence, the confidence score is slightly increased; for aspects that are directly conflicting or uncertain, the confidence score is slightly decreased.
[0086] If vulnerability candidate groups belong to the same code location and have similar vulnerability types but conflicting vulnerability candidate analysis records, then the vulnerability candidate analysis record for each vulnerability candidate in the group is extracted. Logical vulnerability existence verification is then performed on the vulnerability candidate analysis records of multiple vulnerability candidates. For example, LLM agent A believes "the variable is not null," while LLM agent B believes "the variable has been implicitly nulled through a global function." Then, based on the logical vulnerability existence verification processing rules, the confidence scores of multiple vulnerability candidates are adjusted. If the analysis basis of an LLM agent is found to have obvious logical errors or is inconsistent with the code facts, then that evidence and its corresponding vulnerability candidate are directly eliminated. If the analysis basis of both parties is logically consistent, then the weighted model is used again for arbitration, and the judgment of the party with the higher weight is adopted as the preliminary conclusion. Similarly, if the weights are similar, they are marked as "requiring manual confirmation."
[0087] The Collaborative Regularization Module 3, through evidence verification and conflict resolution mechanisms, achieves intelligent fusion and arbitration of multi-agent analysis results, thereby systematically improving the overall accuracy and credibility of vulnerability determination while maintaining automation efficiency. Specifically, when evidence from different experts corroborates each other (such as pattern matching and data flow tracing forming a complete chain of evidence), the system can quantify and increase vulnerability confidence, effectively enhancing the credibility of real vulnerabilities. When expert opinions conflict (such as contradictions in the vulnerability type or judgment criteria for the same code), the system can make intelligent decisions based on a dynamic weight model and logical verification (adopting opinions from experts with high weights, reducing or eliminating evidence with logical errors, or marking evidence as "requiring manual confirmation" when weights are close). This greatly reduces the false positive rate and ensures that the final output is a rigorously reviewed, interpretable, and authoritative collective decision, rather than a simple summary of opinions.
[0088] Since different agents have varying detection capabilities for different vulnerability types, a dynamic weighting model is used to dynamically calculate the agent's weight for specific vulnerability types, thereby achieving differentiated weighting of evidence and improving the accuracy of the fusion results. To this end, a dynamic weighting model maintains a dynamic weight for each "agent-vulnerability type" combination. This model is calculated based on historical detection accuracy, recall, and other metrics, and is optimized through continuous human review and feedback.
[0089] The specific calculation process of the dynamic weight model is as follows: Step 1: Data Initialization. Collect historical audit data, including the detection results of each agent under different vulnerability types (such as SQL injection, XSS, null pointer reference, etc.), specifically covering true positives (real vulnerability cases), false positives (misjudgments, i.e., false vulnerabilities are misjudged as real vulnerabilities), and false negatives (misjudgments, i.e., real vulnerabilities are misjudged as false vulnerabilities).
[0090] Step 2: Calculation of core metrics. For each "agent-vulnerability type" combination, calculate two core metrics: 1. Precision = number of true positives / (number of true positives + number of false positives), used to measure the accuracy of the detection results; 2. Recall = number of true positives / (number of true positives + number of false negatives), used to measure the comprehensiveness of the detection.
[0091] Step 3: Weight Fusion. A weighted average method is used to fuse precision and recall to obtain the initial weights for the "agent-vulnerability type" combination. The formula is: Initial Weights accuracy Recall rate, of which This is a weighting coefficient, with a value ranging from 0.7 to 0.9. It can be flexibly adjusted according to the business's needs for "accuracy priority" or "comprehensiveness priority".
[0092] Step 4: Dynamic Optimization. A continuous learning mechanism is introduced. After each manual review or system verification, historical data is updated based on the review results: if the agent's detection result is confirmed as correct, the weight of the corresponding "agent-vulnerability type" combination will be increased; if the detection result is confirmed as a false positive or false negative, the corresponding weight will be decreased. The optimization formula is: Updated Weight (Review result score -0.5), among which, The learning rate is 0.1-0.2; the scoring rules for the review results are: 1 for a correct result and 0 for an incorrect result.
[0093] Dynamic weight model calculation results: The dynamic weight of each agent for different vulnerability types is obtained, with values ranging from 0 to 1. The higher the weight, the more reliable the detection results of the corresponding agent on that vulnerability type. For example, "Data Flow Tracker B" has a weight of 0.9 on the SQL injection type, while "Logic Auditor C" has a weight of 0.4 on the SQL injection type, which means that the former's SQL injection detection results are more reliable.
[0094] The dynamic weighting model not only accurately reflects the expertise and authority of different agents in various vulnerability detection methods, but more importantly, through a continuous learning mechanism, it can adjust weights in real time and with fine precision based on feedback from each human review. This allows the entire system to continuously optimize itself and become more accurate with use. This fundamentally ensures that the system can make scientific, credible, and dynamically adaptable intelligent decisions that align with business preferences during evidence fusion and conflict resolution, thereby maximizing the accuracy and efficiency of collaborative analysis.
[0095] Specifically, report generation module 4 is used to generate a code security audit report for the source code to be inspected based on the target vulnerability candidate set and structured adjudication information. The code security audit report includes a list of identified vulnerabilities, detailed information for each vulnerability, a complete chain of evidence, and remediation recommendations. The specific determination process is as follows: The confirmed vulnerability list is filtered based on the final judgment results and overall confidence level of the structured adjudication information, removing entries judged as "false positives" to ensure a clean list. Subsequently, the remaining confirmed and suspected vulnerabilities are automatically classified according to their overall confidence level: those with an overall confidence level of 80% or higher are classified as high-confidence vulnerabilities and require priority remediation; those with an overall confidence level between 50% and 80% are classified as vulnerabilities requiring manual review and expert confirmation is recommended. The final generated list clearly displays the standard vulnerability name, vulnerability type, associated file path, and line number for each vulnerability, providing users with a risk priority list that allows for immediate action.
[0096] Detailed information for each vulnerability: In addition to providing the precise location and standard type of the vulnerability, its core innovation lies in the dynamic assessment of severity level. This level is not simply mapped to the vulnerability type, but is generated by the arbitrator LLM intelligent agent during the adjudication process, combining the theoretical harm of the vulnerability with the actual environmental context in which the code exists. This makes the risk assessment results more closely aligned with business realities. For example, a theoretically high-risk vulnerability that only exists in a test environment without external access may have its level dynamically adjusted to "medium risk".
[0097] A complete chain of evidence: A structured traceability chain is organized for each confirmed vulnerability, comprising three levels: the discovery layer records the agent that first flagged the problem, such as "reported by LLM agent A"; the corroboration layer presents the cross-validation results of other LLM agents, such as "LLM agent B confirmed the data flow," "LLM agent C did not report a logical problem"; and the adjudication layer directly cites the reasoning of the arbitrator LLM agent, clarifying the core logic behind the final acceptance or rejection of certain evidence. This chain of evidence fully records the entire process from problem discovery to final adjudication, greatly enhancing the credibility and interpretability of the report.
[0098] Remediation Recommendations: These recommendations are not derived from a static knowledge base, but rather generated in real-time by the arbitrator LLM agent after the final ruling is made, using code snippets specific to the vulnerability. The generation process is context-aware, providing syntactically correct remediation code based on the code language and business logic. It is also highly targeted; for example, if the vulnerability is SQL injection, the recommendation will explicitly state "replace the string concatenation logic on line 45 with parameterized queries"; if the vulnerability is a third-party library risk, the recommendation will explicitly state "upgrade..." Ku Zhi The report will also provide specific instructions for the "version". In addition, it will briefly assess the potential side effects of the fix, such as how new filtering rules might affect legitimate input, helping developers make more informed decisions when making fixes.
[0099] This application provides an open-source code security detection system based on multi-LLM collaborative regularization. Compared with the prior art, this application includes a code preprocessing module, a multi-LLM agent analysis module, a collaborative regularization module, and a report generation module. The code preprocessing module receives the source code to be detected, parses and slices it to obtain multiple code fragments. The multi-LLM agent analysis module inputs the multiple code fragments into a multi-LLM agent for data analysis to obtain a target vulnerability candidate set. The collaborative regularization module performs result clustering on the target vulnerability candidate set, performs evidence weighting and cross-validation on the clustered target vulnerability candidate set, generates a structured weighted evidence package, inputs the structured weighted evidence package into an arbitrator LLM agent for adjudication, and generates structured adjudication information. The report generation module generates a code security audit report for the source code to be detected based on the target vulnerability candidate set and the structured adjudication information. By introducing multiple specialized LLM agents for parallel analysis, this approach simulates scenarios where different experts in a security team collaborate on penetration testing, code auditing, and supply chain security. This overcomes the knowledge blind spots inherent in single models and enables code review from multiple dimensions, including vulnerability patterns, data flow, business logic, and third-party dependencies, significantly improving vulnerability detection rates. Furthermore, the collaborative regularization module employs a three-tiered processing mechanism—including result clustering, evidence weighting and cross-validation, and arbitrator adjudication—to systematically address the instability of AI model outputs, greatly reducing false positives while simultaneously improving accuracy.
[0100] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties.
[0101] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0102] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of this patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.
[0103] In an exemplary embodiment, see Figure 2 The document also provides a computer device including a bus, a processor, a memory, and a communication interface. It may also include an input / output interface and a display device, wherein the various functional units can communicate with each other via the bus. The memory stores computer programs, and the processor executes the programs stored in the memory, performing the open-source code security detection system based on multi-LLM collaborative regularization described in the above embodiments. The memory is divided into four layers from top to bottom: the top layer consists of application software modules that directly address functional requirements; the second layer is the application programming interface (API) that provides the application layer with interfaces to call underlying functions; the third layer is middleware that provides general services and connects software components at the upper and lower layers; and the bottom layer is the kernel, which is the foundational core of the software system and is responsible for underlying functions such as hardware resource management and process scheduling.
[0104] A computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the steps of the open-source code security detection system based on multi-LLM collaborative regularization.
[0105] Through the above description of the embodiments, those skilled in the art can clearly understand that this application can be implemented in hardware or by using software plus necessary general-purpose hardware platforms. Based on this understanding, the technical solution of this application can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as a CD-ROM, USB flash drive, external hard drive, etc.) and includes several instructions to cause a computer device (such as a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments of this application.
[0106] Those skilled in the art will understand that the accompanying drawings are merely schematic diagrams of a preferred embodiment, and the modules or processes shown in the drawings are not necessarily essential for implementing this application.
[0107] Those skilled in the art will understand that the modules in the apparatus of the implementation scenario can be distributed within the apparatus of the implementation scenario as described, or they can be located in one or more apparatuses different from this implementation scenario, with corresponding changes. The modules of the above-described implementation scenario can be combined into one module, or they can be further divided into multiple sub-modules.
[0108] The serial numbers in this application are for descriptive purposes only and do not represent the superiority or inferiority of the implementation scenario.
[0109] The above disclosures are only a few specific implementation scenarios of this application. However, this application is not limited to these. Any variations that can be conceived by those skilled in the art should fall within the protection scope of this application.
Claims
1. An open-source code security detection system based on multi-LLM collaborative regularization, characterized in that, It includes a code preprocessing module, a multi-LLM agent analysis module, a collaborative regularization module, and a report generation module; The code preprocessing module is used to receive the source code to be detected, and to perform parsing and slicing processing on the source code to be detected to obtain multiple code fragments; The multi-LLM agent analysis module is used to input the multiple code snippets into the multi-LLM agent for data analysis to obtain a target vulnerability candidate set; The collaborative regularization module is used to perform result clustering on the target vulnerability candidate set, perform evidence weighting and cross-validation on the clustered target vulnerability candidate set, generate a structured weighted evidence package, input the structured weighted evidence package into the arbitrator LLM agent for adjudication, and generate structured adjudication information. The report generation module is used to generate a code security audit report of the source code to be detected based on the target vulnerability candidate set and the structured adjudication information.
2. The system according to claim 1, characterized in that, The code preprocessing module is used to receive the source code to be detected, which is any one of the following: source code repository address, source code compressed package, or source code fragment; The source code to be detected is preprocessed to parse the preprocessed source code into an abstract syntax tree and / or a control flow graph. The abstract syntax tree and / or control flow graph are cut according to a preset partitioning rule to obtain the multiple code fragments. The preset partitioning rule refers to partitioning based on functions or classes, according to class keywords and function names.
3. The system according to claim 2, characterized in that, The code preprocessing module is used to download all files from the source code repository corresponding to the source code repository address using a version control tool if the source code to be detected is a source code repository address, and filter out non-code files from all files according to preset filtering rules. If the source code to be detected is a source code compressed package, then the source code compressed package is decompressed, and non-code files in the decompressed source code compressed package are filtered out according to the preset filtering rules; If the source code to be detected is a source code fragment, then a format standardization operation is performed on the source code fragment.
4. The system according to claim 1, characterized in that, The multi-LLM agent analysis module includes a first LLM agent, a second LLM agent, a third LLM agent, a fourth LLM agent, and a vulnerability candidate result integration unit. The first LLM agent is used to perform vulnerability pattern detection on the multiple code segments to obtain a first vulnerability candidate record corresponding to each code segment. The second LLM agent is used to perform data flow tracing on the multiple code segments to obtain a second vulnerability candidate record corresponding to each code segment; The third LLM agent is used to perform logical defect review on the multiple code segments and obtain a third vulnerability candidate record corresponding to each code segment. The fourth LLM agent is used to perform third-party library risk assessment on the multiple code segments and obtain a fourth vulnerability candidate record corresponding to each code segment. The vulnerability candidate result integration unit is used to take the first vulnerability candidate record, the second vulnerability candidate record, the third vulnerability candidate record, and the fourth vulnerability candidate record corresponding to each code segment as the target vulnerability candidate set.
5. The system according to claim 4, characterized in that, The first LLM agent is used to input the plurality of code fragments into the first LLM agent. For each code fragment, the first LLM agent performs a semantic comparison between the code fragment and a known vulnerability pattern library, identifies at least one suspected vulnerability pattern in the code fragment that matches preset vulnerability features, and generates a vulnerability candidate analysis record for each suspected vulnerability pattern. The preset vulnerability features include string concatenation query features for identifying SQL injection vulnerabilities, direct output user data features for identifying cross-site scripting vulnerabilities, and system command execution features for identifying command injection vulnerabilities. The vulnerability candidate analysis record for the suspected vulnerability pattern is used to describe the specific pattern information of the suspected vulnerability pattern. Determine the line number of each suspected vulnerability pattern and the first vulnerability type identifier, wherein the first vulnerability type identifier is any one of SQL injection vulnerability, cross-site scripting vulnerability, and command injection vulnerability; Calculate the matching degree of each suspected vulnerability pattern, and use the product of the matching degree of each suspected vulnerability pattern and the confidence degree of the known vulnerability corresponding to each suspected vulnerability pattern as the confidence score of each suspected vulnerability pattern. The matching degree is the ratio of the number of keywords in the code segment that belong to the suspected vulnerability pattern to the total number of keywords in the code segment. The first vulnerability candidate record corresponding to the code segment is generated using the line number of each suspected vulnerability pattern, the first vulnerability type identifier of each suspected vulnerability pattern, the confidence score of each suspected vulnerability pattern, and the vulnerability candidate analysis record of each suspected vulnerability pattern.
6. The system according to claim 4, characterized in that, The second LLM agent is used to input the plurality of code fragments into the second LLM agent. For each code fragment, the second LLM agent identifies external input points in the code fragment as taint source variables to obtain multiple taint source variables. For each pollution source variable, the assignment, transmission and operation process of the pollution source variable is semantically analyzed to construct a data flow diagram corresponding to the pollution source variable. The data flow diagram is used to represent multiple data propagation paths of data starting from the pollution source variable in the code logic. By performing semantic analysis on each of the pollution source variables, multiple data flow diagrams are obtained, and multiple data propagation paths are extracted from the multiple data flow diagrams; For each data propagation path, it is detected whether the data propagation path has undergone security processing operations to obtain the status assessment result of the data propagation path. The security processing operations include encoding operations, verification operations, and filtering operations. Multiple sensitive convergence points were identified in the code snippet, and the sensitive convergence points are any one of database query functions, command execution functions, and file operation functions; Extract the target data propagation path from the multiple data propagation paths whose state assessment result is that it has not undergone effective security processing and whose path endpoint belongs to the multiple sensitive convergence points, to obtain at least one target data propagation path, and generate a vulnerability candidate analysis record for each target data propagation path. The vulnerability candidate analysis record of the target data propagation path is used to describe the complete contamination path of the target data propagation path. For each target data propagation path, the line number of the target data propagation path is determined according to the path start and path end, and a second vulnerability type identifier of the target data propagation path is determined. The second vulnerability type identifier is any one of SQL injection vulnerability, command injection vulnerability, and file operation vulnerability. The clarity and unambiguity of the target data propagation path in the code logic are quantitatively evaluated to obtain the propagation path determinism evaluation value of the target data propagation path. The product of the propagation path determinism evaluation value of the target data propagation path and the operation type confidence score corresponding to the target data propagation path is used as the confidence score of the target data propagation path. The operation type confidence score is determined according to the operation type corresponding to the path endpoint of the target data propagation path. The second vulnerability candidate record corresponding to the code fragment is generated by using the line number of the code for each target data propagation path, the second vulnerability type identifier of each target data propagation path, the confidence score of each target data propagation path, and the vulnerability candidate analysis record of each target data propagation path.
7. The system according to claim 4, characterized in that, The third LLM agent is used to input the multiple code fragments into the third LLM agent. For each code fragment, the conditional branch structure in the code fragment is parsed to construct a conditional branch graph, which contains multiple possible execution paths. At least one possible execution path containing sensitive operations is identified from the plurality of possible execution paths. Among the at least one possible execution path containing sensitive operations, a possible execution path with privilege bypass risk is identified as a privilege bypass path, thus obtaining at least one privilege bypass path. A vulnerability candidate analysis record is generated for each privilege bypass path. The sensitive operations include modifying the database, accessing files, and executing privileged commands. The privilege bypass risk includes bypassing through abnormal branches, bypassing through default branches, and bypassing through condition combinations. The vulnerability candidate analysis record of the privilege bypass path is used to describe the specific logical defects of the privilege bypass path. The ratio of the number of paths of the at least one permission bypass path to the number of paths of the plurality of possible execution paths is used as the first ratio; For each of the privilege bypass paths, the line number and third vulnerability type identifier of the privilege bypass path are determined, and the product of the confidence score of the sensitive operation type corresponding to the privilege bypass path and the first proportion is used as the confidence score of the privilege bypass path. The confidence score of the sensitive operation type corresponding to the privilege bypass path is determined according to the sensitive operation type corresponding to the privilege bypass path. Perform combined condition coverage testing on the conditional branch graph to generate multiple combinations of input boundary values and obtain multiple test outliers. Use the multiple test outliers and the multiple combinations of input boundary values as a test case set. The multiple test outliers include null, empty value, extreme value, and illegal type value. The code snippet is tested using the test case set to obtain a test result set. Based on the test result set, test cases belonging to unhandled boundary cases are extracted from the multiple test cases included in the test case set as unhandled boundary defects, resulting in multiple unhandled boundary defects. A vulnerability candidate analysis record is generated for each unhandled boundary defect. The vulnerability candidate analysis record of the unhandled boundary defect is used to describe the specific logical defect of the unhandled boundary defect. The ratio of the total number of unprocessed boundary defects to the total number of test cases is used as the second ratio; For each unprocessed boundary defect, the line number and fourth vulnerability type identifier of the unprocessed boundary defect are determined, and the product of the confidence score of the sensitive operation type corresponding to the unprocessed boundary defect and the second ratio is used as the confidence score of the unprocessed boundary defect. The confidence score of the sensitive operation type corresponding to the unprocessed boundary defect is determined according to the sensitive operation type corresponding to the unprocessed boundary defect. The third vulnerability candidate record corresponding to the code fragment is generated using the line number of the code for each privilege bypass path, the third vulnerability type identifier of each privilege bypass path, the confidence score of each privilege bypass path, the vulnerability candidate analysis record of each privilege bypass path, the line number of the code for each unprocessed boundary defect, the fourth vulnerability type identifier of each unprocessed boundary defect, the confidence score of each unprocessed boundary defect, and the vulnerability candidate analysis record of each unprocessed boundary defect.
8. The system according to claim 4, characterized in that, The fourth LLM agent is used to input the plurality of code fragments into the fourth LLM agent, and for each code fragment, to identify multiple calls to the third-party library application programming interface in the code fragment; For each of the aforementioned calls, the call is matched against a risky application programming interface knowledge base containing multiple risky application programming interfaces that are obsolete, insecure, or have common vulnerabilities. If a match is successful, the call is treated as a risky call. The line number of the risky call, the name of the risky application programming interface, and the fifth vulnerability type identifier are determined. A vulnerability candidate analysis record is generated for each risky call, and the confidence score of the deprecated insecure API corresponding to the risky call is obtained as the confidence score value of the risky call. The vulnerability candidate analysis record of the risky call is used to describe the specific risky API information of the risky call. The fourth vulnerability candidate record corresponding to the code fragment is generated using the line number of each risky call, the name of the risky application programming interface of each risky call, the fifth vulnerability type identifier of each risky call, the confidence score of each risky call, and the vulnerability candidate analysis record of each risky call.
9. The system according to claim 1, characterized in that, The collaborative regularization module is used to perform result clustering on the target vulnerability candidate set, and add vulnerability candidates that point to the same code location and have similar vulnerability types to the same cluster group to obtain multiple vulnerability candidate groups. The vulnerability candidate is any one of the following: suspected vulnerability pattern, target data propagation path, privilege bypass path, unprocessed boundary defect, and risky call. For each vulnerability candidate group, the vulnerability candidate analysis record of each vulnerability candidate in the vulnerability candidate group is extracted. According to the type of the vulnerability candidate group, the vulnerability candidate analysis records of the multiple vulnerability candidates are subjected to consistency verification and / or contradiction ruling. The confidence scores of the multiple vulnerability candidates are adjusted based on the results of the consistency verification and / or contradiction ruling. For each vulnerability candidate in the vulnerability candidate group, the current confidence score of the vulnerability candidate is obtained. Based on the agent-vulnerability type weight mapping table, the agent identifier and vulnerability type identifier of the vulnerability candidate are used to determine the agent-vulnerability type weight corresponding to the vulnerability candidate. The product of the current confidence score of the vulnerability candidate and the agent-vulnerability type weight corresponding to the vulnerability candidate is used as the weighted evidence score of the vulnerability candidate. The total weighted evidence score of the vulnerability candidate group is calculated using the weighted evidence scores of multiple vulnerability candidates. Specifically, if the vulnerability candidate groups belong to the same code location and have similar vulnerability types, the sum of the weighted evidence scores of the multiple vulnerability candidates is used as the total weighted evidence score of the vulnerability candidate group. If the vulnerability candidate groups belong to the same code location but have different vulnerability types, or if the vulnerability candidate groups belong to the same code location and have similar vulnerability types but conflicting vulnerability candidate analysis records, the LLM agent with the highest weight value is determined among the multiple LLM agents according to the agent-vulnerability type weight mapping table. The sum of the weighted evidence scores of at least one vulnerability candidate belonging to the LLM agent is used as the total weighted evidence score of the vulnerability candidate group. By performing consistency verification and / or conflict resolution on each of the vulnerability candidate groups, the total weighted evidence score of each vulnerability candidate group and the weighted evidence score of multiple vulnerability candidates in each vulnerability candidate group are obtained. The structured weighted evidence package is generated by combining the multiple code snippets, the target vulnerability candidate set, the weighted evidence scores of multiple vulnerability candidates in each vulnerability candidate group, and the total weighted evidence score of each vulnerability candidate group. The structured weighted evidence package is input into the arbitrator LLM agent for adjudication, generating structured adjudication information. The structured adjudication information includes the final judgment result, comprehensive confidence level, and adjudication reason for each vulnerability candidate. The final judgment result is any one of the following: confirmed as a vulnerability, marked as a suspected vulnerability, or determined as a false alarm.
10. The system according to claim 1, characterized in that, The collaborative regularization module is used to detect the type of each vulnerability candidate group. If the vulnerability candidate groups belong to the same code location and have similar vulnerability types, then the vulnerability candidate analysis record of each vulnerability candidate in the vulnerability candidate group is extracted, the vulnerability candidate analysis records of the multiple vulnerability candidates are logically verified, and the confidence scores of the multiple vulnerability candidates are adjusted based on the logical verification processing rules. If the vulnerability candidate groups belong to the same code location but have different vulnerability types, then the vulnerability candidate analysis record of each vulnerability candidate in the vulnerability candidate group is extracted, the vulnerability candidate analysis records of the multiple vulnerability candidates are logically coexisted, and the confidence scores of the multiple vulnerability candidates are adjusted based on the logical coexistence verification processing rules. If the vulnerability candidate groups belong to the same code location and have similar vulnerability types but conflicting vulnerability candidate analysis records, then the vulnerability candidate analysis record of each vulnerability candidate in the vulnerability candidate group is extracted, the existence of logical vulnerabilities is verified on the vulnerability candidate analysis records of the multiple vulnerability candidates, and the confidence scores of the multiple vulnerability candidates are adjusted based on the logical vulnerability existence verification processing rules.