Target code file vulnerability mining system and method and readable storage medium
The intelligent agent group and large language model system enhances PHP code file vulnerability detection by combining static and dynamic analysis, addressing high false positives and low coverage issues in existing methods, thereby improving detection accuracy.
Patent Information
- Application Number
- CN202510365874.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-26
- Publication Date
- 2025-07-15
AI Technical Summary
In the existing PHP code file vulnerability mining scheme, the static analysis scheme has a high false alarm rate and a low code coverage rate of the dynamic analysis scheme, resulting in insufficient accuracy of vulnerability mining.
Using a combination of intelligent group and large language model, we conduct static and dynamic analysis through project management, security experts, developers and tester agents, generate and combine vulnerability mining strategies to improve the accuracy of vulnerability mining.
By combining static and dynamic analysis, we can understand the code operation in a deeper way, improve the accuracy and coverage of vulnerability mining, and effectively improve the accuracy of vulnerability mining.
Smart Images

Figure CN120316772A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer network security technology, and particularly to a vulnerability mining system, method, and readable storage medium for target code files. Background Art
[0002] Currently, vulnerability mining solutions for code files such as PHP can be roughly divided into static analysis solutions and dynamic analysis solutions. Among them, the static analysis solution is based on the static source code analysis tool RIPS, which uses predefined rule patterns and matching methods to detect common vulnerabilities existing in the target code file. However, this static method cannot accurately reflect the real situation of the code file and has a high false positive rate. The dynamic analysis solution is based on the open-source dynamic scanning tool OWASP ZAP. Although it conducts vulnerability mining by observing the running behavior of the code program, this method has a low code coverage rate.
[0003] Therefore, how to effectively improve the accuracy of vulnerability mining for target code files has become a technical problem that needs to be solved urgently by those skilled in the art. Summary of the Invention
[0004] Based on the above problems, in order to improve the accuracy of vulnerability mining for target code files, the embodiments of this application provide a vulnerability mining system, method, and readable storage medium for target code files.
[0005] The embodiments of this application disclose the following technical solutions:
[0006] In a first aspect, the embodiments of this application provide a vulnerability mining system for target code files, including: an agent group and a large language model; the execution logic of each agent in the agent group is determined based on task keywords and the large language model; the agent group includes: a project management agent, a security expert agent, a tester agent, and a developer agent;
[0007] The project management agent is used to deploy an audit task according to the target code file attribute information, so as to send the audit task to the security expert agent;
[0008] The security expert agent is used to determine the vulnerability analysis logic for the target code file according to the audit task, and generate a vulnerability risk analysis result based on the code content of the target code file and the vulnerability analysis logic; the vulnerability risk analysis result includes: the risk function of the target code file;
[0009] The developer agent is used to perform call chain tracing configuration on the risk function through the abstract syntax tree of the target code file and a hash function, and generate first call chain information of the risk function; the first call chain information is used to represent the static call chain of the risk function.
[0010] The tester agent is used to perform dynamic running tests on the target code file according to the first call chain information, obtain second call chain information, and send the second call chain information to the security expert agent and the developer agent to obtain a first vulnerability mining strategy and a second vulnerability mining strategy; the second call chain information is used to represent a call chain whose similarity to the first call chain is greater than a preset first threshold; the first vulnerability mining strategy performs vulnerability mining based on vulnerability characteristics and comes from the security expert agent; the second vulnerability mining strategy performs vulnerability mining based on code logic and comes from the developer agent.
[0011] The tester agent is also used to perform vulnerability mining on the target code file based on the first vulnerability mining strategy or the second vulnerability mining strategy to obtain a vulnerability mining result.
[0012] In a possible implementation manner, the developer agent is specifically used for:
[0013] Perform call chain analysis according to the abstract syntax tree and the vulnerability risk analysis result to determine the static call chain of the risk function.
[0014] Perform code instrumentation on the static call chain of the risk function based on the hash function to obtain the first call chain information.
[0015] In a possible implementation manner, the first call chain information includes: a risk function hash index and entry information of the target code file; the tester agent includes: a dynamic testing unit; the dynamic testing unit is used for:
[0016] Generate test data according to the risk function hash index and the entry information of the target code file.
[0017] Perform a running test on the target code file based on the test data to obtain third call chain information; the third call chain information is used to represent the call chain information and variable values of the risk function during runtime.
[0018] Calculate the link similarity of each call chain between the third call chain information and the first call chain information based on the LCS algorithm, and determine the call chain information with a similarity greater than the preset first threshold in the third call chain information as the fourth call chain information.
[0019] Perform mutation processing on the third call chain information, calculate the link similarity of each call chain between the mutated call chain information and the first call chain information through the LCS algorithm, and determine the call chain information with a similarity greater than the preset first threshold as the fifth call chain information;
[0020] Merge the fifth call chain information with the fourth call chain information to obtain the second call chain information.
[0021] In a possible implementation manner, the dynamic testing unit is further configured to:
[0022] If the sum of the number of call chains of the fifth call chain information and the fourth call chain information is less than a preset second threshold, send a call chain exploration strategy request to the security expert agent and the developer agent;
[0023] Receive the first call chain exploration strategy and the second call chain exploration strategy in response to the call chain exploration strategy request; the first path exploration strategy performs path exploration based on the parameters of the vulnerability characteristics and comes from the security expert agent; the second call chain exploration strategy performs path exploration based on the URL parameters of the code structure and comes from the developer agent;
[0024] Use the first call chain exploration strategy or the second call chain exploration strategy to perform call chain exploration to obtain the second call chain information; the execution priority of the first call chain exploration strategy is higher than that of the second call chain exploration strategy.
[0025] In a possible implementation manner, the security expert agent includes: a first vulnerability mining unit; the developer agent includes: a second vulnerability mining unit;
[0026] The first vulnerability mining unit is specifically configured to:
[0027] Receive the second call chain information sent by the tester agent;
[0028] Generate a vulnerability mining strategy based on vulnerability characteristics and a corresponding first verification parameter according to the second call chain information and the vulnerability risk analysis result;
[0029] Determine the vulnerability mining strategy based on vulnerability characteristics and the first verification parameter as the first vulnerability mining strategy;
[0030] The second vulnerability mining unit is specifically configured to:
[0031] Receive the second call chain information sent by the tester agent;
[0032] Generate a vulnerability mining strategy based on code logic and corresponding second verification parameters according to the second call chain information and the abstract syntax tree;
[0033] Determine the vulnerability mining strategy based on vulnerability characteristics and the second verification parameters as the second vulnerability mining strategy.
[0034] In a possible implementation, the tester agent includes: a third vulnerability mining unit; the third vulnerability mining unit is specifically used for:
[0035] Perform vulnerability mining using the first vulnerability mining strategy to obtain a first mining result;
[0036] Verify the first mining result for vulnerabilities through the first verification parameter to obtain a first verification result;
[0037] When the first verification result fails, perform vulnerability mining using the second vulnerability mining strategy to obtain a second mining result, and verify the second mining result for vulnerabilities through the second verification parameter to obtain a second verification result;
[0038] If the second verification result fails, send a mining strategy reconstruction request to the security expert agent.
[0039] In a possible implementation, the vulnerability mining result includes: a vulnerability risk level; the security expert agent includes: a repair strategy generation unit; the developer agent includes: a vulnerability repair unit; the tester agent includes: a code verification unit;
[0040] The repair strategy generation unit is used for:
[0041] Generate a first vulnerability repair strategy according to the vulnerability mining result and send the first vulnerability repair strategy to the vulnerability repair unit;
[0042] The vulnerability repair unit is used for:
[0043] Generate a second vulnerability repair strategy according to the vulnerability mining result and determine the strategy reference weights of the first vulnerability repair strategy and the second vulnerability repair strategy based on the vulnerability risk level;
[0044] Repair the vulnerabilities of the target code file according to the first vulnerability repair strategy, the second vulnerability repair strategy and the strategy reference weights to obtain a repaired target code file and send it to the code verification unit;
[0045] The code verification unit is used for:
[0046] Perform vulnerability verification on the repaired target code file to determine whether the vulnerabilities in the repaired target code file have been completely repaired.
[0047] In a possible implementation, the intelligent group further includes: an operation and maintenance intelligent agent; the operation and maintenance intelligent agent is specifically used for:
[0048] Receive an environment deployment task from the project management intelligent agent; the environment deployment task includes multiple environment deployment steps;
[0049] Generate operation prompt words for each of the environment deployment steps based on the large language model;
[0050] Predict the expected execution results of each of the environment deployment steps according to the operation prompt words of each of the environment deployment steps;
[0051] When the expected execution result does not meet the preset conditions, iteratively optimize the environment deployment steps whose expected execution results do not meet the preset conditions;
[0052] If the number of optimization times of the iterative optimization reaches a preset third threshold, send an environment deployment assistance request to the developer intelligent agent.
[0053] In a possible implementation, the system further includes: a message middleware; the message middleware is used to implement asynchronous communication between the intelligent agents in the intelligent agent group; the message middleware includes: a message standardization unit, a priority processing unit, and a message tracking unit;
[0054] The message standardization unit is used to receive the communication messages between the intelligent agents and perform formatting processing on the communication messages to obtain multiple formatted communication messages; the formatted communication messages include: message type, processing priority, sender information, recipient information, message content, and tracking identifier;
[0055] The priority processing unit is used to implement asynchronous communication between the intelligent agents according to the processing priorities of the formatted communication messages.
[0056] The message tracking unit is used to perform communication link tracking on the formatted communication messages according to the tracking identifiers of the formatted communication messages.
[0057] Second aspect, an embodiment of the present application provides a method for vulnerability mining of target code files, which is applied to a target code file vulnerability mining system including an intelligent group and a large language model; the execution logic of each intelligent agent in the intelligent agent group is determined based on task keywords and the large language model; the intelligent agent group includes: a project management intelligent agent, a security expert intelligent agent, a tester intelligent agent, and a developer intelligent agent; the method includes:
[0058] Control the project management intelligent agent to deploy an audit task according to the target code file attribute information, so as to send the audit task to the security expert intelligent agent;
[0059] Control the security patent intelligent agent to determine the vulnerability analysis logic for the target code file according to the audit task, and generate a vulnerability risk analysis result based on the code content of the target code file and the vulnerability analysis logic; the vulnerability risk analysis result includes: the risk function of the target code file;
[0060] Control the developer intelligent agent to configure call chain tracing for the risk function through the abstract syntax tree and hash function of the target code file, and generate the first call chain information of the risk function; the first call chain information is used to represent the static call chain of the risk function;
[0061] Control the tester intelligent agent to perform dynamic running tests on the target code file according to the first call chain information, obtain the second call chain information, and send the second call chain information to the security expert intelligent agent and the developer intelligent agent to obtain the first vulnerability mining strategy and the second vulnerability mining strategy; the second call chain information is used to represent a call chain whose similarity to the first call chain is greater than a preset first threshold; the first vulnerability mining strategy performs vulnerability mining based on vulnerability characteristics and comes from the security expert intelligent agent; the second vulnerability mining strategy performs vulnerability mining based on code logic and comes from the developer intelligent agent;
[0062] Control the tester intelligent agent to perform vulnerability mining on the target code file based on the first vulnerability mining strategy or the second vulnerability mining strategy, and obtain a vulnerability mining result.
[0063] Third aspect, an embodiment of the present application provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements any possible method for vulnerability mining of target code files in the first aspect.
[0064] Compared with the prior art, the present application has the following beneficial effects: The embodiments of the present application provide a target code file vulnerability mining system, method and readable storage medium. The system includes an agent group and a large language model. The execution logic of each agent in the agent group is determined based on specific task keywords and the large language model, so as to improve the analysis ability of the dynamic characteristics of the target code file through the semantic understanding ability of the large language model. Among them, the project management agent is used to send an audit task to the security expert agent according to the target code file attribute information. The security expert agent determines the vulnerability analysis logic for the target code file according to the audit task and generates a vulnerability risk analysis result marked with a risk function. The developer agent configures the call chain tracking for the risk functions included in the vulnerability risk analysis result through the abstract syntax tree and hash function of the target code file, and generates the first call chain information for the risk functions, so as to perform a preliminary static analysis of the call chain of the risk functions through the abstract syntax tree. Subsequently, the tester agent performs a dynamic running test on the target code file according to the first call chain information to obtain the actual call chain involved during the running of the risk functions, and obtains the second call chain information. In this way, through a combination of static analysis and dynamic analysis, code coverage is realized at a deeper level on the basis of understanding the actual running situation of the code, thereby effectively improving the accuracy of vulnerability mining for the target code file. Finally, the second call chain information is fed back to the security expert agent and the developer agent to obtain the first vulnerability mining strategy and the second vulnerability mining strategy feedback by both, and the target code file is mined for vulnerabilities based on these, so as to carry out vulnerability mining for the target code file from the dual aspects of vulnerability characteristics and code logic, effectively improving the accuracy of vulnerability mining for the target code file. BRIEF DESCRIPTION OF THE DRAWINGS
[0065] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0066] Figure 1 It is a schematic structural diagram of a target code file vulnerability mining system provided by an embodiment of the present application;
[0067] Figure 2 It is a schematic flow diagram of task keyword generation and optimization provided by an embodiment of the present application;
[0068] Figure 3 It is a schematic flow diagram of a vulnerability mining unit executing a mining strategy provided by an embodiment of the present application;
[0069] Figure 4 It is a schematic flow chart of a dynamic operation test method provided by an embodiment of the present application;
[0070] Figure 5 It is a schematic flow chart of an operation and maintenance intelligent agent deployment environment provided by an embodiment of the present application;
[0071] Figure 6 It is a schematic diagram of a multi-intelligent agent collaboration structure provided by an embodiment of the present application;
[0072] Figure 7 It is a schematic diagram of another target code file mining system provided by an embodiment of the present application;
[0073] Figure 8 It is a schematic diagram of a message middleware provided by an embodiment of the present application;
[0074] Figure 9 It is a schematic flow chart of a target code file vulnerability mining method provided by an embodiment of the present application. Detailed implementation manners
[0075] To make the objectives, technical solutions and advantages of the present application clearer and more understandable, the following further describes the present application in detail with reference to specific embodiments and the accompanying drawings. It should be particularly noted that the embodiments described in the embodiments of the present application are only a part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present application without creative efforts shall fall within the protection scope of the present application.
[0076] It should be noted that unless otherwise defined, the technical terms or scientific terms used in the embodiments of the present application should be of the ordinary meaning understood by those of ordinary skill in the art to which the present application belongs. The "first", "second" and similar terms used in the embodiments of the present application do not indicate any order, quantity or importance, but are only used to distinguish different components. The terms such as "including" or "comprising" mean that the elements or objects appearing before this term cover the elements or objects listed after this term and their equivalents, without excluding other elements or objects. The terms such as "connected" or "coupled" are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. The terms such as "upper", "lower", "left" and "right" are only used to represent relative positional relationships, and when the absolute position of the object being described changes, the relative positional relationship may also change accordingly.
[0077] As described above, the current vulnerability mining solutions for code files such as PHP can be roughly divided into static analysis solutions and dynamic analysis solutions. Among them, the static analysis solution is based on the static source code analysis tool RIPS, which uses predefined rule patterns and matching methods to detect common vulnerabilities existing in the target code file. However, this static method cannot accurately reflect the real situation of the code file and has the problem of high false positive rate. The dynamic analysis solution is based on the open-source dynamic scanning tool OWASP ZAP. Although it conducts vulnerability mining by observing the running behavior of the code program, this method has the problem of low code coverage.
[0078] To solve the above problems, the embodiments of this application provide a vulnerability mining system, method and readable storage medium for target code files. The system includes an agent group and a large language model. The execution logics of the agents in the agent group are determined based on specific task keywords and the large language model, so as to improve the analysis ability for the dynamic characteristics of the target code file through the semantic understanding ability of the large language model. Among them, the project management agent is used to send an audit task to the security expert agent according to the attribute information of the target code file. The security expert agent determines the vulnerability analysis logic for the target code file according to the audit task and generates a vulnerability risk analysis result marked with risk functions. The developer agent configures the call chain tracing for the risk functions included in the vulnerability risk analysis result through the abstract syntax tree and hash function of the target code file, and generates the first call chain information for the risk functions, so as to conduct a preliminary static analysis of the call chain of the risk functions through the abstract syntax tree. Subsequently, the tester agent conducts a dynamic running test on the target code file according to the first call chain information to obtain the actual call chain involved during the running of the risk functions and get the second call chain information. In this way, through the combined processing of static analysis and dynamic analysis, on the basis of understanding the actual running situation of the code, deeper code coverage is realized, so as to effectively improve the accuracy of vulnerability mining for the target code file. Finally, the second call chain information is fed back to the security expert agent and the developer agent to obtain the first vulnerability mining strategy and the second vulnerability mining strategy fed back by both of them, and the vulnerability mining of the target code file is carried out based on this, so as to carry out the vulnerability mining for the target code file from the dual aspects of vulnerability characteristics and code logic, effectively improving the accuracy rate of vulnerability mining for the target code file.
[0079] To enable those skilled in the art to better understand the solution of this application, the following will clearly and completely describe the technical solutions in the embodiments of this application in conjunction with the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all the embodiments. Based on the embodiments in this application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of this application.
[0080] See Figure 1 This figure is a schematic structural diagram of a target code file vulnerability mining system provided by an embodiment of this application. As can be seen from the figure, in the target code file vulnerability mining system provided by the embodiment of this application, there are an intelligent agent group and a large language model. Among them, the intelligent agent group is connected to the large language model, and they together constitute an analysis server in the vulnerability mining system for performing vulnerability analysis. Correspondingly, connected to the analysis server is a test mining server, which is used to run the operating environment for the target code file, so as to facilitate subsequent vulnerability mining by running the target code file.
[0081] One end of the intelligent agent group is connected to the large language model, and the specific execution logic inside the intelligent agent is determined by specific task keywords and the corresponding large language model. For example, taking the target code file as a php code file, the task keyword can be set as "vulnerability risk analysis", and this task keyword is sent to the security expert intelligent agent. The security expert intelligent agent generates specific vulnerability risk analysis methods through the large language model for the keyword "vulnerability risk analysis", and then determines its own execution logic.
[0082] In the actual application scenario, the task keywords can be confirmed through repeated training and iterative optimization based on the large language model. For specific details, see Figure 2 the schematic flowchart of a process for generating and optimizing task keywords disclosed Figure 2 The prompt words in it represent the task keywords in this application. It can be seen from Figure 2 that the basic prompt word template (task keyword generation template) is determined by the role definition of the intelligent agent, the professional field, the permission scope of the intelligent agent, and the specific task description. First, the first round of prompt words (the first round of task keywords) can be generated based on these basic templates, and it is verified that the first round of prompt words enable the intelligent agent to accurately complete the expected work, so as to perform repeated iterative optimization, improve the generation accuracy of the task keywords, and thereby indirectly achieve precise control of the intelligent agent.
[0083] In addition, the task keywords can also be determined in real time by the large language model. Taking the security expert intelligent agent as an example again, the security expert intelligent agent can generate corresponding task keywords according to the code content in the target code file and the large language model, and further confirm the specific vulnerability analysis logic based on this.
[0084] In the embodiments of the present application, the intelligent agent group includes: a project management intelligent agent 100, a security expert intelligent agent 200, a developer intelligent agent 300, and a tester intelligent agent 400. Next, the execution content of each type of intelligent agent will be introduced in sequence.
[0085] The project management intelligent agent 100 is used to deploy an audit task according to the target code file attribute information, so as to send the audit task to the security expert intelligent agent.
[0086] As the task deployment center, based on its application of the large language model, the project management intelligent agent needs to generate an audit task according to a series of basic information in the target code file (such as directory structure, environmental requirements, readme file, etc.), and send the audit task to the security expert intelligent agent, so as to realize the task deployment of the security expert intelligent agent. Similarly, the project manager can also perform task deployments such as operating environment and vulnerability mining method development according to the overall project information (such as a project for vulnerability mining of the target code file) and the basic information of the target code file. The role of the project management intelligent agent is to deploy tasks to other intelligent agents in the intelligent agent group according to the project information, and then initially plan the execution logic of each intelligent agent.
[0087] The security expert intelligent agent 200 is used to determine the vulnerability analysis logic for the target code file according to the audit task, and generate a vulnerability risk analysis result based on the code content of the target code file and the vulnerability analysis logic; the vulnerability risk analysis result includes: the risk function of the target code file.
[0088] The security expert intelligent agent can conduct vulnerability analysis on the vulnerability characteristics themselves. After receiving the audit task, it performs text analysis on the audit task through the large language model, and then determines the vulnerability analysis logic for the target code file from the audit task. On this basis, the potential vulnerability types, risk levels, and existing risk functions in the target code file are evaluated through the specific code content and vulnerability analysis logic in the target code file, so as to generate a vulnerability risk analysis result, and send the vulnerability risk analysis result to the developer intelligent agent, so that the developer intelligent agent can plan further vulnerability mining strategies in combination with the vulnerability risk analysis result.
[0089] In a possible implementation manner, the security expert intelligent agent can use static analysis technology to scan the target code file for code analysis of the target code file, and when a risk function is detected, determine the risk score of the risk function in combination with the danger level of the function itself and the context risk coefficient, so as to construct a vulnerability risk analysis result.
[0090] The developer agent 300 is used to perform call chain tracing configuration on the risk function through the abstract syntax tree and hash function of the target code file, and generate first call chain information of the risk function; the first call chain information is used to represent the static call chain of the risk function.
[0091] After the developer agent receives the vulnerability risk analysis result from the security expert agent, the developer agent constructs an abstract syntax tree (AST) for the target code file. The abstract syntax tree is a structured representation of the source code, which converts the program code into a tree-like data structure to describe the syntax structure of the code without including specific syntax details. Therefore, the abstract syntax tree of the target code file can be regarded as a static structure representation of the target code file. By performing call chain analysis on the risk functions in the target code file through the abstract syntax tree of the target code file and the vulnerability risk analysis result, the static call chains of each risk function in the target code file can be determined.
[0092] Furthermore, by performing code instrumentation on the call chains of each risk function through a hash function, specific hash indexes for each risk function can be established in the risk function, thereby implementing a tracing mechanism for the risk function call chain and generating first call chain information for the risk function.
[0093] The tester agent 400 is used to perform dynamic running tests on the target code file according to the first call chain information, obtain second call chain information, and send the second call chain information to the security expert agent and the developer agent to obtain a first vulnerability mining strategy and a second vulnerability mining strategy; the second call chain information is used to represent a call chain whose similarity to the first call chain is greater than a preset first threshold; the first vulnerability mining strategy is based on vulnerability characteristics for vulnerability mining and comes from the security expert agent; the second vulnerability mining strategy is based on code logic for vulnerability mining and comes from the developer agent.
[0094] In the actual running scenario of the code, there is an essential gap between the static structure analysis of the code and the dynamic behavior during actual operation: static analysis is deduced based on the syntax structure of the code and visible symbol relationships, and cannot fully capture method calls dynamically bound during program operation, such as polymorphic implementation, reflection mechanism, class loading dependent on the environment, etc. Therefore, these execution mechanisms unique to runtime will reconstruct or extend the static call chain of the risk function, resulting in a structural deviation between the pre-statically deduced call chain and the actually executed call chain.
[0095] Therefore, to avoid such problems, the tester agent needs to perform dynamic running tests on the target code file based on the call chain information (the first call chain information) contained in the first call chain information regarding the static call chain, so as to obtain the second call chain information. Among them, the second call chain information is used to represent the call chain whose similarity to the first call chain information (static call chain) is greater than a preset first threshold. During code operation, the higher the similarity between the call chain of the risk function and its static call chain, the more it proves that the acquisition of the call chain of the risk function conforms to the expected static analysis result. In this way, the static analysis and dynamic analysis are deeply integrated to improve the accuracy of obtaining the call path of the risk function.
[0096] After obtaining the second call chain information, the tester agent re-feeds the second call chain information to the security expert agent and the developer agent to obtain the first vulnerability mining strategy and the second vulnerability mining strategy. Among them, the first vulnerability mining strategy is fed back by the security expert agent. As can be seen from the previous text, the security expert agent can perform vulnerability analysis through vulnerability features. Correspondingly, it can also provide the corresponding vulnerability mining strategy (the first vulnerability mining strategy) at the level of vulnerability features. Similarly, the second vulnerability mining strategy is fed back by the developer agent. The developer agent mainly performs vulnerability analysis by constructing the abstract syntax tree of the target code file and combining the actual code structure and code logic of the target code file. Therefore, the developer agent provides the vulnerability mining strategy based on code logic (the second vulnerability mining strategy).
[0097] In the security expert agent and the developer agent, a first vulnerability mining unit and a second vulnerability mining unit are respectively set up. The vulnerability mining units of both are used to receive the second call chain information fed back by the tester agent, and respectively generate corresponding vulnerability mining strategies based on the analysis results of the target code file, so as to provide vulnerability mining solutions at different levels and angles.
[0098] Among them, the first vulnerability mining unit (security expert agent) is used to generate a vulnerability mining strategy based on vulnerability features and the first verification parameter corresponding to this strategy according to the second call chain information and the previously generated vulnerability risk analysis result, and then determine the vulnerability mining strategy based on vulnerability features and the first verification parameter corresponding to this strategy as the first vulnerability mining strategy. Among them, the first verification parameter can be used to verify whether the execution result of the strategy meets the expectation during the subsequent execution of the vulnerability mining strategy.
[0099] Similarly, the second vulnerability mining unit (developer agent) is used to generate a vulnerability mining strategy based on code logic and a second verification parameter according to the second call chain information and the previously constructed abstract syntax tree, and then generate the second vulnerability mining strategy and feed it back to the tester agent.
[0100] The tester agent 400 is further configured to perform vulnerability mining on the target code file based on the first vulnerability mining strategy or the second vulnerability mining strategy to obtain a vulnerability mining result.
[0101] Finally, the tester agent performs vulnerability mining on the target code file based on the first vulnerability mining strategy and the second vulnerability mining strategy fed back by the security expert agent.
[0102] Specifically, the process of the tester agent performing vulnerability mining based on the first vulnerability mining strategy or the second vulnerability mining strategy is carried out by a third vulnerability mining unit in the tester agent. Next, the process of the third vulnerability mining unit executing the vulnerability mining strategy will be introduced in conjunction with the specific flowchart of the embodiment.
[0103] See Figure 3 , which is a flowchart showing the process of a vulnerability mining unit executing a mining strategy provided by an embodiment of the present application, specifically including the following steps:
[0104] S101: Perform vulnerability mining using the first vulnerability mining strategy to obtain a first mining result;
[0105] S102: Verify the vulnerabilities of the first mining result using the first verification parameter to obtain a first verification result;
[0106] S103: When the first verification result is not passed, perform vulnerability mining using the second vulnerability mining strategy to obtain a second mining result, and verify the vulnerabilities of the second mining result using the second verification parameter to obtain a second verification result;
[0107] S104: If the second verification result is not passed, send a mining strategy reconstruction request to the security expert agent.
[0108] In the embodiment of the present application, the adoption priority of the security expert agent is higher than that of the developer agent. Therefore, it is necessary to first perform vulnerability mining based on the first vulnerability mining strategy from the security expert agent to obtain a first mining result. Subsequently, the first verification parameter included in the first vulnerability mining strategy is used to verify the vulnerabilities of the first mining result to determine whether the first mining result obtained by executing the first vulnerability mining strategy meets the expectations. When the first verification result is not passed, the second vulnerability mining strategy provided by the developer agent is used to perform vulnerability mining, and the same verification steps are executed.
[0109] If the vulnerability mining strategies provided by both the security expert agent and the developer agent do not meet the expectations, a mining strategy reconstruction request is sent to the security expert agent to re-perform vulnerability mining, so as to ensure the accuracy of vulnerability mining. In particular, in a possible implementation, when sending a mining strategy reconstruction request to the security expert agent, the tester agent can also provide a vulnerability mining strategy based on historical vulnerability mining experience data. If this strategy also does not meet the expectations, a mining strategy reconstruction request is then sent to the security expert agent.
[0110] The above is the introduction for Figure 3 Next, the process of dynamically running and testing the tester agent will be introduced in combination with the specific embodiment drawings.
[0111] See Figure 4 , which is a schematic flowchart of a dynamic running and testing method provided by an embodiment of the present application, specifically including the following steps:
[0112] S201: Generate test data according to the risk function hash index and the entry information of the target code file.
[0113] The process of dynamically running and testing the tester agent is completed by the dynamic testing unit in the tester agent. Among the first call information previously determined by the tester agent through the abstract syntax tree, it includes the hash index corresponding to the risk function and the entry information of the target code file. The dynamic testing unit is used to generate test data according to the hash index of the risk function and the entry information of the target code file.
[0114] S202: Perform a running test on the target code file based on the test data to obtain third call chain information; the third call chain information is used to represent the call chain information and variable values of the risk function during runtime.
[0115] Subsequently, perform a running test on the target code file based on the generated test data. During the running of the target code file, monitor the call chain information of the dynamic running of the risk function through the hash index of the risk function, analyze the code stubs set in the risk function, and then obtain the call chain path and variable values of the risk function during runtime, and use them as the third call chain information.
[0116] Among them, the variable value refers to the specific data of the key variables captured by the code stub technology during program runtime, which is used to analyze the formation path and exploitability of vulnerabilities. By analyzing its variable values, the propagation process of the input in the call chain can be traced, and the variable usage scenarios correctly filtered can be identified.
[0117] S203: Calculate the link similarity of each call chain between the third call chain information and the first call chain information based on the LCS algorithm, and determine the call chain information with a similarity greater than the preset first threshold in the third call chain information as the fourth call chain information;
[0118] Further, the third call chain information includes multiple call chains in the running state of the risk function. To screen out the call chains that have a certain similarity with the static call chain (the first call chain information), it is necessary to calculate the link similarity between each call chain in the third call chain information and each call chain in the first call chain information based on the LCS algorithm, and then determine the call chain information with a similarity greater than the preset first threshold in the third call chain information as the fourth call chain information.
[0119] Among them, the LCS algorithm is a classic dynamic programming algorithm used to find the longest common subsequence existing in two sequences. The elements in the subsequence do not need to be continuous, but need to maintain the original order. In the embodiment of the present application, the LCS algorithm is used to calculate the similarity between call chain paths (the third call chain information and the first call chain information). After obtaining the third call chain information containing multiple call chain paths, the LCS algorithm is used to convert different call chains into function hash sequences, and calculate the length of the longest common subsequence between the paths, so as to screen and determine the fourth call chain information.
[0120] S204: Mutate the third call chain information, calculate the link similarity of each call chain between the mutated call chain information and the first call chain information through the LCS algorithm, and determine the call chain information with a similarity greater than the preset first threshold as the fifth call chain information.
[0121] In order to improve the coverage rate of code mining as much as possible, on the basis of obtaining the fourth call chain information, it is also necessary to further perform mutation processing in combination with the dynamic call chain information (the third call chain information) of the risk function, so as to derive more call chains where the risk function may exist. Specifically, the way to mutate the dynamic call chain information (the third call chain information) can be to modify the value structure of the HTTP request parameters, such as adding special characters, constructing a memory payload, etc., or through parameter combination, such as parameter order adjustment, parameter redundancy, etc. This embodiment does not limit the way of mutation processing.
[0122] After mutating the third call chain information, the LCS algorithm is also used to calculate the link similarity between each mutated call chain information and the first call chain information, so as to determine the call chain information with a similarity greater than the preset first threshold as the fifth call chain information.
[0123] S205: Merge the fifth call chain information with the fourth call chain information to obtain the second call chain information.
[0124] Finally, merge the fourth call chain information with a similarity greater than a preset first threshold in the third call chain information, and the fifth call chain information obtained after mutation processing and filtered based on similarity, so as to deeply complete the call chain analysis for the risk function, and then more comprehensively obtain the real call chain situation of the risk function, improving the coverage rate of vulnerability mining.
[0125] Correspondingly, if the sum of the number of call chains of the fifth call chain information and the fourth call chain information is less than a preset second threshold, it indicates that the acquisition of the call chain for the risk function does not meet the expectation. At this time, the test unit can send a call chain exploration strategy request to the security expert agent and the developer agent to request the security expert agent and the developer agent to feedback new call path exploration strategies.
[0126] Among them, the security expert agent feedbacks the first call chain exploration strategy, and the developer agent feedbacks the second call chain exploration strategy. The first call chain exploration strategy from the security expert is similar to the aforementioned first vulnerability mining strategy. The first call chain exploration strategy explores based on the parameters of the vulnerability characteristics, and its own execution priority is higher than the second call chain exploration strategy generated by the developer agent. Specifically, the first call chain exploration strategy can be generated based on a specific vulnerability characteristic database (such as CVE, OWASP Top 10) to construct test parameters with attack characteristics.
[0127] Correspondingly, the second call chain exploration strategy from the developer agent explores the path based on the URL parameters of the code structure, so as to cover more code branch paths.
[0128] The execution order of the first call chain exploration strategy and the second call chain exploration strategy is the same as the aforementioned vulnerability mining strategy, and the first call chain exploration strategy provided by the security expert agent is preferentially executed. When the exploration link found by the first call chain exploration strategy does not meet the expectation, the second call chain exploration strategy is executed.
[0129] The above is the introduction to the execution function of the dynamic test unit in the tester agent. Next, the subsequent vulnerability repair process will be introduced.
[0130] After obtaining the vulnerability mining results, it is necessary to repair the vulnerabilities according to the vulnerability risk level indicated in the vulnerability mining results and the actual type of the vulnerabilities. In the embodiments of the present application, the vulnerability repair strategy for the vulnerability mining results is provided by the repair strategy generation unit in the security expert agent and the vulnerability repair unit in the developer agent. The specific execution process of the funnel repair and the subsequent test verification link for whether the target code file has completed the vulnerability repair are completed by the vulnerability repair unit in the developer agent and the code verification unit in the tester agent respectively. Next, these units will be introduced in turn.
[0131] In the security expert agent, its repair strategy generation unit is specifically used for:
[0132] Generate a first vulnerability repair strategy according to the vulnerability mining results, and send the first vulnerability repair strategy to the vulnerability repair unit.
[0133] In the vulnerability repair stage, the security expert agent judges the key positions where the vulnerabilities exist and the causes of the vulnerabilities in the target code file according to the actually generated vulnerability mining results and in combination with the prior vulnerability analysis of the target code file. Furthermore, a first vulnerability repair strategy is generated from the level of vulnerability characteristics and sent to the vulnerability repair unit in the developer agent, so that the developer agent can perform specific vulnerability repair work in combination with the first vulnerability repair strategy provided by the security expert agent.
[0134] In the developer agent, its vulnerability repair unit is specifically used to perform the following two steps:
[0135] Step 1: Generate a second vulnerability repair strategy according to the vulnerability mining results, and determine the strategy reference weights of the first vulnerability repair strategy and the second vulnerability repair strategy based on the vulnerability risk level;
[0136] Step 2: Repair the vulnerabilities of the target code file according to the first vulnerability repair strategy, the second vulnerability repair strategy, and the strategy reference weights to obtain the repaired target code file, and send it to the code verification unit.
[0137] The developer agent itself can generate a second vulnerability repair strategy based on the code structure and code logic, and determine the reference weight for referring to the security expert agent's repair strategy according to the vulnerability risk level indicated in the vulnerability mining results, that is, determine the strategy reference weights of the first vulnerability repair strategy and the second vulnerability repair strategy. If the risk level of the vulnerability is relatively high, it indicates that the vulnerability existing in the target code file belongs to a critical security vulnerability, and thus the reference weight for the first vulnerability repair strategy (security expert agent) is relatively large. Correspondingly, if the risk level of the vulnerability is relatively low, it indicates that the vulnerability existing in the target code file belongs to a regular vulnerability, and thus the vulnerability repair strategies provided by both the security expert agent and the developer agent itself can be comprehensively referred to, so as to ensure that the vulnerability can be completely repaired.
[0138] In the tester agent, its code verification unit is specifically used for:
[0139] Performing vulnerability verification on the repaired target code file to determine whether the vulnerabilities in the repaired target code file are repaired.
[0140] After the developer agent notifies the tester agent that the vulnerability repair task has been completed, the tester agent performs vulnerability verification on the target code file after the vulnerability repair to determine whether the vulnerabilities in the repaired target code file are repaired. Specifically, the tester agent can re-execute the original vulnerability trigger use case and monitor the call chain information of the code instrumentation at the risk function to determine whether the risk function is still triggered, so as to verify whether the vulnerabilities in the target code file are repaired.
[0141] In a possible implementation manner, in order to ensure that the vulnerability mining strategy for the target code file can be run, a running environment that conforms to the target code file needs to be deployed. In the intelligent group of the embodiments of the present application, an operation and maintenance agent is further set up, and the operation and maintenance agent is used to deploy the running environment of the target code file.
[0142] See Figure 5 , which is a schematic flowchart of the deployment environment of an operation and maintenance agent provided by the embodiments of the present application, and specifically includes the following steps:
[0143] S301: Receive an environment deployment task from the project management agent; the environment deployment task includes multiple environment deployment steps.
[0144] S302: Generate operation prompt words for each of the environment deployment steps based on the large language model;
[0145] S303: Predict the expected execution results of each of the environment deployment steps according to the operation prompt words for each of the environment deployment steps;
[0146] S304: When the expected execution result does not meet the preset conditions, iteratively optimize the environment deployment steps where the expected execution result does not meet the preset conditions;
[0147] S305: If the number of optimization times of the iterative optimization reaches a preset third threshold, send a request for environment deployment assistance to the developer agent.
[0148] After receiving the environment deployment task from the project management agent, the operation and maintenance agent performs semantic analysis on each environment deployment step in the deployment task through a large language model, thereby generating an operation prompt word for each environment step, and predicting the expected execution result corresponding to each deployment step according to each operation prompt word. When the expected execution result does not meet the preset conditions, iteratively optimize the steps that do not meet the expected execution result, so as to ensure the smooth deployment of the operating environment. Correspondingly, if there are still situations that do not meet the expectations after the number of iterative optimizations reaches the preset third threshold, a request for environment deployment assistance is sent to the developer agent to build an operating environment for the target code file.
[0149] Specifically, for the collaborative relationship between the agents, reference can be made to Figure 6 the schematic diagram of a multi-agent collaborative structure disclosed, which will not be elaborated in this embodiment.
[0150] In a possible implementation manner, the target code file vulnerability mining system provided by the embodiments of the present application is also provided with a message middleware, specifically, reference can be made to Figure 7 the schematic diagram of another target code file mining system shown. The message middleware is used to realize asynchronous communication between the wholes in the agent group. During operation, the communication interaction between the agents is completed by the message middleware. The communication messages between the agents are not directly sent to the agent side, but sent to the message middleware. The message middleware performs standardized format processing on the communication messages, and clarifies the processing priority corresponding to each communication message, and then transmits the communication messages according to the actual priority to realize asynchronous communication between multiple agents in this system.
[0151] In the message middleware, there are a message standardization unit, a priority processing unit, and a message tracking unit. Next, the units included in the message middleware will be introduced one by one.
[0152] See Figure 8 , which is a schematic diagram of a message middleware provided by the embodiments of the present application.
[0153] Among them, the message standardization unit is used to receive the communication messages between the agents and format the communication messages to obtain multiple formatted communication messages; the formatted communication messages include: message type, processing priority, sender information, recipient information, message content, and tracking identifier;
[0154] The priority processing unit is used to implement asynchronous communication between the agents according to the processing priorities of the formatted communication messages.
[0155] The message tracking unit is used to perform communication link tracking on the formatted communication messages according to the tracking identifiers of the formatted communication messages.
[0156] In this way, by standardizing the communication protocols (message type / priority / tracking identifier) between the agents, the message middleware can ensure cross-role data flow between the agents. Correspondingly, by dividing the processing priorities of the communication messages, the work execution order between the agents can be accurately planned to achieve asynchronous communication and ensure timely response to critical vulnerability tasks. At the same time, each communication message has a corresponding tracking identifier for task backtracking and performance analysis.
[0157] The embodiments of the present application provide a vulnerability mining system, method and readable storage medium for target code files. The system includes an agent group and a large language model. The execution logic of each agent in the agent group is determined based on specific task keywords and the large language model, so as to improve the analysis ability for the dynamic characteristics of the target code file through the semantic understanding ability of the large language model. Among them, the project management agent is used to send an audit task to the security expert agent according to the target code file attribute information. The security expert agent determines the vulnerability analysis logic for the target code file according to the audit task and generates a vulnerability risk analysis result marked with a risk function based on this. The developer agent configures the call chain tracking for the risk functions included in the vulnerability risk analysis result through the abstract syntax tree and hash function of the target code file, and generates the first call chain information for the risk functions, so as to perform a preliminary static analysis on the call chain of the risk functions through the abstract syntax tree. Subsequently, the tester agent performs a dynamic running test on the target code file according to the first call chain information to obtain the actual call chain involved during the running of the risk functions, and obtains the second call chain information. In this way, through a combination of static analysis and dynamic analysis, code coverage is achieved at a deeper level on the basis of understanding the actual running situation of the code, thereby effectively improving the accuracy of vulnerability mining for the target code file. Finally, the second call chain information is fed back to the security expert agent and the developer agent to obtain the first vulnerability mining strategy and the second vulnerability mining strategy fed back by both, and the target code file is mined for vulnerabilities based on this. Thus, the vulnerability mining for the target code file is carried out by combining the vulnerability characteristics and the code logic in two aspects, effectively improving the accuracy of vulnerability mining for the target code file.
[0158] The following introduces a vulnerability mining method for target code files provided by the embodiments of the present application. The vulnerability mining method for target code files described below can be mutually referred to with the vulnerability mining system for target code files described above.
[0159] See Figure 9 , which is a schematic flowchart of a vulnerability mining method for target code files provided by the embodiments of the present application, and specifically includes the following steps:
[0160] S401: Control the project management agent to deploy an audit task according to the target code file attribute information, so as to send an audit task to the security expert agent;
[0161] S402: Control the security patent agent to determine the vulnerability analysis logic for the target code file according to the audit task, and generate a vulnerability risk analysis result based on the code content of the target code file and the vulnerability analysis logic; the vulnerability risk analysis result includes: the risk functions of the target code file;
[0162] S403: Control the developer agent to perform call chain tracing configuration on the risk function through the abstract syntax tree and hash function of the target code file, and generate first call chain information of the risk function; the first call chain information is used to represent the static call chain of the risk function.
[0163] S404: Control the tester agent to perform dynamic running tests on the target code file according to the first call chain information, obtain second call chain information, and send the second call chain information to the security expert agent and the developer agent to obtain a first vulnerability mining strategy and a second vulnerability mining strategy; the second call chain information is used to represent a call chain whose similarity to the first call chain is greater than a preset first threshold; the first vulnerability mining strategy performs vulnerability mining based on vulnerability characteristics and comes from the security expert agent; the second vulnerability mining strategy performs vulnerability mining based on code logic and comes from the developer agent.
[0164] S405: Control the tester agent to perform vulnerability mining on the target code file based on the first vulnerability mining strategy or the second vulnerability mining strategy, and obtain a vulnerability mining result.
[0165] Based on the same inventive concept, corresponding to the method of any of the above embodiments, an embodiment of the present application further provides a computer-readable storage medium. The computer-readable storage medium stores computer instructions, and the computer instructions are used to cause the computer to execute the target code file vulnerability mining method described in any of the above embodiments.
[0166] The computer-readable media of the embodiments of the present application include permanent and non-permanent, removable and non-removable media, and information storage can be implemented by any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory, or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD), or other optical storage, magnetic cassette tapes, magnetic tape magnetic disk storage, or other magnetic storage devices, or any other non-transmission medium that can be used to store information accessible by a computing device.
[0167] The computer instructions stored in the storage medium of the above embodiments are used to cause the computer to execute the target code file vulnerability mining method described in any of the above embodiments, and have the beneficial effects of the corresponding method embodiments, which will not be elaborated here.
[0168] It should be noted that the various embodiments in this specification are described in a progressive manner. For the same or similar parts among the various embodiments, reference can be made to each other, and the key point of each embodiment is to illustrate the differences from other embodiments. In particular, for methods, systems, and media, since they are basically similar to the method embodiments, the description is relatively simple, and reference can be made to the corresponding parts of the method embodiments for the relevant content. The methods, systems, and media described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components referred to as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. A person of ordinary skill in the art can understand and implement it without creative efforts.
[0169] As described above, this is only a specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed in the present application should be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. A target code file vulnerability mining system, characterized in that, Including: An agent group and a large language model; The execution logic of each agent in the agent group is determined based on task keywords and the large language model; the agent group includes: a project management agent, a security expert agent, a tester agent, and a developer agent; The project management agent is used to deploy an audit task according to the target code file attribute information, so as to send the audit task to the security expert agent; The security expert agent is used to determine the vulnerability analysis logic for the target code file according to the audit task, and generate a vulnerability risk analysis result based on the code content of the target code file and the vulnerability analysis logic; the vulnerability risk analysis result includes: the risk function of the target code file; The developer agent is used to perform call chain tracing configuration on the risk function through the abstract syntax tree and hash function of the target code file, and generate the first call chain information of the risk function; the first call chain information is used to represent the static call chain of the risk function; The tester agent is used to perform dynamic running tests on the target code file according to the first call chain information, obtain the second call chain information, and send the second call chain information to the security expert agent and the developer agent, so as to obtain a first vulnerability mining strategy and a second vulnerability mining strategy; the second call chain information is used to represent a call chain whose similarity to the first call chain is greater than a preset first threshold; the first vulnerability mining strategy is based on vulnerability characteristics for vulnerability mining and comes from the security expert agent; the second vulnerability mining strategy is based on code logic for vulnerability mining and comes from the developer agent; The tester agent is also used to perform vulnerability mining on the target code file based on the first vulnerability mining strategy or the second vulnerability mining strategy, and obtain a vulnerability mining result.
2. The system according to claim 1, wherein The developer agent is specifically used for: Performing call chain analysis according to the abstract syntax tree and the vulnerability risk analysis result to determine the static call chain of the risk function; Performing code instrumentation on the static call chain of the risk function based on the hash function to obtain the first call chain information.
3. The system according to claim 1, wherein The first call chain information includes: a risk function hash index and the entry information of the target code file; the tester agent includes: a dynamic test unit; the dynamic test unit is used for: Generating test data according to the risk function hash index and the entry information of the target code file; Performing a running test on the target code file based on the test data to obtain the third call chain information; the third call chain information is used to represent the call chain information and variable values of the risk function during runtime; Calculating the link similarity of each call chain between the third call chain information and the first call chain information based on the LCS algorithm, and determining the call chain information with a similarity greater than the preset first threshold in the third call chain information as the fourth call chain information; Mutate the third call chain information, calculate the link similarity of each call chain between the mutated call chain information and the first call chain information through the LCS algorithm, and determine the call chain information with a similarity greater than the preset first threshold as the fifth call chain information; Merge the fifth call chain information with the fourth call chain information to obtain the second call chain information.
4. The system according to claim 2, characterized in that, The dynamic testing unit is further configured to: If the sum of the number of call chains of the fifth call chain information and the fourth call chain information is less than the preset second threshold, send a call chain exploration strategy request to the security expert agent and the developer agent; Receive the first call chain exploration strategy and the second call chain exploration strategy in response to the call chain exploration strategy request; The first path exploration strategy performs path exploration based on the parameters of the vulnerability characteristics and comes from the security expert agent; the second call chain exploration strategy performs path exploration based on the URL parameters of the code structure and comes from the developer agent; Use the first call chain exploration strategy or the second call chain exploration strategy to perform call chain exploration to obtain the second call chain information; the execution priority of the first call chain exploration strategy is higher than that of the second call chain exploration strategy.
5. The system according to claim 1, characterized in that, The security expert agent includes: a first vulnerability mining unit; the developer agent includes: a second vulnerability mining unit; The first vulnerability mining unit is specifically configured to: Receive the second call chain information sent by the tester agent; Generate a vulnerability mining strategy based on vulnerability characteristics and a corresponding first verification parameter according to the second call chain information and the vulnerability risk analysis result; Determine the vulnerability mining strategy based on vulnerability characteristics and the first verification parameter as the first vulnerability mining strategy; The second vulnerability mining unit is specifically configured to: Receive the second call chain information sent by the tester agent; Generate a vulnerability mining strategy based on code logic and a corresponding second verification parameter according to the second call chain information and the abstract syntax tree; Determine the vulnerability mining strategy based on vulnerability characteristics and the second verification parameter as the second vulnerability mining strategy.
6. The system according to claim 5, characterized in that, The tester agent includes: a third vulnerability mining unit; the third vulnerability mining unit is specifically configured to: Use the first vulnerability mining strategy to perform vulnerability mining to obtain a first mining result; Verify the vulnerability of the first mining result through the first verification parameter to obtain a first verification result; When the first verification result fails, use the second vulnerability mining strategy to perform vulnerability mining to obtain a second mining result, and verify the vulnerability of the second mining result through the second verification parameter to obtain a second verification result; If the second verification result fails, send a mining strategy reconstruction request to the security expert agent.
7. The system according to claim 1, wherein The vulnerability discovery results include: vulnerability risk levels; the security expert agent includes: a repair strategy generation unit; the developer agent includes: a vulnerability repair unit; the tester agent includes: a code verification unit; The repair strategy generation unit is configured to: Generate a first vulnerability repair strategy based on the vulnerability discovery results and send the first vulnerability repair strategy to the vulnerability repair unit; The vulnerability repair unit is configured to: Generate a second vulnerability repair strategy based on the vulnerability discovery results and determine the strategy reference weights of the first vulnerability repair strategy and the second vulnerability repair strategy based on the vulnerability risk levels; Perform vulnerability repair on the target code file according to the first vulnerability repair strategy, the second vulnerability repair strategy, and the strategy reference weights to obtain a repaired target code file and send it to the code verification unit; The code verification unit is configured to: Verify the vulnerabilities in the repaired target code file to determine whether the vulnerabilities in the repaired target code file are repaired.
8. The system according to claim 1, characterized in that, The intelligent group also includes: an operation and maintenance agent; the operation and maintenance agent is specifically configured to: Receive an environment deployment task from the project management agent; the environment deployment task includes multiple environment deployment steps; Generate operation prompt words for each of the environment deployment steps based on the large language model; Predict the expected execution results of each of the environment deployment steps according to the operation prompt words of each of the environment deployment steps; When the expected execution results do not meet the preset conditions, iteratively optimize the environment deployment steps for which the expected execution results do not meet the preset conditions; If the number of optimization times for the iterative optimization reaches a preset third threshold, send an environment deployment assistance request to the developer agent.
9. The system according to claim 1, wherein The system further includes: a message middleware; the message middleware is used to implement asynchronous communication between the agents in the intelligent agent group; the message middleware includes: a message standardization unit, a priority processing unit, and a message tracking unit; The message standardization unit is configured to receive the communication messages between the agents and perform formatting processing on the communication messages to obtain multiple formatted communication messages; the formatted communication messages include: message type, processing priority, sender information, recipient information, message content, and tracking identifier; The priority processing unit is configured to implement asynchronous communication between the agents according to the processing priorities of the formatted communication messages; The message tracking unit is configured to perform communication link tracking on the formatted communication messages according to the tracking identifiers of the formatted communication messages.
10. A method for vulnerability mining of an object code file, characterized in that, Applied to a target code file vulnerability discovery system including an intelligent group and a large language model; The execution logics of the agents in the intelligent agent group are determined based on task keywords and the large language model; the intelligent agent group includes: a project management agent, a security expert agent, a tester agent, and a developer agent; the method includes: Control the project management agent to deploy an audit task according to the target code file attribute information, so as to send the audit task to the security expert agent; Control the security patent agent to determine the vulnerability analysis logic for the target code file according to the audit task, and generate a vulnerability risk analysis result based on the code content of the target code file and the vulnerability analysis logic; the vulnerability risk analysis result includes: the risk function of the target code file; Control the developer agent to configure the call chain tracing of the risk function through the abstract syntax tree and the hash function of the target code file, and generate the first call chain information of the risk function; the first call chain information is used to represent the static call chain of the risk function; Control the tester agent to perform dynamic running tests on the target code file according to the first call chain information, obtain the second call chain information, and send the second call chain information to the security expert agent and the developer agent to obtain the first vulnerability mining strategy and the second vulnerability mining strategy; the second call chain information is used to represent a call chain whose similarity to the first call chain is greater than a preset first threshold; the first vulnerability mining strategy mines vulnerabilities based on vulnerability characteristics and comes from the security expert agent; the second vulnerability mining strategy mines vulnerabilities based on code logic and comes from the developer agent; Control the tester agent to mine vulnerabilities in the target code file based on the first vulnerability mining strategy or the second vulnerability mining strategy, and obtain a vulnerability mining result.
11. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by a processor, it implements the target code file vulnerability mining method described in claim 10.
Citation Information
Cited By
Software defect collaborative detection method and system based on multi-agent dynamic adaptation
CN120560989A
Vulnerability analysis method and device, electronic equipment, medium and program product
CN121615146A