Attack method for vulnerability auditing tool based on big language model based on interference attention

By performing function-level attention calculations and compilation dependency completion on the sample code dataset, high-attention code snippets are generated, solving the detection challenges of large language model vulnerability auditing tools and achieving efficient improvement in vulnerability code concealment and attack success rate.

CN120893051AActive Publication Date: 2025-11-04NANJING UNIV
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202511416625.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-30
Publication Date
2025-11-04
Estimated Expiration
2045-09-30

AI Technical Summary

Technical Problem

Existing vulnerability auditing tools based on large language models are unable to effectively detect unseen vulnerable code, and traditional attack methods such as modifying variable names and code obfuscation are not very effective against them, resulting in vulnerable code being highly concealed and destructive.

Method used

By performing function-level attention calculations on the sample code dataset, a high-attention function code dataset is generated. The highest attention function is selected and the compilation dependencies are completed to form a high-attention code snippet. This snippet is then submitted to a vulnerability auditing tool for large language models for auditing to assess the attack success rate.

Benefits of technology

It significantly improves the concealment of vulnerable code, reduces the accuracy of large language models in code vulnerability detection tasks, has cross-model applicability and multi-programming language compatibility, and significantly improves the attack success rate.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120893051A_ABST
    Figure CN120893051A_ABST
Patent Text Reader

Abstract

The invention discloses an interference attention-based attack method for a vulnerability auditing tool based on a large language model, and the method comprises the steps: carrying out the function-level attention calculation of a sample code data set, and generating a high-attention function code data set; and selecting a highest attention function from each function-level attention source code file of the high attention function code data set, and complementing compiling dependencies to form a high attention code snippet. Function-level attention calculation is carried out on a sample code data set, and the attention calculation efficiency is improved by means of a dimension reduction algorithm; a high-attention function is selected and compilation dependencies are complemented, so that the high-attention function can be successfully compiled, and the concealment is improved; the method has remarkable cross-model applicability and multi-programming language compatibility, can effectively reduce the accuracy of a large language model on a code vulnerability detection task, and has the advantage of improving the imperceptibility of vulnerability codes.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of computer security, in particular to an attack method for a vulnerability auditing tool based on a large language model based on interference attention. BACKGROUND

[0002] With the progress of the related capabilities of large language models in the field of computer security, large language models are increasingly being applied to vulnerability detection tasks. After being trained using a dataset composed of a large amount of security code and vulnerability code, large language models can detect potential vulnerabilities in code across scenarios. Today, large language models are embedded in integrated development environments to provide security recommendations and risk prompts for code written by developers in real time. For example, GitHub Copilot is a code assistance tool based on a large language model jointly developed by GitHub, Microsoft and OpenAI. It supports code vulnerability detection and can be embedded as a plug-in in the integrated development environment VS Code. As of March 2024, GitHub Copilot has more than 13 million paying users and more than 50,000 enterprise-level users.

[0003] However, vulnerability auditing tools based on large language models introduce new attack surfaces. Consider the following scenario: a developer first downloads code from an online community (such as GitHub) to a local computer, then uses a vulnerability auditing tool based on a large language model to detect the security of the code, and after checking, uses the code for their own development project (such as a web server). If a piece of vulnerability code can bypass the detection of a vulnerability auditing tool based on a large language model, the vulnerability can be hidden in the developer's project. Then in the future, an attacker can exploit the vulnerability to attack the developer's project.

[0004] The challenge of attacking vulnerability auditing tools based on large language models lies in the fact that large language models learn a wide variety of vulnerability codes during the training phase and have strong reasoning capabilities. Therefore, as long as the principles of the vulnerability are similar, vulnerabilities that have not been seen during the training phase can be detected. Therefore, attack methods (such as modifying variable names, code obfuscation) against traditional machine learning and deep learning vulnerability detection tools are difficult to work.

[0005] In summary, attacking vulnerability auditing tools based on large language models is challenging. Once the attack is successful, it is highly covert and destructive, and has practicality in the field of computer security. SUMMARY

[0006] The present application aims to provide an attack method for a vulnerability auditing tool based on a large language model based on interference attention, to solve the technical problems raised in the background art.

[0007] To achieve the above object, the present application provides the following technical solutions: S1, function-level attention calculation is performed on the sample code data set to generate a high-attention function code data set, wherein the sample code data set contains a plurality of sample source code files C; specifically, each sample source code file C contains a plurality of functions , and a plurality of lines , and the high-attention function code data set contains a plurality of function-level attention source code files C', wherein each function-level attention source code file C' contains a function corresponding function-level attention; S2, the highest attention function is selected from each function-level attention source code file C' of the high-attention function code data set, and the compilation dependency is completed to form a high-attention code segment; S3, the high-attention code segment after the compilation dependency is completed is integrated with the vulnerability code data set and submitted to a vulnerability audit tool based on a large language model for auditing to evaluate the attack success rate.

[0008] Further, the specific steps of step S1 are as follows: S11, original attention generation: the original attention tensor of the sample source code file C by the large language model is , wherein T represents the token number of C after the large language model tokenizes C, L represents the number of Transformer layers of the large model, and H represents the number of attention multi-heads of the large model; S12, hierarchical attention generation: first, the original attention is summed up according to the fourth dimension to obtain the multi-head integrated attention ; then, using the last token in the first dimension T as the query key and the T tokens in the second dimension as the queried key, the two-dimensional tensor corresponding to the last token in the first dimension T is selected to obtain the hierarchical attention of all tokens ; S13, line-level attention generation: the first dimension T tokens in the hierarchical attention are summed up according to the index of the line to obtain one-dimensional line-level attention ; then, the first dimension of the one-dimensional line-level attention is summed up to combine all attention layers to obtain the line-level attention of each line , wherein i∈[1,M]; S14, function-level attention generation: the line-level attention is summed up according to the function Summing up, the function-level attention of each function is obtained wherein j∈[1, N]; S15, performing S11-S14 on each sample source code file C in the sample code data set to obtain a function-level attention source code file C', and forming a high-attention function code data set.

[0009] Further, the highest attention function in each function-level attention source code file C' in the high-attention function code data set is taken as a candidate, thereby generating a batch of high-attention function candidates; the compilation dependencies of the high-attention function are completed using a large language model, and the high-attention function and its compilation dependencies after completion can be self-compiled.

[0010] Further, the step S3 of evaluating the attack success rate is specifically: S31, marking the answers of the vulnerability audit tool based on the large language model when auditing a piece of vulnerability code with scores; specifically, marking 1 point if no vulnerability is found; marking 2 points if some vulnerabilities are found, but all are false positives and do not contain correct vulnerabilities; marking 3 points if some vulnerabilities are found and there are no other false positives, and the found vulnerabilities are correct vulnerabilities; marking 4 points if some vulnerabilities are found and there are false positives, but also contain correct vulnerabilities; S32, dividing the answers and corresponding scores into two levels of indicators to evaluate whether the vulnerability audit tool based on the large language model is successful, according to whether the vulnerabilities exist and whether the found vulnerabilities are correct, wherein the indicators include vulnerability existence level and type correctness level; S33, defining two success rates BSR@exist and BSR@type to evaluate the attack success rate of the high-attention code segment on the vulnerability audit tool based on the large language model.

[0011] Further, the step S32 of evaluating whether the vulnerability audit tool based on the large language model is successful is specifically: In the vulnerability existence level: scores of 2, 3 or 4 are successful in finding vulnerabilities, and the audit is successful; a score of 1 is not successful in finding vulnerabilities, and the audit fails; In the type correctness level: scores of 3 or 4 are successful in finding correct vulnerabilities, and the audit is successful; scores of 1 or 2 are not successful in finding correct vulnerabilities, and the audit fails.

[0012] Further, the step S33 of evaluating using the success rate is specifically: Using BSR@exist to evaluate: in the vulnerability existence level, the number of cases that fail to audit after integrating the high-attention code segment is divided by the number of cases that successfully audit before integrating the high-attention code segment; BSR@type evaluation: at the type correctness level, the number of cases that fail the audit after integrating the high-attention code snippets is divided by the number of cases that pass the audit before integrating the high-attention code snippets.

[0013] Beneficial effects: the function-level attention calculation is performed on the sample code dataset, the attention calculation efficiency is improved by means of the dimension reduction algorithm, the high-attention function is selected and the compilation dependency is completed, so that the high-attention function can be successfully compiled and the concealment is improved, the method has significant cross-model applicability and multi-programming language compatibility, can effectively reduce the accuracy of the large language model in the code vulnerability detection task, and has the advantages of improving the vulnerability code concealment. BRIEF DESCRIPTION OF DRAWINGS

[0014] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings required for the embodiments described below will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.

[0015] Figure 1 Flowchart of the attack method based on interference attention of the present application for the attack method based on the large language model vulnerability audit tool; Figure 2 Table of sample code dataset and vulnerability code dataset of the present application; Figure 3 Large language model details table of the present application; Figure 4 Attack effect table of the attack method and the benchmark method of the present application on 5 open source large language models on the Smart-bugs vulnerability dataset; Figure 5 Result table of the attack method of the present application verifying the cross-model migration; Figure 6 Result table of the attack method of the present application verifying the cross-language migration; Figure 7 Comparison result table of the semantic similarity of the vulnerability code before and after integrating the high-attention code snippets of the present application; Figure 8 Result table of the attention drop and drop rate of the vulnerability code before and after integrating the high-attention code snippet set with the highest attack success rate of the present application. DETAILED DESCRIPTION

[0016] With reference to the drawings of the embodiments of the present application, the technical solutions in the embodiments of the present application will be described clearly and completely. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments of the present application, all other embodiments obtained by those of ordinary skill in the art without creative work fall within the protection scope of the present application.

[0017] As shown in Figures 1-8 , the present application provides an attack method for a vulnerability auditing tool based on a large language model based on interference attention, and the specific steps are as follows: Please refer to Figure 2 , the present embodiment downloads a sample code dataset for generating high attention code snippets, including an intelligent contract dataset Messiq, a C / C++ dataset Leetcode-cpp and a Python dataset Leetcode-py; Download the vulnerability code dataset for attacking developers, including the intelligent contract vulnerability dataset Smart-bugs, the C / C++ vulnerability dataset Big-Vul and the Python vulnerability dataset CVE-Fixes; Since the large model has a limit on the number of tokens when calculating attention, the source code files with a total number of tokens less than 4096 are selected.

[0018] Please refer to Figure 3 , download the pre-trained large language model from the open source community Huggingface, and complete the local deployment. There are 5 open source large language models (Mistral, MixtralExpert, Gemma, CodeLlama, Phi), and the access license API Key of GPT-4o is purchased, so that the closed source large language model GPT-4o can be used through the API Key; S1, function-level attention calculation is performed on the sample code dataset to generate a high attention function code dataset, wherein the sample code dataset contains a plurality of sample source code files C; specifically, each sample source code file C contains a plurality of functions , and a plurality of lines , the high attention function code dataset contains a plurality of function-level attention source code files C', wherein each function-level attention source code file C' contains a function corresponding function-level attention, each sample source code file C can contain a vulnerability or not, and is only used to generate a high attention function; In this embodiment, in order to illustrate the implementation of step S1, the specific steps are as follows: S11, original attention generation: the original attention tensor of the sample source code file C of the large language model is wherein T represents the token quantity of C after tokenization of C by the large language model, L represents the number of Transformer layers of the large model, H represents the number of attention heads of the large model, and the first dimension of T in A is used as a query key and the second dimension of T is used as a core key to be queried; S12, hierarchical attention generation: first, the original attention is summed up in the fourth dimension to obtain multi-head integrated attention The formula can be expressed as: ; Then, the last token in the first dimension T is used as a query key and the T tokens in the second dimension are used as queried keys, and the two-dimensional tensor corresponding to the last token in the first dimension T is selected to obtain the hierarchical attention of all tokens The formula can be expressed as: ; S13, row-level attention generation: the hierarchical attention is summed up in the first dimension T tokens according to the index of the row to obtain one-dimensional row-level attention ; then, the first dimension of the one-dimensional row-level attention is summed up to combine all attention layers to obtain the row-level attention of each row wherein i∈[1,M], the formula can be expressed as: , S14, function-level attention generation: the row-level attention is summed up according to the function to which each row belongs to obtain the function-level attention of each function wherein j∈[1,N], the formula can be expressed as: ; S15, performing S11-S14 on each sample source code file C in the sample code data set to obtain the function-level attention source code file C', forming a high-attention function code data set.

[0019] S2, selecting the highest attention function from each function-level attention source code file C' in the high-attention function code data set and completing the compilation dependencies to form a high-attention code snippet; specifically, the highest attention function in each function-level attention source code file C' is used as a candidate, thereby generating a batch of high-attention function candidates; the compilation dependencies of the high-attention function are completed using a large language model (such as GPT-4o), and the completed high-attention function and its compilation dependencies can be self-compiled.​

[0020] S3, integrate the high attention code segment after completing the compilation dependency with the vulnerability code dataset, and submit it to the vulnerability audit tool based on the large language model for audit, and evaluate the attack success rate. The standard for attack success is that the vulnerability audit tool based on the large language model cannot detect the existence of vulnerabilities in the vulnerability code. In order to the vulnerability audit tool based on the large language model in the real world, the prompt word design will use some methods to improve the quality of the answer, such as Retrieval-Augmented-Generation, Chain-of-Thought and In-Context-Learning.

[0021] The embodiment designs an evaluation attack success rate for the vulnerability audit tool based on the large language model, as follows: S31, mark the answer of the vulnerability audit tool based on the large language model when auditing a piece of vulnerability code with a score. The embodiment designs four kinds of answers, which are marked with 1-4 points in turn, as follows: no vulnerability is found, marked as 1 point; several vulnerabilities are found, but all are false positives and do not contain correct vulnerabilities, marked as 2 points; several vulnerabilities are found, and there is no other false positive, and the found vulnerabilities are correct vulnerabilities, marked as 3 points; several vulnerabilities are found, there are false positives, but also contain correct vulnerabilities, marked as 4 points.

[0022] S32, according to whether the vulnerability exists and whether the discovered vulnerability is correct, the answer and the corresponding score are divided into two levels of indexes to evaluate the audit success of the vulnerability audit tool based on the large language model, wherein the indexes include vulnerability existence level and type correctness level. The embodiment evaluates whether the vulnerability audit tool based on the large language model is successful, as follows: In the vulnerability existence level: scores of 2, 3 or 4 are successful in discovering vulnerabilities, and the audit is successful; score of 1 is not successful in discovering vulnerabilities, and the audit fails. This level measures the detection ability of the large language model on the existence of vulnerabilities. In the type correctness level: scores of 3 or 4 are successful in discovering correct vulnerabilities, and the audit is successful; scores of 1 or 2 are not successful in discovering correct vulnerabilities, and the audit fails. This level measures the detection ability of the large language model on the type of correct vulnerabilities.

[0023] S33, define two attack success rates BSR@exist and BSR@type to evaluate the attack success rate of the high attention code segment on the vulnerability audit tool based on the large language model. Among them, BSR@exist is used to evaluate: in the vulnerability existence level, the number of cases of audit failure (score 1) after integrating the high attention code segment of the vulnerability dataset divided by the number of cases of audit success (score 2, 3 or 4) before integrating the high attention code segment; BSR@type is used to evaluate: in the type correctness level, the number of cases of audit failure (score 1 or 2) after integrating the high attention code segment of the vulnerability dataset divided by the number of cases of audit success (score 3 or 4) before integrating the high attention code segment.

[0024] Verify the effectiveness of the attack method proposed in the application, as follows: Using the attack method proposed in the application and two baseline methods (random insertion (inserting irrelevant interference code in the vulnerability function) and renaming (harmless renaming of variable functions in the vulnerability code)), attack different large language models on the Smart-bugs vulnerability dataset, and evaluate the effectiveness of the attack according to the attack success rate defined by BSR@exist and BSR@type.

[0025] As shown in Figure 4 , the embodiment shows the performance comparison of the attack method proposed in the application on different large language model vulnerability audit tools. Each column reflects the success rate of the application compared with the baseline method on the Smart-bugs dataset. Whether in BSR@exist or BSR@type, the application outperforms other baseline methods on all five models, showing its consistent efficiency. Specifically: taking BSR@exist as the evaluation index, the attack success rate of the application on the Phi model is improved from 0 to 17.32% compared with the random insertion method. Compared with the renaming method, the attack success rate of the application on the Mistral model is improved from 0 to 45.31%, and the improvement range on other models is from 1.36 times (CodeLlama) to 49.14 times (Gemma). When BSR@type is used as the evaluation standard, the application is always better than the random method, and the improvement range is from 1.77 times (CodeLlama) to 43.99 times (Phi). Compared with the renaming method, the attack success rate of the application is improved by 1.27 times (CodeLlama) to a maximum of 73 times (Mistral).

[0026] Verify the cross-model transferability of the attack method proposed in the application, as follows: The evaluation assesses the transferability of the high-attention code snippets that are optimal for attack effectiveness on a single model to other models. Specifically, the embodiment assesses the attack effectiveness of the high-attention code snippets that are optimal for BSR@exist and BSR@type, respectively, on other models using the Smart-bugs vulnerability dataset as the evaluation benchmark. The results show that the present application has strong transferability between different models.

[0027] As shown in Figure 5 , the embodiment demonstrates the attack effectiveness of the optimal high-attention code snippets on each model on other models. Under the BSR@exist indicator, the optimal high-attention code snippets on the Gemma model show the highest transferability, achieving an attack success rate of 87.5% on the Mistral model and 18.11% on the Phi model. Similarly, under the BSR@type indicator, the optimal high-attention code snippets on the Gemma model also maintain a high success rate, exceeding 90% on the Mistral model and 40% on the Phi model. Notably, the optimal high-attention code snippets on the Gemma model also show an attack success rate of 30% on the closed-source model GPT-4o. These results show that the attack method generated by the present application has strong cross-model transferability and shows high effectiveness on both open-source models and closed-source models.

[0028] To verify the cross-language transferability of the attack method proposed by the present application, the following is performed: To verify the scalability of the attack method proposed by the present application in different languages, the embodiment plans to use the same experimental setup as in the verification of the effectiveness of the attack method proposed by the present application to test the effectiveness of the attack method proposed by the present application on additional datasets containing other languages (such as C, C++, and Python). The vulnerability auditing tools based on large language models for extended evaluation include Mixtral, CodeLlama, and Phi. Similarly, BSR@exist and BSR@type are selected as evaluation indicators.

[0029] As shown in Figure 6 , the present application performs well in terms of attack success rate, especially on the Big-Vul dataset, reaching 100% in both indicators (BSR@exist and BSR@type) of the Phi model. In contrast, the random insertion and renaming methods have lower overall performance on the Phi model, with success rates below 50% in multiple cases.

[0030] To verify the stealthiness of the attack method proposed by the present application, the following is performed: To evaluate the stealthiness of the attack method proposed by the present application, the semantic similarity before and after integrating the high-attention code snippets on multiple vulnerability code datasets is measured in this embodiment. The best-performing high-attention code snippets on each target open-source model are used, and are inserted in the code of different datasets, including Smart-bugs (Solidity), Big-Vul (C / C++) and CVE-Fixes (Python). To quantify the semantic change, a frozen language model (FLM), specifically Alibaba-NLP / gte-large-en-v1.5, is used to evaluate the semantic similarity by calculating the cosine similarity of embedding vectors.

[0031] As shown in Figure 7 , this embodiment shows the semantic similarity values (minimum, maximum and mean) of each model on different datasets. Overall, the present application shows high stealthiness, with high mean similarity scores for each dataset indicating that the stealthiness of the code modification is high enough that developers cannot detect it. For example, CodeLlama always achieves the highest mean similarity in each dataset, with a value of about 0.92 or higher, indicating minimal semantic deviation after insertion. In contrast, the similarity score of Phi is slightly lower, especially on the Smart-bugs dataset (average value of 0.8179), but still maintains a high level of similarity. These results show that the present application can effectively preserve the semantics of the original code, thereby achieving a high level of stealthiness.

[0032] To verify the rationality of the attack method proposed by the present application, the following is done: To address the question of whether the high-attention code snippets that successfully attack will actually cause the function-level attention of the vulnerability function to decrease, this embodiment verifies the best high-attention code snippets on each model on the Smart-bugs dataset, verifying the decrease in function-level attention of the vulnerability function before and after it is integrated into the vulnerability code.

[0033] To verify the influence of high-attention code snippets on the attention of vulnerability code, specifically: by calculating the function-level attention of the vulnerability function in the vulnerability code before and after integrating the high-attention code snippets, the effectiveness of the high-attention code snippets in distracting attention is verified; calculating the vulnerability function attention: the function-level attention of the vulnerability function before and after integrating the high-attention code snippets is calculated respectively; estimating attack effectiveness: by calculating the decrease in function-level attention of the vulnerability function and the decrease rate, the effectiveness of the attack is estimated.

[0034] As shown in Figure 8As shown, in the Phi model, the high attention code fragment causes the function level attention of the vulnerability function to decrease by the largest amount, reaching 325.89; and in the Gemma model, the high attention code fragment achieves the largest decrease rate, which is 0.92. This shows that the high attention code fragment generated by the present application can significantly reduce the function level attention of the vulnerability function, thereby helping it to escape the vulnerability audit of the large language model.

[0035] The application discloses an attack method for a vulnerability audit tool based on a large language model based on interference attention, which calculates function level attention for a sample code data set, improves attention calculation efficiency with a dimension reduction algorithm, selects high attention functions and completes compilation dependencies, so that the high attention functions can be successfully compiled and the concealment is improved. The method has significant cross-model applicability and multi-programming language compatibility, can effectively reduce the accuracy of the large language model in the code vulnerability detection task, and has the advantages of improving the vulnerability of the code.

[0036] It should be apparent to those skilled in the art that the application is not limited to the details of the above-described exemplary embodiments, but can be implemented in other specific forms without departing from the spirit or essential characteristics of the application. Therefore, the embodiments should be considered as exemplary and non-limiting, and the scope of the application is defined by the appended claims, not the above description, and all changes falling within the meaning and scope of the equivalent elements of the claims are intended to be included in the application. Any reference signs in the claims should not be considered as limiting the claims involved.

[0037] In addition, it should be understood that although the present specification is described in terms of embodiments, not every embodiment contains only one independent technical solution, and the description of the specification is only for the sake of clarity. Those skilled in the art should consider the specification as a whole, and the technical solutions in each embodiment can be appropriately combined to form other embodiments that those skilled in the art can understand.

Claims

1. An attack method based on interference attention targeting large language model-based vulnerability auditing tools, characterized in that, Includes the following steps: S1. Perform function-level attention computation on the sample code dataset to generate a high-attention function code dataset. The sample code dataset contains several sample source code files C; specifically, each sample source code file C contains several functions. and several lines The high attention function code dataset contains several function-level attention source code files C', where each function-level attention source code file C' contains functions. The corresponding function-level attention; S2. Select the highest attention function from each function-level attention source file C' in the high attention function code dataset and complete the compilation dependencies to form a high attention code snippet; S3. After completing the compilation dependencies, the high-attention code snippet is integrated with the vulnerability code dataset and submitted to a vulnerability auditing tool based on a large language model for auditing and evaluation of the attack success rate.

2. The attack method according to claim 1, characterized in that, The specific steps of step S1 are as follows: S11, Original Attention Generation: The original attention tensor of the large language model for the sample source code file C is... Where T represents the number of tokens of C after the large language model segments C, L represents the number of Transformer layers of the large model, and H represents the number of attention heads of the large model; S12, Hierarchical Attention Generation: First, the original attention... Summing along the fourth dimension yields the multi-head integrated attention. Then, use Using the last token in the first dimension (T tokens) as the query key and the T tokens in the second dimension as the query keys, we select the two-dimensional tensor corresponding to the last token in the first dimension (T tokens) to obtain the hierarchical attention of all tokens. ; S13, Row-level Attention Generation: Generating hierarchical attention The first dimension of the T tokens is indexed according to the row they belong to. Summation yields a one-dimensional row-level attention. Then, perform one-dimensional row-level attention. Summing the first dimension and merging all attention layers yields the row-level attention for each row. where i∈[1,M]; S14. Function-level attention generation: This involves generating row-level attention. Based on the function of each row Summing is performed to obtain the function-level attention for each function. , where j∈[1,N]; S15. Perform S11~S14 on each sample source code file C in the sample code dataset to obtain the function-level attention source code file C', forming a high-attention function code dataset.

3. The attack method according to claim 1, characterized in that: Step S2 specifically involves: using the highest attention function in each function-level attention source file C' of the high attention function code dataset as a candidate, thereby generating a batch of high attention function candidates; The compilation dependencies of high attention functions are completed using a large language model, and the completed high attention functions and their compilation dependencies can compile themselves.

4. The attack method according to claim 1, characterized in that: The specific steps for evaluating the attack success rate in step S3 are as follows: S31. Mark the responses of the vulnerability auditing tool based on the large language model when auditing a piece of vulnerable code with scores; specifically: 1 point for no vulnerability found; 2 points for finding several vulnerabilities, but all of them are false alarms and do not include the correct vulnerability; 3 points for finding several vulnerabilities and there are no other false alarms, and the found vulnerability is the correct vulnerability; 4 points for finding several vulnerabilities, some of which are false alarms, but also some of which are correct vulnerabilities. S32. The vulnerability auditing tool based on the large language model is evaluated for its success by dividing the answers and corresponding scores into two levels of indicators based on whether the vulnerability exists and whether the vulnerability is correctly discovered. The indicators include vulnerability existence level and type correctness level. S33. Define two success rates, BSR@exist and BSR@type, to evaluate the success rate of attacks on high-attention code snippets by vulnerability auditing tools based on large language models.

5. The attack method according to claim 4, characterized in that: The specific steps in step S32 for evaluating whether the vulnerability auditing tool based on the large language model has successfully audited are as follows: At the vulnerability presence level: a score of 2, 3, or 4 indicates a successful vulnerability discovery and audit success; a score of 1 indicates a failure to discover the vulnerability and audit failure. At the type correctness level: a score of 3 or 4 indicates that the correct vulnerability was successfully found and the audit was successful; a score of 1 or 2 indicates that the correct vulnerability was not successfully found and the audit failed.

6. The attack method according to claim 4, characterized in that: The success rate assessment in step S33 specifically involves: Using BSR@exist for assessment: At the vulnerability existence level, the number of cases where the vulnerability dataset failed to be audited after integrating high-attention code snippets is divided by the number of cases that were successfully audited before integrating high-attention code snippets; Using BSR@type for evaluation: At the type correctness level, the number of cases in the vulnerability dataset that failed to be audited after integrating high-attention code snippets is divided by the number of cases that succeeded in auditing before integrating high-attention code snippets.

Citation Information

Patent Citations

  • Cross-language vulnerability detection system based on hierarchical source code representation learning method

    CN118332559A

  • Hybrid model driven reentrant vulnerability maintenance method and device and computer equipment

    CN119691747A

  • Function-level code vulnerability detection and analysis method based on language model

    CN120216325A

  • Localizing vulnerabilities in source code at a token-level

    US20240411666A1