Code vulnerability processing method
By constructing a false positive detection model and retrieving enhanced generative agents, the challenges of cross-project reuse and maintenance of code vulnerability false positive detection are solved, improving the accuracy and efficiency of security scanning and reducing maintenance costs.
Patent Information
- Application Number
- CN202511516393.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-22
- Publication Date
- 2025-11-21
- Estimated Expiration
- 2045-10-22
AI Technical Summary
In existing technologies, false positive detection of code vulnerabilities relies on the specific code structure of a project, making it difficult to reuse across projects. As the project scales up, the false positive list becomes lengthy and difficult to maintain. Code modifications can cause false positives to reappear, increasing maintenance costs.
Initial vulnerability data is obtained by scanning target code snippets, forming an initial vulnerability feature set and constructing category feature vectors. The vulnerability feature set is filtered using a false positive detection model. Combined with retrieval enhancement, the generated intelligent agent determines the false positive and non-false positive vulnerability feature sets and executes corresponding processing actions to generate actual vulnerability data.
It enables false alarm detection and handling across projects, reduces the length of false alarm lists, improves the accuracy and efficiency of security scanning, and lowers maintenance costs.
Smart Images

Figure CN120995470A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of vulnerability management, and particularly relates to a code vulnerability processing method. BACKGROUND
[0002] In the related art, false positives can be manually maintained in a filter file, i.e., the identified false positives are recorded in a filter list (such as a whitelist), and in subsequent detection using a security scanning tool, these detection results are automatically ignored; or a code pattern can be matched by a preset rule, and a context logic is combined to determine whether it is a valid vulnerability.
[0003] However, in the related art, the false positives cannot be automatically identified and filtered when a new project introduces similar code, and as the project scale expands, the false positive list can become lengthy and difficult to manage, which cannot fundamentally reduce the workload of security audits and increases maintenance costs. In addition, when the code is modified (such as refactoring, version updating, etc.), the originally recorded false positives can no longer match the filter rules due to position offset, resulting in the reappearance of false positives, which needs to be improved. SUMMARY
[0004] The present application provides a code vulnerability processing method to at least solve the problems in the related art that it is difficult to cross-project reuse due to the dependence on project-specific code structure, the false positive list is lengthy and difficult to maintain as the project scale expands, and code modification causes false positives to reappear due to position changes, increasing maintenance costs.
[0005] The present application provides a code vulnerability processing method, including the following steps: scanning a target code segment of a software to be tested to obtain initial vulnerability data corresponding to the target code segment, extracting at least one initial vulnerability feature of the initial vulnerability data to form an initial vulnerability feature set, and obtaining a category feature vector corresponding to the initial vulnerability feature set; inputting the initial vulnerability feature set and the category feature vector into a pre-constructed false positive detection model to filter the initial vulnerability feature set using the false positive detection model to obtain a vulnerability feature set that meets a preset condition; inputting the vulnerability feature set into a pre-constructed search enhancement generation agent to determine a false positive vulnerability feature set and a non-false positive vulnerability feature set in the vulnerability feature set, and a first processing action corresponding to the false positive vulnerability feature set and a second processing action corresponding to the non-false positive vulnerability feature set using the search enhancement generation agent; respectively executing the first processing action and the second processing action to obtain a first processing result of the false positive vulnerability feature set and a second processing result of the non-false positive vulnerability feature set, and determining actual vulnerability data of the target code segment according to the false positive vulnerability feature set, the non-false positive vulnerability feature set, the first processing result, and the second processing result.
[0006] The application further provides a code vulnerability processing apparatus, comprising: an acquisition module, configured to scan a target code segment of a software to be tested to acquire initial vulnerability data corresponding to the target code segment, extract at least one initial vulnerability feature of the initial vulnerability data to form an initial vulnerability feature set, and acquire a category feature vector corresponding to the initial vulnerability feature set; a generation module, configured to input the initial vulnerability feature set and the category feature vector into a false positive detection model constructed in advance to filter the initial vulnerability feature set by using the false positive detection model to obtain a vulnerability feature set meeting a preset condition; a first determination module, configured to input the vulnerability feature set into a retrieval enhancement generation intelligent agent constructed in advance to determine, by using the retrieval enhancement generation intelligent agent, a false positive vulnerability feature set and a non-false positive vulnerability feature set in the vulnerability feature set, a first processing action corresponding to the false positive vulnerability feature set, and a second processing action corresponding to the non-false positive vulnerability feature set; and a second determination module, configured to execute the first processing action and the second processing action respectively to obtain a first processing result of the false positive vulnerability feature set and a second processing result of the non-false positive vulnerability feature set, and determine actual vulnerability data of the target code segment according to the false positive vulnerability feature set, the non-false positive vulnerability feature set, the first processing result and the second processing result.
[0007] The application further provides an electronic device, comprising: a memory configured to store a computer program; and a processor configured to execute the computer program to implement the steps of the code vulnerability processing method.
[0008] The application further provides a computer readable storage medium, wherein the computer readable storage medium stores a computer program, and the computer program is executed by a processor to implement the steps of the code vulnerability processing method.
[0009] The application further provides a computer program product, comprising a computer program, and the computer program is executed by a processor to implement the steps of the code vulnerability processing method.
[0010] Through the present application, by scanning the target code segment of the software to be tested, the initial vulnerability data is obtained, and by extracting the vulnerability features, the initial vulnerability feature set is formed, and then the corresponding category feature vector is obtained, so as to utilize the false positive detection model to screen the initial vulnerability feature set, and then obtain the vulnerability feature set that meets certain conditions, and then utilize the pre-constructed search enhancement to generate an intelligent agent to determine the false positive vulnerability feature set and the non-false positive vulnerability feature set, and the corresponding processing actions, by executing the processing actions to obtain the corresponding processing results, so as to determine the actual vulnerability data, therefore, the technical problems that the project-specific code structure is relied on, it is difficult to cross-project reuse, as the project scale expands, the false positive list is lengthy and difficult to maintain, and the code modification causes the false positives to reappear due to the position change, increasing the maintenance cost can be solved, real-time analysis and processing of the scanned vulnerability data can not only accurately detect the false positive features in the code scanning, but also generate effective processing schemes, thereby improving the accuracy of the security scanning and the work efficiency of the people. BRIEF DESCRIPTION OF DRAWINGS
[0011] In order to more clearly illustrate the embodiments of the present application, the drawings needed in the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor on the basis of these drawings.
[0012] Figure 1 The flowchart of the code vulnerability processing method provided according to the embodiments of the present application; Figure 2 The flowchart of feature selection provided according to the embodiments of the present application; Figure 3 The flowchart of Rag intelligent agent construction provided according to the embodiments of the present application; Figure 4 The flowchart of the intelligent vulnerability management platform executed according to the embodiments of the present application; Figure 5 The block schematic diagram of the code vulnerability processing device provided according to the embodiments of the present application.
[0013] Reference signs: Among them, 10-Code vulnerability processing device; 100-First acquisition module, 200-Generation module, 300-First determination module, 400-Second determination module. DETAILED DESCRIPTION
[0014] With reference to the drawings of the embodiments of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described. Obviously, the described embodiments are only a part of the embodiments of the present application, but not all embodiments of the present application. Based on the embodiments of the present application, all other embodiments obtained by a person of ordinary skill in the art without creative work fall within the protection scope of the present application.
[0015] It should be noted that in the description of the present application, the terms "comprising", "containing" or any other variants thereof are intended to cover non-exclusive inclusion, so that the process, method, article or equipment including a series of elements not only includes those elements, but also includes other elements not explicitly listed or inherent to such process, method, article or equipment. The terms "first", "second" and the like in the present application are used to distinguish similar objects, and are not used to describe a specific order or sequence.
[0016] In order for those skilled in the art to better understand the present application, the present application will be further described in detail below with reference to the drawings and specific embodiments.
[0017] The embodiments of the present application provide a code vulnerability processing method. The method is described in detail in combination with the execution flow of the code vulnerability processing method.
[0018] Specifically, Figure 1 A flowchart of a code vulnerability processing method according to an embodiment of the present application is shown.
[0019] As Figure 1 shown, the code vulnerability processing method includes the following steps: In step S101, the target code segment of the software to be tested is scanned to obtain initial vulnerability data corresponding to the target code segment, at least one initial vulnerability feature of the initial vulnerability data is extracted to form an initial vulnerability feature set, and a category feature vector corresponding to the initial vulnerability feature set is obtained.
[0020] It can be understood that in the embodiments of the present application, the category feature vector can include but is not limited to a false positive vector and a non-false positive vector. The specific setting can be performed by a person skilled in the art according to the actual situation, and the present application does not make specific limitation.
[0021] In some embodiments, the embodiments of the present application can scan the target code segment of the software to be tested, and then obtain the initial vulnerability data corresponding to the target code segment, so as to extract at least one initial vulnerability feature from the initial vulnerability data, and then form an initial vulnerability feature set by summarizing these features, and then obtain a category feature vector corresponding to the initial vulnerability feature set.
[0022] For example, embodiments of this application can collect target code fragments in security scans, which may include, but are not limited to, false alarm data and non-false alarm data from different languages, projects, and products, and the number of false alarm and non-false alarm data is not significantly different.
[0023] Furthermore, in this embodiment of the application, the extracted target code fragments can be represented by an abstract syntax tree and a structured representation of the syntax features of the code to obtain initial vulnerability data, and initial vulnerability features can be extracted from them, including the number of lines of code, cyclomatic complexity, node type, number of loop layers, nesting depth, number of calls, etc. The coding language type and scanning tool type are also used as initial vulnerability features, thereby forming an initial vulnerability feature set, and obtaining the category feature vector corresponding to the initial vulnerability feature set.
[0024] Optionally, in one embodiment of this application, before inputting the initial vulnerability feature set and category feature vector into the pre-constructed false positive detection model, the method further includes: scanning the target code segment of the target software to obtain initial target vulnerability data corresponding to the target code segment; extracting at least one initial target vulnerability feature from the initial target vulnerability data to form an initial target vulnerability feature set to obtain a category target feature vector corresponding to the initial target vulnerability feature set; constructing a first evaluation function based on the correlation between at least one initial target vulnerability feature and the category target feature vector in the initial target vulnerability feature set to calculate a first correlation value between the initial target vulnerability feature and the category target feature vector; constructing a second evaluation function based on the correlation between different initial target vulnerability features in the initial target vulnerability feature set to calculate a second correlation value between different initial target vulnerability features; constructing a third evaluation function based on the first and second evaluation functions to calculate a third correlation value between a subset of initial target vulnerability features and the category target feature vector; and constructing an evaluation function for the false positive detection model based on the first, second, and third evaluation functions.
[0025] It is understood that, in the embodiments of this application, the evaluation function is used to measure the ability of a feature or set to distinguish different categories of data. It can be used to evaluate the merits of candidate subsets and is one of the important factors affecting feature selection performance.
[0026] Furthermore, embodiments of this application may employ symmetric uncertainty (abbreviated as symmetric uncertainty). The evaluation criteria are used to measure the degree of correlation between features and classes, and between features themselves. That is, for variables... , The symmetric uncertainty between them can be, but is not limited to, expressed as: , wherein, refers to the mutual information between two variables; refers to the information entropy of a variable; The value range of is [0-1], and the greater the value is, the stronger the correlation between the two is, otherwise the weaker the correlation is.
[0027] In some embodiments, the embodiments of the present application can construct a first evaluation function based on the correlation between the initial target vulnerability features and the category target feature vector in the initial target vulnerability feature set and the category target feature vector. wherein, is the total number of initial target vulnerability features, is the category feature vector.
[0028] The correlation between the initial target vulnerability feature and the category feature vector can be but is not limited to represented as: , Further, the embodiments of the present application can construct the first evaluation function, and then calculate the first correlation value between the initial target vulnerability feature and the category target feature vector by using the first evaluation function.
[0029] In some embodiments, the embodiments of the present application can construct a second evaluation function based on the correlation between different initial target vulnerability features in the initial target vulnerability feature set.
[0030] wherein, The correlation between the initial target vulnerability feature and the initial target vulnerability feature can be but is not limited to represented as: , Further, the embodiments of the present application can construct the second evaluation function, and then calculate the second correlation value between different initial target vulnerability features by using the second evaluation function.
[0031] In some embodiments, the embodiments of the present application can construct a third evaluation function based on the first evaluation function and the second evaluation function.
[0032] It can be understood that the embodiments of the present application can use the subset correlation (Subset Correlation, abbreviated as SC) to measure the correlation between a feature subset and a class, that is: , wherein, is the number of features; The value range of is [0-1], and the greater the value is, the stronger the correlation between the feature set and the class is, otherwise the weaker the correlation is; an average value of the correlation between all features and the class, an average value of the correlation between features, which can be expressed as, but not limited to: , , Further, the embodiment of the present application can construct a third evaluation function, and then calculate a third correlation value between the initial target vulnerability feature subset and the class target feature vector by using the third evaluation function.
[0033] Before inputting the initial vulnerability feature set and the class feature vector into the false positive detection model, the embodiment of the present application can obtain the initial target vulnerability data and the feature set and the class target feature vector by scanning the target code segment, and then construct the first and second evaluation functions, so as to calculate the correlation values between the initial target vulnerability features and the class target feature vector and different initial target vulnerability features, and construct the third evaluation function according to the first and second evaluation functions, so as to calculate the correlation value between the initial target vulnerability feature subset and the class target feature vector, finally construct the evaluation function of the false positive detection model, improve the reasoning efficiency, avoid model overfitting, improve the model generalization ability, and make the model have project adaptive ability and evolutionary learning ability.
[0034] In step S102, the initial vulnerability feature set and the class feature vector are input into the false positive detection model constructed in advance, so as to filter the initial vulnerability feature set by using the false positive detection model, and obtain a vulnerability feature set meeting the preset condition.
[0035] It can be understood that, in order to improve the efficiency of false positive detection, the embodiment of the present application uses the feature set after feature selection when constructing the false positive detection model. It can be understood that, the embodiment of the present application can perform feature selection based on the clustering mode, and then filter out a vulnerability feature set with lower redundancy and higher correlation with the class, so as to improve the efficiency and accuracy of the false positive detection model. The feature selection is a process of selecting a vulnerability feature set meeting certain conditions from the initial vulnerability feature set according to a certain evaluation function, such as the process of selecting an optimal feature subset. The process of performing feature selection can include, but is not limited to, removing irrelevant features, feature clustering, removing redundant features, etc. The present application does not make specific limitation, and the specific process is shown in Figure 2 .
[0036] In addition, the certain condition can be set by a person skilled in the art according to the actual situation, and the present application does not make specific limitation.
[0037] In some embodiments, the embodiments of the present application can utilize the pre-constructed false positive detection model to screen the initial vulnerability feature set based on the category feature vector, and then obtain a vulnerability feature set that meets certain conditions. In addition, the embodiments of the present application can also preprocess the initial vulnerability feature set, such as deduplication, normalization, etc., which are not specifically limited by the present application.
[0038] For example, the embodiments of the present application can utilize the pre-constructed false positive detection to perform real-time and accurate false positive detection on the initial vulnerability feature set. In addition, in order to construct a stable and efficient false positive detection model, the embodiments of the present application can perform the following operations on the initial vulnerability feature set: Obtain an initial target vulnerability feature set corresponding to the initial vulnerability feature set, and train and verify the initial target vulnerability feature set in a ten-fold cross-validation manner. That is, divide the initial target vulnerability feature set into ten parts, select nine parts for training and one part for verification to obtain the classification accuracy; perform ten times, calculate the average of the ten classification accuracies as the final classification accuracy; and select a variety of classifiers (such as NB (Naive Bayes), SVM (Support Vector Machine), RF (Random Forest), etc.) to classify the initial target vulnerability feature set, and select the classifier with the highest classification accuracy as the classifier of the false positive detection model, thereby obtaining the final false positive detection model.
[0039] Optionally, in an embodiment of the present application, the initial vulnerability feature set and the category feature vector are input into the pre-constructed false positive detection model to screen the initial vulnerability feature set by using the false positive detection model to obtain a vulnerability feature set that meets a preset condition, including: based on a first evaluation function in the false positive detection model, calculating a first correlation value between different initial vulnerability features and the category feature vector; judging whether the first correlation value is less than a preset threshold; if the first correlation value is less than the preset threshold, eliminating the corresponding initial vulnerability feature to determine the vulnerability feature set; and if the first correlation value is greater than or equal to the preset threshold, retaining the corresponding initial vulnerability feature to determine the vulnerability feature set.
[0040] In some embodiments, the embodiments of the present application can utilize the false positive detection model to screen the initial vulnerability feature set, that is, by calculating the first correlation value between different initial vulnerability features and the category feature vector through the first evaluation function, and judging whether the first correlation value is less than a certain threshold, so that when it is less than, the corresponding initial vulnerability feature is eliminated, and then the vulnerability feature set is determined, and when it is greater than or equal to a certain threshold, the corresponding initial vulnerability feature is retained, and then the vulnerability feature set is determined, wherein the certain threshold can be set by a person skilled in the art according to the actual situation, and the present application does not make specific limitations.
[0041] For example, in combination withFigure 2 As shown in the embodiment of this application, when moving irrelevant features, the degree of correlation between each initial vulnerability feature and the category feature vector can be calculated through a first evaluation function to obtain a first correlation value. The initial vulnerability features are then sorted in descending order based on the first correlation value, and a certain threshold is set. The feature set is obtained by retaining features whose first relevance value is greater than a certain threshold. .
[0042] This application embodiment calculates the correlation between the initial vulnerability features and the category feature vector through the first evaluation function in the false positive detection model, and removes features with a first correlation value less than a certain threshold, thereby reducing interference from irrelevant features, improving detection accuracy, reducing computational complexity, and improving model efficiency. It can be flexibly adjusted according to different application scenarios and needs to improve the model's operating efficiency.
[0043] Optionally, in one embodiment of this application, an initial vulnerability feature set and category feature vector are input into a pre-built false alarm detection model to filter the initial vulnerability feature set using the false alarm detection model, thereby obtaining a vulnerability feature set that meets preset conditions. This includes: calculating a second correlation value between different initial vulnerability features based on a second evaluation function in the false alarm detection model, and constructing a feature association set for different initial vulnerability features based on the second correlation value; determining an initial cluster subset and other cluster subsets of the feature association set based on different initial vulnerability features and the feature association set; traversing other cluster subsets with the initial cluster subset as a reference to determine the corresponding remaining feature set and initial cluster set; and calculating the corresponding final cluster set based on the remaining feature set and the initial cluster set, thereby determining the vulnerability feature set using the final cluster set.
[0044] In some embodiments, the present application embodiments may utilize a second evaluation function to calculate a second correlation value between different initial vulnerability features, and construct a feature association set for different initial vulnerability features based on the second correlation value. Thus, based on different initial vulnerability features and feature association sets, an initial cluster subset and other cluster subsets of the feature association set are determined. Using the initial cluster subset as a reference, other cluster subsets are traversed to determine the corresponding remaining feature set and initial cluster set. Then, based on the remaining feature set and initial cluster set, the corresponding final cluster set is calculated, thereby determining the vulnerability feature set.
[0045] For example, in combination Figure 2 As shown, in the embodiment of this application, during feature clustering, the vulnerability feature set can be... Each feature in , and a second evaluation function is used to calculate a second correlation value between the feature and other features, the greater the second correlation value, the greater the correlation, indicating that the two features are more similar; further, the embodiments of the present application can sort the correlation between the features in descending order to obtain a feature association set corresponding to the correlation degree, which can be but not limited to represented as: , wherein, is the second correlation value between the two features (denoted as ) ranked th, and the calculation formula can be but not limited to represented as: , Further, the embodiments of the present application can perform the following processing on the data in : taking corresponding two features as an initial clustering subset , other as other clustering subsets, and taking the initial clustering subset as a reference, traversing the other clustering subsets to determine the corresponding remaining feature set and the initial clustering set , wherein each clustering subset contains at least one feature.
[0046] Further, the embodiments of the present application calculate the corresponding final clustering set based on the remaining feature set and the initial clustering set, thereby determining the vulnerability feature set.
[0047] The embodiments of the present application calculate the correlation value between the initial vulnerability features through the second evaluation function in the false positive detection model, thereby constructing the feature association set, determining the initial clustering subset and the other clustering subset, and taking the initial clustering subset as a reference, traversing to determine the remaining feature set and the initial clustering set, and further obtaining the final clustering set, thereby determining the vulnerability feature set. Through the feature clustering driven by the second evaluation function, the dynamic grouping optimization based on the inherent correlation between the features is realized, the feature redundancy is intelligently eliminated, and the false positive detection accuracy is greatly improved.
[0048] Optionally, in an embodiment of the present application, taking the initial clustering subset as a reference, traversing the other clustering subsets to determine the corresponding remaining feature set and the initial clustering set, comprising: judging whether the initial clustering subset contains at least one other clustering feature in the other clustering subset; if the initial clustering subset contains at least one other clustering feature, determining the remaining feature set based on the other clustering features not contained; if the initial clustering subset does not contain other clustering features, determining the initial clustering set based on the other clustering subset corresponding to the other clustering features.
[0049] In some embodiments of this application, during the process of determining the remaining feature set and the initial cluster set, it can be determined whether the initial cluster subset contains one of the other cluster features of other cluster subsets. If it does, the other cluster features that are not included are added to the remaining feature set, thereby determining the remaining feature set; if it does not, the other cluster subsets are combined into a new initial cluster set, thereby determining the initial cluster set.
[0050] For example, the embodiments of this application are for ,Pick The two corresponding features, if Include If one feature is included, then the other unincluded feature is added to the remaining feature set. In the middle; if None of them include If two features are selected, then these two features are combined to form a new initial cluster set. .
[0051] In this embodiment, the initial cluster subset is used as a reference to traverse other cluster subsets. By determining whether it contains features from other cluster subsets, the remaining feature set and the initial cluster set are determined. The remaining feature set is formed by extracting features that are not included, thus avoiding feature omission. The complete cluster subset is directly used as the initial cluster set, preserving complete semantics and reducing computational complexity.
[0052] Optionally, in one embodiment of this application, calculating the corresponding final cluster set based on the remaining feature set and the initial cluster set includes: calculating a first correlation value between the remaining features and the corresponding category feature vector based on a first evaluation function in the false alarm detection model; sorting the remaining features according to a first target sorting method based on the first correlation value to obtain sorted remaining features, and determining the final remaining feature set based on the sorted remaining features; traversing the final remaining feature set and the initial cluster set to calculate the effective information of different final remaining features in the initial cluster subset; determining whether the effective information meets a preset information condition; if the effective information meets the preset information condition, adding the corresponding final remaining feature to the corresponding initial cluster subset to obtain the final cluster subset, and obtaining the final cluster set based on the final cluster subset.
[0053] In some embodiments, the embodiments of the present application can calculate a first correlation value between the remaining features and the corresponding category feature vectors by using the first evaluation function, and sort them according to a first target sorting manner (such as from large to small, from small to large, etc., which is not specifically limited in the present application), and then obtain the sorted remaining features, and determine the final remaining feature set according to the sorted remaining features, traverse the final remaining feature set and the initial clustering set, calculate the effective information of different final remaining features in the initial clustering subset, and judge whether the effective information meets a certain information condition, if it meets, add the corresponding final remaining feature to the corresponding initial clustering subset to obtain the final clustering subset, so as to determine the final clustering set. The certain information condition can be set by those skilled in the art according to the actual situation, and the present application does not make specific limitations.
[0054] For example, the embodiments of the present application can sort the remaining features in the remaining feature set in descending order according to the first correlation value of the features and the category feature vectors, and then determine the final remaining feature set traverse the final remaining features in , perform the following operations: traverse each initial clustering subset in the initial clustering set , and calculate the effective information brought by in : suppose that the effective information brought by the feature in the clustering subset is the least, then add the feature to , that is, , , After all the features in are traversed, the final clustering set is obtained.
[0055] Among them, the calculation method of the effective information brought by the feature in the clustering subset may be but is not limited to: calculating the correlation between and the category feature vector , denoted as ; calculating the correlation between and the category feature vector , denoted as ; the effective information - .
[0056] This application embodiment uses a first evaluation function to calculate the first correlation value between the remaining features and the category feature vector, and determines the final remaining feature set after sorting. By traversing it and the initial cluster set, the corresponding effective information is calculated, and the final cluster set is obtained when the effective information meets certain information conditions, thereby improving computational efficiency, reducing maintenance costs, and strengthening the ability to suppress false alarms.
[0057] Optionally, in one embodiment of this application, determining the vulnerability feature set using the final cluster set includes: calculating a first correlation value between different cluster features and corresponding category feature vectors in the final cluster set based on a first evaluation function in the false positive detection model; sorting the different cluster features according to a second target sorting method based on the first correlation value to obtain a sorted cluster set; traversing the sorted cluster set and determining the corresponding initial cluster feature set to calculate the effective information of different sorted cluster features in the initial cluster feature set; and determining the vulnerability feature set based on the effective information.
[0058] In some embodiments, the present application embodiments may use a first evaluation function to calculate the first correlation value between different clustering features and corresponding category feature vectors in the final clustering set, and sort them according to a second target sorting method (such as from smallest to largest, from largest to smallest, etc., the present application does not make specific restrictions), thereby obtaining a sorted clustering set, and traversing the sorted clustering set to determine the corresponding initial clustering feature set, and calculating the effective information, thereby determining the vulnerability feature set.
[0059] For example, in combination Figure 2 As shown, embodiments of this application can, when removing redundant features, optimize the final cluster set. For each cluster subset, the following operations are performed: Calculate the first correlation value using the first evaluation function, and then sort the final cluster sets in descending order according to the first correlation value, thus obtaining the sorted cluster sets. .
[0060] Furthermore, embodiments of this application can... The first feature As the initial cluster feature set , ; obtain The first feature , ,calculate exist The effective information brought in, if the effective information is greater than 0, then ; Traversal Repeat the above operation on all clustered subsets until... All sets in the subset are empty, thus obtaining the final clustering feature set. This allows us to determine the vulnerability signature set.
[0061] The embodiment of the present application calculates the first correlation value between different clustering features and corresponding category feature vectors through the first evaluation function, sorts them, and then determines the initial clustering feature level and calculates the effective information, and then determines the vulnerability feature set. This quantitative method avoids subjective judgment and provides an objective and accurate basis for feature screening, ensuring that the screened features are highly related to the vulnerability categories, improving feature screening accuracy, enhancing vulnerability detection effectiveness, improving model running efficiency, not relying on specific vulnerability types or application scenarios, having strong universality, and being able to adapt to various complex actual situations.
[0062] Optionally, in an embodiment of the present application, before inputting the vulnerability feature set into the pre-constructed retrieval enhancement generation agent, it further includes: obtaining a target vulnerability feature set; constructing a vector database based on the target vulnerability feature set; screening a target final candidate feature vector based on the vector database; constructing a large language model based on the target final candidate feature vector; and constructing a retrieval enhancement generation agent based on the vector database and the large language model.
[0063] It can be understood that the pre-constructed retrieval enhancement generation (Rag) agent, knowledge base preparation and vector retrieval of the embodiment of the present application are the core steps of generating high-quality detection information.
[0064] Among them, the embodiment of the present application takes the target vulnerability feature corresponding to the target vulnerability feature set as a retrievable index vector, thereby constructing a vector database, and using the similarity between features and features when retrieving data as an evaluation criterion, and then finding the most similar information to the query data in the knowledge base to obtain the target final candidate feature vector, and then constructing a large language model, and constructing a retrieval enhancement generation agent according to the vector database and the large language model.
[0065] For example, the process of constructing the Rag agent in the embodiment of the present application is as shown in Figure 3 The main content can be: Step S301: Knowledge preparation.
[0066] Among them, the knowledge preparation includes the index vector corresponding vulnerability code snippet, vulnerability information, false positive reason, etc.
[0067] Further, the embodiment of the present application can upload the target vulnerability feature set to the vector database, including vulnerability code snippets, target vulnerability features, vulnerability information, false positive reasons, etc. Among them, the target vulnerability feature is a retrievable index vector, and other information is associated metadata. The feature vector replaces the calculation of the similarity of traditional high-dimensional vectors, significantly improving the retrieval efficiency.
[0068] Step S302: vector retrieval.
[0069] Step S303: result generation.
[0070] In the embodiments of the present application, the prediction result of whether it is a false positive can be predicted based on the target vulnerability feature through the false positive detection model; and according to the prediction result, all records with the same prediction result in the vector database are queried as a candidate set; and then the similarity between the target vulnerability feature and the feature vectors in the candidate set is calculated The most similar TOP-K records are retrieved from the candidate set to obtain the target final candidate feature vector, and the target final candidate feature vector is input into the large language model as a prompt word to generate a more accurate answer.
[0071] It should be noted that the embodiments of the present application adopt The higher the correlation between the feature vectors, the more similar the code snippets, so the most similar TOP-K records can be retrieved.
[0072] The embodiments of the present application can obtain a target vulnerability feature set, construct a vector database, and then use the vector database to filter a target final candidate feature vector, thereby constructing a large language model, and then constructing a retrieval-enhanced generation agent, which can more accurately locate features related to the target vulnerability, avoid possible semantic bias, improve retrieval accuracy and efficiency, help the large language model better understand the context and features of the vulnerability, and thus generate more accurate and targeted answers and suggestions, improve the performance of the model in vulnerability detection and processing, and enable the agent to more comprehensively and accurately handle vulnerability-related problems, adapt to complex and variable vulnerability scenarios.
[0073] In step S103, the vulnerability feature set is input into the pre-constructed retrieval-enhanced generation agent to determine the false positive vulnerability feature set and the non-false positive vulnerability feature set in the vulnerability feature set, and the first processing action corresponding to the false positive vulnerability feature set and the second processing action corresponding to the non-false positive vulnerability feature set.
[0074] It can be understood that the Rag agent in the embodiments of the present application can integrate the retrieved information and the problem as a prompt word input to the large language model, thereby generating high-quality and accurate processing information.
[0075] In some embodiments, the Rag agent constructed in advance can be used to determine the false positive vulnerability feature set and the non-false positive vulnerability feature set in the vulnerability feature set, and the first processing action corresponding to the false positive vulnerability feature set and the second processing action corresponding to the non-false positive vulnerability feature set.
[0076] For the false positive vulnerability feature set, the first processing action can be determined by using a large language model in the Rag agent; for the non-false positive vulnerability feature set, the second processing action can be determined by using a large language model in the Rag agent, and the specific setting can be performed by a person skilled in the art according to the actual situation, and the present application does not make specific limitations.
[0077] Optionally, in an embodiment of the present application, the vulnerability feature set is input into a pre-constructed retrieval enhancement generation agent to determine the false positive vulnerability feature set and the non-false positive vulnerability feature set in the vulnerability feature set, and the first processing action corresponding to the false positive vulnerability feature set and the second processing action corresponding to the non-false positive vulnerability feature set by using the retrieval enhancement generation agent, comprising: determining the index vector corresponding to the vulnerability feature set based on the vector database in the retrieval enhancement generation agent; using the index vector as an index, screening the initial candidate feature vector satisfying the preset detection condition from the vector database; based on the initial candidate feature vector, screening the final candidate feature vector satisfying the preset similarity condition; based on the final candidate feature vector, using the large language model in the retrieval enhancement generation agent to determine the false positive vulnerability feature set, the non-false positive vulnerability feature set, the first processing action and the second processing action.
[0078] In some embodiments, the vulnerability feature set of the present application can be indexed by the vulnerability feature of the vulnerability feature set, and the initial candidate feature vector satisfying a certain detection condition is screened from the vector database, and then the final candidate feature vector satisfying a certain similarity condition is screened through the initial candidate feature vector, so as to determine the false positive vulnerability feature set, the non-false positive vulnerability feature set, the first processing action and the second processing action by using the large language model. The certain detection condition and the certain similarity condition can be set by a person skilled in the art according to the actual situation, and the present application does not make specific limitations.
[0079] For example, the embodiment of the present application can screen the initial candidate feature vector satisfying a certain detection condition from the vector database based on the vulnerability feature, as a candidate set; and then the similarity between the vulnerability feature and the feature vector in the candidate set is calculated , the most similar TOP-K record in the candidate set is retrieved as the final candidate feature vector, and the final candidate feature vector is input into the large language model as a prompt word to determine the false positive vulnerability feature set, the non-false positive vulnerability feature set, the first processing action and the second processing action, and generate the false positive reason for the false positive vulnerability feature set; and generate the repair scheme for the non-false positive vulnerability feature set.
[0080] The embodiment of the application determines the corresponding index vector by searching and enhancing the vector database in the agent, and then screens the initial candidate feature vector that meets certain detection conditions by taking the index vector as an index, and further screens the final candidate feature vector that meets certain similarity conditions based on the initial candidate feature vector, so as to determine the false positive vulnerability feature set, the non-false positive vulnerability feature set, the first processing action and the second processing action by using the large language model in the search and enhancement generation agent, improve the accuracy and comprehensiveness of feature retrieval, reduce manual intervention and tedious analysis steps, shorten the time of vulnerability processing, improve the processing efficiency, can respond to security threats in the system in a timely manner, and have good adaptability and expansibility.
[0081] In step S104, the first processing action and the second processing action are respectively performed to obtain the first processing result of the false positive vulnerability feature set and the second processing result of the non-false positive vulnerability feature set, and the actual vulnerability data of the target code fragment is determined according to the false positive vulnerability feature set, the non-false positive vulnerability feature set, the first processing result and the second processing result.
[0082] As a possible implementation manner, the embodiment of the application can determine the actual vulnerability data in the target code fragment by performing the first processing action and the second processing action to determine the corresponding first processing result and the second processing result.
[0083] For example, in the case that the processing action of the false positive vulnerability feature set is to generate false positive reasons, the false positive reasons can be manually audited, and in the case that it is manually determined that the feature is false positive, it is removed, and the actual vulnerability data is determined.
[0084] Further, in the case that the processing action of the non-false positive vulnerability feature set is to generate a repair scheme, the corresponding repair scheme can be executed, and the actual vulnerability data is determined after the vulnerability repair is completed.
[0085] In addition, the embodiment of the application can combine the false positive detection model, the Rag agent and the code scanning tool together to construct an intelligent vulnerability management platform, realize intelligent processing of vulnerability scanning, false positive detection and vulnerability analysis, and the specific process is as shown in Figure 4 .
[0086] Wherein, 1 represents that the target code segment is scanned by using the code scanning tool to obtain initial vulnerability data; 2 represents that the initial vulnerability data is parsed by the intelligent vulnerability management platform, and features are extracted to form an initial vulnerability feature set, and a category feature vector corresponding to the initial vulnerability feature set is obtained; 3 represents that the initial vulnerability feature set is parsed and reasoned by using the false positive detection model to obtain a vulnerability feature set; 4 represents that the vulnerability feature set is searched and generated by the Rag intelligent agent to obtain a false positive vulnerability feature set, a non-false positive vulnerability feature set, a first processing action corresponding to the false positive vulnerability feature set and a second processing action corresponding to the non-false positive vulnerability feature set; 5 represents that the false positive vulnerability feature set, the non-false positive vulnerability feature set, the first processing result and the second processing result are synchronized to the knowledge base; 6 represents that the processing action is executed to obtain a corresponding processing result, and the processing result is returned to the intelligent vulnerability management platform, and a developer determines actual vulnerability data according to a repair scheme to process the vulnerability.
[0087] The embodiment of the application only processes vulnerability data, and can improve the accuracy of scanning, improve work efficiency and optimize the automatic vulnerability detection process.
[0088] According to the code vulnerability processing method provided in the embodiment of the application, the target code segment of the to-be-tested software is scanned to obtain initial vulnerability data, and the vulnerability features are extracted to form an initial vulnerability feature set, and then a corresponding category feature vector is obtained, so that the initial vulnerability feature set is screened by using the false positive detection model, and then a vulnerability feature set meeting certain conditions is obtained, and then the Rag intelligent agent constructed by searching enhancement is used to determine the false positive vulnerability feature set and the non-false positive vulnerability feature set, and the corresponding processing action is determined, the processing result corresponding to the processing action is obtained by executing the processing action, and the actual vulnerability data is determined, so that the technical problems of relying on the specific code structure of the project, being difficult to reuse across projects, the false positive list being long and difficult to maintain as the project scale expands, and the false positives being reproduced due to the position change caused by code modification and increasing the maintenance cost can be solved, real-time analysis and processing of the scanned vulnerability data can be achieved, the false positive features in the code scanning can be accurately detected, and an effective processing scheme can be generated, so that the technical effects of improving the accuracy of security scanning and the work efficiency of people are achieved.
[0089] Those skilled in the art can clearly understand from the above description of the embodiments that the method according to the above embodiments can be realized by means of software and a general hardware platform as required, of course, and can also be realized by hardware, but in many cases, the former is a better embodiment.
[0090] The embodiment of the application further provides a code vulnerability processing device.
[0091] Figure 5 A block schematic diagram of the code vulnerability processing device provided according to the embodiment of the application is shown.
[0092] As shown in the code vulnerability processing apparatus 10, Figure 5 includes a first acquisition module 100, a generation module 200, a first determination module 300, and a second determination module 400.
[0093] The first acquisition module 100 is configured to scan a target code segment of a software to be tested to obtain initial vulnerability data corresponding to the target code segment, extract at least one initial vulnerability feature of the initial vulnerability data, form an initial vulnerability feature set, and obtain a category feature vector corresponding to the initial vulnerability feature set.
[0094] The generation module 200 is configured to input the initial vulnerability feature set and the category feature vector into a pre-constructed false positive detection model, filter the initial vulnerability feature set by using the false positive detection model, and obtain a vulnerability feature set satisfying a preset condition.
[0095] The first determination module 300 is configured to input the vulnerability feature set into a pre-constructed retrieval enhancement generation agent, determine a false positive vulnerability feature set and a non-false positive vulnerability feature set in the vulnerability feature set, and a first processing action corresponding to the false positive vulnerability feature set and a second processing action corresponding to the non-false positive vulnerability feature set by using the retrieval enhancement generation agent.
[0096] The second determination module 400 is configured to perform the first processing action and the second processing action respectively to obtain a first processing result of the false positive vulnerability feature set and a second processing result of the non-false positive vulnerability feature set, and determine actual vulnerability data of the target code segment according to the false positive vulnerability feature set, the non-false positive vulnerability feature set, the first processing result, and the second processing result.
[0097] Optionally, in an embodiment of the present application, the code vulnerability processing apparatus further includes a second acquisition module, a first construction module, a second construction module, a third construction module, and a fourth construction module.
[0098] The second acquisition module is configured to scan the target code segment of the target software to obtain initial target vulnerability data corresponding to the target code segment, extract at least one initial target vulnerability feature of the initial target vulnerability data, form an initial target vulnerability feature set, and obtain a category target feature vector corresponding to the initial target vulnerability feature set before inputting the initial vulnerability feature set and the category feature vector into the pre-constructed false positive detection model.
[0099] The first construction module is configured to construct a first evaluation function based on a correlation between the at least one initial target vulnerability feature in the initial target vulnerability feature set and the category target feature vector, and calculate a first correlation value between the initial target vulnerability feature and the category target feature vector by using the first evaluation function.
[0100] The second constructing module is configured to construct a second evaluation function based on the correlation of different initial target vulnerability features in the initial target vulnerability feature set, so as to calculate a second correlation value between different initial target vulnerability features by using the second evaluation function.
[0101] The third constructing module is configured to construct a third evaluation function based on the first evaluation function and the second evaluation function, so as to calculate a third correlation value between the initial target vulnerability feature subset and the category target feature vector by using the third evaluation function.
[0102] The fourth constructing module is configured to construct an evaluation function of the false alarm detection model based on the first evaluation function, the second evaluation function and the third evaluation function.
[0103] Optionally, in an embodiment of the present application, the generating module 200 comprises a first calculating unit, a judging unit, a first determining unit and a second determining unit.
[0104] The first calculating unit is configured to calculate a first correlation value between different initial vulnerability features and a category feature vector based on a first evaluation function in the false alarm detection model.
[0105] The judging unit is configured to judge whether the first correlation value is less than a preset threshold.
[0106] The first determining unit is configured to eliminate the corresponding initial vulnerability feature when the first correlation value is less than the preset threshold, so as to determine the vulnerability feature set.
[0107] The second determining unit is configured to retain the corresponding initial vulnerability feature when the first correlation value is greater than or equal to the preset threshold, so as to determine the vulnerability feature set.
[0108] Optionally, in an embodiment of the present application, the generating module 200 comprises a second calculating unit, a third determining unit, a fourth determining unit and a third calculating unit.
[0109] The second calculating unit is configured to calculate a second correlation value between different initial vulnerability features based on a second evaluation function in the false alarm detection model, so as to construct a feature association set of different initial vulnerability features according to the second correlation value.
[0110] The third determining unit is configured to determine an initial clustering subset and other clustering subsets of the feature association set based on different initial vulnerability features and the feature association set.
[0111] The fourth determining unit is configured to traverse the other clustering subsets with reference to the initial clustering subset, so as to determine a corresponding remaining feature set and the initial clustering set.
[0112] The third computing unit is configured to calculate a corresponding final cluster set based on the residual feature set and the initial cluster set, so as to determine the vulnerability feature set by using the final cluster set.
[0113] Optionally, in an embodiment of the present application, the fourth determining unit comprises a first judging subunit, a first determining subunit and a second determining subunit.
[0114] The first judging subunit is configured to judge whether the initial cluster subset contains at least one other cluster feature in the other cluster subset.
[0115] The first determining subunit is configured to determine the residual feature set based on the other cluster features not contained when the initial cluster subset contains at least one other cluster feature.
[0116] The second determining subunit is configured to determine the initial cluster set based on the other cluster subset corresponding to the other cluster features when the initial cluster subset does not contain the other cluster features.
[0117] Optionally, in an embodiment of the present application, the third computing unit comprises a first computing subunit, a third determining subunit, a second computing subunit, a second judging subunit and a generating subunit.
[0118] The first computing subunit is configured to calculate a first correlation value between the residual feature and the corresponding category feature vector based on a first evaluation function in the false positive detection model.
[0119] The third determining subunit is configured to sort the residual feature according to a first target sorting manner based on the first correlation value to obtain a sorted residual feature, and determine a final residual feature set according to the sorted residual feature.
[0120] The second computing subunit is configured to traverse the final residual feature set and the initial cluster set to calculate effective information of different final residual features in the initial cluster subset.
[0121] The second judging subunit is configured to judge whether the effective information satisfies a preset information condition.
[0122] The generating subunit is configured to add the corresponding final residual feature into the corresponding initial cluster subset to obtain a final cluster subset when the effective information satisfies the preset information condition, and obtain the final cluster set according to the final cluster subset.
[0123] Optionally, in an embodiment of the present application, the third computing unit comprises a third computing subunit, a fourth computing subunit and a fourth determining subunit.
[0124] The third calculation subunit is configured to calculate first correlation values between different clustering features and corresponding category feature vectors in the final clustering set based on a first evaluation function in the false detection model, and sort the different clustering features according to a second target sorting manner based on the first correlation values to obtain a sorted clustering set.
[0125] The fourth calculation subunit is configured to traverse the sorted clustering set, determine a corresponding initial clustering feature set, and calculate effective information of different sorted clustering features in the initial clustering feature set.
[0126] The fourth determination subunit is configured to determine the vulnerability feature set based on the effective information.
[0127] Optionally, in an embodiment of the present application, the method further includes a third acquisition module, a fifth construction module, a screening module, a sixth construction module, and a seventh construction module.
[0128] The third acquisition module is configured to acquire a target vulnerability feature set before inputting the vulnerability feature set into the retrieval enhancement generation agent.
[0129] The fifth construction module is configured to construct a vector database based on the target vulnerability feature set.
[0130] The screening module is configured to screen a target final candidate feature vector based on the vector database.
[0131] The sixth construction module is configured to construct a large language model based on the target final candidate feature vector.
[0132] The seventh construction module is configured to construct the retrieval enhancement generation agent based on the vector database and the large language model.
[0133] Optionally, in an embodiment of the present application, the first determination module 300 includes a fifth determination unit, a first screening unit, a second screening unit, and a sixth determination unit.
[0134] The fifth determination unit is configured to determine an index vector corresponding to the vulnerability feature set based on the vector database in the retrieval enhancement generation agent.
[0135] The first screening unit is configured to screen an initial candidate feature vector satisfying a preset detection condition from the vector database by taking the index vector as an index.
[0136] The second screening unit is configured to screen a final candidate feature vector satisfying a preset similarity condition based on the initial candidate feature vector.
[0137] The sixth determination unit is configured to determine the false positive vulnerability feature set, the non-false positive vulnerability feature set, the first processing action and the second processing action based on the final candidate feature vector and by using the retrieval enhancement to generate the large language model in the agent.
[0138] According to the code vulnerability processing apparatus provided in the embodiments of the present application, the target code segment of the to-be-tested software is scanned to obtain initial vulnerability data, and the initial vulnerability feature set is formed by extracting vulnerability features, and then the corresponding category feature vector is obtained, so that the initial vulnerability feature set is screened by using the false positive detection model, and then the vulnerability feature set meeting certain conditions is obtained, and then the false positive vulnerability feature set and the non-false positive vulnerability feature set and the corresponding processing actions are determined by using the retrieval enhancement generated agent, the corresponding processing result is obtained by executing the processing actions, and the actual vulnerability data is determined, so that the technical problems that the project-specific code structure is relied on and it is difficult to cross-project reuse, the false positive list is lengthy and difficult to maintain as the project scale expands, and the false positives are reproduced due to the position change caused by code modification and the maintenance cost is increased can be solved, real-time analysis and processing of the scanned vulnerability data are achieved, not only the false positive features in the code scanning can be accurately detected, but also effective processing schemes can be generated, and the technical effects of improving the accuracy of the security scanning and the work efficiency of people are achieved.
[0139] The features in the embodiments of the code vulnerability processing apparatus can refer to the related descriptions of the embodiments of the code vulnerability processing method, which will not be described herein.
[0140] Embodiments of the present application also provide an electronic device, comprising a memory and a processor, the memory stores a computer program, and the processor is configured to run the computer program to perform the steps in any of the above code vulnerability processing method embodiments.
[0141] Embodiments of the present application also provide a computer readable storage medium, which stores a computer program, wherein the computer program is configured to perform the steps in any of the above code vulnerability processing method embodiments when running.
[0142] In an example embodiment, the above computer readable storage medium can include, but is not limited to, a U disk, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk or an optical disk, and various media that can store computer programs.
[0143] Embodiments of the present application also provide a computer program product, which comprises a computer program, and the computer program is executed by a processor to implement the steps in any of the above code vulnerability processing method embodiments.
[0144] Embodiments of the present application further provide another computer program product comprising a non-transitory computer readable storage medium storing a computer program which, when executed by a processor, implements the steps of any of the above code vulnerability processing method embodiments.
[0145] Those skilled in the art will further appreciate that the units and algorithm steps of the examples described in connection with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or both, and that the implementation decisions are within the skill of an expert in the art to make based on their specific application and design constraints. The examples have been described in general terms in the above description solely for purposes of clarity and understanding and not by limitation. The specific implementation decisions regarding, whether the functions are performed in hardware or software lie within the purview of the person of skill in the art.
[0146] The above has carried out the detailed introduction to the code vulnerability processing method provided by the present application. The principle and implementation mode of the present application are described by applying specific examples in the present text, and the above example description is only for helping to understand the method of the present application and its core idea. It should be pointed out that, for the ordinary skilled in the art, some improvements and modifications can be made to the present application without departing from the principle of the present application, and these improvements and modifications also fall within the protection scope of the claims of the present application.
Claims
1. A method of handling a code vulnerability, characterized by, The method comprises the following steps: scanning a target code segment of a software to be tested to obtain initial vulnerability data corresponding to the target code segment, extracting at least one initial vulnerability feature of the initial vulnerability data to form an initial vulnerability feature set, and obtaining a category feature vector corresponding to the initial vulnerability feature set; inputting the initial vulnerability feature set and the category feature vector into a false positive detection model constructed in advance to filter the initial vulnerability feature set by using the false positive detection model, and obtaining a vulnerability feature set meeting a preset condition; inputting the vulnerability feature set into a retrieval enhancement generation agent constructed in advance to determine a false positive vulnerability feature set and a non-false positive vulnerability feature set in the vulnerability feature set, and a first processing action corresponding to the false positive vulnerability feature set and a second processing action corresponding to the non-false positive vulnerability feature set by using the retrieval enhancement generation agent; respectively executing the first processing action and the second processing action to obtain a first processing result of the false positive vulnerability feature set and a second processing result of the non-false positive vulnerability feature set, and determining actual vulnerability data of the target code segment according to the false positive vulnerability feature set, the non-false positive vulnerability feature set, the first processing result, and the second processing result.
2. The method of claim 1, wherein, Before inputting the initial vulnerability feature set and the category feature vector into the false positive detection model constructed in advance, the method further comprises the following steps: scanning a target code segment of a software to be tested to obtain initial vulnerability data corresponding to the target code segment, extracting at least one initial vulnerability feature of the initial vulnerability data to form an initial vulnerability feature set, and obtaining a category feature vector corresponding to the initial vulnerability feature set; constructing a first evaluation function based on the correlation between at least one initial vulnerability feature in the initial vulnerability feature set and the category feature vector, and calculating a first correlation value between the initial vulnerability feature and the category feature vector by using the first evaluation function; constructing a second evaluation function based on the correlation between different initial vulnerability features in the initial vulnerability feature set, and calculating a second correlation value between the different initial vulnerability features by using the second evaluation function; constructing a third evaluation function based on the first evaluation function and the second evaluation function, and calculating a third correlation value between an initial vulnerability feature subset and the category feature vector by using the third evaluation function; constructing an evaluation function of the false positive detection model based on the first evaluation function, the second evaluation function, and the third evaluation function.
3. The method of claim 2, wherein, The method comprises the following steps: calculating a first correlation value between different initial vulnerability features and the category feature vector based on a first evaluation function in the false positive detection model; determining whether the first correlation value is less than a preset threshold value; If the first correlation value is less than the preset threshold, the corresponding initial vulnerability feature is rejected to determine the vulnerability feature set; If the first correlation value is greater than or equal to the preset threshold, the corresponding initial vulnerability feature is retained to determine the vulnerability feature set.
4. The method of claim 2, wherein, The initial vulnerability feature set and the category feature vector are input into a pre-constructed false alarm detection model to filter the initial vulnerability feature set by using the false alarm detection model to obtain a vulnerability feature set that meets a preset condition, including: Based on a second evaluation function in the false alarm detection model, a second correlation value between different initial vulnerability features is calculated to construct a feature association set of the different initial vulnerability features according to the second correlation value; Based on the different initial vulnerability features and the feature association set, an initial clustering subset and other clustering subsets of the feature association set are determined; With the initial clustering subset as a reference, the other clustering subsets are traversed to determine a corresponding remaining feature set and an initial clustering set; Based on the remaining feature set and the initial clustering set, a final clustering set is calculated to determine the vulnerability feature set by using the final clustering set.
5. The method of claim 4, wherein, With the initial clustering subset as a reference, the other clustering subsets are traversed to determine a corresponding remaining feature set and an initial clustering set, including: determining whether the initial clustering subset contains at least one other clustering feature in the other clustering subsets; If the initial clustering subset contains at least one other clustering feature, the remaining feature set is determined based on the other clustering features that are not included; If the initial clustering subset does not contain the other clustering features, the initial clustering set is determined based on the other clustering subsets corresponding to the other clustering features.
6. The method of claim 4, wherein, The calculation of the corresponding final clustering set based on the remaining feature set and the initial clustering set includes: Based on a first evaluation function in the false alarm detection model, a first correlation value between a remaining feature and a corresponding category feature vector is calculated; The remaining features are sorted according to a first target sorting manner based on the first correlation value to obtain sorted remaining features, and a final remaining feature set is determined according to the sorted remaining features; The final remaining feature set and the initial clustering set are traversed to calculate the effective information of different final remaining features in the initial clustering subset; determining whether the effective information meets a preset information condition; If the effective information meets the preset information condition, the corresponding final remaining feature is added to the corresponding initial clustering subset to obtain a final clustering subset, and the final clustering set is obtained according to the final clustering subset.
7. The method of claim 4, wherein, The final clustering set is determined by using the final clustering set, including: Based on a first evaluation function in the false alarm detection model, a first correlation value between different clustering features in the final clustering set and corresponding category feature vectors is calculated to sort the different clustering features according to a second target sorting manner based on the first correlation value to obtain a sorted clustering set; Traverse the sorted cluster set, and determine the corresponding initial cluster feature set to calculate the effective information of different sorted cluster features in the initial cluster feature set; Based on the effective information, determine the vulnerability feature set.
8. The method of claim 1, wherein, Before inputting the vulnerability feature set into the pre-constructed retrieval enhancement generation agent, further comprising: Obtain a target vulnerability feature set; Based on the target vulnerability feature set, construct a vector database; Based on the vector database, filter to obtain a target final candidate feature vector; Based on the target final candidate feature vector, construct a large language model; Based on the vector database and the large language model, construct the retrieval enhancement generation agent.
9. The method of claim 8, wherein, The inputting the vulnerability feature set into the pre-constructed retrieval enhancement generation agent, to determine the false positive vulnerability feature set and the non-false positive vulnerability feature set in the vulnerability feature set by using the retrieval enhancement generation agent, and the first processing action corresponding to the false positive vulnerability feature set and the second processing action corresponding to the non-false positive vulnerability feature set, comprises: Based on the vector database in the retrieval enhancement generation agent, determine the index vector corresponding to the vulnerability feature set; With the index vector as the index, filter to obtain the initial candidate feature vector satisfying the preset detection condition from the vector database; Based on the initial candidate feature vector, filter to obtain the final candidate feature vector satisfying the preset similarity condition; Based on the final candidate feature vector, use the large language model in the retrieval enhancement generation agent to determine the false positive vulnerability feature set, the non-false positive vulnerability feature set, the first processing action and the second processing action.
10. An electronic device, comprising: Comprise: Memory, processor and computer program stored on the memory and executable on the processor, the processor executes the program to realize the code vulnerability processing method as claimed in any one of claims 1-9.
Citation Information
Patent Citations
Source code security analysis method based on historical optimization feature intelligent learning
CN112148602A
Source code vulnerability detection method and device
CN115510449A
Code vulnerability detection large model construction method and device and electronic equipment
CN118171291A
Vulnerability remediation code detection method and related device
WO2025086604A1