Vulnerability patch existence detection method based on deep learning

By building an equivalent patch set based on deep learning and using a two-way LSTM twin network for detection, the problem of being unable to accurately judge the patching of open source software in the existing technology is solved, and efficient and robust vulnerability patch detection is achieved.

CN116108446BActive Publication Date: 2025-05-13XIDIAN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211557968.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-12-06
Publication Date
2025-05-13
Estimated Expiration
2042-12-06

AI Technical Summary

Technical Problem

Existing vulnerability patch detection methods cannot accurately determine whether the open source software used by the enterprise has been patched, especially when the enterprise version is different from the mainline version, traditional version numbers or static signature methods are not enough to meet the needs of the agile development model.

Method used

Using a deep learning-based method, by obtaining the information of the original patch and downstream OSS warehouse, an equivalent patch set is constructed, and the equivalent patch is version rolled back, code attribute graph construction, slice generation and word vector processing is performed on the equivalent patch. The bidirectional LSTM twin network is used for characterization results analysis to determine the existence of the patch.

Benefits of technology

It realizes accurate detection of the existence of vulnerability patches without manually defining matching rules, which is more robust, can solve the problem of cross-functional vulnerability detection, and greatly reduces time overhead by reducing detection space.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116108446B_ABST
    Figure CN116108446B_ABST
Patent Text Reader

Abstract

The present invention provides a vulnerability patch existence detection method based on deep learning, by comparing the original patch with the potential patch of the downstream OSS warehouse to form an equivalent patch; then select the slice entrance to generate an equivalent slice for the equivalent patch, and then convert it into a word vector, and use it as the input of a bidirectional LSTM twin network; according to whether the output characterization result is similar to the input, it is determined to adjust the network parameters to complete the training. For the OSS project to be detected, its detection space is reduced, and then the common features of the vulnerability are selected as the slice entrance to generate a slice, and it is input into the bidirectional LSTM twin network after training, and the characterization results of the two inputs are obtained, and according to the similarity between the characterization results and the two inputs, it is confirmed whether there is a vulnerability patch. The present invention uses a bidirectional LSTM twin network to detect vulnerabilities, and it is more robust without manually defining matching rules; it can solve the problem of being unable to detect cross-function vulnerabilities, and the time overhead of detection can be greatly reduced by reducing the detection space.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of network security, and in particular relates to a vulnerability patch existence detection method based on deep learning. Background Art

[0002] More and more open source software (OSS) allows enterprise developers to reuse simple functions from reliable OSS projects. However, at the same time, the spread of vulnerabilities brought by the reuse of third-party OSS may threaten the security of the entire system. In principle, keeping the reused OSS code always up to date can prevent the impact of vulnerabilities. However, OSS updates very quickly, and the version used by the enterprise is quite different from the mainline version, which is mainly reflected in two aspects: 1) the enterprise has modified the existing OSS, and the context of the vulnerability has changed the original meaning; 2) the version used by the enterprise is a branch version and has been separated from the mainline. The current vulnerability scanning software is based on version numbers or static signatures, which is more suitable for the era of private software supply and cannot meet the current agile development model. Therefore, when the mainline of the reused third-party OSS project exposes a vulnerability, a method that can accurately determine the existence of the vulnerability patch is needed, which can determine whether the OSS code currently used by the enterprise has been patched based on the given CVE vulnerability information (not simply by version number and code comparison).

[0003] In order to improve the efficiency of software development, Internet companies usually adopt issue tracking and source code control management systems such as GitHub, JIRA or Bugzilla to manage the entire project. According to statistics, as of April 2017, GitHub reported nearly 20 million users and 57 million repositories. According to Atlassian, more than 75,000 companies use JIRA. These tools are very popular in open source projects and are essential for modern software development. Developers deal with problems reported in these systems and then submit corresponding code changes to GitHub (or other source code hosting platforms such as SVN or BitBucket). Bugs and new features are often merged into a central repository and then automatically built, tested and prepared for release to the production environment as part of the DevOps continuous integration (CI) and continuous delivery (CD) practices. There is no doubt that pipeline production helps improve the productivity of developers and enables them to solve problems faster. However, due to the focus on rapid release cycles and lack of manpower and expertise, a large number of security-related issues and errors in software are quietly patched without public disclosure.

[0004] In today's world of agile software development, developers increasingly rely on and extend free open source libraries to get work done quickly. Many people don't even know which open source library they are using, let alone the hidden defects that are often silently patched in the software repository. Therefore, even if the bug appears in the source code from the commit message or error report, it is still possible that users of these software are unaware of the problem and continue to use the old version. By not being aware of these hidden defects or vulnerabilities, they put their products at risk of being attacked by hackers; secondly, when developing software, manufacturers directly use a functional module in the open source code or directly make minor or even no modifications to the open source code before using it in the software they want to develop. When using these reused codes, manufacturers do not fix the published vulnerabilities contained in the code, and do not release patches in a timely manner after the vulnerability is disclosed, so that the software distribution developed by the manufacturer still has the problem of published vulnerabilities. For example, the OpenSSL Heartbleed vulnerability (CVE-2014-0160) affects the OpenSSL software, but as long as this software code is reused, it will be affected, so that in the end this vulnerability has a very wide range of impact. Therefore, locating such patches and knowing the intent of each patch are important as they can be used to help developers decide on component updates.

[0005] Current vulnerability scanning software is based on version numbers or static signatures, which is more suitable for the era of private software supply. Now most software is built based on OSS, and with the promotion and use of DevSecOps within various companies, security issues are more easily spread at the source code level. However, OSS updates very quickly, and the version used by enterprises is quite different from the mainline version, which is mainly reflected in two aspects: 1) The enterprise has modified the existing OSS, and the context of the vulnerability has changed its original meaning; 2) The version used by the enterprise is a branch version and has deviated from the mainline.

[0006] Therefore, when the mainline exposes a vulnerability, it is impossible to accurately determine whether the OSS currently used by the enterprise has been patched (it cannot be simply determined by version number or code comparison). The application of patches is the main means to combat software vulnerabilities. Using patch existence detection to improve the ability to accurately test whether there are security patches in software releases is crucial in maintaining system security. Therefore, a method is needed to accurately determine the existence of patches, give accurate patch determinations, and even automatically generate corresponding patches.

[0007] The existing vulnerability patch detection methods have the following problems:

[0008] (1) Rule-based vulnerability detection method: This method predefines some rules to statically detect program vulnerabilities. This method detects vulnerabilities very quickly, but it requires defining new rules to detect new vulnerabilities, and its accuracy is low.

[0009] (2) Vulnerability detection method based on data mining: This method uses data mining algorithms to automatically mine potential rules from source code, and then uses these rules to detect vulnerabilities. The disadvantage of this method is that it has a high false positive rate, which means that security personnel need to rely on their own knowledge to determine whether the detected vulnerability is a vulnerability, which may be more difficult than discovering a vulnerability.

[0010] (3) Vulnerability detection method based on graph neural network: To use graph neural network for vulnerability detection tasks, the vulnerability code must first be represented as a graph. Based on the abstract syntax tree of the code, the control flow graph and the edges of the control flow graph are added to the abstract syntax tree to form a graph representing the source code. Backward edges are added, that is, by transposing the adjacency matrix, to increase the expressiveness of the code graph, and then the GGNN model is used for vulnerability detection tasks. However, this method will introduce a lot of intermediate process overhead. Converting the source code into a graph and manually adding edge types is a time-consuming and expensive task;

[0011] (4) Vulnerability detection methods based on machine learning / deep learning of a single function: The effectiveness of vulnerability detection models based on machine learning depends on whether the features selected by the model are representative. Models based on deep learning mainly learn features from training data sets. Compared with traditional machine learning models, deep learning models often only consider the features of a single function, while vulnerabilities may be caused by a combination of multiple functions. Therefore, methods based on deep learning will have a relatively high false alarm rate. Summary of the invention

[0012] In order to solve the above problems existing in the prior art, the present invention provides a vulnerability patch existence detection method based on deep learning. The technical problem to be solved by the present invention is achieved through the following technical solutions:

[0013] The present invention provides a method for detecting the existence of vulnerability patches based on deep learning, including:

[0014] Step 1: Obtain the vulnerability disclosure entry CVE information from the common database, and search for the OSS repository of the original patch and the corresponding downstream OSS repository based on the CVE information. By comparing the original patch with the potential patch in the downstream OSS repository, the original patch and the potential patch with consistent patch information are determined as equivalent patch pairs, and all equivalent patch pairs are combined into an equivalent patch set;

[0015] Step 2: for each equivalent patch pair in the equivalent patch set, perform version rollback for each patch in the equivalent patch pair, obtain a source code file, and build a code attribute graph according to the source code file, find a statement node related to the core code line in the code attribute graph, and generate a slice with the statement node as the slice entry. Each equivalent patch pair can generate an equivalent slice pair, and the equivalent slice pairs form a slice set;

[0016] Step 3: normalize all slice codes in the slice set to obtain a processed slice set;

[0017] Step 4: According to different situations of whether the slice pairs in the slice set are from the same equivalent patch, data processing is performed on the equivalent slices in a positive and negative sample manner, and the processed slice pairs are converted into word vector pairs;

[0018] Step 5: Input the word vector pair into the bidirectional LSTM twin network to obtain the representation result of the slice pair, and determine whether they are equivalent according to the similarity between the representation results of each slice pair, so as to train the bidirectional LSTM twin network and obtain a trained bidirectional LSTM twin network;

[0019] Step 6: Obtain an OSS project to be detected that includes multiple source code files to be detected, and narrow the detection space of the OSS project to be detected;

[0020] Step 7: Build a code attribute graph based on each source code file in the OSS project after the detection space is reduced, select the slice entry of the code attribute graph to generate a slice to be detected that carries the core code of the source code to be detected, and repeat steps 3 to 4 for the slice to be detected to convert it into a word vector to be detected;

[0021] Step 8: Based on the trained bidirectional LSTM twin network, the word vector to be detected is identified to obtain the representation result of the detection word vector, and according to the similarity between the representation results and the similarity between the input objects of the trained bidirectional LSTM twin network, it is confirmed whether the source code file to be detected has a patch to fix the vulnerability.

[0022] Beneficial effects of the present invention:

[0023] A vulnerability patch existence detection method based on deep learning provided by the present invention, by obtaining the OSS warehouse and downstream OSS warehouse of the original patch, then comparing the original patch with the potential patch to form an equivalent patch set; then selecting the slice entrance to generate an equivalent slice for the equivalent patch, normalizing the equivalent slice, and then converting it into a word vector, inputting the word vector into a bidirectional LSTM twin network, and determining whether to adjust the parameters of the bidirectional LSTM twin network according to whether the characterization result of the output is similar to the input, and realizing the purpose of training the bidirectional LSTM twin network. After the training is completed, for the OSS project to be detected, its detection space is reduced to improve the detection efficiency, and then the common features of the vulnerability are selected as the slice entrance to generate a slice, and input into the bidirectional LSTM twin network after the training is completed, and the characterization results of the two inputs are obtained, and according to the similarity between the characterization results and the similarity between the two inputs, it is confirmed whether the OSS project to be detected has a vulnerability patch. The present invention uses a bidirectional LSTM twin network to detect vulnerabilities, without manually defining matching rules, and is more robust; the problem that the existing scheme cannot detect cross-function vulnerabilities can be solved; the time overhead of detection can be greatly reduced by reducing the detection space.

[0024] The present invention will be further described in detail below with reference to the accompanying drawings and embodiments. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] Figure 1 This is the MVP model architecture diagram of the existing technology;

[0026] Figure 2 It is a flowchart of a vulnerability patch existence detection method based on deep learning of the present invention;

[0027] Figure 3 The present invention is to establish a code slicing flow chart based on a static analysis tool;

[0028] Figure 4 It is an example diagram of the control flow graph (CFG) of the present invention. DETAILED DESCRIPTION

[0029] The present invention is further described in detail below with reference to specific embodiments, but the embodiments of the present invention are not limited thereto.

[0030] Before introducing the present invention, the closest prior art and technical concept of the present invention are first introduced.

[0031] As the closest prior art to the present invention, reference Figure 1 shown. Figure 1 This is the MVP model architecture diagram, a repeated vulnerability detection framework with function as the detection granularity, which mainly includes the following three steps: (the following uses sig instead of signature)

[0032] Generate sig of target function: Input the target system to be tested and generate sig for each target function in the system;

[0033] Generate vulnerability code and patch code sig: Input the security patch program, generate vulnerability and patch sig, reflect the vulnerability from the perspective of vulnerability generation and vulnerability repair, and obtain a set of to-be-matched vulnerabilities and patch sigs;

[0034] Match the sig of each function in the target system with the vulnerability and patch sig: If a vulnerability sig matching the target function sig is found in the to-be-matched set, but there is no matching patch sig, the target function is considered to have a duplicate vulnerability.

[0035] First, in the first step, MVP defines the function signature. Given a C / C++ function f, the signature of f is defined as a tuple (fsyn, fsem), where fsyn is the set of hash values ​​of all statements in the function; fsem is a set consisting of a series of 3-tuples (h1, h2, type), where h1 and h2 represent the hash values ​​of any two statements, and type∈{data, control} indicates that the statement with hash value h1 and the statement with hash value h2 have data dependency or control dependency. Among them, fsyn captures the statements of the target function as a syntactic signature; fsem captures the data dependency and control dependency between the statements of the target function as a semantic signature. Both provide supplementary information of a function to help improve matching accuracy.

[0036] Subsequently, in the second step, MVP generates sigs of the vulnerability code and patch code. Given a pair of (fv, pv) and PatchPv, the following describes how to generate sigs to capture key statements related to the vulnerability, rather than including all statements in fv and pv, so as to obtain a small and accurate sig for effective matching. MVP identifies the changed files by parsing the header file (diff file) of the security patch, and finds the deleted and added statements and their line numbers by parsing the diff file. At the same time, the start and end addresses of all functions are found. If a statement includes one or more deleted (added) lines of code, the statement is considered to have been deleted (added), and the relationship between the statement line number and the function line number is compared to determine which functions have been changed. In this way, the set of added statements S can be obtained. add , the set of deleted statements S del , the vulnerable function statement S vul , patch function statement S pat . Slicing technology can be used to extract relevant statements and exclude irrelevant statements. MVP performs forward and backward slicing on PDG, using S add and S delAs the criterion for slicing. The traditional program slicing method has certain problems. If the slicing is required to be directly related to the slicing criteria, the slicing may not contain the vulnerability statement. Once it is allowed to be indirectly related to the slicing criteria, there will be too much noise in the slicing. Therefore, a new slicing criterion is proposed in MVP.

[0037] Based on the above, MVP then calculates the vulnerability sig and patch sig at the grammatical and semantic levels, namely (Vsyn, Vsem) and (Psyn, Psem). The vulnerability sig is related to the generation of the vulnerability, and the patch sig is related to the patching of the vulnerability. The first step is to calculate Vsyn. Contains the semantic information of the vulnerability, but the modification process of the vulnerability does not involve deleting statements, only adding statements, so it is necessary to fill in the original vulnerable function statements that depend on the added statement data or control, and then obtain Vsem from Vsyn according to the PDG graph, Then, and S vul The statements that only exist in the patch function are found by doing the difference set, which is Psyn. Then, the triple set F of the vulnerability function is calculated based on the PDG graph. The triple set Psem of the statements that only exist in the patch function (newly added statements) can be obtained by doing the difference set between T and F. Then, Sdel, (Vsyn, Vsem) and (Psyn, Psem) are normalized and hash values ​​are calculated as described above to obtain the vulnerability sig and patch sig. After processing all the vulnerability patch code pairs, the set to be tested is obtained.

[0038] The calculation method is as follows:

[0039]

[0040] V sem ={(s1,s2,type)|s1,s2∈V syn}

[0041]

[0042]

[0043]

[0044]

[0045] Finally, in the third step, MVP matches each function signature in the target system with the vulnerability and patch signatures. In the first and second steps, the sig (fsyn, fsem) of each target function in the target system, as well as the deleted statement Sdel, the vulnerability sig (Vsyn, Vsem), and the patch sig (Psyn, Psem) have been obtained. The target function is judged whether it has a vulnerability (matches the vulnerability sig but does not match the patch sig) according to the following principles. There are 5 rules:

[0046] The target function must contain all the removed statements;

[0047] The signature of the target function matches the vulnerability signature at the syntactic level (the intersection of Vsyn and fsyn is greater than a certain threshold);

[0048] The signature of the target function does not match the patch signature at the syntactic level (the intersection of Psyn and fsyn is less than a certain threshold);

[0049] The signature of the target function matches the vulnerability signature at the semantic level (the intersection of Vsem and fsem is greater than a certain threshold);

[0050] The signature of the target function does not match the patch signature at the semantic level (the intersection of Psem and fsem is less than a certain threshold);

[0051] However, this technology is a rule-based vulnerability detection method, which has a high false alarm rate and cannot detect frequent changes in variable types and function names during OSS migration. Therefore, the present invention proposes a cross-function detection technology with improved accuracy and stronger robustness. The details of the solution involved in the present invention are introduced in detail below.

[0052] Embodiment 1

[0053] refer to Figure 2 The present invention provides a method for detecting the existence of vulnerability patches based on deep learning, including:

[0054] Step 1: Obtain the vulnerability disclosure entry CVE information from the common database, and search for the OSS repository of the original patch and the corresponding downstream OSS repository based on the CVE information. By comparing the original patch with the potential patch in the downstream OSS repository, the original patch and the potential patch with consistent patch information are determined as equivalent patch pairs, and all equivalent patch pairs are combined into an equivalent patch set;

[0055] Step 2: for each equivalent patch pair in the equivalent patch set, perform version rollback for each patch in the equivalent patch pair, obtain a source code file, and build a code attribute graph according to the source code file, find a statement node related to the core code line in the code attribute graph, and generate a slice with the statement node as the slice entry. Each equivalent patch pair can generate an equivalent slice pair, and the equivalent slice pairs form a slice set;

[0056] Step 3: normalize all slice codes in the slice set to obtain a processed slice set;

[0057] Step 4: According to different situations of whether the slice pairs in the slice set are from the same equivalent patch, data processing is performed on the equivalent slices in a positive and negative sample manner, and the processed slice pairs are converted into word vector pairs;

[0058] Step 5: Input the word vector pair into the bidirectional LSTM twin network to obtain the representation result of the slice pair, and determine whether they are equivalent according to the similarity between the representation results of each slice pair, so as to train the bidirectional LSTM twin network and obtain a trained bidirectional LSTM twin network;

[0059] Step 6: Obtain an OSS project to be detected that includes multiple source code files to be detected, and narrow the detection space of the OSS project to be detected;

[0060] Step 7: Build a code attribute graph based on each source code file in the OSS project after the detection space is reduced, select the slice entry of the code attribute graph to generate a slice to be detected that carries the core code of the source code to be detected, and repeat steps 3 to 4 for the slice to be detected to convert it into a word vector to be detected;

[0061] Step 8: Based on the trained bidirectional LSTM twin network, the word vector to be detected is identified to obtain the characterization result of the detection vector, and according to the similarity between the characterization results of the detection vector and the similarity between the input objects of the trained bidirectional LSTM twin network, it is confirmed whether the source code file to be detected has a patch to fix the vulnerability.

[0062] Embodiment 2

[0063] As an optional embodiment of the present invention, step 1 is an equivalent patch set preparation stage, which specifically includes:

[0064] Step 11: Obtain common vulnerability entry disclosure CVE information from the common database;

[0065] This step collects Common Vulnerabilities and Exposures (CVE) entry information from the National Vulnerabilities and Exposures Database (NVD) of the United States. The vulnerability information in this database is strictly reviewed by authoritative personnel before being made public, and has a high degree of credibility. The collection of CVE information is mainly based on the open source tool cve-search, which imports public CVEs into the MongoDB local database for faster search and processing of CVEs. The main purpose is to avoid direct and public searches of the public CVE database, which leads to the problem of Internet restrictions on sensitive queries.

[0066] Step 12: Get the URL information of the patch from the relevant link in the CVE information description;

[0067] It is worth noting that: in addition to the necessary description, there are many reference links under the CVE entry. Select those links pointing to github submissions, which are used as samples of submissions related to bug fixes, and then crawl the logmessage text of the corresponding web page and the commit number that uniquely identifies each submission. Each CVE contains a series of information about the disclosed vulnerability. Analyzing this information can further query the specific content of the patch. Based on the URL information, the OSS repository, modification location, and commit number of the original patch corresponding to the CVE can be located, as shown in the following code.

[0068]

[0069] Step 13: According to the URL information, locate the OSS repository where the original patch is located, the modification location and the commit number;

[0070] Step 14: According to the license content recorded in the OSS warehouse where the original patch is located, obtain the warehouse list of the OSS warehouse of the original patch and the downstream target OSS warehouse;

[0071] The downstream target OSS repository includes the commit number of the patch and the potential patches after the patch is applied;

[0072] After the OSS repository where the original patch is located has been identified in the previous step, you must first determine whether the OSS repository component is used by other OSS projects. Usually, the OSS repository's license will disclose which OSS tools have been ported to this project. Based on this, you can use the git tool to download each project to obtain the OSS repository of the original patch and the list of downstream OSS repositories.

[0073] Step 15: For each commit operation in the downstream target OSS repository, perform string matching on the subject of the commit operation and the message of the original patch; if either the subject or the message matches successfully, it is determined that the potential patch corresponding to the commit operation is an equivalent patch pair with the original patch; if neither the subject nor the message matches successfully, check whether the original patch and the specific patch content of the commit operation match; if they match, it is determined that the potential patch corresponding to the commit operation is an equivalent patch pair with the original patch.

[0074] It is worth noting that various information about the original patch is combined to determine its existence in other OSS repositories. First, for each commit submitted by the downstream target OSS repository, try to perform a string match between the subject submitted and the subject of the original patch. This is because if the developer transplanted the patch from the original project OSS, the project will usually retain the original subject. If there are multiple hits, use the submitted message to identify the true match; if no results are found, it means that when the same patch is applied downstream, the submitted subject and message are not retained. Search the submission history of the corresponding patch file and try to match the complete source code level changes (including added / deleted lines and context lines) with the changes in the original patch. Once the target is matched, it can be considered that the commit number in the target OSS repository is equivalent to the commit number of the original patch, and the two patches are labeled and marked as 1. The following code shows an equivalent patch pair, both of which are equivalent repair patches for CVE-2019-19531.

[0075] Commit-ID is 1041f5092179

[0076]

[0077] Commit-ID is 28318f535030

[0078]

[0079] In addition, the present invention can also be improved as follows based on the above:

[0080] 1) File path / name changes: If no commit is found that changes the patch file, the present invention expands the search area to files with the same name but in different directories (sometimes the downstream kernel decides to rearrange some source files). If we find any commit that renamed the patch file at some point, we also track the evolution of the renamed file.

[0081] 2) Function name changes: Similar to file names, function names may also change over time. The present invention can track their evolution by developing a small script to check related submissions.

[0082] Embodiment 3

[0083] As an optional embodiment of the present invention, step 2 is a static analysis to establish a slicing stage. The main idea of ​​slicing is to decompose the program by finding the relationship between the internal codes of the program, and then analyze the decomposed core code. The slicing technology used in the present invention is well known to relevant practitioners. The present invention only makes adaptive changes in the selection of slicing entry. Figure 3 As shown in the figure, the static analysis slicing stage specifically includes:

[0084] Step 21: for each equivalent patch pair in the equivalent patch set, roll back each patch in the equivalent patch pair to the corresponding project version according to the patch commit number, and obtain the path information in the patch;

[0085] Step 22: extracting the source code file of the patch based on the path information;

[0086] In step 1, pairs of equivalent patches and their corresponding OSS code repositories are collected. According to the commit number of the patch, the corresponding project version is rolled back. Then, the source code file is extracted based on the path information in the patch. The source code file is the code version after the patch is applied.

[0087] Step 23: Input the source code file into a static analysis tool to obtain a control flow graph (CFG), a program dependency graph (PDG), and a function call graph (CallG) of the source code file;

[0088] Among them, the control flow graph (CFG) describes the possible execution path of each line of code in the program, where the nodes represent the code lines and the edges represent the execution path of the code; the program dependency graph (PDG) describes the mutual dependence and mutual influence between code instructions, including control dependency and data dependency; the function call graph (CallG) describes the relationship between functions in the source file, including the calling and called relationship.

[0089] In this step, the extracted source code file is input into the static analysis tool to obtain its control flow graph (CFG), program dependency graph (PDG) and function call graph (CallG). Finally, the three directed graphs are combined to obtain the code attribute graph. Figure 4 As shown in the figure, it describes the possible execution path of each line of code in the program, where the nodes represent the code lines and the edges represent the execution path of the code; PDG mainly describes the interdependence and mutual influence between code instructions, including control dependency and data dependency; CallG mainly describes the relationship between functions in the source file, including the calling and called relationship.

[0090] Step 24: combining the control flow graph (CFG), the program dependency graph (PDG) and the function call graph (CallG) to obtain a code attribute graph;

[0091] Step 25: Using the core code line related to the patch as the slice entry, searching the control dependency node, data dependency node and possible function call relationship related to the core code line in the code attribute graph;

[0092] Step 26: forming equivalent slice pairs with the code lines corresponding to the control dependency nodes, data dependency nodes and possible function call relationships;

[0093] The slice entry is the selected core code line. Through this line of code, the control dependency nodes, data dependency nodes and possible function call relationships related to it are found, and finally the code set of the slice can be determined. The slice entry specified in the present invention is the newly added line in the patch and the three lines before and after the corresponding line in the source code after the patch is applied, as shown in the following slice entry of the source code file after the patch is applied:

[0094]

[0095] The slice entry is marked because the deleted lines of the patch do not exist in the source code after the patch is applied, so the slice cannot be extracted in the patched file.

[0096] As shown in the following code,

[0097]

[0098] This code represents a simple patch file, which consists of two parts: the patch header and the patch block. The details are as follows: The patch header is two lines starting with --- / +++ (corresponding to line 1 and line 2), which are used to indicate the files to be patched. The file starting with --- indicates an old file, and the file starting with +++ indicates a new file. A patch file may contain many sections starting with --- / +++, each section is used to apply a patch. So a patch file can contain multiple patches.

[0099] The patch block is the place to be modified in the patch (corresponding to lines 3 to 10 of the code). It usually starts and ends with a part of the code lines that do not need to be modified. These lines are only used to mark the location to be modified. It usually starts with @@ and ends at the beginning of another block or a new patch header. The number after '-' represents the line number of the block in the old file and the total number of lines in the block before the modification; the number after '+' represents the line number of the block in the new file and the total number of lines in the block after the modification. The block will be indented by one column, and this column is used to indicate whether the line is added or to be deleted.

[0100] The + sign indicates that the line is a new line, and the content of the line will be added to the source file.

[0101] The - sign indicates that the line is to be deleted and the content of the line will be deleted in the source file. No plus sign or minus sign indicates that it is only referenced and does not need to be modified.

[0102] Step 27: Group all equivalent slice pairs into slice sets.

[0103] Embodiment 4

[0104] Step 3 is the normalization phase of the slice code. When developers reuse OSS code or projects, they may rename or change the definition of parameters / variables. To mask the detection errors caused by inconsistent naming with consistent semantics, the slices obtained in the previous step need to be normalized, including identifying formal parameters, local variables, strings, and function names from the slice code, and replacing them with the normalized symbols VAR, PARAM, DTYPE, and FUN in the order in which they appear.

[0105] As an optional embodiment of the present invention, the slice code normalization stage specifically includes:

[0106] Step 31: for each slice pair in the slice set, use the symbol VAR to replace all local variable definitions appearing in the slice, and add the order information of the variable appearance to obtain a slice set after the variables are normalized;

[0107] Step 32: For each slice pair in the slice set after variable normalization, use the symbol PARAM to replace all formal parameters appearing in the function, and add the order information of the formal parameters to obtain the slice code set after parameter normalization;

[0108] Step 33: For each slice code in the slice code set after parameter normalization, use the symbol DTYPE to replace all data types, and add the order information of the data type, so as to obtain the slice code set after data type normalization;

[0109] Step 34: For each slice code in the slice code set after data type normalization, use the symbol FUN to replace all the function names that appear, and add the order information of the function appearance to obtain the processed slice code set.

[0110] The process of the present invention is specifically shown in the following table:

[0111]

[0112] The detailed steps recorded in the above table are explained as follows:

[0113] (1) Variable normalization: Based on the original slice, the symbol "VAR" is used to replace all local variable definitions appearing in the slice, and the order information of the variable appearance is added. For example, in the embodiment of the invention, char*lensdir is replaced with char VAR_1, because the variable lensdir is the first variable appearing in the slice. Note: The same variable will only be replaced with the same symbol.

[0114] (2) Parameter normalization: Based on variable normalization, use the symbol "PARAM" to replace all formal parameters that appear in the function, and add the order information of the formal parameters. For example, as shown in 2 in the above table, replace the first parameter *Tests in the function run_tests() with PARAM_1. Note: The same parameter will only be replaced with the same symbol.

[0115] (3) Data type normalization: Based on parameter normalization, all data types are replaced with the symbol "DTYPE" and the order in which the data type appears is added. For example, as shown in Table 3 above, the structure struct Test is replaced with struct DTYPE_1. Note: The same data type will only be replaced with the same symbol, and modifiers (such as static) will not be replaced. This is because for some types of vulnerabilities, the signedness of the variables in the source code will affect its data flow.

[0116] (4) Function normalization: Based on data type normalization, use the symbol "FUN" to replace all the function names that appear, and add the order information of the function appearance. For example, as shown in Table 4 above, the function run_tests is replaced with FUN_1. Note: The same function will only be replaced with the same symbol.

[0117] Embodiment 5

[0118] As an optional embodiment of the present invention, step 4 is a data set preprocessing stage, which specifically includes:

[0119] Step 41: Match each equivalent slice pair with other slice pairs to construct inequivalent slice pairs;

[0120] Step 42: taking the equivalent slice pairs as positive samples and the non-equivalent slice pairs as negative samples, setting the labels of the positive samples to 1 and the negative samples to 0;

[0121] Step 43: Use the pre-trained word2vec model to convert positive samples and negative samples into word vectors of the same dimension.

[0122] The input of the bidirectional LTSM twin network proposed in this invention is a pair of word vectors of a pair of slices. When the two slices come from the same pair of equivalent patches, they are labeled as 1, otherwise they are labeled as 0. The preprocessing operation is explained in detail as follows:

[0123] (1) Labeling: If n pairs of equivalent patches are obtained during the preparation phase of the equivalent patch set,

[0124] {(p 11 ,p 12 ),(p 21 ,p 22 ),...,(p n1 ,p n2 )},

[0125] The equivalent slice pairs generated during the static analysis slice creation phase are:

[0126] {(slice 11 ,slice 12 ),(slice 21 ,slice 22 ),...,(slice n1 ,slice n2 )},

[0127] Each pair of equivalent slices is matched with other pairs of slices. For example, slice11 in (slice11, slice12) can be matched with slice21 in (slice21, slice22) to construct an inequivalent slice pair (slice11, slice21). Thus, Pairs of slices. Among them, n pairs of equivalent slices are used as positive samples, and the label is 1; 2n(n-1) pairs of non-equivalent slices are used as negative samples, and the label is 0;

[0128] (2) Word vector representation: The processed pre-processed slice data is converted into word vectors of equal dimension through the pre-trained word2vec model, which is represented as vi here. The subscript i is used to distinguish different word vectors. The corresponding sentence containing n words is mapped to a group of multiple vectors (vec1, vec2, vec3, ..., vecn). Therefore, the slice pair (slice11, slice12) can be represented as (slice11, slice12)(vec1, vec2, vec3, ..., vecn).

[0129] (vec1, vec2, vec3, ..., vecm, 0, 0) Since the number of words in the two slices is inconsistent, the dimensions of the mapped word vectors are inconsistent, so a padding operation is required. As shown in the above embodiment, m=n-2, so 0s need to be filled in two dimensions.

[0130] Embodiment 6

[0131] As an optional embodiment of the present invention, step 5 is a bidirectional LSTM twin network training phase, which specifically includes:

[0132] Step 51: For the word vector pair of the negative sample and the word vector pair of the positive sample, one word vector in the word vector pair is input into the left sub-network of the bidirectional LSTM twin network, and the other word vector is input into the right sub-network to obtain the output vectors M1 and M2 corresponding to the left sub-network and the right sub-network;

[0133] Step 52: using the output vectors M1 and M2 as the representation results of the slice pair respectively, and using the similarity function to calculate the similarity between the representation results of each slice pair;

[0134] Step 53: judging whether the input slice pair is consistent with the characterization result according to the similarity, and adjusting the parameters of the bidirectional LSTM twin network if they are inconsistent;

[0135] Step 54: Repeat steps 51 to 53 until all word vector pairs are traversed to obtain a trained bidirectional LSTM twin network.

[0136] The bidirectional LSTM twin network input, output and iterative training process of the present invention are introduced as follows:

[0137] (1) Network input: In the dataset preprocessing stage, we have obtained word vectors for n pairs of equivalent slices, with labels 1; and word vectors for 2n(n-1) pairs of non-equivalent slices, with labels 0. Each pair of slices is input into the twin network separately, with equivalent patch pairs (slice 11 , slice 12 ) as an example, slice 11The corresponding word vector (vec1, vec2, vec3, ..., vec n ) is input into the left sub-network of the bidirectional LSTM twin network, represented as Input1; the right sub-network of the bidirectional LSTM twin network inputs slice 12 The corresponding word vectors (vec1, vec2, vec3, ..., vec m , 0, 0), represented as Input2, and then the final output vectors M1 and M2 of the bidirectional LSTM twin network are taken as the characterization results of each slice.

[0138] (2) Result classification: The present invention introduces a similarity function to calculate the difference between M1 and M2, namely Similarity(M1, M2), so as to determine whether the input slices need to be merged into the same cluster. There are many similarity functions to choose from here, and the common ones are cosine similarity function, Manhattan distance function, geometric distance function, etc. The present invention selects cosine similarity as the similarity function for calculation. The cosine similarity is expressed as follows: given two vectors, A and B, their cosine similarity θ is given by the dot product and the vector length, as shown below:

[0139]

[0140] Among them, A i and B i Here represent the components of vectors A and B respectively.

[0141] The left and right subnetworks of the twin LSTM network share parameters, and the network can be dynamically adjusted according to the length of the vector. LSTM_1 and LSTM_2 represent the structures of the left and right subnetworks respectively. The input of the left subnetwork is a vector of patch function slices, and the input of the right subnetwork is another vector of patch function slices. The two vectors are input into the parameter-sharing twin LSTM network respectively, and the output of the left subnetwork of the last layer is represented as M1, and the output of the right subnetwork is represented as M2. The similarity between the outputs of the two networks is calculated to determine whether the two input log statements can be merged into the same cluster.

[0142] Embodiment 7

[0143] As an optional embodiment of the present invention, step 6 is a stage of reducing the detection space in the real OSS project detection stage, the purpose of which is to reduce the space of the real OSS project and improve the detection efficiency. The stage of reducing the detection space specifically includes:

[0144] Step 61: Preliminarily screen the source code files to be detected based on the subject and log information in the historical submission of the OSS project to be detected, so as to merge the source code files to be detected that have the same open source OSS components;

[0145] Step 62: Screening the source code files to be detected according to the patch path information of the disclosed CVE to merge the source code files to be detected that have the same path;

[0146] Step 63: Screen the source code files to be detected according to the function information of the patch to merge the source code files to be detected that have the same function information, so as to reduce the detection space of the OSS project to be detected.

[0147] In the previous stage, the network training has been completed, and the matching purpose of two slices with high similarity can be achieved. Therefore, the present invention can further complete the existence detection of OSS patches for disclosed vulnerabilities. The input of the detection stage of the present invention includes: the patch Patch1 and source code file source1 of the disclosed CVE, and the OSS project to be detected. The detailed steps are as follows:

[0148] (1) Reducing the detection space: Considering the time cost, it is not realistic to perform a full match operation on all source code files in the OSS project to be detected. Therefore, the present invention proposes a method for reducing the detection space, which mainly includes:

[0149] Perform a preliminary screening based on the subject and log information in the historical submissions of the OSS project to be tested. This is because developers usually announce which versions of open source OSS the project has ported in the version description. This information helps narrow the space to be tested.

[0150] Filter based on the path information of patch Patch1. The effective range of the patch is the path declared in the patch, as shown in the patch example code. The application path of the patch is ". / 0aa6fd109de6_patched uD_cx231xx-cards.c", so it can be matched based on the path of the patch. If the same path and file name exist in the project, the detection space can be greatly reduced;

[0151] Filter according to the function information of the patch. The essence of the patch taking effect is to modify the corresponding function content in the source code, as shown in the code of the patch example. The patch modifies the content of the function static int cx231xx_usb_probe(). If the same function name exists in the project, the purpose of accurate detection can be achieved.

[0152] Embodiment 8

[0153] As an optional embodiment of the present invention, step 7 is the stage of generating source code file slices without patches in the real OSS project detection stage, which specifically includes:

[0154] Step 71: construct a code property graph according to each source code file in the OSS project after the detection space is reduced;

[0155] Step 72: setting the slice entry of the source code file to be detected as a control flow node, an array pointer node and a system call function node, and generating a slice to be detected of the source code file to be detected by using the slice entry;

[0156] Step 73: Repeat steps 3 to 4 for the slice to be detected to convert it into a word vector to be detected.

[0157] After narrowing the detection space, it can be determined that there are multiple source code files in the OSS project to be detected that need to be sliced ​​and then matched.

[0158] (2) Slicing of source code files without patches: In the source code file obtained in the previous step, first perform the static analysis of the present invention to obtain the code property graph (CPG) operation (in the second section of the second stage of the present invention). At this time, since there is no patch in the file to be analyzed, the slice entry needs to be selected. The present invention sets the slice entry of the source code to be detected as a control flow node and an array pointer node, and generates a slice with the detection source code with the slice entry. Thereafter, the normalization operation of the third stage and the preprocessing operation of the fourth stage are performed to obtain the slice word vector corresponding to the code to be detected and the slice word vector that discloses the CVE patch.

[0159] Embodiment 9

[0160] As an optional embodiment of the present invention, step 8 is a twin network matching and false alarm detection stage without patches in the real OSS project detection stage, which specifically includes:

[0161] Step 81: input the word vector to be detected into a sub-network of the trained bidirectional LSTM twin network, input the slice word vector of the patch in the existing CVE information into another sub-network, and output the respective representation results based on the trained bidirectional LSTM twin network;

[0162] (3) Twin network matching

[0163] First, the slice word vector of the code to be detected is input into the left sub-network of the twin LSTM network, and the other network of the twin LSTM network sequentially inputs the slice word vectors of the existing CVE patches to achieve the purpose of comparing the slice of the code to be detected with the semantic pattern of each existing patch slice. Finally, the existence status of the patch in the target function is uniquely determined by the similarity function value.

[0164] Step 82: Calculate the similarity between the characterization results in step 81 and the similarity between the word vectors of the two sub-network inputs, and confirm whether the OSS project to be detected has a patch to fix the vulnerability based on whether the two similarities exceed their respective thresholds.

[0165] (4) False alarm check

[0166] After the previous step, the similarity of the two slices can be obtained. However, as a means of extracting the core lines of the code, slices have the problem of losing some information. Therefore, even if the two slices are similar, it cannot guarantee that the source code is similar. Therefore, it is necessary to calculate the similarity of the two source codes again. If it is greater than a certain threshold, it can be considered that the patch exists.

[0167] A vulnerability patch existence detection method based on deep learning provided by the present invention, by obtaining the OSS warehouse and upstream OSS warehouse of the original patch, then comparing the original patch with the potential patch to form an equivalent patch set; then selecting the slice entrance to generate an equivalent slice for the equivalent patch, normalizing the equivalent slice, and then converting it into a word vector, inputting the word vector into a bidirectional LSTM twin network, and determining whether to adjust the parameters of the bidirectional LSTM twin network according to whether the characterization result of the output is similar to the input, and realizing the purpose of training the bidirectional LSTM twin network. After the training is completed, for the OSS project to be detected, its detection space is reduced to improve the detection efficiency, and then the common features of the vulnerability are selected as the slice entrance to generate a slice, and input into the bidirectional LSTM twin network after the training is completed, and the characterization results of the two inputs are obtained, and according to the similarity between the characterization results and the similarity between the two inputs, it is confirmed whether the OSS project to be detected has a vulnerability patch. The present invention uses a bidirectional LSTM twin network to detect vulnerabilities, without the need to manually define matching rules, and is more robust; the problem that the existing scheme cannot detect cross-function vulnerabilities can be solved; the time overhead of detection can be greatly reduced by reducing the detection space.

[0168] Although the present application is described herein in conjunction with various embodiments, in the process of implementing the claimed application, those skilled in the art may understand and implement other variations of the disclosed embodiments by reviewing the drawings, the disclosure, and the appended claims. In the claims, the word "comprising" does not exclude other components or steps, and "a" or "an" does not exclude a plurality of components or steps.

[0169] The above contents are further detailed descriptions of the present invention in combination with specific preferred embodiments, and it cannot be determined that the specific implementation of the present invention is limited to these descriptions. For ordinary technicians in the technical field to which the present invention belongs, several simple deductions or substitutions can be made without departing from the concept of the present invention, which should be regarded as falling within the protection scope of the present invention.

Claims

1. A vulnerability patch existence detection method based on deep learning, characterized in that: include: Step 1: Obtain the vulnerability disclosure entry CVE information from the common database, and search for the OSS repository of the original patch and the corresponding downstream OSS repository based on the CVE information. By comparing the original patch with the potential patch in the downstream OSS repository, the original patch and the potential patch with consistent patch information are determined as equivalent patch pairs, and all equivalent patch pairs are combined into an equivalent patch set; Step 2: for each equivalent patch pair in the equivalent patch set, perform version rollback for each patch in the equivalent patch pair, obtain a source code file, and build a code attribute graph according to the source code file, find a statement node related to the core code line in the code attribute graph, and generate a slice with the statement node as the slice entry. Each equivalent patch pair can generate an equivalent slice pair, and the equivalent slice pairs form a slice set; Step 3: normalize all slice codes in the slice set to obtain a processed slice set; Step 4: According to different situations of whether the slice pairs in the slice set are from the same equivalent patch, data processing is performed on the equivalent slices in a positive and negative sample manner, and the processed slice pairs are converted into word vector pairs; Step 5: Input the word vector pair into the bidirectional LSTM twin network to obtain the representation result of the slice pair, and determine whether they are equivalent according to the similarity between the representation results of each slice pair, so as to train the bidirectional LSTM twin network and obtain a trained bidirectional LSTM twin network; Step 6: Obtain an OSS project to be detected that includes multiple source code files to be detected, and narrow the detection space of the OSS project to be detected; Step 7: Build a code attribute graph based on each source code file in the OSS project after the detection space is reduced, select the slice entry of the code attribute graph to generate a slice to be detected that carries the core code of the source code to be detected, and repeat steps 3 to 4 for the slice to be detected to convert it into a word vector to be detected; Step 8: Based on the trained bidirectional LSTM twin network, the word vector to be detected is identified to obtain the representation result of the detection word vector, and according to the similarity between the representation results and the similarity between the input objects of the trained bidirectional LSTM twin network, it is confirmed whether the source code file to be detected has a patch to fix the vulnerability.

2. According to the method for detecting vulnerability patch existence based on deep learning in claim 1, it is characterized in that: Step 1 includes: Step 11: Obtain common vulnerability entry disclosure CVE information from the common database; Step 12: Get the URL information of the patch from the relevant link in the CVE information description; Step 13: According to the URL information, locate the OSS repository where the original patch is located, the modification location and the commit number; Step 14: According to the license content recorded in the OSS warehouse where the original patch is located, obtain the warehouse list of the OSS warehouse of the original patch and the downstream target OSS warehouse; The downstream target OSS repository includes the commit number of the patch and the potential patches after the patch is applied; Step 15: For each commit operation in the downstream target OSS repository, perform string matching on the subject of the commit operation and the message of the original patch; if either the subject or the message matches successfully, it is determined that the potential patch corresponding to the commit operation is an equivalent patch pair with the original patch; if neither the subject nor the message matches successfully, check whether the original patch and the specific patch content of the commit operation match; if they match, it is determined that the potential patch corresponding to the commit operation is an equivalent patch pair with the original patch.

3. According to the method for detecting vulnerability patch existence based on deep learning in claim 1, it is characterized in that: Step 2 includes: Step 21: for each equivalent patch pair in the equivalent patch set, roll back each patch in the equivalent patch pair to the corresponding project version according to the patch commit number, and obtain the path information in the patch; Step 22: extracting the source code file of the patch based on the path information; Step 23: Input the source code file into a static analysis tool to obtain a control flow graph (CFG), a program dependency graph (PDG), and a function call graph (CallG) of the source code file; Step 24: combining the control flow graph (CFG), the program dependency graph (PDG) and the function call graph (CallG) to obtain a code attribute graph; Step 25: Using the core code line related to the patch as the slice entry, searching the control dependency node, data dependency node and possible function call relationship related to the core code line in the code attribute graph; Step 26: forming equivalent slice pairs with the code lines corresponding to the control dependency nodes, data dependency nodes and possible function call relationships; Step 27: Group all equivalent slice pairs into slice sets.

4. According to the method for detecting vulnerability patch existence based on deep learning in claim 3, it is characterized in that: The control flow graph (CFG) in step 23 describes the possible execution path of each line of code in the program, wherein the nodes represent the code lines and the edges represent the execution path of the code; the program dependency graph (PDG) describes the interdependence and mutual influence between code instructions, including control dependency and data dependency; the function call graph (CallG) describes the relationship between functions in the source file, including the calling and called relationship.

5. According to the method for detecting vulnerability patch existence based on deep learning in claim 1, it is characterized in that: Step 3 includes: Step 31: for each slice pair in the slice set, use the symbol VAR to replace all local variable definitions appearing in the slice, and add the order information of the variable appearance to obtain a slice set after the variables are normalized; Step 32: For each slice pair in the slice set after variable normalization, use the symbol PARAM to replace all formal parameters appearing in the function, and add the order information of the formal parameters to obtain the slice code set after parameter normalization; Step 33: For each slice code in the slice code set after parameter normalization, use the symbol DTYPE to replace all data types, and add the order information of the data type, so as to obtain the slice code set after data type normalization; Step 34: For each slice code in the slice code set after data type normalization, use the symbol FUN to replace all the function names that appear, and add the order information of the function appearance to obtain the processed slice code set.

6. The method for detecting vulnerability patch existence based on deep learning according to claim 1, characterized in that: Step 4 includes: Step 41: Match each equivalent slice pair with other slice pairs to construct inequivalent slice pairs; Step 42: taking the equivalent slice pairs as positive samples and the non-equivalent slice pairs as negative samples, setting the labels of the positive samples to 1 and the negative samples to 0; Step 43: Use the pre-trained word2vec model to convert positive samples and negative samples into word vectors of the same dimension.

7. The method for detecting vulnerability patch existence based on deep learning according to claim 1, characterized in that: Step 5 includes: Step 51: For the word vector pair of the negative sample and the word vector pair of the positive sample, one word vector in the word vector pair is input into the left sub-network of the bidirectional LSTM twin network, and the other word vector is input into the right sub-network to obtain the output vectors M1 and M2 corresponding to the left sub-network and the right sub-network; Step 52: using the output vectors M1 and M2 as the representation results of the slice pair respectively, and using the similarity function to calculate the similarity between the representation results of each slice pair; Step 53: judging whether the input slice pair is consistent with the characterization result according to the similarity, and adjusting the parameters of the bidirectional LSTM twin network if they are inconsistent; Step 54: Repeat steps 51 to 53 until all word vector pairs are traversed to obtain a trained bidirectional LSTM twin network.

8. The method for detecting vulnerability patch existence based on deep learning according to claim 1, characterized in that: Step 6 includes: Step 61: Preliminarily screen the source code files to be detected based on the subject and log information in the historical submission of the OSS project to be detected, so as to merge the source code files to be detected that have the same open source OSS components; Step 62: Screening the source code files to be detected according to the patch path information of the disclosed CVE to merge the source code files to be detected that have the same path; Step 63: Screen the source code files to be detected according to the function information of the patch to merge the source code files to be detected that have the same function information, so as to reduce the detection space of the OSS project to be detected.

9. The method for detecting vulnerability patch existence based on deep learning according to claim 1, characterized in that: Step 7 includes: Step 71: construct a code property graph according to each source code file in the OSS project after the detection space is reduced; Step 72: setting the slice entry of the source code file to be detected as a control flow node, an array pointer node and a system call function node, and generating a slice to be detected of the source code file to be detected by using the slice entry; Step 73: Repeat steps 3 to 4 for the slice to be detected to convert it into a word vector to be detected.

10. The method for detecting vulnerability patch existence based on deep learning according to claim 1, characterized in that: Step 8 includes: Step 81: input the word vector to be detected into a sub-network of the trained bidirectional LSTM twin network, input the slice word vector of the patch in the existing CVE information into another sub-network, and output the respective representation results based on the trained bidirectional LSTM twin network; Step 82: Calculate the similarity between the characterization results in step 81 and the similarity between the word vectors of the two sub-network inputs, and confirm whether the OSS project to be detected has a patch to fix the vulnerability based on whether the two similarities exceed their respective thresholds.

Citation Information

Patent Citations

  • Vulnerability detection method and device, equipment and storage medium

    CN113297584A

  • Mapping a vulnerability to a stage of an attack chain taxonomy

    US20210367961A1