Inter-procedural Vulnerability Detection Method and Device Based on Hypergraph Convolution

By applying hypergraph convolution technology in the field of software analysis, combining separation logic and graph convolution methods, the problems of high false positive rates and difficult semantic information capture in the existing technology are solved, and more accurate and efficient inter-process vulnerability detection is achieved.

CN115455432BActive Publication Date: 2025-06-27SICHUAN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211167285.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-09-23
Publication Date
2025-06-27
Estimated Expiration
2042-09-23

AI Technical Summary

Technical Problem

Existing static detection methods have high false positive rates when detecting inter-process vulnerabilities in software, making it difficult to accurately define the scope of vulnerability codes, and it is difficult to capture syntactic semantic information across processes.

Method used

The source code inter-process vulnerability detection method based on hypergraph convolution is adopted, and the weak inter-process control flow diagram is initially positioned and reconstructed through tools based on separation logic. Combined with simple graph convolution and hypergraph convolution operations, multi-level information within and between processes is captured and inputted into a multi-layer fully connected network for detection.

Benefits of technology

It effectively reduces the false positive rate, improves the ability to accurately define the scope of inter-process vulnerability code, and can more comprehensively capture the syntactic semantic characteristics of the code, thereby improving the accuracy and effectiveness of vulnerability detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115455432B_ABST
    Figure CN115455432B_ABST
Patent Text Reader

Abstract

The present invention discloses a method and apparatus for inter-procedural vulnerability detection of source code based on hypergraph convolution, which relates to the field of software analysis. The present invention includes potential vulnerability location, reconstruction of a weak inter-procedural control flow graph, initialization of feature representation, multi-stage convolution, and readout detection; for the inter-procedural vulnerabilities spanning multiple functions in a software project, the present invention utilizes a multi-stage graph convolutional network to capture the intra-procedural and inter-procedural syntactic and semantic features of the code, so as to achieve the detection and identification of inter-procedural vulnerabilities in the code; the present invention can not only reasonably delimit the code scope involved in the inter-procedural vulnerabilities, but also fully extract the initial syntactic and semantic information of the code, and use multi-stage graph convolution operations to effectively capture the intra-procedural and inter-procedural high-order information spanning multiple functions, thereby enhancing the effect of inter-procedural vulnerability detection; the present invention can meet the detection requirements of inter-procedural vulnerabilities in open-source code and achieve an improvement in the detection effect of inter-procedural vulnerabilities in the code.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of software analysis, in particular to the technical field of software source code vulnerability detection, and more specifically to a method and device for detecting inter-procedural vulnerabilities in source code based on hypergraph convolution. Background Art

[0002] The explosive growth of software program size poses a severe challenge to software security. On the one hand, software developers spend more than 50% of their time detecting code defects. They rely on automated code auditing tools to improve the security and reliability of software code by detecting and fixing code vulnerabilities. On the other hand, software supply chain attacks have increased at a rate of 650%. Such attacks usually have a wide impact and are difficult to detect because attackers inject vulnerabilities or malicious code into the software source code of trusted vendors. Therefore, it is crucial to identify source code vulnerabilities as early as possible to ensure software security. Static detection has been proven to be an effective measure for vulnerability or error detection. It can be easily applied to vulnerability detection because it does not require code execution and can cover a wider range of code errors. Although static methods for code vulnerability detection already exist, the false positive rate of current static detection methods is still generally high because these methods usually only analyze a small range of vulnerability-related codes in order to ensure the performance and scalability of detection, which makes the detection results inaccurate and does not help improve software security.

[0003] Software vulnerabilities can be divided into intra-procedural vulnerabilities and inter-procedural vulnerabilities according to the scope of the code that triggers the vulnerability. Intra-procedural vulnerabilities are code interactions that involve only a single program. Since the relevant code for this type of vulnerability is only within one process and the processing boundary is relatively simple, a lot of research has focused on handling this type of vulnerability, and there have been multiple research results on this type of vulnerability. However, some statistics show that inter-procedural vulnerabilities are also worthy of attention.

[0004] Meta counted nearly 100 fixes for five vulnerability types, of which 49.9% were cross-process code vulnerabilities. Another study used Infer to analyze 4,002 submissions of the OpenSSL project and found that among 359,192 potential vulnerabilities, only 13,437 involved a single process, while 96.26% of the vulnerabilities were caused by multiple execution processes. These data show that in actual software projects, inter-process vulnerabilities cannot be ignored and seriously threaten the security of the software.

[0005] Inter-procedural vulnerabilities involve multiple procedures, which may cover different functions in a single file or even multiple executing functions across multiple files. Therefore, when detecting inter-procedural vulnerabilities, the detection model needs to have the ability to co-analyze multiple execution procedures. However, this brings two severe challenges. Challenge 1: It is complex to reasonably define the scope of the vulnerable code. Actual programs usually have multiple execution paths spanning multiple execution procedures, and the code related to the vulnerability only involves a part of the execution path. Therefore, it is difficult to reasonably define the scope of the vulnerable code. Challenge 2: It is challenging to fuse and analyze the syntactic and semantic information passed in multiple procedures. Intuitively, there are strong relationships between the codes within an execution procedure, and the mutual interaction of multiple execution procedures can convey richer semantic information.

[0006] For Challenge 1, a common approach is to slice the program using keywords related to the vulnerability. However, this largely depends on the size and quality of the keyword library, making it unable to effectively define the code involved in the vulnerability. The separation logic has the characteristic of performing inter-procedural analysis on the program. The vulnerability detector Infer based on separation logic can better locate the trace of inter-procedural vulnerable code, but its false positive rate is usually high.

[0007] For Challenge 2, the general idea is to train an effective model to capture the syntactic and semantic features of software code. Some works detect code vulnerabilities by mining vulnerable code patterns or testing the similarity between suspicious code and vulnerable code.

[0008] However, both of these two methods can only capture the syntactic features of the code at a coarse-grained level, unable to achieve an understanding of the code semantics, and it is difficult for them to handle process detection and analysis. Some other works use deep learning techniques to learn the abstract syntactic and semantic information of the source code. The form of the input data and the choice of the deep learning model are the keys to this type of work. Some works extract the execution sequence of the code as the input and cooperate with different recurrent neural network models to learn the syntactic and semantic features of the code. However, due to the complex characteristics of the code, these methods usually cannot fully obtain the semantic information of the code. There are complex relationships between code statements themselves. Therefore, some works use graph neural networks to learn high-order code syntactic and semantic features from a more expressive code graph structure. However, the current methods cannot effectively process the information exchanged between multiple execution procedures. In summary, effectively defining the scope of the code involved in inter-procedural vulnerabilities and accurately capturing the cross-procedural syntactic and semantic information is a challenging task, and the existing methods cannot effectively handle the detection of inter-procedural vulnerabilities. Summary of the Invention

[0009] To overcome the defects and deficiencies existing in the above-mentioned prior art, the present invention provides a source code inter-procedural vulnerability detection method and device based on hypergraph convolution. The object of the present invention is to provide a source code inter-procedural vulnerability detection method based on hypergraph convolution (abbreviated as HGIVul for A Code Inter-procedural Vulnerabilities Detection Method based on Hypergraph Convolution), which is used to detect cross-procedural vulnerabilities in software systems, so as to better meet the requirements of detecting complex vulnerabilities in code, expand the scope of code vulnerability detection, improve the effect of source code vulnerability detection, and thus provide support for enhancing the security of code software systems.

[0010] To solve the problems existing in the above-mentioned prior art, the present invention is implemented through the following technical solutions.

[0011] The first aspect of the present invention provides a source code inter-procedural vulnerability detection method based on hypergraph convolution. The detection method specifically includes the following steps:

[0012] S1. Use a separation logic-based tool to perform inter-procedural analysis on the source code to be detected, and use the separation logic-based tool to preliminarily locate the suspicious vulnerabilities existing in the source code to be detected;

[0013] S2. Reconstruct the weak inter-procedural control flow graph according to the Trace of the code corresponding to the suspicious vulnerabilities preliminarily located by the separation logic-based tool in step S1;

[0014] S3. Vectorize the initial feature information in the node code of the weak inter-procedural control flow graph reconstructed in step S2; specifically, extract the initial semantic information of the node code, vectorize the syntax information of the node code, and splice the initial semantic information and syntax information of the node code to form a vectorized representation of the initial weak inter-procedural control flow graph node feature information;

[0015] S4. First, perform simple graph convolution operations on the initialized weak inter-procedural control flow graph in step S3 to capture the features of the intra-procedural code; then apply hypergraph convolution on the initialized weak inter-procedural control flow graph in step S3 to capture the features of the inter-procedural code, so as to achieve fine-grained capture of multi-level information in the weak inter-procedural control flow graph;

[0016] S5. Read out the embedding of the weak inter-procedural control flow graph updated by the convolution operation in step S4 as the feature representation of the entire graph, and then input the obtained feature representation into a detector constructed by a multi-layer fully connected layer for detection, and judge whether there are vulnerabilities according to the detection results output by the detector.

[0017] Further, in step S2, reconstructing the weak inter-procedural control flow graph specifically includes the following sub-steps:

[0018] S201. According to the Trace of the code corresponding to the suspicious vulnerability preliminarily located by the tool based on separation logic in step S1, locate each execution process involved in the code, and reconstruct the control flow graph of each process;

[0019] S202. According to the code execution order in the Trace, determine the inter-procedural call points, and then add two incoming and outgoing edges at each determined call point to connect the control flow graphs of multiple processes to form a weak inter-procedural control flow graph.

[0020] Further, in step S3, vectorize the initial feature information in the node code of the weak inter-procedural control flow graph reconstructed in step S2, specifically including:

[0021] S301. Use a lexical analyzer to extract the basic units in the node code, and perform symbolic processing on the variable names and function names in the basic units;

[0022] S302. For the basic units after symbolic processing, use the pre-trained Word2vec model to extract the corresponding initial embeddings to capture the initial semantic information of the code; calculate the average value of the embeddings corresponding to multiple basic units existing in the node code in the same dimension to form the initial semantic embedding of the node code;

[0023] S303. Extract the canonical LLVM syntax tree basic units for the node code, and multiple LLVM syntax tree basic units corresponding to each node form a set;

[0024] S304. Count the number of different types of basic units in the set of LLVM syntax tree basic units corresponding to each node, normalize the number of each type of basic unit, and then perform One-hot encoding in the order of specific basic unit types to form the initial syntax embedding of the node;

[0025] S305. Concatenate the initial semantic embedding of the node code obtained in step S302 and the initial syntax embedding of the node obtained in step S304 to form the initial feature representation of the node, that is, the vectorized representation of the initialized weak inter-procedural control flow graph node feature information.

[0026] Further preferably, in step S4, the simple graph convolution operation specifically refers to:

[0027] Regard the initialized weak inter-procedural control flow graph in step S3 as a simple graph G=(V, E), where V represents the node set and E represents the edge set; the initial embeddings of all nodes in the initialized weak inter-procedural control flow graph are represented as where \(d\) represents the dimension of the word embedding vector, represents the set of real numbers, and its dimension is \(|V|\times|d|\);

[0028] For node \(v\) i , the initial embedding representation is After \(k - 1\) convolutions, the feature embedding representation of node \(v\) i is Therefore, the feature embedding representation of node \(v\) i after \(k\) convolutions is

[0029] Specifically, the convolution operation on a simple graph is calculated by the following formula:

[0030] In the formula, \(N(i)\) represents the set of neighbor nodes of node \(v\) i , represents any neighbor node \(v\) i of node \(v\) j , and its feature embedding representation after \(k - 1\) convolutions is \(M(\cdot)\) is the aggregation mean aggregation function, \(W\) represents the trainable weight, and \(\sigma(\cdot)\) represents the activation function; the feature embedding representation of the weakly interprocedural control flow graph after simple graph convolution is \(X\) simple .

[0031] Based on the above simple graph convolution operation, the weakly interprocedural control flow graph is regarded as a hypergraph \(G\) h \(=(V, E\) h ), where \(E\) h represents the set of hyperedges. For a hyperedge \(e\in E\) h , \(e = \{v_1,\ldots,v\) p \}, \(v\) i \in V\), \(2\leq p\leq|V|\), \(v\) i represents the \(i\)-th node in the node set \(V\), and \(v\) p represents the \(p\)-th node in the node set \(V\); the co-occurrence matrix on the hypergraph represents the set of real numbers, and its dimension is \(|V|\times|E\) h |, and each element \(H\) (v,e) of the co-occurrence matrix is determined by the following formula:

[0032]

[0033] The degree matrix \(D\) s of the simple graph is a diagonal matrix representing the set of real numbers and its dimension is \(|V|\times|V|\), and each diagonal element is calculated by the following formula:

[0034] Degree matrix D of the hypergraph h is also a diagonal matrix denotes the set of real numbers and its dimension is |E h |×|E h |, and each diagonal element is calculated by the following formula: D h (v,v)=∑ v∈V H (v,e) .

[0035] Hypergraph convolution is performed on the weak inter-procedural control flow graph and is calculated by the following formula:

[0036]

[0037] where H T denotes the transpose of H, and let The above formula can be simplified to:

[0038] In the formula, k represents performing k convolutions on the hypergraph, and X k denotes the set of node feature representations. When k = 0, X 0 = X simple ; The node feature representation updated through the above multi-level graph convolution is X'.

[0039] Furthermore, the S5 specifically includes:

[0040] S501. Read out the embedding of the weak inter-procedural control flow graph updated by the convolution operation in step S4 as the feature representation of the entire graph, and select the maximum value reading method according to the corresponding dimension, which is specifically calculated by the following formula:

[0041] r = MAX({x′1,x′2,…,x′ i}), x′ i ∈X′, where r represents the feature representation of the entire weak inter-procedural control flow graph;

[0042] S502. Input the obtained feature representation r into a multi-layer fully connected network for detection, which is specifically calculated by the following formula:

[0043] In the formula, represents the final detection result, MLP represents the multi-layer fully connected network, and the sigmoid function is used to output the final detection result.

[0044] Furthermore, in the S1 step, the tool Infer based on separation logic is used to perform preliminary inter-procedural analysis on the source code to be detected, and potential vulnerabilities are initially screened out based on the results output by Infer, and the scope involved in the vulnerabilities is initially located.

[0045] The second aspect of the present invention provides a source code inter-procedural vulnerability detection device based on hypergraph convolution, which includes:

[0046] A potential vulnerability location module, which uses a tool based on separation logic to perform inter-procedural analysis on the source code to be detected, and uses the tool based on separation logic to preliminarily locate the suspicious vulnerabilities existing in the source code to be detected;

[0047] A reconstructed weak inter-procedural control flow graph module, which reconstructs a weak inter-procedural control flow graph according to the Trace of the code corresponding to the suspicious vulnerability preliminarily located by the potential vulnerability location module;

[0048] An initialization feature representation module, which is used to extract the initial semantic information of the node code, vectorize the syntax information of the node code, and splice the initial semantic information and syntax information of the node code to form a vectorized representation of the initialized weak inter-procedural control flow graph node feature information;

[0049] A multi-stage convolution module, which is used to perform simple graph convolution operations and hypergraph convolution operations on the initialized weak inter-procedural control flow graph in sequence to achieve fine-grained capture of multi-level information in the weak inter-procedural control flow graph;

[0050] A readout detection module, which reads out the embedding of the weak inter-procedural control flow graph updated by the multi-stage convolution module as the feature representation of the entire graph, and then inputs the obtained feature representation into a detector constructed by a multi-layer fully connected network for vulnerability detection.

[0051] Compared with the prior art, the beneficial technical effects brought by the present invention are as follows:

[0052] 1. The present invention has the ability to reasonably define the code scope involved in inter-procedural vulnerabilities, and uses the processing results of the tool Infer based on analysis logic to better define the code scope of multiple procedures involved in inter-procedural vulnerabilities.

[0053] 2. The present invention has a more comprehensive code initial information extraction ability, converts the initial code into an intermediate representation in the form of a graph structure, and fully extracts the initial syntax and semantic information of the code, providing support for learning the high-order information of the code.

[0054] 3. The present invention can more effectively learn high-order syntax and semantic information among multiple procedures. Performing multi-stage graph convolution can effectively capture the complex relationships within and between code procedures, extract high-value code syntax and semantic information, and thus improve the inter-procedural vulnerability detection effect. BRIEF DESCRIPTION OF THE DRAWINGS

[0055] Figure 1 It is the overall architecture diagram of the method HGIVul of the present invention.

[0056] Figure 2 It is a flow chart example for initializing the feature representation module.

[0057] Figure 3 It is the vulnerability detection effect of HGIVul compared with multiple methods on the dataset Dataset 1.

[0058] Figure 4 It is the confusion matrix of VulDeePecker for classifying 10 types of vulnerabilities on the dataset Dataset 2.

[0059] Figure 5 It is the confusion matrix of SySeVR for classifying 10 types of vulnerabilities on the dataset Dataset 2.

[0060] Figure 6 It is the confusion matrix of C-BERT for classifying 10 types of vulnerabilities on the dataset Dataset 2.

[0061] Figure 7 It is the confusion matrix of DeepWukong for classifying 10 types of vulnerabilities on the dataset Dataset 2.

[0062] Figure 8 It is the confusion matrix of HGIVul for classifying 10 types of vulnerabilities on the dataset Dataset 2. Specific implementation manners

[0063] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings of the present invention specification. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0064] Embodiment 1

[0065] As a preferred embodiment of the present invention, referring to the attached Figure 1 As shown in the figure, this embodiment discloses a method for inter-procedural vulnerability detection of source code based on hypergraph convolution. The detection method specifically includes the following steps:

[0066] S1. Use the tool Infer based on separation logic to perform inter-procedural analysis on the source code to be detected, and use the tool Infer based on separation logic to preliminarily locate the suspicious vulnerabilities existing in the source code to be detected.

[0067] S2. Reconstruct the weak inter - procedural control flow graph according to the Trace of the code corresponding to the suspicious vulnerability preliminarily located by the tool based on separation logic in step S1; First, locate multiple procedures involved in the vulnerability code according to the Trace output by Infer; Then, construct the corresponding control flow graph for each involved procedure; Next, determine the inter - procedural call sites according to the execution order of the code in the Trace; Finally, add two connecting edges, Callin and Call out, at each determined call site, and connect the control flow graphs of multiple procedures to form a weak inter - procedural control flow graph.

[0068] S3. Vectorize the initial feature information in the node code of the weak inter - procedural control flow graph reconstructed in step S2; Specifically, extract the initial semantic information of the node code, and vectorize the syntactic information of the node code, and splice the initial semantic information and syntactic information of the node code to form a vectorized representation of the features of the nodes in the initialized weak inter - procedural control flow graph;

[0069] More specifically, use the pre - trained Word2vec model to extract the initial semantic information of the code. On the other hand, vectorize the syntactic information of the nodes based on the basic units Token of the LLVM syntax tree corresponding to the code. Finally, splice the initial semantic and syntactic information of the code to form a vectorized representation of the feature information of the nodes in the initialized weak inter - procedural control flow graph.

[0070] S4. First, perform simple graph convolution operations on the initialized weak inter - procedural control flow graph in step S3 to capture the features of the intra - procedural code; Then, apply hyper - graph convolution on the initialized weak inter - procedural control flow graph in step S3 to capture the features of the inter - procedural code, so as to achieve fine - grained capture of multi - level information in the weak inter - procedural control flow graph;

[0071] S5. Read out the embedding of the weak inter - procedural control flow graph updated by the convolution operation in step S4 as the feature representation of the whole graph, and then input the obtained feature representation into the detector constructed by the multi - layer fully - connected layer for detection, and judge whether there is a vulnerability according to the detection result output by the detector.

[0072] Embodiment 2

[0073] As another preferred embodiment of the present invention, this embodiment is a detailed elaboration of the specific implementation manner of step S2 in the above - mentioned embodiment 1. In this embodiment, in step S2, the reconstruction of the weak inter - procedural control flow graph specifically includes the following sub - steps:

[0074] S201. According to the Trace of the code corresponding to the suspicious vulnerability preliminarily located by the tool based on separation logic in step S1, locate each execution procedure involved in the code, and reconstruct the control flow graph of each procedure;

[0075] S202. Determine the inter - procedural call points according to the code execution order in the Trace, and then add two incoming and outgoing edges at each call point based on the determined call point positions to connect the control flow graphs of multiple procedures to form a weak inter - procedural control flow graph.

[0076] Embodiment 3

[0077] As another preferred embodiment of the present invention, this embodiment is a detailed elaboration of the specific implementation manner of step S3 in the above - mentioned Embodiment 1. In this embodiment, in step S3, the initial feature information in the node code of the weak inter - procedural control flow graph reconstructed in step S2 is vectorized, which specifically includes:

[0078] S301. Use a lexical analyzer to extract the basic units in the node code, and perform symbolic processing on the variable names and function names existing in the basic units.

[0079] S302. For the basic units after symbolic processing, use the pre - trained Word2vec model to extract the corresponding initial embeddings to capture the initial semantic information of the code; calculate the average value of the embeddings corresponding to multiple basic units existing in the node code in the same dimension to form the initial semantic embedding of the node code.

[0080] S303. Extract the canonical LLVM syntax tree basic units for the node code, and a set composed of multiple canonical basic units corresponding to each node.

[0081] S304. Count the number of different types of basic units in the LLVM basic unit set corresponding to each node, perform normalization processing on the number of each type of basic unit, and then perform One - hot encoding in the order of specific basic unit types to form the initial syntax embedding of the node.

[0082] S305. Concatenate the initial semantic embedding of the node code obtained in step S302 and the initial syntax embedding of the node obtained in step S304 to form the initial feature representation of the node, that is, the vectorized representation of the initialized weak inter - procedural control flow graph node feature information.

[0083] Embodiment 4

[0084] As another preferred embodiment of the present invention, referring to the attached Figure 1 , this embodiment discloses an inter - procedural vulnerability detection method for source code based on hyper - graph convolution, including the following steps:

[0085] S1. Use the tool Infer based on separation logic to perform inter - procedural analysis on the source code to be tested, and use Infer to preliminarily locate the suspicious vulnerabilities.

[0086] S2. Reconstruct the soft inter - procedural control flow graph (Soft ICFG). Reconstruct the inter - procedural control flow graph according to the trace of the suspicious code output by Infer. First, locate multiple procedures involved in the vulnerable code based on the trace output by Infer; then construct the corresponding control flow graph for each involved procedure; next, determine the call sites between procedures according to the execution order of the code in the trace; finally, add two connecting edges, call in and call out, at each call site according to the determined call site positions, and connect the control flow graphs of multiple procedures to form the soft inter - procedural control flow graph (Soft ICFG).

[0087] S3. Initialize the feature representation of Soft ICFG nodes. Initializing the feature representation is to vectorize the initial feature information of nodes, and vectorizing node attributes is an important step in automated vulnerability detection. HGIVul initializes the feature representation of nodes from two perspectives. On the one hand, HGIVul uses the pre - trained Word2vec model to extract the initial semantic information of the code. On the other hand, HGIVul vectorizes the syntactic information of nodes based on the LLVM syntax tree tokens corresponding to the code. Finally, the initial semantic and syntactic information of the code is concatenated to form the initial feature representation of Soft ICFG nodes.

[0088] S4. Multi - stage graph convolution. This step is used to obtain high - order valuable code features. HGIVul first performs simple graph convolution operations on the Soft ICFG to capture the intra - procedural code features, and then applies hyper - graph convolution on the Soft ICFG to capture the inter - procedural code features, thereby achieving fine - grained capture of multi - level information in the Soft ICFG.

[0089] S5. Read out the updated feature representation and perform detection. First, read out the embedding of the updated Soft ICFG as the feature representation of the whole graph. Then input the obtained high - order feature representation into the detector constructed by multiple fully - connected layers for detection. Finally, output the detection result.

[0090] Furthermore, the S2 step specifically includes:

[0091] Step 201: Based on the trace output by Infer, locate each execution procedure (function) involved in the code and reconstruct the control flow graph (CFG, Control Flow Graph) of each procedure.

[0092] Step 202: Determine the inter-procedural call sites according to the code execution order in the Trace, and then connect the CFGs of multiple procedures based on the determined call sites to form a Soft ICFG. Note that two connecting edges, namely Call in and Call out, are added respectively at each call site.

[0093] Furthermore, the initialization of node feature representation in Step 3 includes:

[0094] Step 301: Use a lexical analyzer to extract the basic unit tokens in the node code, and perform symbolic processing on the variable names and function names in the tokens to avoid the impact of different naming habits on the code semantic information;

[0095] Step 302: For the symbolically processed node tokens, use a pre-trained Word2vec model to extract the corresponding initial embeddings to capture the initial semantic information of the code. And calculate the average value of the embeddings corresponding to multiple tokens in the node code in the same dimension to form the initial semantic embedding of the node;

[0096] Step 303: Extract the standard LLVM syntax tree tokens for the initial code of the node, and multiple standard tokens corresponding to each node form a set;

[0097] Step 304: Count the number of different types of tokens in the LLVM token set corresponding to each node, normalize the number of each type of token, and then perform One-hot encoding in the specific token type order to form the initial syntax embedding of the node;

[0098] Step 305: Concatenate the initial semantic embedding of the node obtained in Step 302 and the initial syntax embedding of the node obtained in Step 304 to form the initial feature representation of the node.

[0099] Furthermore, the multi-stage graph convolution in Step S4 specifically includes:

[0100] S401: First, perform a simple graph convolution on the Soft ICFG

[0101] Regard the weakly inter-procedural control flow graph initialized in Step S3 as a simple graph G=(V, E), where V represents the node set and E represents the edge set; the initial embeddings of all nodes in the initialized weakly inter-procedural control flow graph are represented as where d represents the dimension of the word embedding vector, represents the set of real numbers, and its dimension is |V|×|d|;

[0102] For node v i, the initial embedding representation is the feature embedding representation of node v after k-1 convolutions i is Therefore, the feature embedding representation of node v i after k convolutions is

[0103] Specifically, the convolution operation on a simple graph is calculated by the following formula:

[0104] In the formula, N(i) represents the set of neighbor nodes of node v i , denotes the arbitrary neighbor node v i of node v j , and the feature embedding representation of its k-1 convolutions is M(·) is the aggregation mean aggregation function, W represents the trainable weight, and σ(.) represents the activation function; the feature embedding representation of the weakly interprocedural control flow graph after simple graph convolution is X simple ;

[0105] S402. Construct the contribution matrix H for hypergraph convolution

[0106] Based on the above simple graph convolution operation, the weakly interprocedural control flow graph is regarded as a hypergraph G h =(V, E h ), where E h represents the set of hyperedges. For a hyperedge e ∈ E h , e = {v1,..., v p}, v i ∈ V, 2 ≤ p ≤ |V|, v i represents the i-th node in the node set V, and v p represents the p-th node in the node set V; the co-occurrence matrix on the hypergraph represents the set of real numbers, and its dimension is |V| × |E h |, and each element H (v,e) of the co-occurrence matrix is determined by the following formula:

[0107]

[0108] S403. Calculate the degree matrix D of the simple graph s

[0109] The degree matrix D of the simple graph s is a diagonal matrix representing the set of real numbers and its dimension is |V| × |V|, and each diagonal element is calculated by the following formula:

[0110] S404. Calculate the degree matrix D of the hypergraph h

[0111] The degree matrix D of the hypergraph h is also a diagonal matrix denotes the set of real numbers and its dimension is |E h |×|E h |. Each diagonal element is calculated by the following formula: D h (v, v) = ∑ v∈V H (v,e) .

[0112] S405. Perform hypergraph convolution

[0113] Perform hypergraph convolution on the weak inter-procedure control flow graph, which is calculated by the following formula:

[0114]

[0115] where H T denotes the transpose of H. Let The above formula can be simplified to:

[0116] In the formula, k represents performing k convolutions on the hypergraph, and X k denotes the set of node feature representations. When k = 0, X 0 = X simple ; The node feature representation after being updated by the above multi-level graph convolution is X'.

[0117] Furthermore, the specific steps of reading out the updated feature representation and performing detection in step S5 include:

[0118] S501. Read out the embedding of the weak inter-procedure control flow graph updated by the convolution operation in step S4 as the feature representation of the entire graph, and select the maximum value reading method according to the corresponding dimension, which is specifically calculated by the following formula:

[0119] r = MAX({x'1, x'2,..., x' i}), x' i ∈ X', where h represents the feature representation of the entire weak inter-procedure control flow graph;

[0120] S502. Input the obtained feature representation r into a multi-layer fully connected network for detection, which is specifically calculated by the following formula:

[0121] In the formula, denotes the final detection result, MLP represents the multi-layer fully connected network, and the sigmoid function is used to output the final detection result.

[0122] Example 5

[0123] As another preferred embodiment of the present invention, referring to the appended Figure 1 description, this embodiment discloses a source code inter-procedural vulnerability detection device based on hypergraph convolution. The overall architecture is as Figure 1 shown. The device mainly consists of a potential vulnerability location module, a reconstructed weak inter-procedural control flow graph module, an initial feature representation module, a multi-stage graph convolution module, and a readout detection module. Among them, the potential vulnerability location module is used to define the scope involved in the potential vulnerability code. The reconstructed weak inter-procedural control flow graph module transforms the code into a more expressive graphical structure. The initial feature representation module is used to vectorize the initial syntactic and semantic information of the code. The multi-stage convolution module is used to capture high-order information within and between procedures. The readout detection module is used to output the detection results of inter-procedural vulnerabilities.

[0124] The potential vulnerability location module uses the separation logic-based tool Infer to perform inter-procedural analysis on the project code and uses the output results of Infer to initially define multiple procedures (functions) involved in the vulnerability code.

[0125] The reconstructed weak inter-procedural control flow graph module reconstructs the weak inter-procedural control flow graph based on the results output by Infer. First, HGIVul locates multiple procedures involved in the vulnerability code according to the Trace output by Infer. Then, for each involved procedure, a corresponding control flow graph is constructed. Next, according to the execution order of the code in the Trace, the inter-procedural call points (CallSite) are determined. Finally, according to the determined call point positions, two connecting edges, namely Call in and Call out, are added at each call point to connect the control flow graphs of multiple procedures to form a weak inter-procedural control flow graph. The complete inter-procedural control flow graph will connect the control flow graph of a target function at each call point, and there may be cases where the control flow graphs of the same function are connected multiple times, which will result in a large amount of storage and computational overhead. In contrast, the weak inter-procedural control flow graph only connects the position where the target function call point appears for the first time.

[0126] The initialization feature representation module vectorizes the initial feature information in the node code of the weak inter-procedural control flow graph. Vectorizing node attributes is an important step in automated vulnerability detection. The initial attribute of a node is code text. To more comprehensively obtain the initial syntactic and semantic information of the code, HGIVul initializes the feature representation of nodes from two aspects. On the one hand, HGIVul uses a pre-trained Word2vec model to extract the initial semantic information of the code. On the other hand, HGIVul vectorizes the syntactic information of nodes based on the LLVM syntax tree tokens corresponding to the code. Finally, the initial semantic and syntactic information of the code is concatenated to form a vectorized representation of the initial Soft ICFG node feature information.

[0127] Figure 2 Figure 4 shows an example process of initializing feature representation. First, use a lexical analyzer to extract the basic unit tokens in the node code, and symbolize the variable names and function names in the tokens (for example, the variable names "dst", "di", "si" are symbolized as "VAR1", "VAR2", "VAR3" respectively) to avoid the impact of different naming conventions on the code semantic information. Then, use the pre-trained Word2vec model to extract the corresponding initial embeddings to capture the initial semantic information of the code. For the case where a node contains multiple tokens, calculate the average value of each token's corresponding dimension to form a new vector as the initial semantic embedding of the node. Next, extract the LLVM syntax tree token expressions corresponding to the initial code of the node, count the number of different types of tokens in the LLVM token set corresponding to each node, and perform normalization processing. Secondly, perform One-hot encoding in a specific token type order to form the initial syntactic embedding of the node. Finally, concatenate the obtained initial semantic embedding and initial syntactic embedding of the node to form the initial feature representation of the node.

[0128] The multi-stage graph convolution module performs simple graph convolution and hypergraph convolution on the initialized weak inter-procedural control flow graph to learn high-order valuable code feature representations.

[0129] First, as Figure 1 shown in the multi-stage graph convolution in Figure 4, perform simple graph convolution on the Soft ICFG. Consider the initialized weak inter-procedural control flow graph as a simple graph G=(V,E), where V represents the node set and E represents the edge set. The initial embeddings of all nodes in the initialized weak inter-procedural control flow graph are represented as where d represents the dimension of the word embedding vector, represents the set of real numbers, and its dimension is |V|×|d|;

[0130] For node v i, the initial embedding representation is the feature embedding representation of node v after k - 1 convolutions i is Therefore, the feature embedding representation of node v i after k convolutions is

[0131] Specifically, the convolution operation on a simple graph is calculated by the following formula:

[0132] where N(i) represents the set of neighbor nodes of node v i , represents an arbitrary neighbor node v i of node v j , whose feature embedding representation after k - 1 convolutions is M(·) is the aggregation mean aggregation function, W represents the trainable weight, σ(.) represents the activation function; the feature embedding representation of the weakly inter - procedural control flow graph after simple graph convolution is X simple ; then the co - occurrence matrix H for hypergraph convolution, the degree matrix D of the simple graph s , and the degree matrix D of the hypergraph h are constructed from the Soft ICFG.

[0133] Based on the above simple graph convolution operation, the weakly inter - procedural control flow graph is regarded as a hypergraph G h =(V, E h ), where E h represents the set of hyperedges. For a hyperedge e ∈ E h , e = {v1,..., v p}, v i ∈ V, 2 ≤ p ≤ |V|, v i represents the i - th node in the node set V, v p represents the p - th node in the node set V; the co - occurrence matrix on the hypergraph represents the set of real numbers, its dimension is |V|×|E h |, and each element H (v,e) of the co - occurrence matrix is defined by the following formula:

[0134]

[0135] The degree matrix D of the simple graph s is a diagonal matrix represents the set of real numbers and its dimension is |V|×|V|, and each diagonal element is calculated by the following formula:

[0136] The degree matrix D of the hypergraph h is also a diagonal matrix represents the set of real numbers and its dimension is |E h |×|E h |, and each diagonal element is calculated by the following formula: D h (v,v)=∑ v∈V H (v,e) .

[0137] Finally, perform hypergraph convolution on the weak inter-procedure control flow graph, which is calculated by the following formula:

[0138]

[0139] where H T represents the transpose of H, let The above formula can be simplified to:

[0140] In the formula, k represents performing k convolutions on the hypergraph, and X k represents the set of node feature representations. When k = 0, X 0 =X simple ; The node feature representation updated by the above multi-level graph convolution is X'.

[0141] The readout detection module reads out the feature representation of the entire graph and inputs the readout feature representation into the detector for detection. First, read out the embedding of the Soft ICFG updated by the above convolution operation as the feature representation of the entire graph. The present invention selects the maximum value readout method according to the corresponding dimension, which is calculated by the following formula:

[0142] r = MAX({x′1,x′1,...,x′ i}),x′ i ∈X′ h = MAX({x′1,x′2,…,x′ i}),x′ i ∈X′, where r represents the feature representation of the entire weak inter-procedure control flow graph; then, input the obtained feature representation r into a multi-layer fully connected network for detection, which is specifically calculated by the following formula:

[0143] In the formula, represents the final detection result, MLP represents the multi-layer fully connected network, and the sigmoid function is used to output the final detection result.

[0144] The present invention has two different states of training and detection in both the multi-stage convolution module and the readout detection module. In the training state, use the training data to train the feature representation extraction model and the vulnerability detection module. In the actual detection state, directly use the trained detection model to extract the feature representation of the vulnerable code and use the trained detector for detection.

[0145] The inter - procedural vulnerability detection effect evaluation of HGIVul is carried out on the D2A dataset. D2A is a dataset containing inter - procedural information of code and composed of actual open - source projects. It contains multiple vulnerability samples of 6 open - source projects, namely Openssl, Nginx, Libtiff, Httpd, Libav, and FFmpeg. Since there are many duplicate samples in FFmpeg and Libav, the samples of the FFmpeg project are removed in the evaluation, and only the samples of the other 5 projects are retained. After processing, D2A contains 566968 available samples, including 12290 available vulnerability samples and 554678 available normal samples. For the convenience of description, this dataset is named Dataset 1. In addition, in order to verify the vulnerability recognition effect of HGIVul, 10 types of vulnerability samples with the most occurrences are selected from D2A to construct Dataset 2. The 10 types of selected samples include various common vulnerability types such as "integer overflow", "buffer overflow", "null pointer", etc. The constructed Dataset contains 12003 vulnerability samples. In addition, the samples are divided into training set, validation set, and test set according to the ratio of 8:1:1. Six metrics, namely Accuracy (Acc), Precision (Pre), Recall, F1 - score (F1), False positive rate (FPR), and False negative rate (FNR), are selected for evaluation, and 5 related methods are selected for comparative verification.

[0146] Figure 3 What is shown is the vulnerability detection effect of HGIVul and 5 related methods on Dataset 1. Among them, the abscissa represents different methods, the ordinate represents the percentage value of the detection result, and different bars represent the specific percentage values of the corresponding evaluation metrics. Overall, HGIVul has a better detection effect compared to the other 5 methods. In addition, from Figure 3The following findings can be obtained: First, Infer has the worst detection effect, with an FPR as high as 98.01%. Since the D2A dataset is constructed using Infer, its Recall and FNR are numerically the best. Second, the method of using code sequences for model learning has better detection performance than Infer. The detection effects of VulDeePecker and SySeVR are significantly better than that of Infer. The lowest FPR of VulDeePecker reaches 1.17%. The reason may be that the data-driven method can discover high-order information that cannot be obtained by human experts. Then, treating code as natural language sequences for processing helps to improve the detection effect to a certain extent. After fine-tuning the C-BERT model, a better detection effect is obtained. Finally, the detection method based on graph structure has better detection effect, and HGIVul with multi-stage graph convolution operation is better than the other 5 methods. Its Accuracy is 96.87%, Precision is 67.89%, Recall is 64.85%, and F1 is 66.33%, all of which are better than other methods, and both FNR and FPR are relatively low. HGIVul fully utilizes the structural attributes of code and obtains fine-grained syntactic and semantic feature information within and between code procedures. Therefore, it has a better inter-procedural vulnerability detection effect compared with existing methods.

[0147] Figure 4 , Figure 5 , Figure 6 , Figure 7 and Figure 8 respectively show the recognition effects of 5 methods on 10 common vulnerability types in the dataset Dataset 2. To verify the ability of HGIVul to recognize different types of vulnerabilities, the models of VulDeePecker, SySeVR, C-BERT, DeepWukong, and HGIVul are adjusted to multi-classification tasks. Figures 4 to 8The confusion matrix showing five methods for identifying ten types of vulnerabilities is presented. The horizontal axis represents the predicted vulnerability types, and the vertical axis represents the actual vulnerability types. The weighted F1 value is used to comprehensively evaluate the identification effect. Overall, HGIVul has a better identification effect for different types of vulnerabilities than other methods, and its weighted F1 value can reach 79.58%. By comparing the identification results of multiple methods, it can be found that different forms of processing the code have a great impact on the classification results. The classification effect of the method based on code sequences is worse than that of the method based on graphical structured processing. The weighted F1 values of VulDeePecker and SySeVR are 66.14% and 71.30% respectively, which are lower than other methods. The method based on graphical structured processing is better than the method based on sequences. The weighted F1 values of DeepWukong and HGIVul are 76.21% and 79.58% respectively, which are better than other methods. In addition, there are also differences in the ability of different methods to identify different types of vulnerabilities. All five methods have relatively good identification effects for three types of vulnerabilities, namely "INTEGER_OVERFLOW_L5", "NULLPTR_DEREFERENCE", and "INFERBO_ALLOC_MAY_BE_BIG", while the identification of some vulnerabilities is relatively poor, such as "INTEGER_OVERFLOW_U5" and "BUFFER_OVERRUN_U5", because the forms of these vulnerabilities vary greatly. In summary, HGIVul can obtain rich code syntax and semantic information of multiple processes, so it has better vulnerability classification and identification performance than existing methods.

[0148] Note that the above is only a preferred embodiment of the present invention and the technical principles applied. Those skilled in the art will understand that the present invention is not limited to the specific embodiments described here. Various obvious changes, re-adjustments, and substitutions can be made by those skilled in the art without departing from the protection scope of the present invention. Therefore, although the present invention has been described in detail through the above embodiments, the present invention is not limited to the above embodiments. Without departing from the concept of the present invention, more other equivalent embodiments can be included, and the scope of the present invention is determined by the scope of the appended claims.

Claims

1. A method for detecting inter - procedural vulnerabilities in source code based on hypergraph convolution, characterized in that The detection method specifically includes the following steps: S1. Use a tool based on separation logic to perform inter-procedural analysis on the source code to be detected, and use the tool based on separation logic to preliminarily locate the suspicious vulnerabilities existing in the source code to be detected; S2. Reconstruct the weak inter-procedural control flow graph according to the Trace of the code corresponding to the suspicious vulnerabilities preliminarily located by the tool based on separation logic in step S1; S3. Vectorize the initial feature information in the node code of the weak inter-procedural control flow graph reconstructed in step S2; specifically, extract the initial semantic information of the node code, vectorize the syntactic information of the node code, and splice the initial semantic information and syntactic information of the node code to form a vectorized representation of the initial feature information of the weak inter-procedural control flow graph node; S4. First, perform simple graph convolution operations on the initialized weak inter-procedural control flow graph in step S3 to capture the features of the intra-procedural code; then, apply hypergraph convolution on the initialized weak inter-procedural control flow graph in step S3 to capture the features of the inter-procedural code, so as to achieve fine-grained capture of multi-level information in the weak inter-procedural control flow graph; S5. Read out the embedding of the weak inter-procedural control flow graph updated by the convolution operation in step S4 as the feature representation of the whole graph, and then input the obtained feature representation into the detector constructed by the multi-layer fully connected layer for detection, and judge whether there are vulnerabilities according to the detection results output by the detector; In step S2, the reconstruction of the weak inter-procedural control flow graph specifically includes the following sub-steps: S201. According to the Trace of the code corresponding to the suspicious vulnerabilities preliminarily located by the tool based on separation logic in step S1, locate each execution process involved in the code, and reconstruct the control flow graph of each process; S202. Determine the inter-procedural call points according to the code execution order in the Trace, and then add two incoming and outgoing edges at each determined call point, and connect the control flow graphs of multiple processes to form a weak inter-procedural control flow graph.

2. The method for inter-procedural vulnerability detection of source code based on hypergraph convolution according to claim 1, characterized in that: In step S3, the vectorization of the initial feature information in the node code of the weak inter-procedural control flow graph reconstructed in step S2 specifically includes: S301. Use a lexical analyzer to extract the basic units in the node code, and perform symbolic processing on the variable names and function names existing in the basic units; S302. For the basic units after symbolic processing, use the pre-trained Word2vec model to extract the corresponding initial embeddings to capture the initial semantic information of the code; calculate the average value of the embeddings corresponding to multiple basic units existing in the node code in the same dimension to form the initial semantic embedding of the node code; S303. Extract the standard LLVM syntax tree basic units for the node code, and multiple LLVM syntax tree basic units corresponding to each node form a set; S304. Count the number of different types of basic units in the LLVM syntax tree basic unit set corresponding to each node, normalize the number of each type of basic unit, and then perform One-hot encoding in the order of specific basic unit types to form the initial syntactic embedding of the node; S305. Concatenate the initial semantic embedding of the node code obtained in step S302 and the initial syntactic embedding of the node obtained in step S304 to form the initial feature representation of the node, that is, the vectorized representation of the initialized weak inter-procedural control flow graph node feature information.

3. The method for detecting inter - procedural vulnerabilities in source code based on hypergraph convolution according to claim 2, wherein: Step S4 specifically includes the following sub-steps: S401. Regard the weakly inter-procedural control flow graph after the initialization in step S3 as a simple graph G=(V, E), where V represents the set of nodes and E represents the set of edges; the initial embeddings of all nodes in the weakly inter-procedural control flow graph after the initialization are represented as where d represents the dimension of the word embedding vector, represents the set of real numbers, and its dimension is |V|×|d|; For node v i , the initial embedding is represented as After k-1 convolutions, the feature embedding of node v i is represented as Therefore, the feature embedding of node v i after k convolutions is represented as Specifically, the convolution operation on the simple graph is calculated by the following formula: where N(i) represents the set of neighbor nodes of node v i ; denotes an arbitrary neighbor node v i of node v j , and the feature embedding after (k - 1) - th convolution of it is expressed as M(·) is the aggregation mean aggregation function, W represents the trainable weight, and σ(.) represents the activation function; the feature embedding of the weakly - controlled flow graph between processes after simple graph convolution after initialization is expressed as X simple ; S402. On the basis of the above S401 step, regard the weak interprocedural control flow graph as a hypergraph G h =(V, E h ), where E h represents the set of hyperedges. For a hyperedge e ∈ E h , e = {v1,..., v p}, v i ∈ V, 2 ≤ p ≤ |V|, v i represents the i-th node in the node set V, and v p represents the p-th node in the node set V; the co-occurrence matrix on the hypergraph represents the set of real numbers, with a dimension of |V| × |E h |, and each element H (v,e) of the co-occurrence matrix is determined by the following formula: S403. Degree matrix D of a simple graph s is a diagonal matrix represents the set of real numbers and has a dimension of |V|×|V|. Each diagonal element is calculated by the following formula: S404. Degree matrix D of the hypergraph h is also a diagonal matrix represents the set of real numbers and its dimension is |E h | × |E h |, and each diagonal element is calculated by the following formula: D h (v,v) = ∑ v∈V H (v,e) ; S405. Perform hypergraph convolution on the weak inter-procedural control flow graph, which is calculated by the following formula: Among them, H T represents the transpose of H. Let The above formula can be simplified to: where k represents k - times of convolution on the hyper - graph, and X k represents the set of node feature representations. When k = 0, X 0 = X simple ; The node feature representation after the above - mentioned multi - level graph convolution update is X'.

4. The method for inter-procedural vulnerability detection of source code based on hypergraph convolution according to claim 1, wherein: The specific content of S5 includes: S501. Read out the embedding of the weak inter-procedural control flow graph updated by the convolution operation in step S4 as the feature representation of the entire graph, and select the maximum value reading method according to the corresponding dimension, which is specifically calculated by the following formula: r = MAX({x ′ 1, x ′ 2, …, x i ′}), x i ′ ∈ X ′ , where r represents the characteristic representation of the entire inter - weak - procedure control - flow graph; S502. Input the obtained feature representation r into a multi-layer fully connected network for detection, which is specifically calculated by the following formula: In the formula, represents the final detection result, MLP represents a multi-layer fully connected network, and the sigmoid function is used to output the final detection result.

5. The method for detecting inter - procedural vulnerabilities in source code based on hypergraph convolution according to claim 1, wherein: In step S1, use the tool Infer based on separation logic to perform preliminary inter-procedural analysis on the source code to be detected, and based on the results output by Infer, preliminarily screen out potential vulnerabilities and preliminarily locate the scope involved in the vulnerabilities.

6. An inter - procedural vulnerability detection device for source code based on hypergraph convolution, characterized in that The device includes: A potential vulnerability location module, which uses a tool based on separation logic to perform inter-procedural analysis on the source code to be detected, and uses a tool based on separation logic to preliminarily locate the suspicious vulnerabilities existing in the source code to be detected; A reconstructed weak inter-procedural control flow graph module, which reconstructs the weak inter-procedural control flow graph according to the Trace of the code corresponding to the suspicious vulnerability preliminarily located by the potential vulnerability location module; An initial feature representation module, which is used to extract the initial semantic information of the node code, vectorize the syntactic information of the node code, concatenate the initial semantic information and syntactic information of the node code, and form the vectorized representation of the initialized weak inter-procedural control flow graph node feature information; A multi-stage convolution module, which is used to perform simple graph convolution operations and hypergraph convolution operations on the initialized weak inter-procedural control flow graph in sequence to achieve fine-grained capture of multi-level information in the weak inter-procedural control flow graph; A read-out detection module, which reads out the embedding of the weak inter-procedural control flow graph updated by the multi-stage convolution module as the feature representation of the entire graph, and then inputs the obtained feature representation into a detector constructed by a multi-layer fully connected network for vulnerability detection; The specific content of the reconstructed weak inter-procedural control flow graph includes: According to the Trace of the code corresponding to the suspicious vulnerability preliminarily located by the tool based on separation logic, locate each execution process involved in the code and reconstruct the control flow graph of each process; According to the code execution order in the Trace, determine the inter-procedural call points, and then add two in-edge and out-edge connections at each call point according to the determined call point positions to connect the control flow graphs of multiple processes to form a weak inter-procedural control flow graph.

Citation Information

Patent Citations

  • Vulnerability detection method and system based on self-supervised learning and multichannel hypergraph neural network

    CN113609488A

  • Vulnerability detection method and device based on code heterogeneous intermediate graph representation

    CN113868650A