Defect code detection method based on text modality and graph modality and related device
By extracting statement-level source code and abstract code attribute graphs, and interacting with the vector representation of the code sequence to generate multimodal embeddings, the problem of low detection accuracy of defect codes in the prior art is solved, and more accurate graph structure representation and higher detection accuracy are achieved.
Patent Information
- Application Number
- CN202510342480.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-21
- Publication Date
- 2025-06-20
AI Technical Summary
Existing defect code detection methods cannot effectively align the feature representations of code text and graph structures, resulting in insufficient complementarity between semantics and structural information, and it is difficult to capture the multimodal dependence characteristics of defects.
By obtaining the source code and performing abstract processing and arrangement processing, the statement-level source code attribute diagram and abstract code attribute diagram are extracted, combined with the code sequence and vector representation of word tags, and interactively fuse to generate multimodal embeddings, which are used to train defective code classifiers.
It effectively solves the problem of low accuracy of defect code detection. By clearly learning the structural information of abstract code, the graph structure representation ability is enhanced and the impact of noise labels on model training is reduced.
Smart Images

Figure CN120179528A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of defect code detection, and particularly relates to a defect code detection method and related device based on text modality and graph modality. Background Art
[0002] In recent years, software defect detection technology based on deep learning has developed rapidly. Among them, pre-trained language models (PLMs) and graph neural networks (GNNs) have become two mainstream methods, but the limitations of their technologies restrict the actual application effects.
[0003] Currently, most of the methods for defect code detection in the existing technologies are as follows: 1) Implement defect code detection by means of a text semantic analysis method based on PLM; 2) Implement defect code detection by means of a graph structure analysis method based on GNN; 3) Implement defect code detection by adopting a PLM-GNN hybrid method.
[0004] The text semantic analysis method based on PLM is represented by models such as CodeBERT and GraphCodeBERT. The code is regarded as a text sequence, and defect detection is realized through word segmentation, embedding mapping, and a binary classifier. This method is good at capturing the context semantic features of the code (such as variable naming and function call logic), and has a remarkable detection effect on semantic type defects (such as hard-coded credentials and logical errors). However, the lack of structural information in this method leads to the neglect of structured representations such as the control flow graph (CFG) and program dependence graph (PDG) of the code, and it is difficult to detect defects that depend on the execution path (such as buffer overflows and resource competitions). The attention mechanism of the Transformer architecture has low processing efficiency for ultra-long code files and cannot model the non-local dependence relationship between code blocks.
[0005] The graph structure analysis method based on GNN constructs a graph structure by using a code property graph (CPG), aggregates node features (such as operation types and variable attributes) through GNN, and extracts structural patterns for classification. This method is sensitive to structural type defects (such as memory leaks and null pointer references), and performs particularly well in scenarios where the code logic branches are complex. However, this method has the problem of insufficient semantic understanding. The GNN node features usually only contain syntactic types or simple attributes (such as operator types), lacking in-depth parsing of code semantics (such as function intentions and variable scopes). There is a problem of dependence on static analysis tools. The construction of the code graph depends on static analysis tools (such as Joern), and the parsing error of the tool will cause the graph structure to be distorted, thereby affecting the robustness of the model.
[0006] The PLM-GNN hybrid method concatenates the text embeddings generated by the PLM with the graph features extracted by the GNN, or achieves multimodal fusion through one-way embedding transfer (such as initializing the GNN node features with the PLM). This method theoretically takes into account both semantic and structural information, and has better accuracy than single-modal methods in some scenarios. However, this method has the problem of inefficient modal interaction. It uses feature concatenation or simple attention mechanisms, resulting in insufficient alignment of text and graph features in the vector space, being unable to capture fine-grained cross-modal associations (such as the dynamic matching between code lines and corresponding graph nodes), and having a high training complexity. When jointly training the PLM and the GNN, a large amount of computing resources are required, and gradient conflicts between modalities easily lead to difficulties in model convergence.
[0007] In summary, the existing defect code detection methods cannot effectively align the feature representations of code text and graph structures, resulting in insufficient complementarity of semantic and structural information and difficulty in capturing the multimodal dependence characteristics of defects.
[0008] Currently, the existing defect detection datasets generally have problems such as annotation errors or subjective biases, and the existing defect code detection methods (text semantic analysis methods based on PLM, graph structure analysis methods based on GNN, and PLM-GNN hybrid methods) are prone to overfitting on noisy data, which in turn leads to a significant decline in generalization performance, resulting in low accuracy of defect code detection. Summary of the Invention
[0009] The purpose of the present invention is to provide a defect code detection method and related device based on text modality and graph modality to solve the problem of low accuracy in defect code detection in the prior art.
[0010] To achieve the above purpose, the present invention adopts the following technical solutions: In the first aspect, the present invention provides a defect code detection method based on text modality and graph modality, including the following steps: Obtain the source code, and perform abstraction processing and arrangement processing on the obtained source code respectively to obtain abstract code and code sequences; Extract statements from the source code and the abstract code respectively to obtain a statement-level source code attribute graph and a statement-level abstract code attribute graph; Perform word segmentation processing on the code sequences to obtain a list of integer indices of several word tokens, and perform encoding processing on the list of integer indices of several word tokens to obtain vector representations of several word tokens; Perform word segmentation processing on the statement-level source code attribute graph and the statement-level abstract code attribute graph respectively to obtain a list of integer indices of word tokens of nodes in the statement-level source code attribute graph and the statement-level abstract code attribute graph; Encode the list of integer indices of word tokens of nodes in the statement-level source code property graph and the statement-level abstract code property graph to obtain the vector representations of the statement-level source code property graph and the statement-level abstract code property graph; Interact and fuse the vector representations of the statement-level source code property graph and the statement-level abstract code property graph with the vector representations of several word tokens to obtain a multimodal embedding; Use the multimodal embedding to train a defective code classifier and score the training result, and select the defective code classifier with the highest score as the final defective code detection model.
[0011] A further improvement of the present invention lies in that, in the step of respectively performing statement extraction on the source code and the abstract code to obtain the statement-level source code property graph and the statement-level abstract code property graph, the specific steps of statement extraction are as follows: Perform static analysis on the source code and the abstract code respectively to obtain a function-level source code property graph and a function-level abstract code property graph; Perform statement extraction on the function-level source code property graph and the function-level abstract code property graph respectively to obtain the statement-level source code property graph and the statement-level abstract code property graph.
[0012] A further improvement of the present invention lies in that, in the step of respectively performing static analysis on the source code and the abstract code to obtain a function-level source code property graph and a function-level abstract code property graph, specifically use the static analysis tool Joern to perform static analysis on the source code and the abstract code respectively to obtain a function-level source code property graph and a function-level abstract code property graph.
[0013] A further improvement of the present invention lies in that, in the step of encoding the list of integer indices of word tokens of nodes in the statement-level source code property graph and the statement-level abstract code property graph to obtain the vector representations of the statement-level source code property graph and the statement-level abstract code property graph, the specific steps of encoding are as follows: Perform node encoding on the list of integer indices of word tokens of nodes in the statement-level source code property graph and the statement-level abstract code property graph to obtain the node vector representations of the statement-level source code property graph and the statement-level abstract code property graph; Perform graph encoding on the node vector representations of the statement-level source code property graph and the statement-level abstract code property graph to obtain the vector representations of the statement-level source code property graph and the statement-level abstract code property graph.
[0014] A further improvement of the present invention lies in that, in the step of interacting and fusing the vector representations of the statement-level source code attribute graph and the statement-level abstract code attribute graph and the vector representations of several words to obtain a multi-modal embedding, specifically, a CTG-Former modality mixer is used to interact and fuse the vector representations of the statement-level source code attribute graph, the statement-level abstract code attribute graph, and the vector representations of several word tokens.
[0015] A further improvement of the present invention lies in that, before the step of interacting and fusing the vector representations of the statement-level source code attribute graph, the statement-level abstract code attribute graph, and the vector representations of several word tokens to obtain a multi-modal embedding, the vector representations of the statement-level source code attribute graph and the statement-level abstract code attribute graph are also trained to obtain a more accurate multi-modal embedding.
[0016] A further improvement of the present invention lies in that the defective code classifier is a TextCNN model.
[0017] In a second aspect, the present invention provides a defective code detection system based on text modality and graph modality, including a data acquisition module, a statement extraction module, a code sequence tokenization processing module, a statement-level code attribute graph tokenization processing module, a statement-level code attribute graph encoding processing module, a vector interaction and fusion module, and a training module; The data acquisition module is used to acquire source code, and respectively perform abstraction processing and arrangement processing on the acquired source code to obtain abstract code and a code sequence; The statement extraction module is used to respectively extract statements from the source code and the abstract code to obtain a statement-level source code attribute graph and a statement-level abstract code attribute graph; The code sequence tokenization processing module is used to perform tokenization processing on the code sequence to obtain a list of integer indexes of several word tokens, and perform encoding processing on the list of integer indexes of several word tokens to obtain vector representations of several word tokens; The statement-level code attribute graph tokenization processing module respectively performs tokenization processing on the statement-level source code attribute graph and the statement-level abstract code attribute graph to obtain a list of integer indexes of word tokens of nodes under the statement-level source code attribute graph and the statement-level abstract code attribute graph; The statement-level code attribute graph encoding processing module is used to perform encoding processing on the list of integer indexes of word tokens of nodes under the statement-level source code attribute graph and the statement-level abstract code attribute graph to obtain vector representations of the statement-level source code attribute graph and the statement-level abstract code attribute graph; The vector interaction and fusion module is used to interact and fuse the vector representations of the statement-level source code attribute graph, the statement-level abstract code attribute graph, and the vector representations of several word tokens to obtain a multi-modal embedding; The training module is used to train the defect code classifier using multimodal embeddings, score the training results, and select the defect code classifier with the highest score as the final defect code detection model.
[0018] In a third aspect, the present invention provides an electronic device, including a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the steps of the defect code detection method based on text modality and graph modality introduced above are implemented.
[0019] In a fourth aspect, the present invention provides a storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the defect code detection method based on text modality and graph modality introduced above are implemented.
[0020] Compared with the prior art, the present invention has the following beneficial effects: The present invention belongs to an improved invention. Compared with the existing defect code detection methods based on code text or code attribute graphs, on the one hand, the present invention extracts statements from source code and abstract code respectively to obtain a statement-level source code attribute graph and a statement-level abstract code attribute graph. By eliminating code semantic interference, the statement-level source code attribute graph and the statement-level abstract code attribute graph can more clearly learn the structural information of the abstract code, thereby providing a more accurate graph structure representation for the model. On the other hand, the present invention encodes the integer index list of word tags of nodes in the statement-level source code attribute graph and the statement-level abstract code attribute graph, which can enhance the representation ability of the statement-level source code attribute graph and the statement-level abstract code attribute graph, and can also reduce the impact of noisy labels on the subsequent model training, thereby effectively solving the problem of low accuracy in defect code detection in the prior art. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] Figure 1 is a flowchart of the defect code detection method based on text modality and graph modality of the present invention; Figure 2 is a schematic diagram of the defect code detection system based on text modality and graph modality of the present invention; Figure 3 is a flowchart of the defect code detection method based on text modality and graph modality in Embodiment 3 of the present invention; Figure 4 is a schematic diagram of function-level source code, function-level abstract code, and function-level source code attribute graph of the present invention; Figure 5 is a schematic diagram of the statement-level source code attribute graph and the statement-level abstract code attribute graph of the present invention; Figure 6 is a network structure diagram of the defect code classifier of the present invention; Figure 7Schematic diagram of the structure of the electronic device of the present invention. Detailed implementation manners
[0022] To further understand the content of the present invention, the present invention will be described in detail below with reference to the accompanying drawings and specific embodiments. It should be understood that the embodiments are only for explaining the present invention rather than limiting it.
[0023] The defect code detection method based on text modality and graph modality proposed by the present invention interacts and fuses the vector representations of the statement-level source code attribute graph and the statement-level abstract code attribute graph and the vector representations of several word tokens to obtain a multi-modal embedding. The multi-modal embedding is used to train a defect code classifier, and the training result is scored, and the defect code classifier with the highest score is selected as the final defect code detection model. Compared with the prior art, the present invention effectively solves the problem of low accuracy in defect code detection in the prior art.
[0024] Embodiment 1: The flowchart of the defect code detection method based on text modality and graph modality of the present invention is as Figure 1 shown. The defect code detection method based on text modality and graph modality of the present invention includes the following steps: S1. Obtain the source code, and perform abstract processing and arrangement processing on the obtained source code respectively to obtain abstract code and code sequence.
[0025] S2. Extract statements from the source code and the abstract code respectively to obtain a statement-level source code attribute graph and a statement-level abstract code attribute graph.
[0026] S3. Perform word segmentation on the code sequence to obtain a list of integer indices of several word tokens, and perform encoding processing on the list of integer indices of several word tokens to obtain vector representations of several word tokens.
[0027] S4. Perform word segmentation on the statement-level source code attribute graph and the statement-level abstract code attribute graph respectively to obtain a list of integer indices of word tokens of nodes in the statement-level source code attribute graph and the statement-level abstract code attribute graph.
[0028] S5. Perform encoding processing on the list of integer indices of word tokens of nodes in the statement-level source code attribute graph and the statement-level abstract code attribute graph to obtain vector representations of the statement-level source code attribute graph and the statement-level abstract code attribute graph.
[0029] S6. Interact and fuse the vector representations of the statement-level source code attribute graph and the statement-level abstract code attribute graph and the vector representations of several word tokens to obtain a multi-modal embedding.
[0030] S7. Train the defect code classifier using multi-modal embeddings, score the training results, and select the defect code classifier with the highest score as the final defect code detection model.
[0031] Example 2: The schematic diagram of the defect code detection system based on text modality and graph modality of the present invention is as Figure 2 shown. The defect code detection system based on text modality and graph modality of the present invention includes a data acquisition module, a statement extraction module, a code sequence word segmentation processing module, a statement-level code attribute graph word segmentation processing module, a statement-level code attribute graph encoding processing module, a vector interaction and fusion module, and a training module.
[0032] Among them, the data acquisition module is used to acquire the source code, and perform abstraction processing and arrangement processing on the acquired source code respectively to obtain the abstract code and the code sequence.
[0033] The statement extraction module is used to extract statements from the source code and the abstract code respectively to obtain the statement-level source code attribute graph and the statement-level abstract code attribute graph.
[0034] The code sequence word segmentation processing module is used to perform word segmentation processing on the code sequence to obtain a list of integer indexes of several word tokens, and perform encoding processing on the list of integer indexes of several word tokens to obtain vector representations of several word tokens.
[0035] The statement-level code attribute graph word segmentation processing module performs word segmentation processing on the statement-level source code attribute graph and the statement-level abstract code attribute graph respectively to obtain a list of integer indexes of word tokens of nodes under the statement-level source code attribute graph and the statement-level abstract code attribute graph.
[0036] The statement-level code attribute graph encoding processing module is used to perform encoding processing on the list of integer indexes of word tokens of nodes under the statement-level source code attribute graph and the statement-level abstract code attribute graph to obtain vector representations of the statement-level source code attribute graph and the statement-level abstract code attribute graph.
[0037] The vector interaction and fusion module is used to interact and fuse the vector representations of the statement-level source code attribute graph and the statement-level abstract code attribute graph and the vector representations of several word tokens to obtain multi-modal embeddings.
[0038] The training module is used to train the defect code classifier using multi-modal embeddings, score the training results, and select the defect code classifier with the highest score as the final defect code detection model.
[0039] Example 3: The flowchart of the defect code detection method based on text modality and graph modality of the present invention is as Figure 3As shown in the figure, the defect code detection method based on text modality and graph modality of the present invention includes the following steps: S1. Obtain the source code, and perform abstraction processing and arrangement processing on the obtained source code respectively to obtain abstract code and a code sequence.
[0040] First, obtain the source code, and perform abstraction processing and arrangement processing on the obtained source code respectively to obtain abstract code and a code sequence.
[0041] In this embodiment, the construction of the abstract code requires first deleting the comments, redundant spaces and blank lines in the code, as well as the corresponding functions with syntax errors, to obtain a clean version, called the function-level source code.
[0042] Based on the work of predecessors, in this embodiment, the source code is abstracted by mapping each variable and function name to a symbolic representation in the format of "TYPE#". Among them, "TYPE" includes: "VAR" represents a variable, "FUNC" represents a function name, and "#" is the sequential number assigned to each unique instance. For example, the first variable in a function is replaced with "VAR1", and the third unique function name is standardized as "FUNC3". Identifiers that appear multiple times in the same function are always replaced with the same placeholder. These abstractions help to remove elements irrelevant to the code structure features, such as variable names and string contents, thereby reducing the noise that may be introduced when encoding the code structure.
[0043] S2. Perform statement extraction on the source code and the abstract code respectively to obtain a statement-level source code attribute graph and a statement-level abstract code attribute graph.
[0044] In this step, when performing statement extraction on the source code and the abstract code respectively to obtain a statement-level source code attribute graph and a statement-level abstract code attribute graph, the specific steps of statement extraction are as follows: Perform static analysis on the source code and the abstract code respectively to obtain a function-level source code attribute graph and a function-level abstract code attribute graph; Perform statement extraction on the function-level source code attribute graph and the function-level abstract code attribute graph respectively to obtain a statement-level source code attribute graph and a statement-level abstract code attribute graph.
[0045] In this step, the static analysis tool Joern is specifically used to perform static analysis on the source code and the abstract code respectively to obtain a function-level source code attribute graph and a function-level abstract code attribute graph. The schematic diagrams of the function-level source code, the function-level abstract code, and the function-level source code attribute graph in this step are as Figure 4 shown. Figure 4 (a) is the function-level source code, Figure 4 (a) is the function-level abstract code, Figure 4 (b) is the function-level source code attribute graph.
[0046] The code property graph (abstract code property graph and source code property graph) consists of three interconnected subgraphs (Abstract Syntax Tree (AST), Control Flow Graph (CFG), and Program Dependence Graph (PDG)).
[0047] This embodiment uses the code property graph as the analysis basis mainly due to its unique advantages. First, it can provide a unified code representation framework, integrating syntax, control, and data dependence information in one graph. Second, this representation method retains rich semantic information, facilitating in-depth static analysis. However, the code property graph at the function level often has a high complexity. Especially when encoding it into a vector representation, it will significantly increase the computational overhead, which is particularly obvious in large-scale code analysis scenarios.
[0048] The process of statement extraction is described in detail below: Statement extraction is performed on the source code property graph and the abstract code property graph respectively to obtain the statement-level source code property graph and the statement-level abstract code property graph (the statement-level source code property graph and the statement-level abstract code property graph, also called statement-level subgraphs), which can reduce computational complexity while retaining meaningful context information. The extraction criteria for the statement-level source code property graph and the statement-level abstract code property graph are determined by the statement context length in the PDG, and the PDG is a subgraph of the code property graph. Different from MGVD, in this embodiment, the context length is defined as the maximum path length between the connected statements and The statement-level subgraphs are generated in the order in which the statements appear in the function to ensure the consistency of the context length. For example, consider a statement in a function-level code property graph. If the context length is set to 2, and the maximum path lengths from the previous statement and the next statement to the statement are both 2, then the path from to is given priority. Therefore, the statement-level subgraph of the statement not only includes its nodes and edges but also the statement nodes and edges in its context. If the context length exceeds 2, there may be duplicate nodes and edges in the statement-level subgraph. In the code of Figure 4 , setting the context length to 3 will result in the same statement-level subgraph for the statement "int y = 2 * x" and the statement "int x = source()". Therefore, in this embodiment, the context length is set to 2.
[0049] Figure 5 shows the source code and the abstract code property graph of the statement "int y = 2 * x", where Figure 5 (a)is the statement-level source code property graph, Figure 5 (b)is the statement-level abstract code property graph. This method reduces the computational complexity while enhancing the representation of a single statement. Finally, this method generates two ordered sets of statement-level code property graphs, one corresponding to the source code and the other corresponding to the abstract code.
[0050] S3. Tokenize the code sequence to obtain a list of integer indices of several word tokens, and encode the list of integer indices of several word tokens to obtain vector representations of several words.
[0051] Tokenize the code sequence (using the BPE tokenizer) to obtain a list of integer indices of several word tokens, and encode the list of integer indices of several word tokens (text encoding) to obtain vector representations of several word tokens.
[0052] The choice of the BPE tokenizer is mainly based on its advantages in dealing with out-of-vocabulary (OOV) problems and its wide applications in fields such as natural language processing and code analysis.
[0053] BPE is a statistical subword tokenization method. The core idea is to build a vocabulary by iteratively merging the most frequently occurring character pairs. Specifically, BPE decomposes words into a series of subwords or character n-grams, where n-gram refers to a sequence of n consecutive characters or words extracted from the text. This method can not only effectively handle new words outside the vocabulary but also significantly reduce the size of the vocabulary required by the language model, thus achieving a balance between model performance and computational efficiency. The tokenization process of BPE starts with an initial tokenization at the character level and then constructs new lexical units by iteratively merging the most frequently occurring character pairs. This merging process continues until a predetermined number of merges or the vocabulary size is reached. In this way, BPE can control the size of the vocabulary while retaining lexical details. For example, Figure 5The code statement "int y = 2 * x;" shown in will be converted into the following token sequence after BPE tokenization: "[int, y, =, 2, *, x,;]". This tokenization method not only preserves the semantic information of the code but also effectively processes various identifiers and operators that may appear in the code. It is worth noting that the BPE tokenizer has unique advantages when processing programming language texts. Since programming languages usually contain a large number of domain-specific terms and symbols (such as variable names, function names, and operators), traditional space-based tokenization methods often struggle to effectively handle these special tokens. However, BPE, through its adaptive subword segmentation mechanism, can better capture the syntactic and semantic features in programming languages, thereby improving the quality of subsequent code representation learning.
[0054] S4. Tokenize the statement-level source code property graph and the statement-level abstract code property graph respectively to obtain the integer index lists of the word tokens of the nodes under the statement-level source code property graph and the statement-level abstract code property graph.
[0055] Tokenize the source code statement-level subgraph and the abstract code statement-level subgraph respectively (using the BPE tokenizer) to obtain the integer index lists of the word tokens of the nodes under the source code statement-level subgraph and the abstract code statement-level subgraph.
[0056] S5. Encode the integer index lists of the word tokens of the nodes under the statement-level source code property graph and the statement-level abstract code property graph to obtain the vector representations of the statement-level source code property graph and the statement-level abstract code property graph.
[0057] The specific encoding process in this step is as follows: A. Encode the integer index lists of the word tokens of the nodes under the statement-level source code property graph and the statement-level abstract code property graph to obtain the node vector representations of the statement-level source code property graph and the statement-level abstract code property graph; B. Encode the node vector representations of the statement-level source code property graph and the statement-level abstract code property graph to obtain the vector representations of the statement-level source code property graph and the statement-level abstract code property graph.
[0058] The detailed description of the node encoding in step A is as follows: To convert the node code text in the code graph into learnable embedding vectors, existing research mainly adopts general text embedding methods (such as Word2Vec and Sent2Vec). However, these methods have significant limitations in code representation learning. Firstly, they cannot effectively handle the syntax structures and semantic patterns unique to programming languages. Secondly, the distribution difference between the pre-trained corpus and the code domain leads to limited generalization ability. To address these problems, this embodiment adopts the pre-trained language model CodeBERT designed specifically for code understanding as the basic framework for node embedding. This model has acquired rich prior knowledge of programming languages through dual-task pre-training of masked language modeling and sentence relationship prediction on a large-scale code corpus. Specifically, first, a dual-channel embedding mechanism is constructed, and the pre-trained weights of CodeBERT are directly transferred to the initialization parameters of the self-embedding layer. This knowledge transfer strategy enables each code token (such as variable names, operators, and keywords) to obtain an initial vector representation containing syntax and semantics. Secondly, for the hierarchical structure characteristics of the code property graph, type-aware embedding is designed. Through a learnable label encoding layer, a type embedding vector is generated for each node (such as AST node type and control flow type), and feature concatenation is performed with the content embedding. This combined embedding strategy explicitly encodes the structured information of the program while retaining the code semantics. Based on the existing method, redundant code attributes are removed from the non-leaf nodes in the code property graph. This is because these attributes are usually already encoded in the corresponding leaf nodes, so it is unnecessary to include them at a higher level. This simplification reduces redundancy and optimizes the embedding process. Finally, two key components are combined to create a comprehensive initialization representation for each node. The node content embedding vector (derived from the pre-training and fine-tuning representations of CodeBERT) and the node type embedding vector (generated through label encoding). This combined representation ensures a detailed encoding of the semantic and structural attributes of each node, laying the foundation for subsequent tasks.
[0059] The detailed description of the graph encoding in step B is as follows: For the statement-level subgraph, first, the graph attention network (GAT) is used to encode each graph and calculate the feature vector of each graph node. GAT uses the attention mechanism to evaluate the importance of neighbor nodes when updating node features, thereby achieving more detailed feature updates. This ability can improve performance in tasks dealing with graph-structured data. Finally, the node features are aggregated through average pooling to obtain the feature vector of each graph.
[0060] Specifically, the update function and aggregation function of the graph neural network use a dynamic attention mechanism to assign different importance to neighbor nodes. Node at the layer representation is calculated as follows: (1) where, represents the set of neighbor nodes of node , represents the learnable weight matrix applied to the input features of neighbor nodes, represents the non-linear activation function, such as ReLU, represents the attention coefficient, which is used to quantify the importance of neighbor to node . The attention coefficient is calculated using a shared attention mechanism, and the calculation formula is as follows: (2) where, represents the learnable attention vector, represents the concatenation operator, is the activation function used to calculate the attention score. This mechanism enables GAT to dynamically focus on the most relevant neighbors of each node during the message passing process.
[0061] After updating the node representation, a readout operation aggregates the node-level representations into a single graph-level representation. The graph-level representation is calculated by the READOUT function, which for GAT is a complex global weighted operation based on the attention mechanism: (3) where, represents the attention weight, which is defined based on the global context as follows: (4) where, represents the learnable query vector. The attention mechanism dynamically determines the contribution importance of each node when generating the graph-level representation. This method ensures that the graph representation can effectively capture the overall semantic and structural information of the graph. Finally, two sets of graph embeddings are obtained, the statement-level source graph set and the statement-level abstract graph set : (5) (6) where, represents the number of statements in the function-level source code. and They respectively represent the final outputs of the statement-level source graph and the abstract graph after being processed by GAT.
[0062] S6. Interact and fuse the vector representations of the statement-level source code property graph and the statement-level abstract code property graph, and the vector representations of several word tokens to obtain a multimodal embedding.
[0063] In this step, the CTG-Former modality mixer is specifically used to interact and fuse the vector representations of the source code statement-level subgraph and the abstract code statement-level subgraph, and the vector representations of several word tokens.
[0064] The following is a detailed description of this process: CTG-Former is a modality mixer that fuses code data. Different from existing vision-language models (such as the Q-Former of BLIP2), this embodiment enables the translation task from the code graph modality to the code text modality. To better adapt to the source code analysis scenario, this embodiment uses the CodeBERT model specifically tailored for code text to replace the BERT model and enhances its cross-modal attention mechanism to process graph and text modalities.
[0065] The cross-attention mechanism is the core component of CTG-Former, which can align the code text modality with the code graph modality, including the source code property graph and the abstract code property graph. This mechanism can capture the relationships between cross-modal data through multi-head attention operations: (7) (8) Among them, represents Queries, represents Queries from the text modality, represents keys, values, and a learnable weight matrix for projecting the input into an appropriate space. represents the graph or text modality. For each modality, CTG-Former calculates a set of independent cross-attention weights to align the modality to the source modality. This ensures that every detail of the statement-level graph can be accurately mapped to its corresponding text representation. The calculation formula of the attention mechanism itself is: (9) Among them, represents the dimension of Queries and keys.
[0066] In this embodiment, the cross-attention mechanism is adopted, and CTG-Former can effectively capture and encode the complex relationships and dependencies between different modalities. This process emphasizes the alignment between the statement text and multiple code graph modalities, especially with the modality.
[0067] Before this step, the vector representations of the source code statement-level subgraph and the abstract code statement-level subgraph are also trained using the contrastive learning algorithm to obtain more accurate multi-modal embeddings. The following is a detailed description of this process: To effectively align the embedding representations of the source statement graph and its corresponding abstract statement graph while differentiating them from irrelevant graphs, this embodiment adopts the InfoNCE loss function based on the contrastive learning algorithm. The core idea of this loss function is to maximize the similarity of the positive sample pairs (i.e., the statement-level source graph and its corresponding abstract-level statement graph) while minimizing the similarity of the negative sample pairs (i.e., the statement-level source graph and other irrelevant abstract-level statement graphs), thereby learning more discriminative embedding representations. To provide a more stable code graph encoder for the subsequent cross-modal interaction module, first, the loss function is used to train the first stage of the pre-training task. In this process, the graph encoder is trained to a state where it can stably extract graph features using contrastive learning, and the parameters are frozen during the subsequent cross-modal learning process. The mathematical form of the loss function is defined as follows: (10) where, represents the total number of statement graphs in the current batch, and respectively represent the embedding vectors of the th statement-level source code graph and its corresponding statement-level abstract code graph. The similarity metric function uses cosine similarity for calculation, and the calculation formula is: (11) where, represents the L2 norm of the vector, is used to adjust the sharpness of the similarity distribution. A smaller will make the distribution sharper, thereby strengthening the attention to difficult negative samples, while a larger will make the distribution smoother and enhance the generalization ability to the overall samples. It can be seen from the design of the loss function that the numerator part maximizes the similarity between the positive sample pairs (i.e., and ) through exponentiation and normalization operations, thereby prompting the model to learn embedding representations that can accurately reflect the semantic consistency between the source statement graph and its corresponding abstract statement graph. At the same time, the denominator part sums over all negative sample pairs (i.e., and , where the similarities of ) are summed up, forcing the model to make the embedding representations between unrelated graph pairs as far apart as possible, thereby enhancing the discriminative ability of the embedding space. This mechanism of contrasting positive and negative samples enables the model to effectively distinguish between relevant and unrelated graph pairs, and thus improve the quality of the embedding representation.
[0068] During the pre-training stage, the text encoder is kept fixed and the graph encoder is allowed to adaptively adjust. CTG-Former plays a key role in aligning the two modalities with the text through a self-supervised learning process, thus improving the conversion from the code graph to the code text. A cross-attention mechanism is also introduced during the pre-training stage to promote cross-modal interaction and contrastive learning.
[0069] The task of aligning the source modality with the corresponding text is modeled as a binary classification problem to judge whether a given set of cross-modal texts match. To achieve effective interaction between different embeddings, self-attention mechanisms (graph cross-attention mechanism and graph-text cross-attention mechanism) are used. Initially, the Queries are represented as zero tensors of fixed length. Then, the Queries are updated through cross-attention with the source modality embeddings, thereby generating Queries embeddings rich in multi-modal information. These embeddings are passed through a linear classifier to generate logits values, and the final matching score is obtained by averaging the logits of all Queries.
[0070] For both the graph modality and the text modality, first generate embeddings for the two modalities ( and ), and create matching and non-matching pairs to provide contrast examples for the model. This helps the model learn to identify the correct pairings. The relationships between modalities are captured through self-attention and cross-attention mechanisms, and the calculation formula is as follows: (12) where represents the attention function that controls self-attention and cross-attention.
[0071] After obtaining the fused embeddings, apply the linear transformation weight to calculate the prediction score, indicating whether the pair matches (i.e., the actual label is 1) or does not match (i.e., the actual label is 0). The performance of the model is evaluated through the following loss function, and the calculation formula is: (13) Contrastive learning algorithms are used to align different types of information with their corresponding text representations by contrasting the similarities between matching and non-matching pairs. The source modality representation is aligned with the text representation, and the pairs showing the highest similarity are considered correct cross-modal pairs.
[0072] First, CTG-Former is used to obtain the learned representation of the input modality. Text modality features are extracted from this representation. . Then, the mutual information between the text modality features and the graph features is calculated, and the calculation formula is as follows: (14) This mutual information serves as a similarity measure for the source modality-text pairs. Finally, the loss function is calculated using cross-entropy , and the calculation formula is as follows: (15) This loss function encourages the model to correctly align the cross-modal data with its corresponding text representation, thus promoting better learning through the contrastive learning method.
[0073] To integrate the two loss components and into a unified framework, the loss function of the pre-training task is defined as the joint optimization form of these terms. Specifically, the final loss function is expressed as: (16) where, represents the cross-entropy loss function, ensuring the alignment of the text embedding vector and the graph embedding vector through the interaction function , is the text-graph matching loss introduced through self-supervised learning.
[0074] The parameter is a tunable hyperparameter used to balance the contributions of the text-graph contrastive loss and the text-graph matching loss . By adjusting , the model can flexibly emphasize the alignment between the text and graph embeddings, or the matching between the text and graph modalities. This balance enables the model to optimize fine-grained cross-modal consistency and high-level semantic alignment, ultimately improving its generalization ability in different tasks.
[0075] S7. Use the multi-modal embeddings to train the defect code classifier and score the training results. Select the defect code classifier with the highest score as the final defect code detection model.
[0076] The defect code classifier in this embodiment is a TextCNN model.
[0077] The following is a detailed description of this process: During the fine-tuning process, the multi-modal embeddings in the pre-trained Vul-CTG model are used as the input to the TextCNN model. The Vul-CTG model performs excellently in cross-modal learning tasks, and its multi-modal embeddings can effectively capture the deep semantic relationships between text and code. By using these embeddings as the initial weights of the TextCNN architecture, a richer and more accurate feature representation can be provided for the defect detection task. As Figure 6 shown, the TextCNN architecture consists of three one-dimensional convolutional layers, and each convolutional layer is followed by a max-pooling operation. The role of the pooling layer is to reduce the feature dimension and highlight important features. Specifically, the first convolutional layer extracts local text information, the second layer further learns mid-level features, and the third layer helps capture more abstract and complex patterns. The features output by the convolutional layers are concatenated and flattened after the pooling operation, which can better retain the global information across convolutional layers. The flattened features are passed to the fully connected layer, and they enhance the expressive power of the model through non-linear activation functions. Finally, an output layer maps the features learned through the deep network to the class space for the final defect detection task.
[0078] To verify the effectiveness of the defect code detection method based on text modality and graph modality proposed in the present invention, this embodiment uses multiple defect data sets containing various defect types for fine-tuning and verification.
[0079] To ensure the comprehensiveness and reliability of the results, this embodiment uses five publicly available defect data sets, including four mature data sets and a newly released comprehensive data set. These data sets contain different types of defect samples and cover defect problems in various actual development scenarios. To compare with different baseline algorithms, these data sets are all divided into an 80% training set, a 10% validation set, and a 10% test set. All algorithms are trained and tuned on the training set and validation set, and finally verified and compared on the test set.
[0080] The hardware platform used in this embodiment is: the CPU is Intel(R) Core(TM) i9-13700KF with a main frequency of 3.4 GHz and 128 GB of memory, and the GPU is NVIDIA A800. The software platform used in this embodiment is the Linux operating system, the Torch 2.2.0 deep learning framework, and Python 3.8.
[0081] When selecting the dataset, the following factors were particularly considered: the distribution of defective and non-defective code samples, the source of the dataset, and the ability of each dataset to distinguish specific vulnerability (also called defect) types. Specifically, mature datasets usually contain a wide range of representative vulnerability samples, which can reflect various common security vulnerability types. The newly released comprehensive datasets, by integrating vulnerabilities from different fields and of different types, strive to cover a wider range of vulnerability categories and more complex vulnerability patterns. The characteristics of each dataset, such as vulnerability distribution, sample diversity, and discrimination ability, have been detailedly summarized in Table 1. This table summarizes the number of vulnerabilities, non-vulnerable functions, code sources, the number of vulnerability categories, and the ratio of the number of vulnerabilities to non-vulnerabilities (imbalance ratio) in five datasets. By fine-tuning these datasets, the adaptability and accuracy of the proposed method in different vulnerability detection scenarios can be comprehensively evaluated, and then its advantages and disadvantages can be compared with other methods. These datasets not only provide diverse training samples for the fine-tuning of the model, but also simulate to a certain extent the complex and changeable code vulnerability situations in the real world, and also provide relatively fair conditions for verifying the practicality and effectiveness of defect detection technologies on these datasets.
[0082] Table 1 Summary of Fine-tuned Datasets
[0083] This embodiment evaluates the performance of the Vul-CTG model (the method proposed in the present invention) and several current state-of-the-art defect detection models based on deep learning. These benchmark models cover a variety of different architectures and technologies, including frozen pre-trained language models, graph neural network-based models, and convolutional neural network-based models, specifically including: Pre-trained language models: such as BERT, CodeBERT, VulBERTa, and PDBERT. These models are pre-trained using large-scale text data, can learn rich language and semantic representations, and have achieved excellent results in multiple NLP tasks. In the defect detection task, these models can effectively capture potential defect features in the code through transfer learning.
[0084] Graph neural network-based models: such as Devign. This model represents the code as a graph structure and processes it using a graph neural network, thereby being able to capture complex patterns of variables, functions, and their dependencies in the code. Graph neural networks are particularly suitable for processing structured data and perform outstandingly in program analysis and defect detection tasks.
[0085] Model based on convolutional neural network: For example, TextCNN. This model effectively extracts local features in the code through convolutional operations and conducts classification. Although convolutional neural networks are mainly applied to image processing, they can still exhibit excellent performance in text analysis, especially in the classification task of code text.
[0086] In terms of model configuration, all baseline models uniformly set the learning rate to 1e-4 and the batch size to 128. The model was trained for 50 epochs and adopted an early stopping strategy to avoid overfitting and ensure the generalization ability of the model on the validation set. When the model performance did not improve further, the training was automatically stopped. For the Devign model, an abstract syntax tree was used as the graph representation of the code. Since the hyperparameters of Devign were not publicly disclosed, this experiment tried to reproduce its original method as much as possible for a fair comparison.
[0087] To comprehensively evaluate the performance of each model in the defective code detection task, this embodiment adopts two key evaluation metrics: accuracy (ACC) and F1 score, where the calculation of the F1 score depends on the calculation of precision (P) and recall (R).
[0088] The accuracy and F1 scores of various defect detection baseline methods on five benchmark datasets are shown in Table 2. The model with the best F1 score in each method was selected for comparison. The method proposed in the present invention (Vul-CTG) performs excellently in most evaluation datasets and evaluation metrics, especially showing its significant advantages in terms of accuracy and F1 score. This result indicates that Vul-CTG outperforms a variety of existing baseline methods in comprehensive performance and demonstrates its strong potential in the defect detection task.
[0089] Table 2 Accuracy and F1 scores of various defect detection baseline methods on five benchmark datasets (%)
[0090] The data in Table 2 show that in the ReVeal dataset, Vul-CTG achieved the best performance with an accuracy of 93.21%, surpassing other powerful baseline models such as CodeBERT (90.11%) and PDBERT (88.92%). In addition, the F1 score of Vul-CTG is 46.45%, which is also the highest among all methods. This result not only reflects its good balance between precision and recall but also indicates its comprehensiveness in dealing with complex defect detection tasks. This balance is particularly important in scenarios where both prediction accuracy and comprehensive coverage need to be considered, especially in practical applications, where false negatives and false positives can both bring serious consequences.
[0091] In the CrossVul dataset, Vul-CTG once again leads other benchmark models with an accuracy of 95.61% and an F1 score of 24.49%. This performance further validates its adaptability and robustness across different datasets. Similarly, in the CVEFixes dataset, Vul-CTG approaches the highest level with an accuracy of 96.05% and an F1 score of 27.07%, demonstrating its stability in diverse defect detection tasks. This result not only proves its superiority on a single dataset but also shows its ability to adapt to different data distributions and task requirements.
[0092] Vul-CTG's performance on the MVD dataset is particularly outstanding, with an accuracy of 99.48%, the highest among all models. Its F1 score of 99.08% also surpasses competitors such as VulBERTa-CNN (98.80%) and PDBERT (98.32%). These results highlight its excellent performance on structured datasets with high consistency. The high accuracy and F1 score of the MVD dataset indicate that Vul-CTG can make full use of the regularity of the data when processing structured data, thus achieving higher detection accuracy.
[0093] In the DiverseVul dataset, Vul-CTG once again sets a new record with an accuracy of 96.11% and an F1 score of 28.70%, significantly outperforming the second-best method. This performance further demonstrates its strong ability in diverse defect detection tasks. The high complexity and diversity of the DiverseVul dataset pose higher requirements for the generalization ability of the model, and Vul-CTG's excellent performance on this dataset indicates its ability to effectively handle diverse defect patterns.
[0094] Example 4: Please refer to Figure 7 As shown, the present invention also provides an electronic device 100 for a defect code detection method based on text modality and graph modality; the electronic device 100 includes a memory 101, at least one processor 102, a computer program 103 stored in the memory 101 and executable on the at least one processor 102, and at least one communication bus 104.
[0095] The memory 101 can be used to store the computer program 103. By running or executing the computer program stored in the memory 101 and invoking the data stored in the memory 101, the processor 102 implements the steps of the defect code detection method based on text modality and graph modality described in Embodiment 1. The memory 101 may mainly include a program storage area and a data storage area. Among them, the program storage area can store an operating system, application programs required for at least one function (such as a sound playback function, an image playback function, etc.); the data storage area can store data created according to the use of the electronic device 100 (such as audio data, etc.). In addition, the memory 101 may include non-volatile memory, such as a hard disk, a memory, a plug-in hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, at least one magnetic disk storage device, a flash memory device, or other non-volatile solid-state storage devices.
[0096] The at least one processor 102 can be a Central Processing Unit (CPU), or can also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The processor 102 can be a microprocessor or the processor 102 can also be any conventional processor, etc. The processor 102 is the control center of the electronic device 100, and connects various parts of the entire electronic device 100 through various interfaces and lines.
[0097] The memory 101 in the electronic device 100 stores a plurality of instructions to implement the defect code detection method based on text modality and graph modality. The processor 102 can execute the plurality of instructions to thereby implement: Obtain the source code, and perform abstraction processing and arrangement processing on the obtained source code respectively to obtain abstract code and a code sequence; Extract statements from the source code and the abstract code respectively to obtain a statement-level source code attribute graph and a statement-level abstract code attribute graph; Perform word segmentation processing on the code sequence to obtain a list of integer indexes of several word tokens, and perform encoding processing on the list of integer indexes of several word tokens to obtain vector representations of several word tokens; Perform word segmentation on the statement-level source code property graph and the statement-level abstract code property graph respectively to obtain the integer index list of word tokens of the nodes under the statement-level source code property graph and the statement-level abstract code property graph; Perform encoding on the integer index list of word tokens of the nodes under the statement-level source code property graph and the statement-level abstract code property graph to obtain the vector representations of the statement-level source code property graph and the statement-level abstract code property graph; Interactively fuse the vector representations of the statement-level source code property graph and the statement-level abstract code property graph with the vector representations of several word tokens to obtain a multimodal embedding; Use the multimodal embedding to train a defective code classifier and score the training results, and select the defective code classifier with the highest score as the final defective code detection model.
[0098] Example 5: If the modules / units integrated in the electronic device 100 are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, to implement all or part of the processes in the above method embodiments of the present invention, it can also be completed by a computer program instructing relevant hardware. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file or some intermediate form, etc. The computer-readable medium can include: any entity or device that can carry the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disc, computer memory and read-only memory (ROM, Read-Only Memory).
[0099] Those skilled in the art should understand that the embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0100] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present invention. It should be understood that each flow and / or block in the flowchart and / or block diagram, and the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing device produce a means for implementing the functions specified in the flow Figure 1 one flow or multiple flows and / or blocks Figure 1 a block or multiple blocks.
[0101] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory produce a manufactured article including an instruction means that implements the functions specified in the flow Figure 1 one flow or multiple flows and / or blocks Figure 1 a block or multiple blocks.
[0102] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, so that the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in the flow Figure 1 one flow or multiple flows and / or blocks Figure 1 a block or multiple blocks.
[0103] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the above embodiments, those of ordinary skill in the art should understand that: modifications or equivalent replacements can still be made to the specific embodiments of the present invention, and any modifications or equivalent replacements that do not depart from the spirit and scope of the present invention should be covered by the protection scope of the claims of the present invention.
Claims
1. A defect code detection method based on text mode and image mode, characterized in that: The following steps are involved: Acquire source code, and perform abstract processing and arrangement processing on the acquired source code to obtain abstract code and code sequence; Perform sentence extraction on the source code and the abstract code respectively to obtain a sentence-level source code attribute graph and a sentence-level abstract code attribute graph; Perform word segmentation on the code sequence to obtain a list of integer indices of several word tags, and perform encoding on the list of integer indices of several word tags to obtain vector representations of several word tags; Perform word segmentation processing on the statement-level source code attribute graph and the statement-level abstract code attribute graph respectively, and obtain integer index lists of word tags of nodes under the statement-level source code attribute graph and the statement-level abstract code attribute graph; Encoding the integer index list of word tags of nodes under the statement-level source code attribute graph and the statement-level abstract code attribute graph to obtain vector representations of the statement-level source code attribute graph and the statement-level abstract code attribute graph; The vector representations of the statement-level source code attribute graph and the statement-level abstract code attribute graph are interactively fused with the vector representations of several word tags to obtain a multimodal embedding. The defect code classifier is trained using multimodal embedding, and the training results are scored. The defect code classifier with the highest score is selected as the final defect code detection model.
2. The defect code detection method based on text modality and image modality according to claim 1 is characterized in that: In the step of extracting statements from the source code and the abstract code respectively to obtain a statement-level source code attribute graph and a statement-level abstract code attribute graph, the specific statement extraction step is: Perform static analysis on the source code and abstract code respectively to obtain a function-level source code attribute graph and a function-level abstract code attribute graph; Statement extraction is performed on the function-level source code attribute graph and the function-level abstract code attribute graph respectively to obtain a statement-level source code attribute graph and a statement-level abstract code attribute graph.
3. The defect code detection method based on text modality and image modality according to claim 2 is characterized in that: In the step of performing static analysis on the source code and the abstract code respectively to obtain a function-level source code property graph and a function-level abstract code property graph, the static analysis tool Joern is specifically used to perform static analysis on the source code and the abstract code respectively to obtain a function-level source code property graph and a function-level abstract code property graph.
4. The defect code detection method based on text modality and image modality according to claim 1 is characterized in that: In the step of encoding the integer index list of word tags of nodes under the statement-level source code attribute graph and the statement-level abstract code attribute graph to obtain the vector representation of the statement-level source code attribute graph and the statement-level abstract code attribute graph, the specific encoding processing steps are: Perform node encoding on the integer index list of word tags of nodes under the statement-level source code attribute graph and the statement-level abstract code attribute graph to obtain node vector representations under the statement-level source code attribute graph and the statement-level abstract code attribute graph; The node vector representations under the statement-level source code attribute graph and the statement-level abstract code attribute graph are graph-encoded to obtain vector representations of the statement-level source code attribute graph and the statement-level abstract code attribute graph.
5. The defect code detection method based on text modality and image modality according to claim 1 is characterized in that: In the step of interactively fusing the vector representations of the statement-level source code attribute graph and the statement-level abstract code attribute graph with the vector representations of several word tags to obtain a multimodal embedding, a CTG-Former modal mixer is specifically used to interactively fuse the vector representations of the statement-level source code attribute graph and the statement-level abstract code attribute graph with the vector representations of several word tags.
6. The defect code detection method based on text modality and image modality according to claim 1 is characterized in that: Before the step of interactively fusing the vector representations of the statement-level source code attribute graph and the statement-level abstract code attribute graph with the vector representations of several word tags to obtain a multimodal embedding, the vector representations of the statement-level source code attribute graph and the statement-level abstract code attribute graph are also trained to obtain a more accurate multimodal embedding.
7. The defect code detection method based on text modality and image modality according to claim 1 is characterized in that: The defect code classifier is a TextCNN model.
8. A defect code detection system based on text mode and image mode, characterized in that: It includes a data acquisition module, a sentence extraction module, a code sequence word segmentation processing module, a sentence-level code attribute graph word segmentation processing module, a sentence-level code attribute graph encoding processing module, a vector interaction fusion module and a training module; The data acquisition module is used to acquire source code, and perform abstract processing and arrangement processing on the acquired source code to obtain abstract code and code sequence; The sentence extraction module is used to extract sentences from the source code and the abstract code respectively, and obtain a sentence-level source code attribute graph and a sentence-level abstract code attribute graph; The code sequence word segmentation processing module is used to perform word segmentation processing on the code sequence to obtain a plurality of integer index lists of word tags, and to perform encoding processing on the integer index lists of the plurality of word tags to obtain vector representations of the plurality of word tags; The sentence-level code attribute graph word segmentation processing module performs word segmentation processing on the sentence-level source code attribute graph and the sentence-level abstract code attribute graph respectively, and obtains an integer index list of word tags of nodes under the sentence-level source code attribute graph and the sentence-level abstract code attribute graph; The statement-level code attribute graph encoding processing module is used to encode the integer index list of word tags of nodes under the statement-level source code attribute graph and the statement-level abstract code attribute graph to obtain vector representations of the statement-level source code attribute graph and the statement-level abstract code attribute graph; The vector interactive fusion module is used to interactively fuse the vector representations of the statement-level source code attribute graph and the statement-level abstract code attribute graph with the vector representations of several word tags to obtain a multimodal embedding; The training module is used to train the defect code classifier using multimodal embedding, score the training results, and select the defect code classifier with the highest score as the final defect code detection model.
9. An electronic device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the defect code detection method based on text modality and graphic modality described in any one of claims 1 to 7 are implemented.
10. A storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the defect code detection method based on text modality and graphic modality described in any one of claims 1 to 7 are implemented.