Intelligent contract vulnerability detection method based on multi-mode selection state space fusion
Through the multimodal selection state space fusion method, the multi-dimensional features of smart contracts are generated and processed, and the accuracy and coverage of existing detection methods are solved, achieving more efficient vulnerability detection.
Patent Information
- Application Number
- CN202510756827.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-09
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2045-06-09
AI Technical Summary
The existing smart contract vulnerability detection methods have problems such as high false positive rate, low computational efficiency, unreliable coverage rate and difficulty in capturing deep semantic information. In particular, deep learning-based methods are limited by a single perspective, resulting in low detection accuracy.
The multimodal selection state space fusion method is adopted to obtain the abstract syntax tree, source code and syntax element types of smart contracts, generate contract graphs, intermediate representation text and image features, and use residual graph convolution networks, residual attention networks and dual-channel convolution neural networks to process these features, perform multiple rounds of feature cross-over and weighted pooling, and finally perform classification processing to detect vulnerabilities.
It improves the accuracy and automation of smart contract vulnerability detection, and can more comprehensively capture the contract's syntax structure, execution logic and underlying semantic information, realizing efficient integration of multimodal features and vulnerability detection across abstract levels.
Smart Images

Figure SMS_22 
Figure SMS_23 
Figure SMS_37
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of electronic digital data processing, and specifically to an intelligent contract vulnerability detection method based on multi-modal selection state space fusion. Background Art
[0002] In order to address the severe security challenges faced by intelligent contracts, researchers and developers have been working on developing effective intelligent contract detection methods. Currently, the methods used to detect intelligent contract vulnerabilities include: pattern matching, symbolic execution, and fuzz testing. The pattern matching method is a method that relies on pre-set rules, which is relatively simple to implement but prone to false positives and false negatives. The symbolic execution method detects vulnerabilities by exhaustively exploring all possible execution paths, but they are usually computationally inefficient and difficult to apply to complex contracts. The fuzz testing method uses random inputs to evaluate contract behavior, but its efficiency and coverage are still unreliable and often require manual intervention for analysis.
[0003] In recent years, deep learning has made breakthrough progress in various fields with its powerful feature learning and pattern recognition capabilities, and intelligent contract vulnerability detection methods based on deep learning have emerged as the times require. Compared with traditional methods, deep learning-based methods can automatically learn complex vulnerability patterns from data, adapt to changing attack techniques, reduce reliance on expert knowledge, and show significant advantages in improving detection accuracy, automation, and handling complex contracts. However, current deep learning-based detection methods are limited by their single perspective and are difficult to capture the deep semantic information contained in intelligent contract code, resulting in relatively low accuracy in intelligent contract detection. Summary of the Invention
[0004] The purpose of the present invention is to provide an intelligent contract vulnerability detection method based on multi-modal selection state space fusion.
[0005] The technical solution of the present invention is as follows: An intelligent contract vulnerability detection method based on multi-modal selection state space fusion includes the following operations: S1. Obtain the abstract syntax tree of the intelligent contract to be detected, add data flow information after simplification processing to obtain a contract graph; convert the abstract syntax tree of the intelligent contract to be detected into a structure, obtain the function operation instruction sequence in the structure to obtain an intermediate representation text; compile the source code of the intelligent contract to be detected, obtain the runtime bytecode generated by compilation, and convert it into a sequence to obtain a bytecode text; respectively perform color-coding mapping processing on different syntax element types in the intelligent contract to be detected to generate a color image; map each bytecode of the intelligent contract to be detected to a grayscale pixel to generate a grayscale image; S2. The contract diagrams are processed by the residual graph convolutional network and the residual attention network respectively to obtain global semantic features and local semantic features, which are mapped to the same shared space to obtain graph modality embedding features. The text context information of the intermediate representation text and the bytecode text is obtained respectively, and feature alignment is performed based on linear projection and batch normalization to obtain text modality embedding features. The color image and the grayscale image are processed by a dual-channel convolutional neural network and then concatenated to obtain visual modality embedding features. S3. The graph modality embedding features, text modality embedding features, and visual modality embedding features are subjected to feature cross-processing based on the selected state space for several rounds to obtain graph modality embedding cross features, text modality embedding cross features, and visual modality embedding cross features, which are globally fused by weighted pooling to obtain multi-modal fusion features. S4. After the multi-modal fusion features are processed by feature normalization and regularization, classification processing is performed to obtain the vulnerability detection result.
[0006] The operation of the feature cross-processing based on the selected state space in the current round in S3 is as follows: The graph modality embedding cross features, text modality embedding cross features, and visual embedding cross features output in the previous round are used as the inputs of the current round and are respectively normalized to obtain the current-round graph normalized features, current-round text normalized features, and current-round visual normalized features. The current-round graph normalized features, current-round text normalized features, and current-round visual normalized features are respectively processed by a multi-layer perceptron and the SiLU activation function to obtain the current-round graph main path features, current-round text main path features, and current-round visual main path features. The current-round graph normalized features, current-round text normalized features, and current-round visual normalized features are respectively processed by a multi-layer perceptron and convolution to obtain the current-round graph convolution features, current-round text convolution features, and current-round visual convolution features. The current-round graph convolution features, current-round text convolution features, and current-round visual convolution features are respectively processed by the selected state space of the remaining two modalities without themselves through cyclic processing to obtain the current-round graph selection interaction features, current-round text selection interaction features, and current-round visual selection interaction features. The current-round graph selection interaction features, current-round text selection interaction features, and current-round visual selection interaction features are respectively element-wise multiplied by the current-round graph main path features, current-round text main path features, and current-round visual main path features and then processed by a multi-layer perceptron, and are element-wise added to the current-round graph normalized features, current-round text normalized features, and current-round visual normalized features, and then normalized to obtain the current-round graph modality cross features, current-round text modality cross features, and current-round visual modality cross features.
[0007] The operation of processing the current-round graph convolution features, the current-round text convolution features, and the current-round visual convolution features through the cyclic selection state space is specifically as follows: The current-round graph convolution features and the current-round text convolution features are respectively processed through the selection state space to obtain the current-round initial graph selection features and the current-round initial text selection features; after the current-round initial graph selection features and the current-round initial text selection features are processed through the selection state space, they are processed through the selection state space together with the current-round visual convolution features to obtain the current-round graph-text-visual selection features; after the current-round graph-text-visual selection features and the current-round text convolution features are processed through the selection state space, they are processed through the selection state space together with the current-round graph convolution features to obtain the current-round graph selection interaction features.
[0008] The operation of the selection state space can be realized through the following formula: , , is t the output of the selection state space for the round, t is , are respectively t the t round, , , are respectively the state matrix, the control matrix, and the output matrix.
[0009] The operation of obtaining the text modality embedding features in S2 is specifically as follows: The intermediate representation text and the bytecode text are respectively converted into token sequences to obtain the intermediate representation token sequence and the bytecode token sequence, which are respectively processed by the trained BERT model, and the output of the CLS token is used to obtain the intermediate representation context information and the bytecode context information; after the intermediate representation context information and the bytecode context information are respectively subjected to batch normalization and ReLU activation function processing in sequence, they are concatenated to obtain the text modality embedding features.
[0010] The types of syntax elements in the smart contract to be inspected in S1 include: parentheses, keywords, operators, identifiers, currency units, and comments.
[0011] The operation of the residual graph convolutional network processing in S2 is realized through the following formula: , , , are respectively the l layer, the lThe residual graph convolutional features of the -1 layer, is the weight matrix of the l layer convolutional layer, is the adjacency matrix of the contract graph, is the degree matrix of the adjacency matrix, is the last L layer of residual graph convolutional features, is the sorting pooling process, is the max pooling process, is the global semantic feature, is the sigmoid function process.
[0012] An intelligent contract vulnerability detection system based on multi-modal selection state space fusion, used to implement the above-mentioned intelligent contract vulnerability detection method based on multi-modal selection state space fusion, including: Multiple modal data generation modules, used to obtain the abstract syntax tree of the intelligent contract to be detected, after simplification processing, add data flow information to obtain a contract graph; convert the abstract syntax tree of the intelligent contract to be detected into a structure, obtain each function operation instruction sequence in the structure to obtain an intermediate representation text; compile the source code of the intelligent contract to be detected, obtain the runtime bytecode generated by the compilation, and convert it into a sequence to obtain a bytecode text; map different syntax element types in the intelligent contract to be detected through color coding respectively to generate a color image; map each bytecode of the intelligent contract to be detected to a grayscale pixel respectively to generate a grayscale image; Modal embedding processing module, used to process the contract graph through a residual graph convolutional network and a residual attention network respectively to obtain global semantic features and local semantic features, map them to the same shared space to obtain graph modal embedding features; respectively obtain the text context information of the intermediate representation text and the bytecode text, and perform feature alignment based on linear projection and batch normalization to obtain text modal embedding features; after the color image and the grayscale image are processed by a two-channel convolutional neural network, they are spliced to obtain visual modal embedding features; Multi-modal fusion feature generation module, used for graph modal embedding features, text modal embedding features and visual modal embedding features, through several rounds of feature cross-processing based on the selection state space, to obtain graph modal embedding cross features, text modal embedding cross features and visual modal embedding cross features, and perform global fusion through weighted pooling to obtain multi-modal fusion features; Vulnerability detection result generation module, used to perform classification processing on the multi-modal fusion features after feature normalization and regularization processing to obtain vulnerability detection results.
[0013] An intelligent contract vulnerability detection device based on multi-modal selection state space fusion, comprising a processor and a memory. When the processor executes the computer program stored in the memory, the above-mentioned intelligent contract vulnerability detection method based on multi-modal selection state space fusion is implemented.
[0014] A computer-readable storage medium for storing a computer program. When the computer program is executed by a processor, the above-mentioned intelligent contract vulnerability detection method based on multi-modal selection state space fusion is implemented.
[0015] The beneficial effects of the present invention are as follows: The intelligent contract vulnerability detection method based on multi-modal selection state space fusion provided by the present invention. First, based on the abstract syntax tree, source code, and syntax element types of the intelligent contract to be detected, a contract graph that can reflect the syntax structure and execution logic of the intelligent contract to be detected and can provide high-information and fine-grained data, an intermediate representation text and bytecode text that can reflect the source of the underlying execution semantic information of the intelligent contract to be detected, and a color image and grayscale image that can reflect the spatial distribution characteristics of the intelligent contract to be detected in different representation perspectives are obtained; then, the graphic structure information of the contract graph is converted into a high-dimensional graph embedding representation, and the features of both global and local scales are fused to improve the complex structure expression ability of the intelligent contract to be detected, and a graph modality embedding feature is obtained; and the text context information of the intermediate representation text reflecting the high-level logical behavior and the bytecode text of the underlying execution semantics is feature-aligned to form a complementary text representation space, providing a basis for vulnerability detection across different abstraction levels, and a text modality embedding feature is obtained; and the spatial layout features are extracted from the color image and grayscale image through a dual-channel convolutional neural network and spliced to capture the visual pattern information of the intelligent contract to be detected from the perspectives of syntax structure and execution semantics, enhancing the perception ability of the spatial semantics of the intelligent contract to be detected, and a visual modality embedding feature is obtained; then, the graph modality embedding feature, the text modality embedding feature, and the visual modality embedding feature are subjected to feature cross-processing based on the selection state space for several rounds to achieve in-modal depth transformation and cross-modal information interaction, obtaining a graph modality embedding cross feature, a text modality embedding cross feature, and a visual modality embedding cross feature, and performing global fusion through weighted pooling to achieve efficient integration of multi-modal features, obtaining a multi-modal fusion feature; finally, after the multi-modal fusion feature is subjected to feature normalization and regularization processing, classification processing is performed to obtain a vulnerability detection result. Detailed implementation manners
[0016] This embodiment provides an intelligent contract vulnerability detection method based on multi-modal selection state space fusion, including the following operations: S1. Obtain the abstract syntax tree of the smart contract to be inspected. After simplification, add data flow information to obtain a contract graph; convert the abstract syntax tree of the smart contract to be inspected into a structure, obtain each function operation instruction sequence in the structure to obtain an intermediate representation text; compile the source code of the smart contract to be inspected, obtain the runtime bytecode generated by compilation, and convert it into a sequence to obtain a bytecode text; map different syntax element types in the smart contract to be inspected through color coding respectively to generate a color image; map each bytecode of the smart contract to be inspected to grayscale pixels respectively to generate a grayscale image. S2. Process the contract graph through a residual graph convolutional network and a residual attention network respectively to obtain global semantic features and local semantic features, map them to the same shared space to obtain graph modality embedding features; obtain the text context information of the intermediate representation text and the bytecode text respectively, and perform feature alignment based on linear projection and batch normalization to obtain text modality embedding features; process the color image and the grayscale image through a dual-channel convolutional neural network and then splice them to obtain visual modality embedding features. S3. The graph modality embedding features, text modality embedding features, and visual modality embedding features are processed through several rounds of feature cross-processing based on the selected state space to obtain graph modality embedding cross features, text modality embedding cross features, and visual modality embedding cross features, and then globally fused through weighted pooling to obtain multi-modal fusion features. S4. After the multi-modal fusion features are processed by feature normalization and regularization, perform classification processing to obtain the vulnerability detection result.
[0017] The specific operation details of the steps are as follows.
[0018] S1. Obtain the abstract syntax tree of the smart contract to be inspected. After simplification, add data flow information to obtain a contract graph; convert the abstract syntax tree of the smart contract to be inspected into a structure, obtain each function operation instruction sequence in the structure to obtain an intermediate representation text; compile the source code of the smart contract to be inspected, obtain the runtime bytecode generated by compilation, and convert it into a sequence to obtain a bytecode text; map different syntax element types in the smart contract to be inspected through color coding respectively to generate a color image; map each bytecode of the smart contract to be inspected to grayscale pixels respectively to generate a grayscale image.
[0019] Based on the abstract syntax tree, source code, and syntactic element types of the smart contract to be inspected, obtain a contract graph that can reflect the syntactic structure and execution logic of the smart contract to be inspected and provide high-information and fine-grained data, an intermediate representation text and bytecode text that can reflect the source of the semantic information close to the underlying execution of the smart contract to be inspected, and color images and grayscale images that can reflect the spatial distribution characteristics of the smart contract to be inspected from different representation perspectives, and collect multi-dimensional information representations of the smart contract to be inspected, so as to be able to extract more comprehensive information of the smart contract to be inspected.
[0020] Generate a contract graph.
[0021] Obtain the abstract syntax tree of the smart contract to be inspected. After simplification, add data flow information to obtain a contract graph. Specifically: Generate an abstract syntax tree (abbreviated as AST) based on the lexical and syntactic rules of the Solidity source code, which realizes structuring the code into a hierarchical node tree. The abstract syntax tree carefully records the structural details of the smart contract, such as variable declarations, data types, and control flow elements, providing a complete picture of the code logic; Delete the nodes in the abstract syntax tree that are irrelevant to vulnerabilities. For example, deleting attributes such as "initialValue", "isDeclaredConst", and "visibility" will not directly affect the existence of vulnerabilities, and implement the simplification process of the abstract syntax tree to obtain a simplified abstract syntax tree (see the logic execution code in Table 1 for this process). The simplified abstract syntax tree contains rich information about the source code, including syntactic structure, semantic details, and function calls; Add data flow information to the simplified abstract syntax tree to enhance the expressive power of the abstract syntax tree by adding data flow and execution sequences, and obtain a contract graph. See the logic execution code in Table 2 for the specific process.
[0022] Table 1 Logic code for simplifying the abstract syntax tree
[0023] Table 2 Logic code for adding data flow and execution sequences to the simplified abstract syntax tree
[0024] In Table 2, for the given input of the simplified abstract syntax tree (abbreviated as S-AST), first traverse the simplified abstract syntax tree, extract each node and edge, and respectively form the nodes and edges of the contract graph. Then, before adding the child nodes to the node set, record their indices to represent the execution order of the statements; At the same time, create data flow edges between the nodes where variables first appear and subsequent nodes, and add them to the edge set. According to the above steps, the generated contract graph provides a high-information and fine-grained data basis for subsequent feature extraction.
[0025] Generate intermediate representation text.
[0026] Convert the abstract syntax tree of the smart contract to be inspected into a structure, obtain the sequence of function operation instructions in the structure, and get the intermediate representation text. Specifically: Parse the smart contract to be inspected through the slither <contract.sol> command. Slither internally converts the abstract syntax tree of the smart contract to be inspected into a structure friendly for static analysis, and at the same time outputs structured intermediate representation information at the function level. Extract the sequence of operation instructions represented by Slither in each function to form the intermediate representation text of the smart contract to be inspected, which retains the static structural features and logical semantics of the smart contract to be inspected and is convenient for mining high-level semantic features.
[0027] Generate bytecode text.
[0028] Compile the source code of the smart contract to be inspected, obtain the runtime bytecode generated by the compilation, convert it into a sequence, and get the bytecode text. Specifically: Use the Solidity compiler (such as solc --bin) to compile the source code of the smart contract to be inspected, obtain the runtime bytecode (Runtime Bytecode) generated by the compilation, which is used as the semantic modality at the binary layer, and then convert it into a hexadecimal sequence text for subsequent deep text modeling to get the bytecode text. The bytecode text is closer to the real runtime behavior, can reveal the execution semantics after being processed by the compiler in the source code, and helps to mine the hidden security risks and the associations between the underlying execution paths.
[0029] Generate a color image.
[0030] After respectively performing color coding mapping processing on different syntax element types in the smart contract to be inspected, generate a color image. Specifically: Divide the syntax elements in the smart contract to be inspected into brackets, keywords, operators, identifiers, currency units, comments, and other types (the set composed of syntax elements other than the above types). Assign fixed color coding to the syntax element types, thereby generating a color image with the ability to distinguish syntax. It not only retains the structural information of the code itself but also enhances the semantic distribution features in the image, which helps the visual model to understand the contract structure from the spatial dimension.
[0031] Generate a grayscale image.
[0032] Each bytecode of the smart contract to be inspected is mapped to a grayscale pixel respectively to generate a grayscale image. Specifically: each bytecode value of the smart contract to be inspected is mapped to a grayscale pixel, thereby constructing a grayscale image of the smart contract. The width of this image is set to a fixed 256 pixels, while the height of the image is adjusted according to the bytecode length, that is, a strategy with dimensions set to (256, -1) is adopted to automatically expand the vertical dimension of the image. The generated grayscale image can not only depict the byte distribution characteristics of the contract execution layer, but also indirectly reflect the spatial distribution of contract complexity and logical density, providing visual support for the underlying semantics of the vulnerability detection model.
[0033] S2. The contract graph is processed by the residual graph convolutional network and the residual attention network respectively to obtain global semantic features and local semantic features, which are mapped to the same shared space to obtain graph modality embedding features; the text context information of the intermediate representation text and the bytecode text is obtained respectively, and feature alignment is performed based on linear projection and batch normalization to obtain text modality embedding features; the color image and the grayscale image are processed by a dual-channel convolutional neural network and then concatenated to obtain visual modality embedding features.
[0034] Convert the graphical structure information of the contract graph into a high-dimensional graph embedding representation, and fuse the features at both the global and local scales to improve the ability to express the complex structure of the smart contract to be inspected, and obtain graph modality embedding features; at the same time, align the text context information of the intermediate representation text reflecting the high-level logical behavior and the bytecode text of the underlying execution semantics to form a complementary text representation space, providing a basis for vulnerability detection across different abstraction levels, and obtaining text modality embedding features; and process the color image and the grayscale image through a dual-channel convolutional neural network, extract the spatial layout features and concatenate them to capture the visual pattern information of the smart contract to be inspected from the perspectives of syntactic structure and execution semantics, and enhance the perception ability of the spatial semantics of the smart contract to be inspected, and obtain visual modality embedding features.
[0035] Generate graph modality embedding features.
[0036] In order to comprehensively obtain the global and local semantic relationships in the contract graph, in this embodiment, the contract graph is processed by the residual graph convolutional network and the residual attention network respectively to obtain global semantic features and local semantic features; the global semantic features and the local semantic features are mapped to the same shared space to obtain graph modality embedding features.
[0037] Among them, the operation of the residual graph convolutional network processing can be realized by the following formula: , , , are respectively the lThe residual graph convolutional features of layer l -1, when l = 0, is the node feature matrix of the contract graph, is the weight matrix of the l -th convolutional layer, is the adjacency matrix of the contract graph, is the degree matrix of the adjacency matrix, is the residual graph convolutional feature of the last L layer, is the sorting pooling process, is the max pooling process, is the global semantic feature, is the sigmoid function process. In each layer of graph convolution in the residual graph convolutional network, after the node features pass through the convolution operation, the global information of the nodes is retained, and at the same time, the features are enhanced through residual connections to avoid redundant nodes in the graph affecting the global feature expression. The sorting pooling process is introduced for the graph convolutional feature matrix, and the top n nodes are selected as the graph-level representation according to the importance of the node features, so as to extract stable and structure-sensitive global embedding features; in order to extract the most significant features of the graph embedding, we perform a max pooling operation on the global embedding features to obtain the global semantic feature.
[0038] Among them, the operations of the residual attention network process can be implemented by the following formula:
[0039] , , , are the residual attention features of the contract graph in the l -th layer and the l -1 layer respectively, is the node feature of node j in the l- 1-st layer, is the attention projection matrix of the k -th attention head, is the attention coefficient of node i and node j in the k -th attention head, K is the total number of attention heads, , are the node features of node i and node j respectively, is the attention parameter of the k -th attention head, is the residual attention feature of the last layer L The feature of the layer is processed by the ELU function is processed by the ELU function is LeakyReLU processed by the function is processed by the softmax function is processed by sorting pooling is processed by max pooling is the local semantic feature
[0040] Map the global semantic feature and the local semantic feature to the same shared space. This can be achieved by respectively and sequentially performing batch normalization and ReLU activation function processing on the global semantic feature and the local semantic feature, and then concatenating them. It can be realized through the following formula: , is the graph modality embedding feature , are the weight matrices of the graph embedding projection layers at the global scale and the local scale respectively is batch normalization processing. Through the above graph embedding extraction process, the graphical structure information of the smart contract is converted into a high-dimensional graph embedding representation, and the graph features of both the global and local scales are organically integrated. This graph embedding method can more comprehensively represent the complex structure of the smart contract
[0041] Generate the text modality embedding feature
[0042] Respectively obtain the text context information of the intermediate representation text and the bytecode text, and perform feature alignment based on linear projection and batch normalization to obtain the text modality embedding feature. Specifically, convert the intermediate representation text and the bytecode text into token sequences respectively to obtain the intermediate representation token sequence and the bytecode token sequence. After being processed by the trained BERT model respectively, the outputs of the CLS tokens are used to obtain the intermediate representation context information and the bytecode context information. After respectively and sequentially performing batch normalization and ReLU activation function processing, they are concatenated to map the two different levels of embedding features to the shared space to achieve feature alignment, forming a unified text representation embedding, and obtaining the text modality embedding feature. This process can fully extract rich syntax and semantic information from the source code and bytecode, and also significantly enhances the expression and recognition ability of different levels of vulnerability patterns. In particular, the information provided by the multi-level text representation embedding constructs a more comprehensive feature space for vulnerability detection, and at the same time integrates high-level logical semantics and low-level execution features, providing complementary semantic information support for cross-modal interaction
[0043] Generate the visual modality embedding feature
[0044] After the color image and the grayscale image are processed by the dual-channel convolutional neural network, they are spliced to obtain the visual modality embedding features. That is, after the color image and the grayscale image are respectively processed by multi-stage repeated convolution, they are respectively subjected to batch normalization and ReLU activation function processing in sequence, and then spliced to map the two features into a shared space to obtain the visual modality embedding features. In this way, the color image of the source code and the grayscale image of the bytecode respectively provide spatial information from different perspectives of the smart contract. By combining the embeddings of these two perspectives, the understanding of the contract structure and semantics is enhanced, and the perception ability of the spatial structure is effectively expanded.
[0045] The operation of multi-stage repeated convolution processing can be realized by the following formula: , is the output, are the multi-stage repeated convolution features of the color image and the grayscale image, represents the convolutional layer at the m stage repeated L times, M is the total number of stages, is the m th stage with a resolution of and a width of of the color image or grayscale image, is the dot product processing.
[0046] During the operation of multi-stage repeated convolution processing, it follows: , , , , are respectively the convolutional scaling depth d , width w and resolution coefficient r of the output, , , , , are respectively the preset convolution scale, preset convolution times, first and second preset resolution parameters, and preset image width of the convolutional layer at the m stage, , are the output memory and target memory, , are the output floating point number and target floating point number.
[0047] S3. The graph modality embedding features, text modality embedding features, and visual modality embedding features are subjected to feature cross - processing based on the selected state space for several rounds to obtain graph modality embedding cross - features, text modality embedding cross - features, and visual modality embedding cross - features, which are globally fused through weighted pooling to obtain multi - modality fusion features.
[0048] The graph modality embedding features, text modality embedding features, and visual modality embedding features are subjected to feature cross - processing based on the selected state space for several rounds to achieve in - modality depth transformation and cross - modality information interaction, obtaining graph modality embedding cross - features, text modality embedding cross - features, and visual modality embedding cross - features. Through weighted pooling for global fusion, efficient integration of multi - modality features is achieved, and multi - modality fusion features are obtained.
[0049] In the current round, the operation steps of the feature cross - processing based on the selected state space are as follows.
[0050] Step 1. The graph modality embedding cross - features (previous - round graph modality embedding cross - features), text modality embedding cross - features (previous - round text modality embedding cross - features), and visual embedding cross - features (previous - round visual embedding cross - features) output in the previous round are used as the inputs of the current round. After being normalized respectively to stabilize the feature distribution and avoid interference from inconsistent scales on feature fusion, the current - round graph - normalized features, current - round text - normalized features, and current - round visual - normalized features are obtained.
[0051] Step 2. The current - round graph - normalized features, current - round text - normalized features, and current - round visual - normalized features are processed by a multi - layer perceptron and the SiLU activation function respectively to obtain the current - round graph main - path features, current - round text main - path features, and current - round visual main - path features; the current - round graph - normalized features, current - round text - normalized features, and current - round visual - normalized features are processed by a multi - layer perceptron and convolution respectively to obtain the current - round graph convolution features, current - text convolution features, and current - round visual convolution features; the current - round graph convolution features, current - round text convolution features, and current - round visual convolution features are respectively processed with the convolution features of the remaining two modalities excluding themselves through the cyclic selected state space to obtain the current - round graph selection interaction features, current - round text selection interaction features, and current - round visual selection interaction features.
[0052] Taking the current-round graph convolution features as an example, the operations of the current-round graph convolution features, the current-round text convolution features, and the current-round visual convolution features processed by the cyclic selection state space are as follows: The current-round graph convolution features and the current-round text convolution features are respectively processed by the selection state space to obtain the current-round initial graph selection features and the current-round initial text selection features; after the current-round initial graph selection features and the current-round initial text selection features are processed by the selection state space, they are processed by the selection state space together with the current-round visual convolution features to obtain the current-round graph-text-visual selection features; the current-round graph-text-visual selection features and the current-round text convolution features are processed by the selection state space and then processed by the selection state space together with the current-round graph convolution features to obtain the current-round graph selection interaction features.
[0053] Taking the current-round text convolution features as an example, the operations of the current-round text convolution features, the current-round graph convolution features, and the current-round visual convolution features processed by the cyclic selection state space are as follows: The current-round text convolution features are processed by the selection state space to obtain the current-round initial text selection features; after the current-round initial text selection features and the current-round graph convolution features are processed by the selection state space, they are processed by the selection state space together with the current-round initial text selection features to obtain the current-round text semi-cyclic selection features; the current-round text semi-cyclic selection features and the current-round visual convolution features are processed by the selection state space and then processed by the selection state space together with the current-round text convolution features to obtain the current-round text selection interaction features.
[0054] Taking the current-round visual convolution features as an example, the operations of the current-round visual convolution features, the current-round text convolution features, and the current-round graph convolution features processed by the cyclic selection state space are as follows: The current-round visual convolution features and the current-round text convolution features are respectively processed by the selection state space to obtain the current-round initial visual selection features and the current-round initial text selection features; after the current-round initial visual selection features and the current-round text convolution features are processed by the selection state space, they are processed by the selection state space together with the current-round graph convolution features to obtain the current-round visual-graph selection features; the current-round visual-graph selection features and the current-round initial text selection features are processed by the selection state space and then processed by the selection state space together with the current-round visual convolution features to obtain the current-round graph selection interaction features.
[0055] The operation of the selection state space can be implemented by the following formula: , , is t the output of the selection state space at the is t the input at the round. When the number of initial input features is greater than 1, different initial inputs are concatenated and used as one input. , are respectively t the potential state quantities of round - 1, t rounds, , , which are respectively the state matrix, the control matrix, and the output matrix.
[0056] Step 3: The current - round graph selection interaction feature, the current - round text selection interaction feature, and the current - round visual selection interaction feature are respectively multiplied element - by - element with the current - round graph main - path feature, the current - round text main - path feature, and the current - round visual main - path feature, then processed by a multi - layer perceptron, and added element - by - element with the current - round graph normalized feature, the current - round text normalized feature, and the current - round visual normalized feature. After normalization, the current - round graph modality cross - feature, the current - round text modality cross - feature, and the current - round visual modality cross - feature are obtained, which are used to perform the feature cross - processing based on the selection state space in the next round.
[0057] Repeat the operations in Step 1, Step 2, and Step 3 until the last round is reached, and the graph modality embedding cross - feature, the text modality embedding cross - feature, and the visual modality embedding cross - feature are obtained.
[0058] The above - mentioned feature cross - processing based on the selection state space in several rounds can make the information of different modalities be adjusted targeted during the feature transfer process, further ensuring the attention to the features related to vulnerability generation. At the same time, cross - modality information exchange is realized between different modalities through intermediate connections, promoting the interaction and complementarity of features of each modality. Thus, the perception ability of each modality to the information of other modalities is enhanced, and through residual connection, the features after modulation and interaction are fused with the input features, and then after normalization, the updated modality features are obtained. In this way, the original features and the new interaction information are effectively combined, further enriching the feature representation.
[0059] Finally, the graph modality embedding cross - feature, the text modality embedding cross - feature, and the visual modality embedding cross - feature are globally fused through weighted pooling to ensure that the most discriminative features can be used for fusion, gradually enhancing the feature fusion effect, and obtaining a more rich and robust multi - modality feature representation - multi - modality fusion feature.
[0060] The whole process combines in - modality depth transformation, inter - modality information interaction, and weighted fusion, successfully realizing the efficient integration of multi - modality features.
[0061] S4: After the multi - modality fusion feature is processed by feature normalization and regularization, classification processing is performed to obtain the vulnerability detection result.
[0062] Perform feature normalization operations on the multi-modal fusion features to standardize the features. Then, through regularization processing, improve the generalization ability and prevent overfitting. Finally, through linear classification processing based on a single fully connected layer, map the features to a specific output space, and then output the prediction result on whether there are vulnerabilities in the smart contract to be detected.
[0063] The operation of obtaining the vulnerability detection result can be achieved through the following formula: , is the predicted probability for each vulnerability. According to the comparison relationship between the predicted probability and different vulnerability probability thresholds, determine whether there are corresponding vulnerabilities; is the classification weight matrix, is the bias term, is the multi-modal fusion feature, is the feature normalization processing, is the regularization processing.
[0064] To verify the effectiveness of the detection method in this embodiment, the following experiments were conducted.
[0065] Experimental purpose. To deeply evaluate the effectiveness and innovation of the method in this embodiment, the following three core research questions were explored in the experiment: RQ1: What is the detection efficiency of the method in this embodiment for four typical smart contract vulnerabilities such as reentrancy, timestamp dependence, integer overflow, and delegate call? Compared with existing detection methods, what are the differences in its performance indicators? RQ2: In the smart contract vulnerability detection task, does the multi-modal learning strategy significantly improve the key performance indicators such as model accuracy, recall rate, and F1 value?
[0066] Dataset description. The dataset mainly comes from blockchain platforms (accounting for more than 96%), GitHub open source code libraries, and professional contract analysis blogs, and has good representativeness and authenticity. Specifically, this dataset contains a total of 42,910 smart contract samples, and each sample provides dual representations of source code and bytecode. Among them, there are 680 contracts with reentrancy vulnerabilities, 2,242 contracts with timestamp dependence vulnerabilities, approximately 1,368 contracts with integer overflow / underflow vulnerabilities, and 136 contracts with delegate call vulnerabilities. All samples have completed vulnerability type annotation, providing a standardized data basis for the training and evaluation of the detection model. In the data partitioning stage, a stratified random sampling strategy was adopted to divide the dataset into a training set, a validation set, and a test set according to a ratio of 3:1:1 to ensure that the sample distribution of each subset is consistent with the original dataset. To ensure the reliability and stability of the experimental results, each experiment was repeated five times, and finally the mean of the five experimental results was used as the model performance evaluation index.
[0067] Evaluation Metrics. Four metrics, namely accuracy (ACC), recall (RE), precision (PRE), and F1-score (F1), were selected for the experiment to comprehensively evaluate the performance of the detection method. The ACC metric represents the proportion of correctly predicted samples in the total number of samples, reflecting the overall prediction accuracy of the detection method. The RE metric measures the proportion of correctly identified positive samples among all actual positive samples, reflecting the ability of the detection method to capture positive samples. The PRE metric calculates the proportion of samples predicted as positive classes that are actually positive classes, evaluating the reliability of the detection method's prediction of positive class results. The F1 metric, as the harmonic mean of precision and recall, balances the consideration of the recall rate and precision rate of the detection method in positive class identification, avoiding the evaluation bias caused by a single metric.
[0068] Parameter Settings. The method of this embodiment is implemented based on the PyTorch deep learning framework. All experiments were run on a high-performance workstation configured with a 3.3 GHz Intel Core i9 processor, an NVIDIA GeForce RTX 2080 Ti graphics processor, and 64 GB of memory. The cross-entropy loss function was used as the optimization objective during training. Combining the Adam optimization algorithm and the learning rate scheduling strategy, the prediction results were calculated through forward propagation and the network parameters were updated through backward propagation to gradually optimize the model performance. In the hyperparameter tuning phase, a grid search strategy was adopted to systematically explore the optimal parameter combination. Specifically, the learning rate was adjusted within the range of {0.0001, 0.0005, 0.001, 0.002}, the hidden layer dimension was searched within the range of {64, 128, 256, 512}, and the batch size was set to {16, 32, 64, 128} for experiments.
[0069] Existing Detection Methods for Comparison. In the experiment, the method of this embodiment was compared and analyzed with traditional smart contract vulnerability detection tools. The selected comparison tools cover different detection paradigms: (1) Oyente, as an early smart contract analysis tool, pioneered automated detection; (2) SmartCheck, a static analysis tool based on abstract syntax trees; (3) Osiris, a dynamic detection scheme using symbolic execution technology; (4) Mythril, a hybrid detection framework integrating symbolic execution and taint analysis; (5) Slither, a static analysis tool based on a rule engine; (6) sFuzz, a dynamic detection tool based on fuzz testing; (7) ConFuzzius, a detection tool using a hybrid fuzz testing strategy.
[0070] Reply to RQ1. The method of this embodiment is compared with six existing detection methods on the dataset for different vulnerability detection results. See Tables 3, 4, 5, and 6. From the data in Tables 3, 4, 5, and 6, it can be seen that in terms of reentrancy vulnerability detection, the accuracies of rule-based SmartCheck and Slither are only 44.32% and 69.38% respectively, and for Oyente and Osiris that rely on bytecode analysis, the accuracies are 65.57% and 51.87% respectively, and for sFuzz and ConFuzzius that adopt fuzz testing technology, the accuracies are 60.42% and 78.02% respectively. In contrast, the method of this embodiment shows significant technical advantages. In the detection of four types of vulnerabilities, namely reentrancy, timestamp dependence, integer overflow / underflow, and delegate call, the method of this embodiment improves by 17.98%, 25.63%, 17.96%, and 11.71% respectively compared with the existing detection methods in terms of the accuracy index, and achieves the best performance in all evaluation indicators.
[0071] Table 3 Comparison between the method of this embodiment and existing methods in reentrancy vulnerability detection
[0072] Table 4 Comparison between the method of this embodiment and existing methods in timestamp vulnerability detection
[0073] Table 5 Comparison between the method of this embodiment and existing methods in integer overflow / underflow vulnerability detection
[0074] Table 6 Comparison between the method of this embodiment and existing methods in delegate call vulnerability detection
[0075] Meanwhile, in the experiment, the method of this embodiment is further compared and analyzed with current mainstream deep learning detection tools. The comparison objects include representative models such as Vanille-RNN, Rechecker, DR-GCN, TMP, DA-GNN, AME, CBGRU, VulnSense, and TMF-Net. The experimental results shown in Tables 3, 4, 5, and 6 indicate that the method of this embodiment shows significant advantages in comparison with other existing multimodal methods. It achieves performance breakthroughs in the four types of vulnerability detection tasks, which benefits from the multi-dimensional data representation system and adaptive cross-modal fusion mechanism constructed by the method of this embodiment. It can comprehensively capture the information of smart contracts and accurately model the complex interaction relationships between modalities, thereby significantly improving the detection accuracy and verifying the advantages of the method of this embodiment in multimodal information processing.
[0076] Response to RQ2. In the method of this embodiment, a contract graph (graph modality), bytecode text and intermediate representation text (text modality), and dual-perspective code images of color images and grayscale images (visual modality) are used to construct a multi-source heterogeneous data representation system. To quantify the gain in vulnerability detection performance brought by multi-modal fusion, a single-modal detection control experiment was carried out in the experiment, and the relevant results are summarized in Table 7. The data in Table 7 shows that in the task of smart contract vulnerability detection, there are significant differences in the detection efficiency of each single modality. Among them, relying on the deep modeling ability of the graph neural network for the code structure, the graph modality accurately captures key logical features such as function call relationships and data flow dependencies, ranking first with an accuracy of 82.33%; the text modality uses natural language processing technology to parse the code semantic information, achieving a detection accuracy of 82.3%, but limited by the ability to represent code structure features, there is room for performance improvement; the detection accuracy of the visual modality is 79.1%, which is relatively lower than the previous two. This is mainly due to the information loss in the code image when mapping the program syntax structure and context logical relationship, resulting in ineffective expression of some key detection elements. Comparative analysis shows that the method of this embodiment is significantly better than the single-modal model in all-dimensional evaluation indicators such as accuracy, recall, precision, and F1 score. Taking the F1 score as an example, the method of this embodiment is improved from the highest single-modal value of 76.67% to 91.94%, achieving a leapfrog breakthrough in performance. This result fully verifies that the collaborative integration of multi-modal information can effectively expand the context awareness ability of the model, significantly enhance the generalization performance and prediction accuracy of the model in the vulnerability detection task, and highlight the technical advantages of multi-modal learning in the field of smart contract security analysis.
[0077] Table 7 Performance Comparison of Different Single Modalities and Multi-Modal Fusion
[0078] This embodiment also provides a smart contract vulnerability detection system based on multi-modal selection state space fusion, which is used to implement the above-mentioned smart contract vulnerability detection method based on multi-modal selection state space fusion, including: Multiple modal data generation modules, which are used to obtain the abstract syntax tree of the smart contract to be detected, add data flow information after simplification processing to obtain a contract graph; convert the abstract syntax tree of the smart contract to be detected into a structure, and obtain the intermediate representation text by getting the function operation instruction sequence in the structure; compile the source code of the smart contract to be detected, obtain the runtime bytecode generated by compilation, and convert it into a sequence to obtain the bytecode text; perform color-coded mapping processing on different syntax element types in the smart contract to be detected to generate color images; map each bytecode of the smart contract to be detected into grayscale pixels to generate grayscale images; A modal embedding processing module is used to process the contract graph through a residual graph convolutional network and a residual attention network respectively, obtain global semantic features and local semantic features, map them to the same shared space, and obtain graph modal embedding features; respectively obtain the text context information of the intermediate representation text and the bytecode text, and perform feature alignment based on linear projection and batch normalization to obtain text modal embedding features; after the color image and the grayscale image are processed by a dual-channel convolutional neural network, they are concatenated to obtain visual modal embedding features. A multi-modal fusion feature generation module is used to perform feature cross-processing based on the selected state space on the graph modal embedding features, text modal embedding features, and visual modal embedding features for several rounds to obtain graph modal embedding cross features, text modal embedding cross features, and visual modal embedding cross features, and perform global fusion through weighted pooling to obtain multi-modal fusion features. A vulnerability detection result generation module is used to perform classification processing on the multi-modal fusion features after feature normalization and regularization processing to obtain vulnerability detection results.
[0079] This embodiment also provides an intelligent contract vulnerability detection device based on multi-modal selection state space fusion, including a processor and a memory. Among them, when the processor executes the computer program stored in the memory, it implements the above-mentioned intelligent contract vulnerability detection method based on multi-modal selection state space fusion.
[0080] This embodiment also provides a computer-readable storage medium for storing a computer program. Among them, when the computer program is executed by a processor, it implements the above-mentioned intelligent contract vulnerability detection method based on multi-modal selection state space fusion.
[0081] An intelligent contract vulnerability detection method based on multi-modal selection state space fusion provided by this embodiment. First, based on the abstract syntax tree, source code, and syntax element types of the intelligent contract to be detected, a contract graph that can reflect the syntax structure and execution logic of the intelligent contract to be detected and provide high-information and fine-grained data, an intermediate representation text and bytecode text that can reflect the source of the execution semantics close to the underlying layer of the intelligent contract to be detected, and color images and grayscale images that can reflect the spatial distribution characteristics of the intelligent contract to be detected in different representation perspectives are obtained; then, the graphic structure information of the contract graph is converted into a high-dimensional graph embedding representation, and features at both global and local scales are fused to improve the complex structure expression ability of the intelligent contract to be detected, obtaining graph modality embedding features; and the text context information of the intermediate representation text reflecting high-level logical behavior and the bytecode text of the underlying execution semantics is feature-aligned to form a complementary text representation space, providing a basis for vulnerability detection across different abstraction levels, obtaining text modality embedding features; and spatial layout features are extracted from the color images and grayscale images through dual-channel convolutional neural network processing and spliced to capture the visual pattern information of the intelligent contract to be detected from the perspectives of syntax structure and execution semantics, enhancing the perception ability of the spatial semantics of the intelligent contract to be detected, obtaining visual modality embedding features; then, the graph modality embedding features, text modality embedding features, and visual modality embedding features are subjected to feature cross-processing based on the selection state space for several rounds to achieve in-modal deep transformation and cross-modal information interaction, obtaining graph modality embedding cross features, text modality embedding cross features, and visual modality embedding cross features, and global fusion is performed through weighted pooling to achieve efficient integration of multi-modal features, obtaining multi-modal fusion features; finally, after the multi-modal fusion features are subjected to feature normalization and regularization processing, classification processing is performed to obtain the vulnerability detection result.
Claims
1. An intelligent contract vulnerability detection method based on the fusion of multi-modal selection state spaces, characterized in that Including the following operations: S1. Obtain the abstract syntax tree of the smart contract to be inspected. After simplification processing, add data flow information to obtain a contract graph; Convert the abstract syntax tree of the smart contract to be inspected into a structure, and obtain the operation instruction sequence of each function in the structure to obtain an intermediate representation text; Compile the source code of the smart contract to be inspected, obtain the runtime bytecode generated by compilation, and convert it into a sequence to obtain a bytecode text; For different syntax element types in the smart contract to be inspected, generate a color image after color-coding mapping processing; Map each bytecode of the smart contract to be inspected into a grayscale pixel to generate a grayscale image; S2. Process the contract graph through a residual graph convolutional network and a residual attention network respectively to obtain global semantic features and local semantic features, and map them to the same shared space to obtain graph modality embedding features; Obtain the text context information of the intermediate representation text and the bytecode text respectively, and perform feature alignment based on linear projection and batch normalization to obtain text modality embedding features; Process the color image and the grayscale image through a dual-channel convolutional neural network, and then splice them to obtain visual modality embedding features; S3. The graph modality embedding features, text modality embedding features, and visual modality embedding features are processed through several rounds of feature cross-processing based on a selective state space to obtain graph modality embedding cross features, text modality embedding cross features, and visual modality embedding cross features, and then perform global fusion through weighted pooling to obtain multi-modal fusion features; S4. After the multi-modal fusion features are processed by feature normalization and regularization, perform classification processing to obtain a vulnerability detection result.
2. The intelligent contract vulnerability detection method based on multi-modal selection state space fusion according to claim 1, wherein In S3, the operation of the feature cross-processing based on the selective state space in the current round is: The graph modality embedding cross features, text modality embedding cross features, and visual embedding cross features output in the previous round are used as the inputs of the current round, and are respectively normalized to obtain the current round graph normalized features, current round text normalized features, and current round visual normalized features; The current round graph normalized features, current round text normalized features, and current round visual normalized features are respectively processed by a multi-layer perceptron and a SiLU activation function to obtain the current round graph main path features, current round text main path features, and current round visual main path features; The current round graph normalized features, current round text normalized features, and current round visual normalized features are respectively processed by a multi-layer perceptron and convolution to obtain the current round graph convolution features, current round text convolution features, and current round visual convolution features; The current round graph convolution features, current round text convolution features, and current round visual convolution features are respectively processed through a cyclic selective state space with the convolution features of the remaining two modalities that do not include themselves to obtain the current round graph selective interaction features, current round text selective interaction features, and current round visual selective interaction features; The current round of graph selection interaction features, current round of text selection interaction features, and current round of visual selection interaction features are respectively multiplied element-wise with the current round of graph main path features, current round of text main path features, and current round of visual main path features, then processed by a multi-layer perceptron, and added element-wise with the current round of graph normalized features, current round of text normalized features, and current round of visual normalized features, and then normalized to obtain the current round of graph modality cross features, current round of text modality cross features, and current round of visual modality cross features.
3. The intelligent contract vulnerability detection method based on multi-modal selection state space fusion according to claim 2, wherein The operations of processing the current round of graph convolution features, current round of text convolution features, and current round of visual convolution features through the cyclic selection state space are specifically as follows: The current round of graph convolution features and current round of text convolution features are respectively processed through the selection state space to obtain the current round of initial graph selection features and current round of initial text selection features; After the current round of initial graph selection features and current round of initial text selection features are processed through the selection state space, they are processed through the selection state space with the current round of visual convolution features to obtain the current round of graph-text-visual selection features; The current round of graph-text-visual selection features and the current round of text convolution features are processed through the selection state space, and then processed through the selection state space with the current round of graph convolution features to obtain the current round of graph selection interaction features.
4. The intelligent contract vulnerability detection method based on multi-modal selection state space fusion according to claim 3, wherein The operations of the selection state space can be implemented through the following formula: , , is t the output of the selection state space of the round, is t the input of the round, and are respectively t the -1 round, t the potential state quantity of the round, and and are respectively the state matrix, the control matrix and the output matrix.
5. The intelligent contract vulnerability detection method based on multi-modal selection state space fusion according to claim 1, characterized in that The operations of obtaining the text modality embedding features in S2 are specifically as follows: The intermediate representation text and bytecode text are respectively converted into token sequences to obtain the intermediate representation token sequence and bytecode token sequence, which are respectively processed by the trained BERT model, and the outputs with the CLS markers are used to obtain the intermediate representation context information and bytecode context information; The intermediate representation context information and bytecode context information are respectively processed by batch normalization and ReLU activation function in sequence, and then concatenated to obtain the text modality embedding features.
6. The intelligent contract vulnerability detection method based on multi-modal selection state space fusion according to claim 1, wherein The types of syntax elements in the smart contract to be inspected in S1 include: parentheses, keywords, operators, identifiers, currency units, and comments.
7. The intelligent contract vulnerability detection method based on multi-modal selection state space fusion according to claim 1, characterized in that, The operations of the residual graph convolution network processing in S2 are implemented through the following formula: , , , are the residual graph convolutional features of the l -th layer and the l -1-th layer respectively. is the weight matrix of the l -th convolutional layer. is the adjacency matrix of the contract graph. is the degree matrix of the adjacency matrix. is the residual graph convolutional feature of the last L -th layer. is the sorting pooling process. is the max pooling process. is the global semantic feature. is the sigmoid function process.
8. An intelligent contract vulnerability detection system based on multimodal selection state space fusion, which is used to implement the intelligent contract vulnerability detection method based on multimodal selection state space fusion described in claim 1, characterized in that, Including: Multiple modality data generation modules, which are used to obtain the abstract syntax tree of the smart contract to be inspected, and after simplification processing, add data flow information to obtain the contract graph; Convert the abstract syntax tree of the smart contract to be inspected into a structure, obtain the operation instruction sequence of each function in the structure to obtain the intermediate representation text; compile the source code of the smart contract to be inspected, obtain the runtime bytecode generated by compilation, and convert it into a sequence to obtain the bytecode text; The different syntax element types in the smart contract to be inspected are respectively processed through color coding mapping to generate a color image; Each bytecode of the smart contract to be inspected is respectively mapped to a grayscale pixel to generate a grayscale image; The modality embedding processing module is used to process the contract graph through the residual graph convolution network and the residual attention network respectively to obtain the global semantic feature and local semantic feature, and map them to the same shared space to obtain the graph modality embedding feature; Respectively obtain the text context information of the intermediate representation text and the bytecode text, perform feature alignment based on linear projection and batch normalization, and obtain the text modality embedding features; After the color image and the grayscale image are processed by the dual-channel convolutional neural network, they are concatenated to obtain the visual modality embedding features; The multimodal fusion feature generation module is used for the graph modality embedding features, the text modality embedding features and the visual modality embedding features. After several rounds of feature cross-processing based on the selected state space, the graph modality embedding cross features, the text modality embedding cross features and the visual modality embedding cross features are obtained. After weighted pooling for global fusion, multimodal fusion features are obtained; The vulnerability detection result generation module is used for the multimodal fusion features to be classified after feature normalization and regularization processing to obtain the vulnerability detection results.
9. An intelligent contract vulnerability detection device based on the fusion of multi-modal selection state spaces, characterized in that, It includes a processor and a memory. Among them, when the processor executes the computer program stored in the memory, it implements the intelligent contract vulnerability detection method based on multimodal selection state space fusion as described in any one of claims 1-7.
10. A computer-readable storage medium, characterized in that, It is used to store a computer program. Among them, when the computer program is executed by the processor, it implements the intelligent contract vulnerability detection method based on multimodal selection state space fusion as described in any one of claims 1-7.
Citation Information
Patent Citations
Fragment information extraction model training method based on multi-answer loss function
CN112131351A
Smart contract vulnerability detection method and system based on HOG and DRSN-LSTM
CN115937878A
Smart contract security analysis method and system based on multi-modal technology
CN116958767A
Malicious code homology analysis method and device, electronic equipment and storage medium
CN117171746A
Transformer-based Ponare fraud detection method under comparison of multiple views
CN117408698A
Cited By
Intelligent contract vulnerability detection method based on multi-view learning
CN121211465A
A smart contract vulnerability detection method based on multi-view learning
CN121211465B
Intelligent contract vulnerability detection method based on machine learning and dynamic and static combined analysis
CN121834827A